Channels
Though the impact of Zentoshin’s sudden bankruptcy is limited, it’s a setback for Japan, which only recently shed its notorious cash-only image.
Confronting the Chaos and Demons
Download Kleios by Ian Clark on the App Store. See screenshots, ratings and reviews, user tips, and more apps like Kleios.
Cosmos is the easiest way to design a keyboard around your one-of-a-kind hands. Scan your hand using just your phone camera, then fit a keyboard to the scan. The key positions align to your fingers' lengths and movement. Add a trackball, trackpad, encoder, or OLED display. There's support for MX, Choc, NIZ, and Alps switches, and almost every type of keycap. Plus with 15 different microcontrollers, you can mix and match all you like. Cosmos partners with keyboard builders to deliver your customized keyboard with premium materials. If you're handy with 3D printing and soldering, you can save money by building your keyboard yourself. Cosmos also sells specialized PCBs for keyboard building! Choose from 3 types of cases, split or unibody, and many customizations. Cosmos catches errors before you print and automatically fixes common model issues. Cosmos lets you drag, drop, and rotate keys and trackballs into place. Your artisans are now ergonomic. Whatever batch of keycaps you decide to use, Cosmos will arrange them to fit your desired curvature. Cosmos has first-class support for single-key PCBs, which let you easily integrate per-key RGB and hotswap sockets. It also has high-quality built-in-hotswap and PCB-less options if you're on a budget. Every model can export to STLs, which are meant to be sent to your 3D printer or an online printing service, or to STEP models, which can be modified in CAD programs. If you don't like the way your model looks, ask your closest CAD guru to make adjustments. Cosmos is made in the open, and 95% of the code is open-source. It's our firm belief everyone should have free access to technology to relieve and prevent typing pain. Come see the unique keyboards we all are making on the Discord server. Don't have an account? I send a few recaps per year to my newsletter . The other 5% of code? That's for the Pro features, which add extra cosmetic options to your keyboard and help keep this project sustainable. Psst! Come here from my Dactyl generator? You should give Cosmos a try. It's changing a lot but it will give you a much better Dactyl-like case and microcontroller holder.
We articulate a vision of artificial intelligence (AI) as normal technology. To view AI as normal is not to understate its impact—even transformative, general-purpose technologies such as electricity and the internet are “normal” in our conception. But it is in contrast to both utopian and dystopian visions of the future of AI which have a common tendency to treat it akin to a separate species, a highly autonomous, potentially superintelligent entity. 1. Nick Bostrom. 2012. The superintelligent will: Motivation and instrumental rationality in advanced artificial agents. Minds and Machines 22, 2 (May 2012), 71–85. https://doi:10.1007/s11023-012-9281-3; Nick Bostrom. 2017. Superintelligence: Paths, Dangers, Strategies (reprinted with corrections). Oxford University Press, Oxford, United Kingdom; Sam Altman, Greg Brockman, and Ilya Sutskever. 2023. Governance of Superintelligence (May 2023). https://openai.com/blog/governance-of-superintelligence; Shazeda Ahmed et al. 2023. Building the Epistemic Community of AI Safety. SSRN: Rochester, NY. doi:10.2139/ssrn.4641526. The statement “AI is normal technology” is three things: a description of current AI, a prediction about the foreseeable future of AI, and a prescription about how we should treat it. We view AI as a tool that we can and should remain in control of, and we argue that this goal does not require drastic policy interventions or technical breakthroughs. We do not think that viewing AI as a humanlike intelligence is currently accurate or useful for understanding its societal impacts, nor is it likely to be in our vision of the future. 2. This is different from the question of whether it is helpful for an individual user to conceptualize a specific AI system as a tool as opposed to a human-like entity such as an intern, a co-worker, or a tutor. The normal technology frame is about the relationship between technology and society. It rejects technological determinism, especially the notion of AI itself as an agent in determining its future. It is guided by lessons from past technological revolutions, such as the slow and uncertain nature of technology adoption and diffusion. It also emphasizes continuity between the past and the future trajectory of AI in terms of societal impact and the role of institutions in shaping this trajectory. In Part I, we explain why we think that transformative economic and societal impacts will be slow (on the timescale of decades), making a critical distinction between AI methods, AI applications, and AI adoption, arguing that the three happen at different timescales. In Part II, we discuss a potential division of labor between humans and AI in a world with advanced AI (but not “superintelligent” AI, which we view as incoherent as usually conceptualized). In this world, control is primarily in the hands of people and organizations; indeed, a greater and greater proportion of what people do in their jobs is AI control. In Part III, we examine the implications of AI as normal technology for AI risks. We analyze accidents, arms races, misuse, and misalignment, and argue that viewing AI as normal technology leads to fundamentally different conclusions about mitigations compared to viewing AI as being humanlike. Of course, we cannot be certain of our predictions, but we aim to describe what we view as the median outcome. We have not tried to quantify probabilities, but we have tried to make predictions that can tell us whether or not AI is behaving like normal technology. In Part IV, we discuss the implications for AI policy. We advocate for reducing uncertainty as a first-rate policy goal and resilience as the overarching approach to catastrophic risks. We argue that drastic interventions premised on the difficulty of controlling superintelligent AI will, in fact, make things much worse if AI turns out to be normal technology— the downsides of which will be likely to mirror those of previous technologies that are deployed in capitalistic societies, such as inequality. 3. Daron Acemoglu and Simon Johnson. 2023. Power and Progress: Our Thousand-Year Struggle over Technology and Prosperity .PublicAffairs, New York, NY. The world we describe in Part II is one in which AI is far more advanced than it is today. We are not claiming that AI progress—or human progress—will stop at that point. What comes after it? We do not know. Consider this analogy: At the dawn of the first Industrial Revolution, it would have been useful to try to think about what an industrial world would look like and how to prepare for it, but it would have been futile to try to predict electricity or computers. Our exercise here is similar. Since we reject “fast takeoff” scenarios, we do not see it as necessary or useful to envision a world further ahead than we have attempted to. If and when the scenario we describe in Part II materializes, we will be able to better anticipate and prepare for whatever comes next. A note to readers. This essay has the unusual goal of stating a worldview rather than defending a proposition. The literature on AI superintelligence is copious. We have not tried to give a point-by-point response to potential counter arguments, as that would make the paper several times longer. This paper is merely the initial articulation of our views; we plan to elaborate on them in various follow ups. Part I: The Speed of Progress Figure 1. Like other general-purpose technologies, the impact of AI is materialized not when methods and capabilities improve, but when those improvements are translated into applications and are diffused through productive sectors of the economy. 4. Jeffrey Ding. 2024. Technology and the Rise of Great Powers: How Diffusion Shapes Economic Competition. Princeton University Press, Princeton. There are speed limits at each stage. Will the progress of AI be gradual, allowing people and institutions to adapt as AI capabilities and adoption increase, or will there be jumps leading to massive disruption, or even a technological singularity? Our approach to this question is to analyze highly consequential tasks separately from less consequential tasks and to begin by analyzing the speed of adoption and diffusion of AI before returning to the speed of innovation and invention. We use invention to refer to the development of new AI methods—such as large language models—that improve AI’s capabilities to carry out various tasks. Innovation refers to the development of products and applications using AI that consumers and businesses can use. Adoption refers to the decision by an individual (or team or firm) to use a technology, whereas diffusion refers to the broader social process through which the level of adoption increases. For sufficiently disruptive technologies, diffusion might require changes to the structure of firms and organizations, as well as to social norms and laws. AI diffusion in safety-critical areas is slow In the paper Against Predictive Optimization, we compiled a comprehensive list of about 50 applications of predictive optimization, namely the use of machine learning (ML) to make decisions about individuals by predicting their future behavior or outcomes. 5. Angelina Wang et al. 2023. Against predictive optimization: On the legitimacy of decision-making algorithms that optimize predictive accuracy. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency (Chicago, IL, USA: ACM, 2023), 626–26. doi:10.1145/3593013.3594030. Most of these applications, such as criminal risk prediction, insurance risk prediction, or child maltreatment prediction, are used to make decisions that have important consequences for people. While these applications have proliferated, there is a crucial nuance: In most cases, decades-old statistical techniques are used—simple, interpretable models (mostly regression) and relatively small sets of handcrafted features. More complex machine learning methods, such as random forests, are rarely used, and modern methods, such as transformers, are nowhere to be found. In other words, in this broad set of domains, AI diffusion lags decades behind innovation. A major reason is safety—when models are more complex and less intelligible, it is hard to anticipate all possible deployment conditions in the testing and validation process. A good example is Epic’s sepsis prediction tool which, despite having seemingly high accuracy when internally validated, performed far worse in hospitals, missing two thirds of sepsis cases and overwhelming physicians with false alerts. 6. Casey Ross. 2022. Epic’s Overhaul of a Flawed Algorithm Shows Why AI Oversight Is a Life-or-Death Issue. STAT. https://www.statnews.com/2022/10/24/epic-overhaul-of-a-flawed-algorithm/. Epic’s sepsis prediction tool failed because of errors that are hard to catch when you have complex models with unconstrained feature sets. 7. Andrew Wong et al. 2021. External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients. JAMA Internal Medicine 181, 8 (August 2021), 1065–70, https://doi:10.1001/jamainternmed.2021.2626. In particular, one of the features used to train the model was whether a physician had already prescribed antibiotics —to treat sepsis. In other words, during testing and validation, the model was using a feature from the future, relying on a variable that was causally dependent on the outcome. Of course, this feature would not be available during deployment. Interpretability and auditing methods will no doubt improve so that we will get much better at catching these issues, but we are not there yet. In the case of generative AI, even failures that seem extremely obvious in hindsight were not caught during testing. One example is the early Bing chatbot “Sydney” that went off the rails during extended conversations; the developers evidently did not anticipate that conversations could last for more than a handful of turns. 8. Kevin Roose. 2023. A Conversation With Bing’s Chatbot Left Me Deeply Unsettled. The New York Times (February 2023). https://www.nytimes.com/2023/02/16/technology/bing-chatbot-microsoft-chatgpt.html. Similarly, the Gemini image generator was seemingly never tested on historical figures. 9. Dan Milmo and Alex Hern. 2024. ‘We definitely messed up’: why did Google AI tool make offensive historical images? The Guardian (March 2024). https://www.theguardian.com/technology/2024/mar/08/we-definitely-messed-up-why-did-google-ai-tool-make-offensive-historical-images Fortunately, these were not highly consequential applications. More empirical work would be helpful for understanding the innovation-diffusion lag in various applications and the reasons for this lag. But, for now, the evidence that we have analyzed in our previous work is consistent with the view that there are already extremely strong safety-related speed limits in highly consequential tasks. These limits are often enforced through regulation, such as the FDA’s supervision of medical devices, as well as newer legislation such as the EU AI Act, which puts strict requirements on high-risk AI. 10. Jamie Bernardi et al. 2024. Societal adaptation to advanced AI. arXiv: May 2024. Retrieved from http://arxiv.org/abs/2405.10295; Center for Devices and Radiological Health. 2024. Regulatory evaluation of new artificial intelligence (AI) uses for improving and automating medical practices. FDA (June 2024). https://www.fda.gov/medical-devices/medical-device-regulatory-science-research-programs-conducted-osel/regulatory-evaluation-new-artificial-intelligence-ai-uses-improving-and-automating-medical-practices; “Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 Laying down Harmonised Rules on Artificial Intelligence and Amending Regulations (EC) No 300/2008, (EU) No 167/2013, (EU) No 168/2013, (EU) 2018/858, (EU) 2018/1139 and (EU) 2019/2144 and Directives 2014/90/EU, (EU) 2016/797 and (EU) 2020/1828 (Artificial Intelligence Act) (Text with EEA Relevance),” June 2024, http://data.europa.eu/eli/reg/2024/1689/oj/eng. In fact, there are (credible) concerns that existing regulation of high-risk AI is so onerous that it may lead to “runaway bureaucracy”. 11. Javier Espinoza. 2024. Europe’s rushed attempt to set the rules for AI. Financial Times (July 2024). https://www.ft.com/content/6cc7847a-2fc5-4df0-b113-a435d6426c81; Daniel E. Ho and Nicholas Bagley. 2024. Runaway bureaucracy could make common uses of ai worse, even mail delivery. The Hill (January 2024). https://thehill.com/opinion/technology/4405286-runaway-bureaucracy-could-make-common-uses-of-ai-worse-even-mail-delivery/. Thus, we predict that slow diffusion will continue to be the norm in high-consequence tasks. At any rate, as and when new areas arise in which AI can be used in highly consequential ways, we can and must regulate them. A good example is the Flash Crash of 2010, in which automated high-frequency trading is thought to have played a part. This led to new curbs on trading, such as circuit breakers. 12. Avanidhar Subrahmanyam. 2013. Algorithmic trading, the flash crash, and coordinated circuit breakers. Borsa Istanbul Review 13, 3 (September 2013), 4–9. http://doi:10.1016/j.bir.2013.10.003. Diffusion is limited by the speed of human, organizational, and institutional change Even outside of safety-critical areas, AI adoption is slower than popular accounts would suggest. For example, a study made headlines due to the finding that, in August 2024, 40% of U.S. adults used generative AI. 13. Alexander Bick, Adam Blandin, and David J. Deming. 2024. The Rapid Adoption of Generative AI. National Bureau of Economic Research. But, because most people used it infrequently, this only translated to 0.5%-3.5% of work hours (and a 0.125-0.875 percentage point increase in labor productivity). It is not even clear if the speed of diffusion is greater today compared to the past. The aforementioned study reported that generative AI adoption in the U.S. has been faster than personal computer (PC) adoption, with 40% of U.S. adults adopting generative AI within two years of the first mass-market product release compared to 20 % within three years for PCs. But this comparison does not account for differences in the intensity of adoption (the number of hours of use) or the high cost of buying a PC compared to accessing generative AI. 14. Alexander Bick, Adam Blandin, and David J. Deming. 2024. The Rapid Adoption of Generative AI. National Bureau of Economic Research. Depending on how we measure adoption, it is quite possible that the adoption of generative AI has been much slower than PC adoption. The claim that the speed of technology adoption is not necessarily increasing may seem surprising (or even obviously wrong) given that digital technology can reach billions of devices at once. But it is important to remember that adoption is about software use, not availability. Even if a new AI-based product is instantly released online for anyone to use for free, it takes time to for people to change their workflows and habits to take advantage of the benefits of the new product and to learn to avoid the risks. Thus, the speed of diffusion is inherently limited by the speed at which not only individuals, but also organizations and institutions, can adapt to technology. This is a trend that we have also seen for past general-purpose technologies: Diffusion occurs over decades, not years. 15. Benedict Evans. 2023. AI and the Automation of Work. https://www.ben-evans.com/benedictevans/2023/7/2/working-with-ai; Benedict Evans, 2023; Jeffrey Ding. 2024. Technology and the Rise of Great Powers: How Diffusion Shapes Economic Competition. Princeton University Press, Princeton. As an example, Paul A. David’s analysis of electrification shows that the productivity benefits took decades to fully materialize. 16. Paul A. David. 1990. The dynamo and the computer: an historical perspective on the modern productivity paradox. The American Economic Review 80, 2 (1990), 355–61. https://www.jstor.org/stable/2006600; Tim Harford. 2017. Why didn’t electricity immediately change manufacturing? (August 2017). https://www.bbc.com/news/business-40673694. Electric dynamos were “everywhere but in the productivity statistics” for nearly 40 years after Edison’s first central generating station. 17. Robert Solow as quoted in Paul A. David. 1990. The dynamo and the computer: an historical perspective on the modern productivity paradox. The American Economic Review 80, 2 (1990), Page 355. https://www.jstor.org/stable/2006600; Tim Harford. 2017. Why didn’t electricity immediately change manufacturing? (August 2017). https://www.bbc.com/news/business-40673694. This was not just technological inertia; factory owners found that electrification did not bring substantial efficiency gains. What eventually allowed gains to be realized was redesigning the entire layout of factories around the logic of production lines. In addition to changes to factory architecture, diffusion also required changes to workplace organization and process control, which could only be developed through experimentation across industries. Workers had more autonomy and flexibility as a result of the changes, which also necessitated different hiring and training practices. The External world puts a speed limit on AI innovation It is true that technical advances in AI have been rapid, but the picture is much less clear when we differentiate AI methods from applications. We conceptualize progress in AI methods as a ladder of generality. 18. Arvind Narayanan and Sayash Kapoor. 2024. AI Snake Oil: What Artificial Intelligence Can Do, What It Can’t, and How to Tell the Difference. Princeton University Press, Princeton, NJ. Each step on this ladder rests on the ones below it and reflects a move toward more general computing capabilities. That is, it reduces the programmer effort needed to get the computer to perform a new task and increases the set of tasks that can be performed with a given amount of programmer (or user) effort; see Figure 2. For example, machine learning increases generality by obviating the need for the programmer to devise logic to solve each new task, only requiring the collection of training examples instead. It is tempting to conclude that the effort required to develop specific applications will keep decreasing as we build more rungs of the ladder until we reach artificial general intelligence, often conceptualized as an AI system that can do everything out of the box, obviating the need to develop applications altogether. In some domains, we are indeed seeing this trend of decreasing application development effort. In natural language processing, large language models have made it relatively trivial to develop a language translation application. Or consider games: AlphaZero can learn to play games such as chess better than any human through self-play given little more than a description of the game and enough computing power—a far cry from how game-playing programs used to be developed. Figure 2: The Ladder of Generality in Computing. For some tasks, higher ladder rungs require less programmer effort to get a computer to perform a new task, and more tasks can be performed with a given amount of programmer (or user) effort. 19. Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012. ImageNet classification with deep convolutional neural networks. Advances in Neural Information Processing Systems 25 (2012); Harris Drucker, Donghui Wu, and Vladimir N. Vapnik. 1999. Support vector machines for spam categorization. IEEE Transactions on Neural Networks 10, 5 (September 1999), 1048–54. http://doi:10.1109/72.788645; William D. Smith. 1964. New I.B.M, System 360 can serve business, science and government; I.B.M. Introduces a computer it says tops output of biggest. The New York Times April 1964. https://www.nytimes.com/1964/04/08/archives/new-ibm-system-360-can-serve-business-science-and-government-ibm.html; Special to THE NEW YORK TIMES. Algebra machine spurs research calling for long calculations; Harvard receives today device to solve in hours problems taking so much time they have never been worked out. The New York Times (August 1944). https://www.nytimes.com/1944/08/07/archives/algebra-machine-spurs-research-calling-for-long-calculations.html; Herman Hollerith. 1894. The electrical tabulating machine. Journal of the Royal Statistical Society 57, 4 (December 1894), 678. http://doi:10.2307/2979610. However, this has not been the trend in highly consequential, real-world applications that cannot easily be simulated and in which errors are costly. Consider self-driving cars: In many ways, the trajectory of their development is similar to AlphaZero’s self-play—improving the tech allowed them to drive in more realistic conditions, which enabled the collection of better and/or more realistic data, which in turn led to improvements in the tech, completing the feedback loop. But this process took over two decades instead of a few hours in the case of AlphaZero because safety considerations put a limit on the extent to which each iteration of this loop could be scaled up compared to the previous one. 20. Mohammad Musa, Tim Dawkins, and Nicola Croce. 2019. This is the next step on the road to a safe self-driving future. World Economic Forum (December 2019). https://www.weforum.org/stories/2019/12/the-key-to-a-safe-self-driving-future-lies-in-sharing-data/; Louise Zhang. 2023. Cruise’s Safety Record Over 1 Million Driverless Miles. Cruise (April 2023). https://web.archive.org/web/20230504102309/https://getcruise.com/news/blog/2023/cruises-safety-record-over-one-million-driverless-miles/ This “capability-reliability gap” shows up over and over. It has been a major barrier to building useful AI “agents” that can automate real-world tasks. 21. Arvind Narayanan and Sayash Kapoor. 2024. AI companies are pivoting from creating gods to building products. Good. AI Snake Oil newsletter. https://www.aisnakeoil.com/p/ai-companies-are-pivoting-from-creating. To be clear, many tasks for which the use of agents is envisioned, such as booking travel or providing customer service, are far less consequential than driving, but still costly enough that having agents learn from real-world experiences is not straightforward. Barriers also exist in non-safety-critical applications. In general, much knowledge is tacit in organizations and is not written down, much less in a form that can be learned passively. This means that these developmental feedback loops will have to happen in each sector and, for more complex tasks, may even need to occur separately in different organizations, limiting opportunities for rapid, parallel learning. Other reasons why parallel learning might be limited are privacy concerns: Organizations and individuals might be averse to sharing sensitive data with AI companies, and regulations might limit what kinds of data can be shared with third parties in contexts such as healthcare. The “bitter lesson” in AI is that general methods that leverage increases in computational power eventually surpass methods that utilize human domain knowledge by a large margin. 22. Rich Sutton. 2019. The Bitter Lesson (March 2019). http://www.incompleteideas.net/IncIdeas/BitterLesson.html. This is a valuable observation about methods, but it is often misinterpreted to encompass application development. In the context of AI-based product development, the bitter lesson has never been even close to true. 23. Arvind Narayanan and Sayash Kapoor. 2024. AI companies are pivoting from creating gods to building products. Good. AI Snake Oil newsletter. https://www.aisnakeoil.com/p/ai-companies-are-pivoting-from-creating Consider recommender systems on social media: They are powered by (increasingly general) machine learning models, but this has not obviated the need for manual coding of the business logic, the frontend, and other components which, together, can comprise on the order of a million lines of code. Further limits arise when we need to go beyond AI learning from existing human knowledge. 24. Melanie Mitchell. 2021. Why AI is harder than we think. arXiv preprint. Retrieved from http://arxiv.org/abs/2104.12871, April 2021), https://arxiv.org/abs/2104.12871. Some of our most valuable types of knowledge are scientific and social-scientific, and have allowed the progress of civilization through technology and large-scale social organizations (e.g., governments). What will it take for AI to push the boundaries of such knowledge? It will likely require interactions with, or even experiments on, people or organizations, ranging from drug testing to economic policy. Here, there are hard limits to the speed of knowledge acquisition because of the social costs of experimentation. Societies probably will not (and should not) allow the rapid scaling of experiments for AI development. Benchmarks do not measure real-world utility The methods-application distinction has important implications for how we measure and forecast AI progress. AI benchmarks are useful for measuring progress in methods; unfortunately, they have often been misunderstood as measuring progress in applications, and this confusion has been a driver of much hype about imminent economic transformation. For example, while GPT-4 reportedly achieved scores in the top 10% of bar exam test takers, this tells us remarkably little about AI’s ability to practice law. 25. Josh Achiam et al. 2023. GPT-4 technical report. arXiv preprintarXiv: 2303.08774; Peter Henderson et al. 2024. Rethinking machine learning benchmarks in the context of professional codes of conduct. In Proceedings of the Symposium on Computer Science and Law; Varun Magesh et al. 2024. Hallucination-free? Assessing the reliability of leading AI legal research tools. arXiv preprint arXiv: 2405.20362; Daniel N. Kluttz and Deirdre K. Mulligan. 2019. Automated decision support technologies and the legal profession. Berkeley Technology Law Journal 34, 3 (2019), 853–90; Inioluwa Deborah Raji, Roxana Daneshjou, and Emily Alsentzer. 2025. It’s time to bench the medical exam benchmark. NEJM AI 2, 2 (2025). The bar exam overemphasizes subject-matter knowledge and under-emphasizes real-world skills that are far harder to measure in a standardized, computer-administered format. In other words, it emphasizes precisely what language models are good at—retrieving and applying memorized information. More broadly, tasks that would lead to the most significant changes to the legal profession are also the hardest ones to evaluate. Evaluation is straightforward for tasks like categorizing legal requests by area of law because there are clear correct answers. But for tasks that involve creativity and judgment, like preparing legal filings, there is no single correct answer, and reasonable people can disagree about strategy. These latter tasks are precisely the ones that, if automated, would have the most profound impact on the profession. 26. Sayash Kapoor, Peter Henderson, and Arvind Narayanan. Promises and pitfalls of artificial intelligence for legal applications. Journal of Cross-Disciplinary Research in Computational Law 2, 2 (May 2024), Article 2. https://journalcrcl.org/crcl/article/view/62. This observation is in no way limited to law. Another example is the gap between self-contained coding problems at which AI demonstrably excels, and real-world software engineering in which its impact is hard to measure but appears to be modest. 27. Hamel Husain, Isaac Flath, and Johno Whitaker. Thoughts on a month with Devin. Answer.AI (2025). answer.ai/posts/2025-01-08-devin.html. Even highly regarded coding benchmarks that go beyond toy problems must necessarily ignore many dimensions of real-world software engineering in the interest of quantification and automated evaluation using publicly available data. 28. Ehud Reiter. 2025. Do LLM Coding Benchmarks Measure Real-World Utility?. https://ehudreiter.com/2025/01/13/do-llm-coding-benchmarks-measure-real-world-utility/. This pattern appears repeatedly: The easier a task is to measure via benchmarks, the less likely it is to represent the kind of complex, contextual work that defines professional practice. By focusing heavily on capability benchmarks to inform our understanding of AI progress, the AI community consistently overestimates the real-world impact of the technology. This is a problem of ‘construct validity,’ which refers to whether a test actually measures what it is intended to measure. 29. Deborah Raji et al. 2021. AI and the everything in the whole wide world benchmark. In Proceedings of the Neural Information Processing Systems (NeurIPS) Track on Datasets and Benchmarks, vol. 1. https://datasets-benchmarks-proceedings.neurips.cc/paper/2021/hash/084b6fbb10729ed4da8c3d3f5a3ae7c9-Abstract-round2.html; Rachel Thomas and David Uminsky. 2020. The problem with metrics is a fundamental problem for AI. arXiv preprint. Retrieved from https://arxiv.org/abs/2002.08512v1. The only sure way to measure real-world usefulness of a potential application is to actually build the application and to then test it with professionals in realistic scenarios (either substituting or augmenting their labor, depending on the intended use). Such ‘uplift’ studies generally do show that professionals in many occupations benefit from existing AI systems, but this benefit is typically modest and is more about augmentation than substitution, a radically different picture from what one might conclude based on static benchmarks like exams 30. Ashwin Nayak et al. 2023. Comparison of history of present illness summaries generated by a chatbot and senior internal medicine residents. JAMA Internal Medicine 183, 9 (September 2023), 1026–27. http://doi:10.1001/jamainternmed.2023.2561; Shakked Noy and Whitney Zhang. 2023. Experimental evidence on the productivity effects of generative artificial int
Proving you can insure just about anything, the University of Illinois has taken out a policy to mitigate the risk of a decline in Chinese student enrollment.
Context bombing" tricks hacking agents into shutting down before they can do harm.
Ludicity AI Mania Is Eviscerating Global Decision-Making Published on July 18, 2026 Note: This has been cross-posted to my company's blog, in case you think there is some use in sharing with someone in a format that looks more authoritative. Link here. I strongly believe there are entire companies right now under heavy AI psychosis and it’s impossible to have rational conversations with them about it. I can’t name any specific people because they include personal friends I deeply respect, but I worry about how this plays out. – Mitchell Hashimoto, of HashiCorp and Ghostty fame Over the past year, I’ve run point on all of our company’s sales, led the technical components of all but two of our engagements, and over the lifetime of this blog have had something like 300 catchups with professionals from around the world. This has ranged from people on the ground in niche service industries to executives at Fortune 500 companies1. Because of this, I've had a front-row view to our collective institutions across both the private and public sector undergoing breath-taking mass psychosis. This essay is an attempt to describe the bizarre dynamics that are currently at play, as I am in the rare position where my wellbeing is not contingent on paying lip service to madness, and to reassure the people trying to survive amidst all of this that they are not crazy. The reality is thus: the people in charge either have no plan, or see no path forwards other than keeping their heads down. Not at banks, not at hospitals, not in our government institutions. The world’s organisations have been captured by people in the throes of frothing excitement, and saner people who now live in a state of constant commingled fear and frustration. I. AI Investments Are Generally Total Failures Reading this while working for a division that pivoted to provide interfaces for agentic workflows, only to discover that only ten users had ever touched the products we made for agents, only to pivot again to support for agentic workflows, which has a lot of competition because every company has to do something agentic now and there's only like four things you can do in that space, is bracing. – An editor of this essay Are companies actually seeing massive productivity gains from their AI adoption? Does any of this sordid affair make sense? This should be an easy question, but it is surprisingly hard to get a straight answer to it. Executives that tell the press that their company has gone insane will quickly find themselves removed from their positions. Employees who are honest will find themselves fired in short-order, or “randomly” selected for a round of layoffs. In fact, it is in the interests of almost every actor in the space – boards, executives, employees, vendors, consultants – to obfuscate and misrepresent the success rate of AI projects. Many publicly traded companies are putting out announcements about their AI productivity gains when I know for a fact that the businesses have done nothing other than purchase Copilot licenses and declare victory. Yet we need to know if these projects are panning out – if the total focus on AI as a core tenet of business strategy is succeeding at a reasonable rate, then a discussion about the relative risk and reward is warranted. Unfortunately, we live in a dark timeline. All of the AI projects we have observed as a team are failing. Every single one – we have seen 0% success in a year and a half, not only amongst projects we have been asked to participate in2, but even within projects that we have observed in passing while doing totally unrelated work. Even if you grant that AI tooling accelerates specific workloads, the method and scale of the current investments is senseless. Frequently the failure is not related to AI itself, but rather that companies are terminally bad at running software projects effectively, and as I have remarked previously, AI projects are subject to all the failure modes of normal projects plus you can get everything right and then still fail because of the method's novelty. Very few companies are so good at shipping software that they can afford the extra risk profile. Often enough, though, it’s an actual failure in what LLMs can accomplish. The most common version of this, being rolled out across businesses around the world, is the internally-facing chatbot, or for the more daring company, the customer-facing chatbot. The story is always the same. For the former, I’ve never seen substantial internal uptake from inside a business. Employees don’t use internal chatbots because companies tend to have low-quality documentation and an LLM is not psychic – it can only know things that have been written down and made accessible. For the latter customer-facing applications, I have rarely had a pleasant experience as a consumer, with perhaps the exception of live transcription during medical appointments – hardly something worth pivoting an entire organisation around. In both cases, project leaders are very careful to avoid tracking basic metrics, such as whether the tools are being used at all, or they track metrics that are easily gamed. For example, my last consumer interaction was attempting to get help from Mitsubishi following an automotive failure, where a very polite robot asked me to describe the problem and that I’d receive a call back as soon as someone was available. This was the single most competent implementation of such a project I’ve seen in the wild, in that the voice was natural sounding, responded quickly, was clearly “live” in production, and promised a swift resolution. That was six months ago, and I did not, in fact, get a call back. When Mitsubishi did not call me back, what happened? Did that request just go into the void, showing one less incident for the year? Does it appear that the phone bot resolved my query without the need for human intervention? All we know is that it didn’t show up as an error, or I’d have received a call. I’m sure it looks great in all sorts of ways except the one that matters, which is that I was planning to buy a car and decided not to buy another one of theirs. For this reason, our team has quickly learned while on an engagement not to ask anything about ongoing AI projects in any context – by the time that project has started, it is too late for the management team, and intervention is not possible until a crisis point is inevitably reached. There is no conceivable positive outcome. The failure rate is so high that even basic inquiry leaves us in an untenable position. Any coherent question about how it’s going, what the goal is, who is using it, constitutes an inadvertent attack on the chain of command responsible for the work because there are no good answers to anything. Even in rare cases where my interlocutor has stated that things are going well (usually while the project is still mid-flight and failure has not had a chance to manifest), it is generally obvious that they are doomed, but at least in these cases I can simply agree and then go home to scream into a pillow for six hours straight3. All of this is to say that I am very confident that almost every report at a company about “massive AI productivity gains” is untrue as a matter of brute fact. Even if some companies are seeing clear gains, this is the exception, not the norm. With that assumption in place, we can talk about the dynamics at play, and how it has become impossible for many organisations to stay focused on things that actually matter to their long-term (or even short-term) health. II. Heretics Will Be Shot It has become outright dangerous to even raise the possibility that AI might not be the solution to a problem, let alone be the sole focus of a company’s entire strategy. In every sufficiently large business we have observed (say, with 500+ employees), we have noted that continued advancement, and increasingly continued employment, has started to require repeated professions of belief in the transformative power of AI for said business. I am not talking about providing ideas about how to use AI in the business – I mean religious profession, declarations of faith. Overwhelmingly these statements are made by non-technicians, though it is not uncommon for technicians to emit deranged statements to curry favour. There have been several occasions where I have seen someone, apropos of nothing, blurt out almost word-for-word “AI is changing everything”, only to concede moments later that their organisation does not currently use LLMs for anything, and indeed, that they cannot name a single thing that has changed other than they get some use out of ChatGPT (frequently the free-tier). In one extreme case, I have seen an executive confess that they had never even used ChatGPT or any AI tool in their life, immediately after producing a technical strategy for an organisation with $2B+ in revenue which was entirely centered around AI. Initially these statements were so absurd on their face that I thought it was some cynical ploy to achieve thought leader status, and there are certainly some people doing this – I have had it admitted to me. But the broader reality is so much worse: people who have no background in the technology at all actually believe what they are saying. As a general rule you should avoid getting into business with a liar, but if you must, you can at least reason with them even if only in private. A true believer is much more threatening because they are impervious to even inducement by self-interest. The turning point in my belief was watching someone with a spectacular amount of money on the line fire their highest performers because they were achieving that performance without LLMs. When an employer publicly talks about AI innovation, we have to ask ourselves if they’re simply trying to manipulate the market or customers. When they privately commit to strategies like this with their own money at stake, with no attempt to communicate that strategy to external clients, I can only assume they really mean what they’re saying. A while ago, I wrote “Contra Ptacek’s Terrible Article On AI”, which was focused on the fact that many of Ptacek’s points in his own essay “My AI Skeptic Friends Are All Nuts” were internally inconsistent4. But on the crux of the matter, we are actually in total agreement, because he opens his essay with this: Tech execs are mandating LLM adoption. That’s bad strategy. Which is to say that we can sidestep arguments about the precise utility of LLMs entirely and we’re left in a very simple place – it is entirely obvious to both myself and Ptacek, two people that are coming at this from fairly opposed views, that people are being really, really stupid about this, and that organisations are demanding bizarre workflow constraints from their specialist staff.5 These mandates have led to extremely strange places. Several of my peers now “AI-wash” their work, meaning that even when they can perfectly competently execute on their jobs to the satisfaction of their management teams, said managers are unhappy if the engineers haven’t used AI in the work… so now they’re lying about using LLMs even in contexts where their professional judgement is that they aren’t the appropriate tool. They just do the work, the same way they have for decades, and say Claude did it. Others are being measured on their AI bills with “token leaderboards”, where higher is better because I have evidently fallen into the pocket of Hell where the demons torment me by doing elaborate impressions of absolute fucking morons, so the people hired for their freakish ability to perform system optimisation do the obvious thing. They set the LLMs prompting themselves in a semi-plausible loop in case someone inspects the token consumption and then they watch Netflix. Not a single one has been caught, even when their own assessment of the output is that it isn’t suitable for deployment. Checking out a parallel copy of our Go repository and telling the AI to rewrite the whole thing in Zig while I work on something else just so I can keep my job. I hate this shit so much. My job has usage tracking and quotas. I don’t use it for actual work, I just spin it up and disregard the output. – An actual software engineer In fact, the only people I know of to be fired over this whole thing are people that have expressed visible doubt about this organisational strategy, which again, even Ptacek thinks is transparently dumb. The net result is that everyone has learned very quickly to praise executives on their visionary AI prowess, or they will be gunned down in the proverbial streets. III. AI Demos Are The Mind-Killer Bless me, Father, for I have sinned. It has been ∞ days since my last confession. I accuse myself of the following sins: One of the main pieces of infrastructure we deploy at our clients is an analytics-focused database called Snowflake – for a typical business, the bill is tiny because it’s a pay-as-you-go situation and we can process all their data in one minute a day, you get a very hands-off deployment, and in short it has many characteristics that are very pleasant for our work. One of the features in Snowflake that we don’t use is called Cortex. Cortex is their AI chatbot layer, with the ability to plug into metadata (for non-nerds, descriptions of your data, like what a column in a spreadsheet means) and query a company’s database autonomously. In theory, you can ask a question like “What was our revenue for last week?” and it will spit out an answer. It is not really suitable for production usage. From memory, the last time I was given a presentation on it, by actual Snowflake staff, they reported that ideal configuration results in something like ~92% accuracy due to the complexity of data at a large business (see: probably best-in-class for these tools, but imagine your CFO having one in every ten of their numbers be outright wrong) and there were serious issues with managing deployments. Nonetheless, it can be used to produce some very flashy demonstrations. On several occasions, we’ve been exposed to folks that have been sort of lukewarm on our main offerings, but they really, really wanted to use AI to perform a natural language query on their data. And we thought “Okay, if you really want to see it, maybe we can caveat this appropriately and show you what it might look like.” This was a terrible mistake. It backfired in the most predictable way imaginable – every lukewarm client that saw the chatbot in action, even with us telling them that it was not going to accomplish what they wanted, wanted to buy it immediately. Every other consideration, including millions of dollars that we could plausibly help them achieve by non-AI means, was swept aside. It was like a dark and terrible force seized control of their limbs, plunged their hands into their own chests, and presented their still-beating credit cards to us in grim supplication. We were so mortified by the inexplicable shift in energy that we (wisely) declined to take the money and ended the sales process, and soon thereafter removed Cortex from our list of demonstrations. It would have been too irresponsible to exploit this gap in their reasoning, and frankly, it was already irresponsible to have even run the demonstration – doctors don’t walk around showing off cool pills that they’d never prescribe. Watching the total 180°, that shift from ice-cold to red-hot buying frenzy, was a deeply unsettling experience. It was personally uncomfortable to see people that clearly didn’t gel with us interpersonally suddenly dying to enter an ongoing relationship, but more broadly uncomfortable because for a brief moment I began to understand what is happening in sales meetings around the world. There was no warning I could have given that would have made them refuse to buy the damn thing – their appetite was as large as their budget could stretch, and some part of me wonders if this is because they knew that their ravenous hunger would be present in their own customers. They’d just buy it from us, then pivot right to a larger company and mind control their leadership team until the buck finally stops with the loser that needs to justify the expense. The main protection against this seems to be that the median vendor is so bad at their jobs that we had presented the first even somewhat-working products these people had seen, and this included an ASX-listed company that was already bragging about their AI usage. It took our team two hours to produce something that was frankly not that good – basically just typing text descriptions of data into a web browser – and it was still better than anything the leads had seen because they had nothing to show for all the investment. In fact, we have been forced to opt out of every sale where the lead has expressed anything beyond the most fleeting curiosity in the use of AI in their business. I don’t mean that we’ve heard that they’re interested in AI and elected to drop the contract on moral grounds. I mean that, over the course of the engagement, these people have exhibited a pattern of behavior that has made it near-impossible to sell to them without incurring reputational and legal risk, and are furthermore crafting management environments that I can only describe as cultish, ineffective, and “please dear God, do not let it be on earth as it is on LinkedIn”. IV. Executives, Game Theory, and The Emperor’s Clothes The good news is, CISOs are used to having to protect the business from their hare-brained initiatives, and this one isn’t really that different, except that there’s a cult-like atmosphere to it that you didn’t see with, say, the cloud. It almost doesn’t matter whether you embrace the initiative or not; there’s work to be done to manage the risk, so that’s what you do. From talking to CISOs everywhere, I would say most of them are quietly skeptical but afraid to speak up. – Career CISO and well-known speaker that asked to remain anonymous Despite the substantial prevalence of true believers, many of the people running large AI initiatives, or making public statements about them, do not believe what they are saying. There are “heads of AI” who read this blog, at companies with $1B+ in annually recurring revenue, who have written in to say they believe their job is totally fraudulent but it was the only promotion pathway remaining at the organisation. On a trip overseas, I had the privilege of a meeting with one of the Fortune 500 executives mentioned at the beginning of the post, who will remain anonymous so that they are not executed by firing squad by their board. As we were chatting, it became clear that they were very switched-on and technically competent, and they also happened to be at a company that had committed to the usual battery of exorbitant claims about their recent innovations – we’ve 100x’d our productivity, AI is the future of everything, I am but a vessel for OpenAI to make love to my wife. You know, normal things. But since I had them there without any microphones around, I asked why this was being repeated without opposition. Was it just sales fluff? The answer was a lot more interesting. It was partially ridiculous sales material being delivered to an easily excitable audience, but this was not the dominant factor constraining honesty. Executives at their customers were saying absurd things about achieving 100x productivity, and this meant that if any executive at the vendor said that these gains were not plausible, it would undermine the credibility of the customer’s executive, be perceived as an attack (or heresy), and possibly result in an enterprise contract cancellation. And getting enterprise contracts cancelled because you wanted to opine on something that doesn’t really matter to your organisation’s mission is a great way to get fired. But this company was also a major player, of the kind that signs enormous enterprise contracts with other companies. So presumably there is another vendor that has sold to them, and their CEO is worried that saying something sane will contradict this executive, and very quickly we can see how we can have executives around the world nervously pointing guns at each other, not wanting to be shot first but also watching everything gradually spiral out of control6. This is to say that we’re facing a coordination problem around executives being honest around the AI gains they’ve witnessed – if they co-operate, they keep their jobs. If they defect, they will possibly be fired by their embarrassed peers (who have now been implicitly called liars, cowards, or incompetents) and then replaced with someone that will toe the line anyway. If they could all admit the truth at once there might be some hope, but there is no way to coordinate that event. This sounds deeply concerning, but it is worth noting that it means that some executives who are emitting nonsensical statements are not as dull as they might seem at first – they’re in a fraught political environment, where they are surrounded by many people that are gunning for their roles, and subject to the whims of a board that is undergoing similar pressure. Against all the dictates of reason, I have presented on navigating AI hype to people on S&P 500 boards7 and they are in exactly the same situation – the main comments I remember from the session were board members admitting they were skeptical, but expressing anxiety that their positions were contingent on demanding AI investment. One of them commented “investing this early seems like risk without much upside”. About two years later, I can see now that their decade-old multi-billion dollar organisation is now branded as “AI-native”, whatever the hell that means. V. You Must Be This AI-Native To Ride All of the above converges on the state that we find ourselves in now, where effective decisionmaking has ground to a halt. Collectively, what started as a few people undergoing either destabilising psychological events or being caught up in hype has now resulted in an environment where leaders cannot speak honestly about their beliefs on how best to guide organisations, for fear of being removed, creating a sort of distributed government by assassination. This means that the least sensible recommendations are going totally unchallenged, resulting in employees being evaluated on totally gameable metrics such as “money spent on AI”, and those employees must play along to avoid being terminated. This has also created an insatiable appetite for purchasing “AI” solutions, which target both true believers that will believe implausible claims, and also non-believers that cannot decline the purchases without having their commitment to the cause coming into question. This means that all offers that are subject to internal politics at an ideologically captured organisation must include AI alignment, even if the value proposition is patently ambiguous. My assessment of the market so far is that a substantial component of the outburst of AI projects are actually non-AI projects with an AI element slapped on after the fact to pass the purity test. For example, I recently witnessed an organisation handling a database migration from an Oracle database to Snowflake – instead of handling the migration directly, the vendor bolted on a preliminary phase which involved trying to get an LLM to automate the translation of the Oracle-flavored SQL to Snowflake-flavored SQL. When the project failed (due to issues getting enough permissions to automate the work, not because an LLM can’t do something that easy), the vendor simply started handling the translation by hand but the company billed it as an AI-driven success because some inconsequential portion of the SQL had been translated by AI before being pasted over. What was actually purchased? A totally standard database migration to help an executive meet the strategic deliverable of decommissioning a system prior to license renewal. What was sold to their superiors? “I allocated a substantial percentage of my budget to AI and it helped me accomplish my mandate.” True AI projects, of the kind that is driven by an LLM as the sole mechanism underlying it, where the project can clearly fail to deliver specific numbers, are actually very rare. We mostly see them in the context of startups, and frankly we have stopped engaging with them because we kept getting to the end of the sales conversation and finding out they wanted us to build the product that they were marketing as completed. However, some projects simply do not have an easy way to tack on the AI label, or the person advocating for them either does not want to lie or has not understood that lying has become necessary. In all cases, this either kills the request for funding outright, or adds a pervasive and intractable drag on all communications, as every request must be worked and re-worked until it is “AI enough”. Failure to comply will either result in denial or, in many cases, a demand from a true believer to know why the extra work “can’t be done with AI”. Many companies have actively publicized that this is their new hiring policy – when a member of staff requests additional headcount, they must demonstrate that they have tried to use AI first. The part that’s being left out is that if you say you used AI and still need the help, you will be labelled “bad at AI” and potentially laid off. The net result of this is that almost every large organisation that I am aware of is no longer able to focus on anything important, unless they are one of the (very) few organisations where AI happens to address their highest priorities. They cannot buy sensible software, hire competent talent, communicate honestly with executives about the state of projects, or undertake any sort of sensible initiative. VI. Navigating AI Mania An emptiness falls through you As you realize what this means You're starting to feel what I feel Now you've seen what I've seen – So Sick, Domesticated Incels This is an unfortunate situation to be in, but it will pass eventually. I’ve learned a lot about the latent insanity that we have inculcated in our leadership strata, and unfortunately those traits will persist long past the current bubble, merely awaiting another similar reactivation trigger – and some organisations will stay captured until they have totally collapsed, in the way that not everyone has successfully moved away from the dreadful blockchain affair. That’s something to write about for another time. What I wanted to get to were some thoughts on surviving the immediate crisis, either by directly making systemic improvements or by holding onto your sanity. I’ll start with the “making improvements” part, because that’s the situation I find myself in the most frequently. When You Have Another Objective We’re going to do a lot of sucking it up and smiling here. This section assumes that you are trying to achieve some goal that isn't repairing the organisation's manic stance, but either trying to course-correct a specific project (and possibly risk getting fired as either a leader or consultant) or achieve some totally unrelated goal. Where possible, when raising issues, do not have conversations about the state of AI projects in group settings, as this creates a dynamic where each individual member of the group is worried about outing themselves in front of their peers. Arrange for one-on-one settings. Make it clear that you are willing to countenance that the current AI environment is frothy, and that you will keep opinions unidentifiable when raising them elsewhere. Be extremely aware that the most outspoken people can be identified by their peers, so take care to avoid exposing your sources by, e.g. direct quotes. In the event that only a small minority (say, one person in a group of six people) is willing to speak out, it might be worth giving up and moving on to a patient that has better chances. For ongoing projects, an effective trick that I believe I picked up from Secrets of Consulting is the anonymous poll, where you can ask individuals to rate their opinion of an AI project’s success chances on a scale of 1 to 10. The typical split I have observed is half of those involved rating the project at a 3/10 and others at around an 8/10 – a clear bimodal split on a project that was already three years late. Bringing this data to a CEO can be an effective method of pointing out that some information is clearly being hidden from them on the state of the project. Always involve people on the ground. The only source of data on whether projects are succeeding or the investment is going anywhere are the people that use it for their day-to-day activity. Care must be taken to bring them into the environment where they are treated with respect (all sufficiently large companies have people that view subordinates as not-quite-real-people). It is not uncommon to uncover worldview-shaking information in short order – with one client, we uncovered that staff were totally unaware they had been given licenses for AI tooling, which cast into doubt all productivity claims. Do not question the broadest claims about AI. I cannot emphasize this enough. If someone says “AI is changing everything”, just let it pass if your goal is to fix an object-level problem rather than challenge the reality at the institution. The challenge can only come after you have gained the trust of the most senior person involved. Trust is gained over a meal in private where you assuage their anxieties, not by embarrassing them in front of peers. Remember that you do not know what statements have been emitted prior to entering a room. There will sometimes be people that have publicly committed to statements like “I am 100x more productive than I was last year”, and some may even wish they hadn’t said that but are too embarrassed to walk it back. In an untested room, common sense like “LLMs should not
You know how a bunch of people went to a Bored Ape dance party with a bunch of UV lights, but it turns out the UV lights were industrial equipment that radiated UV light at a much higher level than what was safe because the organisers didn't give a fuck, so a bunch of people went temporarily blind and got burns all over their body? I wonder how they're doing. Not in a "point and laugh at the obvious shit-stirrers getting their comeuppance" kind of way, because despite my disdain for the NFT craze and the sneering contempt NTFbros had for everyone outside of their community, there was zero excuse for the negligence shown by the party organisers and nobody deserves to go blind because they went to a dinky fuck-ass little dance party like these guys did. More than anything, I want to know if there were any long-lasting health effects due to the burns and I'd be fascinated to hear about how the affected people feel about the NFT craze and Bored Ape specifically in the aftermath of the incident. Chances are no info has ever come out, but if anyone was gonna know about any follow-up information, it'd be this sub. submitted by /u/Tracy-Jaccs [link] [Kommentare]