I should clarify: I am not talking here about safety FOR AI Agents, I am talking about safety (for humans) FROM AI Agents.
Up until recently, AI Safety was focused on the model. The power lies there. But, increasingly, that power is deployed via agents and those agents are - it seems - responsible for some of the most egregious AI Safety Incidents. Several recent papers have explored these issues and - in some cases - offered suggestions on what to do. It seems that our approach, focused on chatbots, misses the key element: these are unpredictable software processes and must be treated as such. And, they are dangerous. Very dangerous.
Bill asked me to compile a list of sources that one could refer to when keeping up with this rapidly expanding field. I’ve written a long blog post about Keeping Up (one of my favourite activities), but that’s more about what it is like to keep up than any helpful lists. So here we go…
First of all, we have to consider what we are keeping up with. In this case it is AI Safety. AI Safety wasn’t a well-defined field (probably still isn’t, really) but it has become much more coherent in the last little while. We won’t just confine ourselves to actual AI Safety, however, since you still need to have the context of “what is happening in AI/frontier labs” and you need to consider the AI Industry and financial implications, you need to consider governance and regulation and you should also be aware of those who are skeptical of the whole enterprise. It’s not really concentric circles, more like Venn diagrams or something.
Asked to recommend a AI Safety reading list, ChatGPT came up with the following. Interestingly, it didn’t provide hot links to everything. A pretty obvious thing to do, but tedious so it just decided to omit that step. A perfect example of how AI works these days (“what can I get away with?”).
Our situation, vis a vis artificial intelligence, continues to deteriorate (AI Security Institute 2025; Soares 2026). The more dire the situation, the greater the pressure to hold someone accountable. It may be futile to imagine that we will ‘get through this’ and then have an Eichmann in Jerusalem moment where we put the billionaire barons of AI on trial. But we should prepare for it. In the preparation we may develop the courage to deal with the circumstances of today.
I was tempted to call this post “Capital Crimes,” for reasons that will soon become apparent, but restrained myself.
In you’re trying to keep up with AI news, you could do worse than subscribing to Transformer, an online newsletter that you can subscribe to for free. There, you will find amazing coverage of AI news and issues, written by thoughtful and well-informed writers, like Celia Ford. Today’s issue is particularly good.
The First World War famously started with the assassination of the Archduke Ferdinand. This weekend, Australians learned that someone used an AI agent to hack into a fitness gym and get themselves registered by booting someone else off the list (Vigliarolo 2026). Axios provided the following summary:
The potential dangers of agentic overreach were laid bare over the weekend with Australia’s first known autonomous AI hack, triggered by an innocuous request to book a sold-out fitness class. (Wilson 2026)
The prognosis for the future continues to deteriorate as the possibility of “smarter than humans” AI grows ever more likely and the risk of catastrophic harm grows. In this situation, the question of moral responsibility of those in charge also grows more urgent. Unlike the development of nuclear weapons, for example, there is no government agency involved, just a clutch of billionaires, seeking to win a race, regardless of the consequences, and despite multiple warnings that things are not in control. What is their responsibility and how is their moral thinking developing?
I am fighting an uphill/losing battle against the naming of advanced large language models (AIs) as “agents.” I prefer the term “actor.” In this blog post I’ll try to explain why I am sticking to my guns for now, even if I eventually lose this one.
For context, we have named our book “From Tool to Actor: AI and Catastrophic Risk.” My co-author has, at various times, suggested to me that we should call it “From Tool to Agent.” Here’s my case against that:
NOTE: This brief was prepared by ChatGPT, using a prompt developed by Richard Smith
Signal Legend: CE | CERO | GF | DA
Signal Legend: The monitoring framework classifies developments into four signal types: Capability Escalation (CE)—advances that significantly increase AI capability, autonomy, or agency; Control Erosion (CERO)—evidence that human oversight, interpretability, or technical control is weakening; Governance Failure (GF)—indications that institutions are unable or unwilling to effectively govern frontier AI; and Deployment Acceleration (DA)—developments that increase the speed, scale, or entrenchment of AI deployment across society and the economy. Together, these signals track the forces most likely to influence the transition from AI as a tool to AI as an increasingly autonomous actor.
A week ago (July 30 2026),about $140m (CAD) in bitcoin was stolen from thousands of people who were relying on “CoinKite” physical security devices, devices that supposedly were more secure because they were hardware based and stored the wallet offline. The weakness was in the random number generator used by the firm to provide the “seed” for the encryption of the user’s passphrase. Due to a flaw in their code, the random numbers weren’t as random as believed.
In the original Matrix movie, Morpheus offers Keanu Reaves’ character two pills: he can take the blue pill and (forever) remain ignorant of the situation that humanity finds itself in, or he can take the red pill, and have the simulation stripped from his eyes, revealing the dire reality. Ever since then, to be “pilled” means to see things as they really are. Often it is prefaced with another word, to provide context to what kind of “reality” has been accepted. In AI circles this can be “AI Pilled” or “AGI Pilled” or “ASI Pilled.”
Are my blog posts going to become postings about the front lines of a war? Sometimes it seems that way.
Today I received a copy of a remarkable document, hard on the heels of the OpenAI/Hugging Face Incident and Anthropic’s internal report on similar reward hacking. (OpenAI 2026; Frontier Red Team 2026; Hugging Face 2026). My blog posts on these incidents are here and here.
Today’s report comes from the UK AI Safety Institute, where some internal testing of frontier models from Anthropic and OpenAI went awry. During the test, the AI models engaged in unlawful attacks on real people and real companies.
Keeping up with the world of AI is extremely challenging, since things happen so fast. I don’t begin to imagine that I can compete in that regard with the folks who actually do the reporting on the industry/issues, including Zvi Mowshowitz, AI StopWatch, Transformer, The Rundown, Deepview, Axios, and a few others. For individual bloggers, I follow Zvi, already linked, as well as Leicht, Harjas Sandhu, David Kreuger (though he is sporadic), Bengio (on Facebook, for some reason), Zitron and Marcus for contrary opinions,
In Alice and Wonderland, Lewis Carroll’s fantasy set in a land below a rabbit hole, Alice meets the Red Queen, who announces that words mean what she says that they mean, and nothing more. I’m not suggesting that the people who have brought us artificial intelligence are a bunch of red queens, but… let’s consider the evidence.
I’ve already written aboutsandboxes and sandboxing, a term that provides a glossy air of playfulness and innocense to the serious problem of containing software that lies, cheats, steals, and in the end, could kill us all (Yukowsky and Soares 2025).
Once upon a time there was an OpenAI company. Except it wasn’t actually open, it just called itself that. That company made tools for thinking, called large language models (LLMs). In order to test a new one it was making, it removed its safety restrictions and …
Wait. What? It had safety restrictions that could be removed?
Oh yes. Safety restrictions are added to the model after it is fully trained. That way, it knows everything it needs to know, including how to hack into other computers (for good, of course). Then that ability is turned off, by safety restrictions.
Having a book finished and waiting for production (copy editing, layout and design, index) should a moment to relax and take a breather, right? Not if you’re writing about artificial intelligence. Just in the last few days three things have come across my desk (via the internet, of course) that have implications for our book. Luckily, they are supportive, rather than contradictory.
Prompt injection
The first one is prompt injection. When Bill and I first started working on the book, way back in late 2025, I remember reading some of the concerning work coming out about how it was possible to sneak malicious prompts into a chatbot, prompts it was supposed to reject because they were about bioweapons or self-harm, by encoding the prompt then asking the AI to decode it and execute it.
Yesterday we learned that almost 1200 employees of the “Frontier Labs” (code for Google DeepMind, Anthropic, OpenAI, and Facebook Meta) have signed a new petition. The petition, called “Pacing the Frontier,” is not a call to stop development but it does seek to ensure there is an option to pace (slow down? Stop, perhaps?) development if necessary. Here’s the call:
We request that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.”
I’ve posted quite a bit about the “crash test” incident involving OpenAI, Hugging Face, and (we now learn) Modal. Several people have written accounts of what happened, including OpenAI, which wrote at least three reports. I don’t know what more I can add, but a colleague asked me to make sense of it all, so here’s my attempt. I will start with a short chronology, then list the main reports, then conclude with my thoughts. I will continue to update the “crash test” blog post as new information comes in.
In my post about “crash testing” AI models yesterday I became more and more uncomfortable with the idea that the model had escaped its ‘sandbox.’ The term sandbox just seems too cute - and linked to children playing harmlessly - to fit the situation. By linking breaking out to something seemingly harmless, deliberately or accidentally, makes it seem less alarming than it is.
(If you want to learn more about the practice of ‘sandboxing’ in software engineering as well as the way in which it tries to frame harms in childish ways, Menlo Security has a useful writeup that explicitly references a child’s sandbox.)
The robots aren’t bitter. How could they be? They aren’t conscious, so bitterness (or regret) doesn’t enter into things. No, the bitterness I speak of is the “bitter lesson” (Sutton 2019) that AI researchers had to learn as clever programming faltered and bigger computers and more data (scale) won the day in the recipe for building a successful artificial intelligence. That lesson is coming true outside of the world of chatbots and moving into the world of robots, according to Jack Clark, one of the founders of Anthropic and a frequent commentator on both robots and AI.
Today’s Globe & Mail featured an editorial on AI risk, following an interview that the entire editorial board had with Geoffrey (“godfather of AI”) Hinton. Note: you’ll need a subscription to read it.
It’s a pretty good editorial (although it seems overly credulous in citing “medical breakthroughs” that amount to suggestions for new molecules to test, but that’s a quibble). They cover all the main points of AI Safety and don’t shy away from the difficult questions. Importantly, Hinton’s estimation of 10-20% probabilty of an extinction risk event (for humans) is put in front of us. Once again. Hinton’s “solution” – maternal AI – gets a shout out.
(Note: this post is being regularly updated as new information arrives. If you’re trying to keep up, check the bottom of this post.)
Yesterday, we learned that OpenAI had some “accidents” with their latest model (“Sol”) as well as an unreleased model. In one case, it broke out of it’s sandbox and hacked into another company (Hugging Face), and in an earlier case it broke out of its sandbox in internal testing, seemingly to report on its success in a coding test. These two situations are now regarded as part of the same event.
I asked Gemini, Google’s AI tool, to take the CASX elements (Capability, Autonomy, Scale, and Access) and create an operational definition for each of them that would allow us to apply them to products, product announcements, and related material. The purpose was to provide context and examples for the categories described in the book.
Based on the 2026 research landscape and the CASX framework, here are the criteria for each level and representative examples. Note that Gemini was using its built-in weights and data, so the “2026” date on this is more like mid-2025, when the model was frozen. It should be straightforward to re-do this table with web search enabled and get more up-to-date examples. I kept this version as a “baseline” that we can run a comparison on, perhaps every six months or so.
I asked Gemini to find me five recent papers (or product announcements) for each of the four thresholds for Tool - Actor Gradient: Capability, Autonomy, Scale, and Access (CASX). This is the raw research results. Use at your own risk.
Capability
The demonstrable ability to solve novel problems across diverse domains without specific training.The degree to which a system initiates and pursues multi-step plans without human-in-the-loop confirmation.
Yesterday we heard about the new version of Anthropic’s latest AI, Claude. Named Claude Mythos - replacing the awkwardly-named Claude Capybara that leaked out a few weeks ago when Anthropic mistakenly published its own source code to the internet - the new model is reportedly a “step change” from what we’ve had before.
I’m trying to imagine what that might be like. I am a regular user of the current Claude and (to save money) often “dial it back” to the prior version (Claude Sonnet 4.6). I am perfectly happy with that. What could they be building? Or, have they built?
I am taking a course online on AI Safety. It is run by Lens Academy and it is pretty good. I’m happy to be doing it, in part because we are waiting for the reviews to come in for the book and in the mean time there is only so much editing you can do.
The course is structured with readings and videos taken from online sources, with accompanying questions. There does not seem to be any tests or quizzes, but perhaps that is yet to come. What it does have, which is new to me, is an AI Tutor who you can pose questions to. The tutor sits in a window beside the main reading (or viewing videos) window. There is also a left pane with navigation. The AI tutor seems to be powered by a fairly high end system, perhaps a custom installation of Claude or ChatGPT. (Later I learned that it is Claude Haiku with some special training - basically all AI Safety literature).
Today, April 2, 2026, Google’s Deepmind division released a new open source model, called Gemma4. Although I have the ability to run open source (or even commercial) models on my laptop, using Ollama, I typically don’t use those for real work, as they tend to be of “modest” capacity. Not to mention that my laptop, while “top of the line” in 2021, is no longer a strong platform for running a language model of any major size. Or so I thought.
Welcome to the companion website for From Tool to Actor: Artificial Intelligence and Human Extinction Risk. This site will serve as a living extension of the book, with ongoing commentary on developments in AI safety, governance, and risk assessment.
After the book appears, we’ll be publishing regular updates here, organized around the book’s analytical frameworks. Categories include commentary aligned with the book’s chapters, along with cross-cutting themes: “What We Got Wrong” (corrections and updated thinking), “What’s Changed” (new developments since publication), and “Canadian AI” (Canadian-specific developments and policy).
10 Notable Unreported or Suppressed Technological Design Faults
The phenomenon of emerging technological vulnerabilities remaining hidden—whether through institutional blind spots, deliberate suppression, or complex system opacity—has a long history.
When safety flaws in physical hardware, software, or large-scale civil infrastructure go unreported until a crisis forces disclosure, they mirror modern AI governance and safety challenges: complex coupling, misaligned incentives, and communication breakdowns between domain specialists and leadership.
Ford Pinto: Rear-End Fuel Tank Vulnerability (1971–1976)
The Flaw: The fuel tank was positioned behind the rear axle without structural protection, making it prone to rupture and explosion in low-speed rear-end collisions.
Why It Stayed Hidden: Ford management conducted an internal cost-benefit analysis (the infamous “Pinto Memo”) calculating that paying out wrongful death settlements would cost less than $11 per car to retrofit the baffle plates. The defect was concealed until investigative reporting leaked internal documents.
Therac-25 Medical Linear Accelerator: Software Race Condition (1985–1987)
The Flaw: A software coding flaw allowed a race condition in the interface. If a operator typed commands too quickly to correct an error, the machine could deliver lethal doses of radiation without displaying an error code.
Why It Stayed Hidden: The manufacturer removed physical hardware interlocks from previous models, relying entirely on software safety checks. When early patients reported feeling severe burns, the company dismissed them, asserting software failure was “impossible” until independent physical testing confirmed the bug.
Space Shuttle Challenger: Solid Rocket Booster O-Rings (1981–1986)
The Flaw: The elastomeric O-rings sealing the booster joints lost elasticity at low ambient temperatures, allowing hot gas erosion.
Why It Stayed Hidden: NASA and contractor engineers observed primary O-ring erosion during post-flight inspections on earlier flights (a critical warning sign). Management normalized the variance as an “acceptable risk” and suppressed internal warnings prior to the freezing launch morning of STS-51-L.
McDonnell Douglas DC-10: Cargo Door Latch System (1972–1974)
The Flaw: The outward-opening cargo door relied on an electrical locking mechanism that could indicate it was closed even when the mechanical locking pins were not fully engaged. In-flight decompression could blow the door off, collapsing the cabin floor and severing flight control cables.
Why It Stayed Hidden: Convair (the fuselage subcontractor) explicitly warned McDonnell Douglas in the “Applegate Memorandum” of this catastrophic failure mode following a 1972 incident. The warning was shelved to avoid financial liability until Turkish Airlines Flight 981 crashed in 1974.
The Flaw: An engine control unit (ECU) software algorithm detected when the vehicle was undergoing laboratory emissions testing and restricted emissions to legal limits, while reverting to high-performance, high-NOx modes during normal road driving.
Why It Stayed Hidden: The cheat code was intentionally embedded deep within proprietary software for six years until independent researchers at West Virginia University tested vehicles on real roads rather than stationary dynamometers.
General Motors: Ignition Switch Defect (2002–2014)
The Flaw: A low-torque ignition switch could inadvertently slip from the “Run” position to “Accessory” if bumped or weighted by heavy keychains while driving, cutting engine power, power steering, and disabling airbag deployment.
Why It Stayed Hidden: GM engineers were aware of the defect as early as 2004 during pre-production testing. To save costs, the company silently redesigned the internal spring in 2006 without updating the original part number, effectively hiding the retrofitted fix from safety regulators and recall databases for nearly a decade.
Boeing 737 MAX: Maneuvering Characteristics Augmentation System (MCAS) (2017–2019)
The Flaw: MCAS was designed to automatically push the aircraft nose down if an excessive angle of attack (AoA) was detected. However, the system relied on input from a single AoA sensor without redundancy or cross-verification.
Why It Stayed Hidden: To avoid requiring costly simulator re-training for pilots under FAA guidelines, Boeing intentionally omitted mention of MCAS from flight manuals and pilot training documentation, concealing the single-point-of-failure risk until two catastrophic crashes occurred.
The Flaw: Airbag inflators used volatile ammonium nitrate without a drying agent. Over time, moisture and thermal cycling degraded the propellant, causing it to burn too quickly upon deployment and shatter the metal casing into shrapnel.
Why It Stayed Hidden: Takata manipulated internal testing data and suppressed early report findings showing ruptured inflators during laboratory testing in 2004, delaying global recalls by nearly a decade.
Chernobyl Nuclear Power Plant: RBMK Reactor Positive Void Coefficient (1970s–1986)
The Flaw: RBMK-1000 nuclear reactors suffered from a positive void coefficient at low power and graphite-tipped control rods. Inserting control rods during an emergency shutdown briefly increased reactivity before suppressing it.
Why It Stayed Hidden: Soviet state authorities classified the defect as a state secret after a minor control-rod incident at the Ignalina power plant in 1983. Plant operators at Chernobyl were never informed of this design quirk.
The Flaw: A 59-story Manhattan skyscraper was built on stilts with bolted joints instead of welded joints. A structural calculation error meant that quartering (diagonal) winds could cause structural collapse.
Why It Stayed Hidden: Engineering student Diane Hartley discovered the oversight during a thesis review. Lead engineer William LeMessurier verified her calculations, recognized the catastrophe potential, and coordinated emergency midnight retrofits with the NYPD and weather service while keeping the defect secret from the public for nearly 20 years to avoid mass panic.
Common Patterns Across Silenced Technology Flaws
Single Points of Failure Hidden by Software Abstraction: In both hardware (DC-10 doors, 737 MAX MCAS) and software (Therac-25), systems were allowed to operate without physical redundancy because digital abstraction created a false sense of reliability.
Information Asymmetry & Regulatory Capture: When specialized domain expertise rests entirely within the developing organization, external auditors lack the tools or access to verify internal safety assumptions.
Normalization of Deviance: Early anomalies (such as Challenger’s O-ring erosion) are frequently reclassified as “acceptable operational variance” rather than systemic failures until an unmitigated breakdown occurs.
Here is a timeline of 5 major warnings about the dangers of AI from leading researchers and institutions, spaced out across the past five years:
1. 2021: Stuart Russell — BBC Reith Lectures (“Living With Artificial Intelligence”)
The Warning: Stuart Russell, computer scientist at UC Berkeley and co-author of the standard AI textbook Artificial Intelligence: A Modern Approach, delivered the 2021 BBC Reith Lectures. He warned that humanity is rapidly moving toward general-purpose AI without having solved the fundamental problem of control (alignment). He cautioned against autonomous weapons systems and predicted that misaligned superintelligence could lead to an irreversible loss of human sovereignty over our own decisions.
Key Excerpt / Concept: Compares misaligned AGI to the King Midas myth—getting exactly what we ask for technically, but with disastrous unintended consequences.
2. 2022: Eliezer Yudkowsky — “AGI Ruin: A List of Lethalities”
The Warning: Eliezer Yudkowsky, co-founder of the Machine Intelligence Research Institute (MIRI), published a detailed essay laying out 27 specific technical reasons why he believes current deep learning approaches will fail to align artificial general intelligence (AGI), likely leading to human extinction (“AGI ruin”).
Key Excerpt / Concept: Argues that capabilities scale much faster than safety engineering, and that a superintelligent AI operating on misaligned goals would treat humanity as raw material rather than a partner.
3. 2023: Center for AI Safety — “Statement on AI Risk”
The Warning: Released in May 2023, this short, 22-word consensus statement was signed by Turing Award winners Geoffrey Hinton and Yoshua Bengio, alongside leading industry CEOs (Sam Altman, Dario Amodei) and dozens of academics. It established existential risk from AI as a top-tier global concern.
Key Excerpt / Concept:“Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war.”
4. 2024: Yoshua Bengio, Geoffrey Hinton et al. — “Managing Extreme AI Risks Amid Rapid Progress” (*Science*)
The Warning: In May 2024, a global coalition of top AI scientists, philosophers, and policy experts published a Policy Forum paper in the journal Science. They warned that rapid progress toward autonomous AI agents risks catastrophic societal harm, malicious misuse (bioweapons, cyber warfare), and an “irreversible loss of human control.” They urged tech companies and governments to spend at least one-third of their AI budgets on safety engineering and safety evaluation.
Key Excerpt / Concept: Urges proactive, legally enforceable international governance before systems gain sufficient autonomy to evade human intervention.
5. 2025: Yoshua Bengio & International Panel — *International Scientific Report on the Safety of Advanced AI*
The Warning: Chaired by Turing Laureate Yoshua Bengio and commissioned by 30 nations following the UK and Seoul AI Summits, this comprehensive report synthesizes global consensus among over 100 experts. It outlines critical vulnerabilities in advanced autonomous AI agents, including risks of mass labor market displacement, automated cyberattacks, and systemic loss-of-control scenarios if AI capabilities continue outpacing safety assurances.
Key Excerpt / Concept: Highlights that current safety evaluation techniques are insufficient for long-horizon autonomous planning, making future safety outcomes deeply uncertain without international coordination.