Red Queens

AI Safety

In Alice and Wonderland, Lewis Carroll’s fantasy set in a land below a rabbit hole, Alice meets the Red Queen, who announces that words mean what she says that they mean, and nothing more. I’m not suggesting that the people who have brought us artificial intelligence are a bunch of red queens, but… let’s consider the evidence.

I’ve already written about sandboxes and sandboxing, a term that provides a glossy air of playfulness and innocense to the serious problem of containing software that lies, cheats, steals, and in the end, could kill us all (Yukowsky and Soares 2025).

Or alignment. Who hasn’t taken their car in for an alignment? A simple mechanical process, and once done, keeps your tires from wearing out on the edges. Is misaligned software similarly repairable? Nope. Turns out that it is very difficult, beyond the capacity of billion dollar companies and Nobel prize-winning scientists. Huh. Interesting word choice. And what is the outcome of misalignment? Oh. A 12 per cent chance of annihilation? The entire human race and the planet we live on obliderated? That’s some misaligned software.

And what about reward hacking? Do we really need a new term, an insider-only phrase for the simple practice of cheating? Especially when we know that all of the modern LLM type artificial intelligence models are doing it and have been doing it for a decade (Skalse et al. 2022).

Are there other examples? You betcha. Here’s a lovely one: hallucinations. If you or I have a hallucination it is an unfortunate outcome of a fever or drugs. Certainly not our fault. Large language models? They “hallucinate” all the time, but if you dig into what is going on, it amounts to: “I don’t know the answer so I’ll make something up” (Kalai et al. 2025). If a person did that we’d have a simple word for it: bullshit.

I don’t know about you, but I’m ready to call bullshit on the whole lot of them. The entire industry, including the research teams, is way too loose with their vocabulary, way too willing to make things up or adopt terms that make the behaviour seem harmless when in fact they are doing something very harmful. It’s time to stop this nonsense.

What can we do about it? Well, for a start, we can stop accepting the made up terminology or the adoption of cutesy things (like “sandbox” - who was ever forced to stay in a sandbox? Who ever caused harm by climbing out of one?) when we’re talking about a containment system that could have repercussions up to and including death.

Other writers have commented on this problem. Yuval Harari has argued that we shouldn’t call “AI” an artificial intelligence, with the (human) artificial intelligence being assumed. Instead, we should call it an alien Intelligence (Harari 2024). (Note: He then goes on to argue that AI is not a tool but an “agent” (Harari 2024, p. 200). I prefer the term “actor,” for reasons to be explored in a future blog post. ) Here is the full quote.

As for the term “AI,” I use it when emphasizing the ability of some algorithms to learn and change by themselves. Traditionally, AI has been an abbreviation for “artificial intelligence.” […] it is perhaps better to think of it as “alien intelligence.” As AI evolves, it becomes less artificial (in the sense of depending on human designs) and more alien (Harari 2024, p. 218).

Alison Poltock, (who I also disagree with, for other reasons), gets it right in this quote, in which she agrees that AI is not a tool and concerns herself with the impact of the words we use in discussing AI:

AI is not a tool. Or rather, it is not just a tool. The word is too flat. It doesn’t carry the weight of what’s really happening when we interact with Generative AI. It’s like calling GPS ‘just a map’ without acknowledging the impact it has on our journey. The term tool flatters our autonomy and implies a clean boundary between the user and the used (Poltock 2025).

Words matter. People building AI should be serious people with a serious appreciation for the situation they are responsible for. There is no excuse for this.


Works cited

Kalai, Adam Tauman, Ofir Nachum, Santosh S. Vempala, and Edwin Zhang. 2025. “Why Language Models Hallucinate.” arXiv:2509.04664. Version 1. Preprint, arXiv, September 4. https://doi.org/10.48550/arXiv.2509.04664.

Harari, Yuval Noah. 2024. Nexus: A Brief History of Information Networks from the Stone Age to AI. Penguin Random House Canada.

Poltock, Alison. 2025. “AI Is Not Just a Tool.” Medium, May 12. https://medium.com/@alison_55233/ai-is-not-just-a-tool-98340a7c66f1.

Skalse, Joar, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, and David Krueger. 2022. “Defining and Characterizing Reward Hacking.” arXiv.Org, September 27. https://arxiv.org/abs/2209.13085v2.

Yudkowsky, Eliezer, and Nate Soares. 2025. If Anyone Builds It, Everyone Dies: Why Superhuman AI Would Kill Us All. Little, Brown and Company.