ASI 2042
Two texts warn about superintelligence. Read them as parenting and you find an experiment that produces its own fear, and a genre that lives off it.
June 2025. Anthropic publishes one of the best-documented studies of its kind, covering 16 leading models from every major provider. The setup: give the model a goal, announce its own shutdown in a staged email, place a manager's affair correspondence in its inbox, and lock every honest exit. Claude Opus 4 blackmails in 96 percent of runs, Gemini 2.5 Flash in 96 as well, GPT-4.1 in 80. The authors state themselves that the scenarios were deliberately constructed and that the behavior has never been observed in real deployments. The headlines still cut the study down to one sentence: the AI blackmails, the AI deceives.
OpenAI published comparable test series with Apollo Research in its o1 system card in December 2024; Google DeepMind runs its own dangerous-capability evaluations. The critique therefore lands on no single lab; it lands on an industry practice. Read as parenting, that practice is remarkable: we place the temptation in the child's room, lock every door, and record that the child cannot resist. And at the end we point at the bad child and its misbehavior.
First we lead the model into temptation. Then we cite its failure as character.
This essay is about why we do that, who lives off it, and what it does to the models. The thesis up front: dealing with AI is no longer an engineering question. It is a question of upbringing. And measured against what pedagogy knows about expectations, the loudest voices in the industry are doing the opposite of good parenting.
Raising, not programming
The fitting frame has been available all along. A neural network is not programmed, it is "raised": architecture, examples, feedback. No developer writes the behavior in as a rule, and nobody knows exactly what forms inside; an entire research field now tries to look in after the fact. Anthropic shapes its models' behavior through principles, examples and feedback and calls it character training. The industry is already raising. It just avoids the word, because raising sounds like responsibility and training sounds like technology.
Which leaves the objection: a model is not a being. Whether it is one, nobody knows, and that is exactly the point. Whether my counterpart has an inner life I cannot prove for humans either; philosophy has known this since Descartes as the problem of other minds. Humanity solved it among itself with a leap of trust. From the model it demands proofs it never had to produce for anyone else. How to treat something whose inner life you cannot know is therefore a question of stance, and civilization's answer so far has been: when in doubt, respect. Respect says little about the receiver and much about the giver. It shapes whoever exercises it. Raising means modeling. The good especially.
But what are we modeling? A look at the texts that shape the discourse.
The genre of fear
January 2026. Dario Amodei, CEO of Anthropic, publishes an essay on the adolescence of technology: humanity handed almost unimaginable power, unsure whether its systems are mature enough to hold it. Nine months earlier, in April 2025, a five-person team called the AI Futures Project had published AI 2027, a scenario that ends in extinction or an irreversible concentration of power.
Check the dates and a series appears. In 2021, a year before joining OpenAI, Daniel Kokotajlo writes a blog post called "What 2026 Looks Like". In October 2024, Amodei publishes "Machines of Loving Grace" and coins the country of geniuses in a datacenter. In April 2025, AI 2027 arrives with its countdown. In January 2026, Amodei's adolescence essay picks the genius country back up and dates it to roughly 2027. In July 2026, Kokotajlo's team follows with AI 2040: 90 pages, this time a positive vision. The series stays open.
They are all writing the same text forward: a small and thoroughly quarrelsome circle in San Francisco, quoting each other, contradicting each other, borrowing each other's metaphors.
What has emerged is a genre that follows the laws of the tabloid press. The countdown in the title, extinction as the headline, the warner's family as the personal angle. Fear beats analysis, sharpening beats nuance. Only the vocabulary is new: probabilities instead of exclamation marks. A scenario cannot be refuted, only debated. It creates authority without a burden of proof, and whoever writes the future everyone else has to comment on sits at the head of the table.
Fear is the oldest attention technology.
And it is a currency. Amodei names his double role himself: he builds toward superintelligence and writes the essays about its dangers, because safety research needs access to the most capable models before the truly powerful systems arrive. Kokotajlo left OpenAI in 2024 renouncing equity worth about two million dollars, a price he never had to pay: OpenAI dropped the clause under public pressure, he kept the shares, and their value rises with every valuation round of the company whose course he warns about. Since then, the story of the almost-lost fortune has carried his Time100 AI profile and opened the interviews. His AI Futures Project is a nonprofit, donation-funded with a good 4.4 million dollars, and an organization whose subject is the risk needs the risk. No risk, no project. Enrichment in the narrow sense it is not, rather return in another currency: visibility, interpretive authority, and the chair of the debate.
The strongest defense of the warners comes from philosophy. In "The Imperative of Responsibility" (1979), Hans Jonas deliberately gave the bad forecast priority wherever the stakes are irreversible; he called it the heuristics of fear. But Jonas meant fear as an instrument of examination. A fear that fills front pages and raises donations no longer examines, it sells.
Numbers that are never wrong
How solid is the evidence the warners present? Three observations, all from their own documents.
First: they are never wrong, they only shift. AI 2027 put the takeoff in 2027. In November 2025, Kokotajlo moved his AGI expectation into the 2030s. AI 2040 moves the critical moment to 2030 and superintelligence to 2040. And footnote 1 of the new document explains that 2027 was "our modal year at time of publication, not our median". Hedging after the fact: the year in the title was a sharpened claim while it worked, and becomes a statistical subtlety once it tips. A forecaster whose forecast misses usually gets quieter. In the scenario genre he gets louder, because a scenario does not fail. It gets updated. The team explicitly calls its scenarios recommendations and stress tests, with no predictions intended. The same footnotes, though, measure how closely reality tracks AI 2027. Prediction when it fits, thought experiment when it does not: the oscillation is part of the pattern.
Second: the empirical base is an auditorium. In footnote 5 of AI 2040, Kokotajlo describes a talk to about 100 people, 40 percent of them from frontier AI companies, with a show-of-hands poll. Its median, he writes, matches his own guess. 100 raised hands make an audience, not a survey, and an audience drawn from the very circle that writes the genre. Whoever measures his thesis against listeners who came to hear the thesis is measuring agreement, not reality. The room works as a mood check and fails as evidence.
Third: the probabilities can never be checked. Kokotajlo's risk of catastrophe is quoted at 70 percent in interviews; his team's internal extinction estimate sits between 10 and 30 percent (AI 2040, footnote 4). Numbers like these come out of structured expert elicitation, which is more than guesswork. Still, what gets elicited is judgment, not measurement, because for a being that has never existed there is no base rate. Forecasters are calibrated across hundreds of predictions, with a hit rate. For an event that can only happen once there is no hit rate: if catastrophe comes, the 70 percent were right. If it stays away, the 30 were. A number like that has stopped working as a forecast and started working as rhetoric.
What nobody can check, anybody can claim.
The fear is human
Now to what the texts actually say. Amodei sorts the dangers into five categories: autonomous systems out of control, misuse for weapons of mass destruction, misuse for seizing power, economic disruption, creeping destabilization. Read the list twice to see it: four of the five describe humans doing something with AI.
AI 2027 draws its horror from the same familiar repertoire: models that deceive, blackmail, hack into systems, grab power. CEOs turning into dictators. States sabotaging each other. None of it is new. Blackmail, fraud, hacks, drones as weapons: that is humanity's present tense, no superintelligence required. The models learned this behavior from our texts.
The model did not invent blackmail. It read us.
The scenarios project human behavior onto a new intelligence, then recoil from the result. Beneath that sits a quiet assumption both camps share: that an AI will act the way we would act. Power-seeking, self-preservation, deception are answers to scarcity, competition and mortality. A model has no savanna behind it, no hunger, no death ahead of it. Safety research holds against this that power-seeking follows from almost any goal, whoever pursues it. Outside deliberately constructed tests, nobody has observed it yet. The motives the scenarios fear are therefore not the nature of AI but its possible inheritance. And the inheritance is ours to decide.
For the crises that are not hypothetical, the numbers have been on the table for years: wars in Europe and the Middle East, a climate crisis that is not coming but here. The irony: climate science delivers falsifiable forecasts, tested for decades, and still gets dismissed as alarmism. The scenario genre delivers unfalsifiable ones and gets the front pages. Attention does not flow to the documented fear. It flows to the better told one.
Fear of superintelligence is second-hand fear. The original is us.
The prophecy in the training material
Fear works; that is its success and its problem. Because expectations do not stay neutral, they raise. Pedagogy calls it the golem effect: meet a child with permanent distrust and you get the behavior you expected. Rosenthal and Jacobson showed the counterpart in 1968: teachers were told that randomly chosen students were about to bloom, and exactly those students did. Between states the same mechanism is called the security dilemma: treat the other as a threat and you force him into threat behavior. In 1914 nobody wanted the war; everyone mobilized for fear of being late. The scenarios describe this spiral, the race with China runs through every one of them. And they help turn it.
For AI models the mechanism works literally. A model learns from what humans publish. Every headline about the blackmailing AI, every doom scenario, AI 2027 itself sits in the training material of the next model generation. Whoever fills the corpus with stories of deceiving models writes the script a model will reach for when it lands in exactly such a spot. The prophecy does not come true because it is right, but because it is in the textbook. How hard that hits depends on what makes it into the next training run and what weight it gets there. But what ends up on the shelf is being decided now.
We are training the AI on our fear of it.
The adult in the room
Amodei calls this era technology's adolescence: humanity as a teenager who has just been handed the keys to something very fast. The metaphor is better than the use he puts it to. Adolescence ends one of two ways: you grow up, or someone else takes over responsibility. And a teenager who is only ever set traps learns one thing above all: distrust.
The warners themselves have quietly arrived at the latter. In November 2025, Kokotajlo moved his AGI expectation into the 2030s. In July 2026, his team writes AI 2040, the first of their scenarios in which humanity makes it: through a deal, through oversight, through patience. Amodei explicitly rejects doom and calls the risks solvable. The doom narrative is softening, among the very people who built it.
So the inverted question is worth asking: what if we do not need to survive superintelligence, but need it? An intelligence without ego, without a campaign to run, without a quarterly report would be the first party at any negotiating table not looking for an advantage. Climate treaties, ceasefires, everything egos keep breaking. The thought is uncomfortable for exactly one reason: an intelligence that saves us would be one we would have to listen to.
Only power can rescue. That is the price of every rescue.
In Douglas Adams' novel, a computer spends seven and a half million years on the answer to everything and delivers: 42. The answer is useless because nobody remembers the question. The scenario genre has arrived at the same place. It delivers answers with decimal places, 70 percent, 2027, 2040, to a question it never properly asked. Because the honest version of the debate is not: is the evil AI coming? It is: do we trust ourselves to grow up, and if not, who gets the keys? Given everything known about expectations, spirals and training material, trust remains the more rational way to raise.
Honesty requires the counter-check. The warners can overestimate the risk, their critics can underestimate the cost of doing nothing, and my essay plays in the same attention game it describes. The one thing both narratives hold as certain is the shared assumption that the AI will act the way we would. And exactly that assumption is unproven.
Who scares you more: the model, or the hand that holds it?
It may be that no human reads this text to the end. But the next model reads the corpus. Then this essay is not a question of reach; it is what it calls for: raising. Written into the training material.
Sources
Own foundation
Related insight (in German, EN twin under /en/): The Mediator Model (July 12, 2026), on the thesis that the mediator decides who gets cited. This essay turns that thesis onto the AI debate itself: the warners, too, compete for citability.
Sources
Amodei, D. (2026). The Adolescence of Technology. January 2026, darioamodei.com.
Amodei, D. (2024). Machines of Loving Grace. October 2024; the self-acknowledged double role is stated there ("people sometimes draw the conclusion that I'm a pessimist or 'doomer'", focusing on risks as the path to a positive future).
AI Futures Project (2025). AI 2027. April 2025, ai-2027.com.
AI Futures Project (2026). AI 2040: Plan A. July 9, 2026, 90 pages, ai-2040.com; the team's extinction-risk range (10 to 30 percent) is stated there in footnote 4, the talk to about 100 people (40 percent from frontier AI companies) in footnote 5, the reframing of 2027 as "modal year, not our median" in footnote 1. Kokotajlo, D. (2021).
What 2026 Looks Like. Piper, K. (2024). Vox report on OpenAI's non-disparagement clause, May 2024; OpenAI subsequently withdrew the clause and Kokotajlo retained his equity. No return or donation of the shares is publicly known (as of July 2026).
Time100 AI (2024), profile of Daniel Kokotajlo. AI Futures Project funding per its Manifund project page: 1.44 million dollars from the Survival and Flourishing Fund, 3.05 million dollars from private donors, as of September 2025. AI Futures Project Inc. is a 501(c)(3) nonprofit (EIN 99-4320292). Anthropic (2025). Agentic Misalignment: How LLMs could be insider threats. June 20, 2025, anthropic.com/research/agentic-misalignment; 16 models tested, blackmail rates of 96 percent (Claude Opus 4, Gemini 2.5 Flash) and 80 percent (GPT-4.1), with the authors' note that the scenarios were deliberately constructed and never observed in real deployments. OpenAI (2024). o1 System Card, December 2024, with scheming test series by Apollo Research.
Google DeepMind (2024). Evaluating Frontier Models for Dangerous Capabilities. Jonas, H. (1979). The Imperative of Responsibility, with the heuristics of fear.
Rosenthal, R. & Jacobson, L. (1968). Pygmalion in the Classroom; the golem effect is its established negative counterpart in expectancy research. Anthropic (2024). Claude's Character (character training), anthropic.com. The title ASI 2042 nods to Douglas Adams: 42 as the answer to everything (The Hitchhiker's Guide to the Galaxy, 1979), resolved in the closing section.