Part I of five

Counting the OOMs, recounted

Re-grading the 2024 predictions: the arithmetic held, the institutions did not. About 7 min.

Part I: Counting the OOMs, recounted

1.1 The framework

Situational Awareness is 165 pages, but its engine fits in one line. Aschenbrenner argued that what matters is not raw compute but effective compute, and that effective compute decomposes multiplicatively into three terms:

$$ \mathrm{EC}(t) = C(t)\, A(t)\, U(t) $$

where \(C\) is physical training compute, \(A\) is algorithmic efficiency (how much less compute you need for the same loss, year over year), and \(U\) is what he called “unhobbling”: the gains from removing artificial constraints on a model that already has the capability latent inside it. Chain of thought, RLHF, tool use, long context, agentic scaffolding. Things that make a model that can already do X actually do X when you ask.

His estimate was roughly half an order of magnitude per year from \(C\) and half from \(A\), with \(U\) as a large and lumpy bonus on top. Compound that from GPT-4 in 2023 and you get a GPT-2-to-GPT-4 sized jump by 2027. His phrase for the epistemics of this was that it required nothing more exotic than “believing in straight lines on a graph.”

The essay then drew a chain of consequences: an automated AI researcher by 2027; an intelligence explosion following from it; trillion-dollar clusters and a national power-grid problem; a security regime for the labs modeled on nuclear weapons; a US-China race in which open-weight models fade to irrelevance; and eventually “The Project,” a nationalized effort that the essay predicted would arrive around 2027 or 2028.

The straight lines were the strong part. The consequences were the part that had to survive contact with a world that is less linear than a graph.

1.2 The scorecard

Several careful public scorecards now exist. The one-year retrospective on LessWrong in June 2025 found the aggregate compute curve roughly on schedule. The two-year scorecard on the EA Forum in June 2026 graded the major predictions as three on track, one clearly wrong, two open, and two too early to call. A separate May 2026 audit reached similar conclusions with a different weighting on the geopolitical claims. What follows is my own reading, weighted toward what matters for the argument of this essay; where I depart from those scorecards I say so.

Claim (2024) Status (Sept 2026) Notes
~0.5 OOM/yr training compute On track Astra on 100k+ GPUs at Stargate Texas. Aggregate power buildout roughly on the predicted curve; per-site interconnect lagging.
~0.5 OOM/yr algorithmic efficiency On track Epoch AI’s canonical estimate (Ho et al., 2024) is roughly a 3x per year reduction in the pre-training compute needed for a fixed capability, which is 0.48 OOM. Post-training gains since 2023 push the effective figure higher.
Large unhobbling gains Understated See Part II. The single largest capability delta measured this year came from the harness, not the model.
The data wall Deferred, not hit Synthetic data and RL on verifiable tasks pushed it out. Whether it was solved or postponed is not knowable from outside.
Automated AI researcher by 2027 Partial, early OpenAI states Astra is the first model where other models played a significant role in supervising its training. That is the loop starting, not the loop closing.
Intelligence explosion following Not observed No public evidence of the compounding dynamic. Too early to grade honestly.
Open-source / open-weight fades Wrong Near-frontier open-weight models continue to ship months behind the frontier. Distillation complicates the mechanism but not the verdict.
The Project (nationalization ~2027-28) Wrong in form, partly right in spirit No nationalization. But a $200M DoD contract, Claude on classified networks, a government-ordered suspension of Mythos, a formal executive-branch review of Astra before release. The state arrived as a customer and a regulator, not an owner.
US-China arms race Open Contested. Chinese labs remain fast followers via distillation; a decisive race dynamic of the kind described has not materialized.
AI revenue ~$100B run rate by 2026 Behind The two-year scorecard cites a March 2026 evaluation putting the most generous figure near $60B. Behind, but by less than a factor of two.
Superalignment as the hard problem Reframed The problem that actually gated releases in 2026 was cyber misuse, not alignment in the essay’s sense. Both labs slowed launches over exploit capability, not over deception or goal misgeneralization.

Three things stand out.

First, the arithmetic held. The rows that are pure scaling, \(C\) and \(A\), are the rows that graded cleanly. This is the part of the essay that was doing real work and it was not wrong.

Second, every miss is a miss about institutions, not about models. Open weights did not fade because the incentives of a dozen labs and two governments did not point that way. Nationalization did not happen because the state found it could get what it wanted by contract. The China race did not crystallize because the actual Chinese strategy turned out to be distillation, which the essay treated as a security problem rather than as a substitute for a race. Aschenbrenner modeled the labs like physics and the governments like the labs. The labs behaved like physics. The governments did not.

Third, and this is the one the rest of the essay is about: the essay never said what would count as having arrived.

1.3 The definition that was never pinned down

Situational Awareness uses “AGI” throughout and defines it, when it defines it at all, by analogy: a system that can do the work of an AI researcher, a drop-in remote worker, a model that is to GPT-4 what GPT-4 was to GPT-2. These are vivid and they were useful for forecasting. They are not operational. There is no test in the essay that a model could pass or fail.

That did not matter in 2024 because the gap was obviously large. It matters now because the gap is not obviously anything.

Consider the two definitions that collided on September 3:

Definition A (OpenAI, historical): an automated system that can perform all economically valuable work as well as or better than humans. By this definition, Astra’s president believes the model may qualify, and said so.

Definition B (ARC Prize): a system that can acquire any skill a human can, as efficiently as a human can. By this definition ARC, having just watched Astra beat the human action-efficiency baseline on 96 percent of its levels, still declined to call it AGI, on the grounds that the environments are bounded, deterministic, and closed-ended, and do not represent the open-endedness of the real world.

Both definitions are reasonable. They are also almost unrelated. A tests economic output. B tests learning efficiency on novel tasks. A model could saturate A while failing B (a very capable system that cannot learn anything new after deployment) or saturate B while failing A (a fast learner with no hands). Neither definition entails the other, and in September 2026 we have a model that arguably satisfies one and arguably does not satisfy the other, and no agreed-upon way to say which.

This is what I mean by the measurement gap. The capability gap was a distance between what models could do and what humans could do, and you could shrink it by scaling. The measurement gap is a distance between what we can observe about a model and what we would need to observe to settle the question. You cannot shrink it by scaling. You can only shrink it by building better instruments.

1.4 Recounting the OOMs for 2026

If I were to rewrite the effective-compute equation for the world as it is now, I would keep the form and change the emphasis.

$$ \mathrm{EC}(t) = C(t)\, A(t)\, U(t) $$

In 2024 the essay treated \(C\) and \(A\) as the engine and \(U\) as the bonus. The evidence of 2026 is that \(U\) has become the dominant term for measured capability, and the least understood.

Here is the data point. Under ARC’s Standard harness, Astra’s score varies from 17.5 percent at low reasoning effort to 62.7 percent at max. Under the Provider Adapter harness, it varies from 96.7 percent at no reasoning effort to 99.9 percent at high. Same weights. The harness change is worth more than every level of reasoning effort combined.

In July, before Astra shipped, OpenAI had already shown the same thing with GPT-5.6 Sol: 13.3 percent on the ARC-AGI-3 public set with the official harness, 38.3 percent with retained reasoning and compaction turned on. A factor of nearly three from two settings.

Aschenbrenner’s original unhobbling examples were things like “teach it to use a scratchpad” and “let it use tools.” Those were unhobblings of the model. What we are seeing in 2026 is unhobbling of the evaluation: the same model, scored differently depending on how much of its own state the evaluator lets it keep. The \(U\) term has migrated out of the model and into the harness, and once it lives in the harness it stops being a property of the thing being measured.

This has a consequence that I do not think anyone has stated plainly. A benchmark score is no longer a measurement of a model. It is a measurement of a model-harness pair. Reporting a score without reporting the harness is like reporting a sprint time without saying whether it was run downhill.

The rest of the essay takes this seriously.