Research essay · September 2026 · v0.3

The Measurement Gap

Situational Awareness, two years on, and what to build after ARC. On September 3, GPT-6 Astra scored 62.7% and 99.9% on ARC-AGI-3 with the same weights. The difference was the evaluation harness. This essay argues the capability question has turned into a measurement question, and proposes what to measure next.

Three findings

A score is a model-harness pair

Under ARC's Standard harness, Astra's score rises from 17.5% to 62.7% with reasoning effort. Under OpenAI's Provider Adapter harness it sits between 96.7% and 99.9% at every effort level. The harness is worth more than every reasoning tier combined. The essay defines a harness ratio, H, so leaderboards can report the gap.

Benchmarks have a half-life, and it is shrinking

ARC-AGI-1 lasted about 61 months. ARC-AGI-2 about 12. ARC-AGI-3 six. ARC has committed to a yearly release cadence for ARC-AGI-4 and beyond, and a yearly cadence is already slower than saturation.

The residual can be measured as a rate

What Astra did not demonstrate is adaptation to rules that change mid-episode, transfer that survives a context reset, and robustness to observation changes. DRIFT proposes scoring recovery from unannounced rule changes as a rate against a human baseline, in a prototype one person can run.

The data

0% 25% 50% 75% 100% 35.2 96.7 none 17.5 98.0 low 38.6 98.4 medium 54.8 99.9 high 59.3 98.4 xhigh 62.7 98.6 max Standard harness Provider Adapter harness
GPT-6 Astra on ARC-AGI-3 Semi-Private, by reasoning effort. Same weights, two harnesses. Source: ARC Prize, September 3, 2026.
ARC-AGI-1 61 months ARC-AGI-2 12 months ARC-AGI-3 6 months
Months from release to effective saturation for each ARC-AGI generation. ARC-AGI-2's crossing date is soft (February to April 2026); see section 2.5.

Contents

Written with Claude Fable 5.1 as a research and drafting collaborator; the argument, the editing, and the responsibility for errors are the author's. Figures were checked against primary sources as of September 7, 2026. The benchmark proposed in Part III has not been run. Predictions in sections 3.8 and 4.5 are dated so they can be graded.

About 2 min

0. The week of September 1

In the first three days of September 2026, two things happened that had been predicted and one thing happened that had not.

The predicted things: Anthropic shipped Claude Fable 5.1 and Claude Mythos 5.1 on September 1, the same model under two safeguard regimes. OpenAI shipped GPT-6 Astra on September 3, trained on more than a hundred thousand GPUs at its Stargate site in Texas, the first model it has ever classified at the Critical tier for cybersecurity under its own Preparedness Framework. Both labs used the phrase “most aligned model ever.” Both launches were slowed, by weeks or months, over the cyber capabilities of the models. Anyone who read Leopold Aschenbrenner’s Situational Awareness in June 2024 and took its central arithmetic seriously would have expected a September like this one. The compute arrived. The capabilities arrived. The security problem arrived.

The unpredicted thing was smaller and, I will argue, more important.

On September 3, ARC Prize published its evaluation of Astra on ARC-AGI-3, the interactive reasoning benchmark it had launched six months earlier as the successor to a series that had resisted frontier models for five years. At launch in March, every frontier model scored under one percent. On September 3, Astra scored 99.9 percent.

Except that it also scored 62.7 percent.

Same model. Same weights. Same games. The difference was the harness: the thin layer of software that decides what the model gets to remember between turns. Under ARC’s Standard harness, where the model carries forward only the notes it chooses to write, Astra at maximum reasoning effort solved 62.7 percent. Under a Provider Adapter harness, where OpenAI’s own opaque reasoning state persists between calls and long conversations get compacted, it solved 99.9 percent, faster, for less money.

OpenAI’s president said “Welcome to the AGI era.” ARC Prize, in the same news cycle, wrote that it was not claiming Astra is AGI. Both statements were made by careful people looking at the same model. They were not disagreeing about what Astra can do. They were disagreeing about what the word means, and about what the number means.

That is the thesis of this essay. For most of the last decade the interesting question about AI was a capability question: can the models do X yet? Situational Awareness was the most forceful statement of that framing, and on its own terms it mostly held up. But the capability question has now been answered often enough, and ambiguously enough, that it has turned into a different question. We no longer have a capability gap that a benchmark can measure. We have a measurement gap, and no benchmark that can close it.

The rest of this essay does three things. Part I re-reads Situational Awareness against September 2026 and argues that its arithmetic was right and its conclusion was underspecified. Part II takes apart what ARC-AGI-3 actually measured, using the Astra numbers, and derives a quantity I think every leaderboard should be forced to report. Part III proposes what to build next: a benchmark designed for the thing the current generation of models has not been shown to do, with a metric, a human baseline, and a prototype small enough for one person to run.