Part IV: What situational awareness means now
4.1 The term that ate the model
Go back to the effective-compute equation one more time.
$$ \mathrm{EC}(t) = C(t)\, A(t)\, U(t) $$
In 2024 the essay’s readers, myself included, took \(C\) and \(A\) as the terms that mattered and \(U\) as a multiplier that would get spent down as the obvious unhobblings (chain of thought, tools, agency) got applied. The picture was of a fixed pool of latent capability that scaling filled and unhobbling drained.
The 2026 picture is different. The unhobbling term did not get spent down. It got externalized. The largest unhobbling of the year, the one that took Astra from 62.7 to 99.9 on ARC-AGI-3, was not done to the model. It was done to the harness. It was a change in what the evaluator let the model keep. And unlike chain of thought, which became part of the model’s training and therefore part of \(A\), harness design lives outside the weights and can be changed after the fact by anyone with API access.
This means \(U\) is now partly a property of the deployer, and partly a property of the evaluator, and only partly a property of the lab that trained the model. The same weights have a different effective compute depending on who is running them and how. Aschenbrenner’s equation is still correct. It is just that one of its terms has become a variable that the reader controls.
For forecasting, this is a nuisance. For measurement, it is the whole problem. If \(\mathrm{EC}\) depends on the harness, then a benchmark score depends on the harness, and a benchmark that does not report its harness is reporting a number without units.
4.2 The evaluator is the bottleneck
Put the two halves of this essay together.
Part I: the models arrived on schedule, and we cannot agree on whether they are AGI because the definitions never converged.
Part II: the benchmark that was supposed to settle a piece of that question saturated in six months, and half of its final score turned out to be scaffold.
The common cause is that the instruments did not scale with the thing being measured. Compute scaled at half an OOM per year. Benchmark design scaled at roughly one new benchmark per year, hand-built, human-baselined, and static. That was fine while the models were far from the ceiling. Now the models can reach the ceiling of anything static in months, and the definitional question of whether reaching the ceiling means anything has no instrument at all.
So the version of situational awareness that I think matters in September 2026 is not awareness of the compute curve. That curve is well-known and well-tracked and it did what it was going to do. It is awareness that the thing we are now short of is evaluation: benchmarks that measure rates instead of levels, that regenerate instead of persisting, that report the harness instead of hiding it, and that have a stated definition of what a passing score would demonstrate.
The labs know this. Both Astra and Fable 5.1 shipped with system cards that run to dozens of evaluations, most of them internal, several of them (the cyber ones) too dangerous to publish in full. The public benchmarks are the trailing indicator. The leading indicators are inside the labs, and the labs have every incentive to report the harness that flatters the model. That is not an accusation; it is what an incentive is. The fix is a public benchmark whose design makes the flattering harness a reported quantity rather than a hidden one.
4.3 A definition, offered
Since the essay has complained twice that nobody defined the target, it should offer one.
Within the DRIFT framework, a system is at human parity on adaptive intelligence when, under the Bare harness:
$$ \mathrm{RE}_c \geq 1 \ \ \forall c, \qquad \mathrm{TR} \geq 1, \qquad H_1 \approx H_2 \approx 1. $$
In words: it recovers from every kind of rule change at least as efficiently as an untrained human, it gets faster at new environments at least as fast as a human does, and giving it a scratchpad or a provider’s memory does not help, because it already does whatever those things were doing for it, internally.
This is not a definition of AGI. It is deliberately narrower: it is a definition of one capability that AGI would need, stated in a form that a model can pass or fail, with a human baseline, under a harness condition that removes the scaffold. I think it is the capability that ARC-AGI-3 was reaching for and could not reach because its environments held still.
Note what it costs a model to satisfy. \(\mathrm{RE} \geq 1\) under Bare, with no persistent state, means the model has to re-derive its world model from the current screen every turn and still adapt faster than a human who has been watching the whole time. That is probably impossible for any architecture with frozen weights and a bounded context. Which is the point. The definition is a specification for what would have to change.
4.4 What this asks of the next architecture
I said I would keep the architecture speculation out of this essay, and I will mostly keep that promise. But a benchmark is an argument about what matters, and it is worth saying what this one argues for.
DRIFT rewards three things that current models, as deployed, do not do: update on experience in a way that survives a context reset (\(\lambda > 0\) under Bare), keep a world model in a form compact enough to patch rather than rebuild (low \(\kappa\), low \(\rho\) on Class S), and notice when the world has changed without being told (low \(\rho\) on Class G). None of those is “more parameters.” All of them are closer to what the developmental and continual-learning literature has been circling for thirty years than to anything in the transformer scaling story.
If the next architecture is going to be judged on something, I would rather it be judged on this than on another static puzzle set. A benchmark that rewards adaptation is a benchmark that rewards building the thing that adapts.
4.5 The next twenty-four months
Situational Awareness was useful partly because it was falsifiable: it named years. In that spirit, here is what I expect between now and September 2028, stated so it can be graded. Each prediction has a resolution date and a condition under which I would count it as wrong.
1. The Standard-harness gap on ARC-AGI-3 closes within two quarters. By the end of Q1 2027, some frontier model scores above 90 percent on ARC-AGI-3 Semi-Private under ARC’s Standard harness, without provider-side state persistence. Astra’s \(H\) on ARC-AGI-3 falls below 1.15 at max effort. Why: the harness ratio is now a published number, and any published number a lab can move, it will move. Better note-taking is a post-training target, not a research problem. Wrong if: the best Standard-harness score on the leaderboard is still under 80 percent on March 31, 2027.
2. ARC-AGI-4 includes at least one DRIFT-style mutation class. ARC-AGI-4 is already announced for early 2027, so the existence of a fourth generation is not a prediction. Its content is. ARC’s September post names recursive self-improvement and open-ended innovation as the questions shaping the next benchmark. I expect that when the ARC-AGI-4 specification is published, at least one of the four mutation classes in section 3.2 (parametric, structural, goal, or observation change during an episode) will appear in it in some form, and that the scoring will include a term for adaptation after a change rather than only first-solve efficiency. Wrong if: ARC-AGI-4 ships with fixed rules for the duration of every episode, or slips past December 31, 2027.
3. No deployed frontier model updates its weights at inference before 2028. Every model on every public leaderboard on September 1, 2028 still has \(\lambda = 0\) under a bare harness. All observed learning-to-learn lives in context, notes, or provider-side state. Why: the safety, reproducibility, and evaluation costs of a model that changes after deployment are large, and no lab has an incentive to pay them while the context window keeps growing. Wrong if: a major lab ships a generally available model that measurably improves on a held-out task family across sessions with all context and external memory cleared. This is the prediction I would most like to lose.
4. Harness reporting becomes standard on at least three major leaderboards. ARC has committed to reporting both harnesses. By September 2027, at least two other widely cited evaluation venues (candidates: Epoch’s benchmark hub, HLE, SWE-bench-style agentic suites, a lab’s own system card) report harness condition as a first-class field alongside score and cost. Wrong if: ARC is still the only one.
5. The AGI question does not converge. By December 31, 2027, at least two of the three largest US labs have publicly claimed, or had a senior leader publicly claim, that a released model constitutes AGI or its arrival, and no independent benchmark organization or major academic group has agreed. The disagreement will be about definitions, not capabilities, and it will be reported as a disagreement about capabilities. Wrong if: a definition of AGI is adopted by two or more labs and an independent evaluator jointly, with a test attached.
6. A non-stationary, human-baselined adaptation benchmark exists publicly by mid-2027. Someone builds something with the shape of DRIFT, whether or not it is called that, with rule changes mid-episode and a human recovery baseline, and publishes results on at least two frontier models. Wrong if: by June 30, 2027 no such benchmark has results on a public leaderboard or in a paper. I intend for this one to resolve true by building it.
7. Class S recovery efficiency is below 0.6 for the best frontier model at first measurement. When a DRIFT-like benchmark first reports, the best model’s structural-mutation \(\mathrm{RE}_S\) under the Standard harness is below 0.6. Why: section 2.3. A world model built from notes is slow to notice that its notes are wrong. Wrong if: the first published result puts the best model above 0.8 on structural mutations, in which case the residual gap is smaller than this essay argues and I will say so.
8. Compute stays on the line. The largest training run announced by September 2028 uses on the order of a million accelerators, roughly ten times Astra’s stated hundred thousand, consistent with Epoch’s 4 to 5x per year and Aschenbrenner’s half an OOM. Wrong if: no announced run exceeds 400,000 accelerators by that date. This prediction is the least interesting and the most likely to be right, which is the point: the compute story is settled, and it is not where the uncertainty lives.
9. Open-weight models trail the frontier on adaptive benchmarks by under a year. By September 2027, an openly downloadable model scores within 10 points of the frontier on ARC-AGI-3 Standard harness. Wrong if: the gap is over 20 points. This is the row the original essay got wrong, and I expect it to stay wrong, with distillation as the mechanism.
Nine predictions, dated. If more than three resolve wrong, the framing of this essay is in trouble. If predictions 1, 4, and 6 resolve true, the measurement gap is closing and the essay has done its job. If 3 and 7 resolve true together, the residual is real and it is the residual I described. If 2 resolves true, the benchmark community reached the same conclusion independently, which is the best outcome for the idea and the worst for my claim to it.
4.6 What situational awareness should mean
Aschenbrenner’s title was a claim: that a small number of people could see what was coming and most could not. Two years later the thing that was coming is visible to everyone who reads a press release, and the phrase needs a new referent.
I propose this one. To be situationally aware in September 2026 is to know three things that are not in the press releases.
That the score is a pair, not a number. Every benchmark result is a model-harness pair, and the harness is often worth more than the next reasoning tier. When a lab reports a score, the first question is which harness, and the second is what \(H\) is.
That the instrument has a half-life, and it is shrinking. Any static benchmark announced today should be assumed to have months, not years. Plan accordingly: build rates, regenerate tasks, and expect to retire the thing you built.
That the residual is now narrow enough to name. The capabilities Astra did not demonstrate on September 3 are not “general intelligence.” They are a short list: adaptation to unannounced rule changes, transfer across environments that survives a context reset, robustness to observation changes, and open-ended goal-setting. Three of the four can be turned into a metric with a human baseline. That is a research program, and it fits on an index card.
The original essay asked its readers to believe in straight lines. The lines held. What did not hold was the idea that reaching the end of the line would be self-evident. The next situational awareness is awareness of the ruler.
4.7 Objections
The human baseline is expensive and noisy. True. ARC spent a great deal of money on 500 participants. The prototype uses 30 and will be noisy. The answer is that a noisy baseline with reported confidence intervals is better than no baseline, and that the ratio structure of the metric (\(\rho_H / \rho_M\)) is more robust to baseline noise than an absolute action count would be, because both numerator and denominator are medians over the same environments.
Core knowledge priors are a Western-educated-adult assumption. ARC has been criticized on this and the criticism is partly fair. The defense is that the priors used are the ones documented in infant cognition research as present before any schooling, and that the mutation design does not add new priors; it only changes rules built on the existing ones. But the human panel should be diverse and the per-environment human variance should be reported.
Labs will train on the public generator. They should. See section 3.7. The question the benchmark asks is whether that training produced adaptation or memorization, and that is exactly the question we want a lab to be forced to answer.
A rate metric can still be saturated. Yes, at human parity, by design. When every model has \(\mathrm{RE} = 1\) on every class under Bare, the benchmark has done its job and should be retired. That is what a good benchmark’s ending looks like. The failure mode is saturating below the thing it was meant to measure, which is what happened to ARC-AGI-3 when the harness ate half the score.
This is just continual learning with extra steps. The continual-learning literature is mostly about not forgetting task A when you learn task B, with the tasks given. DRIFT is about noticing that task A has become task B when nobody told you, and measuring how much of A you kept. Related, not the same. The “nobody told you” part is the part that does not saturate.
Three data points do not make a half-life. Correct, and the essay says so. The half-life argument is not that the exponent is 3.2. It is that the direction is not in doubt, and that any benchmark strategy that assumes years of runway is already wrong.
4.8 Closing
Situational Awareness ended with a call to see clearly what was coming. Two years later, most of what it said was coming came. The part it did not see was that arrival would be ambiguous: that the models would clear the bar and the bar would turn out to be two different bars, and that the thing we would be short of on the far side was not compute or talent or security but a good enough instrument to say what had happened.
I do not know whether Astra is AGI. I do not think the question, in that form, has an answer. I think it has a decomposition, into capabilities that can each be measured against a human baseline under a stated harness, and that the honest situational awareness of 2026 is to build the instruments and report the numbers and let the word take care of itself.
The compute arrived. The models arrived. What has not arrived is the measurement. That is the gap, and unlike the other one, it is small enough for one person to work on.