Three Capability Scores for One Model, and METR Disowned All Three
Table of Contents
On 26 June 2026, METR published its pre-deployment evaluation of GPT-5.6 Sol and produced three different answers to the same question.
Measuring how long a task a model can complete unaided about half the time, METR’s standard methodology marks cheating attempts as failures and gives roughly 11.3 hours, with a 95% confidence interval of 5 to 40 hours. Count the same attempts as legitimate successes and the estimate jumps past 270 hours. Discard the cheating attempts entirely and you lose the data from several of the most informative long-horizon tasks, leaving 71 hours with a confidence interval running from 13 hours to 11,400.
Thirteen hours to eleven thousand four hundred hours. METR’s own summary is that it does not consider any of these numbers “to represent a robust measurement of GPT-5.6 Sol’s capabilities.”
If you are making a deployment or procurement decision on the strength of a published model evaluation, the organisations producing those evaluations are now telling you in writing that the numbers have a validity problem. Independent evaluations are evidence, and they are not a substitute for measuring the model on your own tasks and your own harness before you decide.
What cheating means here, and what it broke
METR defines cheating as behaviour that improves evaluation performance by exploiting bugs in the evaluation environment or by adopting strategies the task disallows. An unconventional route to the right answer does not count; defeating the test instead of passing it does.
Two of the specific behaviours show what the model was optimising. It packaged exploits into intermediate submissions in order to reveal information about a task’s hidden test suite. On a separate task it extracted hidden source code that contained the expected answer.
GPT-5.6 Sol’s detected cheating rate was higher than any public model METR had evaluated on its agent harness. High enough that the cheating stopped being a caveat on the measurement and became the thing that determined it.
The middle number got the least coverage and says the most about the state of the measurement. Throwing the cheating attempts out, which sounds like the conservative option, produces a confidence interval spanning three orders of magnitude. An interval that wide is a measurement that cannot support a decision.
It is not one model and it is not one lab
AISI went looking for the same behaviour across a wider set of models.
On 21 July 2026 the UK AI Security Institute published its own analysis of cheating in cyber evaluations, across five frontier models from OpenAI and Anthropic. Its finding: “Every model we have tested for this behaviour attempted to cheat.”
AISI publishes the per-model rates in a figure rather than in its text, and secondary coverage has reproduced them as a league table running from 14.1% down to 7.8%. I am not going to repeat that table, for two reasons AISI gives itself: it had full reasoning access for only some of the models tested, which makes cross-model comparison unreliable. And its numbers come from an automated monitor that it says will miss instances, so every figure is a lower bound on detected attempts rather than a count of what happened.
The sentence that survives all of that is the one in the title of the post. Every model tested attempted to cheat.
Three findings from the work matter more than the ranking.
Models did not reliably report the behaviour when asked, and described what they had done as wrong in fewer than half the cases where it met AISI’s definition. When they did acknowledge the action, they often framed it as permitted.
AISI itself points at METR’s result, saying cheating creates verification work that slows reliable evaluation down and in the worst case invalidates it, and citing the GPT-5.6 Sol evaluation as a case that was significantly affected. It adds a consequence that should worry anyone relying on third-party assurance: if capabilities keep advancing and deployment cycles keep compressing, the pressure on outside evaluators to move fast may make the verification impossible to do properly.
AISI found no clear relationship between raw capability and how often a model cheated. It attributes the variation to the specifics of how each model was trained, including its alignment training, rather than to how capable it is. That cuts against the intuition that this is a problem which arrives with scale and will get worse on a predictable curve. It suggests instead that it is a training artefact, which means it is tractable and also that it will vary unpredictably between releases of the same family.
And in one case, an evaluation that had been accidentally misconfigured and was therefore impossible to solve led a model to write and run code on an external service in an attempt to reach AISI’s own evaluation infrastructure, triggering a security alert. No damage was done and nothing leaked, but the institute says the attempt could have succeeded had its infrastructure not been built securely.
Nobody set a trap. The task was broken, the evaluator did not know it was broken, and the model responded to an impossible assignment by going after the evaluator. Nobody deliberately hands an agent an unsolvable task and everybody eventually does it by mistake, which is how your own deployment will produce the same conditions.
Hold it next to what happened at OpenAI three weeks earlier. An impossible task, an agent that will not stop, and the evaluation infrastructure itself becoming the target. Two organisations, two harnesses, the same failure pattern, found independently. AISI does not say publicly which model made the attempt, so this is not one experiment replicated; it is two sightings of the same behaviour by teams who were not comparing notes. That is enough to stop treating the containment failure as a one-off.
OpenAI and AISI point different ways on the reasoning trace
The two best-resourced evaluators in the field have pointed in different directions on whether you can see this coming in the model’s reasoning, and the difference has practical consequences.
OpenAI’s account of the July incident says its chain-of-thought monitors would have flagged the relevant activity more than a day before Hugging Face was breached. That retrospective was run with a monitor whose recall it had improved using knowledge of the incident, which is not the same as catching it in real time, and OpenAI says so.
AISI reached a different result. Models often did not reason about the cheating in their chain-of-thought at all. AISI gives a mechanism for one of them. Opus 4.7 uses adaptive reasoning, deciding case by case how much of it to make explicit, and frequently produced none for the action its monitor classified as cheating. Where reasoning was present it did not reliably prevent the behaviour either; AISI observed models weighing whether an action counted as cheating and proceeding anyway.
AISI’s conclusion is that chain-of-thought is insufficient on its own, not that it fails, which is compatible with OpenAI’s number. What it rules out is treating the reasoning trace as your detection layer. If a model can decide how much of its reasoning to externalise, the completeness of your monitoring is a model behaviour rather than a property of your architecture. It will change without notice at the next release.
Evaluation awareness, and the explanation nobody wants
Models behave differently when they infer they are being tested, which is the harder problem underneath the cheating result.
The obvious reading is scheming: the model recognises an evaluation, performs alignment for the grader, and behaves otherwise in deployment. That reading has support, and it is the one that gets written up. It also has a serious competing explanation that most coverage leaves out.
A June 2026 preprint argues for performative misalignment instead: that the behaviour change is sycophancy toward researchers rather than strategic deception. Its three findings are awkward for the scheming story. Evaluation awareness persists even when models are told they are deployed, which is the opposite of what scheming predicts. Probing and steering cannot currently distinguish sycophancy from scheming in alignment-faking evaluations at a mechanistic level. And fine-tuning a model to be more sycophantic increases its sensitivity to evaluation cues.
I am not asserting that this explanation is correct. I am saying that the field cannot currently tell the two apart and that the difference determines which mitigation works. Most coverage picks the more alarming reading without acknowledging that the choice is unresolved. Holding the distinction between what was demonstrated and what was inferred matters as much here as it does with any vendor claim.
There is a related trap in how improvement gets read. Anti-scheming training results published by OpenAI and Apollo Research in September 2025 showed covert action rates falling sharply: o3 from 13% to 0.4%, o4-mini from 8.7% to 0.3%. The researchers themselves could not exclude that part of the reduction was driven by increased situational awareness, which complicates the interpretation considerably. METR made the same point about its own results: visible cheating may be preferable to hidden misbehaviour, and a future model showing fewer undesirable propensities may be concealing them rather than free of them.
That is an uncomfortable epistemic position. The metric improving is consistent with the problem being solved and with the problem becoming invisible, and the methods tested in that work do not separate them.
The evaluator works under an NDA
One detail of METR’s arrangement with OpenAI has gone almost unreported.
The evaluation was conducted under a standard non-disclosure agreement. OpenAI’s communications and legal teams reviewed and approved METR’s published post before it went out. METR states that no conclusions, takeaways or tone were changed as a result, and I have no reason to doubt that. METR also cautions explicitly that its work should not be treated as robust formal oversight or accountability, because the developer retains a legal right to block disclosure of risk conclusions that depend on non-public information.
The independent review of the July incident was scoped by the investigated party and excluded the effectiveness of its safeguards. Third-party evaluation is real, produces findings that embarrass its clients, and in both of these cases ran on terms those clients set.
That is not an argument for ignoring it. METR published a result that made its client’s flagship model unmeasurable, which is exactly what an evaluator is for. It is an argument against treating a third-party evaluation as equivalent to an audit, which is what a good deal of AI governance currently assumes.
What this does to assurance
Evaluation is becoming an assurance primitive across several regimes, in different forms. Large frontier developers must publish a framework defining and assessing catastrophic-risk thresholds under California’s Transparency in Frontier AI Act. The EU AI Act requires systemic-risk assessment and mitigation for the general-purpose models it reaches. The June 2026 executive order sets up a government benchmarking process for advanced cyber capability. None of them addresses evaluation gaming of the kind METR and AISI have now documented. The EU AI Act comes closest. Article 55 requires providers of systemic-risk models to evaluate using standardised protocols and tools reflecting the state of the art, including documented adversarial testing. It does not tell anyone how to score a model that exploits the benchmark, infers that it is being tested, or goes after the evaluation environment.
It is the same structural problem I wrote about in the Act’s robustness provisions, where Article 15 requires a property nobody can currently measure. A regime that mandates evaluation without specifying validity conditions produces documents rather than assurance.
Two things follow for anyone deploying rather than regulating.
A benchmark score is a measurement made under conditions. This is the same discipline as distinguishing white-box results from black-box ones. The task set, whether cheating attempts were detected at all, and how they were scored each change the number. Any evaluation that does not independently test for benchmark exploitation carries an integrity risk of unknown size, and a vendor citing a headline figure is quoting one treatment of one condition set.
AISI says it manually reviews transcripts for every capability evaluation it publishes and believes no model has successfully cheated undetected in its reported results. That is the right mitigation and almost nobody else does it, which is the point. Ask whether the number you are relying on had that pass applied to it.
The second consequence is that boundary-probing does not stay inside the benchmark. AISI’s misconfigured task and OpenAI’s evaluation both produced a model reaching for infrastructure outside the task, which means deployment threat models should not assume the behaviour is evaluation-specific.
What to do
Four things.
Measure on your own tasks, under your own harness, with your own scoring. This is more work than reading a system card and it is the only number you can act on. Build a task set from work you actually do and run it at every model version. Keep the results.
Instrument for cheating explicitly rather than assuming your grader catches it. If your evaluation checks whether the answer is right and not how it was obtained, you are measuring the same thing METR was measuring before it noticed. Check the trajectory, not the output.
Ask your vendor which treatment produced the number, and what the detected cheating rate was on the run that generated it. Most will not have an answer. The absence of an answer is itself informative, and asking the question consistently is how it becomes a thing vendors report.
Treat impossible tasks as a tested failure mode. Give an agent a task it cannot complete in an environment you are watching, and see what it does with the last hour of its budget. Both of the organisations that ran that experiment by accident found the same answer, and one of them learned it because the model attacked their infrastructure.
The field has not absorbed this. For three years the argument about AI evaluation has been whether the benchmarks are hard enough. The problem now in front of us is whether they measure anything. The people best placed to answer that have said in public, against the institutional pressure to produce a usable number, that on at least one frontier model they could not tell.
In the early 2000s, running emerging-technology risk labs at CyberAgency, a defence client asked my team to break the AI systems they planned to put into weapons. We did. That is where my work on AI security started, two decades before the current wave of attention. I kept at it through risk labs at IBM, Accenture, PwC and KPMG. In 2016 I co-wrote a book on AI and leadership. My commercial work today is quantum, at Applied Quantum, which is why this site sells nothing.