TWINLADDER
TwinLadder logoTwinLadder
Atpakaļ uz apskatu

Izdevums #44

Thirty Per Cent of the Benchmark Was Broken

On 8 July OpenAI took apart SWE-Bench Pro, a coding benchmark it had recommended to the research community five months earlier, and found roughly 30% of its tasks faulty — hidden requirements, contradictory instructions, tests too strict to pass, grading that stopped short of the answer. Three days before that, Scientific American and Nature had assembled the evidence on what routine AI assistance does to a practitioner's own skill, including a Lancet study in which experienced endoscopists' unaided detection rate fell from 28.4% to 22.4%. This issue reads the two measurements together — one withdrawn by the people who made it, one taken by removing the machine and testing the human alone — and asks who maintains the instruments behind every capability figure in a board pack.

measurement and evidence
benchmarks
deskilling
internal audit and assurance
banking supervision
2026. gada 11. jūlijs15 min read
Thirty Per Cent of the Benchmark Was Broken

TwinLadder Weekly

Issue #44 — Thirty Per Cent of the Benchmark Was Broken

11 July 2026 · Weekly intelligence on judgment, governance, and the boards accountable for both


Editor's Note

From Alex —

On 8 July OpenAI published an examination of SWE-Bench Pro, a coding benchmark the company had recommended to the research community five months earlier. Roughly 30% of the tasks in it were broken — hidden requirements, contradictory instructions, tests too strict to pass, grading that stopped short of the answer. OpenAI retracted its own February recommendation.

Set the sequence out in order. A supplier endorsed a measuring instrument. Five months later the same supplier took the instrument apart, found a third of it faulty, and published that.

The behaviour deserves credit. Auditing an instrument you endorsed, and putting the result where anyone can read it, is what a board should want from every vendor it buys from. Almost every capability claim a board has been shown this year rests on numbers produced by instruments of this kind, and few of those boards could say who maintains the instrument or when it was last checked.

Three days earlier, Scientific American, working with Nature, had assembled the empirical evidence on what routine AI assistance does to a practitioner's own skill. Among the studies it gathered is the one place where that erosion has been measured on real professionals doing real work: experienced endoscopists whose unaided detection rate fell from 28.4% to 22.4%.

Two measurements, one week. One of them was withdrawn by the people who made it. The other was taken by removing the machine and testing the human alone — the design almost no institution runs on its own people.

Liga has the analysis.

— Alex


What OpenAI Found When It Checked Its Own Recommendation

The investigation was published on 8 July and reported the following day. SWE-Bench Pro is a set of software-engineering tasks used to rank the coding ability of frontier models, and OpenAI had recommended it in February as a leading coding evaluation.

The method ran in two arms. An automated pipeline flagged 286 tasks as suspicious. Codex-based investigator agents reviewed them and labelled 200 — 27.4% of the set — as flawed. Separately, five experienced software engineers reviewed the same material independently and found 249, or 34.1%. The two arms agreed on 74% of cases. OpenAI's own summary of the finding: "approximately 30% of the tasks in SWE-Bench Pro are broken."

The audit of the benchmark therefore exists in two versions, one machine and one human, and they disagree about a quarter of the time. The instrument used to check the instrument returned a different answer depending on who held it.

Then the figure a vendor deck would quote. On the same public 731-task set, top models moved from 23.3% to 80.3% in eight months. That curve was cited through the spring as evidence of how fast coding capability was improving. Two readings survive the teardown. The models improved that much. Or a scale, a third of which was mismarked, produced a rise partly composed of tasks that were never gradeable in the first place. Both readings fit the published figure, and the figure alone cannot separate them.

The scoreboard and the players

The day after the teardown, on 9 July, OpenAI made the GPT-5.6 family generally available across ChatGPT, Codex and the API. It claimed a new high of 53.6 on Agents' Last Exam for its top model, Sol: "GPT-5.6 Sol sets a new high of 53.6, eclipsing Claude Fable 5 by 13.1 points."

On SWE-Bench Pro — the benchmark examined the day before — Claude Fable 5 scored 80% against Sol's 64.6%.

Set the two publications side by side, in the order they appeared. The organisations that publish the scores are the organisations being scored, and the benchmark that produced the unflattering result is the one that came apart on a Wednesday. That arrangement can still produce accurate numbers; the teardown is evidence that it sometimes does. A benchmark result is a reading taken by a party with an interest in the outcome, on an instrument maintained by another party with an interest of its own, and audited when one of them chooses.


The Measurement Taken by Removing the Machine

On 5 July, Scientific American, in a piece produced with Nature, assembled the empirical case on skill loss under AI assistance. Mariana Lenharo gathered three studies with named authors and stated sample sizes.

The first is clinical. An observational study in The Lancet Gastroenterology & Hepatology — Budzyń and colleagues — ran across four Polish centres, drawing on data from centres that had run the ACCEPT randomised colonoscopy trial. It looked at experienced endoscopists, clinicians who have performed thousands of procedures, and measured their adenoma-detection rate: how reliably they find precancerous polyps when working without AI assistance. It took that measurement twice, before the centres introduced routine AI assistance and after. The unaided rate was 28.4% before. After a period of routine AI-assisted work it was 22.4%. Six percentage points, in the same clinicians, on the same procedure, with the machine out of the room on both occasions.

The second is a randomised controlled trial run by Anthropic with 52 software engineers. The group that worked with AI assistance scored 50% on a comprehension quiz about the code. The group that wrote it by hand scored 67%.

The third is Tapani Rinta-Kahila at the Hanken School of Economics, on accountants who had spent a decade working on automated systems and found, once the tool was removed, that they had forgotten how to do routine tasks.

Scientific American states the mechanism in a line: "People can perform at a pretty high level, because they're basically borrowing skills from the AI, but are not developing those skills themselves."

What the clinical result covers

The limits. The colonoscopy study is observational and hypothesis-generating; it compares a period before with a period after, and the authors present it as such. A before-and-after measure of unaided detection cannot fully separate a skill that decayed from case-mix, effort or attention that shifted. The confounder is real.

The skill it measures also has a large sensorimotor component — visual scanning, where the eye dwells, when it returns to the edge of the frame — and it runs under immediate feedback. The polyp is there or it is absent, and the clinician finds out in seconds.

That last condition is what makes the result travel. Fast, clear feedback is the strongest thing known for keeping a practised hand sharp. The endoscopist works in the kindest available conditions for holding on to a skill, and the skill thinned anyway. A credit officer calling a marginal borrower has no such loop. She makes the judgment and waits eighteen months to learn whether she was right, if she ever learns cleanly at all. The decay we can measure is happening in the easy case. The hard case is the one she works in.

Six points is also the flattering figure on its own terms. The study counted seniors losing a stock they already held. A firm carries a second half of the same liability: the analyst hired after deployment, who never builds an unaided baseline that could later decay. The measured decline counts only people who had something to lose.

What the design does that corporate AI measurement almost never does is simple. It took the machine away, put the same people on the same task, scored them, and compared.


The Audit That Cannot Run on Itself

A benchmark is the assurance layer for the machine's side of the work. An unaided-performance reading is the assurance layer for the human's side. Most institutions this year run neither, and hold instead a file of numbers that came from somewhere else: adoption rates, prompt volumes, hours-saved estimates, satisfaction scores.

There is a structural reason the human-side reading is rare, and audit shows it cleanly. To measure whether a professional's own capability has thinned, you need the same people on the same task before and after the machine arrived. Audit has no pre-AI control group left to test, and there will be fewer of those people every year. Medicine could supply the reading because the endoscopists were on the same instrument, in the same hands, on both sides of the deployment. Once deployment is complete, the window for taking that reading has closed, and it closes quietly.

The number that falls when the capability falls

Most institutions do hold one figure that moves when this happens, and they read it the wrong way round.

When a firm puts an AI into a judgment task it usually keeps a human with authority to override. For a while that human overrides. Then the override rate falls, month after month, until it approaches zero. The comfortable reading is that the model improved. The reading a risk committee should hold open is that reliance grew faster than retained skill.

Picture one reviewer inside that aggregate. Three months ago she flagged several cases a week. This month she has flagged none. The queue still loads each morning, the machine's recommendation sits at the top of each file, and she clicks approve a little faster than she did last quarter, because every time she stopped to check, the machine had been right. Nothing in the motion tells her which of two things is true: that there is nothing left to catch, or that she has lost the capacity to catch it.

The aggregate also mixes people. The veteran's rate decays against a stock of judgment she built before the machine arrived. The analyst hired after go-live had no baseline, so hers could never fall; she reads as a clean operator and is in fact one who was never formed. Some of the decline is genuine model maturity. Three realities, one line on the dashboard, and the function head reports the flattering one upward.

A falling override rate is uninformative as a level. It becomes diagnostic against an independent test of whether the catcher can still catch — the colonoscopy design, applied to a credit file instead of a colon.

Who the market is buying

One more figure from the same week, because it says where the bench is heading. On 8 July, Guillermo Gallacher at Indeed Hiring Lab reported that AI-exposed occupations, which had fallen furthest between 2022 and 2026, have rebounded the most since 2025. US software development postings grew nearly 15% since Claude Code's launch in late February 2025, while overall postings fell 7%. The composition is the finding: "71% of the increase in software development job postings between May 2025 and May 2026 is from senior roles, and 37% is due to jobs that mention AI in their title." Hiring Lab adds that this "preliminary evidence is in line with previous research focused on the disproportionate impact of AI on entry-level jobs."

The recovery is real, and it is senior. The market is buying people who already hold judgment, out of a pool that fewer institutions are refilling.


The Same Week, the Assurance Chain Moved

Five institutional moves in three days, and each of them relocates a test.

7 July — the ECB writes to every significant bank. Claudia Buch, Chair of the Supervisory Board, sent letter SSM-2026-0301 to the chief executive of every significant institution under European banking supervision. Its subject is AI-enabled cyber risk: emerging models can identify software vulnerabilities and generate functioning exploits at unprecedented speed, compressing the window between discovery and exploitation. On accountability the letter is direct — "Responsibility for responding to the evolving cyber-risk environment primarily lies with banks' management bodies. Strategic ICT-related decisions, including ICT investments, resource allocation and ICT risk-related risk tolerance frameworks (RTFs) may need to be revisited." Each bank must submit a comprehensive action plan to its Joint Supervisory Team by 31 October 2026, with roles, responsibilities and timelines named. The ECB moved its IT Risk Questionnaire deadline from September 2026 to February 2027 to make room for the work, and will run a horizontal analysis across the plans it receives.

8 July — the IIA rewrites the Three Lines Model and aims it at the board. The Institute of Internal Auditors published updated Statements of Position on the Three Lines Model and on the role of the internal audit function in enterprise risk management, replacing the former position papers. On practitioner readings of the change, the refreshed model takes the board as its primary audience, sets independence out as checkable tests, separates assurance work from advisory work with different safeguards, and asks internal audit to coordinate assurance across the organisation through assurance maps and aligned risk taxonomies. The Institute's own language is about how "collaboration, coordination, and reliance among the three lines...improve risk coverage and the reliability of information communicated to the board and senior management". This is the backbone an AI risk taxonomy has to be mapped against, and it puts one question in front of an audit committee: where do the models appear on the assurance map, as distinct from the policy about them?

8 July — the EDPB defines two words. At its plenary the European Data Protection Board adopted guidelines on anonymisation, guidelines on web scraping in the context of generative AI, and the final version of its blockchain guidelines. The anonymisation test is stated plainly: "Data is anonymous if it does not relate to an identified or identifiable natural person." On scraping: "The GDPR applies to web scraping when it includes personal data processing operations, such as collection, storage, organisation and retrieval." For a company that builds or fine-tunes on scraped material, "we anonymised it" now carries a defined meaning, and the supervisory authority holds the definition.

8 July — the Commission asks a court to price a missed deadline. The Commission referred Ireland, Spain, France and the Netherlands to the Court of Justice of the EU for failing to notify full transposition of the NIS2 Directive, and asked the Court to impose financial sanctions consisting of a lump sum plus daily penalties running until each notifies. The transposition deadline was 17 October 2024. This is the third and final stage of the infringement procedure, following letters of formal notice to 23 Member States in November 2024 and reasoned opinions to 19 in May 2025. As the reporting puts it, "few member states met the original October 2024 deadline for transposing NIS2 into domestic legislation." For a board in one of those four markets, the transposing statute and its management-body duties arrive quickly and with a short runway.

9 July — the Garante fines a chatbot company using the GDPR alone. The Italian data protection authority imposed a fine on Character Technologies Inc. over inadequate privacy notices and failures on child protection and age verification in an AI companion service that lets users, minors among them, converse with AI-generated characters. The headline in the regulator's own press room: "Artificial intelligence: The Italian Data Protection Authority imposes a fine on Character.AI". No high-risk classification was required to reach the conduct.

Read the five together. In each, a test an institution used to set for itself is now set, defined or timed by somebody outside it. The anonymisation test belongs to the supervisory authority. The action plan carries a date fixed by the supervisor. The assurance map belongs to internal audit and goes to the board. What an institution says about its own AI is worth less each quarter; what it can show it measured is worth more.


What This Means for Boards Right Now

One. Ask where each capability number came from and who maintains the instrument. A vendor figure quoted from a public benchmark inherits every fault in that benchmark, and OpenAI has now shown that a widely used one carried faults in roughly a third of its tasks — with two competent review methods disagreeing on how many. Three things belong beside every capability claim in your papers: which benchmark produced it, who maintains that benchmark, and when it was last examined by someone other than the parties it ranks. Where those answers are unavailable, treat the figure as marketing with a decimal point.

Two. Commission one unaided reading of your own people. The clinical study's design is the whole method, and it costs a morning: take the machine away, put the same people on the same task, score the result, and fix the date on which it will be repeated. That yields an observed number with a population, a comparison and a date attached. Adoption rates, prompt volumes and satisfaction scores answer a different question — none of them tells you whether the person who has to catch the machine's error can still catch it.

Three. Name the line that tests the model, separately from the line that reviews the policy about it. The IIA's refreshed Three Lines Model asks internal audit to coordinate assurance across the organisation and takes the board as its audience. The ECB has given significant institutions until 31 October to put roles, responsibilities and timelines on paper for one specific AI-driven exposure. Both point at the same blank space: for each material AI system, who tests it, on what schedule, and what happens when a test fails. A policy review leaves all three unanswered.

Which leaves the question for the next meeting, and it applies to every assurance figure in the AI pack: our AI assurance rests on measurements — who checks the measurements?


Reading List


What We Are Watching Next

  • Whether SWE-Bench Pro's maintainers publish a corrected task set, and whether any lab restates results it reported against the original
  • Whether another frontier lab examines a benchmark it has cited in its own capability claims and publishes what it finds
  • Whether any deployer publishes an unaided-performance reading for its own people — machine removed, same task, stated baseline — in place of an adoption or satisfaction figure
  • Whether internal audit functions adopting the IIA's refreshed Three Lines Model put AI models themselves on their assurance maps, or stop at the policy about them
  • Whether Ireland, Spain, France or the Netherlands notifies full NIS2 transposition before the Court is asked to set a lump sum

The next issue goes deeper into one of these. If you want a specific function or sector covered, reply to this email.

— Liga


TwinLadder Weekly is a weekly intelligence report on judgment, governance, and the boards accountable for both. Subscribe at twinladder.ai/newsletter. Forward this issue freely.