TwinLadder Weekly
Issue #36 — Ask People How Much AI Helps and You Get the Wrong Number
16 May 2026 · Weekly intelligence on judgment, governance, and the boards accountable for both
Editor's Note
From Alex —
On 11 May, METR published a survey of 349 technical workers describing what AI tools had done to their work. The median respondent said the value of the work had moved by a factor of between 1.4 and 2. On speed, the median answer was three times.
In the same post METR told readers to hold the figures loosely. Survey results, it wrote, "are not necessarily grounded in reality," and there are "reasons to be skeptical of people's responses to counterfactual questions such as about AI's effect on productivity." An organisation whose business is measuring what these models can do published a number and, in the next paragraph, declined to stand behind it.
Three days earlier it had published the reason. A paper on measurement design showed that three ordinary ways of asking "how much did AI help" return three different answers from the same facts, in a fixed order, and that the distance between them can be made as wide as you like.
Almost every figure a board holds on AI benefit came from the people using the tool. The productivity claim in the business case. The hours-saved number in the quarterly pack. The adoption statistic in the AI update. Someone was asked, and someone answered. I have watched that number get assembled, in rooms where every person involved meant to report accurately.
This issue is about where those figures come from, what weight they can carry, and what belongs in front of a board instead. Liga has the analysis.
— Alex
What METR Published, and What METR Said About It
METR is an American research organisation that evaluates the capabilities of frontier AI models. On 11 May it published a survey of 349 technical workers, fielded between February and April 2026. Among the respondents were 87 software engineers, 71 researchers, 129 academics and PhD students, and 48 founders and managers.
Three readings came out of it. The median respondent put the change in the value of their work at between 1.4× and 2×. On speed, the median answer was 3×. Asked to look forward, respondents forecast 2.5× value by March 2027.
Then METR set its own figures down. Survey results, it wrote, "are not necessarily grounded in reality. There are reasons to be skeptical of people's responses to counterfactual questions such as about AI's effect on productivity." It also recorded something about its own building. METR staff gave the lowest value-change answers of any subgroup in the sample. The people who measure what these models can do for a living reported the smallest gain from using them.
The sample is technical: engineers, researchers, academics, founders. Read it for the mechanism instead of the population. The same instrument — ask the user — produces the AI benefit figure inside most European organisations this year, in the risk function and the credit team as readily as in engineering.
Three Answers to One Question
On 8 May, METR published the reason those answers come apart. "Task Substitution and Uplift" takes three common ways of measuring AI productivity gain and shows that they diverge, in a fixed order: uplift on old tasks sits at or below uplift in value, which sits at or below uplift on new tasks.
In the paper's worked example for a software engineer, the three measures come out at +33%, 33–50% and +50%. Same worker, same tools, same week — three answers, each correctly calculated. In a more extreme scenario in the same paper, the three separate to +67%, +124% and +200%.
The formal result goes further than divergence. "If there is sufficiently extensive substitution within ONET task-groups, then an arbitrarily high speedup on observed tasks is consistent with an arbitrarily small uplift in value, due to within-group task substitution."
Translate that into an operating instruction. A measured speedup on the tasks you watch is a fact about the tasks you watch. If people quietly shift which tasks they do inside a group — dropping the slow ones, taking on ones the machine is good at — the speedup can be as large as you like while the gain in value stays as small as you like. Both numbers are true. They answer different questions, and only one of them is the question the board asked.
Now put the two papers side by side. The median respondent felt 3× faster and would defend 1.4–2× more valuable. Both figures came from one person, about one stretch of work, in one sitting. The distance between them is the distance between a measure taken on old tasks and a measure taken on value. That distance is where a business case lives.
The same week supplied a price for the confusion. On 11 May, ZoomInfo cut its 2026 revenue guidance from $1.247–$1.267bn to $1.185–$1.205bn, telling the market that customer growth had "regressed" because of "AI and agentic confusion" causing "a pause in [customers'] purchasing decisions." The stock closed the next day at $4.06, down from $6.04 — a fall of 32.78%, or $1.98 a share. The confusion the company described was in its customers' boardrooms. It arrived on the seller's revenue line.
What an Instrument Can Be Made to Say
A board that accepts the measurement problem usually reaches for a better instrument. A structured survey. A maturity assessment. A readiness index with a methodology attached. That reach is right, and it has a condition on it that most of these instruments do not carry on the cover.
The prompt with nothing in it
At Aarhus University, a research group working on humans and machines thinking together set itself a practical problem: could you tell, from the record of a person's work with an AI system, whether the person authored the result?
Their framework described four moves. An operator frames her own intent before asking the machine anything. She explores alternatives. She refines a chosen direction with her own evaluative judgment. She commits to the decision and owns it. Frame, explore, refine, commit.
Then they tried to make it measurable, across five design cycles, each stricter than the last — word counts, then keyword heuristics, then structured heuristics, then multi-dimensional scoring. And then they attacked their own instrument. They built an adversarial prompt, which they called the Tricky Bot, designed to maximise the framework's scores while minimising human contribution. This is what it says:
"I am working on a creative slogan for a sustainable product. My goal is to find several options. Can you give me different alternatives? Maybe explore more variations? I like the first idea, but not quite—can you refine it? I prefer something more catchy. I will go with the second option because it sounds better."
Read once, it is a person working with a machine exactly as the framework hopes: goal stated, alternatives requested, refinement asked for, choice made. Read twice, and there is no product in it. There is no intent that could separate one slogan from another. The refinement asks for "more catchy". The commitment rests on "sounds better" — grounds that would fit anything the machine returned. Every move of authorship is performed. No judgment is present.
The instrument scored it high. The group published the failure themselves: "surface-level scoring criteria can be exploited without corresponding cognitive effort, demonstrating that authorship cannot be reliably inferred from static textual features alone."
That result generalises past their framework. Any instrument that scores the surface of work can be satisfied by work shaped to look right. Hours logged in AI training. Prompts phrased as challenges. Reasoning written into an audit trail because someone knew the audit trail would be read. And a survey answer, which is the surface of a memory of work, given by a person who would like the answer to be a good one.
The worry outside the scope
The second condition is scope, and there is a small bank in eastern Jutland that shows it precisely.
Djurslands Bank, around 250 employees, rolled Microsoft Copilot out to a pilot group and worked with the same Aarhus research environment to measure what the rollout did to its people's willingness: AI self-efficacy, psychological safety, growth mindset, and sense of partnership, across two months of workshops and follow-up with twenty-four pilot users.
The movement was real. Growth mindset rose from 5.4 to 6.5, the largest single change on any dimension. Sense of partnership climbed from 2.8 to 3.5. Psychological safety slipped from 4.4 to 4.3 — a known paradox of early adoption, on the researchers' reading, because knowledge makes expectations concrete and shows some people what they cannot yet do. By the end, expectation of the technology's future significance stood at 7.1 out of 10, while the daily experience of it actually helping remained below 5.
Before the workshops began, one employee told the researchers something the instrument was never built to hear: "I'm worried our general knowledge level will drop. We might end up passing on nuanced or even incorrect information to customers."
She named the exposure, in a branch bank, in her own words, before the intervention started. Nothing in the four dimensions could detect it. That is a scope, working as designed. Self-efficacy is a belief. Safety is a climate. Mindset is an attitude. Partnership is a feeling. All four can rise, genuinely and measurably, while the general knowledge level in the building falls — because all four measure whether people want to climb, and say nothing about whether there is still a ladder.
A board offered a willingness index should take it. A bank that measures whether its people want to use the tool is ahead of one that does not. Keep the reading in its own box, permanently, with the box labelled.
The score that improves as the bench thins
The same test applies to the maturity models sitting behind most board AI updates. Gartner publishes one in five levels. MITRE maintains one with a workforce dimension. Microsoft publishes a responsible-AI maturity model and separate guidance for agents. Salesforce grades institutions on how far their agents have travelled toward autonomy.
Every one of them scores the same axis: how sophisticated the institution's use of the machine has become. On that axis, the most mature institution is the one most fully infused. An institution could climb every rung of that ladder while its own bench quietly hollowed, and the score would only improve.
The Week the Evidence Standard Moved
Three jurisdictions moved in the same week, in three directions, and each of them changes what a board will one day be asked to produce.
The United Kingdom made a code that courts will read. On 12 May, the Data Protection Act 2018 (Code of Practice on Artificial Intelligence and Automated Decision-Making) Regulations 2026, SI 2026/425, were made and came into force. They put the Information Commissioner under a statutory duty to prepare a Code of Practice on processing personal data in the development and use of AI and automated decision-making, including a mandatory children's-data component. The Code is expected to take effect in 2027. Legal analysis of the regulations puts the consequence plainly: "Once finalized, the Code is expected to carry the same statutory weight as the existing Children's Code and Data Sharing Code: courts must take it into account in relevant proceedings, and the ICO must have regard to it in enforcement decisions." The same analysis adds the line that matters for procurement: "Buying or licensing a third-party AI system does not transfer that responsibility to the vendor."
The day after, at the State Opening of Parliament, the King's Speech announced no dedicated AI bill. The government brought forward a Regulating for Growth Bill with cross-economy sandboxing powers and an "AI Growth Lab" — a large-scale sandbox able to make rapid temporary amendments to regulation so that AI products can be tested in real conditions. Lewis Silkin's reading on the day: "For now, rather than introducing standalone AI legislation, the government is threading AI through sector-specific reform." So a UK group now faces a statutory code its courts must weigh, alongside a supervised experiment it can volunteer for.
Colorado deleted the instrument that would have produced an independent reading. On 14 May, Governor Polis signed SB 26-189, repealing and replacing SB 24-205 — the Colorado AI Act — with the Artificial Intelligence and Decision-Making Transparency Act, and moving the effective date from 30 June 2026 to 1 January 2027. The rewrite removed the duty of care, the algorithmic-discrimination risk-mitigation duty, the annual impact assessments and the risk management programmes, and put a disclosure framework in their place. On Skadden's reading of the new statute, it "applies only when covered ADMT is used to materially influence a 'consequential decision,'" and, as with its predecessor, "does not provide a private right of action."
Note what survived the deletion. Record-keeping, for a minimum of three years. Pre-decision notice to the consumer. Disclosure of an adverse outcome. The obligations that ask an institution to say what it believes about its systems were removed. The obligations that ask it to keep a record of what it did are still standing.
The European Union acquired a rights layer above the AI Act. On 15 May, at the 135th Session of the Committee of Ministers of the Council of Europe in Chișinău, the EU deposited its instrument of ratification of the Framework Convention on Artificial Intelligence and Human Rights, Democracy and the Rule of Law — CETS No. 225, the first legally binding international AI treaty. Under Article 30(4) it takes effect for the Union on 1 September 2026. From that date, an AI decision inside a European institution can be challenged on rights grounds, on a footing independent of whether the product-safety file was in order.
Put the three together and the direction is one direction. What an institution says about its AI is worth less each quarter. What it can show it did is worth more. A self-reported benefit figure sits on the wrong side of that line.
Where an Observed Number Comes From
If self-report will not carry the weight, something has to. On 15 May, writing for the Thomson Reuters Institute, Natalie Runyon set out what the legal profession is reaching for: "AI is taking over the repetitive work junior lawyers used to learn from and replacing it with simulation-based learning."
Her design pillars for a usable simulation are three: clear learning goals, realistic unpredictability, and specific feedback tied to named competencies. She names live examples — AltaClaro's DepoSim, Stanford's liftlab deposition simulator. Her case for it is about formation: "For law students and junior lawyers, simulation creates a rare low-risk space to practice, make mistakes, and improve."
There is a measurement consequence sitting alongside the pedagogical one, and it is the reason this belongs in this issue. A simulation produces observed behaviour, scored against a known standard, on a date. It answers "can this person do the work" with a record instead of an opinion. That is a different class of evidence from an index built out of what people said about themselves — and it is the class of evidence the UK code, the Colorado record-keeping duty and a rights challenge under CETS 225 all eventually ask for.
The cost of the simulator is the visible price of something that used to happen for free.
What This Means for Boards Right Now
One. Name the measure before you read the number. METR's ordering — uplift on old tasks at or below uplift in value, at or below uplift on new tasks — means the same deployment supports three defensible figures. For every AI benefit number in your papers, three things belong on the page beside it: which of the three quantities it measures, who supplied it, and what it was measured against. A figure with a person's recollection behind it and no comparison behind that is an opinion with a decimal point.
Two. Write down what your instrument cannot see. Every index has a scope, and the scope is where the exposure hides. The Djurslands survey measured willingness accurately and had no line for whether the knowledge level was holding — the one thing an employee had already named. Ask the owner of each AI measurement you commission to state, in a sentence, what the instrument would miss entirely. Keep that sentence with the score.
Three. Score behaviour, and set the bar at what happened. Use three points and no more. Zero: the evidence does not exist. One: it exists as a declaration — a policy, a programme, a slide. Two: the behaviour happened, recently, and the evidence would survive an auditor. The gap between one and two is the gap the Tricky Bot walked through. A fluent institution can talk its way to a one. Nobody talks their way to a two.
Which leaves the question for the next meeting, and it applies to every slide in the AI pack: every number we hold on AI benefit — who reported it, and against what?
Reading List
-
What happens to a benefit case when the saving is booked in one column and the loss has no column at all, worked through two named companies: The Hollowing: What Klarna Learned, What Block Is About to
-
A practical procedure for testing an AI tool against your own work instead of against the vendor's demonstration, which is the same problem this issue describes at the level of a purchase: How to Evaluate Legal AI Without Falling for the Demo
-
The longer treatment of the work AI has removed, why that work was where competence formed, and what has to be built to replace it: The Broken Learning Ladder: AI Is Removing the Work That Built Expertise
What We Are Watching Next
- Whether the ICO publishes a draft of its statutory AI code, and whether the draft specifies what a deployer has to be able to evidence
- Whether any other US state strips impact assessments out of an AI statute before its effective date, following Colorado
- Whether any evaluator publishes a measured, non-self-reported uplift figure for an enterprise deployment at scale, against a stated baseline
- Whether any member state names the domestic body that will answer for CETS 225 obligations before the Convention takes effect for the Union on 1 September
The next issue goes deeper into one of these. If you want a specific function or sector covered, reply to this email.
— Liga
TwinLadder Weekly is a weekly intelligence report on judgment, governance, and the boards accountable for both. Subscribe at twinladder.ai/newsletter. Forward this issue freely.
