TWINLADDER
TwinLadder logoTwinLadder
Back to Archive
TwinLadder Intelligence
Issue #39

TwinLadder Weekly

June 2026

TwinLadder Weekly

Issue #39 — How to Tell Who Can Still Do the Work Without the Machine

6 June 2026 · Weekly intelligence on judgment, governance, and the boards accountable for both


Editor's Note

From Alex —

Every board I have sat with this year has been shown an adoption number. Seats deployed. Prompts issued. Hours saved. A maturity rating with an arrow on it. I have yet to see one shown the number underneath: how much of the work would come out right if the tools were switched off on Monday.

The question sounds theatrical until you count how often it gets asked by accident. An outage. A vendor deprecating a model version. A regulator asking a firm to reconstruct, in November, a decision it made in March. Each of those is the same examination, administered without warning, by somebody outside the building, on a day the institution did not choose.

On 4 June a paper appeared on arXiv about competitive programming. It sets out a way to ask the question on purpose. Song Yao's argument is that an examination which prohibits AI, sat in a room with a proctor, does something a usage dashboard cannot: it tells you which of your people used the machine to build skill and which used it in place of skill. The two look the same in the telemetry. They come apart under the gate.

The same week, BCG's fourth annual survey of workers found that 72% say AI has already considerably changed the skills expected in their role, and that 47% now spend more time managing and directing AI than doing the underlying work. And on 3 June, the ECB's Frank Elderson told a conference in Zurich that the challenges of new-generation AI models "should not be viewed solely as a cybersecurity issue – they are a firm-wide strategic challenge."

A firm-wide strategic challenge is something a board owns. Owning it requires a reading. The instrument for taking that reading was built decades ago by medicine, by law and, before either of them, by aviation — and Yao's paper names them.

Liga has the analysis.

— Alex


What the Paper Did

The paper is "When the Scaffold Stays On: AI, Practice Style, and Screening in Elite Skill Formation", by Song Yao, posted to arXiv on 4 June as a preprint. It has not been peer-reviewed, and it should be read on that footing.

Its setting is competitive programming, chosen because the field runs two regimes side by side. Codeforces contests are online, unproctored and open to anyone. The International Collegiate Programming Contest and the International Olympiad in Informatics prohibit AI, run under proctoring, and admit entrants through qualification rounds. One population, one skill, two rooms.

From Codeforces submission histories Yao builds what he calls an AI-prompt signature: more first-attempt acceptances, fewer attempts, fewer debugging retries. The trace of someone whose code arrives already working.

Three findings follow. Practice across entry cohorts shifted toward that signature over two AI rollouts. In the open, unproctored contests, a stronger signature predicts smaller rating gains among entrants with no ICPC or IOI affiliation, and does not do so among those who qualified. And inside the AI-prohibited environment, a shift toward AI-style practice predicts higher unassisted scores for AI-era entrants.

Hold the second finding against the third. The abstract states the result plainly: "The same practice input carries opposite signs depending on whether the environment screens for it."

The same behaviour, measured the same way, points in two directions depending on which room the person is standing in. Which is what a screening instrument is for. The paper puts the mechanism in terms a board will recognise: "The sharper question is whether selection mechanisms can screen apart two coexisting types: substitute-users, who use AI in place of deliberate practice, and complement-users, who use it to accelerate skill development."

Two types, one payroll. Both of them show up in the seat-utilisation report.

Then the paper reaches past its own field, in a sentence addressed to institutions. The two levers it identifies are how AI is integrated into training and "the design of AI-prohibited evaluation gates as a type-separating institution" — and, it continues, "Both extend beyond programming to credentialing systems (medical and legal boards, professional certification) that certify skill in a workforce increasingly shaped by AI."

That is a design specification for an instrument most companies could build this quarter.


Why the Telemetry Cannot Separate the Two Types

A substitute-user and a complement-user produce identical logs. Same seat, same prompt volume, same throughput, same satisfaction score. On the measures a firm currently collects, the person who is compounding capability and the person who is spending it look like the same employee.

METR spent May demonstrating how far that gap runs.

On 8 May it published Task Substitution and Uplift, showing formally that three common measures of AI productivity gain diverge, and establishing the ordering: uplift on old tasks ≤ uplift in value ≤ uplift on new tasks. In its worked software-engineering example the three came out at +33%, 33–50% and +50%. In a more extreme scenario they diverged to +67%, +124% and +200%. The paper also establishes that "an arbitrarily high speedup on observed tasks is consistent with an arbitrarily small uplift in value, due to within-group task substitution."

A board reading a single productivity figure is therefore reading one of three incompatible things, and the choice of which was made further down the org chart.

On 11 May METR published a survey of 349 technical workers — 87 software engineers, 71 researchers, 129 academics and PhD students, 48 founders and managers — conducted between February and April. Median self-reported speed gain: 3x. Median self-reported change in the value of the work: 1.4–2x. Respondents forecast 2.5x value by March 2027. In the same post, METR warned against its own findings: "survey results are not necessarily grounded in reality. There are reasons to be skeptical of people's responses to counterfactual questions such as about AI's effect on productivity." METR's own staff returned the lowest value-change answers of any subgroup.

So a board holds two readings. The dashboard measures the machine. The survey measures what people believe about the machine. Neither one measures the person.

The BCG survey supplies the surface of the same thing at scale. Across 11,749 workers in 14 markets, 72% report that AI has already considerably changed skills expectations in their role and 47% spend more time directing AI than doing the work. In the same study, 42% of regular users save at least a full working day a week — and 66% receive little or no guidance on how to reinvest the recovered time.

Read those two together. The hours came back and nobody said where to put them. A day a week returned to a professional with no instruction attached is a day a week that will be spent on more throughput, because throughput is what gets counted.

Seed the test rather than wait for a real one

There is a way to get a reading, and it does not require a research programme.

A watcher tested only when the machine happens to fail is tested too rarely and too late — until, once, expensively, it is not. So the test gets seeded. Take a review function where the people see only what the model surfaces: a fraud desk, a claims queue, a transaction-monitoring team. Each month, drop a handful of known-bad cases back into the stream — cases the machine has already cleared, with the correct answer recorded in advance. Nobody on the desk knows which ones they are. At month end there is a number no dashboard would otherwise produce: of the planted misses, how many did the human catch?

That is a fire drill for judgment, and it measures the eye directly instead of inferring it. Eight of ten planted misses caught last quarter, with the log attached, is a fact a board can hold. "The reviewers are experienced" is not.

One counting rule makes the difference between an instrument and a comfort. Count the re-judgments, not the overrides.

An override is the case where the independent judgment happened to differ. It is the outcome. The control is the independent judgment itself — the decision someone formed again, from the file, without seeing the machine's answer first. A desk with a falling override rate may have a better machine, or it may have a blind desk. The two produce the same number, and only one of them is good news. What separates them is whether anyone is re-judging at all, sampled blind, against a benchmark held aside.


The Professions That Already Run the Gate

Yao's paper names medical and legal boards because those professions settled this question long before the current tools existed. They decided that competence had to be demonstrated on a schedule, unassisted, in a room, by every practitioner — including the ones nobody doubted.

Aviation built the same instrument earlier and runs it more often. On 20 May, at the Tulum Air Show, Sereya Schotborgh, Regional Officer for Safety Implementation at ICAO's North America, Central America and Caribbean office, sat on a panel titled "Technology and Automation and How They Affect Manual Flight." She argued that automation supports pilots rather than replacing them, that system failure remains a realistic operational scenario, and — this is the trade press's account of spoken remarks, so take it as one source — that "manual flight remains aviation's foundational safety layer, enabling crews to stabilize operations when technology encounters limitations."

A pilot who has flown an automated approach every day for two years is checked, on a schedule, on the manual approach she has not flown in two years. The industry prices that check as a cost of operating. It argues about the interval and holds the check.

Law has begun improvising toward the same answer. On 15 May, Natalie Runyon wrote in the Thomson Reuters Institute that "AI is taking over the repetitive work junior lawyers used to learn from and replacing it with simulation-based learning." She sets out three design pillars for a usable simulation — clear learning goals, realistic unpredictability, and specific feedback tied to named competencies — and names live examples: AltaClaro's DepoSim, Stanford's liftlab deposition simulator. "For law students and junior lawyers," she writes, "simulation creates a rare low-risk space to practice, make mistakes, and improve."

A simulator is what an institution buys once the real repetitions have stopped arriving for free. Its invoice is the first time the cost of the vanished apprenticeship appears anywhere in the accounts.


What It Costs When the Examination Is Administered From Outside

Two rulings in the ten days before this issue show the check running downstream.

On 28 May the California Court of Appeal reversed and remanded in H.C. v. Contreras. The trial judge, Irene A. Luna, had issued a ruling incorporating language nearly identical to the father's brief — including a citation to a case that does not exist and a misstatement of Family Code §6203. Opposing counsel had already flagged both errors before the ruling issued. The appellate court named the fabricated authority in its own opinion: "The trial court cited and relied on a fictitious case, i.e., Enrique M. v. Angelina V. (2005) 15 Cal.App.5th 788." And it set out the principle: "Reliance on fake cases is fundamentally incompatible with an informed exercise of discretion controlled by genuine principles of law."

The check existed. It arrived on time. It was read by nobody with the authority to act on it, and then it was copied into an order.

On 3 June the Ninth Circuit published an order in Malkeet Lnu; Sunita Rani Lnu; Jaivin Lohan v. Todd Blanche, Acting Attorney General, No. 24-4790. It sanctioned counsel $5,000 and suspended him from practice before the court for six months, with a notification requirement. The court found non-existent cases, quotations attributed to Kamalthas and Avendano-Hernandez that appear nowhere in those opinions, and further fabricated citations in the same attorney's other matters. It described the citations as ones that "were the result of hallucinations by generative AI," and set out an affirmative duty: on discovering such a hallucination, the lawyer "should immediately alert the court."

The fabrications ran across the attorney's caseload. This was a working method, and the first competent reader to encounter it was a federal appellate court. His firm had every opportunity to administer that examination first, at a cost of one afternoon.

An institution that declines to test its people's unassisted work has still scheduled the test. It has scheduled it for a court, a supervisor or a customer, at whatever price they set.


The Score That Rises While the Bench Thins

There is no shortage of AI maturity models. They score one axis: how sophisticated the institution's use of the machine has become. On that axis the most mature organisation is the most fully infused, and an institution can climb every rung while the capability of its own people quietly drains, with the score improving all the way up.

The other axis asks a different question. Can this institution still form and hold the judgment that governs the machine? Read against that axis, a firm's position is set by the weakest dimension it has, and the dimensions interlock rather than averaging out.

Which yields a rule worth importing into any board's reading of its own AI scorecard: no rung may be claimed while any dimension stands at zero. A zero means a part of the exposure is unwatched, and unwatched is where this risk has done all of its damage so far. An institution with disciplined thresholds and no idea who could rebuild its output with the tools off has fenced the wrong field. A firm with an excellent pairing programme and no override anyone has ever exercised is transmitting judgment that nobody is permitted to use.

The question that scores that dimension takes a form a board can put in a meeting and get an answer to before the coffee arrives. Take one consequential output from last month — a credit memo, a remediation plan, an opinion that went to a client. Name the person who could reconstruct and defend it, line by line, with the tools off. Then say when that was last actually tried on anything.

A name and a date. A policy, a programme and a training completion rate answer a different question.

Two things happened in the same week that make the question harder to defer. On 1 June the Commission announced that it had recruited 60 independent experts and published the full list, establishing the AI Act's Scientific Panel under Article 68 and Commission Implementing Regulation (EU) 2025/454, alongside the Advisory Forum. Members serve two-year terms, and the two bodies "will advise the Commission's AI Office and national authorities on applying rules." Their remit runs to model classification, evaluation methodologies and cross-border market surveillance. The AI Office now has independent technical capacity to second-guess a provider's own account of its systems.

And Elderson, in the Zurich speech, tied improved bank preparedness to governance arrangements and awareness "particularly among banks' management bodies."

The outside examiner is being staffed. The inside one has to be built by the people who would have to sit it.


What This Means for Boards Right Now

One. The two numbers already in the pack measure the machine, and a third one is missing. METR's ordering — uplift on old tasks ≤ uplift in value ≤ uplift on new tasks — means a single productivity figure has already settled an ambiguity, and somebody further down the org chart settled it. The self-reports run wider still: 3x speed felt, 1.4–2x value defended, by the same people in the same survey. The reading that answers the board's actual question is unassisted performance on real work, scored and archived, taken on a date the board sets in advance. A first round with nothing to compare it to still counts, because it is the baseline every later round is read against.

Two. Count the re-judgments, not the overrides, and seed the cases rather than waiting for real ones. A falling override rate is produced equally by a better machine and a deskilled desk. Sampling blind — the decision formed again from the file, before the machine's answer is seen — separates them, and planted known-bad cases give a monthly catch rate that no live workflow will hand over on its own. Set the pass rate in advance, so the result is a claim checked against a standard.

Three. A maturity score with a zero underneath it certifies nothing. Before accepting the next AI readiness rating, ask which of its dimensions scores zero and what the plan is for that one dimension. Then put the tools-down question to the executive who owns the largest automated process: one output from last month, the name of the person who could rebuild it, and the date it was last attempted.

Which leaves the question this week's evidence puts on the table, and it belongs on the agenda while it is still cheap to answer: if we withdrew the tools tomorrow, what would still get done correctly?


Reading List


What We Are Watching Next

  • Whether the AI Act Scientific Panel's first published work names evaluation methodologies specific enough for a deployer to run in-house
  • Whether the Solicitors Regulation Authority's open consultation on continuing competence ends with a duty to record how a solicitor established their competence, and whether any other European professional regulator follows
  • Whether any large professional firm publishes a tool-free proficiency result — a score, on a date, for work its people did with the tools switched off
  • Whether Yao's screening result survives peer review, and whether anyone replicates the type separation outside competitive programming

The next issue goes deeper into one of these. If you want a specific function or sector covered, reply to this email.

— Liga


TwinLadder Weekly is a weekly intelligence report on judgment, governance, and the boards accountable for both. Subscribe at twinladder.ai/newsletter. Forward this issue freely.