TwinLadder Weekly
Issue #48 — Ten Minutes of Help, and the Unassisted Score Falls
8 August 2026 · Weekly intelligence on judgment, governance, and the boards accountable for both
Editor's Note
From Alex —
On 5 August a research group posted a revised paper carrying a number I have kept coming back to all week. Across a series of randomised controlled trials with 1,222 participants, AI assistance improved performance while it was present. Afterwards, working alone, the same people scored significantly worse and gave up sooner. The effects appeared after roughly ten minutes of interaction.
Ten minutes. Set that against the clocks a board actually runs. The annual training cycle. The quarterly risk review. The half-yearly people committee. The onboarding programme measured in weeks, the competency framework refreshed every second year. Every instrument we own for watching capability samples at an interval far longer than the interval at which the thing being watched appears to move.
The day before that paper went up, a judge in Brooklyn described what the same mechanism looks like once it arrives in a courtroom. Nineteen submissions. Twenty-three cases that do not exist, cited forty-seven times. Eighty-three genuine decisions whose holdings were misstated. And when opposing counsel challenged the citations, the lawyer went back to the same generative systems and asked them to verify their own work.
He used the machine to check the machine.
I want to hold that at its proper size. He was a solo practitioner who told the court he could not afford Lexis or Westlaw, and it would be easy to file the case under one man's poor judgment and move on. The structure is more general than the man. A great many organisations have arranged their AI review exactly this way without anyone deciding to. The draft is produced by a model. The summary of the draft is produced by a model. The policy check is evaluated by a model. And the person at the end, whose signature makes it the institution's, has had less and less occasion to do the underlying work since the tools arrived.
Liga has the analysis, and the question I would take into your next risk meeting sits at the end of it.
— Alex
What the Trial Measured
Grace Liu, Brian Christian, Tsvetomira Dumbalska, Michiel A. Bakker and Rachit Dubey ran a series of randomised controlled trials on human–AI interaction with 1,222 participants. The task families were mathematical reasoning and reading comprehension. Participants were assigned to work with AI assistance or without it, and the outcome measured was how they performed afterwards, on their own — together with persistence, meaning how long a person kept working at a problem before abandoning it. Version four of the paper was posted on 5 August. It is a preprint, and it has not been through peer review.
The finding, in the authors' words:
"Across a variety of tasks, including mathematical reasoning and reading comprehension, we find that although AI assistance improves performance in the short-term, people perform significantly worse without AI and are more likely to give up. Notably, these effects emerge after only brief interactions with AI (approximately 10 minutes)."
Read the second clause with the first still in view. During the assisted phase, output improved. That is real, it is what an adoption programme measures, and it is why the tools are being bought. The unassisted score afterwards fell, and persistence fell with it. Two things moved in opposite directions inside a single sitting.
The authors give a mechanism for the persistence half: "persistence is reduced because AI conditions people to expect immediate answers, thereby denying them the experience of working through challenges on their own." They frame the paper by contrasting these systems with a mentor, who "doesn't just answer questions, but also scaffolds learning, tracks progress, and prioritizes the other person's growth over immediate results" — against systems "optimized for providing instant and complete responses, without ever saying no (unless for safety reasons)."
Now the limits, stated before anyone else states them for us. This is a laboratory study of two task families, on a timescale of minutes, in a preprint. It says nothing directly about a credit officer, a radiologist or a tax partner. Whether the effect persists for days, whether it reverses with practice, whether it holds for expert users on their own subject matter — the paper does not settle any of that, and a board should treat a single preprint as one reading rather than a verdict.
What the design does give you is a direction and an interval. The direction is causal, because people were assigned to conditions instead of being asked how they felt afterwards. The interval is minutes. And the authors are explicit about why the persistence result carries weight: persistence "is foundational to skill acquisition and is one of the strongest predictors of long-term learning."
Which produces a measurement problem before it produces a competence problem. An organisation that assesses capability once a year is sampling a variable that appears to move inside a single working session. The reading will always look stable, because the instrument was built for a slower world.
Nineteen Submissions, and the Same System Called to Check Them
On 4 August, Justice Heela D. Capell of the Supreme Court of the State of New York, Kings County, issued a decision and order after hearing in Kleyman Law Group, P.C. v. Kaloidis, Index No. 502644/2025. She imposed the maximum sanction available under 22 N.Y.C.R.R. § 130-1.2 — $10,000 — together with $36,511.35 in attorneys' fees and costs, $46,511.35 in all, and referred counsel to the Attorney Grievance Committee.
The court's own summary of the conduct: "Indeed, Kleyman filed nineteen separate submissions that collectively cited twenty-three fictitious cases on forty-seven occasions and misstated actual authority at least eighty-three times."
Alex I. Kleyman testified that he had used Claude and other generative AI platforms between April and December 2025 because he could not afford Lexis or Westlaw, and that this was his first time using generative AI for legal research. The court recorded that he "conceded that he could have attempted to verify the cited decisions through a free Internet search engine or research service and did not do so." After opposing counsel challenged the citations, he returned to the same platforms and asked them to confirm the authorities. The court's word for that was "puzzlingly".
Defence counsel spent roughly eighty hours identifying the false authorities. That is the part of the bill that appears in nobody's productivity figure. One practitioner's saved drafting hours were converted into another firm's eighty hours of forensic reading, then into a judge's time, then into a disciplinary referral. The saving stayed with the person who made it. The cost travelled.
The verification failure has a shape worth stating plainly, because it generalises well beyond a solo practice. A model cannot reliably grade its own work. Where it invented a citation, inventing and reporting feel identical from the inside — the same process produced both, and there is no internal signal distinguishing them. Asking that system to check is asking the source of the error to serve as the detector of the error.
A laboratory calibrates a balance against a reference weight held by an independent standards body. Calibrate the balance against itself and every reading agrees perfectly, on the same day the readings are wrong. The agreement is the product of the arrangement, and it contains no information about the world. That is what happened across forty-seven citations.
The Controls That Shipped in the Same Week
While that was going on, the platforms shipped real governance machinery, and it deserves to be described fairly before anything is said about its limits.
On 5 August, Anthropic put inference hooks into beta for Claude Enterprise organisations. "Point Claude at your organization's AI security server," the release note reads, "and each governed prompt across claude.ai, Cowork and Claude Code is held for the server's allow or deny verdict before inference proceeds." Requests are signed. Failure handling is configurable. Every denial is written to the compliance activity feed.
Two days later, three more primitives arrived for managed agents. Session budgets cap what a session may spend at public list rates: "A session that reaches its budget pauses with the budget_reached stop reason instead of starting new model requests; changing or removing the budget resumes it." An advisor model — one at least as capable as the agent's own — can be consulted mid-turn for strategic guidance. And inference_geo pins where inference physically runs, per agent or per session, for data residency.
On 4 August, Mistral released Shieldstral, a three-billion-parameter open-weights multimodal safety classifier under Apache 2.0. It takes the institution's policy at inference time: "you write the policy as a plain-language question at inference time, and the model returns a calibrated safety score." It runs on a single 16GB GPU, and it evaluates text and images through one interface.
These are useful, and a risk committee should want them. A hard spend cap is a circuit breaker that a firm can actually set, in a number, in advance — the first of its kind to ship as a product feature. A pre-inference veto returns to the enterprise a decision the enterprise had been unable to make: this prompt, in this context, does not run here. Geographic pinning answers a data-residency question that has been costing European deployments months. And Shieldstral moves the policy from a retraining exercise into a sentence somebody in the building has to write and own.
Every one of them governs what the machine may do. The hook decides whether inference happens. The budget decides how far an agent may go. The classifier decides whether an output crosses a line the institution drew. Each requires a named person to set a threshold and leaves a record of it, which is what a control is.
And all of it sits before, or around, the inference. The moment these controls leave untouched is the one after the output arrives, when a person reads it and decides whether it is right. That judgment has no hook, no budget and no classifier. It has a human being, and this week's trial reported that the capacity behind it is measurable, movable, and moves inside ten minutes.
Regulation reached a similar boundary in the same week from the other direction. On 2 August the AI Act's transparency obligations came into application: AI-generated or manipulated content must be "clearly and visibly labelled and include machine-readable marks", and people must be told "when they are not interacting with a real person, but an AI system, for example a chatbot, AI agent, and avatar." The national penalty framework became operative alongside it, with fines running to €15 million or 3% of worldwide annual turnover, whichever is higher. That duty tells a reader where the text in front of them came from. What the reader is able to do with it next belongs entirely to the reader.
One more item from the same days, because it names the metric most boards are relying on. On 4 August a securities fraud class action was filed in the US District Court for the Western District of Washington, docket 26-cv-02071, captioned City of St. Clair Shores Police and Fire Retirement System, et al. On the plaintiff firm's account of the complaint — which we have read, where the complaint itself we have not — Microsoft "consistently touted Copilot's best-in-class capabilities, which purportedly drove widespread and growing user adoption", when "in truth, Copilot suffered from severe functionality issues that caused user adoption to decline." The pleading puts Microsoft 365 Copilot premium customers at 15 million, below analyst expectations, disclosed at the 28 January earnings announcement that preceded a 10% one-day fall.
Leave the merits to the court. The board-relevant part sits under the allegation. Adoption is a seat count. It records how many people have the tool open. It carries no information about whether the person in the seat can still evaluate what the tool returns. Most organisations report the first number monthly and hold no instrument capable of producing the second.
The Second Skill, and Why It Decays While the Dashboard Improves
Picture a review process of the kind most large institutions now run in some form. A model triages five hundred items and surfaces twenty for human review. The twenty are read, argued over, occasionally overturned. Every control metric lives on those twenty — the override rate, the agreement rate, the audit sample, the quality score in the pack. The four hundred and eighty the model left alone have gone out of the human loop entirely. They are decided.
Two competences are running in that process, and they decay on different schedules.
The first is the skill of the work itself: the credit judgment, the clinical read, the reading of a clause against the deal it sits in. That is what the trial measured, and what a thinning caseload starves.
The second is the skill of overseeing the machine — telling whether it stayed inside the boundaries the institution set, whether it answered from the material it was given or invented a figure, whether the case it put in front of you is the one that mattered. That skill is real, it is scarce, and it carries a turn of the screw the first one lacks. The better the surrounding layer runs, the fewer exceptions reach a human, so the less the overseer practises the exact judgment the next exception will demand. An institution can automate its way to a clean dependability dashboard and a workforce that can no longer read it.
The framing effect arrives before the reading even starts. Surfaced cases come to a reviewer in an order, and an order carries a claim. The case at the top reads as the answer; the case at the bottom reads as an afterthought. The machine has framed the judgment before the person has formed one, and part of the second skill is the discipline to read against that frame — to remember that the model has said what to look at first, which is a different statement from what is true.
And here the older evidence on human oversight becomes uncomfortable, because it says the danger sits in the direction of the advice more than in its source. Saar Alon-Barkat and Madalina Busuioc, in pre-registered experiments published in the Journal of Public Administration Research and Theory in 2023, tested how administrators respond to algorithmic advice and reported: "We do not find evidence for automation bias." What they found instead was selective adherence. People adopted advice that agreed with what they already believed, and it made no significant difference whether the advice came from an algorithm or from a human expert. Advice that flatters the position you were already holding gets waved through, whatever produced it.
Set the Kleyman verification beside that. He asked the platforms to confirm citations he had already filed and had already been challenged on. He wanted a particular answer. He asked a system with no independent access to the record whether the answer was right. It agreed with him, at length and in the register of authority, across nineteen submissions.
That is the general form. A verification step drawing on the same source as the thing being verified will return agreement, and the agreement will feel like evidence. Confidence has no connection to accuracy, and a fluent confirmation is the most persuasive form confidence takes.
Meanwhile the bench that supplies the sceptical reader is thinning at its entry point. Indeed Hiring Lab's UK mid-year report, published on 3 August, found that "graduate job postings are at their lowest level for this time of year since the pandemic", with summer job postings at a four-year low. In the same report, "AI mentions in job postings have reached a record 9.4%, while jobseeker searches for AI roles have risen sevenfold since the launch of ChatGPT." Employers are asking for AI capability in record numbers at the moment the rung where people learned to evaluate anyone's work is at its narrowest in years.
What This Means for Boards Right Now
One. Match the sampling interval to the mechanism. A capability reading taken annually is measuring a quantity that a controlled trial reports moving inside a single session. Pick the two or three decisions in this house where a wrong machine output would cost most, and take a reading on the unassisted judgment behind those decisions at an interval measured in weeks. The reading is a blind one: the person forms the judgment from the file, before seeing the machine's answer. Anything else re-approves the machine and records the re-approval as a control.
Two. Require the checker and the checked to have different sources. For each material AI use, three lines belong on one page: who reads the output, what they check it against, and where that reference comes from. When the third line names the same system, or a second instance of it, the check will produce agreement on a schedule and will measure nothing. Grounding — did this output confine itself to the material it was given — is the check a model genuinely cannot run on itself, and it is the one the Kleyman citations needed.
Three. Test the watchers with seeded cases. Take a function where people see only what the model surfaces — a fraud desk, a claims queue, a transaction monitoring team — and each month drop a small number of known-bad cases, already cleared by the machine, quietly back into the stream, with the answer recorded in advance and nobody on the desk knowing which they are. At month end you have a figure no dashboard would otherwise give you: of the planted misses, how many the human caught. The board sets the pass rate before the test runs. A desk that catches eight of ten can still see. A desk that waves them through has told you, before a real one costs anything, that the review has become a signature at the bottom of a screen.
Which leaves the question this week put on the table, and it takes one meeting to answer badly and a quarter to answer properly: where in this organisation is the AI checking the AI?
Reading List
-
The verification problem set out as a working method, including where the independent reference has to come from and what a check is worth when it does not have one: A 5-Step Citation Verification Protocol for AI-Generated Research
-
The longer treatment of what happens to a profession when the tasks people learned from are the first ones automated, and what has to be built to replace them: The Broken Learning Ladder: AI Is Removing the Work That Built Expertise
-
On advice that agrees with you, why a room that converges on one source stops generating alternatives, and what a decision record has to show for the challenge to count: When AI Enters the Room, Your Best Thinking Leaves
What We Are Watching Next
- Whether the ten-minute result survives peer review, and whether anyone replicates it on professional tasks with a real deadline attached
- Whether any national market surveillance authority opens the first Article 50 transparency case, and what it accepts as sufficient labelling
- Whether any enterprise buyer publishes what it configured its pre-inference deny rules and session budgets to, now that both primitives exist as products
- Whether a court distinguishes, in sanction, between a verification step that was skipped and one performed by the same system that produced the error
- Whether UK graduate postings recover in the September cycle, or hold at the level Indeed recorded in July
The next issue goes deeper into one of these. If you want a specific function or sector covered, reply to this email.
— Liga
TwinLadder Weekly is a weekly intelligence report on judgment, governance, and the boards accountable for both. Subscribe at twinladder.ai/newsletter. Forward this issue freely.

