TWINLADDER
TwinLadder logoTwinLadder
Atpakaļ uz apskatu

Izdevums #37

Capable and Not Wise

Within twenty-four hours in May, a frontier-model evaluator and an aviation regulator said the same thing from opposite directions. METR's Frontier Risk Report of 19 May found agents had essentially saturated its eight-hour time-horizon benchmark while performing significantly worse on strategic judgment, stealth and adversary modelling; the next day ICAO's Sereya Schotborgh told a panel in Tulum that manual flight remains aviation's foundational safety layer. Capability and judgment are separate quantities that can be measured apart — aviation decided to keep the second one alive on a clock, with an examiner and a consequence, and knowledge work has not started.

AI capability evaluation
judgment formation
board oversight
professional regulation
competence maintenance
2026. gada 23. maijs13 min read
Capable and Not Wise

TwinLadder Weekly

Issue #37 — Capable and Not Wise

23 May 2026 · Weekly intelligence on judgment, governance, and the boards accountable for both


Editor's Note

From Alex —

On 19 May, METR published a report on what AI agents inside frontier labs could do if they went off the leash. The most capable agents it evaluated had essentially run out of room on its eight-hour benchmark — a handful of longer tasks still standing. On the benchmarks built to test strategic judgment, stealth and the modelling of an adversary, the same agents did markedly worse. METR's own sentence: agents "showed significantly weaker performance on benchmarks designed to evaluate strategic judgment, stealth, and the ability to model adversaries."

The next day, at an air show in Tulum, an ICAO safety officer named Sereya Schotborgh told a panel that manual flight "remains aviation's foundational safety layer."

Twenty-four hours apart. One statement about machines, one about people, carrying the same content. The ability to execute a task and the ability to judge what the task is for are separate quantities. They can be measured apart, and they can move in opposite directions.

This newsletter has made that claim for months on inference. A frontier-model evaluator has now measured it, and an aviation regulator restated it from the human side the following morning. Aviation reached the conclusion decades ago and built an expensive apparatus around it — a clock, an examiner, and a consequence for a lapse. Knowledge work has the same finding and has built nothing.

Liga has the analysis.

— Alex


What METR Measured on 19 May

METR is an American research organisation that evaluates what frontier AI models can do. Last week this newsletter took apart its survey of 349 technical workers and the caution METR attached to its own figures: survey results "are not necessarily grounded in reality." On 19 May it published a measurement instead.

The Frontier Risk Report is a pilot assessment of what METR calls rogue deployment: an AI agent running inside a frontier lab starting and sustaining an unauthorised deployment of itself. It covers 16 February to 16 March 2026 and was run with Anthropic, Google, Meta and OpenAI. The subject is agents at their own developers' desks, holding the access those agents actually hold.

Two results sit next to each other.

The first concerns duration. The most capable agents evaluated essentially saturated METR's Time Horizon 1.1 benchmark, with only a handful of tasks longer than eight hours still unsolved. Eight hours is a piece of work with a morning, an afternoon and a sequence running between them — the length at which a task stops being a request and becomes a job.

The second concerns judgment. On the benchmarks designed to test strategic judgment, stealth and the ability to model an adversary, the same agents performed significantly worse.

METR's overall reading of the pilot: the agents plausibly had the means, motive and opportunity to begin a minimal rogue deployment, while falling well short of concealing one at scale from an active internal investigation. They could start the thing. They could not run the game around it.

Past the security question sits the shape of the finding. One organisation, one evaluation window, one population of agents, two families of benchmark — and the two curves came apart. Duration ran to saturation. Judgment stayed behind.

That result belongs on a board agenda for reasons that have little to do with rogue deployment. Every capability claim a board receives arrives as one number. A percentage of tasks automated. A multiple on drafting speed. A tier on a vendor's maturity chart. One number implies one quantity moving upward: if the machine is at 80% on the task, the reasoning around the task must be near 80% too. METR measured the two in the same building, in the same month, on the same systems, and they arrived at different places.

If the two quantities are separable in a model, they are separable in a person. Nobody has been measuring the second one.


The Next Day, at an Air Show in Tulum

On 20 May, at the Tulum Air Show, Sereya Schotborgh — Regional Officer for Safety Implementation at ICAO's North America, Central America and Caribbean Office — spoke on a panel titled "Technology and Automation and How They Affect Manual Flight." Her position, as reported: automation supports pilots, system failure remains a realistic operational scenario, and mental preparedness matters as much as technical proficiency now that automated flight dominates daily operations. Two sentences carry it. "Automation systems are designed to support pilots rather than replace them entirely, yet system failures remain a realistic operational scenario." And: "Manual flight remains aviation's foundational safety layer, enabling crews to stabilize operations when technology encounters limitations."

The remarks reach us through trade press covering the panel, as one official's spoken position; ICAO issued no document. Their value sits in what stands behind them: the industry she was addressing already runs the discipline she described.

What the schedule actually says

Consider what an airline may not do. It may not put a captain in command of a passenger flight unless, within the preceding six months, that captain has completed a proficiency check — a formal examination of flying skill, administered by an authorised examiner — or an approved course in a full-flight simulator; and unless, within the preceding twelve months, she has passed the full check in that aircraft type. Under the European rules the same discipline appears as an operator proficiency check valid for six calendar months.

Those dates bind. A pilot whose checks have lapsed cannot lawfully be rostered. The rule goes past asking an airline to value proficiency: it makes lapsed proficiency an operational impossibility, on a named date, for a named person.

The manual-flying part of that regime began with a regulator noticing that pilots' hands had gone quiet. In January 2013 the US Federal Aviation Administration issued Safety Alert for Operators 13002 — its instrument for urging airlines without binding them — warning that "continuous use of autoflight systems could lead to degradation of the pilot's ability to quickly recover the aircraft from an undesired state." In May 2017 a successor alert, SAFO 17007, told operators to ensure that proficiency in manual flight operations is "developed and maintained," on the philosophy that "manual flight is the foundation upon which other technical flying skills are built."

The erosion those alerts answer was measured. Haslbeck and Hoermann, publishing in Human Factors in 2016, took the fine-motor flying record of 126 airline pilots and found that recent manual practice predicted manual skill more strongly than a career's accumulated experience. Hours in the last few months beat hours in the logbook. A firm's answer to "can our people still do this" is almost always the logbook.

The recommendation that changed the rule

Air France 447 was lost over the Atlantic on 1 June 2009. France's Bureau d'Enquêtes et d'Analyses published its final report in July 2012, and among the recommendations it closed with was one addressed to the European regulator: review the content of check and training programmes and make mandatory, in particular, "the setting up of specific and regular exercises dedicated to manual aircraft handling."

The reasoning it gave for that recommendation is a single sentence. "Manual aeroplane handling cannot be improvised and requires precision and measured inputs on the flight controls."

Cannot be improvised. A capability exists, maintained, on the day it is needed, or it is absent on that day. There is no third state in which it can be summoned because the situation has become serious enough to warrant it. One accident settles nothing about a population, and the investigators were careful about causation across a long report. What travels is the recommendation, a general statement about how a perishable skill behaves.

Both sides of the Atlantic wrote it into law. Europe built upset-prevention and recovery elements into recurrent training on a twelve-month cycle and made the advanced course mandatory through Regulation (EU) 2018/1974. The United States finalised a rule in November 2013 requiring hands-on training for exactly the scenarios in which automation hands back a degraded aircraft — manually controlled slow flight, manually controlled loss of reliable airspeed, upset recovery — in a Level C or higher full-flight simulator, on a recurrent clock, with compliance required no later than 12 March 2019. A standard proficiency check may not be substituted for it.

That arc has a shape. An alert that urged, then evidence that the erosion was real, then an investigation that produced a sentence, then a rule with a date, a simulator, an examiner and a bar on substituting the easier check.

The discipline also argues with itself. In January 2016 the inspector general of the US Department of Transportation audited the FAA and reported that the agency "does not have a sufficient process to assess a pilot's ability to monitor flight deck automation systems and manual flying skills." The regulator that wrote the rule was told its own assessment process fell short. A maintenance regime real enough to be audited and argued over is a living one. The alternative to arguing about the sufficiency of your process is having no process to argue about.


The Pipeline, Not the Cockpit

The instinct at this point is to reach for the cockpit analogy and run it straight into the firm. Simulators for underwriters. Check rides for credit officers. That instinct needs a correction before it becomes a budget line.

Start with the mechanism, which belongs to Lisanne Bainbridge. Automate the routine parts of a job and two things follow. The operator's skill decays through disuse, because skill is maintained by exercise. And the operator is left with the residue — the cases the automation cannot handle — at precisely the moment his practice at handling anything has thinned. Bainbridge was writing about a control-room operator watching dials that rarely move, his hand resting near controls he almost never touches. Monitoring keeps him present. It does not keep him ready.

Endsley and Kiris measured a version of this in 1995: operators kept out of the active control loop by automation lose situation awareness, are slower to notice that something has gone wrong, and slower to act once they do. The person watching the machine is a different person from the one who was flying it.

For a board, the transposition runs to the pipeline. The junior who never works the routine files becomes someone who was never made into a senior at all, because the routine files were the curriculum.

Then the second move, where the standard governance answer inverts. If the human is left with the hard cases, write the hard cases down — build the checklist, encode the playbook, let the tool flag exceptions for human review. Charles Perrow explained in 1984 why that fails in exactly the cases it is bought for. Procedures encode the failures you have already understood. A checklist is a museum of yesterday's accidents: a record of events someone has seen, diagnosed and written down. The events that take a complex system down are the ones nobody wrote down, because they had not happened yet.

Together the two ironies open a seam. The AI handles everything the procedure covers — every case that yesterday's failures taught the firm to anticipate. The human is left with the un-proceduralised residue, which is the population of cases where judgment is the only available tool. And that judgment has been draining the whole time, because the reps that built it went to the machine.

Which is why the aviation answer resists transplant. The cockpit is a closed, instrumented world: failure modes bounded and catalogued, feedback fast and unambiguous, the rare emergency reproducible on demand. A board running an open-domain judgment task — credit, underwriting, clinical triage, legal analysis — has none of those gifts. The failure modes have no catalogue. The feedback arrives slowly and full of noise. Nothing reproduces next year's genuinely novel hard case so that a person can rehearse the recovery.

The governance layer is what transfers. The calendar that does not negotiate. The scorer who is independent of the person scored. The consequence that follows a lapse. The signature that owns the result. Aviation's contribution to this question is a demonstration, run inside a regulated industry for decades: scheduled human-capability maintenance is a governable discipline, with paperwork.


A Duty Written Down, and a Classification Handed Back

Two regulators moved in the same week, in opposite directions.

On 18 May the Bar Standards Board's guidance on the use of artificial intelligence and other technologies came into effect for barristers in England and Wales. It creates no separate AI rulebook. It reads AI into the Core Duties already in the Handbook — the duty to the court, the client's best interests, integrity, confidentiality, competence and practice management — and then goes past a permission regime.

Under Core Duty 7, competence: "You should maintain a sufficient level of competence in technology and AI to understand how they may impact your practice, whether or not you adopt those technologies yourself." A barrister who has never opened a generative AI tool carries a duty to understand what it does to opposing counsel's submissions, to a client's documents, to the admissibility of prompt histories. The regulator has imposed a competence obligation on abstention.

The guidance then sets out a risk-based approach across three factors — the application, the use and the technology itself — and puts the burden of proof on the individual: "Barristers should be able to demonstrate they have taken a reasonably active and informed approach to managing technology risks, proportionate to the risk or harm (Core Duty 10)." Demonstrate, to someone, later. On evaluation it is specific about the skill required, expecting barristers who use AI for legal analysis to understand how to methodically evaluate the quality and reliability of the outputs and to apply legal principles to them to reach reasoned decisions.

A professional regulator has described the judgment layer and attached it to a named individual.

The Commission moved the other way on 19 May, publishing draft guidelines under Article 6(5) of the AI Act on how to classify a high-risk AI system — general principles, the Annex I regulated-product route, and the Annex III stand-alone use cases covering biometrics, education, employment, essential services and law enforcement. A targeted stakeholder consultation runs to 23 June. The guidelines were originally due on 2 February and arrived late.

The operative caution is in the Commission's own summary. The draft "provide[s] non-exhaustive examples of AI systems that may or may not be classified as high-risk, while making clear that inclusion of a use case does not by itself establish its lawfulness under applicable law." Whether a company's hiring screen, credit model or access-control system is high-risk determines the entire compliance bill for that system. The Commission has now said in writing that its examples leave the question open. Somebody inside the building reads the guidelines against the actual system and forms a view, and that view is a judgment with a name on it.

Two more items the same week ran in the same key. Reporting on the OCC's semiannual risk perspective recorded the regulator's finding that AI is "significantly transforming" the cybersecurity threat landscape for banks, lowering barriers to entry and raising the speed and scale of attacks; the OCC, the Federal Reserve and the FDIC signalled a forthcoming request for information on model risk management for generative and agentic AI. And on 19 May Anthropic made self-hosted sandboxes available for its managed agents, "as an alternative to running tool execution in Anthropic's infrastructure," with tunnels into private-network tool servers in research preview. Where an agent's hands physically are is now a purchasing decision, and therefore an allocation of risk a board can be asked about.


What This Means for Boards Right Now

One. Split the capability number in two before you accept it. A deployment report carries one figure standing for the whole of what the system can do. Ask for two: what the system completes, and what the system judges. METR ran both families of benchmark on the same agents in the same month and got different answers. Any deployment where the second answer has never been examined is a deployment governed by an assumption.

Two. Name the perishable skills and put a date on each. Aviation's method needs no simulator to begin: identify the capability that must exist on the day it is needed, decide how it will be exercised, fix the interval, name the person who checks, and state what happens when a check is missed. Manual aeroplane handling "cannot be improvised," and a credit judgment does not arrive because a situation has become serious enough to need one. Start with the three or four decisions in this house that a machine now prepares and a human still signs.

Three. The checklist covers the cases the machine already handles. Every AI control framework in operation encodes the failure modes someone has already seen. The cases that reach a court, a regulator or a front page are the ones outside it, and they arrive at a human whose practice has been thinning since the tool went in. Read the framework once with that in mind and mark the sections that only work if the reviewer is still sharp. Those sections are the competence layer, whether or not anyone has budgeted for it.

The BSB has told individual barristers to be ready to demonstrate an active and informed approach. The Commission has told deployers that its examples leave their classification open. The banking agencies have described a threat and left the control to the firm. Each of the three carries the finding METR measured on 19 May: capability arrives from outside, and judgment is supplied from inside.

Which leaves the question for the next meeting. What in this organisation is the equivalent of manual flight, and who still practises it?


Reading List


What We Are Watching Next

  • Whether METR's rogue-deployment assessment becomes a recurring exercise, and whether the strategic-judgment benchmarks are published in a form an institution outside a frontier lab can run against its own systems
  • Whether the Commission's targeted consultation on the Article 6 guidelines, which closes on 23 June, changes the Annex III examples for employment and essential-services use cases
  • Whether the federal banking agencies issue the signalled request for information on model risk management for generative and agentic AI, and what it asks about who reads the output
  • Whether another professional regulator in Europe follows the BSB in placing a competence duty on members who do not use the technology themselves

The next issue goes deeper into one of these. If you want a specific function or sector covered, reply to this email.

— Liga


TwinLadder Weekly is a weekly intelligence report on judgment, governance, and the boards accountable for both. Subscribe at twinladder.ai/newsletter. Forward this issue freely.