TwinLadder Weekly
Issue #34 — The Rulebook Stops Where the Models Begin
2 May 2026 · Weekly intelligence on judgment, governance, and the boards accountable for both
Editor's Note
From Alex —
On 1 May the Federal Reserve published remarks given four days earlier by Michelle W. Bowman, its Vice Chair for Supervision. In them she confirmed that the revised interagency guidance on model risk management leaves generative and agentic AI outside its formal scope. Fast-moving technologies may need a different supervisory approach, she said, and supervisory guidance "should not be a barrier for banks to engage with new and evolving tools and technologies."
Banking has governed model risk since 2011. It is among the most mature control disciplines any industry has built. And the models now drafting the credit memo, the client letter and the board pack sit outside it, by design, with the supervisor saying so out loud.
A word about a change here. For two years this newsletter tracked AI competence, governance and the workforce. From this issue the subject is judgment — where it sits inside an organisation, how it forms, what drains it, and why the board answers for it. The evidence base stays the same. The question we put to it is different.
The week supplies the case. A rulebook built for models that calculate has been asked to govern models that write, and the supervisor has declined to pretend it reaches. The European board that would have to close that gap in its own house mostly has no committee in which to do it: 10% of STOXX Europe 600 boards run a formal innovation-and-technology committee, against 93% of the S&P 100.
Liga has the analysis.
— Alex
What the Federal Reserve Confirmed
Bowman's remarks were delivered on 27 April and published by the Board on 1 May. Their subject was AI in the financial system, and the operative passage concerns scope. The revised interagency model risk management guidance — the successor to the 2011 framework, issued by the federal banking agencies in April — deliberately holds generative and agentic AI outside its formal reach. The stated reasoning is that novel and fast-moving technologies may warrant a different supervisory approach.
The practical consequence is specific. A bank that puts an agent into a workflow can no longer point to model-validation practice as the controlling standard and rest there. It has to construct a control framework of its own and then defend it to an examiner who has been told, in advance, that the existing guidance does not settle the question.
That is a transfer of authorship. When the supervisor writes the standard, the board's task is conformity. When the supervisor stands back, the board's task is design — and design is a judgment, made by named people, that someone will later test.
The UK arrived at the same place by a different route. The FCA's AI Lab opened a second cohort of AI Live Testing in late April, with applications closed on 24 March and successful firms notified mid-month. The regulator describes it as "a safe place to test AI systems in real-world conditions, with appropriate regulatory support and oversight." The FCA, the PRA and the Bank of England have been consistent that AI will be supervised through existing frameworks. So the evidence a UK firm can put in front of a supervisor is operational: what it tested, what it found, what it changed. Two supervisors, two methods, one implication. The institution supplies the standard.
What the 2011 Rules Assumed
On 4 April 2011 the Federal Reserve and the Office of the Comptroller of the Currency issued joint supervisory guidance on model risk management — SR Letter 11-7 at the Fed, Bulletin 2011-12 at the OCC. It named the hazard plainly: the "possible adverse consequences (including financial loss) of decisions based on models that are incorrect or misused."
Fifteen years later that sentence reads like a document about AI. A regulator, in 2011, locating the danger in the unexamined trust an institution places in a model.
What the regulators then demanded was wiring. Three pieces of it carried the weight.
Inventory. Every model in the firm on one list, so that everything deciding things inside the house appeared somewhere a director could see it.
Effective challenge. Every model examined by people independent of those who built it and those who benefit from its use.
An unbroken line. Accountability running from the desk that uses the model, through the executives who manage it, to the board — which the guidance named as ultimately responsible for the framework.
Model risk stopped being one specialist's hobby and became a line item in many people's jobs. That is the general mechanism. A class of risk becomes real inside a company on the day it is inventoried, owned, challenged and reported on a clock. Cyber made the same climb, acquiring named officers, board-oversight disclosure and, in Europe, statute. Operational resilience made it after that.
Now hold each of the three against a model that writes.
The inventory assumed a countable thing
A model in 2011 was an artefact with a home. It had an owner, a version, a documented purpose and a place in a system. Large banks run inventories in the thousands, and the count is meaningful because each entry has edges.
A general-purpose model has different edges inside the firm. It arrives in a browser tab, an email client, a code editor, a slide deck. The same weights draft a credit narrative on Monday and a supplier letter on Tuesday. The unit a 2011 inventory counts — one model, one purpose, one owner — dissolves into a population of uses. A use has no version number and no owner unless somebody assigns one.
Validation assumed an outcome
A probability-of-default model makes a claim the world eventually settles. The validator pulls five years of predictions against five years of what actually happened and back-tests one against the other. Disagreement is measurable. Drift is measurable. The discipline works because reality returns a score, on a schedule, in a form a committee can read.
A model that writes produces text: a memo, a clause, a summary of a counterparty's position, a paragraph of reasoning about a statute. Ground truth exists in there somewhere. It arrives late, in fragments, and only where somebody goes and looks. Back-testing settles a number. Prose is settled by a reader who can tell — and the reader is the control.
Effective challenge assumed a challenger who could still do the work
The 2011 principle put a second, independent set of eyes on the model, and it assumed those eyes could see. For a quantitative model that assumption held cheaply, because the challenger's skill was maintained by the work itself. Validators build competing models. Building keeps the eye sharp.
For a model that writes, the equivalent skill is the capacity to read a fluent, correctly formatted, plausibly reasoned document and find the place where it is wrong. That capacity is built by doing the work the model now does. Which is where the model-risk framework and the judgment question meet: independent challenge is a person with retained capability, and retained capability has a maintenance schedule whether or not anyone has written one.
What it looks like when the reader is downstream
On 27 April — four days before Bowman's remarks were published — the US District Court for the Eastern District of North Carolina publicly reprimanded an attorney over AI-fabricated material in filings, on docket 2:25-CV-00041-M in Fivehouse v. U.S. Department of Defense. Norton Rose Fulbright's survey of 2026 generative-AI sanctions records that the court found the attorney's initial responses "not credible" and determined that he "knowingly and intentionally submitted fabricated materials." A public reprimand sits a step above the admonition imposed in comparable cases, and the court got there because his account of how the fabrications entered the filing did not survive scrutiny.
The control sat downstream. The institution that caught the fabrication was the court, after the document had been filed and relied upon. That is the recurring shape for a model that writes. The check happens in the forum, in the audit, in the regulator's file — somewhere outside the firm, and later than the firm would choose.
The Committee That Would Have to Close It
On 30 April, ecoDa published the second edition of its European Corporate Governance Barometer with Ethics & Boards, drawn from index composition and disclosure for the STOXX Europe 600, Eurozone 300, FTSE 100 and S&P 100 across 2018–2026.
The technology findings are these. Only 10% of STOXX Europe 600 boards have a formal innovation-and-technology committee. A further 24% handle the topic through another committee or at the full board — 34% combined. The S&P 100 figure is 93%, of which 76% runs through audit and risk and 17% through a dedicated committee. "Digital" experience is held by an average of 16.8% of European board members and is present on 64% of boards. The average age of European independent directors has risen from 60.2 in 2018 to 61.9 in 2026, and average tenure runs around six years against eight-plus in the US.
One caution on the figures. The report's committee-oversight charts carry the source note "companies' data disclosed in 2024 annual report", while the key-findings page presents the 10% as a 2026 figure. The skills data is likewise drawn from board information in 2024 annual reports. Treat the direction as firm and the vintage as approximate.
Look at how the American number is composed. The S&P 100's 93% is overwhelmingly audit-and-risk coverage, and only 17% of it is a dedicated committee. So the US answer to "where does technology sit" is mostly an existing committee whose charter was amended to say so. That is a cheap move. It costs a charter revision and a calendar slot. One large European board in three covers technology in a committee of any kind.
The pressure from shareholders reads the same way. At IBM's annual meeting on 28 April, a stockholder proposal asking the board to "issue a report within one year on methods used to eliminate bias from IBM's AI models" drew 2.4% support, on Boardroom Alpha's proxy tracking. It was the first AI proposal of the 2026 season to reach a vote. The proposal asked for a report. Roughly one vote in forty was willing to ask for it.
Put the three together. The supervisor has stepped back from writing the standard. The board mostly has no chartered place to write one. And the shareholder base is registering something close to indifference on the record. Whatever governs a model that writes inside a European listed company this year, it will not arrive from outside the building.
A Reading Taken Once Is a Diagnosis
The instinctive response to a gap like this is to commission a score. An AI readiness assessment, a maturity model, a rating against a framework, delivered as a slide with a number on it.
A score is useful. It is also a diagnosis, and diagnoses age. What converts a reading into governance is the arrangement around it: a date on which it must be taken again, a room in which it is taken, and a person in that room entitled to ask why the number moved. Without those three, a maturity assessment records the state of an organisation on a Tuesday in May and quietly stops describing it by August.
Europe already has an instrument that demonstrates the arrangement, and most boards reading this are already inside its scope. Regulation (EU) 2022/2554 — DORA — has applied since 17 January 2025. Its Article 5 puts the management body at the top without ambiguity: the board "shall define, approve, oversee and be responsible for the implementation" of the ICT risk management framework and shall "bear the ultimate responsibility" for managing the entity's ICT risk. Then Article 5(4) goes somewhere unusual. Members of the management body "shall actively keep up to date with sufficient knowledge and skills… including by following specific training on a regular basis."
A financial regulator, writing rules about machine resilience, finished by scheduling the board's own capability maintenance. Knowledge kept current, by training, on a clock, as a personal legal obligation of each member of the management body. The object of that regulation is the technology estate. The mechanism — a named duty, a recurring obligation, a record — is portable to any risk a board decides to run it against.
The gap Bowman described closes inside the building, this year, or it stays open. It closes with an instrument the board builds and keeps: a place where the question is asked, a date on which it is asked again, and a name against the answer.
What This Means for Boards Right Now
One. The instrument question comes before the policy question. Name which committee's charter carries AI, and the date it was last amended to say so. If AI reaches the board through a management deck that appears when management chooses to bring it, the board has a topic. A topic has no clock, no owner and no record. The S&P 100 route — amend the audit-and-risk charter — costs a charter revision and a calendar slot, and it accounts for 76% of technology oversight in that index.
Two. Validation for a model that writes is a reading discipline, and reading has to be staffed. Back-testing settles a number, and the discipline that settles a paragraph is a competent human who reads it before it leaves the building. For each material use, three answers belong on the page: who reads, what they check, and what happens when they find something. Any of the three left blank moves the control downstream, to the court or the examiner — which is where the Fivehouse reprimand was issued.
Three. Commission the clock before you commission the score. A readiness rating with no scheduled re-reading, no owning committee and no independent challenger is a diagnosis in a drawer. Fix the date, the room and the challenger first; the score is then worth taking, because someone is entitled to ask why it moved.
Which leaves one question for the next meeting, and it is the question this week put on the table: which instrument in this house actually governs a model that writes?
Reading List
-
The longer treatment of what happens when a control document predates the technology it now has to govern, and how to tell which of your procedures are in that position: Your Compliance SOPs Were Written for a World Without AI
-
A banking case study in what it costs when the people who run a control function are thinned out, and what the regulator did about it: ING Was Fined 775M for Understaffing Compliance. Now They're Cutting 1,250 Jobs.
-
The practical build: how to assemble an AI governance framework from the instruments an institution already holds, and where the charter language sits: Building an AI Governance Framework: From US Rules to Article 4 Readiness
What We Are Watching Next
- Whether the federal banking agencies open a formal consultation on model risk management for generative and agentic systems, and what scope they propose for it
- Whether any STOXX Europe 600 issuer amends a committee charter this AGM season to name AI, and which committee takes it
- Whether the FCA's second AI Live Testing cohort produces published findings that name control expectations firms can plan against
- Whether any AI-oversight shareholder proposal in the 2026 season clears double figures, after IBM's 2.4%
The next issue goes deeper into one of these. If you want a specific function or sector covered, reply to this email.
— Liga
TwinLadder Weekly is a weekly intelligence report on judgment, governance, and the boards accountable for both. Subscribe at twinladder.ai/newsletter. Forward this issue freely.
