TWINLADDER
TwinLadder logoTwinLadder
Atpakaļ uz apskatu

Izdevums #40

The Vendor Chain Is Where Judgment Gets Delegated

In one week OWASP catalogued real CVEs across agentic AI and reported that a model cannot reliably separate instruction tokens from data tokens, Anthropic released Claude Fable 5 with a thirty-day data-retention floor that removed zero retention as a procurement option, and Waymo filed a recall covering 3,871 automated driving systems that could enter a closed freeway construction zone and keep going at speed. Each was settled in a contract long before an operator met the consequence. This issue reads the vendor chain as the place where a board's AI judgment is either exercised or handed away, sets the FTC's Illuminate Education order and the new generative-AI insurance exclusions beside it, and asks which obligations have been contracted away, to whom, and whether anyone can still see them.

third-party risk
procurement
agentic AI security
board oversight
AI Act transparency
2026. gada 13. jūnijs13 min read
The Vendor Chain Is Where Judgment Gets Delegated

TwinLadder Weekly

Issue #40 — The Vendor Chain Is Where Judgment Gets Delegated

13 June 2026 · Weekly intelligence on judgment, governance, and the boards accountable for both


Editor's Note

From Alex —

Three things happened this week, in three different industries. Each of them was settled in a contract long before anyone had to live with it.

OWASP published its 2026 report on agentic AI security, and the change from last year's edition is that it now catalogues real CVEs, vendor advisories and breach reports where it used to catalogue plausible threats. Anthropic released its new frontier model with a condition attached: thirty-day data retention, and zero retention withdrawn as an option. And an automated driving fleet was recalled — 3,871 systems — because the software could take a vehicle into a closed freeway construction zone and keep going at speed.

File those as a security story, a procurement story and an engineering story and they go to three different desks, which is how this class of risk usually gets filed. Read them together and they are one story: what an organisation agreed to when it bought something, and whether anybody senior enough can still see the agreement.

I have spent much of this year in rooms where AI arrives on the agenda as a technology item. It arrives more often as a contract item. The model's terms, the vendor's indemnity, the scope of what an agent is allowed to reach, the class of road a machine is permitted on — every one of those is a judgment somebody exercised on the organisation's behalf, at a desk, months before an operator met the consequence of it.

Liga has the analysis. The question I would take into your next meeting sits at the end of it, and it is short: which of our AI obligations have we contracted away, to whom, and can we still see them?

— Alex


Three Purchases, Read Together

The security report

On 11 June OWASP published State of Agentic AI Security and Governance, version 2.01. The 2025 edition described a threat landscape. This one documents an incident record: CVEs, vendor advisories and breach reports across nearly every category of agentic risk. It tracks 53 agentic projects, 28 of them coding agents. It maps 42 regulatory instruments across 10 jurisdictions. For the workflow-automation tool n8n alone it counts 57 security advisories. And 37% of the organisations surveyed reported having a policy to detect shadow AI, which leaves a majority of them with no method of finding the tools already in use.

The report's central technical finding concerns prompt injection:

"Large language models treat the system prompt, the user's request, and any text retrieved from external sources as a single stream of tokens. There is no reliable way to mark some of those tokens as commands and others as data."

That describes a property of the architecture, and architectural properties outlive patch cycles. This settles a question many procurement teams are still treating as open. If the boundary between instruction and data cannot be enforced inside the model, it has to be enforced outside it: by scoping what the agent can reach, and what it is permitted to do unaided. Both of those are settled in a contract and a configuration file, by people who are usually several levels below the board.

The model release

On 9 June, Anthropic released Claude Fable 5 and Claude Mythos 5. Fable 5 carries a one-million-token context window by default and always-on adaptive thinking, priced at $10 and $50 per million input and output tokens. Mythos 5 is the same underlying model with safeguards lifted, deployed through Project Glasswing and restricted to authorised cyber-defenders, infrastructure providers and selected biomedical researchers.

The release material carries one sentence that functions as a procurement instruction:

"Claude Fable 5 requires 30-day data retention and is not available under zero data retention."

For any firm whose data policy required zero retention on client files, patient material or privileged work product, the frontier capability and the policy stopped being compatible on 9 June. Nobody negotiated that. It arrived in a release note. A board that had been assured its AI arrangements ran under zero retention now has an assurance with a date on it, and the date has passed.

The recall

On 12 June Waymo filed a Part 573 safety recall report covering 3,871 units of its 5th Generation Automated Driving System, campaign number 26E035. The defect statement is plain:

"Waymo LLC (Waymo) is recalling certain 5th Generation Automated Driving Systems (ADS). The software may allow the vehicle to enter a closed freeway construction zone and continue driving at speed."

The consequence, in the report's own words: "Driving through a closed construction zone increases the risk of a crash." A remedy was still under development at the time of the report. The interim measure was to restrict freeway driving altogether.

Look at what the interim measure concedes. The operating envelope had been drawn wider than the demonstrated competence, and the correction was to pull the envelope back in until the competence caught up. That is a machine doing the wrong thing confidently, at fleet scale, with a documentary record — and the record exists because vehicle safety is one of the few domains where somebody is obliged to file one.


The Risk That Arrives in Fragments

This class of risk announces itself in pieces, and each piece looks like an operational nuisance belonging to somebody else's function. A security advisory. A change of terms. A software recall. The pieces never assemble themselves; someone has to assemble them.

Lay out the risks an institution takes on when it adopts AI at scale and the family has a known shape. There is the behaviour of the model: bias in the training data, drift as the world moves away from it, brittleness on inputs it never saw, the difficulty of explaining a decision. There is security: a new surface to attack, a new channel to leak through. There is the data: privacy, provenance, the rights attached to what the model was fed. There is dependency: the vendor whose model now sits inside your process, and what happens when that service degrades or its terms change. And there is the law, arriving to meet all of the above.

Every one of those families watches the machine. The dependency line is where this week's three events actually sit. In most houses it exists as a paragraph in a supplier questionnaire, and nowhere on a register with a name against it.

Frank Elderson put the general point to the Goldman Sachs European Financials Conference in Zurich on 3 June, speaking as ECB Executive Board member and Vice-Chair of the Supervisory Board: "the challenges posed by new generations of AI models should not be viewed solely as a cybersecurity issue – they are a firm-wide strategic challenge." He tied bank preparedness to governance arrangements and to awareness "particularly among banks' management bodies", and flagged concentration risk in cloud, telecoms and payments infrastructure. The supervisor is describing a risk that arrives through infrastructure the institution rents from somebody else.

There is a settled mechanism for making a risk of that kind real inside a company, and every board reading this has watched it operate at least twice. A class of risk becomes real on the day it is inventoried, owned, challenged and reported on a clock. Before that it is a topic, and a topic has no owner and no record. Model risk made that climb in banking in 2011. Cyber made it next, acquiring named officers, board-oversight disclosure and, in Europe, statute. Operational resilience made it after that, and Europe's own resilience regulation already obliges in-scope financial entities to govern their critical technology vendors, report serious incidents and test themselves under stress.

So the wiring exists. The question is whether AI dependency has been run through it, or whether it is still being handled as a topic — reaching the board when management chooses to bring it.

The warning that came from inside the chain

On 5 June the Federal Trade Commission gave final approval, on a 2-0 vote, to a modified order against Illuminate Education Inc. over a breach exposing the personal data of 10.1 million students, including addresses, dates of birth, student records and health-related information. One line of the complaint concerns the two years before the breach:

"The FTC's complaint further alleges that, despite being alerted almost two years before the breach by its third-party vendor about numerous security vulnerabilities on its network, Illuminate failed to take steps to adequately address the problems."

The warning came from inside the supply chain. It came in writing. It came two years early. It reached nobody with the authority or the appetite to act on it. The order now requires deletion of unnecessary personal data, a published retention schedule and a comprehensive information security programme — the wiring, installed by a regulator, after the loss.

Now reverse the seat. Illuminate was somebody's vendor. Ten million students' records sat inside a school district's arrangement with a supplier whose own supplier had already documented the problem. Somewhere in your chain a similar letter may have been written already, to somebody other than you, about a system you depend on.

Gartner's forecast for the year ahead describes the same geometry from the buying side: "By next year, 40% of enterprises will have their autonomous AI efforts in part derailed by gaps in governance discovered only after production incidents." Its account of the mechanism runs like this. Enterprises apply uniform governance across agents regardless of autonomy level and trust boundary, and get two symmetrical failures: over-restricted agents that deliver nothing, and over-trusted agents that act beyond their scope. Both are failures to distinguish what a system is able to do from what it has been given access to do — a delegation judgment, made before deployment, or made by default.


The Two Clocks a Contract Runs On

A board reads its risks on two clocks, and every procurement decision writes on both.

The fast clock carries what the arrangement can cost this quarter. A retention term that conflicts with a data policy. An injection surface reachable from a customer email. An agent holding a permission it never needed for the task it was given. These are present-tense exposures, and they belong to the chief risk officer's line, where they can be watched with an override rate, a sampled challenge, an alarm.

The slow clock carries something the same contract also decides and almost never minutes: what the arrangement does to the formation of the people who would catch the fast-clock failure. When a firm buys a system that produces the first draft, the fast question is whether the draft is good. The slow question is who, in five years, will be able to tell when it is wrong. The associate who would have written that draft was being made by writing it. The reviewer who catches a fabricated citation at eleven at night was built out of years of unglamorous files. Remove the reps and the capacity thins on a clock too slow for any quarterly dashboard to show the loss — and it thins invisibly, while every reported figure improves.

The two clocks explain what a vendor contract settles, and why so little of it gets examined. A contract sets the fast-clock exposure directly: what the system may reach, what it may do alone, who pays when it fails. It sets the slow-clock exposure indirectly, by deciding which steps of the work leave the building. The first is negotiated by people trained to negotiate it. The second goes through on the same signature, with nobody in the room whose job it is to ask the question.

The indemnity that may have no money behind it

There is a third thing a contract decides, and this year it moved. Analysis published in late May by the law firm Honigman describes what it calls the AI insurance gap. The Insurance Services Office has introduced optional generative-AI exclusions for 2026 commercial general liability policies, several major carriers have adopted them or similar language, and parallel exclusions are appearing in directors' and officers' policies. The consequence is stated in one sentence:

"When a vendor agrees to indemnify, it may have no insurance to fund that obligation."

Read that beside the indemnity clause in your own AI supplier agreements. A board that treats the clause as its protection is relying on a promise whose funding it has never checked, from a counterparty whose own carrier may have excluded exactly this loss — while the board's own policy may exclude it too. Where all three hold, the residual exposure sits on the company's balance sheet. It can sit there priced and accepted by name. It is sitting there either way.


Three European Instruments Reached Down the Chain This Week

The same week produced three of them, and each makes a company answerable for something a supplier does.

Marking and labelling. On 10 June the Commission published the final Code of Practice supporting Articles 50(2) and 50(4) of the AI Act, ahead of the transparency obligations that apply from 2 August 2026. The Code is voluntary and layered: machine-readable metadata plus watermarking, with fingerprinting and logging as supporting measures. Section 1 addresses providers of generative AI. Section 2 addresses deployers — and deepfakes and AI-generated or AI-manipulated text published on matters of public interest must be clearly labelled. The deployer half of that Code rarely reaches a board. If an agency, a communications team or a customer-service function publishes AI-assisted text under the company's name on a matter of public interest, the company is the deployer and the labelling duty is the company's.

A German enforcer with an address. On 11 June the Bundestag passed the KI-MIG, the AI market-surveillance and innovation-promotion act. It designates the Bundesnetzagentur as Germany's central market-surveillance authority, notifying authority and single point of contact and complaints office. It assigns financial-sector AI supervision to BaFin, leaves data protection with the BfDI and AI security with the BSI. As the analysis of the bill puts it: "The AI Act supplies the principal substantive obligations, while the KI-MIG assigns authorities, enforcement powers and German administrative-offence rules." For a buyer, the practical change is that complaints about a supplier's system now have a named destination and a body with powers behind it.

Due diligence down the chain. On 12 June the Commission opened a public consultation on the guidelines that will support implementation of the Corporate Sustainability Due Diligence Directive, closing on 24 July. The guidelines will cover risk identification and prioritisation, appropriate measures, responsible disengagement, stakeholder engagement, data sources, digital tools, model contractual clauses and risk-factor assessment. Adoption is planned for the first quarter of 2027, ahead of a statutory deadline of 26 July 2027. The directive itself "will require in-scope companies to take steps to identify, prevent, mitigate, and bring to an end adverse impacts relating to human rights and the environment."

Note the phrase model contractual clauses. The language that in-scope companies will push down their supply chains is being drafted now, in a consultation that closes in six weeks. A board that waits for transposition will inherit clauses it had the opportunity to comment on and did not.


What This Means for Boards Right Now

One. Inventory the terms, not only the tools. Most AI inventories list systems and owners. The obligation lives one level down, in the clauses: the retention term, the indemnity, whether an insurer would fund that indemnity, what the system may reach, and what it may do without a human in the path. Ask for one sheet carrying those five columns for every material AI arrangement in the house. The Fable 5 retention floor shows how quickly a column changes value — it changed in a release note, on a Tuesday in June.

Two. Enforce outside the model what cannot be enforced inside it. OWASP's finding is that the model itself cannot separate an instruction from a piece of data. So the usable controls are all perimeter controls: the scope of what an agent can reach, the actions it may take unaided, the threshold at which it stops, and the ability to roll back what it did. Gartner's forecast prices the cost of skipping that work: governance gaps found only after the production incident. Waymo's interim measure is the same move made correctly: narrow the envelope to the demonstrated competence, and record who narrowed it and on what evidence.

Three. Put the slow clock on the procurement paper. Every AI purchase is minuted against the fast clock — cost, capability, security, term. Add one question to the approval template and make it answerable by the function owner who signs: which steps of this work are we handing over, and if the machine does every instance of those steps, does anyone here still learn to catch it when the machine is wrong? Where the answer is nobody, the arrangement is buying this year's margin with next decade's bench, and the board should at least know it is making that trade.

Which leaves the question to put on the agenda for the next meeting: which of our AI obligations have we contracted away, to whom, and can we still see them?


Reading List


What We Are Watching Next

  • Whether any frontier vendor restores a zero-retention tier for its most capable model, or whether a retention floor becomes the standard condition of frontier capability
  • Whether OWASP's CVE catalogue is converted into control language a buyer can put into a contract schedule, by an industry body or by a supervisor
  • Whether the KI-MIG completes its passage into force, and what the Bundesnetzagentur publishes about how it intends to supervise
  • Whether the CSDDD implementation guidelines, open to comment until 24 July, carry model contractual clauses that reach AI systems sitting inside a supply chain
  • Whether the automated-driving recall produces a remedy, and what the operator does about freeway driving until it has one

The next issue goes deeper into one of these. If you want a specific function or sector covered, reply to this email.

— Liga


TwinLadder Weekly is a weekly intelligence report on judgment, governance, and the boards accountable for both. Subscribe at twinladder.ai/newsletter. Forward this issue freely.