TWINLADDER
TwinLadder logoTwinLadder
Back to Newsletter

Issue #42

The Payroll Data and the Self-Report Disagree

On 26 June Anthropic's Economic Index reported that the people who delegate the most work to the model say they are learning at the same rate as everyone else — the strongest available counter-evidence to the deskilling argument, taken with a survey. The next day Stanford's Digital Economy Lab and ADP opened a dashboard built on payroll records covering 4.6 million workers and more than 730 occupations, showing 22- to 25-year-olds in the most AI-exposed jobs contracting at 3.8% a year while the same cohort grows 2% in the least-exposed jobs. This issue puts the counter-argument at full strength, sets it beside what employers actually did, and asks which AI benefit claims in a board pack would survive being checked against an administrative record.

measurement and evidence
entry-level hiring
deskilling
board oversight
executive incentives
June 27, 202614 min read
The Payroll Data and the Self-Report Disagree

Listen to this article

0:000:00

TwinLadder Weekly

Issue #42 — The Payroll Data and the Self-Report Disagree

27 June 2026 · Weekly intelligence on judgment, governance, and the boards accountable for both


Editor's Note

From Alex —

On 26 June, Anthropic published the June edition of its Economic Index. One finding in it cuts against the argument this newsletter has been making for two years. The people who delegate the most work to the model report learning at the same rate as everyone else.

I want to spend this issue on that finding. And I want the first half of it to make the case for the finding, at full strength, in the terms a reasonable person who believes it would use.

A newsletter that only meets the evidence agreeing with it is a newsletter you can stop reading. We have argued, repeatedly, that handing the formative work to a machine drains the capacity to judge the machine's output. If that argument is right, the heaviest delegators are exactly where you would see it first. Anthropic looked there and reported no difference.

The morning after, on 27 June, Erik Brynjolfsson's group at Stanford and ADP Research opened a public dashboard built on payroll records covering 4.6 million workers across more than 730 occupations. Among 22- to 25-year-olds, employment in the most AI-exposed occupations is contracting at 3.8% a year, and the rate has risen. The same age band, in the least-exposed occupations, is growing at 2%.

Two readings, one day apart. One asks people how they are getting on. The other counts what employers did about them. I have sat through several technology transitions in markets where the data was thin, and the rule I came out with is this: when a survey and an administrative record disagree, the administrative record is the one that turns up in the hearing.

The two are also measuring different things, and Liga sets out where they can both be right. She then puts the board question that follows, which is the question I would take into my next audit-and-risk meeting.

— Alex


The Case Against This Newsletter

Anthropic's June report pairs usage telemetry from Claude with a survey of the people producing it. The headline numbers are favourable to AI adoption and clear about it. 68% of respondents report learning more with AI. 57% say AI has made their skills more valuable. Over 35% expect AI to handle most of their work tasks within twelve months.

Then the finding that matters here. Heavier delegators report learning at the same rate as everyone else. And the heaviest delegators are the most optimistic about their own labour-market outcomes: "people who delegate to Claude the most are the most optimistic about their future labor market outcomes."

The argument that follows from this is the best one available against everything written in this newsletter. Here it is, in the terms someone who holds it would use.

It is a prediction failing on the sample where it should be easiest to see. The deskilling thesis says that offloading the formative work thins the underlying capability. Run that forward and heavy delegators should report noticing something — more friction when the tool is away, more uncertainty about their own work, less sense of growth. Anthropic went to the population with the highest delegation rate in the economy and the reported learning curve was flat against everyone else's. A theory that survives every result is worth less than a theory that risked one, and this was a risk the theory took and lost.

The respondents are practitioners, not subjects. Most of the deskilling literature works by removing the tool and measuring what happens. That is a clean design and an artificial one, because in working life the tool stays. These respondents are describing their own practice as it actually runs, at a scale no laboratory reaches.

A skill displaced by a tool has often relocated instead of vanishing. This is the standing position behind the finding, and it is respectable. The claim is that what looks like erosion is substitution: the capability moves up a level, from producing the draft to directing and checking it, and the person doing the directing is learning something real. On this reading, "I am learning at the same rate" is accurate, and the thing being learned has changed.

People have some purchase on their own competence. If capability were draining at scale, some sizeable fraction would feel it. Dismissing millions of self-assessments as false consciousness is a heavy move, and anyone making it owes an account of why the self-assessment fails in this particular case and holds in others.

That is the case at full strength. Two things belong alongside it.

The first is method. The learning measure is self-reported. Respondents were asked how much they are learning, and their answers were recorded. That is a legitimate instrument, widely used, and what it captures is a perception of learning. What a person can still do with the tool switched off is a second reading, and it takes a second instrument.

The second is that the same report carries findings pointing the other way. Early-career workers report that AI can do the highest share of their work, and they express the most concern about job loss. More than a third of respondents believe junior colleagues face a greater than 60% probability of losing their job within a year. So the document reports in both directions. It contains a workforce saying it is learning fine, and saying the rung below it is in trouble.


What Employers Did

The Canaries Dashboard opened on 27 June, built by the Stanford Digital Economy Lab with ADP Research and updated continuously. Its base is payroll: the records generated as a by-product of paying people, for 4.6 million workers across more than 730 occupations.

For workers aged 22 to 25, employment in highly AI-exposed occupations is now shrinking at 3.8% per year. In April 2024 that figure was 2.8%. In the least-exposed occupations, the same age band is growing at 2% a year. Workers aged 31 to 34 contracted 1.7% year on year.

Payroll is a different kind of instrument. Nobody was asked a question. There is no response rate, no framing effect, no gap between what a person believes about their week and what their week contained. A payroll record exists because money moved. It is the same class of evidence a court asks for, and the same class an examiner asks for.

Payroll has its own limits. "AI exposure" is a constructed measure, assigned to occupations by researchers, and reasonable people build it differently. Payroll counts jobs, so it is silent on whether the people still in those jobs can do anything they could do three years ago. And it cannot say why a firm hired fewer 23-year-olds — interest rates, a hiring correction, a wage effect and an AI decision all leave the same footprint in a headcount series. This is a signal about employment, with the cause argued rather than demonstrated.

Two events in the same week narrow the cause a little, from opposite directions.

On 23 June, Oracle's annual report was reported as stating that the "adoption and deployment of AI technologies across our operations have resulted, and may continue to result, in reductions to our workforce." Headcount fell from roughly 162,000 to about 141,000, with $1.84bn of severance and related charges. Set aside the number and look at where the sentence sits. It sits in a statutory filing. A filing carries legal exposure, so the language was chosen by people who expected to be held to it. A company put the causal claim on the record itself.

Also on 23 June, AWS chief executive Matt Garman told Casey Newton the opposite: "This is the reason we're hiring 11,000 interns and new college grads this year at Amazon. They come in with an energy and excitement, a new view on things." Fortune carried his reasoning the next day. He expects roughly half of white-collar jobs to change and rejects elimination — "I do think that half of white-collar jobs may change, but wipe out and change are different" — and says Amazon currently employs more software developers than it did two years ago. His hiring criterion is capacity to learn: "They haven't learned bad habits, you can teach them the culture, they're willing to learn the new tools."

Garman is the counter-example, on the record, at scale. He is also buying the input to an apprenticeship that somebody at Amazon now has to run. Hiring for capacity to learn pays off only where the organisation still contains the situations that teach.


Two Instruments, Two Questions

The disagreement resolves once you notice that the two readings answer different questions.

Anthropic asked working adults, most of them already in a job, how their own learning feels. Stanford and ADP counted whether 23-year-olds were hired. A person can be learning a great deal in a seat while the number of such seats falls. Both findings can be true at once, in the same economy, in the same month, about different people.

The population reporting flat learning is the population that already got in. The population the payroll series is counting is the one at the door. If the mechanism worth worrying about operates by closing the entry rung, then a survey of incumbents is the last place it would appear — and the flat learning curve is exactly what the mechanism predicts.

The tuition was inside the chore

The work now being handed to the model is the first-draft, first-pass, low-status work: the data pull, the reconciliation, the summary of a counterparty's position, the diligence read. That work was never only output. It was the route by which a graduate acquired the sense that a number is wrong before the evidence is in. The formation happened inside the chore, which is why removing the chore removes the formation, and why nobody writing the business case noticed. Drudgery was carrying the apprenticeship on its back, and no cost model has a line for it.

Notice how the loss shows up in the accounts. A firm that slows graduate hiring records a saving. The junior who was never hired leaves no empty seat on the org chart, because the seat was never drawn. The gain is owned, immediate, and sits on a line someone is measured against. The loss is unowned, deferred, and sits on no line at all until the day it sits on all of them.

Every number improves while the capability erodes

There is a version of an institution in which throughput rises, cost falls, cycle time shortens, and every published number moves the right way — while the population trained to know when those numbers are lying gets thinner. The two run together, in the same building, at the same time, and the institution's own instruments record only the first one.

This is the condition the disagreement describes. The improving numbers are collected continuously, automatically, and reported to the board on a schedule. The eroding capability sits on no instrument, has no owner, and appears on no dashboard. So the board receives a monthly reading of one and never a reading of the other, and reasonably concludes that only the first one is happening.

The correction is a measurement decision, and it can be made this quarter. A firm already holds an administrative record of its own entry cohort: hires by age band and function, offers, retention at twelve months, the internal work now performed by a model that a first-year performed in 2023. None of that requires a survey, a vendor or a study. It requires somebody to be asked for it by a date.


The Room That Has to Decide

Whoever adjudicates between a favourable survey and an unfavourable payroll series does it in a boardroom, so it is worth asking who is in the room.

On 25 June, Deloitte published an analysis of the last six leadership roles held by each Fortune 100 director. CEO experience appears on 100% of boards. Finance executives on 96. Operations on 89. Marketing on 67. Senior HR on 56 or more. Nonprofit, military or government on 65. The study publishes no equivalent count for technology or AI leadership experience. That is an absence in the study rather than proof of an absence in the rooms. It is worth sitting with, because the study is framed as a baseline for board-refreshment discussions. Deloitte's own framing: "Having broader functional experience in the boardroom could help enterprises adapt more effectively to shifting market conditions."

The duty is arriving anyway. On 24 June, FCA chief executive Nikhil Rathi told a techUK conference on agentic AI in financial services that accountability for regulated activities and outcomes must remain clear, with the right human oversight, and put the obligation on named people: "Boards and leadership teams must understand the risks. Dependencies—particularly on model providers and third parties—must be properly mapped and governed."

A day later, a guest post on The D&O Diary made the argument in fiduciary terms, under the title "AI Governance Is a Fiduciary Duty." Its point about examination is the operative one for anyone reading this in a regulated firm: "The SEC has identified AI-based systems as a 2026 examination priority for registered advisers and broker-dealers, with examiners directed to assess whether automated tools operate consistently with regulatory expectations, specifically whether governance and validation structures exist." That is commentary from a single source, and the site cannot be read by automated retrieval, so treat it as a considered legal reading and confirm it against the SEC's own priorities before quoting it internally.

The measurement question has travelled less far than the duty. Pearl Meyer examined roughly 2,500 public-company proxy statements filed in 2026. About 2% include a formal AI metric in executive incentive programmes. Of those, only 12% use an explicit AI metric; 60% embed AI inside a broader technology or transformation objective; 28% handle it through individual performance assessment. And 77% of companies report they have not yet scaled AI enterprise-wide. Pearl Meyer's conclusion is that the first question is measurement, not plan design: the issue is "not simply whether to incorporate AI into incentive plans, but whether organizations are prepared to measure and govern it in a way that supports sound pay decisions."

Read those four together. The claimed benefits of AI are being reported to boards continuously. Roughly one company in fifty has attached a formal metric to any of them. And the duty to understand what is being reported is being stated aloud by a regulator and argued as fiduciary by counsel, in the same week.


What This Means for Boards Right Now

One. Put the instrument beside every AI benefit claim in the next management pack. For each claimed gain, write down where the number came from: a user survey, a self-assessment, tool telemetry, an administrative record, or an observed test with the tool switched off. Those five carry different weights and are routinely reported in one column, in the same typeface, as though they were the same kind of evidence. Anthropic's flat learning curve is a real finding taken with a real instrument, and the instrument is a survey. Stanford and ADP's 3.8% is a real finding taken with a different one. A pack that does not distinguish them is asking the board to average a perception with a payroll record.

Two. Decide what would falsify the claim before the claim is made. This is the same discipline a credit committee applies to a model, and it costs a sentence per claim. If the productivity case for a deployment is right, some administrative series in your own house moves — cycle time, error rate on a sampled review, headcount by cohort, external spend on the work that was insourced. Name the series and the date it gets read. A benefit claim with no series attached is a perception, and perceptions are what the survey already measured.

Three. Run the entry-cohort series whether or not you believe the mechanism. Payroll data covering 4.6 million workers shows 22- to 25-year-olds contracting at 3.8% a year in the most AI-exposed occupations and growing at 2% in the least exposed. Amazon is hiring 11,000 interns and graduates on the opposite bet. Both are defensible strategies, and the board that can defend either is the one that knows its own numbers: hires by age band, by function, year on year, and what the first-year work now consists of. A board that cannot produce that series is choosing between the two bets without knowing which one it already made.

Which leaves the question this week put on the table, and it is a short one for the next meeting: which of our AI benefit claims would survive being checked against payroll rather than survey?


Reading List


What We Are Watching Next

  • Whether the Canaries Dashboard's next update shows the 22–25 contraction rate continuing past 3.8%, and whether the 31–34 band widens from 1.7%
  • Whether any large employer publishes its AI benefit claims alongside an administrative series that would test them
  • Whether Amazon follows its 11,000-strong intake with a published account of how those graduates are actually trained, and on what work
  • Whether any issuer attaches an explicit AI metric to an executive incentive plan for the next proxy season, after Pearl Meyer's 2%
  • Whether the SEC's 2026 examination priority for AI-based systems produces a published finding that names what a governance and validation structure has to contain

The next issue goes deeper into one of these. If you want a specific function or sector covered, reply to this email.

— Liga


TwinLadder Weekly is a weekly intelligence report on judgment, governance, and the boards accountable for both. Subscribe at twinladder.ai/newsletter. Forward this issue freely.