Interactive explainer · AI & Work

After Automation

Dan Shipper's argument for why the more we automate, the more expert human work there is to do — and why benchmark progress doesn't change that.

Source: "After Automation" by Dan Shipper · Every · May 21, 2026

01 The paradox

Shipper opens with a contradiction from inside his own company:

"There is a paradox at the heart of AI. At Every, we've automated everything we can… And yet it seems like, for us, there's more human work to do than ever." — Dan Shipper, "After Automation"

Every runs on AI as aggressively as any company can: coworker agents in Slack, an AI handling most of customer support, Claude Code writing the software. The standard prediction says headcount should be collapsing. Instead, the work changed shape — and the new shape needs humans at least as much as the old one did.

The rest of the essay is Shipper explaining why that isn't a temporary lag before the machines finish the job. He argues it's a structural loop that regenerates expert human work every time automation succeeds.

The one big idea

Automation doesn't subtract human work from a fixed pile. It commodifies yesterday's expertise, which floods the world with sameness, which creates demand for difference — and producing difference is expert human work. The loop never terminates.

02 Two modes of working with AI

First, the ground truth from inside Every. Shipper sorts their AI usage into two modes:

1. Agent employees — async delegation. Coworker agents like Claudie (writes sales proposals, drafts training decks) and Andy (collects story "nuggets" into digests) get tagged in Slack like colleagues. Embedded agents like Fin run inside a product surface: in one week in May, Fin participated in 65% of 202 support conversations and closed 81 of them without a human.

2. Human–agent collaboration — shared operating systems like Claude Code and Codex, where human and AI work the same problem simultaneously. Here Shipper names the shape that recurs all through the essay:

"The human sandwich: a human sets the frame, AI collapses the task, and a human judges and extends." — the essay's core working pattern

// the human sandwich

HUMAN sets the frame AI collapses the task HUMAN judges & extends

And both modes carry a cost that rarely makes the demo reel — every agent needs humans to point it at the right problems, evaluate its output, catch its errors, and turn results into real decisions. Plus maintenance:

Gotcha · automation is not free

"One of our PowerPoint automations includes 24 skills and 18 scripts and costs $62 in tokens to make a single deck." Even trivial-sounding automations become systems someone has to own.

03 The flywheel: how automation manufactures new expert work

This is the engine of the essay. Shipper traces a four-stage loop. Click each stage — or press ▶ run the loop — and notice that stage 4 feeds straight back into stage 1.

// the after-automation flywheel — click a stage

STAGE 1 AI makes yesterday's human competence cheap STAGE 2 Cheap competence gets rapidly adopted STAGE 3 Abundance creates sameness — slop STAGE 4 Sameness creates a demand for difference
Analogy

Think of it like fashion. The moment a once-exclusive look can be mass-produced, everyone wears it — and the moment everyone wears it, it stops signaling anything, and the premium shifts to whatever can't be mass-produced yet. Expertise behaves the same way: slopShipper's term: "Slop is visible sameness, repeated ad nauseam" — the commodity-grade default output of models trained on the same data. is last season's expertise worn by everyone at once.

Now operate the loop yourself. Drag the automation level and watch what happens to the total amount of expert human work:

// simulate: what happens to expert work as automation rises

15%
cost of yesterday's expertise
sameness (slop)
demand for difference
executing yesterday's expertise framing, judging, differentiating (new expert work)
Gotcha · trigger it yourself

Drag automation to 100%. The bar refuses to shrink — it recomposes. The blue work (doing what experts used to do) collapses, but the purple work (deciding what's worth doing, judging output, producing difference) more than replaces it. That recomposition is the essay's whole claim, in one bar.

04 "But the benchmarks are exponential!" — Zeno's paradox of AI

The obvious objection: model capability charts go up and to the right. Shipper doesn't dispute the numbers — he cites them. Humanity's Last ExamA deliberately brutal academic benchmark. The essay notes models went from single-digit scores to 44%. climbed from single digits to 44%. GDPvalA benchmark of economically valuable work tasks cited in the essay; frontier models jumped to roughly 85% performance. hit ~85%. METRAn AI evaluation org measuring how long a task a model can complete autonomously; the essay cites Claude Mythos at an 80% success rate on 4-hour expert tasks. measured an 80% success rate on 4-hour expert tasks.

His counter: benchmarks measure performance inside a chosen frame — and the moment a benchmark saturates, the frame shifts. Step through what that looks like:

// the frame keeps moving — step through it

model capability

Why does the frame always move? Because a benchmark is, by definition, a problem that has already been written down. And Shipper's sharpest line is about exactly that:

"Once a situation has been reduced to text, once it has become corpus, it is a corpse." — the essay's central epigram

Models train on the corpusThe recorded text of past human competence — the only thing a model can learn from. Shipper's point: by the time something is corpus, the live problem has moved on. — the frozen record of problems humans already solved and framed. Humans are alive to the current situation before it has been reduced to text. The framer is not the same kind of thing as the frame, so measuring the frame ever-more-impressively never closes the gap.

Key insight

The gap between AI and humans isn't a fixed distance being crossed — it's regenerated by the act of crossing it. Every saturated benchmark forces someone to ask "okay, what actually matters next?" — and answering that question is human expert work that didn't exist before.

05 Smuggled intelligence

There's a second reason the charts mislead. When a model "achieves" something — say, a security firm using Anthropic's Mythos to find a macOS kernel exploit in 5 days instead of weeks — the headline credits the model. Shipper says the achievement smuggles in an unmeasured ingredient:

"There is an immense amount of smuggled intelligence — the hidden layer of human judgment." — on what benchmarks don't count

Flip the toggle to see the same task two ways — as the benchmark scores it, and as it actually happened:

// same result, two accountings

what the benchmark sees what actually happened

Every purple box is human expert work that the score attributes to the model. And note: the work in those boxes — choosing the target, framing the prompt, judging the output, deciding what to do about it — is exactly the work the flywheel in section 03 keeps manufacturing more of.

06 The AGI caveat: autonomy is not agency

Final objection: "fine, but what about real AGI?" Shipper's answer is a distinction. Today's models have autonomy — they can execute long tasks independently. What they lack is agency:

"Agency is the ability to act independently; an agent is someone who acts on behalf of another." — the distinction the AGI debate skips

His illustration is a toddler. A toddler is hopeless at tasks — and yet: "The toddler has ends… he invents games constantly… he is alive inside a field of desire." A toddler wants things nobody prompted. A model, however capable, waits. It acts on behalf of someone — and that someone has to exist, care about something, and choose what matters now.

So even an AGI that optimizes brilliantly toward goals reproduces the loop one level up: humans still choose which goals matter now. The frame problem doesn't dissolve; it climbs.

// check your understanding

1. Fin participated in 65% of Every's support conversations and closed 81 without a human. According to the essay, what happened to the human support role?
2. Why, per Shipper, won't exponential benchmark progress close the human–AI gap?
Context · not from the essay

The autonomy/agency distinction Shipper draws has a long history in philosophy of action (intentions and desires vs. mere goal-directed behavior). The essay doesn't cite that literature — this note is added background, not his argument.

07 Where am I myself?

Shipper closes with a Hasidic parable, quoted here in full because the essay's ending depends on it:

"There was once a man who was very stupid. When he got up in the morning, it was so hard for him to find his clothes that at night he almost hesitated to go to bed for thinking of the trouble he would have on waking. One evening he finally made a great effort, took paper and pencil and as he undressed noted down exactly where he put everything he had on. The next morning, very well pleased with himself, he took the slip of paper in his hand and read: 'cap' — there it was, he set it on his head; 'pants' — there they lay, he got into them; and so it went until he was fully dressed. 'That's all very well, but now where am I myself?' he asked in great consternation."

The notes are the corpus. Following them flawlessly is automation. And the question at the end — where am I myself? — is the one thing the notes can never answer. That, for Shipper, is the work that's left after automation: not executing what's written down, but being the one who's alive to this morning, recognizing what matters in this moment.

Recap — the one sentence

The more we automate, the more expert human work there is to do — because automation commodifies yesterday's expertise into sameness, sameness creates demand for difference, and only someone alive to the present moment can frame what difference matters now.