Home
May 23, 2026
View All

What AI Did to Headcount: Dan Shipper's After Automation Report

The Short Version#

Dan Shipper's "After Automation" report gives us one of the clearest first-person accounts yet of what AI actually does to a small company's headcount and skill mix. Meanwhile, Andrej Karpathy joining Anthropic is the talent signal of the year, Cursor keeps shipping at a pace that's hard to keep up with, and OpenAI's Virgin Atlantic case study shows what "AI-accelerated shipping" looks like with real stakes.

Dan Shipper / Every — After Automation#

Source: https://every.to/p/after-automation Credibility: High (first-party account from Every's CEO about their own company's operations)

What happened: Dan Shipper published "After Automation," a report on what happened to Every's team structure as AI tools improved from GPT-3 to the present. The short version: the team went from 4 to 30 people. But the growth wasn't despite automation — it was partly because of it. AI made expert-level competence cheap enough that Every could produce more, which created demand for people who could apply judgment, framing, and differentiation that AI can't supply on its own.

Key patterns:

  • AI made "default outputs" cheap and commoditized — good enough writing, analysis, and research became table stakes
  • That commoditization created a new scarcity: people who can frame problems well, make editorial calls, and differentiate output
  • Every grew headcount because cheaper execution unlocked more ambition, not because they needed fewer people
  • The jobs that expanded were judgment-heavy; the ones that compressed were process-heavy
  • This is a live experiment with a company that builds AI tools and uses them internally, making it more credible than most case studies

Why it matters for PMs: This directly challenges the "AI reduces headcount" assumption that a lot of teams are operating under. What Shipper describes is closer to what economists call the "rebound effect" — efficiency gains increase output appetite rather than reducing input. For PMs thinking about team composition, the implication is that AI shifts which skills matter, not how many people you need. If you're evaluating your team's capabilities through that lens, the question isn't "what can we automate?" but "what judgment calls are we making that only humans can make — and do we have enough of those people?"

Critical questions:

  • Every is a media and software company with a strong editorial identity. Does this pattern hold for companies where output quality is harder to define?
  • The report says AI created demand for human judgment — but is that demand coming from users, or from Shipper's editorial philosophy? Would a growth-at-all-costs operator see the same outcome?
  • At what point does commoditized AI output become good enough that the premium for human judgment collapses too?
  • How do you actually hire for "framing" and "differentiation" as skills when most job descriptions and interviews aren't built to evaluate them?

Action you could take today: Audit one recurring deliverable on your team — a weekly update, a spec, a research synthesis — and ask whether AI is producing the "default output" version and a human is adding the frame. If the answer is no and a human is still doing the whole thing, that's a workflow ripe for restructuring.

Andrej Karpathy — Joins Anthropic#

Source: https://x.com/karpathy/status/2056753169888334312 Credibility: High (direct post from Karpathy)

What happened: Andrej Karpathy announced he's joining Anthropic. He framed it as returning to R&D at "the frontier of LLMs" during what he called an especially formative few years. He also said he intends to resume his education work in time — suggesting this isn't a full pivot away from that mission, just a pause.

Key details:

  • Karpathy left OpenAI in 2023, spent time on Eureka Labs (an AI education startup he founded), and is now moving to Anthropic
  • This is the most high-profile AI researcher hire in recent memory — Karpathy is the person who made neural nets legible to a generation of practitioners
  • Anthropic already pulled ahead of OpenAI in enterprise spending according to Ramp's May 2026 AI Index (per Ravi Mehta's analysis earlier this week)
  • The combination of Claude's enterprise momentum and now Karpathy's R&D focus suggests Anthropic is playing a long game on model quality and researcher credibility

Why it matters for PMs: Talent signals often precede product signals by 12-18 months. Karpathy at Anthropic is the kind of hire that shifts where the best model research happens next. If you're making build-vs-buy decisions or choosing which foundation model to build on, this is a reason to watch Anthropic's model releases more closely over the next year. It also reinforces that Anthropic is differentiating on researcher-brand trust, which matters in enterprise sales.

Critical questions:

  • Karpathy's stated passion is education — does Anthropic bring that lens to how they train models or explain them, or is this purely an R&D hire?
  • How does this affect OpenAI's researcher retention and perception? Does it accelerate or dampen the narrative that Anthropic is "winning" the talent war?
  • Is the Eureka Labs work on pause indefinitely, or is there a signal here about Anthropic's interest in AI-native education products?

Action you could take today: If you haven't read Ravi Mehta's OpenAI vs. Anthropic post from this week (https://blog.ravi-mehta.com/p/openai-vs-anthropic), read it alongside this news. The talent move and the enterprise spending shift are telling the same story from two different angles.

OpenAI / Virgin Atlantic — Codex in Production#

Source: https://openai.com/index/virgin-atlantic Credibility: High (first-party case study from OpenAI, co-published with Virgin Atlantic)

What happened: OpenAI published a case study on how Virgin Atlantic used Codex to ship a revamped mobile app on a fixed holiday travel deadline. The headline results: near-total unit test coverage and zero P1 defects at launch. The constraint was real — holiday travel deadlines don't move — so this isn't a "we used AI and things got better" story. It's a "we had a hard date and we hit it" story.

Key details:

  • Virgin Atlantic was shipping a mobile app revamp with a non-negotiable holiday travel deadline
  • Codex was used to accelerate development and testing, not just code generation
  • Near-total unit test coverage suggests Codex was used for test writing, which is often the first thing cut under deadline pressure
  • Zero P1 defects at launch is a meaningful quality signal — this isn't just shipping faster, it's shipping cleaner
  • OpenAI and GitHub were both named Leaders in the 2026 Gartner Magic Quadrant for Enterprise AI Coding Agents (separate announcement same day)

Why it matters for PMs: Fixed deadlines are where AI-assisted development is most legible as a business case. When the alternative is "we miss the deadline or cut scope," Codex as a forcing function for test coverage is a concrete ROI argument. For PMs evaluating AI coding tools for their teams, this is the kind of case study that answers the "but does it actually ship?" question with a name and a date. The test-writing angle is worth flagging — test coverage is usually the first casualty of deadline pressure, and AI-assisted test generation is an underrated use case.

Critical questions:

  • How much of the team was using Codex vs. traditional development? Is this a full-team adoption story or a targeted use case?
  • What did the Virgin Atlantic team look like before Codex adoption — senior-heavy, junior-heavy, outsourced? The baseline matters for how transferable this is.
  • "Near-total unit test coverage" is qualitative. What percentage, and what types of tests? Unit tests alone don't catch integration failures.
  • Did Codex change the scope of what they shipped, or just the speed? A smaller app shipped on time isn't the same as the full app shipped on time.

Action you could take today: If your team has a hard deadline coming up in the next quarter, document now what scope you'd cut if you ran out of time. Then ask whether AI-assisted test writing could protect that scope. It's a more concrete evaluation than "should we try Codex" in the abstract.

Quick Hits#

  • Cursor: Cursor Automations now available in the Agents Window; multiple repos can be attached per automation; new automation runs are 50% off for the first 7 days. Also: Cursor is now in Jira — assign work items to Cursor directly or @mention it in comments to kick off cloud agents. (May 19-20, 2026): https://cursor.com/changelog/05-20-26

  • Dan Shipper / Every: "After Automation" report — Every grew from 4 to 30 people as AI improved, with demand shifting toward human judgment and framing. Full analysis above, but worth bookmarking the source directly. (May 21, 2026): https://every.to/p/after-automation

  • Ravi Mehta: "OpenAI has the smarter model. Anthropic is winning anyway." — Analysis of Ramp's May 2026 AI Index showing Anthropic overtook OpenAI in enterprise spending, with a breakdown of why product trust and developer experience matter more than benchmark scores. (May 19, 2026): https://blog.ravi-mehta.com/p/openai-vs-anthropic

  • Hugging Face / Thomas Wolf: "Specialization Beats Scale: A Strategic Variable Most AI Procurement Decisions Overlook" — Argues that smaller, specialized models consistently outperform larger general models in domain-specific tasks, and most enterprise procurement decisions underweight this. Direct input for build-vs-buy AI decisions. (May 22, 2026): https://huggingface.co/blog/Dharma-AI/specialization-beats-scale

  • Microsoft Research: "Smarter AI agents, built to run on smaller models" — MagenticLite and MagenticBrain are research systems optimizing agentic task performance for small models, not just frontier ones. Worth watching as a signal that agentic AI is moving toward on-device and cost-efficient deployment. (May 22, 2026): https://www.microsoft.com/en-us/research/blog/magenticlite-magenticbrain-fara1-5-an-agentic-experience-optimized-for-small-models/

The Thread#

The "AI changes the job mix, not just the job count" pattern is becoming clearer. Dan Shipper's report, the Virgin Atlantic case study, and the Karpathy hire all point in the same direction: AI doesn't flatten teams, it reshapes them. Headcount shifts toward judgment and differentiation; the constraint moves from execution speed to framing quality. The open question for PMs is whether their teams are actually building the judgment-heavy skills that remain scarce, or just getting faster at the things that are becoming commoditized.

Sit With This#

Dan Shipper reports that as AI made execution cheap, Every's demand for human judgment, framing, and differentiation increased — growing the team from 4 to 30 rather than shrinking it.

For your team: Which of your team's current outputs are "default outputs" that AI could produce adequately? And if AI handled those, would you hire for more judgment-intensive work — or would your stakeholders just expect the same team to produce more volume?