Where We Actually Are: AI, Generalists, and Everything Else
The code got easier and the job got bigger. On token bills with nothing to show, the apprenticeship we're deleting, and what's scarce now.
Fourteen months ago I wrote that AI was changing software engineering, and that generalists were becoming viable again.
I want to update that, because it’s no longer a prediction.
If you know what you want to build, and you put enough rails around the agents, the implementation is not where your value comes from anymore. That part is settled, even if plenty of people are still in denial about it.
That isn’t the same as saying implementation was easy. It was hard, and it was hard for a reason worth naming. I’ll come back to that.
So the interesting question isn’t whether this happened. It’s what it does to everyone whose job was the implementation.
The Bill Came Due
2026 is the year companies found out what tokens cost.
Uber burned through its entire 2026 AI budget in four months. Not four months of careless experimentation. Four months of an official rollout, complete with an internal leaderboard ranking teams by how much they used the tools.
Their COO, Andrew Macdonald, put it plainly: “If you’re not actually able to draw a direct line to how [many] useful features and functionality you’re shipping to your users, that trade becomes harder to justify.” On whether the productivity showed up: “That link is not there yet.”
Meta ran the same play and got the same result. It also ran an internal leaderboard tracking token spend, and shut it down once internal AI use put the company on track to spend billions in 2026. Adam Mosseri’s description of where some of that money went: “token incinerators.” His projection for where this lands: “the burn rate of a strong engineer might be the same as their salary, or their cost of employment. And in that world, you’re going to probably need to put in some caps.”
Microsoft cancelled most internal Claude Code licenses in its Experiences and Devices division, the group behind Windows, Office and Teams, effective at the end of its fiscal year in June. The trigger was familiar: costs ran past the annual AI budget months ahead of schedule. Engineers were moved onto GitHub Copilot CLI.
That last one is at least half politics. Microsoft would rather not rent intelligence from a competitor by the token, and it has its own tool to sell. But the budget still ran out first, and it ran out for the same reason it ran out everywhere else.
These are not companies that struggle with technology. They are some of the best engineering organizations on the planet, and they went through a year of budget without being able to point at what they got for it.
And here’s the part that should worry you more than the bills. None of them paid full price.
Inference is being sold below what it costs to serve. OpenAI, Anthropic, Google and Meta are all doing it, buying market share with investor money. Sam Altman has said plainly that OpenAI loses money on its $200 a month subscriptions. GitHub spent a period quietly absorbing up to eight times the subscription value for its heaviest users.
That isn’t a stable arrangement and it’s already unwinding. Anthropic moved enterprise customers to usage-based billing in April. Estimates for what normalized pricing looks like start at 30 to 50% higher and climb from there.
That gap probably closes eventually. Open models keep getting better and cheaper to run, and plenty of what we send to a frontier model today will run on something local and unremarkable in a few years. But eventually isn’t a plan. Between here and there is a stretch where the bill is real, it arrives every month, and it’s attached to work you have already made structural.
So be careful which word you reach for. Agents make you faster. Whether they make you cheaper is a question almost nobody has answered, because almost nobody has seen the real price yet. Build your operating model on the speed. Don’t build it on a discount someone else is funding.
Nobody Complains About a Bet That Worked
This is what I keep coming back to.
If that spending had produced outcomes, nobody would be complaining. They would be doubling down. A CFO who watches a line item triple and revenue follow doesn’t cap the line item. They ask why it isn’t bigger.
The complaint is the signal. It tells you the money went out and nothing came back.
This isn’t limited to the biggest spenders. MIT’s NANDA study found that 95% of enterprise GenAI pilots delivered zero measurable impact on the P&L. A survey of 200 finance chiefs at companies with $500M to $10B in revenue found that 14% could point to clear, measurable impact from their AI investments. Fourteen percent. Meanwhile 48% of those same CFOs said they were the ones ultimately accountable for delivering it.
That is not a technology failure. The models work fine. It’s a failure to know what problem you were solving before you started spending.
AI for the sake of AI doesn’t produce anything. It never did. We just found an expensive new way to prove it.
We Went Back to Measuring Activity
Look at what both leaderboards actually did. They counted tokens consumed, then ranked people by the count.
We have done this before. We built Jira and started measuring velocity instead of asking whether anyone used what we shipped. Story points went up. Nothing changed.
Token spend is the new story point. Activity dressed up as progress, with the same flaw: easy to measure, and not the thing you care about.
Meta told employees that demonstrating AI-driven results would be a core performance requirement, with bonuses attached. People optimized for the metric. They always do. That isn’t a character flaw, it’s what happens when you tell people which number matters.
If you want outcomes, measure outcomes. If you measure activity, activity is exactly what you’ll get.
Though I should complicate my own advice, because “measure outcomes” is one sprint planning away from becoming the next thing we cargo cult.
Outcomes get gamed too, and in lending they get gamed quietly. You can lift conversion by accepting worse risk, and the loss shows up two quarters later in someone else’s number. You can cut fraud to nearly zero by rejecting good customers. You can reduce support tickets by making support harder to reach. Every one of those is a team hitting its target and damaging the company.
So a target on its own isn’t enough. If conversion is the target, loss rate is the number that isn’t allowed to move, and somebody has to still be around in two quarters when it does. Otherwise “own the outcome” just means own the metric, and people are extremely good at owning metrics.
How We All Became Cogs
Specialization isn’t a software invention. We inherited it.
The industrial era worked out that if you break a complex job into small repeatable pieces, you can hire cheaper, train faster, and scale further. Nobody builds the whole thing. Everybody builds their part. It was enormously effective and it built the modern world.
Adam Smith opened The Wealth of Nations with the pin factory, the example that made the division of labour famous. Later in the same book, he wrote down the cost: “The man whose whole life is spent in performing a few simple operations… generally becomes as stupid and ignorant as it is possible for a human creature to become.”
That was 1776. We built the next 250 years on the first half of Smith’s observation and quietly skipped the second.
He was ignored, and reasonably so. The economics were firmly on the other side.
Software did the same thing for the same reason. Early software came from small teams who handled everything: architecture, code, deployment, operations. Then the stack got taller. Frontend, backend, mobile, data, platform, SRE, security. Each layer got deep enough to need someone who only did that. Cheap capital did its part: when money is free, a specialist for every layer reads as ambition rather than overhead.
The complexity was real (most of it). This wasn’t a mistake.
But it went further than the complexity required. We ended up with roles defined by activity rather than by depth. A ticket arrives. Someone else described it, someone else prioritized it, someone else specified it. You turn it into a React component. It goes into a system you have never seen end to end, solving a problem for a user you have never met.
A person who only knows their slice is easy to direct and easy to replace. That was a feature for the org. It was never a feature for the person.
The Cog Job Is Going Away
I want to be careful here, because this is the point where it stops being an interesting argument and starts being someone’s career.
If your job is to receive a well-specified ticket, produce the component, and hand it back, the machine does that now. Not badly. Not eventually. Now.
That work is disappearing, and it is not the fault of the people doing it. This usually gets skipped. We built an industry that asked people to work this way. We hired for it. We wrote job ladders around it. We told people that staying in their lane and closing their tickets was the professional thing to do, and many of them did it well for fifteen years.
Now we’re telling them the thing we asked for has no value. That’s a hard message and it deserves more honesty than it normally gets.
The way through isn’t learning a new framework. It’s a change in what you’re responsible for: from the piece to the problem. What are we building, who is it for, and how will we know it worked. Those questions were always the job. Most people were never allowed anywhere near them.
Some will make that jump quickly. Others will find it genuinely hard, not for lack of ability, but because nobody ever asked them to think that way. Unlearning fifteen years of being handed the answer is real work.
Ninety and Ten
Here’s my prediction, specific enough to be wrong. Call it 2029, and hold me to it.
Roughly ninety percent of product teams in technology will be small groups that own outcomes end to end. Three, four, five people who talk to users, decide what to build, build it, ship it, and carry the responsibility for whether it worked. No handoffs. No translation layers. Nobody whose entire role is converting one document into another document.
I should be precise about handoffs, because I work in a regulated business and “no handoffs” is the sort of thing people say when nobody has ever asked them to prove a control.
Some handoffs are waste. A ticket passed between four people who each understand a quarter of it is waste. Other handoffs are controls. Separation of duties, independent risk review, someone other than the author approving a change to credit policy: those exist because the cost of being wrong is not paid by the person who was wrong.
The test isn’t how many people touch the work. It’s whether the second pair of eyes is there to translate or to challenge. Delete the translators. Keep the challengers, and give them enough authority to actually stop something.
The other ten percent will be real specialists, and they’ll be worth more than they are today, not less. People who go deep enough that the depth is the product: distributed systems, cryptography, compilers, performance at scale, serious ML, hard security. They won’t sit inside every product team. They’ll build the blocks the other ninety percent stand on, and they’ll be scarce and expensive.
I wrote that ratio down before I heard Netflix’s CPTO describe the same shape.
Elizabeth Stone went on Lenny’s podcast this month. On what Netflix wants less of: “The days of very narrow, deep specialization feel more limited to me.” On the direction of travel: “compared to 5 or 10 years ago, I would believe we have fewer specialists and more people who are generalists or adaptable in multiple directions.”
On what they want more of: “We need more systems thinkers in a world with AI.” In practice that means hiring “people who can look across all the business domains and abstract that to here’s the building blocks we’re going to need in a world with AI.”
She holds the other side firmly. For the places where only a few people in the world understand how something works, Netflix’s encoding and playback systems among them: “I still believe we need specialized practitioners in those spaces.”
Fewer, not none. Building blocks underneath, systems thinkers on top. That’s the same structure, arrived at independently, by someone running it at a scale I don’t operate at.
None of this is a new job description. It’s the CEO’s job description, handed to more people.
The top job has always been the generalist job. Enough finance to read the numbers, enough technology to know what’s possible, enough of the market to know what matters, and expert in none of it. Nobody ever found that strange, because it was the only role where breadth counted as a qualification instead of a lack of focus. Everyone below it was told to pick a lane.
What changed is that the shape is moving down the org chart. Owning an outcome end to end used to require a title. Now it requires a small team and decent tools.
What disappears is the middle. Specialization by activity instead of by depth. Being the frontend person on a team not because frontend is a deep discipline you’ve mastered, but because that’s the box you were put in.
Depth still pays. Lanes don’t.
The Edge Is in the Intersection
Dan Koe has the sharpest line on this: your edge lies more in intersection than in expertise.
That’s what a generalist actually is, and it’s why the word gets misread. A generalist is not someone who knows a little about a lot. Breadth on its own is worth very little. A generalist is someone with enough real depth in several places to see connections nobody standing in a single lane can see.
Martin Fowler and his colleagues at Thoughtworks call these people Expert Generalists, and argue the industry has it backwards with its obsession over narrow tool expertise. Their definition is the useful one: a broad general skill plus several areas of deeper knowledge, picked up as the work demanded it. Not breadth instead of depth. Breadth on top of several depths.
Their point about AI is the one worth taking. Expert generalists get more out of LLMs because they know enough to interrogate what comes back. If you can’t tell a good suggestion from a merely plausible one, a faster suggestion machine doesn’t help you.
Knowing the business model and the database. Knowing why a customer abandoned the checkout and what the risk model scored them. Knowing the regulation and the deploy pipeline. Those combinations are where non-obvious solutions come from, and they always were. We just spent twenty years organizing against them.
It also needs high agency to matter. Seeing the connection is worth nothing if you wait for permission to act on it.
And it needs one more thing, which I underrated for years.
Breadth shows you the connection. Agency gets you to act on it. Critical thinking is what tells you the connection was real in the first place.
That has always mattered. It matters more now, because the machine is agreeable. Ask an agent whether your approach is sound and it will usually say yes. Ask it to challenge you and it will produce three polite objections and then agree with you anyway. Confident, plausible, well-formatted output at volume will carry you a long way in the wrong direction before anyone notices.
So the generalist’s advantage isn’t knowing more. It’s having enough context from enough places to feel that something is off when everything about the answer looks right.
I’ve written about this before: question the hype, question the best practices, and question your own instincts, especially the ones that feel good. That was true when the bad advice arrived in conference talks. It’s more true now that it arrives instantly, in your editor, sounding certain.
It’s September Again
In 1993, Usenet had a rhythm. Every September a wave of university freshmen got network access, showed up, behaved badly, and slowly learned the norms. Then commercial providers, AOL most famously, opened access to everyone and the wave stopped receding. The regulars called it the September that never ended.
Tobi Lütke reached for that phrase to describe now. Responding to someone arguing that pre-2015 software beat anything shipping today, he pointed out that the products of 2015 were built by people who joined in the early 2000s and spent the 90s as teenagers honing the craft while everyone told them they’d amount to nothing. “No one was writing code because it was trendy, they all showed up because they could not stay away. Eternal September 2.”
I have sympathy for that. I also think it’s partly nostalgia. There was a mountain of terrible software in 1995 and we’ve forgotten all of it, and mass access produced an enormous number of excellent engineers who would never have found a way in otherwise. Survivorship bias is doing quiet work in that argument.
But the useful part survives the caveat. Programming used to have a filter, and the filter selected for obsession. Mainstream salaries removed part of it. AI removes most of what’s left, because you can now produce working software without spending years understanding why it works.
So yes, there’s an advantage right now in having been here a while. Not because the old code was better. Because if you’ve spent twenty years solving problems with technology, mostly because you couldn’t help yourself, you’ve built a sense for which problems are worth solving and what a good solution looks like. That’s the part the tools don’t hand you.
It’s a small advantage and it won’t last. The newcomers learn, they always do. But it’s real today, and if you have it you should be using it.
Where the Next Ones Learn
There’s a hole in what I just said, and I don’t have a clean answer for it.
I claimed the advantage comes from twenty years of solving problems with technology. Then I said the work that used to teach that is going away. Both of those can’t sit comfortably at the same time.
Nobody arrives with judgment. You get it from debugging things you didn’t write, from carrying a pager, from shipping something that broke and having to explain why, from watching a customer use your product wrong and realizing it was your fault. The ticket queue was a bad career. It was also a real apprenticeship.
If we delete the apprenticeship and keep the expectation, we get a generation who can direct agents and can’t tell when the agents are wrong.
That’s a leadership problem, not a personal one. It means buying deliberately what used to happen by accident. Put people where things break: on call, in incidents, on the support queue, owning one small system all the way through. And have them review what the agents got wrong, not just what the agents produced.
Most companies haven’t worked this out. Neither have we.
Code Still Has Value. Reading It Might Not.
None of this means implementation stopped mattering. That’s the wrong lesson, and I watch people draw it constantly.
Uncle Bob Martin, who started writing code in the late 60s, described his current approach: he doesn’t read the code his agents write. That’s the only way he can take the productivity. Instead he surrounds them with extreme constraints. Unit tests, gherkin tests, QA procedures, quality metrics, mutation testing, coverage. “In the end, I have very high confidence in the code they produce because they’ve had to run the gauntlet of all of my constraints and tests.”
Read that twice, because it’s easy to hear as “quality stopped mattering.” It’s the exact opposite. He can stop reading the code precisely because he built a machine that won’t let bad code through. Designing that gauntlet takes more engineering judgment than writing the code ever did.
This is what I meant in the earlier post about deterministic building blocks around non-deterministic agents. Simon Willison calls the professional version of this “vibe engineering”: senior people accelerating with LLMs while staying accountable for what ships. His conclusion is that AI amplifies existing expertise. It amplifies the absence of it just as well.
None of that is resolved. Elizabeth Stone is openly uncomfortable with it. On the code the agents are writing: “they’re very hard to follow. It’s like I know I’m getting better performance from this but I have no idea why and if this thing breaks I’m going to have no idea how to fix it. That makes me uncomfortable.”
Both of them are right, and the distance between them is the useful part. A gauntlet proves the code does what you told it to. It says nothing about what happens when the system does something nobody thought to test for, at three in the morning, with real customers’ money. Strong tests make you confident, and confidence is the thing you least want to be wrong about.
So that heading is narrower than it sounds. The question was never whether humans read every line. It’s what evidence is enough for a person to put their name on something they didn’t write. For most code, a good gauntlet is enough. For the parts where being wrong is expensive and slow to discover, somebody still has to understand how the thing works.
The value moved. It didn’t evaporate. It went from producing the code to defining what correct means and building the system that enforces it.
Cheap For Now, Expensive Forever
There’s a second thing that moved, and it gets less attention than it should.
The cost of producing software has collapsed, at least at today’s subsidized prices. The cost of owning it has not moved at all, and nobody is subsidizing that.
Every service you generate still needs somewhere to run. It still has a security surface, a dependency tree that rots, an upgrade path, a migration waiting for it in three years, an on-call rotation, and a retention policy for whatever data it touches. None of that got cheaper this year. Some of it got more expensive, because there is more of it.
This is the part I’d put in front of a CFO. Tokens per engineer is not the number. Neither is features shipped. The number is what it costs to run and keep trustworthy everything you decided to build, for as long as customers depend on it.
Cheap production plus unchanged ownership costs has an obvious failure mode: you generate far more than you can carry, and you find out eighteen months later.
The discipline this asks for is unfashionable. Build less. Delete more. Be much more suspicious of things that are easy to start.
The way out isn’t building less of everything. It’s deciding upfront which category something is in.
Some software exists to answer a question. Build it fast, put it in front of real users, learn, then delete it. Cheap implementation makes this genuinely better than it used to be: you can now afford to disprove an idea properly instead of arguing about it in a document for three weeks.
Other software is infrastructure. It gets tests, ownership, an on-call rotation, a migration plan, and a person whose name is on it.
The failure is the middle: a prototype that quietly became production because it worked and nobody decided anything. That’s always been a problem. Cheap generation makes it a much bigger one, because the prototypes now arrive faster than the decisions about them.
The middle disappears for people, as I said. For software it does the opposite. It grows, unless someone keeps deciding.
Building Was the Bottleneck. Now It’s Being Chosen.
If everyone can build, working software stops being scarce. Something else becomes the constraint, and it’s worth being precise about what.
John Burn-Murdoch asked how much value AI is really creating in the FT, and charted the answer about as plainly as it can be put. Since agentic tools arrived, iOS app releases have climbed to roughly 180% of their 2024 average. Over the same stretch, the number of apps with significant usage has not moved, and app reviews have fallen.
Nearly twice as many apps. The same number of people using any of them.
The study underneath that chart is worth more than the chart. Demirer, Musolff and Yang tracked more than 100,000 GitHub developers against their actual AI usage telemetry. Autonomous coding agents raised commits by 180%. That 180% became 50% by the time you count projects, and 30% by the time you count real releases.
Line the whole thing up. Code written: up 180%. Software released: up 30%. Software anyone uses: unchanged.
The gains are real, and they shrink at every step away from the keyboard. The authors call it a weak-link problem, which is a polite way of saying the chain moves at the speed of whatever isn’t the typing.
So when I say the question is what’s worth building, here is the specific version. Can you reach the person. Will they trust you with their money and their data. Does it survive contact with their actual workflow. Is there any reason they would still be using it in a year. In a regulated business, add one more: will anyone let you deploy it at all.
None of those are engineering problems. All of them are now the engineer’s problem, because there is nobody left in the chain to hand them to.
This Isn’t Faster. It’s Different.
One last thing, because it’s the most common and most expensive mistake I see.
Most companies are using AI to run their existing process faster. Same discovery, same specs, same tickets, same sprints, same handoffs, with an agent bolted onto the implementation step. That’s why the money goes out and nothing comes back. You’ve made the cheapest part of a slow process slightly cheaper.
The way we build software and products is going to change shape, not speed. The sequence, the artifacts, the size of teams, what a specification even is, how you validate an idea before committing to it. Almost everything we settled into over the last fifteen years was designed around one constraint: implementation was expensive and slow. Remove the constraint and the process built around it stops making sense.
We’re changing it at SeQura right now. I’ll write about what that actually looks like in another post, once I’m confident it’s working and not just interesting.
What I’d Do Now
If you’re an engineer: stop optimizing to be the best at your slice. Get closer to the problem. Find out who the user is and what they need. Learn enough of the adjacent layers to see the whole system. Build the constraints that let you trust output you didn’t write.
If you’re running a team: stop measuring activity. Tokens, story points, tickets closed, it’s the same mistake with a new dashboard each decade. Give small groups whole problems, hold them to outcomes, then stay out of the delivery and in the guardrails. This is the same thing our product manifesto puts first, and it hasn’t changed.
If you’re choosing where to specialize: go deep enough that the depth is the product, or go wide enough to own an outcome. The middle is where the jobs go.
It Was Never the Typing
Forty years ago Fred Brooks split the difficulty of building software in two. Essential complexity is the problem itself, the conceptual structure of the thing you are trying to make. Accidental complexity is everything that comes from expressing it: the languages, the tooling, the coordination, the number of people you needed in a room to get a simple idea into production.
Implementation was hard. I’m not going to pretend otherwise, and anybody who shipped something before this decade knows it. But most of what made it hard was accidental. Engineers were scarce, so you needed a lot of them. A lot of people needed structure, so we built handoffs and specialisms and tickets around them. Then the structure became the job.
That’s the part that’s collapsing. Not the essence. The accidents.
The hard part was never the part that took all the time.
Now that the accidents are getting cheap, what’s left is the thing underneath them all along: what’s worth building, and how will we know it worked?
Stay in the loop
New posts in your inbox when I have something worth saying. No schedule. Just writing.
Related posts
Agile Was About People. Then We Made It About Jira
The Agile Manifesto valued people over process. Then we built Jira and made two-week sprints mandatory for everything.
How AI Is Changing Software Engineering
A CTO's take on what's actually shifting: context over vibes, generalists over specialists, and why understanding problems matters more than writing code.
Profitable, Pragmatic, and Still Standing
How we built a €190M revenue fintech by being contrarian, what broke when we scaled, and why critical thinking beats both rebellion and conformity.