86 Comments
User's avatar
Scenarica's avatar

Everyone has been asking which professions AI replaces. The data says the question is wrong. A lawyer who's never written code was as effective as a software engineer working in an unfamiliar domain. Profession didn't predict success. Understanding the problem did. AI doesn't replace professions. It makes professions porous. The walls between "coder" and "non-coder" dissolve. The walls between "expert" and "novice" get higher, because the thing that makes the agent useful is knowing whether the output is right, and knowing whether the output is right comes from ten years of doing the work by hand.

The closing observation about institutions moving at committee speed while capability moves at exponential speed is the reason every AI policy written before winter 2025 is already obsolete. But there's a human version of the same problem. Individuals also move at human speed. The person who spent six months building a workflow around a chatbot just learned it can run for fourteen hours autonomously, which means the workflow they built is already the wrong workflow. The exponential doesn't just outrun institutions. It outruns the people trying to keep up with it, and the gap between what the tool can do today and what the person learned to do with it last quarter is the new form of obsolescence nobody prepared for.

ChiefofStaff's avatar

David Maister showed us that law firms are made up of 3 different businesses: efficiency, experience and expertise. I've grouped legal activities into 12 categories and mapped those to Maister's 3 layers. Turns out each layer has a very different sensitivity to AI impact. Read the white paper here: https://www.chiefofstaff.pro/maister-paper

Christian Lundgren's avatar

The workflow-obsolescence point is the one I feel most in client work, and the hedge I've found is to write down the judgment rather than the steps. A workflow that encodes "click here, paste this prompt" dies the moment the tool changes, whereas one that encodes what good output looks like and how you'd catch it being wrong survives the model swapping underneath it. That's usually why the people getting whiplash built tight procedures, while the calmer ones built standards and let the tool churn below them.

Anne A’herran's avatar

When talking about “real work” I see that you all mean coding, more or less. Whereas “Real work” for me and other humans is gardening, Planting, pulling weeds, harvesting our veges, cooking, walking the dog, attending committee meetings, baking and putting out the supper, so on. For my adult children, female, it’s breast feeding the Bub, dressing the kids for school, so on and so on, you get the drift. Oh, robots can breast feed now? Wipe up the dog piddle? Eat the cake? TAste test it? Enjoy the cake? Put the leftover cake in the fridge under gladwrap for tomorrow and put it in the lunchbox? No, but they can code.

Jared Harris's avatar

I think this is right and speaks to our idea of what counts as "work". The focus on coding is too narrow. Paving roads, allocating funds, diagnosing bodily pain, etc. also count as work -- things we get paid to do that will be affected more or less quickly by AI. But we need to start appreciating much more the value of things we *want* to do that may not be economically valued. Those need to be our focus going forward.

Susan Knopfelmacher's avatar

A good idea - all those minions ‘released’ by AI will need something to fill their time.

Early Signal's avatar

Even for software engineers, the coding part is not the whole job, and there will be likely be a shift to higher-level work just like there were laundry machines that prevented us from hand washing clothes. Some things will remain uniquely human and hopefully people will learn to emphasize relationships and connection but I think what is considered 'real work' is always changing and adapting

Dr Jen | Syringa Wellness's avatar

For a woman running her solo business, if she taps into AI and agents, she can now have a team that handles a lot of the pieces that sap her attention, maybe make some marketing strategies more feasible, so that she can spend more time in the garden or with the kids or doing the actual craft that she's selling.

Will S Johnston's avatar

It would be interesting to see a chart showing token costs when an agent runs for days comparing the open vs. frontier models. It seems that if the agentic capability is growing this fast, the costs could get out of control quickly. This in turn seems like it would be a big push toward open and undermine frontier unless it's a critical or unique project. Am i missing something?

Alex Tolley's avatar

Indeed. Using more tokens costs more, like paying workers overtime for the extra hour they work above 40. Neither effective productivity nor ROI is improving, apparently, which suggests that these frontier models are not performing as well as the benchmarks suggest they ought. Burning more tokens on a pay-per-use model just lines the pockets of the BigAI companies hoping to cash out with IPOs. Just like restaurants opt for quantity vs quality, so do AI companies push thinking->anegntic->Loop (Anthropic) use to increase consumption spending.

Christian Lundgren's avatar

You're not missing anything, and in production it's already playing out that way. On client automation I route the long, boring, high-volume steps to a cheaper or open model and reserve the frontier model for the few calls where judgment actually matters, because a step that fires constantly has to be cheap or it stops being worth running. It ends up being a per-step decision, not per-project, since most of an agent's runtime is grunt work that doesn't need the expensive brain.

The Writer in the Loop's avatar

The costs are already out of control and the tangible return is not being shown yet.

Andy Barkett's avatar

For software engineers and some others, it's clear and obvious that it's worth tens of thousands of dollars a year - easily.

Nobody who uses it seriously doubts this.

The Writer in the Loop's avatar

I've worked in AI for years, most recently at Meta. The actual productivity gains are yet to be found in the majority of companies.

And the cost is not tens of thousands of dollars a year at these companies - it's in the billions.

I'm not saying that at one point that productivity gains won't be found. I am saying that they haven't been found yet compared to the investment - not even close - and the quality of what is being produced is simultaneously being deteriorated.

If you find data that points otherwise, I'm all ears.

The Writer in the Loop's avatar

specifically, I am talking about generative AI. The ML models themselves (recommendation systems, etc) are huge revenue drivers.

Greg G's avatar

There are a lot of variables at play here. Frontier models mostly work better but are more expensive. Tokens at a given performance point get cheaper over time, but since performance improves frontier-level tokens are getting more expensive. Many people and companies are somewhat locked into frontier labs, but also model routing is a smaller but booming business that seeks to make it easier for people to use different models for different tasks.

Exactly how all of this will play out remains to be seen.

Antti-Juhani Kaijanaho's avatar

You do not consider Gemma 4 an open weight model released by a non-Chinese company, specifically Google? Or is April too far in the past to count? ... maybe it is, Anthropic released three since then.

My own expectation is that the frontier becomes too expensive for most jobs and we end up running local models, or open models in our own cloud, for most use cases.

Ethan Mollick's avatar

I clarified the difference between frontier (all Chinese companies since OpenAI's one-time release last year) and non-frontier open weights models (small models, Mistral, etc) in the texxt.

Ethan Mollick's avatar

Gemma is good but not frontier.

Antti-Juhani Kaijanaho's avatar

True, but it is an open source model that did not come from China.

Ethan Mollick's avatar

There are lots of non-frontier models from lots of companies, including Mistral from France.

Alec Pritzos's avatar

If the Claude Code result holds, that expertise beat profession, then the scarce input just moved from writing the code to knowing what good looks like. A manager who can't tell a right answer from a confident wrong one gets less out of an agent, not more, no matter how many they run at once. That's the part orgs planning agent rollouts keep underpricing: the bottleneck shifts to judgment, and judgment doesn't scale by adding seats.

The Writer in the Loop's avatar

I wrote an article a few days ago about how management skills are now being applied to agents (https://substack.com/home/post/p-203914727), and it's fascinating to see the shift.

AI in some ways is both smarter and more knowledgeable than employees and also more unruly and rigid. For example, I was scrolling through substack the day after I wrote the above article and just saw soooo many articles that looked the same and sounded the same.

I have used Clause plenty and even found myself, someone with a journalism degree and a master's in English, over-relying on it and questioning my own writing and thinking while also increasing my own AI usage. I got so tired of hearing that my drafts had "good bones." I know AI doesn't get the actual emotion side of things and works in probability, and I also have the real-life experience of knowing a professor telling me something had "good bones" was just another way to say it was terrible. Then, Claude would re-write it to sound exactly like, well... Claude... and the dozens of other articles I scrolled past in an effort to find something more engaging, meaningful, and useful. I took a hard look and made a hard stop.

I am still using Agents (and sometimes teams of agents) for automation, fact checking, doing deep research (leveraging other agents to check), and double checking my writing (with very specific prompts and skills on how to do this), but there are parts of the human experience that AI can't understand.

Yes, management skills (assessing the problem, laying out a plan, supervising the agent, watching it through to completion) are and will continue to be important AND I think anyone who says they know what work will look like actually does NOT know (I'm not referring to you or your article, which was thoughtful and well-researched). What I see so often is a "credibility" play (something I have been guilty of doing myself) or a "marketing" play but no one really knows how this will play out. The researchers building these systems don't even fully understand how they work. Company leaders rolling it out are still figuring it out and it is changing constantly.

And a lot of the hype I feel is companies needing to justify the huge amount of money they are spending on powerful technology that is, at best, in beta trials at the companies trying to role it out.

Many of us are going along for the ride feeling like we need to "keep up" but out-sourcing our thinking to AI (whether it is through AI Agents or by allowing recommendation systems to determine where and how we spend our time), is where we get into trouble. And the hallmark of good management is to listen to others and then to actually think and exercise discernment. So, perhaps, this is the most important management skill to take seriously - discernment in where it makes sense to leverage this technology, where it doesn't, and the humility to acknowledge when we just aren't sure yet.

Nesibe Kiris Can's avatar

The twilight framing is exactly right and slightly uncomfortable for anyone who has been bullish on chat interfaces as the primary AI interaction model. Agents do not replace chatbots for everything, but for any task that requires sustained action, external tool use, or multi-step reasoning across time, the chatbot model was always a workaround rather than a solution. The shift is less about capability and more about what the interface is optimized for.

Colin Gautrey's avatar

The power of expertise is being restructured and will never be the same again.

Rajesh Achanta's avatar

You have a rare gift for turning this blur into something legible: the exponential-from-the-inside framing is the best account of why this keeps feeling like leaps rather than a curve.

What caught my attention is how this post ends. The AI conversation until now—your corpus, and most comments here—runs on two registers: how good the models are, and how work reshapes around them. This is the first piece I've seen where a third thing surfaces, even unnamed. Human-speed institutions trying to track a curve that isn't human is exactly that.

I called it the society clock in an essay on Monday (four in all: customer, model, infrastructure, society: https://rajeshachanta.substack.com/p/no-time-for-pit-stops). A thing worth adding about this clock: it doesn't behave like the others. The capability clocks are exponental curves, accelerating. This one is about consent — and consent doesn't lag, it withdraws, suddenly and non-linearly. Which is why the interventions that stopped Fable and Mythos weren't slowness, they were legitimacy withdrawing at once. Not a lag, a lurch. The gap may not close even if the humans speed up, because that clock was never running on speed.

Kevin McLeod's avatar

The glass ceiling is the symbolic. This doesn't even reach the level of the brain's compute which is multi-dimensional. It's an interesting threshold, but binary is limited.

Paul Wilkinson 🧢's avatar

After working with 90-120 students a day plus colleagues and family and friends, managing 2-4 agents seems easy. Reporting news with dozens of sources and answering to multiple editors was good prep, too.

Eric Porres's avatar

My experiments with Fable 5 led me to say that it’s AGI. I haven’t moved from that position.

John Blanch's avatar

I thrashed Fable in Claude Code from pretty much the moment it dropped to when they took it back a few short days later. Subjectively for me and the work I was doing it was quite a bit better than Opus 4.8, but not a staggering leap. But that's just like, my opinion man.

The Writer in the Loop's avatar

And how would you define AGI?

Eric Porres's avatar

See this coming Tuesday’s newsletter

Oddish's avatar

I'm devastated that this had nothing to do with chat bots.

Ethan Mollick's avatar

But it does! We are moving from chatbots to agentic systems, which is what the adoption graphs are showing

Oddish's avatar

But babe I wanted poetry to smarterchild

James Holt's avatar

The expertise finding is the one that matters most here. If domain experience is what makes someone good at directing an agent, that raises a question this piece doesn't answer: where does the next generation's domain experience come from, if the reps that build it are exactly the tasks agents now handle? Right now the org chart still has people who built judgment the slow way, by doing the small stuff badly a few hundred times before they got good at it. The management model you're describing works great as long as that supply keeps arriving. I'm less sure it does, five years out.

Susan Knopfelmacher's avatar

Love the way these authors disappear, when the questions get interesting!

Omri Ben-Shoham's avatar

The "think of yourself as a manager" framing matches what I see leading an R&D team building voice agents. The skill that transfers is not prompting, it's the same judgment a manager needs: knowing when to trust a result versus when to check it closely, and building the checks that catch the difference before a customer does. Domain expertise seems to matter more than technical skill for exactly that reason - you know what wrong looks like faster.

Sergey Morev's avatar

Same stochastic LLM.

Better execution architecture.

Random errors have less impact on the final result.

Rafael Abreu's avatar

If they don't compound.

Andy Barkett's avatar

humans are stochastic.