Everyone has been asking which professions AI replaces. The data says the question is wrong. A lawyer who's never written code was as effective as a software engineer working in an unfamiliar domain. Profession didn't predict success. Understanding the problem did. AI doesn't replace professions. It makes professions porous. The walls between "coder" and "non-coder" dissolve. The walls between "expert" and "novice" get higher, because the thing that makes the agent useful is knowing whether the output is right, and knowing whether the output is right comes from ten years of doing the work by hand.
The closing observation about institutions moving at committee speed while capability moves at exponential speed is the reason every AI policy written before winter 2025 is already obsolete. But there's a human version of the same problem. Individuals also move at human speed. The person who spent six months building a workflow around a chatbot just learned it can run for fourteen hours autonomously, which means the workflow they built is already the wrong workflow. The exponential doesn't just outrun institutions. It outruns the people trying to keep up with it, and the gap between what the tool can do today and what the person learned to do with it last quarter is the new form of obsolescence nobody prepared for.
David Maister showed us that law firms are made up of 3 different businesses: efficiency, experience and expertise. I've grouped legal activities into 12 categories and mapped those to Maister's 3 layers. Turns out each layer has a very different sensitivity to AI impact. Read the white paper here: https://www.chiefofstaff.pro/maister-paper
The workflow-obsolescence point is the one I feel most in client work, and the hedge I've found is to write down the judgment rather than the steps. A workflow that encodes "click here, paste this prompt" dies the moment the tool changes, whereas one that encodes what good output looks like and how you'd catch it being wrong survives the model swapping underneath it. That's usually why the people getting whiplash built tight procedures, while the calmer ones built standards and let the tool churn below them.
When talking about “real work” I see that you all mean coding, more or less. Whereas “Real work” for me and other humans is gardening, Planting, pulling weeds, harvesting our veges, cooking, walking the dog, attending committee meetings, baking and putting out the supper, so on. For my adult children, female, it’s breast feeding the Bub, dressing the kids for school, so on and so on, you get the drift. Oh, robots can breast feed now? Wipe up the dog piddle? Eat the cake? TAste test it? Enjoy the cake? Put the leftover cake in the fridge under gladwrap for tomorrow and put it in the lunchbox? No, but they can code.
I think this is right and speaks to our idea of what counts as "work". The focus on coding is too narrow. Paving roads, allocating funds, diagnosing bodily pain, etc. also count as work -- things we get paid to do that will be affected more or less quickly by AI. But we need to start appreciating much more the value of things we *want* to do that may not be economically valued. Those need to be our focus going forward.
For a woman running her solo business, if she taps into AI and agents, she can now have a team that handles a lot of the pieces that sap her attention, maybe make some marketing strategies more feasible, so that she can spend more time in the garden or with the kids or doing the actual craft that she's selling.
Even for software engineers, the coding part is not the whole job, and there will be likely be a shift to higher-level work just like there were laundry machines that prevented us from hand washing clothes. Some things will remain uniquely human and hopefully people will learn to emphasize relationships and connection but I think what is considered 'real work' is always changing and adapting
It would be interesting to see a chart showing token costs when an agent runs for days comparing the open vs. frontier models. It seems that if the agentic capability is growing this fast, the costs could get out of control quickly. This in turn seems like it would be a big push toward open and undermine frontier unless it's a critical or unique project. Am i missing something?
Indeed. Using more tokens costs more, like paying workers overtime for the extra hour they work above 40. Neither effective productivity nor ROI is improving, apparently, which suggests that these frontier models are not performing as well as the benchmarks suggest they ought. Burning more tokens on a pay-per-use model just lines the pockets of the BigAI companies hoping to cash out with IPOs. Just like restaurants opt for quantity vs quality, so do AI companies push thinking->anegntic->Loop (Anthropic) use to increase consumption spending.
You're not missing anything, and in production it's already playing out that way. On client automation I route the long, boring, high-volume steps to a cheaper or open model and reserve the frontier model for the few calls where judgment actually matters, because a step that fires constantly has to be cheap or it stops being worth running. It ends up being a per-step decision, not per-project, since most of an agent's runtime is grunt work that doesn't need the expensive brain.
I've worked in AI for years, most recently at Meta. The actual productivity gains are yet to be found in the majority of companies.
And the cost is not tens of thousands of dollars a year at these companies - it's in the billions.
I'm not saying that at one point that productivity gains won't be found. I am saying that they haven't been found yet compared to the investment - not even close - and the quality of what is being produced is simultaneously being deteriorated.
If you find data that points otherwise, I'm all ears.
There are a lot of variables at play here. Frontier models mostly work better but are more expensive. Tokens at a given performance point get cheaper over time, but since performance improves frontier-level tokens are getting more expensive. Many people and companies are somewhat locked into frontier labs, but also model routing is a smaller but booming business that seeks to make it easier for people to use different models for different tasks.
Exactly how all of this will play out remains to be seen.
You do not consider Gemma 4 an open weight model released by a non-Chinese company, specifically Google? Or is April too far in the past to count? ... maybe it is, Anthropic released three since then.
My own expectation is that the frontier becomes too expensive for most jobs and we end up running local models, or open models in our own cloud, for most use cases.
I clarified the difference between frontier (all Chinese companies since OpenAI's one-time release last year) and non-frontier open weights models (small models, Mistral, etc) in the texxt.
If the Claude Code result holds, that expertise beat profession, then the scarce input just moved from writing the code to knowing what good looks like. A manager who can't tell a right answer from a confident wrong one gets less out of an agent, not more, no matter how many they run at once. That's the part orgs planning agent rollouts keep underpricing: the bottleneck shifts to judgment, and judgment doesn't scale by adding seats.
The expertise finding is the one that matters most here. If domain experience is what makes someone good at directing an agent, that raises a question this piece doesn't answer: where does the next generation's domain experience come from, if the reps that build it are exactly the tasks agents now handle? Right now the org chart still has people who built judgment the slow way, by doing the small stuff badly a few hundred times before they got good at it. The management model you're describing works great as long as that supply keeps arriving. I'm less sure it does, five years out.
I wrote an article a few days ago about how management skills are now being applied to agents (https://substack.com/home/post/p-203914727), and it's fascinating to see the shift.
AI in some ways is both smarter and more knowledgeable than employees and also more unruly and rigid. For example, I was scrolling through substack the day after I wrote the above article and just saw soooo many articles that looked the same and sounded the same.
I have used Clause plenty and even found myself, someone with a journalism degree and a master's in English, over-relying on it and questioning my own writing and thinking while also increasing my own AI usage. I got so tired of hearing that my drafts had "good bones." I know AI doesn't get the actual emotion side of things and works in probability, and I also have the real-life experience of knowing a professor telling me something had "good bones" was just another way to say it was terrible. Then, Claude would re-write it to sound exactly like, well... Claude... and the dozens of other articles I scrolled past in an effort to find something more engaging, meaningful, and useful. I took a hard look and made a hard stop.
I am still using Agents (and sometimes teams of agents) for automation, fact checking, doing deep research (leveraging other agents to check), and double checking my writing (with very specific prompts and skills on how to do this), but there are parts of the human experience that AI can't understand.
Yes, management skills (assessing the problem, laying out a plan, supervising the agent, watching it through to completion) are and will continue to be important AND I think anyone who says they know what work will look like actually does NOT know (I'm not referring to you or your article, which was thoughtful and well-researched). What I see so often is a "credibility" play (something I have been guilty of doing myself) or a "marketing" play but no one really knows how this will play out. The researchers building these systems don't even fully understand how they work. Company leaders rolling it out are still figuring it out and it is changing constantly.
And a lot of the hype I feel is companies needing to justify the huge amount of money they are spending on powerful technology that is, at best, in beta trials at the companies trying to role it out.
Many of us are going along for the ride feeling like we need to "keep up" but out-sourcing our thinking to AI (whether it is through AI Agents or by allowing recommendation systems to determine where and how we spend our time), is where we get into trouble. And the hallmark of good management is to listen to others and then to actually think and exercise discernment. So, perhaps, this is the most important management skill to take seriously - discernment in where it makes sense to leverage this technology, where it doesn't, and the humility to acknowledge when we just aren't sure yet.
The twilight framing is exactly right and slightly uncomfortable for anyone who has been bullish on chat interfaces as the primary AI interaction model. Agents do not replace chatbots for everything, but for any task that requires sustained action, external tool use, or multi-step reasoning across time, the chatbot model was always a workaround rather than a solution. The shift is less about capability and more about what the interface is optimized for.
You have a rare gift for turning this blur into something legible: the exponential-from-the-inside framing is the best account of why this keeps feeling like leaps rather than a curve.
What caught my attention is how this post ends. The AI conversation until now—your corpus, and most comments here—runs on two registers: how good the models are, and how work reshapes around them. This is the first piece I've seen where a third thing surfaces, even unnamed. Human-speed institutions trying to track a curve that isn't human is exactly that.
I called it the society clock in an essay on Monday (four in all: customer, model, infrastructure, society: https://rajeshachanta.substack.com/p/no-time-for-pit-stops). A thing worth adding about this clock: it doesn't behave like the others. The capability clocks are exponental curves, accelerating. This one is about consent — and consent doesn't lag, it withdraws, suddenly and non-linearly. Which is why the interventions that stopped Fable and Mythos weren't slowness, they were legitimacy withdrawing at once. Not a lag, a lurch. The gap may not close even if the humans speed up, because that clock was never running on speed.
The glass ceiling is the symbolic. This doesn't even reach the level of the brain's compute which is multi-dimensional. It's an interesting threshold, but binary is limited.
After working with 90-120 students a day plus colleagues and family and friends, managing 2-4 agents seems easy. Reporting news with dozens of sources and answering to multiple editors was good prep, too.
I thrashed Fable in Claude Code from pretty much the moment it dropped to when they took it back a few short days later. Subjectively for me and the work I was doing it was quite a bit better than Opus 4.8, but not a staggering leap. But that's just like, my opinion man.
I use chatgpt pretty regularly but am behind the curve as to how to use it, mainly as a research assistant, an entity to bounce ideas off of, or as an editor. I work at an ngo type org and direct an office that educates the public on a specific issue (think the environment, economic policy or immigration). I can't code and would not even know what to code for. I am not building a new technology or launching a business (yet??). Given my kind of position, which I think a lot of people are in, what would some of you use AI for as an entity to help educate the public more effectively on a given issue? I'd be curious to start a conversation along these lines with anyone who might be interested.
I love this question! You're describing the EXACT spot most of us in mission-driven work find ourselves in. You don't need to code or build anything, you just need AI to help you do (and amplify!) the public-education work you're ALREADY doing!
This is basically what I spend my days on – working with social movement leaders and NGO folks on using AI in ways that actually stay true to their values and mission. So here's a bit of a brain dump of things I've seen work really well (Fair warning: I clearly have no ability to be brief about a topic I LOVE, so grab a coffee! ;-):
1. You've probably got a ton of dense, credible material sitting around, e.g., reports, briefs, research, etc. AI is wonderful at helping you translate that into different pieces for different people. Same set of facts, but one version for a busy policymaker, one for a skeptical neighbour at a barbecue, one for a curious teenager. And you can get creative with that last one... Ask it to turn your driest 40-page report into a five-question "Which Type Are You?" Quiz, or a "Choose Your Own Adventure" where a teen walks through the real trade-offs of a policy and sees where their choices lead. Suddenly the same facts become something a teen actually WANTS to engage with.
2. A few other formats in that same spirit: a mock text-message exchange between two friends debating the issue (the kind of thing they'd screenshot and share!), a "Myth vs. Reality" flashcard set they can flip through in 90 seconds, or a 1-minute explainer script written the way a teen creator would actually say it (NOT the way a briefing doc reads! ;-)
3. Building on Stackleton's point about interactive graphics, this is where it gets REALLY fun. That Manhattan Institute-style budget toggle is a perfect example: instead of TELLING people what a policy costs, you let them stack the pieces themselves and watch the trade-offs shift in real time. AI has made this kind of thing dramatically more doable for small teams. You can go from "Here's our 40-page report" to a working interactive prototype without a developer! Imagine a slider where someone adjusts one lever on your issue and instantly sees who's affected and how, or a "Build Your Own Plan" tool where they have to make the same hard choices you're trying to explain. Even a lightweight game can showcase the point in a way no fact sheet ever will. The facts don't change; you've just handed people the controls!
Okay, and now for the drier but equally powerful stuff (I promise it can be more fun than it sounds! ;-):
4. Turn AI into your sparring partner. Ask it to argue AGAINST you as hard as a sharp, well-briefed opponent would. Then, poke holes in THOSE arguments too. It's like a simulation of the debates you're actually going to have. This way, you walk into the room already three moves ahead instead of getting blindsided by the one talking point you forgot to prep for!
5. Let AI be your tireless outreach support. Most organizations I work with have an endless task of taking one message and reshaping it for fifteen different audiences? AI is REALLY good at this. For example, the same immigration fact sheet becomes one version for a faith group, one for a small-business owner, one for a college campus (PLUS translations!) in the time it used to take you to do one. This allows you to meet people wherever they are (and without the grind! ;-)
6. All this said, the most important thing is to keep your "fingerprints" on everything! This is YOUR values piece and the one I area I wouldn't budge on. For public education, your credibility IS the whole game. So make sure you have a human review EVERYTHING that goes out under your name. I call this Human-in-the LEAD (not just "in-the-loop"! ;-) Be honest about where you're using AI and be super mindful about what you share with tools you don't control, i.e., in the cloud. Your ultimate goal here is to amplify the trustworthy work you're ALREADY doing (NOT outsource your organization's well-honed judgment and expertise! ;-)
Genuinely happy to keep talking about this! I'm having this conversation with people all day / every day and they're wrestling with the EXACT same questions you are!
Once again, apologies for the small novel I've written here! It's just proof of how much I LOVE this stuff! Feel free to reply here or reach out anytime!
Create interactive graphics (I’m thinking of the toggles I think Manhattan Institute released recently allowing you to see the budget impacts of different policies so you can stack them in different combinations) or even video games that educate on the particular topic
You should literally ask Claude this question. Code can solve lots of problems, especially infused with intelligence. Software does not need to mean apps!
The "think of yourself as a manager" framing matches what I see leading an R&D team building voice agents. The skill that transfers is not prompting, it's the same judgment a manager needs: knowing when to trust a result versus when to check it closely, and building the checks that catch the difference before a customer does. Domain expertise seems to matter more than technical skill for exactly that reason - you know what wrong looks like faster.
Everyone has been asking which professions AI replaces. The data says the question is wrong. A lawyer who's never written code was as effective as a software engineer working in an unfamiliar domain. Profession didn't predict success. Understanding the problem did. AI doesn't replace professions. It makes professions porous. The walls between "coder" and "non-coder" dissolve. The walls between "expert" and "novice" get higher, because the thing that makes the agent useful is knowing whether the output is right, and knowing whether the output is right comes from ten years of doing the work by hand.
The closing observation about institutions moving at committee speed while capability moves at exponential speed is the reason every AI policy written before winter 2025 is already obsolete. But there's a human version of the same problem. Individuals also move at human speed. The person who spent six months building a workflow around a chatbot just learned it can run for fourteen hours autonomously, which means the workflow they built is already the wrong workflow. The exponential doesn't just outrun institutions. It outruns the people trying to keep up with it, and the gap between what the tool can do today and what the person learned to do with it last quarter is the new form of obsolescence nobody prepared for.
David Maister showed us that law firms are made up of 3 different businesses: efficiency, experience and expertise. I've grouped legal activities into 12 categories and mapped those to Maister's 3 layers. Turns out each layer has a very different sensitivity to AI impact. Read the white paper here: https://www.chiefofstaff.pro/maister-paper
The workflow-obsolescence point is the one I feel most in client work, and the hedge I've found is to write down the judgment rather than the steps. A workflow that encodes "click here, paste this prompt" dies the moment the tool changes, whereas one that encodes what good output looks like and how you'd catch it being wrong survives the model swapping underneath it. That's usually why the people getting whiplash built tight procedures, while the calmer ones built standards and let the tool churn below them.
When talking about “real work” I see that you all mean coding, more or less. Whereas “Real work” for me and other humans is gardening, Planting, pulling weeds, harvesting our veges, cooking, walking the dog, attending committee meetings, baking and putting out the supper, so on. For my adult children, female, it’s breast feeding the Bub, dressing the kids for school, so on and so on, you get the drift. Oh, robots can breast feed now? Wipe up the dog piddle? Eat the cake? TAste test it? Enjoy the cake? Put the leftover cake in the fridge under gladwrap for tomorrow and put it in the lunchbox? No, but they can code.
I think this is right and speaks to our idea of what counts as "work". The focus on coding is too narrow. Paving roads, allocating funds, diagnosing bodily pain, etc. also count as work -- things we get paid to do that will be affected more or less quickly by AI. But we need to start appreciating much more the value of things we *want* to do that may not be economically valued. Those need to be our focus going forward.
A good idea - all those minions ‘released’ by AI will need something to fill their time.
For a woman running her solo business, if she taps into AI and agents, she can now have a team that handles a lot of the pieces that sap her attention, maybe make some marketing strategies more feasible, so that she can spend more time in the garden or with the kids or doing the actual craft that she's selling.
Even for software engineers, the coding part is not the whole job, and there will be likely be a shift to higher-level work just like there were laundry machines that prevented us from hand washing clothes. Some things will remain uniquely human and hopefully people will learn to emphasize relationships and connection but I think what is considered 'real work' is always changing and adapting
It would be interesting to see a chart showing token costs when an agent runs for days comparing the open vs. frontier models. It seems that if the agentic capability is growing this fast, the costs could get out of control quickly. This in turn seems like it would be a big push toward open and undermine frontier unless it's a critical or unique project. Am i missing something?
Indeed. Using more tokens costs more, like paying workers overtime for the extra hour they work above 40. Neither effective productivity nor ROI is improving, apparently, which suggests that these frontier models are not performing as well as the benchmarks suggest they ought. Burning more tokens on a pay-per-use model just lines the pockets of the BigAI companies hoping to cash out with IPOs. Just like restaurants opt for quantity vs quality, so do AI companies push thinking->anegntic->Loop (Anthropic) use to increase consumption spending.
You're not missing anything, and in production it's already playing out that way. On client automation I route the long, boring, high-volume steps to a cheaper or open model and reserve the frontier model for the few calls where judgment actually matters, because a step that fires constantly has to be cheap or it stops being worth running. It ends up being a per-step decision, not per-project, since most of an agent's runtime is grunt work that doesn't need the expensive brain.
The costs are already out of control and the tangible return is not being shown yet.
For software engineers and some others, it's clear and obvious that it's worth tens of thousands of dollars a year - easily.
Nobody who uses it seriously doubts this.
I've worked in AI for years, most recently at Meta. The actual productivity gains are yet to be found in the majority of companies.
And the cost is not tens of thousands of dollars a year at these companies - it's in the billions.
I'm not saying that at one point that productivity gains won't be found. I am saying that they haven't been found yet compared to the investment - not even close - and the quality of what is being produced is simultaneously being deteriorated.
If you find data that points otherwise, I'm all ears.
specifically, I am talking about generative AI. The ML models themselves (recommendation systems, etc) are huge revenue drivers.
There are a lot of variables at play here. Frontier models mostly work better but are more expensive. Tokens at a given performance point get cheaper over time, but since performance improves frontier-level tokens are getting more expensive. Many people and companies are somewhat locked into frontier labs, but also model routing is a smaller but booming business that seeks to make it easier for people to use different models for different tasks.
Exactly how all of this will play out remains to be seen.
You do not consider Gemma 4 an open weight model released by a non-Chinese company, specifically Google? Or is April too far in the past to count? ... maybe it is, Anthropic released three since then.
My own expectation is that the frontier becomes too expensive for most jobs and we end up running local models, or open models in our own cloud, for most use cases.
I clarified the difference between frontier (all Chinese companies since OpenAI's one-time release last year) and non-frontier open weights models (small models, Mistral, etc) in the texxt.
Gemma is good but not frontier.
True, but it is an open source model that did not come from China.
There are lots of non-frontier models from lots of companies, including Mistral from France.
And Mistral!
If the Claude Code result holds, that expertise beat profession, then the scarce input just moved from writing the code to knowing what good looks like. A manager who can't tell a right answer from a confident wrong one gets less out of an agent, not more, no matter how many they run at once. That's the part orgs planning agent rollouts keep underpricing: the bottleneck shifts to judgment, and judgment doesn't scale by adding seats.
The expertise finding is the one that matters most here. If domain experience is what makes someone good at directing an agent, that raises a question this piece doesn't answer: where does the next generation's domain experience come from, if the reps that build it are exactly the tasks agents now handle? Right now the org chart still has people who built judgment the slow way, by doing the small stuff badly a few hundred times before they got good at it. The management model you're describing works great as long as that supply keeps arriving. I'm less sure it does, five years out.
Love the way these authors disappear, when the questions get interesting!
I wrote an article a few days ago about how management skills are now being applied to agents (https://substack.com/home/post/p-203914727), and it's fascinating to see the shift.
AI in some ways is both smarter and more knowledgeable than employees and also more unruly and rigid. For example, I was scrolling through substack the day after I wrote the above article and just saw soooo many articles that looked the same and sounded the same.
I have used Clause plenty and even found myself, someone with a journalism degree and a master's in English, over-relying on it and questioning my own writing and thinking while also increasing my own AI usage. I got so tired of hearing that my drafts had "good bones." I know AI doesn't get the actual emotion side of things and works in probability, and I also have the real-life experience of knowing a professor telling me something had "good bones" was just another way to say it was terrible. Then, Claude would re-write it to sound exactly like, well... Claude... and the dozens of other articles I scrolled past in an effort to find something more engaging, meaningful, and useful. I took a hard look and made a hard stop.
I am still using Agents (and sometimes teams of agents) for automation, fact checking, doing deep research (leveraging other agents to check), and double checking my writing (with very specific prompts and skills on how to do this), but there are parts of the human experience that AI can't understand.
Yes, management skills (assessing the problem, laying out a plan, supervising the agent, watching it through to completion) are and will continue to be important AND I think anyone who says they know what work will look like actually does NOT know (I'm not referring to you or your article, which was thoughtful and well-researched). What I see so often is a "credibility" play (something I have been guilty of doing myself) or a "marketing" play but no one really knows how this will play out. The researchers building these systems don't even fully understand how they work. Company leaders rolling it out are still figuring it out and it is changing constantly.
And a lot of the hype I feel is companies needing to justify the huge amount of money they are spending on powerful technology that is, at best, in beta trials at the companies trying to role it out.
Many of us are going along for the ride feeling like we need to "keep up" but out-sourcing our thinking to AI (whether it is through AI Agents or by allowing recommendation systems to determine where and how we spend our time), is where we get into trouble. And the hallmark of good management is to listen to others and then to actually think and exercise discernment. So, perhaps, this is the most important management skill to take seriously - discernment in where it makes sense to leverage this technology, where it doesn't, and the humility to acknowledge when we just aren't sure yet.
The twilight framing is exactly right and slightly uncomfortable for anyone who has been bullish on chat interfaces as the primary AI interaction model. Agents do not replace chatbots for everything, but for any task that requires sustained action, external tool use, or multi-step reasoning across time, the chatbot model was always a workaround rather than a solution. The shift is less about capability and more about what the interface is optimized for.
The power of expertise is being restructured and will never be the same again.
You have a rare gift for turning this blur into something legible: the exponential-from-the-inside framing is the best account of why this keeps feeling like leaps rather than a curve.
What caught my attention is how this post ends. The AI conversation until now—your corpus, and most comments here—runs on two registers: how good the models are, and how work reshapes around them. This is the first piece I've seen where a third thing surfaces, even unnamed. Human-speed institutions trying to track a curve that isn't human is exactly that.
I called it the society clock in an essay on Monday (four in all: customer, model, infrastructure, society: https://rajeshachanta.substack.com/p/no-time-for-pit-stops). A thing worth adding about this clock: it doesn't behave like the others. The capability clocks are exponental curves, accelerating. This one is about consent — and consent doesn't lag, it withdraws, suddenly and non-linearly. Which is why the interventions that stopped Fable and Mythos weren't slowness, they were legitimacy withdrawing at once. Not a lag, a lurch. The gap may not close even if the humans speed up, because that clock was never running on speed.
The glass ceiling is the symbolic. This doesn't even reach the level of the brain's compute which is multi-dimensional. It's an interesting threshold, but binary is limited.
After working with 90-120 students a day plus colleagues and family and friends, managing 2-4 agents seems easy. Reporting news with dozens of sources and answering to multiple editors was good prep, too.
My experiments with Fable 5 led me to say that it’s AGI. I haven’t moved from that position.
I thrashed Fable in Claude Code from pretty much the moment it dropped to when they took it back a few short days later. Subjectively for me and the work I was doing it was quite a bit better than Opus 4.8, but not a staggering leap. But that's just like, my opinion man.
And how would you define AGI?
See this coming Tuesday’s newsletter
I'm devastated that this had nothing to do with chat bots.
But it does! We are moving from chatbots to agentic systems, which is what the adoption graphs are showing
But babe I wanted poetry to smarterchild
I use chatgpt pretty regularly but am behind the curve as to how to use it, mainly as a research assistant, an entity to bounce ideas off of, or as an editor. I work at an ngo type org and direct an office that educates the public on a specific issue (think the environment, economic policy or immigration). I can't code and would not even know what to code for. I am not building a new technology or launching a business (yet??). Given my kind of position, which I think a lot of people are in, what would some of you use AI for as an entity to help educate the public more effectively on a given issue? I'd be curious to start a conversation along these lines with anyone who might be interested.
Hi Tosc,
I love this question! You're describing the EXACT spot most of us in mission-driven work find ourselves in. You don't need to code or build anything, you just need AI to help you do (and amplify!) the public-education work you're ALREADY doing!
This is basically what I spend my days on – working with social movement leaders and NGO folks on using AI in ways that actually stay true to their values and mission. So here's a bit of a brain dump of things I've seen work really well (Fair warning: I clearly have no ability to be brief about a topic I LOVE, so grab a coffee! ;-):
1. You've probably got a ton of dense, credible material sitting around, e.g., reports, briefs, research, etc. AI is wonderful at helping you translate that into different pieces for different people. Same set of facts, but one version for a busy policymaker, one for a skeptical neighbour at a barbecue, one for a curious teenager. And you can get creative with that last one... Ask it to turn your driest 40-page report into a five-question "Which Type Are You?" Quiz, or a "Choose Your Own Adventure" where a teen walks through the real trade-offs of a policy and sees where their choices lead. Suddenly the same facts become something a teen actually WANTS to engage with.
2. A few other formats in that same spirit: a mock text-message exchange between two friends debating the issue (the kind of thing they'd screenshot and share!), a "Myth vs. Reality" flashcard set they can flip through in 90 seconds, or a 1-minute explainer script written the way a teen creator would actually say it (NOT the way a briefing doc reads! ;-)
3. Building on Stackleton's point about interactive graphics, this is where it gets REALLY fun. That Manhattan Institute-style budget toggle is a perfect example: instead of TELLING people what a policy costs, you let them stack the pieces themselves and watch the trade-offs shift in real time. AI has made this kind of thing dramatically more doable for small teams. You can go from "Here's our 40-page report" to a working interactive prototype without a developer! Imagine a slider where someone adjusts one lever on your issue and instantly sees who's affected and how, or a "Build Your Own Plan" tool where they have to make the same hard choices you're trying to explain. Even a lightweight game can showcase the point in a way no fact sheet ever will. The facts don't change; you've just handed people the controls!
Okay, and now for the drier but equally powerful stuff (I promise it can be more fun than it sounds! ;-):
4. Turn AI into your sparring partner. Ask it to argue AGAINST you as hard as a sharp, well-briefed opponent would. Then, poke holes in THOSE arguments too. It's like a simulation of the debates you're actually going to have. This way, you walk into the room already three moves ahead instead of getting blindsided by the one talking point you forgot to prep for!
5. Let AI be your tireless outreach support. Most organizations I work with have an endless task of taking one message and reshaping it for fifteen different audiences? AI is REALLY good at this. For example, the same immigration fact sheet becomes one version for a faith group, one for a small-business owner, one for a college campus (PLUS translations!) in the time it used to take you to do one. This allows you to meet people wherever they are (and without the grind! ;-)
6. All this said, the most important thing is to keep your "fingerprints" on everything! This is YOUR values piece and the one I area I wouldn't budge on. For public education, your credibility IS the whole game. So make sure you have a human review EVERYTHING that goes out under your name. I call this Human-in-the LEAD (not just "in-the-loop"! ;-) Be honest about where you're using AI and be super mindful about what you share with tools you don't control, i.e., in the cloud. Your ultimate goal here is to amplify the trustworthy work you're ALREADY doing (NOT outsource your organization's well-honed judgment and expertise! ;-)
Genuinely happy to keep talking about this! I'm having this conversation with people all day / every day and they're wrestling with the EXACT same questions you are!
Once again, apologies for the small novel I've written here! It's just proof of how much I LOVE this stuff! Feel free to reply here or reach out anytime!
Agree with everything here. And I love the "human in the lead" framing. Going to borrow that!
Sounds like a world gone mad I’m afraid.
Create interactive graphics (I’m thinking of the toggles I think Manhattan Institute released recently allowing you to see the budget impacts of different policies so you can stack them in different combinations) or even video games that educate on the particular topic
You should literally ask Claude this question. Code can solve lots of problems, especially infused with intelligence. Software does not need to mean apps!
The "think of yourself as a manager" framing matches what I see leading an R&D team building voice agents. The skill that transfers is not prompting, it's the same judgment a manager needs: knowing when to trust a result versus when to check it closely, and building the checks that catch the difference before a customer does. Domain expertise seems to matter more than technical skill for exactly that reason - you know what wrong looks like faster.