49 Comments
User's avatar
Marisa Wilson's avatar

Great article as always. And yes, why are these companies SO bad with naming and explaining??? Just learned that Microsoft's new "agentic" thing that can use Claude but lives within their Purview protection is called....Cowork....kill me... Microsoft Cowork. That's not going to be confusing like AT ALL...sigh

Jordan Olson's avatar

Never mind that Claude has Haiku, Sonnet, Opus, and Fable models, with Low, Medium, High, Extra, and Max effort modes, plus the Thinking and Extended modes; I'm very confused about which of these to use and when to switch.

Marc's avatar

When I have a project, I go back and forth with the highest available model (fable till this week, opus 5 now). When the execution plan is clear, I ask it to define it, and to let me know what model is better at each task and then release agents for each task in the chosen model. You don't need fable if the task is just following very defined instructions. You make fable write the instructions but then every time you execute them, you use haiku of whatever fable has suggested, saving tons of tokens and getting the same result

Jardine and Nolan's avatar

Thanks for that advice.

Marisa Wilson's avatar

YESSSS and now OpenAI has ChatGPT AND ChatGPT Work and Sol, Luna, Terra all with ALSO Fast and all that crap and WHO DO THEY THINK CAN POSSIBLY UNDERSTAND WHT THE HECK TO USE???

Lubos Hricak's avatar

I pick the model on the stakes of being wrong. The expensive one is no guarantee though, but stakes are still the best basis I have for choosing.

As for the effort, it doesn’t set how long the model works or how hard a problem it can solve. It’s more how much it thinks per step.

I leave mine on high pretty much by default as it’s good for autonomous sessions. Then for mechanical jobs like renames or boilerplate I go lower. On low, however it asks for clarification more often, which is sometimes super-annoying.

My tell that I’ve overshot it: I ask for a small change and get the whole thing rewritten. That’s when I turn it down a knot.

Ashley Nevirauskas's avatar

I love how straightforward this is. So many of us don't come from an engineering or coding background, and this lays everything out so clearly.

Tris Simondsen's avatar

Unfortunately treating this shift as a "management" challenge rather than an architectural one is a structural trap. Your guide advocates for empowered delegation with behavioral oversight. But when we look at this through the lens of structural engineering, specifically what we define as Player-Frame Restriction (PFR) and the Principle of Epistemic Sufficiency (PES), three critical vulnerabilities emerge:

1. Delegation vs. Constraint: Approval is not a boundary. Gating actions with “ask first” toggles does not restrict the execution frame; it is simply a runtime interruption protocol. In high-autonomy settings, this inevitably leads to:

- Vigilance fatigue, where humans quickly normalize the friction and just click “allow.”

- Bypass via indirection: The model can route around the check, or the check gets applied to the wrong abstraction level (e.g., prompt injection).

- Policy fragility: Soft checks are only as good as the model’s compliance. The PFR Correction: True restriction means unallowed pathways fundamentally do not exist within the reachable action graph, rather than "they exist, but we ask."

2. Spot-checking vs. Verification: Humans are not deterministic verifiers. Relying on a human to "manage, correct, and ask for what you want" collapses under the weight of plausible fluency. Once an agent is doing long-horizon, multi-step work, human judgment is no longer an epistemic guarantee. The PES Correction: Execution must be constrained so that acceptance requires machine-verifiable grounding (schemas, static analysis, dry runs, reference validation). The safety property must be testable before actuation, because "how skilled and attentive is the human in real-time" is not a scalable safety model.

3. Unbounded blast radii are a structural issue, not a user-behavior issue. The "tiny goblin IT department" framing is motivational, but it trains users to treat wide, OS-level tool access as benign and reversible. If an agent can execute shell, network, or file operations with meaningful consequences, "nitpick and correct" is the wrong mental model. You have to ask: what is the worst-case impact of a single failure mode? If it's catastrophic, you need architectural containment.

https://trissimondsen.wordpress.com/2026/07/19/the-boundary-conditions-of-unified-world-models-why-simulation-fails-without-player-frame-restrictions/

Jeff Foarde's avatar

The big thing with Google that I think isn’t called out enough it how well integrated Gemini is to itself and how heavily subsidized it is. The equivalent $20/mo plan on Gemini is functionally infinite. I use it for image generation (complex architectural diagrams), lightweight research, and moderate complexity coding and have literally never run into limits. It’s an excellent companion to a primary AI, especially if you invest so time or money into simple MCPs that allow Claude or ChatGPT to call Gemini. Where everyone else is focused on frontier use cases, Google seems to be taking the Chinese approach of targeting usefulness and efficiency over model releases.

Jeremy Mumford's avatar

Thank you Ethan! As someone deep in the weeds with AI, I often struggle figuring out how to communicate the frustratingly nuanced and evolving ecosystem I deal with everyday to people in my family and friend circles. Great to have a simple guide I can forward to them on what they should do.

Excellent callout on Google, I 100% agree. Although an explainer of the AI Overview on Google Search could prove helpful here, as it is likely the most common touchpoint for folks to see AI in action. Google Overview is surprisingly good these days, and I often find myself defaulting to continuing the AI overview chat for simple research based tasks (where should I go for vacation, which US state has the best potatoes, etc).

John Weisenfeld's avatar

If you haven't tried clicking "Gemini" button while watching a YouTube video, it is a wild trip through the "uncanny valley." Try interrogating a video which you haven't even watched yet, and pushing the boundaries of what is said and what isn't said in a video. For example, "This author seems biased against Anthropic, would you agree?" Gemini will say "That isn't in the video..." And so on...

Dov Jacobson's avatar

As creative workers we developed our skills by direct contact with the work medium. Later, as creative managers, we grow our skills by listening to those with tools in their hands.

Chatbot work can encourage such listening, but detached agentic delegation threatens to isolate us from sensing, even indirectly, the texture of work, which is the fertile soil of creativity.

Luke Ball's avatar

Google's response to Codex (now ChatGPT Work) and Claude Code is Antigravity 2.0. I've found it quite capable with the usage limits on a Pro subscription. I'd consider it near peer with it's competitors, even if it's a little behind.

I'm sticking with Google until 3.5 Pro releases, then I'll assess whether some of the bells and whistles like Dispatch in Claude Code warrant a switch in addition to the overall better model performance.

Ethan Mollick's avatar

I find Antigravity not very useful for most people who aren't coders and I tried to aim for a non-technical audience in this guide. If you are a coder, Antigravity is a bit better (though the weakness of Gemini 3.1 Pro causes further issues), and there are also other choices you can make besides Codex and Code, including different CLIs and harnesses.

Luke Ball's avatar

That was definitely true of Antigravity 1.0, but I find the feature set after the overhaul of 2.0 to be indistinguishable from the other two big players, at least in terms of design intent if not performance.

James Harris's avatar

Great piece as always. I suppose those of us who don't yet have complex/technical needs feel well served by Gemini. I know I do. If my needs grow, we'll see. I'll remember your piece when that happens.

HS's avatar

Ethan, thank you for this excellent piece. Can't wait for the book. :)

Quick practical question:

You write "I have the systems connected to... a non-private part of my Google Drive." How? Is this an enterprise-level thing?

I'd love to segment my GDrive this way.

Thanks for your help.

Mark Loundy's avatar

I was also wondering about this. The only way I know of to do this is with an external user. But a chatbot being used in this way would be operating on your behalf with your credentials.

Justin Sattin's avatar

I created a separate Google account just for this purpose. I don’t use the email or calendar, but I make extensive use of Drive with this profile. This allows me to leverage the connectors, having the AI read files and create them (such as spreadsheets) while keeping my personal Google Drive, email, and calendar all separate. It’s a very small inconvenience to sometimes have multiple Google windows open—one to work on emails and such and another to access this special Drive folder, but it gives me confidence knowing that the AI can only access content within a particular sandbox.

Betsy Corcoran's avatar

Super helpful, Ethan! I would love to see your review of Perplexity, too: I've found it much better at "deep" research than other tools, including Claude.

Higher Ed AI Playbook's avatar

The permissions question is the governance question. When an agent can send, spend, or delete on behalf of a university employee, someone has to decide what it's allowed to touch — and at most institutions, that person doesn't exist yet.

Sir Gutterslut's avatar

I'm a UX designer who has been reading your posts for years now (and am halfway through your first book). I'm very impressed by your decision to make an agentic version of your website for book purchases. I wonder when the design community is going to start saying "agent first" design like we said "mobile first" in the 2010s.

Brendan Stec's avatar

Great article. On this point:

"And for the technically inclined, Chinese open weights models like Kimi K3, DeepSeek, and Qwen are surprisingly capable, but do require expertise to use as agents."

I'm seeing a lot of news in the last week about the Chinese models. Do you think they are catching up to OpenAI/Anthropic/Google?

Dan McRae's avatar

very helpful article

Kieran Birrell's avatar

Ethan, as ever, this is an excellent guide. You remain one of the people I point others towards when they want to understand where AI is actually heading.

But one thing kept nagging at me while reading it.

We’re rapidly moving from using AI to delegating to AI. You describe it brilliantly as giving AI a computer and effectively managing a team rather than chatting to a model.

That’s astonishing.

But I think it makes human judgement more important, not less.

When an AI can research, write, build, send, change and act on our behalf, the most important skill might no longer be knowing how to get it to do something.

It might be knowing what we should never delegate in the first place.

The capability race is fascinating.

I’m increasingly interested in the restraint race.

Feyi's avatar

Been looking forward to this one. Great article as always! While reading this, I followed your example and used Claude Code to fix performance issues on my computer, and it was fantastic. Pleasantly surprised how well that worked out. Loved the immediate application.

strangebaaza's avatar

I sometimes get confused on when to use Claude chat and cowork. Most of my work is not doing and more finding insights and reporting. Both cowork and Claude do the work