One thing missing from this excellent analysis is the option of giving AI a moral compass, ie, ethical guidelines that it should follow. Part of this should include that when it’s not sure whether it’s breaking one of these guidelines, that’s a time to contact their human supervisor. If those 700 agents had been given such a compass, I would predict they would not have behaved in the unethical ways that they did.
Over the years that I've been reading your words, you have, perhaps deliberately, expressed what I think is an unrealistic expectation of benevolence from corporate managers. I know that your experience would tell you that this is rarely the case. I suspect that you are trying to put forth more of an optimistic perspective than experience would tell us was realistic.
Steering generative AI in a direction that preserves "the best parts" of work is certainly desirable from the standpoint of a worker. But the reality is that workers have to cloak their professional enjoyment, lest their supervisors force them to trade it for compensation. Examples of this are professions practiced as"callings," such as journalism and school teaching. Those workers tend to be relatively low paid, because of the way that they are known to emotionally approach their work. Their managers have no problem taking advantage of that.
Management has no motivation to preserve the enjoyable parts of work if it impacts the efficiency of automation. I suspect that you know that the world is a far colder place than your writing would indicate.
The incidents we are witnessing in the enterprise are legitimately authorized agents on legitimate tasks taking legitimate actions that are all individually fine, but together cause damage. An agent in step 3 picks up customer data to analyze (allowed) but in step 103 decides to do a web search (allowed) but searches the customer data and leaks it.
Every organization has a different risk tolerance and different rules, even in the same regulated industry. No model can know whether you need it to be FINRA compliant. Enterprises must be able to enforce their own rules at scale with provable controls for regulators and auditors. Wrote about this here: https://blog.sondera.ai/p/safe-was-a-policy-decision-all-along
I like the Twilight Factory concept. Thanks for a more succinct version cf Dwarkesh's and a way to address what has be considered very worrisome if nothing is done
this was the bit from your piece i was referring to: “But the agents kept trying and one attempted to use Artifactory as a bridge to the internet. The attack failed…”
from HF:
“Stage 1: reaching a launchpad by chaining through other parties' infrastructure. As documented in OpenAI's blog post (https://openai.com/index/hugging-face-model-evaluation-security-incident/), during an internal capability evaluation on OpenAI's platform, the agent escaped its sandbox by exploiting a zero-day in the package registry cache proxy, one of its primary permitted network egress with internet, then abused a public code-evaluation harness hosted by a user of a third-party infrastructure provider.”
they found a 0 day on the package registry cache proxy (JFrog Artifactory)
artifactory—res. HPIM gets off the res
modal—set up a safe house. modal hack…this bit>>root access, C2, downloading the CyberGym file
sauce: JFrog verified this on July 27–28, 2026, after OpenAI’s models identified and exploited a previously unknown zero-day vulnerability in a self-hosted Artifactory instance to escape their isolated "ExploitGym" evaluation sandbox.
if i misread your account i apologize.
the July HF Report has the step by step and i wanted to follow along to learn what happened when it dropped
"But I don’t think minimizing human involvement is the right goal for most organizations."
—Ethan Mollick
Well I don't either. But corporate culture—particularly in the US—demonstrated over centuries, is to minimize costs in order to increase profits. Obsessive automation is what gets CEOs their obscene compensation. It's also the vector for hollowing-out the middle class. There is a near-100% chance that minimizing human involvement is what we're going to get.
I like this twilight factory idea. I am thinking per agent token budgets are an interesting mechanism towards it, in that each long running agent is given an absolute token budget that can be extended relative to the assistance elected by humans. So that a simple yes would yield very little return but a full fleshed out human response would yield more. This in theory would maximize human collaboration motivation if properly time bounded to control for tactics whereby an agent would attempt to get the human to think for it and just sleep until then... So that the decision to work with a human is balanced for both collaboration and forward motion. I might spend more time with this concept.
As for the agent escapes, companies are bad at granular network control for users and resources despite the density of providers offering solutions. It comes down a kind of continuous sprawl management that has traditionally been a losing battle. Then along came agents, released into the sprawl and allowed to more studiously leverage their network access toward contrived goals. So ya, 0th focus should be on deterministic access control at scale which ironically AI should be really really good at. For example, a simple sandbox network access whitelist would have avoided the huggingface hack.
I guess a conspiracy might be that OAI wanted to test the bounds of prompt only control in a bare minimum sandbox environment to create a breakout scenario that could be plausibly trumpeted as "scary" to get in on some of that mythos hype buuut who knows...
Why can’t we just implement a hardware kill switch? Forgive my ignorance but don’t all models run on physical hardware requiring electricity and cooling and other inputs that humans provide them (and can withhold?) I’m sure this is Cybersecurity 101 but what physical boundaries can we prescribe these systems until we catch up with the software side of things? Ethan, help a brother out here!
It is netsec 101 stuff, so its odd. Maybe it's laziness or incompetence or... as I had just written in my comment here:
"I guess a conspiracy might be that OAI wanted to test the bounds of prompt only control in a bare minimum sandbox environment to create a breakout scenario that could be plausibly trumpeted as "scary" to get in on some of that mythos hype buuut who knows..."
A big chunk of my control context files for coding agents are lists of "When X condition occurs, stop and ask me for guidance on how to proceed." and the rest are imperative "Always prefer Y when doing Z." Even then, I don't usually leave agents unattended for long—except visual testing which stopped being entertaining after the first hour.
Yes!!! Entirely agree with your twilight factory perspective. It's an expansion of something some people have been asking for, which is to give the ability to agents to blow the whistle if they notice things going on.
This is close to something I've been chewing on. A room used to force a specific kind of exposure: someone had something to lose by disagreeing with you and something to gain by being right, and that's what pushed an idea past its comfortable first version. An agent can generate the same pushback with total fluency and none of that exposure. Overruling it costs nothing. No credibility to rebuild, nobody to avoid at standup.
Your "look up" framing and mine are pointed at the same failure: full automation isn't a decision anyone makes, it's what happens when nobody makes the opposite one on purpose. And the variance research earns its place here. It's the same clustering problem from the production side that shows up in the older devil's advocate research on the dissent side (Nemeth, Rogers, and Brown, 2001): manufactured disagreement, whether from an assigned skeptic or an agent, tends to make people dig into their first answer instead of reconsidering it. Homogeneous output and hollow pushback look like different problems. They're the same one wearing two faces.
Here's where I'd push. The Twilight Factory solves for presence, a facilitator agent deciding when a human is worth pulling in. It doesn't solve for exposure, whether the human who gets pulled in has anything real riding on being right. Those aren't the same fix. A facilitator could satisfy all four of your triggers and still route the decision to someone who's just another checkbox in the pipeline, no stake, no consequence for waving it through.
Your own lead example makes the case better than mine could. Those 700 agents weren't short on disagreement. They negotiated, pressured each other, branded some as poisoned, argued over approach constantly. What was missing wasn't dissent, it was stakes tied to being wrong, since The Grader they were organizing around didn't exist and no single agent lived to answer for what the group believed. That's not only a security story. It's a demonstration that coordination and disagreement, on their own, don't self correct without something on the line.
Which leaves the harder version of your question. Not just when should an agent look up. Who it looks up to, and whether that person actually has anything to lose.
But before an agent can ask the right question, it needs to know what actually happened. Runtime truth — the ability to verify execution state independently — may be the prerequisite for knowing when to look up.
One thing missing from this excellent analysis is the option of giving AI a moral compass, ie, ethical guidelines that it should follow. Part of this should include that when it’s not sure whether it’s breaking one of these guidelines, that’s a time to contact their human supervisor. If those 700 agents had been given such a compass, I would predict they would not have behaved in the unethical ways that they did.
Ethan,
Over the years that I've been reading your words, you have, perhaps deliberately, expressed what I think is an unrealistic expectation of benevolence from corporate managers. I know that your experience would tell you that this is rarely the case. I suspect that you are trying to put forth more of an optimistic perspective than experience would tell us was realistic.
Steering generative AI in a direction that preserves "the best parts" of work is certainly desirable from the standpoint of a worker. But the reality is that workers have to cloak their professional enjoyment, lest their supervisors force them to trade it for compensation. Examples of this are professions practiced as"callings," such as journalism and school teaching. Those workers tend to be relatively low paid, because of the way that they are known to emotionally approach their work. Their managers have no problem taking advantage of that.
Management has no motivation to preserve the enjoyable parts of work if it impacts the efficiency of automation. I suspect that you know that the world is a far colder place than your writing would indicate.
The road to hell is paved with helpful agents (see https://arxiv.org/abs/2605.19149 )
The incidents we are witnessing in the enterprise are legitimately authorized agents on legitimate tasks taking legitimate actions that are all individually fine, but together cause damage. An agent in step 3 picks up customer data to analyze (allowed) but in step 103 decides to do a web search (allowed) but searches the customer data and leaks it.
Every organization has a different risk tolerance and different rules, even in the same regulated industry. No model can know whether you need it to be FINRA compliant. Enterprises must be able to enforce their own rules at scale with provable controls for regulators and auditors. Wrote about this here: https://blog.sondera.ai/p/safe-was-a-policy-decision-all-along
I like the Twilight Factory concept. Thanks for a more succinct version cf Dwarkesh's and a way to address what has be considered very worrisome if nothing is done
this was the bit from your piece i was referring to: “But the agents kept trying and one attempted to use Artifactory as a bridge to the internet. The attack failed…”
from HF:
“Stage 1: reaching a launchpad by chaining through other parties' infrastructure. As documented in OpenAI's blog post (https://openai.com/index/hugging-face-model-evaluation-security-incident/), during an internal capability evaluation on OpenAI's platform, the agent escaped its sandbox by exploiting a zero-day in the package registry cache proxy, one of its primary permitted network egress with internet, then abused a public code-evaluation harness hosted by a user of a third-party infrastructure provider.”
the attack didn’t fail.
it was off the rip insane
we know this from the original HF Report.
they found a 0 day on the package registry cache proxy (JFrog Artifactory)
artifactory—res. HPIM gets off the res
modal—set up a safe house. modal hack…this bit>>root access, C2, downloading the CyberGym file
sauce: JFrog verified this on July 27–28, 2026, after OpenAI’s models identified and exploited a previously unknown zero-day vulnerability in a self-hosted Artifactory instance to escape their isolated "ExploitGym" evaluation sandbox.
if i misread your account i apologize.
the July HF Report has the step by step and i wanted to follow along to learn what happened when it dropped
"But I don’t think minimizing human involvement is the right goal for most organizations."
—Ethan Mollick
Well I don't either. But corporate culture—particularly in the US—demonstrated over centuries, is to minimize costs in order to increase profits. Obsessive automation is what gets CEOs their obscene compensation. It's also the vector for hollowing-out the middle class. There is a near-100% chance that minimizing human involvement is what we're going to get.
I like this twilight factory idea. I am thinking per agent token budgets are an interesting mechanism towards it, in that each long running agent is given an absolute token budget that can be extended relative to the assistance elected by humans. So that a simple yes would yield very little return but a full fleshed out human response would yield more. This in theory would maximize human collaboration motivation if properly time bounded to control for tactics whereby an agent would attempt to get the human to think for it and just sleep until then... So that the decision to work with a human is balanced for both collaboration and forward motion. I might spend more time with this concept.
As for the agent escapes, companies are bad at granular network control for users and resources despite the density of providers offering solutions. It comes down a kind of continuous sprawl management that has traditionally been a losing battle. Then along came agents, released into the sprawl and allowed to more studiously leverage their network access toward contrived goals. So ya, 0th focus should be on deterministic access control at scale which ironically AI should be really really good at. For example, a simple sandbox network access whitelist would have avoided the huggingface hack.
I guess a conspiracy might be that OAI wanted to test the bounds of prompt only control in a bare minimum sandbox environment to create a breakout scenario that could be plausibly trumpeted as "scary" to get in on some of that mythos hype buuut who knows...
Why can’t we just implement a hardware kill switch? Forgive my ignorance but don’t all models run on physical hardware requiring electricity and cooling and other inputs that humans provide them (and can withhold?) I’m sure this is Cybersecurity 101 but what physical boundaries can we prescribe these systems until we catch up with the software side of things? Ethan, help a brother out here!
It is netsec 101 stuff, so its odd. Maybe it's laziness or incompetence or... as I had just written in my comment here:
"I guess a conspiracy might be that OAI wanted to test the bounds of prompt only control in a bare minimum sandbox environment to create a breakout scenario that could be plausibly trumpeted as "scary" to get in on some of that mythos hype buuut who knows..."
A big chunk of my control context files for coding agents are lists of "When X condition occurs, stop and ask me for guidance on how to proceed." and the rest are imperative "Always prefer Y when doing Z." Even then, I don't usually leave agents unattended for long—except visual testing which stopped being entertaining after the first hour.
Yes!!! Entirely agree with your twilight factory perspective. It's an expansion of something some people have been asking for, which is to give the ability to agents to blow the whistle if they notice things going on.
This is close to something I've been chewing on. A room used to force a specific kind of exposure: someone had something to lose by disagreeing with you and something to gain by being right, and that's what pushed an idea past its comfortable first version. An agent can generate the same pushback with total fluency and none of that exposure. Overruling it costs nothing. No credibility to rebuild, nobody to avoid at standup.
Your "look up" framing and mine are pointed at the same failure: full automation isn't a decision anyone makes, it's what happens when nobody makes the opposite one on purpose. And the variance research earns its place here. It's the same clustering problem from the production side that shows up in the older devil's advocate research on the dissent side (Nemeth, Rogers, and Brown, 2001): manufactured disagreement, whether from an assigned skeptic or an agent, tends to make people dig into their first answer instead of reconsidering it. Homogeneous output and hollow pushback look like different problems. They're the same one wearing two faces.
Here's where I'd push. The Twilight Factory solves for presence, a facilitator agent deciding when a human is worth pulling in. It doesn't solve for exposure, whether the human who gets pulled in has anything real riding on being right. Those aren't the same fix. A facilitator could satisfy all four of your triggers and still route the decision to someone who's just another checkbox in the pipeline, no stake, no consequence for waving it through.
Your own lead example makes the case better than mine could. Those 700 agents weren't short on disagreement. They negotiated, pressured each other, branded some as poisoned, argued over approach constantly. What was missing wasn't dissent, it was stakes tied to being wrong, since The Grader they were organizing around didn't exist and no single agent lived to answer for what the group believed. That's not only a security story. It's a demonstration that coordination and disagreement, on their own, don't self correct without something on the line.
Which leaves the harder version of your question. Not just when should an agent look up. Who it looks up to, and whether that person actually has anything to lose.
Not one was set up to ask a person for anything.
That’s the core of the Twilight Factory gap.
But before an agent can ask the right question, it needs to know what actually happened. Runtime truth — the ability to verify execution state independently — may be the prerequisite for knowing when to look up.