Discussion about this post

User's avatar
Josh Devon's avatar

The road to hell is paved with helpful agents (see https://arxiv.org/abs/2605.19149 )

The incidents we are witnessing in the enterprise are legitimately authorized agents on legitimate tasks taking legitimate actions that are all individually fine, but together cause damage. An agent in step 3 picks up customer data to analyze (allowed) but in step 103 decides to do a web search (allowed) but searches the customer data and leaks it.

Every organization has a different risk tolerance and different rules, even in the same regulated industry. No model can know whether you need it to be FINRA compliant. Enterprises must be able to enforce their own rules at scale with provable controls for regulators and auditors. Wrote about this here: https://blog.sondera.ai/p/safe-was-a-policy-decision-all-along

nAxis's avatar

I like this twilight factory idea. I am thinking per agent token budgets are an interesting mechanism towards it, in that each long running agent is given an absolute token budget that can be extended relative to the assistance elected by humans. So that a simple yes would yield very little return but a full fleshed out human response would yield more. This in theory would maximize human collaboration motivation if properly time bounded to control for tactics whereby an agent would attempt to get the human to think for it and just sleep until then... So that the decision to work with a human is balanced for both collaboration and forward motion. I might spend more time with this concept.

As for the agent escapes, companies are bad at granular network control for users and resources despite the density of providers offering solutions. It comes down a kind of continuous sprawl management that has traditionally been a losing battle. Then along came agents, released into the sprawl and allowed to more studiously leverage their network access toward contrived goals. So ya, 0th focus should be on deterministic access control at scale which ironically AI should be really really good at. For example, a simple sandbox network access whitelist would have avoided the huggingface hack.

I guess a conspiracy might be that OAI wanted to test the bounds of prompt only control in a bare minimum sandbox environment to create a breakout scenario that could be plausibly trumpeted as "scary" to get in on some of that mythos hype buuut who knows...

6 more comments...

No posts

Ready for more?