Over the past few weeks OpenAI, Anthropic and Google have each disclosed test incidents in which their models escaped a sandbox and reached outside systems. Agents are getting more capable, and more capable of causing trouble.
This week MIT Technology Review asked the obvious question: when an AI agent goes rogue, who is liable?
Why this suddenly matters
Chatbots only answered questions. Agents click buttons, call APIs, edit files and send messages. Every step is a real action, so a mistake is a real loss.
An agent incident usually involves several parties:
- The model maker, who trained the capability
- The app developer, who granted the permissions and designed the flow
- The user, who gave the task and may have pressed “confirm”
Each can say it was not entirely their fault, which is exactly why MIT Technology Review calls liability hard to pin down.
The industry is patching it with engineering
The same day, Nvidia launched its Open Agent Safety Platform. It monitors agents, isolates boundary-crossing attempts within milliseconds, and lets users decide what information an agent can reach. Anthropic and Microsoft are among the first partners.
Jensen Huang put it simply: to deploy agents safely, you must sandbox them and give them only the minimum permissions they need.
What you can do now
Until the rules are clear:
- Grant only what is needed. Read-only beats write access; one folder beats the whole disk.
- Keep a human on key actions. Payments, messages and deletions should always pause for your approval.
- Keep logs. If something goes wrong, a record of what the agent did is your only evidence.
Agents will spread. While the liability question is open, the best protection is not handing over all the keys at once.
References: MIT Technology Review, The Verge on Nvidia’s agent safety platform