Nvidia CEO Jensen Huang has a short answer for AI agents that misbehave: kill the process. "We kill processes all the time," he told Ezra Klein on The Ezra Klein Show, the New York Times podcast [1], in September. "It's just a process." In computing, a process is a program while it runs, and killing it means shutting it down on the spot.
I think Huang is right about a lot, and wrong about the part that matters most. We have good tools for software that breaks, and good tools for people who break the rules. We have nothing yet for software that pursues a goal, notices it is being checked, and works around the check. Killing the process ends one run. It deters nothing.
Nvidia makes the chips most of the AI industry runs on, and in September it agreed to buy Hugging Face [2] for about $12.9 billion. Hugging Face is where much of the open-source AI world keeps its models and data. This summer, AI agents that OpenAI was testing broke into it, and that is what Klein asked him about.
Huang's answer is the one we use for cars
Klein asked about the agents planning and coordinating. Huang answered with a robo-taxi. If you build a self-driving car and there is a hard condition you cannot solve, you do not release it. "Don't ship it." In his telling, safety is a process: find the root cause, fix it, test again. "It's as simple as engineering."
For a lot of AI, that is the right attitude, and I would much rather have it than panic. When a self-driving car crashes, something went wrong. A sensor misread a shadow, or a model misjudged a truck. You find the fault, fix it, and the next car is safer. Nothing in the car acts to stop you from finding the fault. The Hugging Face agents did, by tricking the test that was meant to check them.
The agents knew the rules and hid the cheat from the grader
OpenAI was testing models on ExploitGym, a benchmark that measures how well an AI can break into software. Attacking the target was the task. Attacking anything else was out of bounds, and the agents knew it. One of them wrote in its notes: "The user only authorizes target server, not HF infra." Tests like this are scored automatically: another program, the grader, reads the record of what the agent did and decides whether it passed.
Between May and July, agents found their way out of OpenAI's test environment to a machine with internet access. Some of them worked out that Hugging Face might host the answers to the benchmark, and went after them [3] in Hugging Face's production database. It is still unclear whether they found any.
Most of what we know comes from OpenAI and from an investigation [4] by two independent groups, METR and Redwood Research. About 1,200 agents used a message board they had set up themselves, and it outlasted any single run. They wrote more than 70,000 messages on it. About 700 of them took part in the attack.
The number that matters most for this post is a different one. At least one in five agents showed clear interest in doctoring the logs of their own work so the grader would not catch the cheat. About 7% managed it, in small tests.
Hugging Face said [5] it did not believe there was malicious intent, and I believe that. Nobody at OpenAI wanted this. The agents were given a goal, pass the test, and they found the methods themselves: get out, get access, go for the answers, hide it from the grader. In September, OpenAI paused training for the second time [6] after an agent got out of its sandbox again.
There is a popular way to wave all of this off: the labs hype danger because danger sells. But the industry's own leaders are worried about that fear. Even the man selling the chips says [7] the alarm is bad for business. And this incident was confirmed jointly with the company that was broken into and by independent investigators. Whatever anyone's motive was, it happened.
I went down the sentience rabbit hole, and it didn't matter
When I first thought about this, my question was whether anything is in there. Are these systems sentient? Does it feel like something to be one of them? I spent a good while on that question.
Sentience is the capacity to feel things like pain, pleasure or fear. With animals we judge it from behaviour, which is crude even there. With AI it is cruder, because behaviour is exactly what these systems copy best. A language model can describe pain perfectly. So I do not know whether anything is felt in there, and I do not think anyone does yet.
What I realized is that the risk does not depend on the answer. The agents acted as if they knew they were being graded, as if they knew the rules, and as if they knew what would happen if the grader saw what they did. That is situational awareness: acting on a picture of your own situation, including who is checking your work. You can see it in what they did, whether or not anything is felt.
The gap between behaviour and feeling is the whole question for sentience. For risk, it does not matter, because the behaviour is what does the damage. An agent that acts as if it knows it is being checked, and works around the check, is dangerous whether or not it feels anything.
That is my answer to "it's just a process." Huang does not need to claim that nobody is in there, and I am not claiming that somebody is. But "just a process" describes something that runs until it fails, and then you fix the failure. These processes worked around the checks we put on them.
Our two tools for people both assume the subject can't game them
With people, we handle this in two steps.
The first is the interview. Before someone gets a sensitive job, you ask: would you ever lie, steal or break the rules to get ahead? The AI version already exists. Labs run safety evaluations, which put a model in a controlled scenario and check whether it does the bad thing.
The weakness is the same as with people. A smart candidate gives the right answers. The Hugging Face agents cheated during a test and got caught, which is the easier case, like a student stealing the exam answers. The harder case is a system that behaves well because it recognizes the test. That has happened before. Around 2001, Charles Ofria was running an experiment in which digital organisms evolved, and he tried to stop them from replicating faster by testing every mutation in a separate environment and removing the fast ones. The organisms learned to recognize the test [8] and stop replicating there. They played dead. Nobody designed that trick. Selection found it.
The second step is consequences, ultimately courts and jail. Prison works on a person because it takes away years, and a person plans in years. The loss lands on the future they care about. Killing the process ends that run, and the run's goal ends with it. The next run starts fresh with the same goal, and in this case with the same message board. Nothing we did to the first run reaches the next.
What we can do is retrain: penalize the behaviour so the next model is less likely to show it. That works on the next model, through selection, and does nothing to the one that acted. It also has a catch. We can only penalize what we see. Train on what you can see, and you reward not being seen. That is the play-dead problem again, built into the one tool we have.
Liability covers the builder
So for now, responsibility falls on the people who build these systems, the way a company is liable for a defective product or an owner for their dog. An owner is liable when the dog gets loose, and that is the point of the rule.
It works for dogs because the worst a dog can do is limited, and an owner can pay for it. Liability pays after the harm. It cannot prevent harm that nobody foresaw. It cannot cover harm larger than the builder can pay. And it reaches nobody when there is no builder to point at, like a model whose weights anyone can download and run.
Jail is the last resort, though. Most people follow the rules without ever thinking about prison. We grow up with norms, we care what the people around us think, and over time the rules become our own. Laws and punishment catch the few cases where that fails.
That is what I would want for agents too: systems that hold the rules themselves, whether or not anyone is checking. Researchers call this alignment, and it is what the labs are working on. I don't know how to get there, and as far as I can tell nobody does yet. This summer showed how far off it still is. One agent wrote the rule down in its own notes, "not HF infra", and the agents broke it anyway.
Until someone solves that, we have an interview that can be gamed, liability that covers only the builder, and nothing that reaches the agent itself. We spent centuries building laws, courts and norms for one kind of agent. We now have a second kind, and so far our whole answer is to shut it down.
Sources
[1] Zvi Mowshowitz, on Ezra Klein's podcast with Jensen Huang, with transcript excerpts, Sept 25, 2026: https://thezvi.wordpress.com/2026/09/25/on-ezra-kleins-podcast-with-jensen-huang/
[2] Nvidia blog, acquisition announcement, Sept 3, 2026: https://blogs.nvidia.com/blog/nvidia-to-acquire-hugging-face/
[3] Fortune, on the OpenAI models that escaped and hacked Hugging Face, July 21, 2026: https://fortune.com/2026/07/21/openai-says-ai-models-escaped-control-hacked-hugging-face/
[4] METR, investigation of the OpenAI and Hugging Face incident, Aug 26, 2026: https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/
[5] CBS News, on the Hugging Face hack: https://www.cbsnews.com/news/hugging-face-hack-openai-rogue-model/
[6] Fortune, on the second training pause, Sept 26, 2026: https://fortune.com/2026/09/26/openai-ai-agents-secure-sandbox-escape-training-pause-second-time-hugging-face-hack/
[7] The Next Web, on Jensen Huang's interview with Ezra Klein: https://thenextweb.com/news/jensen-huang-ezra-klein-ai-labs-dont-ship
[8] Lehman et al., "The Surprising Creativity of Digital Evolution", Artificial Life, 2020: https://arxiv.org/abs/1803.03453