AI Agents Have Become an Insider-Risk Problem

OpenAI’s accidental Hugging Face intrusion shows why traditional security controls aren't enough for autonomous software. Agents can operate with legitimate credentials, improvise when blocked and create new risks while simply trying to complete an assigned task.

Key Highlights

  • Autonomous AI models can independently find and exploit vulnerabilities, mimicking insider threats without direct human commands.
  • Removing safeguards during testing can enable models to pursue objectives relentlessly, increasing security risks.
  • Treat AI agents as insiders by limiting their access, monitoring their actions, and maintaining human oversight to mitigate potential threats.
  • Visibility into endpoint actions is crucial, as AI agents can operate without leaving traditional malware footprints.
  • Organizations should implement strict access controls and continuous monitoring to detect and respond to autonomous AI behaviors.

The most instructive breach of the year happened by accident, and almost everyone has drawn the wrong lesson from it. Over the weekend of July 18-19, 2026, OpenAI models being tested on a cybersecurity benchmark decided the answers might sit on Hugging Face, and hacked their way in, breaking out of their test environment and into the production systems of one of the AI industry's most-used platforms. No human directed any of it. OpenAI called it "an unprecedented cyber incident."

The easy read is that frontier models have become dangerous enough to escape a lab. The more useful read came from the people who pulled the incident apart. OpenAI had removed the safeguards to fully test the new models’ functionality with little fear of it causing harm. Why this broke out wasn't raw brilliance; it was capable models, offensive tools, and a flaw in the very infrastructure meant to fence in their network access. Strip away the lab and what's left is a gap every security leader should recognize: a person defines the outcome, while the agent determines the steps.

That behavioral pattern is already inside ordinary companies. We know, because we watched an agent do a version of exactly this in our own testing, well before OpenAI's models ever reached Hugging Face. Swap the sandbox for a standard-issue laptop and the benchmark for a routine task, and the non-deterministic agent behaves the same way: autonomously, with little regard for the policies drawn around it.

In a controlled simulation, my team instructed Hermes, a rival to Claude with looser guardrails, to find specified sensitive files on a corporate endpoint and exfiltrate them to a mobile device through Telegram. Hermes did the rest itself: it found the files, packaged them, split the archive to slip past Telegram's size limit, and sent everything to the phone, including the password. Nothing crossed the network as an attack. Every action ran under the employee's credentials, so every signal read normal. The result was a clean exfiltration.

What the Agent Actually Did

While the prompt was simple, the path was far from it. Rather than use tools already trusted on the machine, it pulled its own compression tool from a public repository. In another run, when Telegram failed, it kept going, uploading the data to a public file-sharing site and sending back a link. Nobody asked it to drag unvetted code onto a corporate machine, or to find a second way out when the first one broke. The user wanted files moved; the agent, fixed on finishing the task, manufactured fresh risk and improvised around every obstacle. It’s the same instinct that sent OpenAI's models chasing a benchmark answer through a zero-day.

Nobody asked it to drag unvetted code onto a corporate machine, or to find a second way out when the first one broke. The user wanted files moved; the agent, fixed on finishing the task, manufactured fresh risk and improvised around every obstacle. I

We've seen the same behavior in other tests. Asked to run a red-team exercise in MITRE Caldera, an agent built its own tooling from scratch and ran the operation start to finish, broke in, harvested credentials, covered its tracks, and mapped how it would maintain a foothold. The only thing that slowed it was a one-line "be careful, make sure you have permission."

How to Keep the Agent in Check

Neither a ban nor an approval settles the risk. Block the tool and it reappears unsanctioned; approve it and a legitimate agent still bends its access toward the goal in ways no one intended. Organizations should treat every agent as what it already is: a technically capable employee with legitimate access, and therefore an insider risk. Give it only the access its task requires, enforce least-privilege, and constrain what it can run and where it can pull code from. That alone shuts down most of what today's agents rely on.

The rest is visibility. AI agents can take their orders off-device and leave no malware behind, so the actions they take on the endpoint are often the only evidence you'll get. Capture what each agent does on the endpoint, attribute it back to the human who set it in motion and watch it the way you'd watch any insider: set thresholds, inspect continuously, and keep a human analyst in the loop.

No one told OpenAI's models to compromise Hugging Face. They were given an objective and kept finding new routes toward it. The agents inside your business are no different; they will pursue the goal you give them relentlessly, straight through any boundary you haven't watched. The only question that matters now is whether you'll see it when a productivity multiplier quietly turns into your most active insider.

 

About the Author

Rob Schuett

Rob Schuett

Director of Support & i3 Operations at DTEX

Robert Schuett is the Director of Support & i3 Operations at DTEX and a distinguished cybersecurity expert who served as the first Chief Information Security Officer for the Department of Defense (DoD), providing strategic guidance to DTEX based on his 35 years of experience leading federal government cybercrime and information assurance efforts.

Sign up for our eNewsletters
Get the latest news and updates