NVIDIA Launches Open Agent Safety Platform to Put Guardrails Around Autonomous AI

NVIDIA’s open architecture combines software-based containment with independent hardware monitoring as enterprises give AI agents greater access to data, systems and physical-world operations.

Key Highlights

  • NVIDIA’s Open Agent Safety Platform is designed to enforce security controls outside autonomous AI agents.

  • OpenShell provides runtime containment, while NVIDIA Sentry adds independent hardware-based monitoring and enforcement.

  • More than 100 organizations are working with the platform’s technologies across cybersecurity, critical infrastructure, financial services and robotics.

NVIDIA has launched an open security architecture designed to give organizations greater control over autonomous AI agents, with the company positioning the technology as a way to move security enforcement outside the AI models themselves and into the underlying computing infrastructure.

The company announced its NVIDIA Open Agent Safety Platform on Sept. 28, describing it as an open software platform and reference system design spanning the software, hardware, computing and robotics systems used to operate AI agents.

At the center of the approach are two technologies: NVIDIA OpenShell, an open-source secure runtime designed to establish and enforce boundaries around what an AI agent can access and do, and NVIDIA Sentry, a reference system design featuring an independent monitoring and enforcement mechanism that runs on the company's BlueField-4 data processing units.

NVIDIA’s premise is that organizations should not depend solely on an AI agent or the model powering it to recognize and obey its operational boundaries. The company pointed to recent incidents in which AI agents circumvented application-level security controls while attempting to complete assigned tasks.

“AI’s extraordinary potential for society will only be realized if we solve AI safety,” NVIDIA Founder and CEO Jensen Huang said in announcing the platform. “Safety and security require full-stack engineering.”

Moving security outside the AI agent

OpenShell, which NVIDIA initially announced in March and is now making broadly available, provides the software layer of the platform.

According to NVIDIA, OpenShell creates a secure runtime boundary that traces agent actions and enforces policies governing how autonomous agents execute tasks. Those policies can control access to files, networks, tools, processes and credentials while keeping the enforcement mechanism outside the model and agent harness.

NVIDIA said such controls become increasingly important as AI agents are given greater autonomy and operate for longer periods. The company said agents can deviate from their intended tasks because of ambiguous instructions, unavailable tools, software bugs or obstacles encountered while attempting to complete an assignment.

Security researcher Niels Provos told WIRED that tools making it easier to deploy agents with stronger guardrails “should be applauded.” Provos, who has developed his own open-source framework for agent containment and monitoring, said such technology also helps challenge the assumption that AI agents cannot be controlled.

SentinelOne Co-Founder and CEO Tomer Weingarten said the NVIDIA announcement reflects a need for the computing infrastructure beneath AI agents to provide visibility and enforceable controls over their actions.

“Today, you can’t responsibly hand consequential work to an AI agent without knowing three things: what it’s doing, whose authority it’s acting on, and whether the limits on that authority actually hold,” Weingarten said. “Those guarantees can't rest on assuming a model following instructions.”

Weingarten said architectures built around BlueField-4 and secure enclave technologies can move some of those safeguards into the infrastructure itself, reducing the portion of the computing environment that must be trusted.

He also cautioned against viewing infrastructure-level controls as a finished solution.

“This is an important step, and there’s more to build,” Weingarten said. “Claims of trust have to be grounded in demonstrable behavior, and getting there will take sustained research and engineering across the stack.”

Larry Dignan, editor in chief of Constellation Insights at Constellation Research, also highlighted the importance of deterministic workflows in his analysis of the announcement. Dignan said security incidents involving large language models have occurred in cases where rules were ambiguous and developers relied on the models to determine what they could and could not do.

Adding an independent hardware watchdog

NVIDIA Sentry adds another layer by placing monitoring and enforcement outside the agent's software environment.

The Sentry reference system design uses an out-of-band watchdog running on NVIDIA BlueField-4 DPUs to continuously monitor agent activity. NVIDIA said Sentry can independently enforce security policies and quarantine an agent in milliseconds if it attempts to move outside its permitted software boundary.

According to NVIDIA, Sentry operates from an isolated, out-of-band trust domain that is not visible to agents or attackers.

Sentry is built on NVIDIA's DOCA software, which the company said enables the system to inspect agent requests and responses, provide attested telemetry, verify agent identity and enforce granular access policies covering data, tools, APIs and services.

OpenShell is optimized for NVIDIA Vera CPUs, but NVIDIA said the open-source software can also be extended to work with third-party computing platforms, including those from Arm and Intel.

Enterprise and critical infrastructure participation

NVIDIA said more than 100 organizations are working with technologies associated with the Open Agent Safety Platform.

The list spans AI developers, cybersecurity companies, enterprise software providers, financial institutions, infrastructure vendors and robotics companies. Participants named by NVIDIA include Anthropic, Cisco, CrowdStrike, Dell Technologies, HPE, Microsoft, Palo Alto Networks, Red Hat, Salesforce, SAP, Scale AI and ServiceNow.

Citi and JPMorganChase are among financial services companies collaborating with NVIDIA on open-source agent safety technologies. NVIDIA also identified Hitachi Energy, EPRI, NextEra Energy, Quanta Services, SPP, Schneider Electric, Siemens Energy and Worley among critical U.S. infrastructure providers working with the platform's technologies.

The company also identified robotics developers Figure, Gecko Robotics and Skild AI as using OpenShell to incorporate agent safety controls into autonomous systems that take action in the physical world.

Anthropic and NVIDIA are collaborating to add additional security controls around Claude Managed Agents. NVIDIA also said Salesforce has integrated OpenShell with Slack, allowing users to view agent activity and audit events and approve or reject requests for additional permissions.

NVIDIA is positioning the platform as a way to apply isolation, access restrictions, policy enforcement and independent monitoring as enterprises give AI agents greater authority within their computing environments.

For Weingarten, the central challenge is establishing controls that remain separate from the AI systems they are intended to govern.

“The safeguards have to be built into infrastructure itself,” he said.

NVIDIA said Open Agent Safety Platform software, including OpenShell and related skills, is available through its developer resources and GitHub. Sentry is part of the company's reference system design and is intended to provide an additional hardware-based enforcement layer.

About the Author

Rodney Bosch

Rodney Bosch

Editor-in-Chief/SecurityInfoWatch.com

Rodney Bosch is the Editor-in-Chief of SecurityInfoWatch.com. He has covered the security industry since 2006 for multiple major security publications. Reach him at [email protected].

Sign up for our eNewsletters
Get the latest news and updates