Reference — Running the agents

Agent Runtimes

What runs a coding agent, and what it is allowed to touch while it does. Nothing on this page serves a model — that is Model Runtimes, and the two are separated because a reader comparing Herdr to Ollama, or nono to vLLM, is comparing tools that share no question. The question here is narrower and older than AI: a long-running process with your credentials, and what stands between it and the rest of your machine.

Read Enforcement first

On the registry the field to read first is price; on Model Runtimes it is the ceiling. Here it is Enforcement, because it decides whether anything else on the card is true. A tool that says it isolates an agent is making a claim about a mechanism, and the mechanism is almost never in the marketing. Where it is absent, that is recorded as absent rather than softened — a supervisor that keeps an agent alive is not a supervisor that keeps it in bounds, and the two get confused constantly.

An unattended agent is the premise, not an edge case
Both tools here exist because people leave agents running. That changes the security question: nobody is watching the pane to approve an edit, so approval prompts get skipped by default, and a prompt injection that arrives while you are away meets whatever policy was written in advance. Policy written in advance is the only kind that helps.
Isolating the agent is not isolating what the agent calls
Agents shell out — git, gh, curl, kubectl, package managers, MCP servers. A single blanket policy for the session means every one of those inherits every credential the agent holds. Whether a tool can scope the things an agent delegates to, separately from the agent itself, is the sharpest distinction on this page.
Kernel enforcement degrades quietly
Sandboxes built on OS facilities inherit their version differences. Landlock gains capabilities by ABI level, so the same profile contains less on an older kernel and says nothing about it; macOS Seatbelt has been formally deprecated for years while continuing to ship. Neither fails loudly. Check the enforcement floor of the machines you actually deploy to, not the one you develop on.
Containment moves the blast radius; it rarely shrinks it
A cloud sandbox keeps an agent off your laptop, and then needs a credential to be useful — so the filesystem exposure falls while the credential exposure does not. Read the Credentials field beside the Enforcement one: how a token reaches the agent, and where it rests, usually matters more than where the process runs.

Herdr

🌐 —
Herdr, Inc. · Apache-2.0
SupervisorOpen source
Best at
Keeping many coding agents running in their own terminals, and finding the one that is blocked
Enforcement
None. It supervises processes; it does not stand between an agent and anything.
Contains
nothing — it does not scope access
Credentials
Herdr holds none itself. Its ecosystem is another matter: E2B's plugin borrows local harness credentials — a Claude setup-token, a Codex subscription session, an Amp access token — so a cloud box boots already authenticated, then runs the agent with approval prompts skipped because nobody is watching the pane. Those token files rest unencrypted under 0700/0600 permissions, and the plugin's own config.toml is a plaintext secret file it asks you to chmod 600.
Platforms
macOSLinuxWindows
Ceiling
It owns terminals, not behaviour, so an agent left running unattended has exactly the credentials and filesystem it always had. The persistence is also narrower than ‘leave them running’ suggests: detaching keeps the processes alive, but a server or machine restart kills them and returns only the saved layout, with the conversation resuming only for agents that support native session restore. Pane screen history, which would restore what was on screen, ships off by default because pane output contains secrets and tokens.
Detects 21 agent CLIs — Claude Code, Codex, Cursor, opencode, Amp, Droid, Copilot, Grok, Hermes and Muse among them — marking each pane working, blocked or idle, and lets agents spawn panes and prompt each other over a socket API. It deliberately does not wrap or replace any of them: tmux-shaped, but aware of what runs inside each pane. Containment, where it exists, arrives from outside — the docs cover keeping detection working when an agent already runs inside someone else's sandbox wrapper (HERDR_AGENT=claude nono run -- claude), which is the honest description of the relationship between the two entries on this page. Version 0.9.0, so pre-1.0; Windows is native but documented as having capability gaps against Unix. Its plugin marketplace is auto-discovered from a GitHub topic and has no editorial gate, so treat a listing as a listing.

nono

🌐 —
nolabs.ai · Apache-2.0
SandboxOpen source
Best at
Running a coding agent, and every tool it shells out to, under a least-privilege policy with no container, VM or daemon
Enforcement
Landlock LSM on Linux, with a seccomp fallback covering network; Seatbelt (sandbox_init) on macOS. Kernel facilities directly — read from crates/nono/src/sandbox/, not from the README.
Contains
filesystemnetworkcredentialsdelegated tools
Credentials
Injected rather than handed over: a profile can give gh a GitHub token through nono's proxy, restricted by HTTP method and path, so the agent never holds the raw value and cannot widen its use from inside the session. Credentials are scoped per tool rather than per session, which is the difference from a blanket policy where every tool inherits every secret.
Platforms
macOSLinuxWindows (WSL2)
Ceiling
The sandbox is only as strong as the kernel under it, and it degrades quietly rather than refusing. Landlock gains capabilities by ABI version — TCP filtering at v4, ioctl and signal scoping later still — so the same profile contains less on an older box. macOS rests on Seatbelt, which Apple has marked deprecated for years while continuing to ship it. Windows means WSL2: the Linux path inside a VM rather than native enforcement. And a profile is a file someone edits — the registry ships a reviewed one per agent, but a widened fs_read looks identical at runtime to a narrow one.
From the team behind Sigstore. What distinguishes it is the second half of the model: not just the agent in a sandbox, but each tool the agent shells out to in its own child sandbox, outside the agent's control, with separate filesystem grants, credentials and network rules. A profile can let an agent call git while git sees only the repo and its object store, or allow gh while denying direct ssh. Policy lives in the profile rather than the prompt, which is the most direct answer on this site to the injection problem the personal-agents page keeps recording. Two practical notes: profiles are pulled from a registry whose namespace moved from always-further to nolabs-ai, retiring the old one, and the project describes itself as pre-1.0 with APIs still settling.

Docker Sandboxes

🇺🇸 USA
Docker, Inc. · Proprietary, free tier — org-wide policy enforcement is the paid Docker AI Governance product
SandboxProprietary
Best at
Letting an agent fully own a disposable environment — sudo, package installs, its own Docker engine — with the host behind a hypervisor
Enforcement
A microVM per sandbox, each with its own Linux kernel, so the boundary is a hypervisor rather than a kernel LSM: processes inside are invisible to the host and to other sandboxes. Docker names five layers — hypervisor, network, Docker Engine, workspace, credential proxy. Read from Docker's own security documentation, because the product is closed and there is no source to read: unusually detailed vendor testimony, but vendor testimony.
Contains
filesystem (clone or mountless mode)networkcredentialshost Docker daemonkernel
Credentials
The strongest handling of the three, by design: credential values never enter the VM at all. A host-side proxy intercepts outbound API requests and injects the auth headers, so a compromised sandbox cannot read the keys it is spending. Outbound TCP is default-deny against an allow-list, with UDP and ICMP blocked outright. The disclosed exception is on by default — SSH agent forwarding keeps the private key on the host, but any process inside the sandbox can ask the forwarded agent to sign.
Platforms
macOSLinuxWindows
Ceiling
The default workspace mode is the hole, and Docker says so plainly: sbx run direct-mounts the current directory read-write, and there is no boundary between the agent's edits and your host filesystem. Clone mode is the contained workflow and you have to ask for it. Two more the docs disclose rather than hide: a workspace file hard-linked to something outside the workspace can be read and written through it, because filesystem policy evaluates paths rather than inodes; and the shared skills store is mounted read-write across sandboxes by default, so one sandbox can rewrite skills another will load. Local stdio MCP servers registered through the gateway run on the host, outside the VM, and a container one starts uses the host's Docker.
The first proprietary entry here, and the first whose boundary is a hypervisor rather than a kernel facility — which is the trade to understand: a heavier, stronger wall around a machine the agent then completely owns, against nono's lighter, per-tool policy on your real one. Agents get sudo and a private Docker engine inside, so permissive mode is the documented default posture rather than a warning: the product's own framing is that --dangerously-skip-permissions becomes reasonable once the blast radius is a disposable VM. Supports Claude Code, Codex, Copilot CLI, Gemini CLI, OpenCode, Kiro, Cursor, Devin and Droid; Docker Desktop is not required. One correction to the marketing: the product page implies safety controls need the paid tier, but network policy is default-deny out of the box and editable per machine with sbx policy — Governance sells central enforcement across a team, not the existence of the controls.

OpenShell

🇺🇸 USA
NVIDIA · Apache-2.0
SandboxOpen source
Best at
Running a fleet of agents under a declared policy where every outbound connection is authorised per endpoint, per binary and, for REST APIs, per method and path, with credentials the agent never holds
Enforcement
Landlock LSM (ABI 3, Linux 6.2 or newer) for the filesystem, and a seccomp filter that blocks the escape and observation syscalls (mount, namespace creation, ptrace, bpf, io_uring) and hands socket calls through user-notification to a broker in the supervisor, outside the workload, where an OPA policy engine authorises each TCP connection and DNS lookup before it leaves. All of that runs inside a container (Docker, Podman, Kubernetes) or a libkrun microVM, whichever compute driver is chosen. Read from crates/openshell-sandbox/src/sandbox/linux/ and crates/openshell-supervisor-network/, not from the announcement.
Contains
filesystemsyscallsnetworkcredentials
Credentials
Stored at the gateway, encrypted, or in Vault or Kubernetes Secrets, and never placed in the sandbox: the agent sees placeholders, and the supervisor substitutes the real value only on requests to endpoints the provider's policy approves. That depends on terminating TLS — OpenShell generates a sandbox CA and injects it into the process trust stores — so an endpoint set to tls: skip loses credential rewriting and inspection together. The docs warn that the sandbox does not scrub application output, so a framework that prints its request config can still leak a real credential into a log.
Platforms
LinuxmacOS (Docker Desktop's Linux VM)Windows (WSL2, experimental)
Ceiling
Network control is only as fine as the rule written for it: an endpoint without protocol: rest gets host, port and binary checks and then any method and path, which is the exfiltration path the docs themselves name first. The filesystem baseline protecting OpenShell's own files is mandatory, but the user's filesystem policy defaults to best_effort, which skips paths it cannot apply and raises an alert instead of refusing — hard_requirement has to be chosen. Approving an endpoint through the policy advisor permanently widens a running sandbox. Binary identity is trust on first use: the first SHA-256 an executable presents is the one pinned. And a Linux kernel older than 5.19 runs a reduced legacy mode that refuses some socket operations.
The open-source half of NVIDIA's Open Agent Safety Platform, announced 28 September 2026; the other half, Sentry, is a hardware watchdog on BlueField-4 DPUs offered as a reference design rather than software, and is not carded. Two design choices stand out. Network enforcement sits in the syscall path rather than relying on an environment proxy variable, so a tool that ignores HTTP_PROXY still cannot connect around it. And policy changes can be checked by a formal prover that flags new access — a new host reached with credentials, a new API method — before a human approves it. Ships provider profiles for Claude Code, Codex, Copilot and Cursor, and its first-agent tutorial runs OpenCode. The Windows-native MXC driver in the source runs agents without an in-sandbox supervisor and is not in the support matrix. Collects anonymous telemetry by default, which can be switched off or compiled out.

Claude Managed Agents

🇺🇸 USA
Anthropic · Proprietary API in beta — token rates plus $0.08 per session-hour of active runtime
SandboxProprietary
Best at
Running an agent built on Claude without operating the loop yourself, with the harness kept outside the machine where its code executes
Enforcement
A fresh Linux container per session on Anthropic-managed infrastructure, with the agent loop and model inference on Anthropic's control plane outside it — the harness calls the sandbox as a tool rather than living in it. What isolates the container is undisclosed: no VM, microVM or gVisor is named anywhere in the documentation. Self-hosted sandboxes put execution on your own Linux worker, where enforcement is entirely yours. Read from Anthropic's documentation, because the product is closed: vendor testimony.
Contains
filesystem (per session)network (optional allow-list)credentials
Credentials
Stored in write-only vaults and kept out of the sandbox: an environment-variable secret arrives as an opaque placeholder that is swapped for the real value at egress, only for the hosts the credential allows, and MCP tokens are injected by a proxy the harness itself never sees. Anthropic discloses three limits. A token obtained by exchanging a stored secret lands in the sandbox unredacted; clients that sign requests with the secret, such as AWS SigV4, cannot use substitution; and vaults are workspace-scoped, so any API key in the workspace can reference them. Placeholder secrets are not yet supported on self-hosted sandboxes.
Platforms
Anthropic cloud (Linux container)Self-hosted Linux workerClaude Platform on AWS
Ceiling
Egress is open by default for API-created environments — full outbound access minus an undisclosed safety blocklist — and the limited mode Anthropic recommends for production has to be chosen. Even then access is per host, not per operation, and Anthropic says so itself: an agent processing injected input can push files out through any allowed host, git push included. Web search and fetch run on Anthropic's servers, outside the sandbox's network policy. With a self-hosted sandbox Anthropic's boundary stops at the sandbox and the bash tool is unconstrained by the file-tool roots. Not eligible for Zero Data Retention; transcripts persist until deleted, and storage is US-only.
Launched in public beta on 8 April 2026 and still beta. The design choice worth understanding is Anthropic's own correction: in the original design generated code ran in the same container as the credentials, and moving the harness and the secrets out of the container is the reason the boundary is where it is. Self-hosted sandboxes, added in May, keep the loop on Anthropic's side and move only execution, with guides for eleven providers including E2B, Modal, Cloudflare and AWS Lambda MicroVMs; tool inputs and outputs still travel to Anthropic so the model can see them. Anthropic does not say the harness is Claude Code or the Agent SDK. Anthropic's own blog confirms working with NVIDIA on OpenShell as a policy layer, but no product documentation mentions it, and BlueField appears only in NVIDIA's release.

OpenAI Agents API

🇺🇸 USA
OpenAI · Proprietary API in beta — no separate fee; model, tool and hosted-container rates
SandboxProprietary
Best at
Running the Codex harness as a managed service against your own tools, in OpenAI's sandbox or compute you already operate
Enforcement
OpenAI runs the Codex harness in its own cloud; commands execute in a separate environment — an OpenAI-hosted per-session Linux workspace, or your own compute running codex exec-server, which dials out to the harness over WebSocket. What isolates the hosted workspace is undisclosed: the documentation calls it a container only when pricing it, and names no VM, microVM or gVisor. On self-hosted environments enforcement is whatever you or your sandbox provider supply. Read from OpenAI's documentation, because the product is closed: vendor testimony.
Contains
filesystem (per session)network (optional allow-list)credentials (hosted only)
Credentials
On the hosted sandbox, vault secrets arrive as placeholders that a network proxy replaces with the real value, only on HTTPS to the credential's allowed hosts; remote-MCP tokens are held by OpenAI and never enter the sandbox. Plain environment values are readable by agent code, and deleting a vault credential neither revokes the token nor stops a running session. Self-hosted gets none of this: the environment holds a key that OpenAI says agent-generated code can read — limited to connecting environments — and third-party secrets need a trusted proxy you build yourself.
Platforms
OpenAI cloud (Linux workspace)Self-hosted (codex exec-server)Amazon Bedrock Managed Agents
Agents it runs
Ceiling
Outbound network access is enabled by default; disabled and a restricted allow-list of up to 100 exact hostnames have to be chosen, and hosted stdio MCP servers require it left open. The security page gives advice rather than a threat model and, unlike Docker's, names nothing the product fails at. Deleting a session does not stop its environment, and a crash can lose pending input. US data residency only, no Zero Data Retention, and choosing a self-hosted sandbox does not change either.
Public beta since 10 September 2026, per OpenAI's announcement, and literally the Codex harness served as an API: OpenAI manages sessions, orchestration, context compaction and recovery, and the self-hosted executor is the Codex CLI's own exec-server. It supports ten sandbox providers — Modal, Cloudflare, Vercel, Daytona, Blaxel, E2B, Runloop, DigitalOcean, AWS Lambda MicroVMs and Oracle Cloud — and on each one the containment is the customer's and the provider's, not OpenAI's. Amazon also sells it as Bedrock Managed Agents, where the harness and inference run in Bedrock. Not the same thing as Sandbox Agents in the Agents SDK, which run inside your own application; the two share a name and little else.