Repo risk ratings
29 repositories received a full written assessment using the same procedure: establish identity and state, hunt for recurring failure families, cluster complaints, look outside the repository for firsthand reports, measure maintainer behaviour, normalise against apparent usage, and write down the reasoning. No numeric score, no composite, no fake precision.
The five categories
There is a sixth label, "not an isolation layer", which is deliberately not a category: it marks software that is useful and does not contain the agent at all.
Recommended / Mature
Appears suitable for serious use based on available evidence: an organisation behind it, a release cadence, longitudinal use at scale, and complaints proportionate to real adoption.
Good, but test it first
Strong enough to investigate and deploy, with meaningful caveats: platform asymmetry, a young component, a default you must change, or a CVE history that is disclosed rather than absent.
Experimental / Early
Interesting technology, currently risky to depend on. Frequently genuinely good engineering with insufficient longitudinal evidence, or an explicit beta label the maintainers are honest about.
Code / ideas worth studying
Do not build directly on the repository, but parts of its implementation or documentation are extremely useful. This is the category for "the idea is better than the dependency".
Probably skip for now
The expected hassle and risk look greater than the likely benefit — usually because development stopped, there is no release, or the licence blocks the reuse you were imagining.
Not an isolation layer
Useful software that does not contain the agent. Included so that the distinction between "gives the agent access" and "contains the agent" stays visible.
Raw issue counts are close to meaningless in this space. A project with tens of thousands of users generates complaints as a function of usage; a project with two hundred stars looks pristine because almost nobody has tried it. Every assessment here therefore asks the same questions: does the same class of failure recur across versions, do users need destructive workarounds, how fast and how completely did maintainers respond, and are the same symptoms reported independently on the issue tracker and elsewhere. Where a trustworthy denominator does not exist, the assessment says so rather than inventing one.
| Project | Rating | Family | ★ | Last push | Open issues |
|---|---|---|---|---|---|
| denoland/deno | WASM / language isolate | 108k | 2026-09-17 | 1.6k | |
| moby/moby | Container / namespace | 72k | 2026-09-19 | 3.9k | |
| mitmproxy/mitmproxy | Network / egress control | 45k | 2026-09-10 | 491 | |
| aquasecurity/trivy | Adjacent tooling | 37k | 2026-09-18 | 271 | |
| firecracker-microvm/firecracker | MicroVM / lightweight VMM | 36k | 2026-09-18 | 100 | |
| tailscale/tailscale | Network / egress control | 36k | 2026-09-19 | 4.6k | |
| hashicorp/vault | Secrets & credential brokering | 36k | 2026-09-18 | 1.4k | |
| utmapp/UTM | Full VM / hypervisor | 35k | 2026-09-19 | 1.1k | |
| podman-container-tools/podman | Container / namespace | 32k | 2026-09-19 | 1k | |
| goharbor/harbor | Adjacent tooling | 29k | 2026-09-18 | 902 | |
| gitleaks/gitleaks | Adjacent tooling | 29k | 2026-09-09 | 483 | |
| trufflesecurity/trufflehog | Adjacent tooling | 27k | 2026-09-18 | 559 | |
| cilium/cilium | Network / egress control | 25k | 2026-09-19 | 1.1k | |
| getsops/sops | Secrets & credential brokering | 23k | 2026-09-18 | 447 | |
| lima-vm/lima | Full VM / hypervisor | 21k | 2026-09-19 | 526 | |
| containerd/containerd | Container / namespace | 21k | 2026-09-19 | 471 | |
| google/gvisor | Application kernel | 19k | 2026-09-19 | 814 | |
| bytecodealliance/wasmtime | WASM / language isolate | 18k | 2026-09-18 | 848 | |
| coder/coder | Hosted sandbox API / CDE | 15k | 2026-09-19 | 1k | |
| kubernetes-sigs/kind | Container / namespace | 15k | 2026-09-04 | 242 | |
| qemu/qemu | Full VM / hypervisor | 13k | 2026-09-19 | 0 | |
| opencontainers/runc | Container / namespace | 13k | 2026-09-18 | 339 | |
| anchore/grype | Adjacent tooling | 12k | 2026-09-18 | 401 | |
| simonw/llm | Adjacent tooling | 12k | 2026-09-08 | 699 | |
| krallin/tini | Adjacent tooling | 11k | 2025-05-08 | 45 | |
| canonical/multipass | Full VM / hypervisor | 9.2k | 2026-09-18 | 418 | |
| containers/bubblewrap | Process sandbox (OS primitives) | 8.8k | 2026-09-18 | 196 | |
| cloudflare/workerd | WASM / language isolate | 8.7k | 2026-09-19 | 730 | |
| bats-core/bats-core | Adjacent tooling | 6.3k | 2026-09-16 | 129 | |
| cloud-hypervisor/cloud-hypervisor | MicroVM / lightweight VMM | 6.2k | 2026-09-19 | 228 | |
| lxc/incus | Full VM / hypervisor | 6.2k | 2026-09-18 | 46 | |
| devcontainers/spec | Container / namespace | 5.7k | 2026-03-20 | 190 | |
| ossf/scorecard | Adjacent tooling | 5.7k | 2026-09-18 | 457 | |
| lxc/lxc | Container / namespace | 5.3k | 2026-09-07 | 151 | |
| google/security-research | Adjacent tooling | 4.6k | 2026-09-18 | 91 | |
| Yelp/detect-secrets | Adjacent tooling | 4.6k | 2026-04-02 | 183 | |
| containers/crun | Container / namespace | 4.1k | 2026-09-17 | 49 | |
| google/nsjail | Process sandbox (OS primitives) | 4.1k | 2026-08-27 | 39 | |
| ddev/ddev | Container / namespace | 3.9k | 2026-09-19 | 176 | |
| opencontainers/runtime-spec | Container / namespace | 3.7k | 2026-04-24 | 94 | |
| buildpacks/pack | Adjacent tooling | 3k | 2026-09-14 | 202 | |
| seccomp/libseccomp | Process sandbox (OS primitives) | 934 | 2026-07-01 | 58 | |
| containers/conmon | Container / namespace | 500 | 2026-09-18 | 27 | |
| google/minijail | Process sandbox (OS primitives) | 384 | 2026-09-15 | 0 | |
| microsoft/artifacts-keyring | Secrets & credential brokering | 45 | 2026-08-19 | 9 | |
| anomalyco/opencode | Agent sandbox wrapper | 208k | 2026-09-19 | 6k | |
| anthropics/claude-code | Agent sandbox wrapper | 146k | 2026-09-19 | 12k | |
| openai/codex | Agent sandbox wrapper | 125k | 2026-09-19 | 17k | |
| browser-use/browser-use | Computer-use / desktop automation | 115k | 2026-09-18 | 452 | |
| google-gemini/gemini-cli | Agent sandbox wrapper | 107k | 2026-09-19 | 841 | |
| modelcontextprotocol/servers | Agent orchestrator / GUI | 90k | 2026-09-03 | 547 | |
| aaif-goose/goose | Agent sandbox wrapper | 54k | 2026-09-19 | 395 | |
| Aider-AI/aider | Agent sandbox wrapper | 49k | 2026-05-22 | 1.9k | |
| ziglang/zig | Adjacent tooling | 43k | 2025-11-27 | 2.8k | |
| microsoft/WSL | Windows containment | 33k | 2026-09-19 | 981 | |
| openai/openai-agents-python | Agent orchestrator / GUI | 29k | 2026-09-18 | 93 | |
| Infisical/infisical | Secrets & credential brokering | 29k | 2026-09-19 | 793 | |
| keepassxreboot/keepassxc | Secrets & credential brokering | 28k | 2026-09-18 | 901 | |
| QwenLM/qwen-code | Agent sandbox wrapper | 27k | 2026-09-19 | 1.5k | |
| trycua/cua | Computer-use / desktop automation | 24k | 2026-09-19 | 1k | |
| wasmerio/wasmer | WASM / language isolate | 21k | 2026-09-19 | 267 | |
| e2b-dev/E2B | Hosted sandbox API / CDE | 13k | 2026-09-19 | 72 | |
| gitpod-io/gitpod | Hosted sandbox API / CDE | 13k | 2026-09-18 | 450 | |
| microsoft/wslg | Windows containment | 11k | 2026-07-06 | 719 | |
| microsoft/WSL2-Linux-Kernel | Windows containment | 10k | 2026-08-01 | 134 | |
| falcosecurity/falco | Policy, approval & audit | 9.4k | 2026-09-18 | 42 | |
| kata-containers/kata-containers | MicroVM / lightweight VMM | 8.8k | 2026-09-19 | 1.2k | |
| draios/sysdig | Network / egress control | 8.3k | 2026-04-13 | 117 | |
| wasm3/wasm3 | WASM / language isolate | 8k | 2026-09-19 | 19 | |
| netblue30/firejail | Process sandbox (OS primitives) | 7.7k | 2026-09-16 | 523 | |
| openbao/openbao | Secrets & credential brokering | 7.4k | 2026-09-18 | 320 | |
| rancher-sandbox/rancher-desktop | Container / namespace | 7.4k | 2026-09-18 | 1.1k | |
| openai/tart | Full VM / hypervisor | 6.8k | 2026-09-16 | 73 | |
| anthropics/sandbox-runtime | Process sandbox (OS primitives) | 5.3k | 2026-09-19 | 212 | |
| cilium/tetragon | Policy, approval & audit | 5k | 2026-09-19 | 283 | |
| nolabs-ai/nono | Secrets & credential brokering | 4.1k | 2026-09-18 | 213 | |
| dagger/container-use | Agent orchestrator / GUI | 4k | 2026-09-14 | 62 | |
| nestybox/sysbox | Container / namespace | 3.9k | 2026-09-15 | 213 | |
| virtio-win/kvm-guest-drivers-windows | Full VM / hypervisor | 2.7k | 2026-09-15 | 180 | |
| libkrun/libkrun | MicroVM / lightweight VMM | 2.7k | 2026-09-16 | 98 | |
| superfly/flyctl | Hosted sandbox API / CDE | 1.7k | 2026-09-18 | 221 | |
| google/crosvm | MicroVM / lightweight VMM | 1.3k | 2026-09-18 | 2 | |
| lahfir/agent-desktop | Computer-use / desktop automation | 1.3k | 2026-09-17 | 22 | |
| step-security/harden-runner | Adjacent tooling | 1.3k | 2026-08-31 | 56 | |
| RchGrav/claudebox | Agent sandbox wrapper | 1.2k | 2026-09-17 | 18 | |
| cloudflare/sandbox-sdk | Hosted sandbox API / CDE | 1.1k | 2026-09-19 | 45 | |
| trailofbits/claude-code-devcontainer | Agent sandbox wrapper | 943 | 2026-08-28 | 4 | |
| firecracker-microvm/firecracker-go-sdk | MicroVM / lightweight VMM | 673 | 2026-02-10 | 53 | |
| modal-labs/modal-client | Hosted sandbox API / CDE | 514 | 2026-09-19 | 26 | |
| docker/sbx-releases | MicroVM / lightweight VMM | 389 | 2026-09-18 | 330 | |
| archlinux/arch-boxes | Full VM / hypervisor | 265 | 2025-12-17 | 6 | |
| imbue-ai/sculptor | Agent orchestrator / GUI | 233 | 2026-09-19 | 20 | |
| vercel/sandbox | Hosted sandbox API / CDE | 201 | 2026-09-19 | 28 | |
| omnigent-ai/omnigent | Agent orchestrator / GUI | 10k | 2026-09-19 | 1.4k | |
| NVIDIA/OpenShell | Agent sandbox wrapper | 8.7k | 2026-09-19 | 530 | |
| superradcompany/microsandbox | MicroVM / lightweight VMM | 8.3k | 2026-09-19 | 93 | |
| smol-machines/smolvm | MicroVM / lightweight VMM | 6.1k | 2026-09-19 | 82 | |
| hyperlight-dev/hyperlight | MicroVM / lightweight VMM | 4.7k | 2026-09-18 | 220 | |
| nanovms/nanos | Full VM / hypervisor | 3.2k | 2026-09-19 | 87 | |
| boxlite-ai/boxlite | MicroVM / lightweight VMM | 2.3k | 2026-09-19 | 294 | |
| earendil-works/gondolin | MicroVM / lightweight VMM | 2.2k | 2026-07-06 | 47 | |
| eugene1g/agent-safehouse | Agent sandbox wrapper | 2.1k | 2026-09-18 | 24 | |
| e2b-dev/desktop | Computer-use / desktop automation | 1.5k | 2026-09-18 | 11 | |
| microsoft/mxc | Windows containment | 1.3k | 2026-09-19 | 84 | |
| clawkwork/clawk | MicroVM / lightweight VMM | 1k | 2026-08-13 | 4 | |
| fencesandbox/fence | Agent sandbox wrapper | 969 | 2026-09-10 | 34 | |
| instavm/coderunner | MicroVM / lightweight VMM | 894 | 2026-08-13 | 9 | |
| Katakate/k7 | MicroVM / lightweight VMM | 808 | 2026-09-14 | 0 | |
| Katakate/k7 | MicroVM / lightweight VMM | 808 | 2026-09-14 | 0 | |
| finbarr/yolobox | Agent sandbox wrapper | 643 | 2026-08-28 | 2 | |
| jingkaihe/matchlock | Agent sandbox wrapper | 621 | 2026-07-26 | 9 | |
| manuelschipper/nah | Policy, approval & audit | 486 | 2026-09-15 | 0 | |
| nikvdp/cco | Agent sandbox wrapper | 425 | 2026-09-12 | 2 | |
| webcoyote/sandvault | Agent sandbox wrapper | 417 | 2026-09-14 | 1 | |
| landlock-lsm/island | Process sandbox (OS primitives) | 334 | 2026-05-26 | 21 | |
| GreyhavenHQ/greywall | Agent sandbox wrapper | 298 | 2026-08-13 | 24 | |
| eqtylab/cupcake | Policy, approval & audit | 293 | 2026-03-02 | 22 | |
| nanvix/nanvix | MicroVM / lightweight VMM | 282 | 2026-09-19 | 244 | |
| butter-dot-dev/bVisor | Application kernel | 217 | 2026-02-23 | 1 | |
| kstenerud/yoloai | Agent sandbox wrapper | 212 | 2026-08-21 | 9 | |
| webcoyote/clodpod | Full VM / hypervisor | 187 | 2026-08-30 | 4 | |
| postrv/forgemax | Agent sandbox wrapper | 151 | 2026-09-10 | 3 | |
| jgbrwn/vibebin | Container / namespace | 108 | 2026-08-10 | 0 | |
| LuD1161/agentjail | Agent sandbox wrapper | 94 | 2026-09-08 | 2 | |
| swelljoe/flar | Agent sandbox wrapper | 55 | 2026-08-18 | 6 | |
| pjlsergeant/byre | Agent sandbox wrapper | 32 | 2026-09-17 | 1 | |
| agentic-dev3o/sandbox-shell | Agent sandbox wrapper | 30 | 2026-09-18 | 1 | |
| recodelabs/lima-devbox | Full VM / hypervisor | 28 | 2026-07-03 | 1 | |
| corv89/shannot | Policy, approval & audit | 26 | 2026-04-13 | 6 | |
| binwiederhier/sandclaude | Agent sandbox wrapper | 23 | 2026-05-30 | 0 | |
| agentcage/agentcage | Agent sandbox wrapper | 22 | 2026-09-19 | 29 | |
| jskswamy/aide | Agent sandbox wrapper | 17 | 2026-09-17 | 2 | |
| tarsgate/skynot | Agent sandbox wrapper | 16 | 2026-06-18 | 6 | |
| ashishgituser/bunkervm | Agent sandbox wrapper | 1 | 2026-08-18 | 0 | |
| jart/cosmopolitan | WASM / language isolate | 21k | 2026-07-20 | 228 | |
| agent-infra/sandbox | Hosted sandbox API / CDE | 6k | 2026-09-14 | 72 | |
| unikraft/unikraft | Full VM / hypervisor | 3.9k | 2026-09-19 | 367 | |
| M2Team/NanaBox | Windows containment | 1k | 2026-09-13 | 18 | |
| bytecodealliance/cap-std | Process sandbox (OS primitives) | 821 | 2026-08-20 | 28 | |
| jamesstringer90/appsandbox | Windows containment | 704 | 2026-09-17 | 45 | |
| restyler/awesome-sandbox | Survey / list / field guide | 585 | 2026-08-12 | 6 | |
| microsoft/Windows-Sandbox | Windows containment | 568 | 2024-09-20 | 57 | |
| rust-vmm/vm-memory | MicroVM / lightweight VMM | 366 | 2026-07-13 | 0 | |
| bureado/awesome-agent-runtime-security | Survey / list / field guide | 116 | 2026-09-13 | 26 | |
| codesandbox/codesandbox-sdk | Hosted sandbox API / CDE | 111 | 2026-06-18 | 25 | |
| buildkite/cleanroom | MicroVM / lightweight VMM | 67 | 2026-09-19 | 1 | |
| microsoft/quicksand | MicroVM / lightweight VMM | 52 | 2026-09-18 | 6 | |
| kkovacs/vmtree | Agent orchestrator / GUI | 48 | 2026-07-26 | 0 | |
| webcoyote/awesome-AI-sandbox | Survey / list / field guide | 33 | 2026-06-11 | 5 | |
| pjlsergeant/ai-sandboxes | Survey / list / field guide | 10 | 2026-07-21 | 1 | |
| daytonaio/daytona | Hosted sandbox API / CDE | 71k | 2026-07-24 | 457 | |
| codesandbox/codesandbox-client | Hosted sandbox API / CDE | 13k | 2026-09-07 | 615 | |
| FerretDB/FerretDB | Adjacent tooling | 11k | 2026-06-05 | 449 | |
| stackblitz/webcontainer-core | WASM / language isolate | 4.6k | 2025-04-22 | 817 | |
| strongdm/leash | Agent sandbox wrapper | 592 | 2026-04-06 | 21 | |
| genuinetools/bpfd | Network / egress control | 484 | 2021-05-07 | 5 | |
| BIGPPWONG/EdgeBox | Computer-use / desktop automation | 213 | 2026-04-22 | 2 | |
| robcholz/vibebox | Agent sandbox wrapper | 188 | 2026-02-18 | 0 | |
| obra/packnplay | Agent sandbox wrapper | 174 | 2026-03-21 | 9 | |
| deepclause/deepclause-sdk | Policy, approval & audit | 56 | 2026-09-10 | 4 | |
| Kiln-AI/Kilntainers | Agent sandbox wrapper | 51 | 2026-03-03 | 8 | |
| HQarroum/microbox | Agent sandbox wrapper | 47 | 2025-10-09 | 0 | |
| cirruslabs/chamber | Full VM / hypervisor | 46 | 2025-12-10 | 2 | |
| colony-2/shai | Agent sandbox wrapper | 43 | 2026-07-02 | 1 | |
| akshayaggarwal99/boxed | Agent sandbox wrapper | 13 | 2026-09-11 | 0 | |
| divmain/treebeard | Filesystem / copy-on-write | 12 | 2026-01-06 | 0 | |
| matheusmoreira/virtdev | Full VM / hypervisor | 8 | 2026-09-04 | 0 | |
| PunkGo/punkgo-jack | Policy, approval & audit | 7 | 2026-07-02 | 0 | |
| ligon/sucoder | Agent sandbox wrapper | 6 | 2026-09-19 | 1 | |
| paulux84/codex-lockbox | Agent sandbox wrapper | 5 | 2026-02-04 | 0 | |
| PredicateSystems/predicate-secure | Policy, approval & audit | 5 | 2026-02-28 | 0 | |
| Sarthak30/agentsafe | Agent sandbox wrapper | 4 | 2025-09-21 | 0 | |
| ashishgituser/nervos | Agent sandbox wrapper | – | – | – | |
| CursorTouch/Windows-MCP | Computer-use / desktop automation | 7k | 2026-09-16 | 26 |
RECOMMENDED / MATURE
Suitable for serious use on the available evidence.
Docker Engine
--dangerously-skip-permissions inside a Docker container is still the most common agent setup in the wild.docker group as equivalent to root for this reason. Rootless Docker exists and narrows the gap, but keeps a per-user daemon plus rootlesskit in the path.mitmproxy
Firecracker
--internal network can. On enforcing SELinux hosts, bind mounts without :z/:Z produce EACCES that people misdiagnose as a Podman bug. Overlay mounting falls back to fuse-overlayfs or vfs on older kernels. Bind mounts are the real footgun: mounting your home directory hands the container everything your user can reach and people then blame the container engine.podman machine giving Windows and macOS users a real Linux VM boundary for free, Quadlet for declarative services, Kube YAML support, and an API/socket that lets a GUI drive it without shelling out.Lima
gVisor
--runtime=runsc, no nested virtualisation requirement (so it works on VPSs where microVMs do not), --network=none and DNS-scoped egress modes, and a public continuous fuzzing dashboard you can watch.Wasmtime
QEMU
Cloud Hypervisor
Incus
GOOD, BUT TEST IT FIRST
Strong enough to investigate, with meaningful caveats.
Claude Code
OpenAI Codex CLI
Everyone-writable directories and NTFS hard links, which is why the implementation reports partial enforcement. PowerShell AST inspection has opaque regions (Invoke-Expression, -EncodedCommand, dynamic ScriptBlock creation) that were only fixed after being reported. Elsewhere, Landlock is port-level only for network, so domain policy still needs a proxy.WSL2
Infisical / Agent Vault
c/ua (Cua)
E2B
Anthropic Sandbox Runtime (@anthropic-ai/sandbox-runtime)
nono (nolabs)
container-use (Dagger)
agent-desktop
Sculptor (Imbue)
EXPERIMENTAL / EARLY
Interesting technology, currently risky to depend on.
Microsandbox
BoxLite
pivot_root isolation, privilege dropping and cgroups v2 on Linux, sandbox-exec on macOS. Storage is a per-box QCOW2 disk with copy-on-write, so a box can be snapshotted, cloned and rolled back. Network egress can be restricted with an allow_net list, and secrets are injected as placeholders so that real values are not placed inside the guest. That is a well-shaped design; every word of it is the vendor's own description, and none of it has been independently measured.allow_net egress control, placeholder secret injection with environment sanitisation, and per-box metrics.Microsoft Execution Containers (MXC)
CODE / IDEAS WORTH STUDYING
Do not build directly on the repo, but parts of it are excellent.
NanaBox
App Sandbox
PROBABLY SKIP FOR NOW
Expected hassle and risk look greater than the likely benefit.
Daytona
EdgeBox
NOT AN ISOLATION LAYER
Useful, but it does not contain the agent.
Windows-MCP
What the assessments have in common
Patterns that keep showing up
- Engine quality and product quality are different things. The independent security
study found that engine classes separate cleanly on every axis while products inside a class do not —
and the worst findings were product-level decisions (a
Privileged: truedefault, a frozen engine pin, nested virtualisation left on) rather than engine bugs. - Pin policy is an operator variable, and it dominates. Engine-side patch latency was effectively zero for coordinated disclosures; the observed delay ranged from zero days to 471 days to "opaque", and every day of it came from downstream pinning. Whatever you choose, decide how you receive engine updates.
- The projects with the best documentation of their own limits are the ones worth reading. nono publishes a security model that enumerates what it does not protect; byre's README says outright that a container is not a microVM; Fence documents that its network allowlisting is content-blind. That honesty correlates with the quality of everything else.
- Abandonment is the most common failure, and it is measurable. The reliable signals are last-commit date, release count, maintainer count and whether the issue tracker has recent non-bot activity — not stars.
- Beta labels are usually accurate. microsandbox calls itself beta and has the bug list to prove it; MXC says its profiles are not security boundaries; Gondolin documents the trust assumptions it relies on. Take the label seriously and integrate as an option rather than a dependency.
It cannot tell you whether a project is secure. Adoption-risk assessment measures reliability, maintenance and honesty, and a project can score well on all three while having an architectural flaw that matters for your threat model. Read the security section of each assessment for that, and read Reality Check for the boundaries that no amount of good maintenance changes.
Every assessment would carry a confidence level if this were a formal report. The short version: confidence is high for projects with public issue trackers and thousands of users, medium for those with active repositories but thin usage evidence, and low for anything measured in weeks of history or single-author projects — which is precisely why several of those are rated "experimental" rather than condemned.