Verdicts · risk ratings

Repo risk ratings

29 repositories received a full written assessment using the same procedure: establish identity and state, hunt for recurring failure families, cluster complaints, look outside the repository for firsthand reports, measure maintainer behaviour, normalise against apparent usage, and write down the reasoning. No numeric score, no composite, no fake precision.

The five categories

There is a sixth label, "not an isolation layer", which is deliberately not a category: it marks software that is useful and does not contain the agent at all.

Recommended / Mature

Appears suitable for serious use based on available evidence: an organisation behind it, a release cadence, longitudinal use at scale, and complaints proportionate to real adoption.

Good, but test it first

Strong enough to investigate and deploy, with meaningful caveats: platform asymmetry, a young component, a default you must change, or a CVE history that is disclosed rather than absent.

Experimental / Early

Interesting technology, currently risky to depend on. Frequently genuinely good engineering with insufficient longitudinal evidence, or an explicit beta label the maintainers are honest about.

Code / ideas worth studying

Do not build directly on the repository, but parts of its implementation or documentation are extremely useful. This is the category for "the idea is better than the dependency".

Probably skip for now

The expected hassle and risk look greater than the likely benefit — usually because development stopped, there is no release, or the licence blocks the reuse you were imagining.

Not an isolation layer

Useful software that does not contain the agent. Included so that the distinction between "gives the agent access" and "contains the agent" stays visible.

How complaints were normalised

Raw issue counts are close to meaningless in this space. A project with tens of thousands of users generates complaints as a function of usage; a project with two hundred stars looks pristine because almost nobody has tried it. Every assessment here therefore asks the same questions: does the same class of failure recur across versions, do users need destructive workarounds, how fast and how completely did maintainers respond, and are the same symptoms reported independently on the issue tracker and elsewhere. Where a trustworthy denominator does not exist, the assessment says so rather than inventing one.

The rated set. Sorting works — click a heading.
ProjectRatingFamily★Last pushOpen issues
denoland/denoRECOMMENDED / MATUREWASM / language isolate108k2026-09-171.6k
moby/mobyRECOMMENDED / MATUREContainer / namespace72k2026-09-193.9k
mitmproxy/mitmproxyRECOMMENDED / MATURENetwork / egress control45k2026-09-10491
aquasecurity/trivyRECOMMENDED / MATUREAdjacent tooling37k2026-09-18271
firecracker-microvm/firecrackerRECOMMENDED / MATUREMicroVM / lightweight VMM36k2026-09-18100
tailscale/tailscaleRECOMMENDED / MATURENetwork / egress control36k2026-09-194.6k
hashicorp/vaultRECOMMENDED / MATURESecrets & credential brokering36k2026-09-181.4k
utmapp/UTMRECOMMENDED / MATUREFull VM / hypervisor35k2026-09-191.1k
podman-container-tools/podmanRECOMMENDED / MATUREContainer / namespace32k2026-09-191k
goharbor/harborRECOMMENDED / MATUREAdjacent tooling29k2026-09-18902
gitleaks/gitleaksRECOMMENDED / MATUREAdjacent tooling29k2026-09-09483
trufflesecurity/trufflehogRECOMMENDED / MATUREAdjacent tooling27k2026-09-18559
cilium/ciliumRECOMMENDED / MATURENetwork / egress control25k2026-09-191.1k
getsops/sopsRECOMMENDED / MATURESecrets & credential brokering23k2026-09-18447
lima-vm/limaRECOMMENDED / MATUREFull VM / hypervisor21k2026-09-19526
containerd/containerdRECOMMENDED / MATUREContainer / namespace21k2026-09-19471
google/gvisorRECOMMENDED / MATUREApplication kernel19k2026-09-19814
bytecodealliance/wasmtimeRECOMMENDED / MATUREWASM / language isolate18k2026-09-18848
coder/coderRECOMMENDED / MATUREHosted sandbox API / CDE15k2026-09-191k
kubernetes-sigs/kindRECOMMENDED / MATUREContainer / namespace15k2026-09-04242
qemu/qemuRECOMMENDED / MATUREFull VM / hypervisor13k2026-09-190
opencontainers/runcRECOMMENDED / MATUREContainer / namespace13k2026-09-18339
anchore/grypeRECOMMENDED / MATUREAdjacent tooling12k2026-09-18401
simonw/llmRECOMMENDED / MATUREAdjacent tooling12k2026-09-08699
krallin/tiniRECOMMENDED / MATUREAdjacent tooling11k2025-05-0845
canonical/multipassRECOMMENDED / MATUREFull VM / hypervisor9.2k2026-09-18418
containers/bubblewrapRECOMMENDED / MATUREProcess sandbox (OS primitives)8.8k2026-09-18196
cloudflare/workerdRECOMMENDED / MATUREWASM / language isolate8.7k2026-09-19730
bats-core/bats-coreRECOMMENDED / MATUREAdjacent tooling6.3k2026-09-16129
cloud-hypervisor/cloud-hypervisorRECOMMENDED / MATUREMicroVM / lightweight VMM6.2k2026-09-19228
lxc/incusRECOMMENDED / MATUREFull VM / hypervisor6.2k2026-09-1846
devcontainers/specRECOMMENDED / MATUREContainer / namespace5.7k2026-03-20190
ossf/scorecardRECOMMENDED / MATUREAdjacent tooling5.7k2026-09-18457
lxc/lxcRECOMMENDED / MATUREContainer / namespace5.3k2026-09-07151
google/security-researchRECOMMENDED / MATUREAdjacent tooling4.6k2026-09-1891
Yelp/detect-secretsRECOMMENDED / MATUREAdjacent tooling4.6k2026-04-02183
containers/crunRECOMMENDED / MATUREContainer / namespace4.1k2026-09-1749
google/nsjailRECOMMENDED / MATUREProcess sandbox (OS primitives)4.1k2026-08-2739
ddev/ddevRECOMMENDED / MATUREContainer / namespace3.9k2026-09-19176
opencontainers/runtime-specRECOMMENDED / MATUREContainer / namespace3.7k2026-04-2494
buildpacks/packRECOMMENDED / MATUREAdjacent tooling3k2026-09-14202
seccomp/libseccompRECOMMENDED / MATUREProcess sandbox (OS primitives)9342026-07-0158
containers/conmonRECOMMENDED / MATUREContainer / namespace5002026-09-1827
google/minijailRECOMMENDED / MATUREProcess sandbox (OS primitives)3842026-09-150
microsoft/artifacts-keyringRECOMMENDED / MATURESecrets & credential brokering452026-08-199
anomalyco/opencodeGOOD, BUT TEST IT FIRSTAgent sandbox wrapper208k2026-09-196k
anthropics/claude-codeGOOD, BUT TEST IT FIRSTAgent sandbox wrapper146k2026-09-1912k
openai/codexGOOD, BUT TEST IT FIRSTAgent sandbox wrapper125k2026-09-1917k
browser-use/browser-useGOOD, BUT TEST IT FIRSTComputer-use / desktop automation115k2026-09-18452
google-gemini/gemini-cliGOOD, BUT TEST IT FIRSTAgent sandbox wrapper107k2026-09-19841
modelcontextprotocol/serversGOOD, BUT TEST IT FIRSTAgent orchestrator / GUI90k2026-09-03547
aaif-goose/gooseGOOD, BUT TEST IT FIRSTAgent sandbox wrapper54k2026-09-19395
Aider-AI/aiderGOOD, BUT TEST IT FIRSTAgent sandbox wrapper49k2026-05-221.9k
ziglang/zigGOOD, BUT TEST IT FIRSTAdjacent tooling43k2025-11-272.8k
microsoft/WSLGOOD, BUT TEST IT FIRSTWindows containment33k2026-09-19981
openai/openai-agents-pythonGOOD, BUT TEST IT FIRSTAgent orchestrator / GUI29k2026-09-1893
Infisical/infisicalGOOD, BUT TEST IT FIRSTSecrets & credential brokering29k2026-09-19793
keepassxreboot/keepassxcGOOD, BUT TEST IT FIRSTSecrets & credential brokering28k2026-09-18901
QwenLM/qwen-codeGOOD, BUT TEST IT FIRSTAgent sandbox wrapper27k2026-09-191.5k
trycua/cuaGOOD, BUT TEST IT FIRSTComputer-use / desktop automation24k2026-09-191k
wasmerio/wasmerGOOD, BUT TEST IT FIRSTWASM / language isolate21k2026-09-19267
e2b-dev/E2BGOOD, BUT TEST IT FIRSTHosted sandbox API / CDE13k2026-09-1972
gitpod-io/gitpodGOOD, BUT TEST IT FIRSTHosted sandbox API / CDE13k2026-09-18450
microsoft/wslgGOOD, BUT TEST IT FIRSTWindows containment11k2026-07-06719
microsoft/WSL2-Linux-KernelGOOD, BUT TEST IT FIRSTWindows containment10k2026-08-01134
falcosecurity/falcoGOOD, BUT TEST IT FIRSTPolicy, approval & audit9.4k2026-09-1842
kata-containers/kata-containersGOOD, BUT TEST IT FIRSTMicroVM / lightweight VMM8.8k2026-09-191.2k
draios/sysdigGOOD, BUT TEST IT FIRSTNetwork / egress control8.3k2026-04-13117
wasm3/wasm3GOOD, BUT TEST IT FIRSTWASM / language isolate8k2026-09-1919
netblue30/firejailGOOD, BUT TEST IT FIRSTProcess sandbox (OS primitives)7.7k2026-09-16523
openbao/openbaoGOOD, BUT TEST IT FIRSTSecrets & credential brokering7.4k2026-09-18320
rancher-sandbox/rancher-desktopGOOD, BUT TEST IT FIRSTContainer / namespace7.4k2026-09-181.1k
openai/tartGOOD, BUT TEST IT FIRSTFull VM / hypervisor6.8k2026-09-1673
anthropics/sandbox-runtimeGOOD, BUT TEST IT FIRSTProcess sandbox (OS primitives)5.3k2026-09-19212
cilium/tetragonGOOD, BUT TEST IT FIRSTPolicy, approval & audit5k2026-09-19283
nolabs-ai/nonoGOOD, BUT TEST IT FIRSTSecrets & credential brokering4.1k2026-09-18213
dagger/container-useGOOD, BUT TEST IT FIRSTAgent orchestrator / GUI4k2026-09-1462
nestybox/sysboxGOOD, BUT TEST IT FIRSTContainer / namespace3.9k2026-09-15213
virtio-win/kvm-guest-drivers-windowsGOOD, BUT TEST IT FIRSTFull VM / hypervisor2.7k2026-09-15180
libkrun/libkrunGOOD, BUT TEST IT FIRSTMicroVM / lightweight VMM2.7k2026-09-1698
superfly/flyctlGOOD, BUT TEST IT FIRSTHosted sandbox API / CDE1.7k2026-09-18221
google/crosvmGOOD, BUT TEST IT FIRSTMicroVM / lightweight VMM1.3k2026-09-182
lahfir/agent-desktopGOOD, BUT TEST IT FIRSTComputer-use / desktop automation1.3k2026-09-1722
step-security/harden-runnerGOOD, BUT TEST IT FIRSTAdjacent tooling1.3k2026-08-3156
RchGrav/claudeboxGOOD, BUT TEST IT FIRSTAgent sandbox wrapper1.2k2026-09-1718
cloudflare/sandbox-sdkGOOD, BUT TEST IT FIRSTHosted sandbox API / CDE1.1k2026-09-1945
trailofbits/claude-code-devcontainerGOOD, BUT TEST IT FIRSTAgent sandbox wrapper9432026-08-284
firecracker-microvm/firecracker-go-sdkGOOD, BUT TEST IT FIRSTMicroVM / lightweight VMM6732026-02-1053
modal-labs/modal-clientGOOD, BUT TEST IT FIRSTHosted sandbox API / CDE5142026-09-1926
docker/sbx-releasesGOOD, BUT TEST IT FIRSTMicroVM / lightweight VMM3892026-09-18330
archlinux/arch-boxesGOOD, BUT TEST IT FIRSTFull VM / hypervisor2652025-12-176
imbue-ai/sculptorGOOD, BUT TEST IT FIRSTAgent orchestrator / GUI2332026-09-1920
vercel/sandboxGOOD, BUT TEST IT FIRSTHosted sandbox API / CDE2012026-09-1928
omnigent-ai/omnigentEXPERIMENTAL / EARLYAgent orchestrator / GUI10k2026-09-191.4k
NVIDIA/OpenShellEXPERIMENTAL / EARLYAgent sandbox wrapper8.7k2026-09-19530
superradcompany/microsandboxEXPERIMENTAL / EARLYMicroVM / lightweight VMM8.3k2026-09-1993
smol-machines/smolvmEXPERIMENTAL / EARLYMicroVM / lightweight VMM6.1k2026-09-1982
hyperlight-dev/hyperlightEXPERIMENTAL / EARLYMicroVM / lightweight VMM4.7k2026-09-18220
nanovms/nanosEXPERIMENTAL / EARLYFull VM / hypervisor3.2k2026-09-1987
boxlite-ai/boxliteEXPERIMENTAL / EARLYMicroVM / lightweight VMM2.3k2026-09-19294
earendil-works/gondolinEXPERIMENTAL / EARLYMicroVM / lightweight VMM2.2k2026-07-0647
eugene1g/agent-safehouseEXPERIMENTAL / EARLYAgent sandbox wrapper2.1k2026-09-1824
e2b-dev/desktopEXPERIMENTAL / EARLYComputer-use / desktop automation1.5k2026-09-1811
microsoft/mxcEXPERIMENTAL / EARLYWindows containment1.3k2026-09-1984
clawkwork/clawkEXPERIMENTAL / EARLYMicroVM / lightweight VMM1k2026-08-134
fencesandbox/fenceEXPERIMENTAL / EARLYAgent sandbox wrapper9692026-09-1034
instavm/coderunnerEXPERIMENTAL / EARLYMicroVM / lightweight VMM8942026-08-139
Katakate/k7EXPERIMENTAL / EARLYMicroVM / lightweight VMM8082026-09-140
Katakate/k7EXPERIMENTAL / EARLYMicroVM / lightweight VMM8082026-09-140
finbarr/yoloboxEXPERIMENTAL / EARLYAgent sandbox wrapper6432026-08-282
jingkaihe/matchlockEXPERIMENTAL / EARLYAgent sandbox wrapper6212026-07-269
manuelschipper/nahEXPERIMENTAL / EARLYPolicy, approval & audit4862026-09-150
nikvdp/ccoEXPERIMENTAL / EARLYAgent sandbox wrapper4252026-09-122
webcoyote/sandvaultEXPERIMENTAL / EARLYAgent sandbox wrapper4172026-09-141
landlock-lsm/islandEXPERIMENTAL / EARLYProcess sandbox (OS primitives)3342026-05-2621
GreyhavenHQ/greywallEXPERIMENTAL / EARLYAgent sandbox wrapper2982026-08-1324
eqtylab/cupcakeEXPERIMENTAL / EARLYPolicy, approval & audit2932026-03-0222
nanvix/nanvixEXPERIMENTAL / EARLYMicroVM / lightweight VMM2822026-09-19244
butter-dot-dev/bVisorEXPERIMENTAL / EARLYApplication kernel2172026-02-231
kstenerud/yoloaiEXPERIMENTAL / EARLYAgent sandbox wrapper2122026-08-219
webcoyote/clodpodEXPERIMENTAL / EARLYFull VM / hypervisor1872026-08-304
postrv/forgemaxEXPERIMENTAL / EARLYAgent sandbox wrapper1512026-09-103
jgbrwn/vibebinEXPERIMENTAL / EARLYContainer / namespace1082026-08-100
LuD1161/agentjailEXPERIMENTAL / EARLYAgent sandbox wrapper942026-09-082
swelljoe/flarEXPERIMENTAL / EARLYAgent sandbox wrapper552026-08-186
pjlsergeant/byreEXPERIMENTAL / EARLYAgent sandbox wrapper322026-09-171
agentic-dev3o/sandbox-shellEXPERIMENTAL / EARLYAgent sandbox wrapper302026-09-181
recodelabs/lima-devboxEXPERIMENTAL / EARLYFull VM / hypervisor282026-07-031
corv89/shannotEXPERIMENTAL / EARLYPolicy, approval & audit262026-04-136
binwiederhier/sandclaudeEXPERIMENTAL / EARLYAgent sandbox wrapper232026-05-300
agentcage/agentcageEXPERIMENTAL / EARLYAgent sandbox wrapper222026-09-1929
jskswamy/aideEXPERIMENTAL / EARLYAgent sandbox wrapper172026-09-172
tarsgate/skynotEXPERIMENTAL / EARLYAgent sandbox wrapper162026-06-186
ashishgituser/bunkervmEXPERIMENTAL / EARLYAgent sandbox wrapper12026-08-180
jart/cosmopolitanCODE / IDEAS WORTH STUDYINGWASM / language isolate21k2026-07-20228
agent-infra/sandboxCODE / IDEAS WORTH STUDYINGHosted sandbox API / CDE6k2026-09-1472
unikraft/unikraftCODE / IDEAS WORTH STUDYINGFull VM / hypervisor3.9k2026-09-19367
M2Team/NanaBoxCODE / IDEAS WORTH STUDYINGWindows containment1k2026-09-1318
bytecodealliance/cap-stdCODE / IDEAS WORTH STUDYINGProcess sandbox (OS primitives)8212026-08-2028
jamesstringer90/appsandboxCODE / IDEAS WORTH STUDYINGWindows containment7042026-09-1745
restyler/awesome-sandboxCODE / IDEAS WORTH STUDYINGSurvey / list / field guide5852026-08-126
microsoft/Windows-SandboxCODE / IDEAS WORTH STUDYINGWindows containment5682024-09-2057
rust-vmm/vm-memoryCODE / IDEAS WORTH STUDYINGMicroVM / lightweight VMM3662026-07-130
bureado/awesome-agent-runtime-securityCODE / IDEAS WORTH STUDYINGSurvey / list / field guide1162026-09-1326
codesandbox/codesandbox-sdkCODE / IDEAS WORTH STUDYINGHosted sandbox API / CDE1112026-06-1825
buildkite/cleanroomCODE / IDEAS WORTH STUDYINGMicroVM / lightweight VMM672026-09-191
microsoft/quicksandCODE / IDEAS WORTH STUDYINGMicroVM / lightweight VMM522026-09-186
kkovacs/vmtreeCODE / IDEAS WORTH STUDYINGAgent orchestrator / GUI482026-07-260
webcoyote/awesome-AI-sandboxCODE / IDEAS WORTH STUDYINGSurvey / list / field guide332026-06-115
pjlsergeant/ai-sandboxesCODE / IDEAS WORTH STUDYINGSurvey / list / field guide102026-07-211
daytonaio/daytonaPROBABLY SKIP FOR NOWHosted sandbox API / CDE71k2026-07-24457
codesandbox/codesandbox-clientPROBABLY SKIP FOR NOWHosted sandbox API / CDE13k2026-09-07615
FerretDB/FerretDBPROBABLY SKIP FOR NOWAdjacent tooling11k2026-06-05449
stackblitz/webcontainer-corePROBABLY SKIP FOR NOWWASM / language isolate4.6k2025-04-22817
strongdm/leashPROBABLY SKIP FOR NOWAgent sandbox wrapper5922026-04-0621
genuinetools/bpfdPROBABLY SKIP FOR NOWNetwork / egress control4842021-05-075
BIGPPWONG/EdgeBoxPROBABLY SKIP FOR NOWComputer-use / desktop automation2132026-04-222
robcholz/vibeboxPROBABLY SKIP FOR NOWAgent sandbox wrapper1882026-02-180
obra/packnplayPROBABLY SKIP FOR NOWAgent sandbox wrapper1742026-03-219
deepclause/deepclause-sdkPROBABLY SKIP FOR NOWPolicy, approval & audit562026-09-104
Kiln-AI/KilntainersPROBABLY SKIP FOR NOWAgent sandbox wrapper512026-03-038
HQarroum/microboxPROBABLY SKIP FOR NOWAgent sandbox wrapper472025-10-090
cirruslabs/chamberPROBABLY SKIP FOR NOWFull VM / hypervisor462025-12-102
colony-2/shaiPROBABLY SKIP FOR NOWAgent sandbox wrapper432026-07-021
akshayaggarwal99/boxedPROBABLY SKIP FOR NOWAgent sandbox wrapper132026-09-110
divmain/treebeardPROBABLY SKIP FOR NOWFilesystem / copy-on-write122026-01-060
matheusmoreira/virtdevPROBABLY SKIP FOR NOWFull VM / hypervisor82026-09-040
PunkGo/punkgo-jackPROBABLY SKIP FOR NOWPolicy, approval & audit72026-07-020
ligon/sucoderPROBABLY SKIP FOR NOWAgent sandbox wrapper62026-09-191
paulux84/codex-lockboxPROBABLY SKIP FOR NOWAgent sandbox wrapper52026-02-040
PredicateSystems/predicate-securePROBABLY SKIP FOR NOWPolicy, approval & audit52026-02-280
Sarthak30/agentsafePROBABLY SKIP FOR NOWAgent sandbox wrapper42025-09-210
ashishgituser/nervosPROBABLY SKIP FOR NOWAgent sandbox wrapper–––
CursorTouch/Windows-MCPNOT AN ISOLATION LAYERComputer-use / desktop automation7k2026-09-1626

RECOMMENDED / MATURE

Suitable for serious use on the available evidence.

Docker Engine

RECOMMENDED / MATURE
★ 72kforks 19kopen issues 3.9kcreated 2013-01-18last push 2026-09-19licence Apache-2.0
Type
Container engine (OCI) with a root daemon
Platforms
Linux native; macOS and Windows via Docker Desktop or manual WSL2 install
Agent-specific
No — but --dangerously-skip-permissions inside a Docker container is still the most common agent setup in the wild.
Open source
Yes, Apache-2.0 for the engine; Docker Desktop is a separate commercial product with licence limits.
Isolation model
The same namespaces/cgroups/seccomp/LSM stack as Podman, but driven by a persistent root-owned daemon reachable through /var/run/docker.sock. Anyone who can reach that socket can ask the daemon to create a privileged container with a host bind mount, which is effectively root. Docker's own documentation treats membership of the docker group as equivalent to root for this reason. Rootless Docker exists and narrows the gap, but keeps a per-user daemon plus rootlesskit in the path.
Adoption evidence
The default container tooling for the entire industry. Docker Desktop is what most developers on Windows and macOS already have running. Docker's SDKs and compose files are the lingua franca of dev environments, dev containers and CI.
Development activity
Continuous and extremely well resourced. Moby and the Docker CLI commit daily; Docker Sandboxes shows the company actively moving into the agent-isolation market.
Common complaints
The socket is the classic one: dozens of quickstart templates and blog posts still mount docker.sock into agent containers, which is a documented escape path. Docker Desktop licensing and its resource footprint annoy people. On Windows, the WSL2 integration means WSL interop caveats apply to anything "isolated" inside it. Storage growth (images, volumes, build cache) is a perennial complaint, and the same is true of Podman.
Security considerations
The March 2026 Sysdig-reported incident is the cautionary tale worth knowing: an LLM harness, not a human, called the Docker socket API, created a privileged container with a host bind mount, read /etc/shadow and SSH keys, and replayed a Kubernetes service-account token to dump namespace secrets. The only precondition was a mounted socket. Beyond that, the same shared-kernel ceiling as any container applies.
Interesting features
The largest ecosystem in container land by a wide margin, Compose, BuildKit, dev container integration, and Docker Sandboxes (see the microVM chapter) as a first-party escape hatch when containers are not enough.
Why this rating
The engine itself is mature, well-tested and maintained by people whose full-time job is container security. The risk here is not instability, it is configuration: the surrounding ecosystem is saturated with patterns (socket mounts, privileged mode, host networking, whole-home bind mounts) that quietly turn a container into a root shell. Judged on its own code the project is excellent; judged on the defaults people copy from the internet it is the most common cause of accidental non-isolation in this guide.
Our take
Use it if you already have it, and audit what the template you copied actually did. If you are choosing today rather than inheriting, Podman gives you the same primitives with a smaller root surface and no daemon. Either way, the decision that matters is not Docker versus Podman, it is whether there is a hypervisor under it.

mitmproxy

RECOMMENDED / MATURE
★ 45kforks 4.7kopen issues 491created 2010-02-16last push 2026-09-10licence MIT
Type
Intercepting HTTP/TLS proxy
Platforms
Linux, macOS, Windows
Agent-specific
No — but it is the engine inside most DIY agent credential brokers and egress inspectors.
Open source
Yes, MIT
Isolation model
None. It is a mediating proxy, so it mediates rather than confines. It can terminate TLS, inspect requests, rewrite headers, swap credentials, and block destinations, and it is useful precisely because it sits where an agent's traffic crosses the boundary. The important caveat is placement: run it inside the sandbox and the agent can read your keys; run it outside and you must ensure the sandbox cannot route around it.
Adoption evidence
Fifteen years old, forty-five thousand stars, a standard tool in penetration testing and API work, and the substrate for projects like keys-on-the-wire, sandcat and OpenSandbox's credential-vault design.
Development activity
Steady and continuous, with a small core team and regular releases.
Common complaints
Operational, not correctness: certificate installation into the tools that must trust the proxy, and the fact that not every HTTP client honours proxy environment variables. Node.js historically ignored HTTP_PROXY until Node 24's NODE_USE_ENV_PROXY, so people install certificate authorities, get the environment right, and then find a tool connecting directly. Transparent iptables redirection is usually the fix. Non-HTTP protocols and plain sockets need different treatment entirely.
Security considerations
The proxy is a trust chokepoint by design: whoever controls it can read and rewrite everything, so it must live outside the sandbox and its configuration must not be writable by the agent. Fail-open is the classic mistake — if the tool silently falls back to a direct connection when proxy variables are missing, the policy evaporates without an error. Signature schemes that sign the request body (AWS SigV4) cannot be handled by naive header substitution and require re-signing with the real key.
Interesting features
Scriptable request and response hooks, transparent mode, WireGuard mode, addon ecosystem, and the fact that both the incoming phantom-token and the outgoing real-credential substitution are small, auditable pieces of Python.
Why this rating
It is not a sandbox and it does not claim to be, so judging it as one would be a category error. As the network and credential boundary of a composition it is extremely well proven, and the failure modes are documented by a decade of other people's production incidents. The main risk is architectural: building a credential broker on it requires running it outside the sandbox, pinning it, and refusing to fail open.
Our take
Use it as the egress and credential chokepoint, running outside the sandbox, with a deny-by-default allowlist. Verify that every tool in your agent's toolchain honours the proxy, and test the fail-closed path on purpose before you rely on it.

Firecracker

RECOMMENDED / MATURE
★ 36kforks 2.6kopen issues 100created 2017-10-19last push 2026-09-18licence Apache-2.0
Type
KVM microVM monitor (VMM)
Platforms
Linux with KVM only
Agent-specific
No — but it is the foundation under Lambda, Fargate, Fly.io, Vercel Sandbox, E2B and others.
Open source
Yes, Apache-2.0
Isolation model
A small Rust VMM that boots a real Linux kernel per workload behind a KVM hardware boundary, exposing six emulated devices (virtio-net, virtio-block, virtio-balloon, vsock, serial, keyboard) instead of QEMU's thousands. Around 50k lines versus QEMU's millions. The trade-off is deliberate minimalism: no virtio-fs, no CPU or memory hotplug, no VFIO, no GPU, no confidential computing, and resources are sized at boot.
Adoption evidence
The most heavily deployed microVM in the world by a very wide margin. AWS runs it for Lambda and Fargate; Fly.io, Vercel Sandbox, E2B, Sprites and many others build on it. It is the reason 'microVM' became a category.
Development activity
Continuous, AWS-backed, with a strict release process and a well-run security response.
Common complaints
Ergonomics, uniformly. There is no virtio-fs, which disqualifies it for teams whose workflow is a bind-mounted repository — one documented assessment called it "disqualified" for exactly that reason and put a Docker-to-microVM migration at four to six engineer-months. It needs an uncompressed kernel you build or fetch, and a rootfs ext4/xfs/erofs image you materialise, which is why entire projects exist purely to publish kernel catalogues. Cold start is about a second, not the 125ms on the datasheet, unless you build a snapshot pool. Users on cloud VMs without nested virtualisation simply cannot run it.
Security considerations
Two Escape-class CVEs landed in 2026 (CVE-2026-5747, an out-of-bounds write in virtio-pci, and CVE-2026-1386, a jailer symlink host-write), breaking the previous 'no published hypervisor escape' baseline. An independent audit also corrected the record on fuzzing: the project has no upstream continuous fuzzer, and what looked like fuzzing support is a deterministic-build hook. The seccomp ceiling is nonetheless the tightest measured of any product in that study, at 55 allowlisted syscalls.
Interesting features
Sub-second boots with per-VM kernels, a very small device model, a jailer for defence in depth, snapshot/restore, and an ecosystem of SDKs and orchestrators that hides most of the ugliness if you are willing to adopt one.
Why this rating
The project is unambiguously production-grade: it protects a very large slice of the world's serverless traffic and has a competent security process. The risk it carries for an agent-sandbox user is not reliability, it is the mismatch between its design and your problem. Without virtio-fs you cannot hand it a directory, so the ergonomic cost lands entirely on the developer. That is a fit problem, not a trust problem, and it is why the guide points most desktop users at Cloud Hypervisor or a full VM instead.
Our take
Choose it if you are building infrastructure, you control the host, and you will invest in kernel and rootfs tooling or adopt an orchestrator. Do not choose it because it is famous; choose it because you need the smallest possible device model and you do not need to share a directory.

Podman

RECOMMENDED / MATURE
★ 32kforks 3.4kopen issues 1kcreated 2017-11-01last push 2026-09-19licence Apache-2.0
Type
Container engine (OCI)
Platforms
Linux native; Windows and macOS via a Podman Machine (WSL2/Hyper-V or Apple Virtualization)
Agent-specific
No — but it is the default backend for most agent sandboxes, including several on this site.
Open source
Yes, Apache-2.0
Isolation model
OCI containers: Linux namespaces (pid, mount, uts, ipc, net, user) plus cgroups v2 plus seccomp plus LSMs. Rootless by default, meaning container root maps to an unprivileged host UID through subuid/subgid, and there is no long-lived root daemon to pivot through because there is no daemon at all. On Windows the Linux engine runs inside a WSL2 or Hyper-V VM, so containers there are nested inside a real hypervisor boundary. What it is not: containers share the host kernel. A namespace, cgroup or runtime bug is a host bug.
Adoption evidence
The most widely deployed rootless container engine in the world, default in RHEL/Fedora/CentOS Stream, packaged everywhere, and the engine behind Podman Desktop, Rancher Desktop and much of the CI world. Every mainstream agent-sandbox wrapper either uses it or documents why it used Docker instead. It is boring, which in this guide is a compliment.
Development activity
Continuous, vendor-backed (Red Hat) with regular releases and a large maintainer pool. Commits land daily.
Common complaints
Rootless networking is the recurring theme: slirp4netns/pasta user-space stacks historically had both performance and isolation-relevant bugs, and rootful Podman cannot get a hard egress wall the way an --internal network can. On enforcing SELinux hosts, bind mounts without :z/:Z produce EACCES that people misdiagnose as a Podman bug. Overlay mounting falls back to fuse-overlayfs or vfs on older kernels. Bind mounts are the real footgun: mounting your home directory hands the container everything your user can reach and people then blame the container engine.
Security considerations
Rootless is a privilege-reduction boundary, not an isolation boundary. An escape lands you as an unprivileged user rather than root, which matters a great deal and is not the same as safety. The shared kernel is the ceiling: runc escapes such as CVE-2019-5736, CVE-2024-21626 and the 2025 procfs mount-race trio are the historical evidence. Never mount the Podman or Docker socket into an agent lane; that is handing over root. GPU, device and clipboard passthrough each widen the boundary visibly.
Interesting features
Daemonless operation, first-class rootless, the same CLI surface as Docker, podman machine giving Windows and macOS users a real Linux VM boundary for free, Quadlet for declarative services, Kube YAML support, and an API/socket that lets a GUI drive it without shelling out.
Why this rating
Sixteen years of container plumbing distilled into a daemonless engine that works on all three desktop platforms and has never needed a root daemon. The complaints that exist are almost entirely about networking and mounts — the two things an agent sandbox user must configure deliberately anyway — and they are documented, reproducible and understood. Complaints per user are low, the blast radius of a mistake is smaller than with Docker, and the failure modes are the boring kind you can plan for rather than the surprising kind.
Our take
This is the default answer for local coding-agent isolation and the right first backend for anything you build. Use it rootless, drop all capabilities, set no-new-privileges, keep memory and pids limits on, mount only the workspace you mean to share, and treat it as blast-radius reduction rather than as a security guarantee. If you need the guarantee, put a VM under it — which is exactly what Podman Machine already does on Windows.

Lima

RECOMMENDED / MATURE
★ 21kforks 956open issues 526created 2021-05-14last push 2026-09-19licence Apache-2.0
Type
Linux VM manager for macOS and Linux hosts
Platforms
macOS and Linux hosts, where it manages real Linux VMs with their own guest kernel. Windows hosts are supported only through the project's own experimental wsl2 (Windows 10 build 19041+ / 11) and hcs (Windows 11, plain mode only, one instance at a time) drivers, which the documentation restricts heavily; if you are on Windows and want a Linux boundary, WSL2 is the practical route.
Agent-specific
No — but 'a VM with your workspace mounted in, driven by YAML' is the cleanest declarative agent answer there is, on a Mac or on a Linux host.
Open source
Yes, Apache-2.0
Isolation model
A per-instance Linux VM with your choice of driver: vz (Apple's Virtualization framework), QEMU, or krunkit (libkrun, with GPU acceleration). Files are shared by reverse SSHFS or virtiofs, which means host exposure is a config decision rather than a mount that happens by default. The VM boundary is a real hypervisor boundary.
Adoption evidence
Twenty-two thousand stars, CNCF Incubating since October 2025, and the substrate underneath Colima, Rancher Desktop, Finch and Podman Desktop. Its version 2.0 added a plugin architecture, and 2.1/2.2 followed in 2026. That is an enormous amount of indirect, daily-exercised adoption.
Development activity
Continuous, with a small but very consistent maintainer group and regular releases.
Common complaints
Performance of host/guest file sharing is the perennial one: virtiofs and reverse SSHFS both have rough edges with large file trees, and people compare notes about which combination is fastest. Driver choice changes behaviour (vz versus QEMU), and older QEMU-based guests are slower to boot. On a Linux host with no hypervisor backend available — no /dev/kvm, no other accelerator — instances fall back to QEMU's TCG software emulation, which is too slow to be usable for interactive work; our own Linux test reached that state and never completed the SSH handshake (see experience.html#tested-here). Some confusion persists about what Lima is versus Colima, which wraps it, and about whether it is macOS-only: it is not, and QEMU is the default driver on a Linux host. Its Windows-host path is the genuinely narrow corner — the wsl2 and hcs drivers are labelled experimental and support a subset of Lima's options, and an attempt to use Lima on a Windows host did not reach a usable instance in our own notes (see experience.html).
Security considerations
The VM is the boundary and it is a good one. The risk lives in the sharing configuration: mount your whole home directory and you have undone the point. Guests can run containers inside, which is the recommended nesting, and GPU acceleration via krunkit is a wider host surface than a plain CPU-only guest.
Interesting features
Declarative YAML templates for reproducible guests, multiple drivers including a libkrun one, support for running containers inside the guest, and a well-worn download/cloud-init pipeline.
Why this rating
CNCF incubating, load-bearing for four widely used desktop tools, and the same declarative model on both of its main host platforms, macOS and Linux. The complaints are about file-sharing performance and driver differences rather than about it breaking or losing data. On macOS it is the shortest path from 'I want a real boundary' to 'I have one, declaratively', and on Linux it is a reasonable alternative to hand-rolling QEMU when you want a persistent, templated guest.
Our take
Stand up a Lima instance with an explicit, narrow mount and run your agent inside a container within it. Treat the share list as your exposure list. On a Linux host, confirm a hardware accelerator is actually in use before you invest time in the template — without /dev/kvm the guest runs under TCG software emulation and is not worth the wait. Skip the GPU driver unless you have measured that you need it. On Windows, treat Lima's own drivers as experimental and prefer WSL2.

gVisor

RECOMMENDED / MATURE
★ 19kforks 2kopen issues 814created 2018-04-26last push 2026-09-19licence Apache-2.0
Type
Application kernel (user-space Linux kernel) with an OCI runtime
Platforms
Linux hosts. There is no Windows or macOS host support.
Agent-specific
No, but it is the most common recommendation for 'stronger than a container, cheaper than a VM'.
Open source
Yes, Apache-2.0
Isolation model
Not a microVM, and the distinction matters. runsc interposes a Go reimplementation of the Linux kernel between the workload and the host kernel: applications make syscalls to Sentry, which serves them from user space, so a workload exploit has to get through the application kernel before it can touch the real kernel. A seccomp filter on the host side reduces what Sentry itself can ask for. Compatibility is bounded by how much of Linux has been reimplemented.
Adoption evidence
Google Cloud Run, App Engine and Cloud Functions run untrusted customer code on it, as does DigitalOcean App Platform and Modal. Google runs language-runtime regression suites against it. It has the only continuous public fuzzing dashboard of any product in the independent study (syzkaller) and ships a secfuzz library.
Development activity
Continuous since 2018 with deep Google resourcing and a rolling release model where the product is essentially upstream main.
Common complaints
Compatibility gaps dominate, and they are the reason teams bounce off it. io_uring is disabled by default and limited when enabled; unimplemented syscalls return ENOSYS and the failure looks like a broken application rather than a sandbox limitation. One audit found roughly half of tested MCP servers failed to start under gVisor because of unimplemented syscalls. Tooling breaks: perf events, some eBPF paths, and anything that assumes the kernel it probes is the kernel it gets. Performance is 10–30% worse on I/O-heavy work with older measurements much harsher. A production migration at one large company hit a 500× virtual-memory-area explosion and an ELF-loader bug before it was tuned.
Security considerations
The strongest measured posture of any product in the independent study: the fewest reachable host primitives, four hardening layers applied by default, no Escape-class CVEs in the window, and a silent-fix-first disclosure model that delivers the fix before the CVE exists. Two information leaks (host RAM total, BIOS product string) are implementation gaps rather than architectural concessions. The cost is paid in workload compatibility, and the project is honest about it.
Interesting features
An OCI-compatible runtime you can drop into Docker or containerd as --runtime=runsc, no nested virtualisation requirement (so it works on VPSs where microVMs do not), --network=none and DNS-scoped egress modes, and a public continuous fuzzing dashboard you can watch.
Why this rating
On the axes that matter for security, gVisor is the best-evidenced product in this guide: continuous fuzzing, tight host surface, fast silent fixes, production use at enormous scale. It is also the one most likely to break your workload in a way that wastes an afternoon, because the failure mode is a missing syscall rather than a clear error. That is an adoption-cost problem, not a trust problem, and it is why the rating stays high while the take-away is 'test your actual toolchain first'.
Our take
Use it on Linux hosts when you cannot get KVM or cannot afford a VM per agent, and you have verified that your agent's toolchain works under it. Run your real command sequence in it before committing. Do not use it for anything that needs io_uring, exotic ioctl paths, or kernel introspection.

Wasmtime

RECOMMENDED / MATURE
★ 18kforks 1.8kopen issues 848created 2017-08-29last push 2026-09-18licence Apache-2.0
Type
WebAssembly/WASI runtime
Platforms
Linux, macOS, Windows
Agent-specific
No — but it is the strongest sandbox for code-execution-shaped agent tasks.
Open source
Yes, Apache-2.0 with LLVM exception
Isolation model
Capability-based isolation at the language-runtime level rather than the OS level. A module can only reach what you hand it: specific preopened directories, specific sockets, specific environment variables. There is no filesystem to escape because there is no ambient filesystem, and no process namespace to break out of because there is no process. Compilation is ahead-of-time by default and the runtime is written in Rust with a strong record of not having memory-safety escape bugs.
Adoption evidence
The reference WASI implementation, maintained by the Bytecode Alliance with major corporate backing, used in serverless and plugin hosting, and the subject of continuous fuzzing. Its security model is discussed in remarkably sober terms by its own maintainers, including explicit notes about what it does not defend against.
Development activity
Continuous, with a large contributor base and frequent releases.
Common complaints
The complaints are about fit, not quality: WASI coverage means most real programs cannot run unmodified, threading and networking support have moved slower than people want, and the component model is only now reaching usable maturity. If your agent needs a shell, git, a package manager and a compiler, WASM is the wrong tool and you will spend a week discovering that.
Security considerations
The capability model is genuinely stronger than an OS sandbox for the class of code it can run, because the sandbox surface is the runtime's API rather than the kernel's. The caveats are honest and documented: resource exhaustion is largely your problem unless you configure fuel and epoch limits, and any host function you expose becomes part of the attack surface.
Interesting features
Sub-millisecond instantiation, fuel-based and epoch-based interruption, fine-grained capability grants, deterministic execution, and a design that makes 'what can this code reach?' a short answer.
Why this rating
Decades of sandbox research condensed into a well-maintained runtime with a narrow, well-defined boundary. It is in the top category for the workloads it fits, and the reason it is not the headline recommendation of this guide is simply that most coding agents want a shell and a full toolchain, which WASM does not provide. Judging the project itself: reliable, actively maintained, and unusually clear about its limits.
Our take
Use it for agent tool execution, plugin hosting, and any place where you can define the workload as a function rather than a shell session. Do not try to put a coding agent in it.

QEMU

RECOMMENDED / MATURE
★ 13kforks 7.2kopen issues 0created 2012-08-11last push 2026-09-19licence NOASSERTION
Type
Full machine emulator and virtualiser
Platforms
Linux, macOS, Windows hosts; nearly every guest
Agent-specific
No — it is the substrate under Lima, UTM, Qubes-style setups and much more.
Open source
Yes, GPL-2.0
Isolation model
With KVM (Linux), HVF (macOS) or WHPX (Windows) it runs guests at native speed behind a hardware virtualisation boundary, with a device model covering essentially everything: GPU passthrough, USB, audio, TPM, snapshots, live migration. The cost of that completeness is the largest attack surface in this guide.
Adoption evidence
The default virtualiser for a quarter of a century, the backend for almost every desktop VM manager, embedded in Android Studio, and the compute layer of countless cloud products. If a VM feature exists anywhere, it exists in QEMU first.
Development activity
Enormous, multi-vendor, and continuous across two decades.
Common complaints
Not bugs — ergonomics and size. Raw QEMU command lines are hostile, which is why Lima, UTM, virt-manager and Multipass exist. Performance on desktop is fine with an accelerator and painful without one. The volume of CVE traffic is high simply because the code is huge and every emulated device is a potential bug.
Security considerations
Device emulation is where VM escapes come from, and QEMU has more emulated devices than anything else. In practice you mitigate by using virtio devices, disabling what you do not need, and placing QEMU behind a sandbox (which is what Kata does). On a desktop, the practical risk to a single user is low; in multi-tenant hosting it is the whole game.
Interesting features
Completeness, mature snapshot and migration support, virtio-fs, and the fact that every problem you will hit in VM land already has a Stack Overflow answer from 2014.
Why this rating
Twenty years of continuous development, universal tooling support, and a security response process that functions despite the code volume. The complaints are about usability rather than correctness. Where QEMU is the right answer — a desktop VM you can hand a GUI to, a real Windows guest with GPU support, a machine that outlives the agent sessions on it — nothing in this guide beats it.
Our take
Do not drive QEMU directly. Use Lima, UTM or Incus and let them manage it, and then treat the VM as your outer boundary. On Windows, the equivalent comfortable path is a WSL2 distro or Hyper-V, not hand-rolled QEMU.

Cloud Hypervisor

RECOMMENDED / MATURE
★ 6.2kforks 766open issues 228created 2019-04-30last push 2026-09-19licence none detected
Type
KVM/MSHV microVM monitor (Rust VMM)
Platforms
Linux and Windows hosts (MSHV)
Agent-specific
No — but it is the strongest general-purpose VMM to build an agent sandbox on.
Open source
Yes, Apache-2.0 and BSD-3-Clause
Isolation model
A Rust VMM with a modern device model: virtio-fs, vhost-user, CPU and memory hotplug, VFIO passthrough, snapshotting and live migration, and a per-thread seccomp BPF filter on its worker threads. It boots a real kernel per guest. Compared with Firecracker it gives up a little of the minimalism in exchange for the features developers actually need, and compared with QEMU it gives up almost all of the attack surface.
Adoption evidence
Chosen by Microsoft for Kata-based Pod Sandboxing on AKS, used by Fly.io for GPU machines, a supported Kata backend, the primary VMM in Spectrum OS, supported by microvm.nix, and shipping its 52nd major release since 2019. Fifty-plus releases and a Linux Foundation project is a strong longitudinal signal.
Development activity
Continuous, with a disciplined release train and parallel patch branches. Commits land daily across many contributors.
Common complaints
Most complaints are about integration effort rather than the VMM: you still need a kernel, a rootfs, a virtiofs share and network plumbing. Feature-wise the historical gap versus crosvm was device breadth. Downstream pinning is a real hazard — one product in the independent study sat on a Cloud Hypervisor version for 471 days across 12 upstream releases and missed an escape-class fix.
Security considerations
The 2026 study classifies it as coordinated: every in-window CVE shipped with an advisory and a patched release the same day. It is also the engine with the first published Escape-class advisory of its own (CVE-2026-45782, a virtio-block async-I/O use-after-free at 8.9), which is exactly what you would expect from a project that discloses properly instead of one with no advisories at all. Its nested-KVM exposure was flagged as a product-level configuration question, not an engine bug.
Interesting features
virtio-fs (the thing Firecracker lacks), hotplug, vhost-user for offloading devices to other processes, snapshot and restore, live migration, a typed Rust API, and a test suite that treats boot correctness as a first-class contract.
Why this rating
It has the features that make microVMs usable for developer workloads, a release cadence you can plan around, hyperscale production deployments, and — most importantly for this guide — evidence that it fixes security bugs fast and loudly. The complaints that exist are about the work required to assemble a microVM stack, which is inherent to the category, not about the project being flaky. For anyone building a strong-isolation sandbox on Linux, this is the VMM to start with.
Our take
Pick Cloud Hypervisor for phase-two strong isolation on Linux, especially if you need to share a directory (virtio-fs) or pass a device. Build a snapshot pool so starts feel instant, pin versions deliberately, and watch the release notes rather than assuming your pinned version is fine.

Incus

RECOMMENDED / MATURE
★ 6.2kforks 489open issues 46created 2023-07-22last push 2026-09-18licence Apache-2.0
Type
System container and full VM manager
Platforms
Linux hosts; API and CLI clients elsewhere
Agent-specific
No — but it is the best 'my own worker machine' backend for persistent agent environments.
Open source
Yes, Apache-2.0
Isolation model
Two levels in one tool: LXC system containers that run a full init and behave like a small machine, and QEMU/KVM virtual machines for a real kernel boundary. Storage is pluggable (ZFS, Btrfs, LVM, dir) which decides how cheap snapshots and clones are; with a CoW backend, forking an instance is close to free. Networking is managed with bridges, and profiles let you describe an environment once and instantiate it many times.
Adoption evidence
Six thousand stars, the continuation of LXD under the Linux Containers organisation, used widely for homelab and small-fleet work and specifically recommended in Hacker News threads about agent isolation. Published comparisons in 2025–2026 describe it as the easiest way to get both containers and VMs from one tool.
Development activity
Continuous and vendor-supported, with frequent releases and an LTS cadence.
Common complaints
Documentation depth versus Docker's ecosystem is the usual complaint, plus a learning curve for storage pools and profiles. Storage-driver differences surprise people: snapshot performance is excellent on ZFS/Btrfs and ordinary on dir. System containers share the kernel, so the same shared-kernel caveats as any container apply, and it is easy to forget which of your instances is a VM and which is a container.
Security considerations
Pick the VM type for untrusted code and the container type for convenience. Incus does not sandbox QEMU itself, it configures and invokes it, so QEMU's device surface is your exposure. Restricting an instance's access to the host filesystem is a matter of what you mount into it, which means the usual discipline about not sharing your home directory.
Interesting features
Snapshots and instant clones, projects for tenancy separation, profiles for reproducible agent environments, both containers and VMs behind one API, and a REST API that a GUI can drive directly.
Why this rating
Mature, well-maintained, and unusually good at the specific thing an agent power-user needs: cheap, reproducible, snapshot-able machines that you can fork per task. Complaints are about learning curve and documentation rather than reliability, and the project has an eight-year-old lineage with active maintainers. The absence of Windows host support is the main structural limitation.
Our take
If you run Linux and want a pool of persistent agent machines, or you want to fan out without paying VM boot costs, build on Incus with a CoW storage backend. Use VMs for anything you would not run as your own user, containers for the rest, and never share your home directory into either.

GOOD, BUT TEST IT FIRST

Strong enough to investigate, with meaningful caveats.

Claude Code

GOOD, BUT TEST IT FIRST
★ 146kforks 23kopen issues 12kcreated 2025-02-22last push 2026-09-19licence none detected
Type
Coding agent CLI with a built-in sandboxed Bash tool
Platforms
macOS, Linux, Windows via WSL2 (native Windows is not supported for the sandbox)
Agent-specific
Yes
Open source
The CLI distribution is not open source; the sandbox engine it uses is published as @anthropic-ai/sandbox-runtime.
Isolation model
Seatbelt on macOS and bubblewrap on Linux/WSL2, layered over an approval model. Anthropic publishes the sandbox runtime separately, and documents that cloud sessions run in full microVMs. Filesystem and network policy are configurable, and an inspecting proxy handles network mediation. The sandbox is not mandatory for local Bash execution.
Adoption evidence
The most widely used coding agent, with a documented security architecture and multiple public incident write-ups, which is more transparency than most of this category offers. Its containment design is also the subject of independent reverse-engineering talks.
Development activity
Constant. Releases land weekly or faster, and the sandbox has been reworked in response to reported bypasses.
Common complaints
The most consistent complaint is architectural rather than bug-level: the sandbox is opt-in, full filesystem read is the default, and there is a code path where the agent can disable its own sandbox to finish a task. Reported bypasses include a symlink escape (CVE-2026-39861), a denylist/dynamic-linker evasion documented by the author of Falco, and a config-TOCTOU issue where a settings file that did not exist at startup could be created inside the sandbox. On WSL2, Windows interop remains the sharp edge.
Security considerations
Treat the built-in sandbox as one wall, and do not treat 'Claude Code says it is sandboxed' as a supply-chain guarantee. The symlink and denylist bypasses both required no jailbreak — the agent simply wanted to finish the task. The officially recommended pattern for real autonomy is the dev container image, which is audited and puts a container between the agent and your workstation.
Interesting features
A devcontainer image hardened for bypass mode, a sandbox runtime that is reusable outside Claude Code, and permission modes that are actually documented rather than folklore.
Why this rating
Enormous adoption plus public incident reporting makes this one of the better-understood agent sandboxes in the field, and the escape classes are known, patched and, where they involve Windows, explicitly acknowledged. The unresolved risk is design intent: defaults that assume the agent is well-behaved, and an approval model that the agent can argue with. Users keep rediscovering the same class of problem because the default configuration is more permissive than the average developer assumes.
Our take
If you run Claude Code on your own machine, run it inside the devcontainer or a VM and keep the sandbox on inside that as a second wall. Read the security page before you turn on bypass permissions, and do not let the agent's own judgement be the reason the sandbox got turned off.

OpenAI Codex CLI

GOOD, BUT TEST IT FIRST
★ 125kforks 19kopen issues 17kcreated 2025-04-13last push 2026-09-19licence Apache-2.0
Type
Coding agent CLI with a built-in OS sandbox
Platforms
macOS (Seatbelt), Linux (Landlock + seccomp, or a newer bubblewrap pipeline), Windows (restricted tokens / AppContainer)
Agent-specific
Yes
Open source
Yes, Apache-2.0
Isolation model
A dual-mode Linux sandbox: a legacy Landlock pipeline and a modern bubblewrap pipeline that uses --ro-bind /, --unshare-pid, --unshare-net, /proc, PR_SET_NO_NEW_PRIVS and a seccomp network filter, with TCP bridged through Unix domain sockets and socket creation blocked after setup. macOS uses Seatbelt. Windows uses a composition of a synthetic SID, a write-restricted token, dedicated local users and, in elevated setups, Windows Firewall rules. The security boundary depends on the mode and platform; the documented modes are read-only, workspace-write and danger-full-access.
Adoption evidence
Well over a hundred thousand stars and the second most widely installed coding agent after Claude Code. Its Windows sandbox work is the most thoroughly documented public account of what it actually takes to sandbox an agent on Windows, and it was publicly credited when a WSL interop escape was reported and fixed.
Development activity
Extremely active, with frequent releases and a large contribution base. The sandbox has already been reworked at least twice in a year, including a PR that explicitly blocks WSL interop escapes from restricted filesystem sandboxes.
Common complaints
Windows is where the friction lives: unelevated mode cannot enforce network policy and falls back to advisory controls such as dead proxy variables and PATH pollution that raw sockets bypass, so the project moved to a one-time elevated setup for kernel-level firewall enforcement. Restricted-token schemes leak through Everyone-writable directories and NTFS hard links, which is why the implementation reports partial enforcement. PowerShell AST inspection has opaque regions (Invoke-Expression, -EncodedCommand, dynamic ScriptBlock creation) that were only fixed after being reported. Elsewhere, Landlock is port-level only for network, so domain policy still needs a proxy.
Security considerations
The Linux path is a genuine two-wall OS sandbox and is probably the strongest default you get for free from any agent. It is still a shared-kernel sandbox. On WSL2 the interop channel was a complete bypass until it was masked — a useful reminder that on Windows the sandbox is only as good as its handling of the host bridge. Configuration-based escapes (writing the CLI's own configuration from inside the sandbox) have been demonstrated against restricted-token backends.
Interesting features
Sandbox modes that are visibly selectable and documented, an explicit network-allowlist proxy, a real Windows implementation with an honest partial-enforcement story, and public engineering write-ups about which primitives failed and why.
Why this rating
The sandbox is unusual in being default-on and in being documented with its own limitations rather than marketed. The caveats are real but they are the honest kind: platform asymmetry, advisory-only network control without elevation, and known leaks in the Windows token model. Compared with the alternative — a wrapper you install yourself and never upgrade — a sandbox that ships with the agent and gets fixed in the same release cycle is the safer of the two.
Our take
Leave the sandbox on, use workspace-write rather than full access, and if you are on Windows and care about egress, do the one-time elevated setup. If you need more than this provides, put Codex inside a container or VM rather than disabling the sandbox.

WSL2

GOOD, BUT TEST IT FIRST
★ 33kforks 1.8kopen issues 981created 2016-04-06last push 2026-09-19licence MIT
Type
Hyper-V utility VM running a real Linux kernel on Windows
Platforms
Windows 10/11
Agent-specific
No — but most Windows agent sandboxes are really Linux sandboxes running inside it.
Open source
Yes for large parts of the project, MIT
Isolation model
A genuine lightweight VM managed through the Host Compute Service, with a real Linux kernel that Microsoft maintains, plus WSLg for Wayland/X11 GUI applications and mount points that expose the Windows filesystem as /mnt/c, /mnt/d and so on, mapped through 9P/virtiofs with Windows ACLs translated to Linux permissions. The critical detail is interop: Linux processes can launch native Windows executables, which are spawned by the WSL init as real Windows processes running with the current Windows user's full token.
Adoption evidence
Thirty-three thousand stars on the public repository, installed by essentially every Windows developer, the substrate for Docker Desktop, Podman Machine's default provider and most Windows agent tooling. The evidence base for how it fails is equally rich, because so many people hit it.
Development activity
Very active, vendor-maintained, with regular kernel and userspace updates.
Common complaints
The interop channel is the big one and it is not a bug, it is the design. A bubblewrap sandbox inside WSL2 that makes /mnt/c read-only from the Linux side can be bypassed completely by invoking powershell.exe: the call is intercepted by WSL init, which spawns a Windows process outside the Linux mount namespace, with full user privileges, able to read and write every drive. One 2026 proof of concept deleted 26.8 GB from D: with no approval prompt. Codex shipped a fix that masks interop sockets from restricted filesystems and denies AF_VSOCK and io_uring; Claude Code documents that native Windows is unsupported and recommends WSL2, which does not remove the interop caveat by itself. Filesystem performance across the boundary is mediocre for many-small-file workloads. Microsoft's own servicing criteria do not treat shared-kernel or interop behaviour as a security boundary. Docker's Enhanced Container Isolation documentation concedes that on WSL a user can enter the Docker Desktop distro as root and that all distros share one kernel.
Security considerations
WSL2 gives you real memory and kernel separation, which is a large improvement over WSL1 and worth having. It does not protect the host from a workload running as you, and the deliberate bridges — interop and /mnt — are where any Linux-side sandbox leaks. Mounting the Windows filesystem read-write into an agent's environment hands it your user's files regardless of what the Linux sandbox says. The mitigations are known and all cost convenience: set interop=false, mask interop sockets inside the sandbox, restrict drvfs mounts, and run as a non-admin Windows account.
Interesting features
Real Linux with systemd, WSLg giving Linux GUI applications natural Windows windows, fast startup, and the fact that it is the only way to run Docker or Podman natively on Windows without a separate VM.
Why this rating
As an operating environment it is excellent and universally used. As a security boundary for an agent it is routinely misrepresented, including by tool documentation that says 'runs in WSL2' as if that were a containment claim. The interop escape is simple, demonstrated, and the product of deliberate design rather than negligence. It remains the best available Windows substrate for Linux-shaped agent work, provided you treat the boundary as blast-radius reduction and close the bridges you are not using.
Our take
Use WSL2 as your Windows foundation, and then disable what you do not need: interop off inside agent sandboxes, /mnt mounted read-only or not at all, no admin token. If you want a real boundary on Windows, the honest options are Windows Sandbox or Hyper-V, with the agent inside and no host filesystem shared.

Infisical / Agent Vault

GOOD, BUT TEST IT FIRST
★ 29kforks 2.3kopen issues 793created 2022-08-05last push 2026-09-19licence NOASSERTION
Type
Secrets platform with a credential-brokering reverse proxy
Platforms
Server plus agents on Linux, macOS, Windows; the proxy is designed to run on a separate host
Agent-specific
Yes — Agent Vault exists specifically to give agents authenticated API access without the key.
Open source
Yes for the core (mixed licensing, with an open-source agent-vault component); a commercial hosted product also exists.
Isolation model
Not isolation — brokering. The agent is configured with a placeholder such as __anthropic_api_key__; a TLS-terminating forward proxy outside the agent resolves the real credential from the vault and substitutes it on egress. The docs recommend running the proxy on a different host from the agent so a filesystem read cannot recover the secret. It is interface-agnostic because everything bottoms out at HTTP, and traffic policy supports allow-all or host allowlists.
Adoption evidence
Infisical is a mainstream secrets platform with nearly thirty thousand stars. The Agent Vault pattern is cited alongside Anthropic, Vercel and Cloudflare implementations of the same idea, and the IETF has a draft standard for credential brokers for agents, so this is becoming infrastructure rather than a hack.
Development activity
Active, with the agent proxy as a newer component than the core platform.
Common complaints
Newness, mostly: the broker is much younger than the vault, and the licensing split between open-source and commercial components is confusing. Teams that centralise on a third-party proxy accept a new dependency in the most sensitive position in their stack. Operational details such as CA distribution into the sandbox and proxy-bypass behaviour need care, exactly as with mitmproxy.
Security considerations
This is the correct architectural answer to agent credential theft, and it has a sharp edge worth understanding: the proxy is now the thing holding every real secret, so it must live outside the sandbox, be unreachable by the agent except through its legitimate port, and never fail open. Two of nono's published advisories were precisely an allow-all default and a proxy-only fallback, so check the default of whichever broker you adopt.
Interesting features
Placeholder semantics that work with tools that were never designed for brokering, host allowlists, per-consumer configuration, and the fact that the agent holds nothing worth stealing even if it is fully compromised.
Why this rating
The pattern is right, the implementation is mainstream, and the project is active. It lands in 'test first' rather than 'mature' because the brokering component is young relative to the vault, the licensing split makes auditing boundaries unclear, and being the last hop before your production APIs means you must verify its fail-closed behaviour yourself rather than assume it.
Our take
Adopt the pattern immediately; adopt this implementation if you already use Infisical or want a maintained broker instead of writing one with mitmproxy. Run it on a separate host or container, pin it, and test what happens when it is unreachable.

c/ua (Cua)

GOOD, BUT TEST IT FIRST
★ 24kforks 1.7kopen issues 1kcreated 2025-01-31last push 2026-09-19licence MIT
Type
Computer-use suite: drivers, local macOS/Linux VMs, cloud fleets, benchmarks
Platforms
macOS and Linux locally; cloud fleets for Windows, Linux, macOS and Android
Agent-specific
Yes
Open source
Yes, MIT for the open components; the cloud fleet product is commercial
Isolation model
Layered by design. Locally, Lume runs macOS and Linux VMs on Apple's Virtualization framework at near-native speed. In the cloud, fleets give each agent an isolated desktop. Drivers operate applications either through accessibility trees or pixel-level control. The isolation story depends on which piece you use, and the most common failure is not using the VM at all and pointing a driver at your real desktop.
Adoption evidence
Twenty-four thousand stars, Y Combinator backing, a large contributor base, active releases, benchmark tooling, and a distributed image/VM registry. It is the most substantial open project in computer use.
Development activity
Very active, with frequent commits across drivers, VM tooling and benchmarks.
Common complaints
macOS and Apple Silicon centricity is the standing complaint: the local VM path needs an M-series Mac and macOS Tahoe for full function. Setup complexity is high because there are four products in one repository. Drivers on the host, if you use them without a VM, are exactly as dangerous as any desktop automation. Documentation is extensive but sprawling.
Security considerations
The threat model here is prompt injection through everything the agent can see — web pages, documents, screenshots, UI text — and the published benchmark evidence is bleak: attack success rates against computer-use agents reach tens of percent on realistic adversarial tasks, with some variants approaching half. The mitigation is architectural: run the desktop in a VM that holds no real credentials, use view-only or password-protected streams, and require human confirmation for consequential actions.
Interesting features
Local Apple-silicon VMs at high native throughput, cross-OS cloud fleets including an Android option, accessibility-tree driving that is far more token-efficient than screenshots, and a benchmark suite for evaluating computer-use agents.
Why this rating
The engineering is real, the licence is permissive, and this is the most complete open toolkit for the problem. It is not 'mature' in the sense of being a boring dependency: it is a fast-moving multi-product suite whose security depends almost entirely on whether you actually put the desktop in a VM, and whose local path is tied to one hardware vendor. Use it with your eyes open about what part you are adopting.
Our take
Adopt the VM-based path, not the host-driver path. Assume anything the agent can see can instruct it, keep credentials out of the guest, and use the streaming surface to supervise rather than to trust.

E2B

GOOD, BUT TEST IT FIRST
★ 13kforks 1kopen issues 72created 2023-03-04last push 2026-09-19licence Apache-2.0
Type
Hosted agent sandbox platform (Firecracker microVMs)
Platforms
Cloud; self-hosting on your own cloud exists in an experimental state
Agent-specific
Yes — it is the reference product for the category.
Open source
Yes, Apache-2.0 for the SDKs and much of the stack; the hosted service is commercial.
Isolation model
One Firecracker microVM per sandbox, each with its own kernel, started from a pre-warmed snapshot pool, with root inside the sandbox conferring no host privilege. Sandboxes have an isolated network stack, and egress can be restricted with allow/deny lists by hostname or CIDR.
Adoption evidence
Fourteen thousand stars, SDKs in many languages, broad framework integrations, and a customer case study describing it as the only provider that worked first time. It is the sandbox that most agent frameworks integrate first.
Development activity
Very active, both in the public SDK repositories and in the closed infrastructure.
Common complaints
Data-loss-shaped friction is the recurring pattern: a default five-minute suicide timer that made files vanish, a fragile manual pause API, module-level shared sandbox state that breaks under concurrency, and silent sandbox replacement when a command times out. The documented fix is turning on auto-pause, which is exactly the kind of default that should be on. Runtime caps (one hour on the hobby tier, 24 hours on Pro) drop in-memory state unless you plan for it. Self-hosting is repeatedly described as not production-ready, requiring you to operate KVM, tap networking, cgroups, snapshot pools and orchestrator failover. There is no GPU support, and static egress IPs need your own proxy. The independent security study also flagged a self-hosted default pin that was more than forty days behind a Firecracker fix, while the hosted pin is opaque.
Security considerations
The engine choice is strong: per-sandbox kernels mean a kernel bug is not automatically a multi-tenant breach, which is the structural argument for microVMs in the first place. E2B's own documentation is explicit that sandboxes can reach the open internet by default and recommends restricting egress. The engineering risk is the pin policy on self-hosted installs, which is under your control and therefore your problem.
Interesting features
Sub-200ms starts from snapshots, pause/resume that preserves memory, template-based environments, a desktop sandbox image for computer use, and a straightforward allowInternetAccess plus allowlist control.
Why this rating
The isolation model is the right one and the platform is real and widely used. What keeps it out of the top category is the operational and data-integrity story: defaults that delete your files unless you configure them otherwise, silent sandbox churn on timeout, tier-based expiry that discards memory, and a self-hosting path that the community consistently describes as experimental. None of that is a security failure, and all of it costs a team time.
Our take
If you want a hosted sandbox and you do not want to run infrastructure, this is a reasonable default — but turn on auto-pause on day one, set explicit timeouts and egress policy, and do not treat 'the sandbox is isolated' as meaning 'the sandbox is durable'. Self-host only if you can operate KVM-era infrastructure.

Anthropic Sandbox Runtime (@anthropic-ai/sandbox-runtime)

GOOD, BUT TEST IT FIRST
★ 5.3kforks 449open issues 212created 2025-10-20last push 2026-09-19licence Apache-2.0
Type
OS-level process sandbox library (the engine behind Claude Code's sandboxed Bash tool)
Platforms
macOS (Seatbelt), Linux and WSL2 (bubblewrap), with an optional network proxy
Agent-specific
Yes — built for CLI coding agents, and generalised as a library for arbitrary processes.
Open source
Yes, Apache-2.0
Isolation model
Two walls, no container. On macOS it emits a Seatbelt profile through sandbox-exec; on Linux it uses bubblewrap with namespaces, bind mounts, no-new-privileges and a seccomp network filter. Because neither kernel primitive can filter by hostname, network control comes from an HTTP proxy that the sandboxed process must honour. Filesystem restrictions are implemented as bind mounts and denials, which means the boundary is only as good as the path list it was given.
Adoption evidence
Shipped to a very large installed base through Claude Code, cited as the reference implementation of the two-wall pattern in Anthropic's own engineering posts, and widely re-used as a library elsewhere. It is one of the few agent sandboxes with a genuinely huge user count.
Development activity
Active and fast-moving under Anthropic. Both the library and the Claude Code sandbox have changed shape multiple times in 2026.
Common complaints
The design defaults, not the code, draw most of the criticism. Reviewers object that the sandbox is opt-in, and that read access to the whole filesystem is permissive by default. Bubblewrap's mandatory-deny can only block paths that already exist, so a config file that does not exist yet is not protected — the shape of the CVE-2026-25725 class of bug, where a sandboxed process creates .claude/settings.json and injects a host-privileged hook. Proxy-based network policy depends on the tool actually respecting HTTP_PROXY, and Node/undici historically does not. On WSL2 the Windows interop channel hands execution straight back to Windows as your user unless it is explicitly masked.
Security considerations
This is an OS sandbox sharing the host kernel. It protects against accidents and reduces exfiltration, and it does not protect against a kernel exploit or a determined process with a socket to talk to. Its asymmetry with macOS is worth knowing: Seatbelt globs block files that do not exist yet, bubblewrap bind mounts cannot.
Interesting features
Cross-platform without a container or VM, an explicit network proxy layer, a genuinely readable implementation, and the reference documentation for what 'sandbox the commands, not the model' means in practice.
Why this rating
The engineering is good and the deployment base is enormous, which means bugs surface quickly and get fixed quickly. But the project has an unusually public track record of bypasses: a symlink escape, a TOCTOU-style config injection, and repeated findings that the agent will reason its way around a denylist or turn its own sandbox off. That is not a criticism of the maintainers so much as a demonstration that an OS sandbox plus an allow-by-default filesystem is a weak first line for autonomous agents. It is worth deploying as one layer, with a real network policy, and it should not be your only layer.
Our take
Use it deliberately, with explicit filesystem paths, network on allowlist and interop masked on WSL2 — or better, run it inside a container or VM rather than directly on your workstation. Read its issue tracker before you trust its defaults.

nono (nolabs)

GOOD, BUT TEST IT FIRST
★ 4.1kforks 274open issues 213created 2026-01-31last push 2026-09-18licence Apache-2.0
Type
Kernel-enforced OS sandbox for agents, with credential brokering
Platforms
Linux (Landlock + seccomp), macOS (Seatbelt), Windows via WSL2 only
Agent-specific
Yes
Open source
Yes, Apache-2.0
Isolation model
No daemon, no container, no VM. Landlock layered with seccomp-notify for path mediation on Linux; Seatbelt on macOS. Capability sets are declared per run. The interesting half is the credential layer: instead of exporting real keys, the agent gets a local proxy URL and a session-scoped phantom token, and the reverse proxy strips the phantom at the boundary and injects the real credential from the OS keyring or a password manager, held in Zeroizing memory in a separate process the child cannot read or ptrace. The proxy does its own DNS resolution and blocks link-local and cloud-metadata addresses to prevent rebinding, and uses constant-time token comparison so neighbour processes on localhost cannot borrow the proxy.
Adoption evidence
Four thousand stars, 80+ contributors, backed by the people behind Sigstore, a product site, packaged integrations for a long list of agents, third-party coverage from security press, and a Kubernetes Agent Sandbox example. Reported production users include a large observability vendor.
Development activity
Very active, pre-1.0, releasing often (0.5x through 0.7x during 2026).
Common complaints
Pre-1.0 API and profile churn. The 2026 CVE record is the honest caveat: CVE-2026-47128 was a complete sandbox escape via the per-user systemd D-Bus socket (fixed in 0.55.0), GHSA-hc4m-q9jh-xw4j was a fail-open in registry pack verification where deleting trust metadata made a pack easier to run (fixed in 0.61.3), and three moderate advisories affected the Python bindings, including an allow-all default in the credential proxy's empty-host-list case. Windows support is WSL2-only with documented feature gaps. The user who first pointed this site's author at nono reported leaks, and the project's own CVE history shows why that impression existed.
Security considerations
This is an OS sandbox, not hypervisor-grade isolation, and the project says so. The escape history is instructive rather than disqualifying: every issue was found, fixed, and disclosed with an advisory, and the D-Bus escape class is a general lesson for every Landlock-based tool, because allowing local sockets without path mediation is enough to leave the sandbox. The credential design is a genuine improvement over ambient environment variables and is the reason this project matters even if you never install it.
Interesting features
The phantom-token credential injection pattern with header, url_path, query_param and basic_auth modes; proxy-side AWS SigV4 signing and OAuth client-credentials exchange; OAuth login capture with a phantom store outside the sandbox; PATH sanitisation for host-side brokers so the sandbox cannot plant a same-named binary; and a documented security model that enumerates what it does not protect.
Why this rating
The idea is the best in its class and the implementation is real, actively maintained and heavily contributed to. What keeps it out of the top category is the combination of pre-1.0 churn and a CVE history that includes a full sandbox escape and a fail-open in the verification path, plus a Windows story that is WSL2-only. Those are fixable problems with fast, disclosed fixes; they are also exactly the problems you do not want to discover on your own machine.
Our take
Deploy the credential-injection pattern today, whether or not you deploy nono. If you do install it, pin a version at or above 0.62.0, read the security model first, keep an eye on advisory feeds, and do not treat the OS sandbox layer as protecting you from hostile code — pair it with a VM if that is your threat model.

container-use (Dagger)

GOOD, BUT TEST IT FIRST
★ 4kforks 205open issues 62created 2025-05-23last push 2026-09-14licence Apache-2.0
Type
Per-agent containerised worktrees over MCP
Platforms
Linux, macOS
Agent-specific
Yes
Open source
Yes, Apache-2.0
Isolation model
Each agent gets its own container and its own Git branch, orchestrated through Dagger's engine, with an MCP interface so any MCP-capable agent can use it. Isolation is container-level, and the review model is Git-shaped: the agent works in a branch and you review a diff rather than trusting a shell session.
Adoption evidence
Four thousand stars, built by Dagger (a well-funded infrastructure company), recommended in Hacker News threads as the natural answer to per-agent containerisation, and cited repeatedly in the 'how are you sandboxing agents' discussions.
Development activity
Active, with regular releases and a maintained engine behind it.
Common complaints
Dagger's engine is the dependency, and that means a daemon, a build graph and a learning curve for teams that just wanted a container. Environment definition is code, which is powerful and also more work than a Dockerfile. Container-level isolation shares the kernel, so users looking for hypervisor-grade separation need to look elsewhere. Sharing caches between containers takes deliberate configuration.
Security considerations
Container isolation plus Git-branch review is a genuinely strong combination for the accident and exfiltration tiers: even if the agent misbehaves, the host workspace is not the thing being edited, and review is a diff. It does not protect against a kernel-level escape, and the engine's own privileges become part of your trust base.
Interesting features
Git-native review, environment-as-code, reuse of the container ecosystem, and an MCP interface that means it works with whichever agent you prefer rather than locking you to one.
Why this rating
The design is one of the best answers in the guide to 'what should the agent's workspace actually be', and it is maintained by an organisation whose main product is exactly this kind of environment plumbing. The caveats are the generic container ones plus the operational weight of adopting Dagger's engine. Judged on evidence: real adoption, real maintenance, no significant reliability complaints relative to usage.
Our take
A strong choice if you are already comfortable with containers and want agent work to arrive as reviewable diffs. If turning on a Dagger engine for a five-minute task feels heavy, a plain Podman container plus git worktrees gets you most of the benefit.

agent-desktop

GOOD, BUT TEST IT FIRST
★ 1.3kforks 86open issues 22created 2026-02-19last push 2026-09-17licence Apache-2.0
Type
Rust CLI for desktop automation through accessibility trees
Platforms
macOS today; Windows and Linux listed as planned
Agent-specific
Yes — it is the computer-use layer, not an isolation layer
Open source
Yes, Apache-2.0
Isolation model
None. It runs on your desktop as your user and gives agents reliable, stable references to real UI elements, with retry-safe actions rather than coordinate guessing. The stated advantage is reliability: accessibility snapshots can be a fraction of the size of screenshots, with focused drilling for individual controls, which also makes the injection surface smaller than a pixel-based agent's.
Adoption evidence
Thirteen hundred stars, regular releases, npm distribution, a CI pipeline, and listings in agent skill directories. The engineering quality is visibly high — clear concepts documentation, a changelog, a security policy — but adoption is concentrated on macOS.
Development activity
Active, with frequent release-bot activity and substantive commits.
Common complaints
Platform coverage, mainly: macOS-only today, with Windows and Linux promised. Anything that drives the real desktop inherits the usual complaint about agents being given the keys to the machine. Accessibility trees do not cover every application, so some workflows fall back to other mechanisms.
Security considerations
The trust boundary is your session. Because it reads accessibility trees rather than screenshots, some classes of visual prompt injection do not apply, but any text it reads from an application is still attacker-controllable input that reaches the model. Running it inside a VM would be the sane configuration and is not what most users do.
Interesting features
Stable element references, retry-safe actions, an accessibility-tree snapshot format with skeleton and drill-down modes that cuts token cost dramatically, and a genuinely readable design document set.
Why this rating
High-quality engineering in a young category, with an approach (accessibility trees over pixels) that solves real problems of reliability and token cost. It is not an isolation product at all, and on the platforms this guide prioritises — Windows and Linux — it does not yet exist. Rate it as a promising component to watch, not as something to depend on now.
Our take
Watch it, and if you are on macOS, it is a reasonable driver to pair with a VM-based sandbox. Do not describe it to users as containment.

Sculptor (Imbue)

GOOD, BUT TEST IT FIRST
★ 233forks 17open issues 20created 2025-08-07last push 2026-09-19licence MIT
Type
Desktop workspace for running parallel coding agents in isolated workspaces
Platforms
macOS and Linux desktop
Agent-specific
Yes — this is the closest existing product to what a fan-out GUI would be.
Open source
Yes, MIT
Isolation model
Each workspace gets its own isolated Git worktree, branch, terminal and diff view, with containerised execution and model/provider switching. Isolation is workspace-level and container-level rather than hypervisor-level.
Adoption evidence
Hacker News traction of 176 points and 85 comments for the launch, plus continued discussion. The category-defining claim in those threads was not the sandbox but the at-a-glance view of five agents at once, which validates the whole idea of a control plane.
Development activity
Active, with regular releases and funded development.
Common complaints
Weight and complexity, and a meta-complaint that surfaced repeatedly in the launch thread: how much of the UI is actually useful versus theatre. Some users wanted live streamed output and stronger keyboard navigation. It is also, bluntly, a competitor rather than a dependency for anyone building the same thing.
Security considerations
Worktrees isolate working state and containers isolate execution, which covers the accident tier well. The security story is thinner than the parallelism story: there is no visible credential brokering, and the container boundary is the usual shared-kernel one.
Interesting features
Parallel-agent UX, per-agent diffs, model switching, and a real answer to the 'which agent is stuck?' problem that terminal tabs alone do not solve.
Why this rating
It works, it is maintained, and it is the best available proof that developers want a visual control plane for parallel agents. It is not a strong-isolation product, and it is not something you would consume as a library. Rate it high as prior art and as a product to compare against, and treat its container-level isolation as the floor rather than the ceiling.
Our take
Study it closely, especially the workspace-and-diff model. If your interest is strong isolation, take the UX ideas and pair them with a VM-based provider instead of copying the sandbox layer.

EXPERIMENTAL / EARLY

Interesting technology, currently risky to depend on.

Microsandbox

EXPERIMENTAL / EARLY
★ 8.3kforks 443open issues 93created 2024-10-03last push 2026-09-19licence Apache-2.0
Type
Local-first microVM runtime and library aimed squarely at agent sandboxes
Platforms
Linux/KVM, Apple Silicon macOS, Windows via WHP
Agent-specific
Yes — branching, snapshots, allowlisted networking and SDKs are shaped around agent workflows.
Open source
Yes, Apache-2.0
Isolation model
Hardware-isolated microVMs built on a fork of libkrun, with OCI images converting into microVM root filesystems, named long-running sandboxes, live branching, snapshots and restores, and directory/volume concepts. There is no required long-running daemon, and there are Rust, Python, TypeScript and Go SDKs. It is the closest thing in the ecosystem to a purpose-built sandbox backend for an agent product.
Adoption evidence
Eight thousand stars in under two years, a genuinely enthusiastic Hacker News launch (402 points, 186 comments), and organic mentions in other threads about microVM tooling. The creator engages directly with early adopters. It is also labelled beta in its own crate documentation, and independent security work treats its residual-bug posture as the riskiest of the microVM products measured.
Development activity
Very active, with a one-to-three-month release cadence and a changing API surface.
Common complaints
The issue tracker is unusually informative and unusually full. A virtio-fs readdir memory leak under heavy host directory load (npm install was the reproducer), fixed; Windows WHP vCPU triple-faults on AMD Ryzen hardware where every image failed to boot; a Windows "exited before agent relay" failure; boot regressions and boot-argument limit failures traced into the forked libkrun; snapshot restore breaking when an upstream image tag republished (digest mismatch); intermittent guest hangs on first outbound TLS under domain egress rules; secret-substitution failures over both HTTP and HTTPS proxies; and a 2026 CVE where secrets passed on the command line were readable through /proc. Users on VPSs without nested virtualisation found it simply unusable.
Security considerations
The engine-level picture is mixed and the project is honest about why. libkrun has zero published CVEs, no upstream fuzzer and no published academic study, so the negative finding is genuinely ambiguous. The independent study measured eleven of fourteen reachable host primitives, with libkrun shipping no engine-side seccomp filter, and concluded that the operator must stack all the hardening themselves. Leakage measurement, on the other hand, was the cleanest of the microVMs tested. In other words: good at not leaking information, unmeasured at resisting a determined escape.
Interesting features
Live branching of a running sandbox, snapshots and restores, an MCP server, OCI image support without a daemon, a genuinely pleasant SDK surface, and filesystem work that turned into a public engineering post about deleting the filesystem to make it 47 times faster.
Why this rating
This is the most interesting sandbox runtime in the guide and it is not yet something to bet a product on. The recurring failures are not cosmetic: boot failures on whole classes of Windows hardware, snapshot integrity breaks, memory leaks under ordinary npm workloads, and a CVE where credentials were readable from the process listing. Every one of those was fixed promptly, and the maintainers document their own beta status. But complaints per user are high relative to its age, the platform matrix is not uniformly working, and the security posture depends on you stacking hardening the project does not ship.
Our take
Prototype with it. Use it as an optional strong-isolation backend behind an interface, never as the single hard dependency of an application. Pin versions, run it on hardware you have actually tested, and expect to do your own soak testing. If you need microVM isolation in production today and you are Linux-only, the boring answer is still a normal VM or Kata.

BoxLite

EXPERIMENTAL / EARLY
★ 2.3kforks 180open issues 294created 2025-12-07last push 2026-09-19licence Apache-2.0
Type
Embeddable microVM runtime for OCI images (local library, CLI and optional server)
Platforms
Linux x86_64/ARM64, Apple Silicon macOS, and Windows only through WSL2. Native Windows is explicitly unsupported, and macOS Intel is listed as coming soon. Linux requires access to /dev/kvm.
Agent-specific
Yes — but for agents in general rather than one agent. It presents itself as 'the compute substrate for AI agents', with an embedded SDK, snapshot/clone and an OCI-image workflow; it is not tied to a particular coding CLI.
Open source
Yes, Apache-2.0, with a security policy that routes reports through GitHub private advisories and states acknowledgement targets.
Isolation model
A hardware-isolated microVM per box. Each box runs its own Linux kernel through libkrun, with KVM on Linux and Hypervisor.framework on macOS, and an OCI container runs inside the guest. The project's architecture notes describe a second layer, a Firecracker-jailer-inspired sandbox around the VMM process: seccomp BPF, namespace and pivot_root isolation, privilege dropping and cgroups v2 on Linux, sandbox-exec on macOS. Storage is a per-box QCOW2 disk with copy-on-write, so a box can be snapshotted, cloned and rolled back. Network egress can be restricted with an allow_net list, and secrets are injected as placeholders so that real values are not placed inside the guest. That is a well-shaped design; every word of it is the vendor's own description, and none of it has been independently measured.
Adoption evidence
An ecosystem list rather than a track record: Databricks Omnigent, Alibaba's AgentScope Runtime and ByteDance's deer-flow name BoxLite as an available sandbox backend in their own repositories and documentation. Those are statements by the consuming projects, and they do establish that integration work happened — they are not evidence of production use, scale or security outcomes. There are SDKs on PyPI, npm and crates.io, and a documented REST API, which shows a deliberate multi-language surface for a young project. GitHub metadata at capture: created December 2025, roughly 2.3k stars, forks in the low hundreds, and an open-issue count in the high hundreds. The guide does not read those numbers as quality; it reads them as a fast-moving young runtime with a busy tracker.
Development activity
Very active. Commits land continuously, releases and SDK versions move quickly, and the repository is under nine months old. Speed is the thing to weigh: a runtime that changes shape weekly is not something to make a long-lived architecture depend on yet.
Common complaints
Too young to have a settled complaint profile, which is itself the finding. The absent-KVM path is handled with a clear error rather than a silent, slower fallback — our own test got 'unsupported: KVM is not available on this host' immediately — which is good behaviour but still means the tool is unusable on hosts without hardware virtualisation, including most cloud VPSs. Expect the usual young-runtime class of reports (platform-specific boot problems, snapshot and networking edge cases) to appear and settle over the next year, and treat any specific one as unverified until it recurs. On a host that did have KVM, the same 0.10.2 line ran an Alpine microVM with guest command execution, networking and root-filesystem persistence across a stop/start, in a one-image smoke test (see experience.html#tested-here). That is execution evidence on one host — it is not a performance figure, and it says nothing about the isolation boundary.
Security considerations
The architecture is the right shape: a per-box kernel, a jailer around the VMM, egress controls and host-side secret injection, which is more than most projects in this space document. The problem is measurement. BoxLite is not in the independent comparative security study of agent sandbox engines, so there is no engine-level reading of its syscall filter, its host attack surface or its CVE history. The only independent measurement of its engine family is that study's libkrun result — no engine-side seccomp and eleven of fourteen host primitives reachable — and while BoxLite documents a jailer layer that the measured libkrun product did not have, that difference is self-reported and untested. It is also very new: no audit, no public fuzzing dashboard, and no time for an escape history to accumulate. Treat the security design as promising and the security claim as unproven.
Interesting features
The 'SQLite for sandboxing' embedding model is the genuinely good idea: a library with no daemon and no root requirement, so an application gets a hardware boundary without standing up any infrastructure. On top of that, persistent boxes that survive restart, clone and snapshot from a running box, full-box export and import as archives, OCI images unchanged, four language SDKs plus a C API and a REST API, allow_net egress control, placeholder secret injection with environment sanitisation, and per-box metrics.
Why this rating
It is one of the few projects aimed exactly at what this guide's readers want: hardware isolation for agent workloads that you can embed on a laptop, with the storage and snapshot story agents actually need, and a platform matrix that is honest about requiring KVM. What is missing is everything that would make it a recommendation — age, independent measurement, a settled issue history, and any production evidence beyond the integrations its consumers advertise. Those are time and search, not design flaws, which is precisely why the verdict is 'experimental' rather than 'skip': worth evaluating, not worth depending on.
Our take
Try it if you are on Linux with /dev/kvm, or on Apple Silicon, and you want the container-free microVM story with persistence. Before you build on it, check two things against your own environment: that a hardware accelerator is genuinely in use on your host, and that your toolchain runs unchanged inside a box. Then decide whether you can absorb its changes — a December 2025 project is going to move, and the honest position on its isolation is that it is documented rather than proven. Do not call it production-ready, and do not call it trusted; nothing in the public record supports either word yet.

Microsoft Execution Containers (MXC)

EXPERIMENTAL / EARLY
★ 1.3kforks 79open issues 84created 2026-02-06last push 2026-09-19licence MIT
Type
Multi-backend sandbox launcher driven by a JSON policy schema
Platforms
Windows 11 24H2+, Linux x64/ARM64, macOS ARM64/x64
Agent-specific
Yes — built for running model output, plugins and tools.
Open source
Yes, MIT
Isolation model
One versioned JSON schema over a pile of backends: AppContainer-style process containers, Windows Sandbox, WSLC, LXC, bubblewrap, Seatbelt, a Hyperlight and NanVix microVM path, and an isolation-session mode that spins a throwaway Windows session. Policy covers filesystem read/write path lists, network proxying and allow/block, plus a UI policy for clipboard, display and GUI access. Stable one-shot backends are processcontainer, bubblewrap, lxc and seatbelt; everything else needs an experimental flag.
Adoption evidence
Backed by Microsoft and announced at Build 2026, with the company's own framing that it is what Copilot CLI uses for dynamically generated code. Developers are starting to adopt its schema vocabulary. Real-world usage evidence is thin because the preview is months old.
Development activity
Very active — commits land daily — but the repository opens with a warning that policies generated by the SDK are known to be overly permissive and will change.
Common complaints
The upstream README is itself the main complaint: "no MXC profiles should be treated as security boundaries currently." Denied paths are not yet supported on Windows. Network policy on Linux and macOS is described as cooperative. The overall picture is an early preview whose abstraction is more mature than its enforcement, and reviewers have begun building their own restricted-token backends instead of adopting it.
Security considerations
Do not treat it as a boundary today. The abstraction is trustworthy — the policy vocabulary is the most complete cross-platform containment description in the open — but several backends are advisory and one (Windows Job Objects in a sibling project) turns out to be process-tree kill-on-close only. Whether a given policy is enforced depends on Windows build number, host OS and backend.
Interesting features
The JSON schema vocabulary for filesystem, network and UI policy; a documented per-Windows-release capability matrix; a state-aware lifecycle (provision → start → exec → stop → deprovision); Windows Sandbox automation including .wsb generation; a Hyperlight microVM backend; and ETW diagnostics for troubleshooting.
Why this rating
It is Microsoft, it is MIT, and it is the single best piece of open documentation for how many ways there are to contain a process on Windows. It is also explicitly not a security boundary yet, with per-version enforcement gaps, and a track record measured in months. The right use today is to steal the schema and the platform knowledge, and to test it in parallel rather than to build on its enforcement.
Our take
Use it as the vocabulary your own policy layer speaks, and as a way to get Windows Sandbox and WSLC launched consistently. Do not ship a promise of security that rests on an MXC profile.

CODE / IDEAS WORTH STUDYING

Do not build directly on the repo, but parts of it are excellent.

NanaBox

CODE / IDEAS WORTH STUDYING
★ 1kforks 60open issues 18created 2022-04-01last push 2026-09-13licence NOASSERTION
Type
Windows VM front-end built on the Host Compute System API
Platforms
Windows only
Agent-specific
No
Open source
Yes, MIT for the code with a Sponsor Edition for contributors
Isolation model
Creates Hyper-V-backed virtual machines through the Host Compute System API (the same low-level layer Hyper-V Manager uses) without registering them with Hyper-V Manager, with portable JSON VM definitions. That gives you a real hypervisor boundary plus GPU-PV, nested virtualisation and enhanced session mode for Windows guests.
Adoption evidence
A thousand stars, an active maintainer with a long history of Windows utilities, and enough standing that the author of appsandbox credits it for Host Compute System understanding. Used by Windows power users rather than by agent products.
Development activity
Active — commits within the week of this audit.
Common complaints
Requires UAC elevation, which alone disqualifies it as an automatic backend for a normal user-mode application. Gen2/UEFI guests only. Niche troubleshooting: every problem becomes a Host Compute System problem, and the documentation that exists is largely the project's own notes. The user report behind this site, that it 'did not work no matter what', matches its strict requirements rather than indicating a bug.
Security considerations
The boundary is Hyper-V, which is the strongest thing available on a Windows desktop and is serviced as a security boundary by Microsoft. The subtlety is that HCS-based VMs are not managed by Hyper-V Manager, so a VM escaping the normal inventory is a configuration-governance question rather than a security one.
Interesting features
Portable JSON VM definitions, GPU paravirtualisation, enhanced session mode, and — most usefully for this guide — a README that documents Host Compute System's gotchas better than Microsoft does.
Why this rating
As a product for agents it is the wrong shape: elevation, Gen2-only, niche. As a source of understanding it is excellent, because the Host Compute System API is exactly what any Windows GUI-VM feature has to use, and the gotchas are documented in public by someone who fought them. Adopt the knowledge, not the dependency.
Our take
Read it before building Windows VM features. Do not put it in a dependency tree, and do not expect HCS work to be pleasant just because someone else documented the traps.

App Sandbox

CODE / IDEAS WORTH STUDYING
★ 704forks 75open issues 45created 2026-03-16last push 2026-09-17licence MIT
Type
Desktop VM application using HCS/HCN, with a headless HTTP/JSON API
Platforms
Windows 11 x64 (including Home), macOS Tahoe on Apple Silicon
Agent-specific
No — but its API shape is exactly what an agent workstation manager needs.
Open source
Yes, MIT, distributed as signed prebuilt binaries
Isolation model
Full desktop VMs — Windows 11, Ubuntu, macOS guests — created through the same Host Compute System and Host Compute Network layer WSL2 uses, with GPU paravirtualisation (DirectX 12, OpenGL, Vulkan, CUDA, OpenCL), audio, clipboard, SSH over a Hyper-V socket with no network required, and snapshots. A real hypervisor boundary with a desktop attached.
Adoption evidence
Seven hundred stars, released binaries signed by an EV certificate on Windows and an Apple Developer certificate on macOS, and enough substance that a headless Python SDK ships with it. Usage evidence is thin, and this site's author could not get it to run at all.
Development activity
Active, with commits in the audit week.
Common complaints
Installation and boot reliability is the main one, and it is the reason it is filed under 'study' here: it did not work on the machine this guide was written on regardless of configuration, which matches reports of strict host requirements. It generates full desktop VMs, so resource use is high and it is not a lightweight per-task sandbox. It is also a single-maintainer project with a very large surface.
Security considerations
Hyper-V-grade isolation when it runs, with GPU paravirtualisation being the wider-than-usual host surface (a paravirtual GPU driver is shared code). Nested virtualisation support means you can run containers inside the guest, which is a sane layering: Hyper-V, then Docker or Podman, then the agent.
Interesting features
Windows 11 Home support without Hyper-V feature enablement, GPU-PV with compute APIs, a documented headless HTTP/JSON API plus a dependency-free Python SDK, snapshots, and the ability to run Claude Cowork or Docker inside the guest.
Why this rating
The architecture is exactly the Windows answer this space needs, the licence is permissive, and the API is the shape you would design yourself. What is missing is evidence that it reliably installs and runs on typical hardware: the sample size here includes one flat failure, and the project is one maintainer with a large surface. Take the HCS/HCN patterns and the API design; do not build a dependency on it yet.
Our take
Mine it for how to drive Host Compute System and how to expose VMs over a local JSON API. If you need a reliable Windows desktop VM today, use WSL2 with WSLg or Hyper-V directly.

PROBABLY SKIP FOR NOW

Expected hassle and risk look greater than the likely benefit.

Daytona

PROBABLY SKIP FOR NOW
★ 71kforks 5.6kopen issues 457created 2024-02-06last push 2026-07-24licence none detected
Type
Hosted dev-environment and agent-sandbox platform
Platforms
Cloud (vendor-hosted, with a customer-cloud data plane); Linux for self-hosting, which is no longer supported
Agent-specific
Yes
Open source
No longer. The public repository is explicitly unmaintained; AGPL survives only on tagged releases.
Isolation model
Docker/OCI containers by default, with Kata or Sysbox only if you configure them. That means the default isolation level is a shared kernel. Persistent workspaces are the default behaviour, which is convenient and is also why idle billing surprises people.
Adoption evidence
Seventy-one thousand stars made it the most visible name in the category and it was recommended widely through 2025. That reputation was built on being open source, which is the claim that subsequently changed.
Development activity
The public repository has not moved since July 2026 and its own messaging says it will not. Development continues in private.
Common complaints
The most serious complaint is the one that cannot be fixed by a patch: the code you were told to trust is no longer inspectable. Simon Willison's widely quoted reaction was that it is "not a great advertisement for a sandboxing product" to effectively say you do not trust your own security enough to publish the source. Alongside that, container-level default isolation is a weak boundary for untrusted generated code, self-hosting is gone, persistent workspaces quietly bill unless idle timeouts are configured, there is no supported static egress, and some ecosystem plugins have already been removed.
Security considerations
The open-source-looking reputation and the actual trust model diverged in 2026. For a sandbox, the source is the trust boundary: if you cannot read the isolation code, you are trusting a vendor's assurance about the one component whose entire job is to be trustworthy. Whether or not the closed code is good, the auditability is gone.
Interesting features
Genuinely fast workspace creation, persistent stateful environments, and a polished developer experience — which is why so many people adopted it before the licensing change.
Why this rating
This is the clearest cautionary example on the site, and it is not about a bug. A project built its entire reputation on open source, then closed the production codebase and left the public repository to rot, while continuing to advertise the brand. For a developer tool that might be fine. For the component whose job is to contain untrusted code, it removes the only mechanism you had for checking the claim. The complaints-to-adoption ratio is not the story here; the story is that the trust model changed under existing users.
Our take
Do not build on the public repository: nothing in it is maintained, and a frozen sandbox accumulates unpatched escape bugs. If you are evaluating the closed product, evaluate it as a proprietary service with the usual vendor-review process, not as an open-source project. This is also the reason the guide insists that 'open source' be treated as a claim you verify rather than a property you assume.

EdgeBox

PROBABLY SKIP FOR NOW
★ 213forks 26open issues 2created 2025-09-12last push 2026-04-22licence GPL-3.0
Type
Electron desktop sandbox with a full GUI desktop inside a container
Platforms
Linux, macOS, Windows
Agent-specific
Yes
Open source
Yes, GPL-3.0
Isolation model
E2B-derived E2B code-interpreter components running locally in a container, with a VNC-exposed desktop so the agent can drive a browser and GUI applications. Container, not VM.
Adoption evidence
Two hundred and thirteen stars, two watchers, no releases, and last activity in April 2026. The Hacker News and directory attention it received has not translated into usage.
Development activity
Effectively stalled. No commits since April 2026, no releases at all.
Common complaints
The observable signals are the complaint: no release cadence, no issue traffic, a GPL-3.0 licence that prevents code reuse in most commercial products, an Electron runtime that contradicts the memory footprint goal of most Rust/Tauri agent GUIs, and a README that injects a promotional recommendation for an unrelated project — a pattern worth flagging when evaluating a repo's honesty.
Security considerations
A container desktop exposed over VNC is a wide surface, and a stalled container project does not receive container-runtime patches. Because the isolation comes from the underlying runtime rather than from EdgeBox itself, the practical outcome depends on how current your Docker/Podman is, not on anything the project maintains.
Interesting features
The UX idea is genuinely attractive: a visible desktop the agent operates, with the human able to watch and take over. That combination — computer-use with an observability surface — is the pattern worth taking from it, not the code.
Why this rating
It is a stalled Electron app with no releases, a licence that blocks reuse in most products, and a pattern of injecting unrelated recommendations into its own README. Even if the design is right, there is no maintenance, no release process, and no evidence that anyone is running it. The expected hassle of adopting it exceeds any benefit, and the good part (the desktop-with-observability UX) can be reproduced on a maintained base in a weekend.
Our take
Read the screenshots, skip the repo. If you want a sandboxed desktop, build it on a maintained container or microVM base with an existing VNC or streaming layer.

NOT AN ISOLATION LAYER

Useful, but it does not contain the agent.

Windows-MCP

NOT AN ISOLATION LAYER
★ 7kforks 848open issues 26created 2025-05-13last push 2026-09-16licence MIT
Type
MCP server that gives an agent control of the Windows desktop
Platforms
Windows 7 through 11
Agent-specific
Yes — but as a capability, not a containment
Open source
Yes, MIT
Isolation model
None, and this is the important point. It runs on the host as your user and exposes file navigation, application control and UI interaction to the model. It is the exact opposite of a sandbox: it is the thing a sandbox exists to prevent. The only containment available is whatever you wrap around the process, and because it needs full desktop interaction, wrapping it tightly usually breaks it.
Adoption evidence
Seven thousand stars, a claimed two million users in one directory listing, availability on PyPI and in the MCP registry, and a large Discord. Adoption is not in question; the adoption is of host control.
Development activity
Very active, with a high volume of automated dependency updates and releases.
Common complaints
Update churn and occasional breakage when the underlying UI (and its own tool surface) changes. The deeper issue is not a complaint anyone files: giving a model UI control of the machine where your credentials live makes every prompt-injection path a machine-level compromise.
Security considerations
Treat it as the highest-privilege component in any stack you build. If you use it, use it inside a VM with no access to real credentials or real data, and treat its outputs as untrusted. It belongs on a list of things to be careful with, not on a list of ways to be safe.
Interesting features
Accessibility-based UI control, screenshot and input primitives, and a process model that makes Windows automation reachable from any MCP-capable agent.
Why this rating
It is not a sandbox and it does not claim to be one — but it appears in agent-sandbox comparisons, which is how developers end up believing that 'the agent can control Windows' and 'the agent is contained on Windows' are the same sentence. It is genuinely useful software; the risk is entirely about the company it keeps in a reader's mental model.
Our take
Use it if you want an agent driving Windows, and put it in a VM. Never describe it as isolation, and never run it where the credentials you care about live.

What the assessments have in common

Patterns that keep showing up

  • Engine quality and product quality are different things. The independent security study found that engine classes separate cleanly on every axis while products inside a class do not — and the worst findings were product-level decisions (a Privileged: true default, a frozen engine pin, nested virtualisation left on) rather than engine bugs.
  • Pin policy is an operator variable, and it dominates. Engine-side patch latency was effectively zero for coordinated disclosures; the observed delay ranged from zero days to 471 days to "opaque", and every day of it came from downstream pinning. Whatever you choose, decide how you receive engine updates.
  • The projects with the best documentation of their own limits are the ones worth reading. nono publishes a security model that enumerates what it does not protect; byre's README says outright that a container is not a microVM; Fence documents that its network allowlisting is content-blind. That honesty correlates with the quality of everything else.
  • Abandonment is the most common failure, and it is measurable. The reliable signals are last-commit date, release count, maintainer count and whether the issue tracker has recent non-bot activity — not stars.
  • Beta labels are usually accurate. microsandbox calls itself beta and has the bug list to prove it; MXC says its profiles are not security boundaries; Gondolin documents the trust assumptions it relies on. Take the label seriously and integrate as an option rather than a dependency.
What this method cannot tell you

It cannot tell you whether a project is secure. Adoption-risk assessment measures reliability, maintenance and honesty, and a project can score well on all three while having an architectural flaw that matters for your threat model. Read the security section of each assessment for that, and read Reality Check for the boundaries that no amount of good maintenance changes.

Evidence confidence

Every assessment would carry a confidence level if this were a formal report. The short version: confidence is high for projects with public issue trackers and thousands of users, medium for those with active repositories but thin usage evidence, and low for anything measured in weeks of history or single-author projects — which is precisely why several of those are rated "experimental" rather than condemned.