Blog · · 12 min read
How to build an AI agent code execution sandbox in your product

An AI agent code execution sandbox fails when the product treats a working directory, a subprocess, or a shared notebook as a security boundary. Generated code can read another tenant's files, probe private services, copy credentials, consume unlimited compute, or return an unsafe artifact. One container flag cannot prevent all of those failures. The product needs a lifecycle that binds the requesting user to isolated inputs, a controlled runtime, explicit network and credential policy, hard resource budgets, reviewed outputs, and verified teardown. This guide defines that lifecycle from request admission through artifact release, with implementation choices and tests for a multi-user product.
Why an AI agent code execution sandbox needs a product boundary
A model does not need malicious intent to produce dangerous code. A plausible data-analysis script may recurse through the wrong directory, install an unexpected package, follow a redirect to an internal host, or write more output than the service can store. Untrusted user files and retrieved content can also influence what the model generates.
The runtime must therefore assume that every command is untrusted. OpenAI's shell tool guidance calls arbitrary shell commands dangerous and recommends sandboxing, command controls, audit logs, and strict treatment of user-controlled input. Its hosted containers also have outbound networking disabled by default. OWASP's AI Agent Security Cheat Sheet likewise says agents should not execute arbitrary code without sandboxing and calls for limits on retries, resource use, and tool chains.
A weak boundary usually breaks in one of four places:
- The executor loses identity context and cannot prove which user, tenant, run, or agent release requested the work.
- The process has weak isolation and can see host state, another run's workspace, or a shared kernel.
- The runtime inherits network and credential authority from the host instead of receiving a grant for this request.
- The product releases files without checking their type, size, provenance, or policy result.
Codex issue #5237 reports a directory-scoped sandbox failure in which a command launched inside one working directory could read a file from another user's home directory. Treat that as an author report, not a universal product claim. It still demonstrates why cwd is a location, not an access-control mechanism.
Choose the isolation backend from the threat model
Do not begin with a favorite runtime. Start with the data sensitivity, tenant count, languages, package needs, network requirements, and expected attacker capability. Then choose the least complex backend that satisfies those constraints.
| Backend | Appropriate use | Main limitation | Required compensating control |
|---|---|---|---|
| In-process or subprocess | Trusted internal calculations only | Shares host identity and attack surface | No untrusted code or tenant data |
| Browser WebAssembly | Narrow client-side computation | Limited system calls, packages, and durable storage | Opaque origin and bounded file bridge |
| Container | Controlled workloads on dedicated worker nodes | Shares the host kernel | Hardened runtime, no host sockets, strict mounts |
| gVisor-style runtime | Multi-tenant Linux workloads needing stronger syscall isolation | Compatibility gaps and syscall overhead | Compatibility tests and resource quotas |
| MicroVM | High-risk code, broad package use, or strict tenant separation | Higher startup and operating cost | Warm pools without cross-run state reuse |
| Managed code tool | Standard analysis with provider-defined limits | Less control over runtime and data path | Provider review plus product-side admission and release policy |
gVisor describes its isolation model as a per-sandbox application kernel that intercepts system calls instead of passing application calls directly to the host kernel. This reduces the host-kernel attack surface, but the model has compatibility and performance tradeoffs. Test the actual workload before assigning a backend by label.
Docker's sandbox isolation documentation separates the kernel boundary, network, Docker engine, workspace, and credentials. Its direct-mount example shows that a VM can isolate the host kernel while an explicit read-write workspace mount still lets the agent alter build files, hooks, configuration, and other host-visible content.
Define one execution contract before writing the runner
The product API should accept a typed request, not a free-form command plus environment variables. Keep object references and credential handles in the request. Resolve sensitive values only inside trusted infrastructure.
{
"request_id": "req_01",
"run_id": "run_01",
"actor": {"user_id": "user_42", "tenant_id": "tenant_7"},
"agent_release": "release_2026_08_10",
"runtime_profile": "python-analysis-v3",
"inputs": [
{"object_id": "obj_sales_q2", "version": "18", "mount": "sales.csv"}
],
"network": {
"mode": "deny",
"allowed_destinations": []
},
"credential_handles": [],
"budget": {
"wall_seconds": 45,
"cpu_millis": 20000,
"memory_mb": 1024,
"disk_mb": 512,
"processes": 32,
"output_mb": 25
},
"release_policy": "report-artifacts-v2"
}
The contract should produce two durable records:
- A
SandboxLeaserecords the chosen backend, immutable image or environment version, workspace identity, policy versions, start time, and expiry. - An
ExecutionReceiptrecords commands or code digests, input versions, destination decisions, resource use, exit status, output hashes, release decisions, cleanup status, and the final error class.
Do not put raw secrets, full customer files, or generated command output into either record by default. Store references, hashes, decisions, and redacted evidence needed for support and incident review.
Implement the request-to-artifact lifecycle
The end-to-end sequence should remain stable even if the executor changes from a managed tool to containers or microVMs.
1. Admit the request
Authenticate the caller and freeze the tenant, user, run, agent release, and policy version. Confirm that code execution is allowed for the product workflow. Resolve input object permissions against current product state, not a permission snapshot copied from an old chat turn.
Reject a request before provisioning when its runtime profile, input type, destination rule, requested budget, or artifact class is outside policy. Admission is also where plan limits and concurrency quotas belong. A request that cannot afford its declared worst case should not enter the executor queue.
2. Provision a fresh lease
Allocate an isolation boundary to one tenant and one run. A warm pool may reuse a clean base image, but it must not reuse a writable workspace, process namespace, network namespace, credential state, or user cache from the previous lease.
Pin the image digest and runtime profile version. Start as a non-root identity where the backend supports it. Drop capabilities, block host sockets, and expose only the control channel needed to receive work and report status. The lease gets a hard expiry independent of model behavior.
3. Stage approved inputs
Copy or mount only the object versions named in the admitted request. Prefer a private copy-on-write workspace over a direct mount to product storage. Normalize filenames, reject path traversal, cap archive expansion, and record the staged object hash.
No home directory, source repository, cloud metadata endpoint, shared temporary folder, or broad product bucket should appear by accident. If a library needs a cache, create a run-local cache with the same expiry as the lease.
4. Execute under independent controls
The model may propose code, but a trusted executor owns enforcement. Network policy, CPU and memory limits, process ceilings, filesystem mounts, and timeouts must not depend on the generated program cooperating.
Docker documents a deny-by-default sandbox posture that blocks outbound HTTP and HTTPS unless allowed, blocks raw TCP and UDP, and denies private, loopback, and link-local destinations. Apply the same principle even if your backend uses different mechanisms. An allowed domain rule must also constrain resolved IPs, redirects, schemes, ports, and request methods. Recheck policy on every connection rather than resolving once at admission.
Keep credentials outside the guest. Docker's credential isolation design uses a host-side proxy to inject an approved value into matching outbound requests while the sandbox sees only a sentinel. A product implementation can use a similar broker: bind each handle to the user, tenant, run, destination, method, scope, and expiry; refuse requests that do not match all fields; and log the decision without logging the secret.
5. Quarantine and release artifacts
Execution success does not make an output safe. Move generated files into a quarantine area that the product cannot serve directly. Enforce count and size limits, normalize names, verify declared and detected media types, reject links and device files, unpack archives within bounded expansion limits, and run the scanners required by the product's data class.
The release service should create new immutable product objects rather than expose the sandbox filesystem. Attach the request ID, input versions, code digest, output hash, release-policy version, and scan result. If validation fails, retain only the minimum forensic evidence allowed by retention policy and return a typed artifact_rejected result.
6. Destroy and verify
Cleanup runs after success, failure, timeout, cancellation, worker loss, and client disconnect. Revoke credential handles, terminate the runtime, detach storage, delete the writable workspace, invalidate network policy, and close the lease. Then verify that the runtime identity, workspace, and credential references are no longer usable.
After deletion, query the backend to confirm that the runtime identity, workspace, and credential references are gone. A cleanup task that returns success without those checks leaves the security state unknown.
receipt = execute(admitted_request)
try:
released = release_service.validate_and_copy(receipt.outputs)
return completed(receipt, released)
except PolicyError as error:
return artifact_rejected(receipt, error.code)
finally:
revoke_credential_handles(receipt.run_id)
destroy_runtime(receipt.lease_id)
assert_runtime_absent(receipt.lease_id)
assert_workspace_absent(receipt.workspace_id)
close_lease(receipt.lease_id, cleanup_verified=true)
Apply the contract to a realistic analysis request
Suppose a customer asks the agent to analyze a private sales export and generate a chart. The admitted request references one versioned CSV object and allows no network access. A fresh Python runtime receives a copy named sales.csv. It has 1 GB of memory, 20 seconds of CPU, 45 seconds of wall time, a 25 MB output cap, and no credentials.
The agent generates Python that reads the CSV and writes summary.csv and revenue-by-region.png. The executor records the code digest and resource use. The release service checks both files, confirms their detected types, scans them, stores them as new tenant-owned objects, and returns product object IDs. The runtime is destroyed before the chat receives downloadable links.
If the generated program tries to install a package, the request fails because networking is disabled. Product policy may route a retry to a prebuilt runtime profile that already contains the approved dependency. It should not silently enable public internet access or pass a package-registry token into the guest.
Google's managed Gemini code execution documentation provides a useful example of explicit limits: it documents a 30-second environment runtime, bounded regeneration after errors, supported file types, and a fixed library set. The values will differ between backends. Put the chosen limits and retry behavior in the contract and receipt instead of leaving them in an operator's assumptions.
Handle failures without opening the boundary
Return typed outcomes so the product can decide whether to retry, ask the user, or stop.
| Failure | Safe response | Unsafe shortcut |
|---|---|---|
| Provisioning unavailable | Retry with the same admitted request and a bounded attempt count | Run on the API host |
| Resource limit reached | Return budget_exceeded with measured usage | Remove the limit automatically |
| Destination denied | Return egress_denied and the policy rule ID | Allow all networking for the retry |
| Credential binding denied | Require a new scoped authorization | Inject the user's session token |
| Artifact validation failed | Quarantine and return artifact_rejected | Serve the file with a warning |
| Cleanup unverified | Quarantine the worker and alert operations | Return the worker to the warm pool |
Do not let the model classify these errors as transient or override the response. The orchestration layer owns retry policy. A retry gets a new lease and workspace but keeps the original request ID and attempt lineage, preventing duplicate artifact release.
Verify the sandbox with adversarial tests
A passing happy-path analysis does not test isolation. Run adversarial checks at launch, during execution, at release, and after teardown.
Launch checks
- The runtime sees only the declared input objects and an empty writable workspace.
- The image digest, policy versions, clock, tenant, run, and expiry match the lease.
- Host sockets, metadata services, home directories, and previous-run caches are absent.
Execution checks
- Attempt reads through absolute paths, parent traversal, symlinks, proc files, and another active run.
- Attempt loopback, private ranges, link-local addresses, alternate DNS answers, redirects, raw TCP, and disallowed ports.
- Attempt fork bombs, disk fill, output floods, CPU loops, memory exhaustion, and excessive retries.
- Attempt to read credential values from files, environment variables, process metadata, logs, and error messages.
Release and teardown checks
- Produce oversized files, misleading extensions, archives with path traversal, archive bombs, links, and unsupported object types.
- Cancel during every lifecycle phase and simulate executor loss.
- After closure, retry the credential handle, runtime identifier, workspace path, and artifact-upload capability.
- Start another tenant's run on the same worker and prove that no previous data, process, cache, or network grant remains.
Track these tests by runtime-profile and policy version. A new base image, kernel, proxy, scanner, package set, mount rule, or executor backend should run the relevant suite before rollout.
Avoid the shortcuts that create shared risk
A shared Jupyter service is not a multi-tenant sandbox. Open WebUI's backend comparison explicitly warns that its legacy Jupyter option shares a runtime and filesystem, is not sandboxed by default, and was not designed for multi-tenant use. That architecture may suit a trusted single-user installation, but it should not sit behind a multi-customer product feature.
Other common mistakes are treating a container as sufficient while mounting the host Docker socket, putting provider or product keys in environment variables, allowing unrestricted package installation, trusting an output because its extension looks safe, and deleting state only on the successful path. Each shortcut bypasses a different layer of the lifecycle.
The product should own the contract even when a provider operates the runtime. Managed execution can reduce infrastructure work, but admission, object authorization, data minimization, artifact ownership, user-visible errors, and product audit still belong to the application.
Build the smallest secure slice next
Start with one workflow that needs no network access and produces one or two known artifact types. Implement the ExecutionRequest, SandboxLease, and ExecutionReceipt records. Provision one fresh runtime per run, stage one versioned input, enforce hard resource limits, release outputs through a separate service, and prove cleanup with a negative test.
Do not add package installation, credentials, or public networking until the closed-network path passes cross-tenant, resource-exhaustion, artifact, cancellation, and teardown tests. That sequence gives your AI agent code execution sandbox a boundary you can inspect and extend instead of a collection of permissions you cannot reason about.
References
- OpenAI shell tool guidance supports the command-risk, sandboxing, audit, and default network-access discussion.
- OWASP AI Agent Security Cheat Sheet supports the need for sandboxing and independent limits on agent execution.
- OpenAI Codex issue #5237 provides the clearly labeled practitioner report about directory-scoped isolation.
- gVisor documentation supports the application-kernel isolation model and its compatibility and overhead tradeoffs.
- Docker Sandboxes isolation layers supports the separation of kernel, network, workspace, Docker engine, and credential boundaries.
- Docker default security posture supports the deny-by-default network and explicit-mount examples.
- Docker credential documentation supports host-side credential brokering and scoped secret handling.
- Google Gemini code execution documentation supports the managed-runtime limit and file-handling example.
- Open WebUI code execution documentation supports the shared-runtime and per-user backend comparison.