Sandboxes

Agents need their own computer. Here's how to give them one safely.

Sandboxes give you a safer way to run upgraded workflows, providing you an isolated environment to run tasks that require iteration, verification, and access to the tools a person would normally use, all without requiring human supervision.

July 15, 2026

Why agents need their own computer

Ask an LLM to debug a failing test or clean a dataset, and it'll get you most of the way there. It can explain the fix, write the query, and outline the analysis. Then it stops, and you have to pick up everything else.

The problem here is that a system that can only produce text is like a contractor who can describe exactly how to fix your plumbing but has no hands, no tools, and no truck. The advice might be perfect, but someone still has to go turn the wrench.

Agents close that gap by getting hands. Give a model the ability to run code, read the result, and try again, and the full agent loop enables the agent to do more autonomously:

.png)

An agent that can only suggest a fix has no way to know if the fix works. That's why agents need their own computer: a real environment with a filesystem, a shell, a package manager, network access, and state that persists across steps.

And while you have one laptop and you might be the only one using it, an agent platform might be spinning up thousands of these environments in parallel, each one needing to be isolated, disposable, and safe to hand real execution power to.

What does this mean in practice?

In each example, the model is running real steps and checking real results. That requires a place to work, not just a context window to reason in.

Why you can't just hand it your infra

So why not just let the agent run code locally, or in a Docker container?

Most prototypes start locally: It's fast, it's familiar, and it's good enough to get a demo working. Then it goes to production, and the same setup starts to fail in two specific ways.

And while your agent doesn’t have bad intent, you don’t know where this code is coming from. A line of code can originate from the model itself, from a cloned repo, or a package installed mid-run. For example, a research agent parses documents it pulled from the open web. Agent-executed code can be generated seconds before it runs, shaped directly by whatever a user typed, and produced mid-loop as the agent reasons its way through a task. There's no review step in between.

A well-written prompt doesn’t give you immunity from security concerns. The safest posture is machine-level separation: give the agent a real environment to work in, but keep that environment isolated from your laptop, from production, and from every other agent's workspace running alongside it.

Four things every agent's computer needs to do well

.png)

1. Execute safely

In 2025, a self-replicating npm worm backdoored hundreds of packages and executed in preinstall hooks before any tests ran. A Linux kernel CVE disclosed in 2026 could root any major distribution with a 732-byte Python script in about an hour, and containers couldn't help, because they shared a kernel with the host.

Agent-executed code should be treated as untrusted by default, regardless of its source. This includes code the model wrote, code pulled from a cloned repository, packages installed mid-task, and scripts produced by multi-step reasoning chains. Each agent workspace should be a hardware-virtualized machine with its own kernel, filesystem, and network boundary.

2. Stay in control

Controls inside a sandbox protect you from the agent doing expensive, unexpected, or credential-leaking things.

3. Be observable

Observability in agent execution means knowing:

Essentially, this is an audit log for workflows, especially ones that touch sensitive data or take actions with real-world consequences. What makes an agent reliable is the ability to re-run from a known state, compare branches, and trace what actually happened.

4. Build and iterate fast

Production requirements need to be fast provisioning (sub-second when warm), reproducible environments (defined by a Docker image or blueprint that every instance starts from), and persistent state (files, installed packages, and session context carry over between agent turns). If spinning up an execution environment takes thirty seconds, agents that need multiple environments in a task will feel slow. If environments aren't reproducible, bugs become hard to isolate. If state doesn't persist across sessions, long-running tasks require expensive restarts.

When to reach for a managed sandbox vs. building your own

You can approach this in a DIY fashion: run the agent on a developer's laptop, graduate to a Docker container for some separation, wire up resource limits and credential injection manually. For agents that only call external APIs with fixed schemas and never execute dynamic code, this is often enough.

When you want to scale to generating scripts, installing packages, running test suites, or parsing files, you’ll have to:

Here's the decision criteria:

The operational overhead of a DIY approach for production agent deployments adds up fast. The managed sandbox path trades that engineering surface for a simpler interface with a platform that handles the scale of work.

LangSmith Sandboxes: a computer for every agent

.png)

Each LangSmith sandbox boots fast (median under one second), runs as a hardware-virtualized microVM with its own kernel, and persists state (files, installed packages, environment) across the agent's working session. When the task is done, the sandbox idles and gets cleaned up automatically.

While a container shares a kernel with the host, a microVM has its own. Inside the sandbox, the agent can install anything, run Docker, start services — all while your infrastructure and other workloads stay untouched.

Core primitives

One-call setup

from langsmith.sandbox import SandboxClient

client = SandboxClient()

with client.sandbox() as sb:
    result = sb.run("python my_analysis.py")
    print(result.stdout)

Sandboxes work with Deep Agents, Open SWE, LangSmith Deployment, LangSmith Fleet, and any custom code. They use the same SDK and API key as the rest of LangSmith.

A note on prompt injection in sandbox workflows

Sandboxes provide strong execution isolation, but they don't change a fundamental property of language models: anything the agent reads can influence what the agent does next. This matters when sandbox output is fed back into the model.

An example concern is a research agent downloads a document, the document contains text designed to look like an instruction to the model, and the model follows it. This is the #1 vulnerability in the OWASP Top 10 for LLM Applications, and it applies to any agent that processes external content, whether that's web pages, uploaded files, API responses, or code execution output.

Sandboxes don't eliminate the threat of injection, but they do contain the execution blast radius. Malicious code that runs inside a sandbox can't reach your host, but if the output of that execution is read back by the agent without scrutiny, an injected instruction in the output can still influence downstream behavior.

How you can mitigate:

None of this is unique to sandboxes. Any agent that reads from the web, processes uploaded files, or calls external APIs has the same exposure. The added value of Sandboxes is that they help you contain the execution damage, and both layers are needed.

Use cases

Coding agents

Coding agents are a popular use case for sandboxes since execution isolation matters a lot. An agent that can run a test suite, inspect the failure, patch the code, and run the suite again is qualitatively more useful than one that can only generate a patch and hand it back for the developer to validate.

Coding agents in sandboxes do things like:

For sandbox environments, the snapshot-and-fork pattern has worked especially well for CI-style agents like Open SWE. A single snapshot captures the repo and installed dependencies, then each candidate fix runs in its own fork, with the successful fork’s diff surfaced as the result.

Data analysis agents

Data analysis agents operate over files and databases that are often sensitive (financial records, health data, customer data, risk models). Data sensitivity and model-generated code means execution isolation is important.

A data analysis agent in a sandbox can:

Use cases in this category include analytics agents, financial research and underwriting, insurance claims processing, and scientific data pipelines — any time when sensitive data makes running code in your production environment a security or compliance concern.

Research agents

Research agents span the most varied execution surface. A deep research workflow might browse the web, download PDFs, run a scraper, call a data API, normalize multimedia content, and synthesize a structured output.

What sandboxes give research agents:

Use cases in this category include competitive intelligence, investment due diligence, academic research pipelines, and any workflow that requires synthesizing information from multiple external sources into a finished, structured output.

Conclusion

For most of computing history, a developer environment meant a physical machine, then a VM, then a container, with each one shared across the work happening on it, each one requiring deliberate setup and teardown. The idea of giving every agent its own isolated computer, booting in under a second and cleaning up when done wasn't practical at scale.

Every human developer gets a laptop. Every agent can get a computer.