beam-logo
← All posts
Tutorials
Engineering

How to Host an MCP Server and Run Tools in a Sandbox

Eli MernitEli Mernit
September 29, 20266 min read
How to Host an MCP Server and Run Tools in a Sandbox

An MCP server is a process that speaks the Model Context Protocol, either over stdio on your own machine or over HTTP for remote clients. To host one without handing tool code the keys to your laptop, run it inside a sandbox: an isolated cloud container with its own filesystem and network. On Beam that is a Sandbox, a background process, and a single expose_port call.

Point an MCP client at that URL and the server is live. When traffic stops, the sandbox scales to zero, so an idle server costs nothing.

What running an MCP server actually involves

MCP has two common transports, and the transport decides how much you have to worry about isolation. A stdio server runs locally and talks to one client through standard input and output; it is fine for a tool you wrote and trust. A remote server speaks Streamable HTTP (or the older SSE transport) and listens on a port, which is what you want when several clients, or a hosted agent, need the same tools.

The catch is what those tools do. A tool that runs a shell command, executes model-generated Python, or pulls a third-party MCP server off the registry is running someone else's code with whatever access the host process has. On your own machine that means your files, your SSH keys, and your network. A sandbox draws a hard line around it: the server and its tools get a fresh filesystem, a separate process space, and only the network you allow. Beam runs each sandbox under gVisor, a userspace kernel that intercepts guest syscalls before they reach the host, which shrinks the surface an escaping tool could reach.

If you would rather keep that boundary inside your own cloud account, Beam is open source (beta9, AGPL) and you can self-host the sandbox or run it BYOC.

How to host a remote MCP server in a sandbox

Hosting a remote server is four steps: build an image with your dependencies, get your code into the sandbox, start the server as a background process, and expose the port. The hero snippet above does all four. A few things worth knowing once you move past the first run:

  • Dependencies live in the image. Image().add_python_packages([...]) bakes your requirements in so they are cached; a warm custom image cold-starts in under a second, a cold one in one to three seconds. You are not billed while the image pulls.
  • `blocking=False` keeps the process alive. sb.process.exec(..., blocking=False) returns immediately and the server keeps running. Read process.logs to stream its output, or check process.exit_code, which stays -1 while the process is up.
  • The exposed URL is authenticated and SSL-terminated. expose_port(8000) hands back something like https://<id>-8000.app.beam.cloud; your MCP client appends the server's path (often /mcp) and connects.
  • Control the lifetime. sb.update_ttl(3600) keeps a sandbox up for an hour of idle time; sb.terminate() shuts it down on demand.

That is enough to run one server. Because sandboxes scale to zero on their own, you can leave a rarely-used server registered and only pay for the seconds it actually serves a request.

How to run MCP tool calls in isolation

The second pattern is narrower and more common than people expect: your server is trusted, but one of its tools executes untrusted code. A "run this Python" tool, a code interpreter for an agent, or a verifier that runs a model's output all fall here. The safe shape is to give each call its own throwaway sandbox and tear it down afterward, so nothing survives between calls and nothing touches your host.

run_code returns the captured stdout on out.result and the process status on out.exit_code, so your tool can hand a clean result (or a clean error) back to the model. If the tool needs a GPU, say for a tool that runs a small model as part of its work, ask for one on the sandbox: Sandbox(gpu="A10G"). The same isolation and per-call teardown applies. This is the same execution model behind stateful sandboxes for code execution, pointed at a single tool call.

What to look for in a sandbox for MCP servers

Not every sandbox fits the job, and the right one depends on which of the two patterns above you are running. A few axes that actually matter:

  • Isolation boundary. Namespaces (plain Docker) are the weakest; a userspace kernel like gVisor is stronger; a microVM like Firecracker is stronger still. For untrusted tool code, do not settle for namespaces alone.
  • A real public endpoint. Hosting a remote server means you need TLS, a stable URL, and ideally auth without standing up your own ingress.
  • Cold start and scale-to-zero. An MCP server that sits idle most of the day should not bill you for idle time, and it should answer the first call in seconds, not minutes.
  • A GPU path. If a tool runs a model, the sandbox has to attach a GPU without a separate service.
  • Self-host or BYOC. When the tools touch private data, running the sandbox in your own account is the difference between a demo and something security will sign off on.

Beam covers these through gVisor isolation, expose_port for an authenticated HTTPS URL, scale-to-zero by default, per-sandbox GPUs, and an open-source self-host path. Where it does not win: E2B's Firecracker microVMs give a stronger hardware-level boundary than gVisor, so if microVM isolation is a hard requirement, weigh that honestly against Beam's faster custom-image starts and GPU support.

MCP server hosting options compared

ApproachIsolationPublic HTTPS endpointScales to zeroGPUSelf-host / BYOC
Local stdio serverNone (your machine)Non/aYour machinen/a
Docker on a VMContainer namespacesYou set up ingress + TLSNoIf the VM has oneYes (your box)
E2BFirecracker microVMYesEphemeral sessionsNoYes (heavy setup)
ModalgVisorYesYesYesNo (managed only)
BeamgVisor + runcYes (expose_port)YesYesYes (beta9, BYOC)

Pricing and plan floors move around, so confirm them before you commit: E2B's Pro tier carries a monthly fee and no GPU, while Beam's Developer plan has no monthly fee and bills compute per second. For a broader field, these E2B alternatives and this rundown of code execution environments for AI agents go deeper than a single table can.

FAQ

What is an MCP server sandbox? It is an isolated cloud container that runs a Model Context Protocol server, and the tools it exposes, away from your host machine. The server gets its own filesystem, process space, and network, so a tool that executes code cannot reach your files or credentials.

Is it safe to run untrusted MCP tools? Safer, if you isolate them properly. Give each tool call a fresh sandbox, run it under a strong boundary like gVisor or a microVM, and terminate it afterward so nothing persists. The risk is running that code directly on a host that has access to your data.

stdio or HTTP transport — which do I use? Use stdio for a local server that talks to one client on your own machine. Use Streamable HTTP when you need a remote server that several clients or a hosted agent can reach over a URL. Hosting in a sandbox is about the HTTP case.

Do I need Docker to sandbox an MCP server? No. Docker is one way to get container-level isolation, but you still have to host it, add TLS, and manage scaling. A sandbox platform gives you the isolation, a public endpoint, and scale-to-zero without running the infrastructure yourself.

Can an MCP tool run GPU code in a sandbox? Yes. Request a GPU when you create the sandbox, for example Sandbox(gpu="A10G"), and the tool can run a model as part of its work with the same isolation and per-call teardown as a CPU tool.

How is this different from running the server locally? A local server shares your machine's filesystem, network, and permissions with every tool it runs. A sandboxed server runs in a clean, disposable environment reachable over an authenticated URL, which is what makes it safe to expose to other clients and agents.

Host your MCP server on Beam

Spin up an isolated sandbox, start your server, and get an authenticated URL in seconds — with a GPU when a tool needs one, and scale-to-zero when it doesn't.

Get started free

Eli Mernit
Eli Mernit
Published September 29, 2026
Pay as you gobilled by the millisecond

Start shipping on infra
you won’t outgrow.

Run sandboxes and GPU workloads on your cloud, and scale out to ours when you need to. No infra to manage.