E2B vs Beam for RL Environments
Eli Mernit
Pick E2B for CPU-only RL environments where you want microVM isolation and one-call forks of a running sandbox. Pick Beam when the environment needs a GPU, when episodes run past E2B's 24-hour session cap, or when you need around 1,000 rollouts at once without paying for a concurrency add-on. Both bill per second and both restore environments from snapshots.
Add gpu="RTX4090" (or "H100") to the Sandbox(...) call and the same loop runs environments that need a GPU. That one argument is the biggest practical difference between the two platforms for RL.
Which one to pick
E2B is the better fit for CPU-bound agent environments where isolation strength and fast forking matter most. Beam is the better fit when rollouts need GPUs, run for a long time, or share a platform with policy training. The table below is the short version; the rest of the page shows the numbers behind it.
| If your RL environment... | Better fit |
|---|---|
| Runs untrusted agent code and you want a hardware-virtualized boundary | E2B |
| Needs to clone a live, warmed-up environment many times in one call | E2B |
| Needs a GPU for a reward model, simulator, or in-environment inference | Beam |
| Runs episodes or curricula longer than 24 hours | Beam |
| Needs about 1,000 concurrent rollouts on a self-serve plan | Beam |
| Is CPU-only and short, and the lowest per-second CPU rate is what matters | E2B |
| Must run inside your own cloud account in production without an enterprise contract | Beam |
What an RL environment needs from a sandbox
An RL environment sandbox has to give every episode a clean, identical starting state, keep the policy's actions away from other episodes, and scale out to the batch size a policy update consumes. After those three, the requirements that separate providers are GPU access, session length, and what concurrency costs.
Code and agent RL is the demanding case. The policy writes programs, runs shell commands, and edits files, and early in training most of that is wrong: infinite loops, deleted directories, network calls it shouldn't make. If one episode leaks state into the next, the reward stops meaning anything. That is why each rollout gets its own sandbox, and why the pillar on sandboxes for reinforcement learning treats isolation as the first requirement rather than a nice-to-have.
The practical checklist:
- Deterministic resets. Every episode starts from the same filesystem, and ideally the same memory state.
- Isolation. A crashed or hostile episode can't touch the host or its neighbors.
- Parallelism. Rollout throughput caps how fast the policy improves, so the concurrency ceiling matters more than a headline price.
- GPU access. Reward models, simulators, and agents that run inference inside the environment all want a GPU next to them.
- Session length. Long-horizon tasks and multi-day curricula break on a hard session cap.
E2B and Beam compared, feature by feature
E2B runs each sandbox as a Firecracker microVM with memory-level snapshots and forking, but it is CPU-only and caps sessions at 24 hours on Pro. Beam runs gVisor-isolated containers that can attach a GPU, has no session cap, and includes 1,000 concurrent CPU containers on its $89/month Team plan. Figures below come from both vendors' pricing pages and docs.
| E2B | Beam | |
|---|---|---|
| Isolation | Firecracker microVM per sandbox | gVisor-isolated container |
| GPU in the environment | No, CPU-only | Yes, RTX 4090 through H100 and up |
| Max session length | 1 hour on Hobby, 24 hours on Pro, multi-day on Enterprise | No cap, runs until you terminate it or the timeout you set |
| Concurrent sandboxes | 20 on Hobby; 100 on Pro, 600 or 1,100 with paid add-ons | 30 CPU and 5 GPU on Developer; 1,000 CPU and 50 GPU on Team |
| Snapshot including memory | Yes | Yes |
| Fork a running sandbox in one call | Yes, up to 100 forks per request | No, boot N sandboxes from a snapshot |
| Sandbox rate for 2 vCPU and 8 GiB | About $0.23/hr | About $0.32/hr |
| Monthly platform fee | $0 Hobby, $150 Pro | $0 Developer, $89 Team |
| Policy training on the same platform | No | Yes, GPU functions in the same SDK |
| Open-source runtime | Apache-2.0 | AGPL-3.0 |
Sources: E2B pricing, E2B fork docs, Beam pricing, and Beam sandbox configuration docs. For the full E2B tier breakdown, see E2B pricing explained.
Isolation: Firecracker microVMs vs gVisor
E2B has the stronger isolation boundary. Each E2B sandbox is a Firecracker microVM with its own kernel under hardware virtualization, while Beam runs containers under gVisor, which intercepts system calls with a user-space kernel so agent code never talks to the host kernel directly.
For most RL work the adversary is a confused policy, not an attacker, and both boundaries hold. If you are training on tasks where the policy is actively rewarded for escaping, or you run third-party environments you don't trust, the microVM boundary is the more conservative choice.
Snapshots and forking
Both platforms restore environments from snapshots, but E2B's fork is the more convenient primitive. One fork(count=N) call pauses a live sandbox, captures its filesystem and memory once, and boots up to 100 running copies that start with the same processes and loaded variables.
Beam offers two snapshot types. create_image_from_filesystem() captures the disk and turns it into an image that any sandbox can boot from, which is what the hero snippet uses. snapshot_memory() captures running processes and exposed ports, and create_from_memory_snapshot() starts a new sandbox from that state. To get N copies on Beam you start N sandboxes from the snapshot in a loop or thread pool, which is more code than E2B's single call but has no per-request limit.
GPUs inside the environment
This is where the two split. E2B sandboxes are CPU-only, and E2B says so on its pricing page. A Beam sandbox takes a gpu argument, so a reward model, a GPU-accelerated simulator, or an agent that runs its own inference can sit inside the same isolated environment as the rollout.
On E2B, the workaround is to host the GPU model somewhere else and call it over the network from every step, which adds latency to each action and a second bill to manage. If your environment is a CPU-only Gym task or a coding sandbox scored by unit tests, none of this matters. For the GPU side of RL, including policy updates, see serverless GPU for reinforcement learning.
Session length and long-horizon episodes
E2B caps a sandbox session at 1 hour on Hobby and 24 hours on Pro; multi-day sessions need the Enterprise plan. Beam has no maximum: keep_warm_seconds sets an idle timeout, and keep_warm_seconds=-1 keeps the sandbox up until you terminate it.
E2B's pause and resume softens its cap. A paused sandbox keeps its filesystem and memory with no expiry, so a long episode can be paused and picked up later. That works for episodes that wait on something. It doesn't help an episode that has to keep computing for 30 hours straight.
Concurrency ceilings
For high-volume rollouts, the concurrency tier is usually the real limit. E2B Pro includes 100 concurrent sandboxes; 600 costs an extra $500 a month and 1,100 costs an extra $1,000 a month on top of the $150 plan fee. Beam's Team plan includes 1,000 concurrent CPU containers and 50 GPU containers for $89 a month.
At the low end, E2B's free Hobby tier allows 20 concurrent sandboxes and Beam's free Developer tier allows 30 CPU and 5 GPU containers. Both vendors sell larger ceilings on enterprise contracts: E2B advertises tens of thousands of concurrent sandboxes on Enterprise, which has a $3,000 monthly minimum.
What 1,000 parallel rollouts cost on each
For a CPU-only batch of 1,000 five-minute rollouts at 2 vCPU and 8 GiB each, E2B's compute costs about $19 and Beam's about $27. Running all 1,000 at the same time changes the picture: E2B needs the 1,100-sandbox add-on, which brings the monthly fee to $1,150, while Beam's Team plan covers it for $89.
The per-second math, from each pricing page:
| Line item | E2B | Beam |
|---|---|---|
| CPU rate | $0.000014 per vCPU-second | $0.0000375 per core-second (1 core = 2 vCPU) |
| RAM rate | $0.0000045 per GiB-second | $0.0000064 per GiB-second |
| 2 vCPU + 8 GiB sandbox | $0.000064/s, about $0.23/hr | $0.0000887/s, about $0.32/hr |
| 1,000 rollouts x 5 minutes (83.3 sandbox-hours) | About $19.20 | About $26.60 |
| Plan to run 1,000 at once | Pro with the 1,100 add-on, $1,150/month | Team, $89/month |
If you run a few batches a month, E2B's lower per-second rate wins on compute but the fixed fee dominates the bill. If you run many batches every day, the per-second gap compounds, and E2B's compute savings can overtake the fee difference; plug your own volume into both calculators before committing. The more your batch can be spread out over time, the fewer concurrent sandboxes you need, and the less the concurrency tier matters.
GPU rollouts only have one side to price. On Beam a serverless H100 bills at $0.000972 per second, about $3.50 an hour, and the policy update can run on an on-demand H100 machine from $1.83 an hour.
How to port an E2B rollout loop
Moving an RL loop from E2B to Beam is mostly a rename, because both SDKs follow the same create, run, snapshot, and terminate lifecycle. The main changes are the fork call, which becomes a loop over sandboxes started from a snapshot, and the session timeout, which becomes an idle timeout with no upper bound.
A typical E2B fan-out forks a warmed sandbox:
The Beam version is the hero snippet at the top of this page. The call-by-call mapping:
| Step | E2B | Beam |
|---|---|---|
| Define the environment | Template | Image, or create_image_from_filesystem() |
| Start a sandbox | Sandbox.create(template) | Sandbox(image=...).create() |
| Run a command | sandbox.commands.run(cmd) | sb.process.exec(...) |
| Run Python code | Code Interpreter run_code | sb.process.run_code(code) |
| Checkpoint with memory | sandbox.create_snapshot() | sb.snapshot_memory() |
| Start from a checkpoint | Sandbox.create(snapshot_id) | Sandbox().create_from_memory_snapshot(id) |
| Clone a live sandbox N times | sandbox.fork(count=N) | Start N sandboxes from a snapshot |
| Session lifetime | timeout, up to 24 hours on Pro | keep_warm_seconds, -1 for no limit |
| Shut down | sandbox.kill() | sb.terminate() |
You don't have to move everything at once. A common split is to keep CPU-only coding tasks on E2B and run the environments that need a GPU on Beam, since the rollout function returns a reward either way. The best E2B alternatives roundup covers the other options if neither fits.
Self-hosting and running in your own cloud
Both runtimes are open source, but they differ on what you can run in production yourself. E2B's runtime is Apache-2.0 and ships a single-machine package called Embed, which E2B describes as an evaluation package rather than a production deployment; production in your own cloud is a dedicated deployment through E2B Enterprise. Beam's runtime, beta9, is AGPL-3.0 and you can self-host it for free.
Beam also offers BYOC on AWS, GCP, or Azure: Beam runs the control plane, the instances run in your account on your cloud credits, and Beam charges a management fee of $0.019 per vCPU-hour and $0.009 per GB-hour of RAM. That matters for RL teams whose reward functions or task data can't leave their account. The guide to self-hosting a code execution sandbox walks through the tradeoffs of running either stack yourself.
FAQ
Does E2B support GPUs?
No. E2B sandboxes are CPU-only, per E2B's own pricing page. If an RL environment needs a GPU for a reward model, a simulator, or in-environment inference, you have to host that model elsewhere and call it over the network, or use a sandbox provider that attaches GPUs directly, such as Beam.
How many sandboxes can E2B run at once?
E2B's Hobby plan allows 20 concurrent sandboxes and Pro allows 100. Pro can be raised to 600 for an extra $500 a month or 1,100 for an extra $1,000 a month. Enterprise, with a $3,000 monthly minimum, advertises tens of thousands. Beam's Team plan includes 1,000 concurrent CPU containers for $89 a month.
Can I fork a running RL environment on Beam like E2B's fork?
Not in a single call. Beam captures a running sandbox with snapshot_memory() or its filesystem with create_image_from_filesystem(), and you then start as many sandboxes from that snapshot as you need. E2B's fork(count=N) does both steps in one request, up to 100 copies at a time.
Is Firecracker safer than gVisor for RL rollouts?
Firecracker gives each sandbox its own kernel under hardware virtualization, which is the stronger boundary. gVisor keeps agent code off the host kernel by handling system calls in a user-space kernel. For typical RL, where the risk is a policy doing something broken rather than malicious, both are adequate. For adversarial tasks or untrusted third-party environments, prefer the microVM.
Which is cheaper for RL rollouts?
For short CPU-only rollouts at modest concurrency, E2B's per-second CPU and RAM rates are lower. Once you need hundreds to a thousand rollouts at the same time, Beam's $89 Team plan usually costs less than E2B's concurrency add-ons, and for anything needing a GPU, E2B isn't an option at all.
Can I use E2B and Beam together?
Yes. The rollout function only has to return a reward, so the trainer doesn't care which provider ran the episode. Teams often keep CPU-only coding tasks on one provider and send GPU environments to another. The RL sandbox provider comparison ranks more platforms if you want a third option.
Run your RL environments on Beam
Boot rollouts from a snapshot, attach a GPU where the environment needs one, and pay per second only while episodes run. See how it fits your loop on the RL environments page, or start on Beam with the free Developer plan.


