Scaling RL Rollouts with Parallel Sandboxes
Eli Mernit
To scale RL rollouts, give every episode its own isolated sandbox and drive the batch from one concurrency pool. Restore each sandbox from a snapshot so it boots in seconds, run the agent's actions inside it, pull back the reward, and terminate. The batch size you can run at once sets how fast your policy improves, and the sandboxes scale to zero the moment the batch drains.
There is no scheduler to write here. The pool submits the work, Beam boots a container per rollout, and results come back as each episode finishes.
What scaling RL rollouts actually demands
The number you are trying to move is rollout throughput: how many episodes you can generate per second. A policy update consumes a batch of rollouts, so the trainer sits idle until the batch is ready. Widen the batch and the policy sees more reward signal per step and improves faster; starve it and the expensive training GPU waits on the environment half of the loop. Throughput is the knob, and concurrency is how you turn it.
Each rollout needs at least one sandbox, because the actions are untrusted. In code and agent RL the policy writes programs, runs shell commands, and edits files, and mid-training most of that is wrong — infinite loops, deleted directories, calls to places it shouldn't reach. One episode leaking a file or a process into the next corrupts the batch, so the isolation is what makes the reward a clean signal in the first place. The pillar on sandboxes for reinforcement learning covers why that isolation is non-negotiable and where the reward comes from.
That turns a scaling problem into a concurrency problem. Run a thousand rollouts at once and you need a thousand sandboxes that spin up fast, stay isolated, and disappear when the episode ends — thousands of times per sweep. The two things that break at that scale are spin-up cost and stragglers, and the rest of this page is about both.
How to run thousands of parallel sandbox rollouts
Fan the batch out across sandboxes, restore each one from a prepared snapshot, and put a timeout and a retry around every episode so one bad rollout can't stall the batch. The hero snippet is the whole path; these are the pieces worth expanding for a real training loop.
Snapshot once, clone into every rollout
Reinstalling dependencies or re-cloning a repo on every episode multiplies by the episode count and dominates the runtime. Prepare the environment once, then snapshot it and restore the snapshot into each rollout so every episode starts from an identical, ready state:
Every Sandbox(image=Image.from_id(TEMPLATE)) now boots from that captured filesystem in one to three seconds instead of rerunning setup. When warm process state matters — a server already listening, a REPL mid-session — snapshot_memory() captures running processes and restores with create_from_memory_snapshot() so a warmed environment comes back without rebooting.
Fan out with a concurrency pool
The fan-out itself is a pool over a function that owns one sandbox for the length of one episode, as in the hero snippet. A thread pool works because each rollout spends its time waiting on the sandbox, not on your CPU. Raise the ceiling by widening the input list, not by provisioning machines — the same code runs 16 rollouts or 16,000.
If your rollout is already wrapped in a Beam @function, you can fan it out with .map() instead of a pool and let queue-depth autoscaling size the fleet. That path, and how to attach a GPU to the environment for a reward model or accelerated simulator, is covered in serverless GPU for reinforcement learning. The same .map() throughput pattern behind batch inference on serverless GPU applies to rollouts unchanged.
Put a timeout and a retry around every episode
At a thousand concurrent episodes, a few will hang — a generated program spins forever, a task deadlocks. Without a bound, the slowest rollout holds the whole batch hostage while the trainer waits. Cap each episode and treat a timeout as a zero-reward sample rather than a crash:
Setting keep_warm_seconds on the sandbox is the platform-side backstop: the container self-terminates after that many seconds even if your process leaks, so a runaway episode can't keep billing.
Keep the trainer saturated
The loop is only as fast as its slowest half. If the trainer consumes rollouts faster than the pool produces them, the training GPU idles; produce far faster and you pay for experience the trainer hasn't reached yet. Size the batch to what one policy update consumes, and when episodes arrive continuously instead of in fixed batches — an actor generating trajectories as the policy shifts — switch the pool for a task queue that autoscales on queue depth and streams results back as they land. Because the sandboxes scale to zero between updates, oversizing the ceiling costs nothing when the queue is empty; you only pay while episodes actually run.
What to look for in RL rollout infrastructure
Whatever platform you fan rollouts out on, throughput lives or dies on the same handful of properties. Evaluate against these rather than a headline GPU price:
- Concurrency ceiling. How many sandboxes run at once? This is the cap on rollout throughput, which is the cap on how fast the policy improves. A hard limit in the low hundreds throttles a large sweep.
- Cold start, and whether you pay for it. Rollouts start and stop constantly, so spin-up latency compounds across tens of thousands of episodes. Being billed for container boot or image-pull time is a tax on churn.
- GPU inside the sandbox. Can a rollout touch a GPU for a reward model or an accelerated simulator? Without it, every in-environment inference step round-trips to separate infrastructure.
- Session cap. A fixed maximum session length cuts off long-horizon episodes and multi-day curricula midway.
- Per-second billing that scales to zero. Rollout volume is spiky. Paying by the second and dropping to zero between updates beats a fixed fleet you rent around the clock.
- Self-host and BYOC. Reward functions and training data are often proprietary. An open-source runtime you can run in your own cloud keeps them in your account.
Beam clears these because it fans one function out to thousands of concurrent sandboxes, runs GPU sandboxes with no session cap, doesn't bill for spin-up, charges per second down to zero, and runs on an open-source runtime you can self-host or point at your own VPC.
Parallel sandbox platforms compared
| Platform | Concurrent sandbox rollouts | GPU in sandbox | Session cap | Billing | Self-host / BYOC |
|---|---|---|---|---|---|
| Beam | Thousands per fan-out | Yes | None | Per second, scales to zero | Yes, open-source runtime |
| E2B | 100, up to 1,100 on Pro | No | 24 hours | Per second | Yes, Firecracker (heavy) |
| Modal | Autoscaled, no public cap | Yes | Configurable | Per second | No, managed only |
| Daytona | Autoscaled for RL | Yes | Configurable | Per second | Yes, open-source |
On raw GPU rate, an H100 is $1.83/hr on Beam versus $3.95/hr on Modal and $2.27/hr on Daytona's preemptible tier, and an A100 80GB is $1.36/hr on Beam; E2B offers no GPU sandboxes at all. For a full ranking of these platforms against RL-specific criteria, see the best sandbox providers for reinforcement learning guide.
FAQ
How many rollouts can I run in parallel?
As many as your batch needs. A single fan-out spreads one function across thousands of concurrent sandboxes, and you raise the ceiling by widening the input list rather than provisioning machines. Match it to the batch size one policy update consumes — the bigger the batch you run at once, the more reward signal each update sees.
How do I keep the trainer fed?
Size rollout throughput to what the trainer consumes per update. If the pool produces slower than the trainer drains, the training GPU idles waiting on experience; much faster and you generate rollouts the policy hasn't reached. For episodes that stream in continuously, use a task queue that autoscales on queue depth instead of fixed batches.
What happens when a rollout hangs?
Bound every episode with a timeout and score a timeout as a zero-reward sample. Kill the process, terminate the sandbox, and optionally retry once before giving up, so the slowest rollout can't stall the batch. A keep_warm_seconds limit on the sandbox is the backstop that stops a leaked process from billing forever.
Do parallel rollouts need a GPU?
Often, but not always. When a model lives inside the environment — a reward model scoring outputs, an agent running inference, a GPU-accelerated simulator — the rollout needs a GPU or every step round-trips elsewhere. Classic tabular or lightweight-simulator environments run fine on CPU sandboxes, which are cheaper to fan out by the thousand.
What does a batch of 1,000 rollouts cost?
Roughly the compute each episode actually uses, billed per second, since the sandboxes scale to zero between updates. A thirty-second CPU rollout at two cores costs a fraction of a cent; add an A100 at $1.36/hr and that same thirty seconds of GPU time is about a cent. You pay for the seconds the batch runs, not for idle capacity between policy updates.
What isolation do the sandboxes use?
Each rollout runs in its own gVisor-isolated container, so untrusted agent code — infinite loops, stray file writes, network calls — can't touch the host or the rollouts running next to it. That hard boundary is what lets you run thousands of them from one process without episodes corrupting each other.
Run your next sweep on parallel sandboxes
Fan your rollouts out across sandboxes that boot in seconds and turn off between policy updates. Get started on Beam — pay only for the compute you use.


