A Guide to Sandboxed Code Execution for RLVR Rewards
Eli Mernit
In RLVR (reinforcement learning with verifiable rewards), the reward is not a learned model — it is the result of running the policy's code. You execute the completion against hidden tests inside an isolated sandbox, then score a clean pass as 1.0 and a failed assertion, a non-zero exit, or a timeout as 0.0. The sandbox is the reward function, so it has to run untrusted, model-generated code safely and return a signal in the training loop. Here is that function on Beam, where each check runs in a gVisor-isolated sandbox with a hard time and memory ceiling.
sb.process.exec(...).wait() returns the exit code, so the grader is one comparison: pytest exits 0 when every test passes and non-zero otherwise. Swap the toy add example for the completion your policy sampled and the tests your dataset ships, and you have a reward you can call from a GRPO or PPO step.
What RLVR needs from a code sandbox
RLVR works because code has a ground truth: it either passes the tests or it doesn't. That is the whole appeal — a rule-based reward instead of a reward model that can be gamed, which is the approach behind DeepSeek-R1 and the reasoning models that followed. But turning "run the code and see" into a training-loop reward puts three demands on the execution layer that a plain subprocess.run cannot meet.
- The reward has to be deterministic and pure. A reward is a function of
(completion, ground_truth)and nothing else. Shared state between checks — a leftover file, a mutated package, a reused port — leaks signal and quietly shifts the objective. Every rollout needs a clean environment. - The code is untrusted. It was written by the model, and during training the model will produce code that hangs, forks, allocates all your memory, or reaches for the network to exfiltrate the test answers. The sandbox has to contain that on the host, not just catch exceptions.
- The throughput is brutal. One RLVR step scores a batch of rollouts; a run scores millions. If each reward takes a container you boot and tear down by hand, the grader becomes the bottleneck and your GPUs sit idle waiting for the CPU to finish judging.
A sandbox that gives you per-execution isolation, real resource caps, and cheap parallelism turns the reward from a liability into a parameter. That is the same set of properties that makes sandboxes the substrate for RL environments generally — the pillar on sandboxes for reinforcement learning walks through the environment side; here the sandbox is narrowed to one job, computing the reward.
How to run model-generated code as a verifiable reward
The hero snippet is the whole pattern: a fresh sandbox per check, the completion and the tests written in, the test runner's exit code returned as the reward. The steps below are what each line is doing and where to tighten it for real training.
Write the solution and the graders' tests, then run them
Write the model's completion to one file and the dataset's tests to another, both inside the sandbox, then run the test runner against them. Using pytest's exit code keeps the grader honest — you are not parsing stdout or trusting the model to self-report, you are asking the interpreter whether the assertions held. For a partial-credit reward, run the tests one at a time and return the fraction that pass instead of a hard 0/1; the SERP-standard shape is a small bonus for "compiles and runs" and full reward for "passes every test."
Treat a non-zero exit or a timeout as failure
Bugs, exceptions, and infinite loops all have to collapse to reward 0.0, not crash the trainer. A non-zero exit code already does that. Timeouts are why update_ttl(30) matters: it is a wall-clock ceiling the platform enforces, so a completion that spins forever is reclaimed and scored as a failure instead of stalling the batch. Set it to a small multiple of how long a correct solution should take.
Cap memory and drop the network
memory="512Mi" caps how much a single check can allocate, so one greedy completion can't take down the node. For adversarial code the other lever is egress: model-generated code with open internet access is the most common way a verifier leaks — it fetches the answers, or worse, exfiltrates secrets from your environment. Run the grader with network egress disabled unless a specific test needs it. Isolation protects the host; the egress policy protects your data. The self-hosting guide covers both layers in more depth.
Cache by content hash
Because the reward is pure, two identical completions always score the same. Hash (completion, tests) and memoize, and re-scored rollouts cost nothing — worth it when a batch contains duplicates or you resume a run. Pin the sandbox image and the test runner version alongside the cache key so an upgrade doesn't silently invalidate old rewards.
How to scale RLVR rollouts to thousands of parallel sandboxes
A reward you call once is a demo; a reward you call for a whole batch every training step is the real workload. Because each verify is independent, you fan them out — a thread pool over the batch is enough, since the work happens in the sandboxes, not in your process.
Each sandbox terminates the moment its check returns, so nothing idles between training steps and you pay per second of actual grading rather than for a pool of graders kept warm. When the grader environment is heavy — a compiler toolchain, a big dependency tree, a dataset to mount — build it once, snapshot the filesystem into a reusable image with create_image_from_filesystem(), and restore that template per rollout instead of reinstalling. That is the throughput lever RL runs live on; the serverless-GPU-for-RL guide covers the case where the reward itself needs a GPU — a model-judge or a completion that runs inference to be scored.
Choosing a sandbox for RLVR reward functions
The reward runs untrusted code, millions of times, on the critical path of a training loop. Four things decide whether a sandbox fits that job:
- Isolation tier. How much separation stands between model-generated code and the host: a microVM boots its own kernel (strongest), a user-space kernel like gVisor mediates syscalls before they reach the host (strong, lighter), a plain container shares the host kernel (weakest). Beam uses gVisor plus runc — a mediated kernel that is stronger than a container and faster to start than a microVM, which matters when you are booting one per reward.
- Throughput and scale-to-zero. Grading is bursty: thousands of sandboxes for a few seconds each step, then nothing. You want cheap parallel fan-out and billing that drops to zero between steps, not a fixed fleet.
- GPU inside the sandbox. If your reward is a model-judge, or the completion runs inference to be scored, the grader needs a GPU. Most lightweight sandboxes are CPU-only; on Beam a GPU is a parameter on the same API.
- Determinism and self-host. A pinnable image keeps the reward reproducible across a run, and an open runtime you can self-host keeps the verifier — and the data it sees — inside your own perimeter. Beam's runtime, beta9, is open source (AGPL-3.0) and runs managed, self-hosted, or bring-your-own-cloud with the same SDK.
RLVR sandbox options compared
Honest picture for the reward-function job specifically — not general code execution. Concede where a tool wins: E2B's Firecracker gives the strongest isolation tier, which is the right call for genuinely adversarial, multi-tenant code.
| Platform | Isolation | GPU in sandbox | Parallel + scale-to-zero | Self-host | License |
|---|---|---|---|---|---|
| Beam | gVisor + runc (user-space kernel) | Yes, same API | Yes, fan-out + per-second | Yes (beta9, BYOC) | AGPL-3.0 |
| E2B | Firecracker microVM (own kernel) | No | Concurrency caps by plan | Yes (Nomad/Firecracker, heavy) | Apache-2.0 SDK |
| Modal | gVisor | Yes | Yes, .map() + scale-to-zero | No (managed only) | Proprietary |
| Daytona | Docker default (Kata opt-in) | Yes | Persistence-first | Yes | AGPL-3.0 |
| DIY (Docker/Firecracker) | Your choice | Your build | You build it | Yes | — |
Architectural facts verified against each vendor's docs and source (see the self-hosting comparison for the full sourcing). For a broader RL-specific roundup, see best sandbox providers for reinforcement learning. The short version: pick E2B when own-kernel isolation is a hard requirement and you will operate a Nomad cluster; pick Beam when you want gVisor isolation, GPU-in-sandbox, high-throughput fan-out, and one runtime you can self-host without assembling orchestration yourself.
FAQ
What is RLVR? Reinforcement learning with verifiable rewards: the reward comes from a deterministic check — running code against tests, matching a math answer — instead of a learned reward model. It is the training signal behind reasoning models like DeepSeek-R1, and for code the check is executing the model's output and reading the result.
Is gVisor enough isolation for untrusted model code? For most RLVR setups, yes. gVisor puts a user-space kernel between the code's syscalls and the host, a much smaller attack surface than a shared container. If you are running genuinely adversarial code from untrusted third parties and need own-kernel separation, a microVM (Firecracker, libkrun) is the stronger tier — weigh it against the higher boot cost per reward.
How do I stop the model from hacking the reward? Keep the reward a pure function of the completion and hidden tests, run the tests the model never sees, and pin the checker and its parser so the objective can't drift. Rule-based rewards cut reward hacking sharply, but models still find format tricks, edge cases, and test memorization — hold out tests and inspect high-reward samples.
Do I need a GPU for the verifier? Not for the common case — running code against unit tests is CPU work. You need a GPU when the reward is a model-judge or the completion runs inference to be scored. On Beam that is the same sandbox API with a GPU parameter, so the grader doesn't become a separate system.
Can I self-host the reward sandbox? Yes. Beam's runtime, beta9, is open source under AGPL-3.0 and runs locally, in your Kubernetes cluster via Helm, or bring-your-own-cloud on AWS, GCP, Azure, or Hetzner — with the same SDK, so the grader you prototype on managed moves in-house without a rewrite.
How many rewards can I run in parallel? As many as your batch needs — each check is independent, so you fan out with a thread pool and every rollout gets its own sandbox that scales to zero when done. Grading throughput becomes a concurrency setting, not a fixed fleet you keep warm.
Turn your reward function into a sandbox
RLVR only works if the thing computing the reward can run untrusted code safely, deterministically, and thousands of times per step. Beam gives you that in one API: a gVisor-isolated sandbox with hard time and memory caps, GPU when the reward needs it, parallel fan-out that scales to zero between steps, and an open runtime you can self-host. Build the verifier on managed Beam, then move it into your own cloud when the run scales.


