beam-logo
← All posts
Engineering

Building RL Environments with Verifiers and OpenEnv

Eli MernitEli Mernit
September 29, 20268 min read
Building RL Environments with Verifiers and OpenEnv

If you are building an RL environment with Prime Intellect's verifiers library or Meta's OpenEnv standard, the framework hands you the scaffold — a dataset, a rollout loop, and a rubric that turns a completion into a reward. What it does not hand you is a safe place to run the model's tool calls. When your environment executes model-generated code to score it, that execution needs strong isolation, hard resource ceilings, and enough parallelism to keep a training batch moving. Here is that backend as a verifiers tool, where each call runs in a gVisor-isolated Beam Sandbox with a time and memory ceiling.

vf.ToolEnv takes your tools as plain Python functions and calls them when the policy emits a tool call, so the sandbox slots in without touching the rest of the environment. Swap the toy run_python for whatever your tasks execute — a unit-test harness, a SQL query, a shell command — and the rollout runs untrusted code inside an isolated sandbox instead of on your trainer host.

What building on verifiers or OpenEnv actually gives you

Both frameworks separate the shape of an environment from the execution inside it, and that split is the whole reason a sandbox belongs in the picture. Knowing which half is yours saves you from either reinventing the scaffold or, worse, running model code with no isolation.

  • `verifiers` gives you the environment abstraction. An environment is a dataset plus a rollout plus a Rubric that combines weighted reward functions into a scalar (SingleTurnEnv for one-shot reasoning, MultiTurnEnv for conversations and games, ToolEnv for function calling). Its GRPOTrainer and the external prime-rl trainer consume that environment directly. The library shipped built-in SandboxEnv and PythonEnv classes, but those are wired to Prime Intellect's own hosted Sandboxes; a ToolEnv tool, by contrast, is a plain function whose backend you choose. That is the seam where you bring your own sandbox.
  • OpenEnv gives you a transport standard. Announced by Meta's PyTorch team and Hugging Face in October 2025, OpenEnv wraps each environment as a FastAPI server in an isolated Docker container and talks to it over HTTP with a Gymnasium-style reset() / step() / state() API. Isolation is part of the spec, which is a real strength — but the reference runs those containers locally, and running thousands of them for an actual RL run is an infrastructure problem the standard leaves open.
  • The execution is the untrusted part. In both cases the code you run mid-rollout was written by the model under training, which means it will eventually hang, fork, exhaust memory, or reach for the network. That is not a place for subprocess.run. The pillar on sandboxes for reinforcement learning covers why per-rollout isolation is the substrate for RL environments generally; this page is about wiring that substrate into the two frameworks people actually build on.

How to back a verifiers environment with a secure sandbox

The hero snippet is the pattern in full: a fresh sandbox per tool call, the model's code executed inside it, the output returned to the rollout. The steps below are what each part is doing and where to tighten it before you point a training run at it.

Define the sandboxed tool

A ToolEnv tool is an ordinary Python function, so the sandbox lives entirely inside it. Create the sandbox, run the model's code with sb.process.exec(...) or sb.process.run_code(...), read result.stdout, and terminate. For a code-verification environment where the reward is pass-or-fail, run the dataset's hidden tests instead and return the test runner's exit code — result.wait() gives you that exit code directly, so the reward is one comparison rather than a stdout parse you have to trust.

Enforce time and memory ceilings, and drop the network

memory="512Mi" caps how much a single call can allocate, so one greedy completion can't take down the node, and update_ttl(30) is a wall-clock ceiling the platform enforces — a completion that spins forever gets reclaimed and scored as a failure instead of stalling the batch. Set the TTL to a small multiple of how long a correct rollout should take. The other lever is egress: model-generated code with open internet access is the most common way an environment leaks, by fetching answers or exfiltrating secrets from the run. Disable network egress unless a specific task needs it. Isolation protects the host; the egress policy protects your data. The self-hosting guide goes deeper on both layers.

Scale the rollout batch

One tool call is a demo; a training step fires the tool across a whole batch, and a run does that millions of times. Because each call is independent, you fan out with a thread pool — the work happens in the sandboxes, not in your process.

Each sandbox terminates the moment its call returns, so nothing idles between steps and you pay per second of actual execution rather than for a warm pool. When the environment is heavy — a compiler toolchain, a large dependency tree, a dataset to mount — build it once and snapshot the filesystem into a reusable image with create_image_from_filesystem(), then restore that template per rollout instead of reinstalling. That throughput lever is the same one RL rollouts live on; the serverless-GPU-for-RL guide covers the case where the rollout itself needs a GPU.

How to run OpenEnv environments in production

OpenEnv already gives you container isolation, so the question is not whether to sandbox but where those containers run. Locally, the reference spins up each environment server in Docker on your machine. A real run needs those servers somewhere with stronger-than-container isolation, a GPU when the environment does inference, and the ability to scale to thousands and back to zero between steps. A sandbox platform hosts the environment server — or just the code-execution inside a step() — behind the same HTTP interface OpenEnv already speaks.

Because OpenEnv can import a verifiers environment, the same sandboxed execution you wired into ToolEnv carries over — you are not rebuilding the backend to move between the two standards. Keep the environment logic in the framework and let the sandbox own isolation, resource caps, and fan-out.

What to look for in an execution backend for RL environments

The backend runs untrusted code, on the critical path of a training loop, at high throughput. Four properties decide whether it fits that job:

  • Isolation tier. How much separation stands between model-generated code and the host: a microVM boots its own kernel (strongest), a user-space kernel like gVisor mediates syscalls before they reach the host (strong, lighter, faster to start), a plain container shares the host kernel (weakest). Beam uses gVisor plus runc, which starts fast enough to boot one per tool call.
  • Throughput and scale-to-zero. Rollouts are bursty: thousands of sandboxes for a few seconds each step, then nothing. You want cheap parallel fan-out and billing that drops to zero between steps, not a fixed fleet held warm.
  • GPU inside the sandbox. When the reward is a model-judge, or the environment runs inference to produce an observation, the backend needs a GPU. Most lightweight sandboxes are CPU-only; on Beam a GPU is a parameter on the same API.
  • Determinism and self-host. A pinnable image keeps rollouts reproducible across a run, and an open runtime you can self-host keeps the environment — and the training data it sees — inside your own perimeter. Beam's runtime, beta9, is open source under AGPL-3.0 and runs managed, self-hosted, or bring-your-own-cloud on the same SDK.

Execution backends compared

Honest picture for the RL-environment execution job specifically. Concede where a tool wins: E2B's Firecracker microVM gives the strongest isolation tier, the right call for genuinely adversarial, multi-tenant code.

PlatformIsolationGPU in sandboxParallel + scale-to-zeroSelf-hostLicense
BeamgVisor + runc (user-space kernel)Yes, same APIYes, fan-out + per-secondYes (beta9, BYOC)AGPL-3.0
E2BFirecracker microVM (own kernel)NoConcurrency caps by planYes (Nomad/Firecracker, heavy)Apache-2.0 SDK
ModalgVisorYesYes, .map() + scale-to-zeroNo (managed only)Proprietary
DaytonaDocker default (Kata opt-in)YesPersistence-firstYesAGPL-3.0
DIY (Docker/Firecracker)Your choiceYour buildYou build itYes—

Architectural facts verified against each vendor's docs and source. For a broader RL-specific roundup see best sandbox providers for reinforcement learning, and for the stateful-session angle see best stateful sandboxes for code execution. The short version: pick E2B when own-kernel isolation is a hard requirement and you will operate a Nomad cluster; pick Beam when you want gVisor isolation, GPU-in-sandbox, high-throughput fan-out, and one runtime you can self-host without assembling orchestration yourself.

FAQ

What is the `verifiers` library? It is Prime Intellect's open-source framework for building RL environments and evals for LLMs. An environment is a dataset, a rollout loop, and a Rubric that combines reward functions into a scalar; classes like SingleTurnEnv, MultiTurnEnv, and ToolEnv cover reasoning, multi-turn, and tool-calling tasks. Its GRPOTrainer and the prime-rl trainer consume the environment directly.

What is OpenEnv? An open standard for agentic RL environments from Meta's PyTorch team and Hugging Face, announced in October 2025. Each environment runs as a FastAPI server in an isolated Docker container and is driven over HTTP with a Gymnasium-style reset() / step() / state() API, so trainers and environments stay decoupled.

Does `verifiers` run the code for me? For a ToolEnv, no — your tools are plain Python functions and you own what runs inside them. The library ships SandboxEnv and PythonEnv that target Prime Intellect's hosted Sandboxes, but if you want a different isolation tier, a GPU in the sandbox, or to keep execution in your own cloud, you supply the backend. That is exactly the seam a sandbox platform fills.

Do I need a GPU for the execution backend? Usually not — running code against tests or a shell is CPU work. You need a GPU when the reward is a model-judge or the environment runs inference to produce an observation. On Beam that is the same sandbox API with a GPU parameter, so the backend doesn't become a separate system.

Can I self-host the sandbox backend? Yes. Beam's runtime, beta9, is open source under AGPL-3.0 and runs locally, in your Kubernetes cluster via Helm, or bring-your-own-cloud on AWS, GCP, Azure, or Hetzner — with the same SDK, so an environment you prototype on managed Beam moves in-house without a rewrite.

How do `verifiers` and OpenEnv relate? They are separate projects with an interop path: OpenEnv ships an import command that wraps a verifiers environment (and other sources) as an OpenEnv environment, and Prime Intellect sits on OpenEnv's technical committee. You can build on verifiers and export to the OpenEnv transport, and the same sandboxed execution backend serves both.

Back your RL environment with a sandbox built for it

verifiers and OpenEnv give you the environment; they leave the execution backend open on purpose. Beam fills it in one API: a gVisor-isolated sandbox with hard time and memory caps, a GPU when the reward needs it, parallel fan-out that scales to zero between steps, and an open runtime you can self-host. Build the backend on managed Beam, then move it into your own cloud when the run scales.

Get started free

Eli Mernit
Eli Mernit
Published September 29, 2026
Pay as you gobilled by the millisecond

Start shipping on infra
you won’t outgrow.

Run sandboxes and GPU workloads on your cloud, and scale out to ours when you need to. No infra to manage.