beam-logo
← All posts
Tutorials

Fast Cold Starts for Serverless GPU Inference

Eli MernitEli Mernit
September 29, 20268 min read
Fast Cold Starts for Serverless GPU Inference

A cold start on a serverless GPU is the time between a request arriving and your model being ready to answer it: pulling the container image, booting the container, and loading the model weights into VRAM. For anything larger than a toy model, the weight load is the part that hurts. You shrink it by loading the model once per container with an on_start loader, caching the weights on a volume, snapshotting the warmed container, and keeping a warm pool for the routes that can't wait. On Beam, containers cold-boot in one to three seconds, and under a second when the image is already cached.

Deploy it with beam deploy app.py:generate. The weights download to the volume on the first boot and are read from the mount after that; on_start puts the model on the GPU once per container, so the cold start you actually pay in latency is the boot plus that single load, not a reload on every request.

What a cold start is on a serverless GPU

A serverless GPU scales to zero when it's idle, so the first request after a quiet period has to bring the whole environment back before any inference runs. That cold start is four costs stacked on top of each other:

  1. Scheduling — the platform finds a GPU and assigns your container to it.
  2. Image pull — the container image is fetched to the node. A fat CUDA image with the wrong layering can be gigabytes.
  3. Container boot — the runtime starts the container and your process comes up.
  4. Weight load — the model is read from disk and moved into VRAM. A 7B model in fp16 is roughly 14GB to shuffle onto the card; a 70B model is far more.

For small models the first three dominate. For anything real, the weight load is the tax that matters — and it's the one people forget, because it doesn't show up when you test locally with the model already in memory. That reframes the whole problem: the fastest provider in a benchmark still stalls if your code reloads a 14GB checkpoint on every request. The fixes below attack the load and the boot separately. Both live on top of a serverless GPU that turns off between requests, which is what makes the cold start exist in the first place — and also what saves you money the rest of the time.

How to reduce cold starts on a serverless GPU endpoint

Cutting cold starts is a handful of levers, and most workloads need two or three of them, not all five. Start by killing the reload, then decide how much warmth you're willing to pay for.

Load the model once with an on_start loader

The single biggest win is not loading the model inside the request handler. An on_start function runs when the container boots and its return value is handed to every request through context.on_start_value, so the weights hit the GPU once per container instead of once per call. That's the pattern in the hero snippet above. Without it, a container that serves a thousand requests loads the model a thousand times, and every one of those looks like a cold start to the caller.

Cache weights on a Volume so they don't re-download

Point cache_dir at a mounted Volume and the model downloads from Hugging Face exactly once, on the first container that ever runs. After that every cold start reads the checkpoint from the volume instead of pulling it across the network again. For a multi-gigabyte model, that's the difference between a network download and a local disk read on the critical path.

Snapshot the container to skip the reload

Loading a model into VRAM takes seconds even from a warm disk. Beam can snapshot a container after the weights are resident and restore from that snapshot on the next boot, so a new container comes up with the model already on the GPU. Beam uses memory snapshots and GPU checkpoint-restore for this — restoring a GPU container from a snapshot is, by Beam's own numbers, up to 35× faster than booting it from scratch. It's the closest thing to a free cold-start cut, because you're not paying to keep anything running between requests.

Keep a warm pool for latency-critical routes

Snapshots make a cold start fast; a warm pool removes it. keep_warm_seconds sets how long a container stays alive after its last request — the default is 180 seconds, and you raise it (or set a floor of always-on containers) for a route where even a one-second cold start is too much. This is the honest trade in the whole exercise: a warm container is billed while it waits, so warmth is a per-route decision, not a global switch. Put it on the checkout endpoint, not the nightly batch job. Beam's autoscaling handles the scale-up beyond the warm pool when traffic spikes.

Keep the image small and let the runtime lazy-load it

The image pull only bites when the image is big and cold. Beam's open-source runtime, beta9, lazy-loads image layers from a distributed cache rather than pulling the whole thing up front the way a Docker-based platform does, so the first request doesn't block on a full download. You help it by not bloating the image — pin the packages you need, keep build artifacts out, and let the cache do the rest.

What to look for in a low-cold-start serverless GPU platform

Whatever platform you land on, cold-start behavior comes down to the same short list. Judge it on these rather than a single headline latency number, because the number depends entirely on the model and whether the container was warm.

  • Container-boot mechanism. Is the runtime pulling a full Docker image every cold start, or lazy-loading from a cache? This sets the floor for the boot itself. Beam runs its own runtime and lazy-loads layers; most providers sit on Docker.
  • Snapshot or checkpoint support. Can the platform restore a container with the model already in VRAM, or does every cold start reload weights? For large models this is the difference between seconds and tens of seconds.
  • Warm-pool control. Can you keep a route warm and set exactly how long, per endpoint? A global setting forces you to overpay everywhere to protect one path.
  • Billing shape. Are you billed per second and only for compute, or per minute and for idle time too? Per-minute rounding and idle charges quietly punish bursty, scale-to-zero workloads.
  • GPU selection. Can you match the model to the right card — a T4 or A10G for a small model, an A100 or H100 for a large one — instead of overpaying on hardware you don't need? Beam covers T4, A10G, A100, H100, and RTX 4090.

Serverless GPU cold starts compared

The providers split by how they boot a container and how granular the bill is. Beam runs a custom runtime built for fast boots and snapshotting; RunPod and Google Cloud Run hand you a Docker-based container; Baseten wraps deployment in its Truss packaging; Replicate optimizes for its public model library and treats custom deployments differently. The numbers below are from Beam's own cold-start benchmark, which ran 100 requests per provider across several days on matched hardware.

PlatformCold startBilling granularityBoot mechanism
Beam1–3s, sub-second cachedPer second, scale to zeroCustom runtime, lazy image load, snapshot restore
RunPod6–12s, sub-second warmPer secondDocker, FlashBoot for active workers
Google Cloud Run20–30sPer secondDocker, L4 GPUs only
Baseten16–60sPer minuteTruss packaging, autoscaling
ReplicateInstant public, 60s+ customPer second, idle billed on customManaged containers

Those competitor figures come from Beam's top serverless GPU providers benchmark, ranked by cold start, and reflect 2025 testing — treat them as directional and re-measure for your own model. On raw GPU rate, an H100 on Beam is $1.83/hr and an A100 80GB is $1.36/hr, billed per second with the endpoint scaling to zero between requests. For a real example of what the boot mechanism buys you, Gepetto cut its cold starts while dropping infrastructure cost after moving to Beam. If you're serving an LLM specifically, the vLLM on Beam guide covers the serving layer, and the Modal pricing breakdown compares one of the managed alternatives.

FAQ

What is a cold start in serverless GPU inference?

It's the delay on the first request after a container has scaled to zero, spent bringing the environment back: scheduling a GPU, pulling the image, booting the container, and loading the model weights into VRAM. Once a container is warm, later requests skip all of that and answer immediately. Cold starts are the price of scaling to zero, which is also what stops you paying for idle GPUs.

Why are GPU cold starts slower than CPU ones?

Because of the weights. A CPU service comes up as soon as its process starts, but a GPU inference container isn't ready until the model is resident in VRAM, and a multi-billion-parameter model is many gigabytes to read from disk and move onto the card. That single load usually dominates the cold start, which is why loading the model once per container and snapshotting it matter far more than shaving milliseconds off the container boot.

How do I avoid cold starts entirely?

Keep a warm pool. Setting keep_warm_seconds high, or keeping a floor of always-on containers, means a container is waiting when the request lands so there's no cold start at all. The catch is that a warm container is billed while it idles, so it's worth doing on latency-critical routes and not on everything. For bursty or background work, a fast cold start plus scale-to-zero is usually the cheaper trade.

Does keeping containers warm cost money?

Yes. A warm container is running, so you're billed for the GPU time it spends waiting for the next request, even when it's doing nothing. That's the trade-off against cold starts: you either pay in latency when a request has to boot a container, or you pay in idle compute to keep one ready. The right split is per route — warm the endpoints users wait on, let the rest scale to zero.

What's the fastest serverless GPU platform for cold starts?

In Beam's 2025 benchmark, Beam had the fastest cold starts at one to three seconds, with RunPod next at six to twelve (and sub-second once a worker is warm via FlashBoot), then Google Cloud Run, Baseten, and Replicate's custom deployments trailing. The real answer depends on your model size and whether you use snapshots and a warm pool, so benchmark with your own checkpoint rather than trusting a headline number.

Are cold starts billed?

On Beam you're billed per second for the time your code actually runs, and the endpoint scales to zero when idle, so you're not paying for a GPU that isn't serving. Billing granularity varies by provider — some bill per minute and some charge for idle time on custom deployments, which makes bursty, scale-to-zero inference more expensive than the headline GPU rate suggests. Check the billing unit, not just the per-hour price.

Get started

Serve your model on GPUs that boot in seconds and scale to zero between requests. Get started on Beam — pay only for the compute you use.

Eli Mernit
Eli Mernit
Published September 29, 2026
Pay as you gobilled by the millisecond

Start shipping on infra
you won’t outgrow.

Run sandboxes and GPU workloads on your cloud, and scale out to ours when you need to. No infra to manage.