Multi-GPU Training Without Kubernetes
Eli Mernit
You can train across several GPUs without Kubernetes by renting one container with multiple GPUs attached and running standard PyTorch DDP inside it. On Beam you ask for the GPUs with gpu="H100", gpu_count=4, the platform boots the node, your torchrun launcher spawns one process per GPU, and the machine scales to zero when the run finishes. No Slurm, no cluster to keep alive between jobs.
torchrun --standalone --nproc_per_node=4 starts four DDP workers on the four GPUs the container was given, train.py is an ordinary distributed script, and train.remote() ships the whole job to Beam. The only Beam-specific lines are the decorator and the packages. gpu_count is available on request — you message the team once and it's switched on for your account.
What multi-GPU training without Kubernetes actually requires
Splitting a training job across GPUs needs four things, and none of them is an orchestrator. Once you have them on a single machine, the job runs.
- Several GPUs on one node. Data-parallel training keeps a full copy of the model on each GPU and averages gradients across them, so the GPUs have to see each other. The simplest way to guarantee that is to put them in the same container — no cross-machine networking to configure.
- A fast interconnect and NCCL. The GPUs sync gradients every step over NCCL. Inside one node that traffic stays on NVLink or the local PCIe bus, which is why single-node multi-GPU is faster per dollar than spreading the same job thin across machines.
- A launcher.
torchrun(oraccelerate launch) starts one process per GPU, assigns ranks, and wires up the process group. This is the piece Kubernetes and Slurm usually own; on a serverless node you just call it yourself. - Storage that outlives the run. The container disappears when the job ends, so checkpoints have to land on a mounted volume or object storage, not the local disk.
Most fine-tuning fits this shape. A LoRA run on Llama 3 or a full fine-tune of a 7–13B model is a single-node multi-GPU job, and single-node is where the setup stays simple. If you have outgrown one GPU on Unsloth or a single-card fine-tune, this is the next step up, not a jump to distributed systems.
How to run multi-GPU training on a serverless GPU node
The walkthrough expands the hero snippet into the five moves that take a single-GPU script to four GPUs. The container is your whole "cluster," so most of the work is in the training script, not the infrastructure.
- Define the image. List your training dependencies with
Image(...).add_python_packages([...]). Beam builds it once and caches it, so the second run starts from a warm image instead of reinstalling PyTorch. - Ask for the GPUs. Set
gpu="H100"andgpu_count=4on the decorator. The available multi-GPU types are A10G, RTX 4090, and H100; H100s are the usual pick for training because 80GB of memory per card leaves room for optimizer state and activations. - Make `train.py` distribution-aware. A DDP script initializes the process group, moves the model onto its local rank, wraps it, and shards the data:
- Launch it inside the function. Shell out to
torchrun --standalone --nproc_per_node=4 train.pyfrom the Beam function, as in the hero snippet.--standalonemeans single-node, so there are no master addresses or node ranks to set. - Checkpoint to a volume. Attach a
Volumeand write checkpoints there so they survive after the node scales to zero. Save only from rank 0 to avoid four processes writing the same file.
Kick the whole thing off with train.remote(). Beam provisions the four-GPU container, runs the job, streams the logs back, and releases the hardware when train.py exits — the same scale-to-zero model that avoids running a standing Kubernetes cluster for GPU work.
Data parallel vs model parallel: choosing a multi-GPU strategy
Which strategy you use depends on one question: does the model fit on a single GPU? That answer decides whether you are replicating a model or splitting one.
| Strategy | When to use it | What it does |
|---|---|---|
| DDP (data parallel) | Model fits on one GPU | Full copy per GPU; each processes a different data shard; gradients averaged every step |
| FSDP / DeepSpeed ZeRO | Model too big for one GPU | Shards parameters, gradients, and optimizer state across GPUs; each holds a slice |
| Tensor / pipeline parallel | Very large models, expert use | Splits individual layers or the layer stack across GPUs |
DDP is the default and the fastest to set up — it is the code in the snippet above, and it is what most fine-tuning uses. Reach for FSDP when the model plus its optimizer state no longer fits in one card's memory; it trades some communication overhead for the ability to train models several times larger than a single GPU could hold. Hugging Face Accelerate exposes both behind one launcher, so switching DDP to FSDP is a config change, not a rewrite. On eight H100s, FSDP with LoRA comfortably reaches into the 70B range — past that, you are into the multi-node tensor-parallel territory that genuinely needs an HPC cluster.
What to look for in a platform for multi-GPU training
The platform's job is to hand you a multi-GPU node and then get out of the way. Judge the options on how much infrastructure you have to run and how you pay for idle time.
- No cluster to operate. A managed GPU node means you never provision, patch, or autoscale a Kubernetes or Slurm cluster to run one training job. On Beam the "cluster" is a decorator; the alternative is managing a GPU cluster yourself.
- Per-second billing and scale to zero. Training is bursty — hours of work, then nothing. A reserved multi-GPU cluster bills whether or not it is training, and an idle 8×H100 box is where budgets quietly disappear. Beam bills GPU time by the second and drops to zero between runs.
- GPU price at the sizes you use. Multi-GPU multiplies the hourly rate, so the per-card price matters more, not less. Beam's H100 is $1.83/hr versus $3.95/hr on Modal; across four cards that gap compounds every hour of every run.
- Reproducible environments. The image is defined in code and cached, so the run that trained last week rebuilds bit-for-bit today. It is the same serverless GPU model used for inference, pointed at a training job.
The honest limit: gpu_count attaches multiple GPUs to one container, so Beam covers single-node multi-GPU, not training spread across separate physical machines. That single node scales to eight modern GPUs, which is enough for the large majority of fine-tuning and mid-size training. If your model only trains across dozens of machines, a dedicated HPC scheduler is still the right tool.
Multi-GPU training approaches compared
| Approach | Cluster to manage | Multi-GPU setup | Billing | Scale to zero | Best for |
|---|---|---|---|---|---|
| Self-managed k8s / Slurm | Yes — you run it | You wire up device plugins and launchers | Reserved nodes | No | Teams already operating a cluster; true multi-node |
| Raw cloud GPU VM | You patch and babysit the box | Manual driver, NCCL, torchrun setup | Per-hour while on | No (you stop it) | One-off runs if you like managing servers |
| Managed serverless (Beam) | None | gpu_count + torchrun in a function | Per-second | Yes | Single-node multi-GPU fine-tuning and training |
| Notebook (Colab) | None | Limited; often single GPU | Per-hour / subscription | No | Prototyping, not production runs |
Slurm and Kubernetes still win when you already run the cluster or your job truly spans many machines. For everything short of that — the single-node multi-GPU jobs that cover most fine-tuning — a serverless node skips the standing infrastructure. If you are weighing hosted notebooks for this, the ceiling shows up fast; Colab alternatives covers where they stop.
FAQ
Can I run multi-node distributed training across separate machines on Beam? Not today. gpu_count attaches multiple GPUs to a single container, so Beam handles single-node multi-GPU training — up to eight GPUs on one node. That covers most fine-tuning and mid-size training. If a job genuinely has to span many physical machines, a dedicated HPC scheduler like Slurm or SkyPilot is still the right tool, and it is fair to say so.
How many GPUs can I attach to one training job? As many as fit on a single node — commonly two to eight. gpu_count is enabled on request, so you message the Beam team once and it is switched on for your account, then you set the number on the decorator.
Do I have to rewrite my training code for multi-GPU? No. A single-node DDP script — init_process_group, wrap the model in DistributedDataParallel, use a DistributedSampler — is the standard PyTorch pattern, and it runs unchanged. The only difference on a serverless node is that you call torchrun yourself instead of a cluster scheduler calling it for you.
When should I use FSDP instead of DDP? Use DDP when the model fits on one GPU and you just want to process more data in parallel. Switch to FSDP (or DeepSpeed ZeRO) when the model plus optimizer state no longer fits in a single card's memory — it shards them across GPUs so you can train a much larger model on the same hardware. Accelerate lets you flip between the two from config.
What does a multi-GPU run cost? You pay per second for the GPUs while the job runs and nothing when it is idle. Four H100s at $1.83/hr each is about $7.32/hr of training time, and the node bills zero the moment train.py exits — no reserved-cluster charge between runs.
Spin up a multi-GPU training run in a few minutes. Get started on Beam — pay only for the compute you use.


