beam-logo
← All posts
Engineering

How to Run a Python Function on a Cloud GPU

Eli MernitEli Mernit
September 29, 20269 min read
How to Run a Python Function on a Cloud GPU

To run a Python function on a cloud GPU without setting up a server, put a decorator on the function that names the GPU, then call it with .remote(). Beam syncs your code, starts a container with that GPU, runs the function, streams the logs and return value back to your terminal, and shuts the container down. You pay for the seconds it ran.

Save it as app.py and run it the way you run any script:

The notes/ folder is read on your machine, the model runs on an RTX 4090, and the summaries print locally. Nothing keeps running after the function returns.

What running Python on a cloud GPU involves

Running Python on a cloud GPU means getting four things to a machine you don't own: your code, its dependencies (including a CUDA-compatible PyTorch or similar), your input data, and a GPU big enough for the model. Then you need the results back and the billing to stop when the work is done. Most of the friction sits in those last two parts, not in the GPU itself.

The questions worth answering before you pick a tool:

  • Dependencies. Do you write a Dockerfile, or can you list packages in Python?
  • Code transfer. Does the platform sync your working directory, or do you build and push an image for every change?
  • GPU choice. Which cards can you get on demand, and how much memory do they have?
  • Results. Do return values come straight back, or do you fetch files from storage?
  • Stopping. Does billing end when the function returns, or when you remember to shut the machine down?

How to run a Python function on a cloud GPU

On Beam, the whole workflow is a decorator and a normal python command. The @function decorator holds the hardware (gpu, cpu, memory), the environment (image), and limits such as timeout. Calling .remote() runs the function in the cloud, and .local() runs the same function on your machine, which is useful for quick tests on small inputs.

Install the CLI and add your token

Create an account at platform.beam.cloud, copy an API key from Settings, and install the client:

The token is saved to ~/.beam/config.ini. In CI or other short-lived environments, exporting BEAM_TOKEN works instead of a config file.

Pick a GPU

Pass the GPU type as a string. Beam's docs list these serverless types: T4 (16 GB), A10G (24 GB), RTX4090 (24 GB), and RTX5090 (32 GB). Larger cards, including H100, H200, A100 80GB, and L40S, are available as dedicated on-demand machines. Run beam machine list to see live availability.

If you don't mind which card you get, pass a list in priority order, and Beam uses the first one with capacity:

A model's weights in 16-bit precision take roughly two bytes per parameter, so a 7B model needs about 14 GB of GPU memory before activations and batch size. The 24 GB cards fit that; anything larger needs quantization or a bigger GPU. The RTX 4090 price guide covers what that card costs to rent elsewhere.

Get arguments in and results out

Arguments and return values work like a normal function call. Anything you pass to .remote() is sent to the container, and whatever the function returns is sent back, so keep both small: strings, numbers, lists, dicts, or modest arrays.

For large inputs or outputs, such as datasets, checkpoints, or generated images, use a volume instead. A volume is persistent storage mounted into the container at a path you choose, and you can copy files in and out from the CLI:

Beam also syncs the files in your working directory to the container on each run, which is how app.py above can import local modules. Add a .beamignore file to keep large or private files out of the upload.

Run an existing script without rewriting it

You don't need to restructure a working script to move it onto a GPU. Because the working directory is synced, a small wrapper can import the script and call its entry point inside the container:

Two details matter here. Import heavy packages such as torch inside the function, not at the top of the file, because the top of the file also runs on your laptop, which may not have them installed. And write outputs (checkpoints, logs you want to keep) to a volume, since the container's disk is gone once the function returns.

Cache model weights on a volume

A function that downloads a large model from Hugging Face on every call pays for that download each time. Beam doesn't bill for starting the machine or pulling the container image, but it does bill for the time your code runs, and a download counts. Point the Hugging Face cache at a volume and only the first run downloads:

Set HF_HOME before importing transformers, since the library reads it at import time.

Keep a long job running after you close the terminal

By default, a remote function stops when your local Python process exits, so closing the laptop or losing Wi-Fi ends the run. Set headless=True and the function keeps going on Beam after you disconnect. Pair it with a timeout that covers the whole job: the SDK default is one hour, and timeout=-1 removes the limit.

Check on it later with beam logs, and have the job write checkpoints to a volume at regular intervals, so a crash at hour ten doesn't cost you the first nine. If the job has become a regular part of your week, the GPU training use case covers longer runs.

Debug on the GPU with a shell

When something works locally and fails in the container, open a shell in the same environment the function uses:

You get a terminal inside a container with the same image, GPU, and synced files, so you can run nvidia-smi, check package versions, or step through the script by hand.

Run the same function on many GPUs at once

If you have many independent inputs, call .map() instead of .remote(). Beam starts a container for each input and returns results as they finish, up to your plan's GPU concurrency limit (5 GPU containers on the free Developer plan, 50 on Team). The batch inference on serverless GPU guide walks through chunking a large dataset this way.

What it costs to run Python on a cloud GPU

A serverless GPU function is billed only while it runs, so a five-minute job costs five minutes of GPU time. The price per hour is higher than renting a whole machine, though. A serverless function costs less when the GPU would otherwise sit idle most of the time, and a rented machine costs less when you keep it busy for most of the hours you have it.

The figures below use Beam's published rates. The serverless column is an RTX 4090 with 2 physical cores and 16 GiB of RAM at the GPU-attached rates, which Beam's pricing page puts at $1.77 an hour ($0.69 GPU, $0.76 CPU, $0.32 RAM). The on-demand column is a dedicated RTX 4090 machine at $0.44 an hour, which includes 8 vCPU and 64 GB of RAM and is billed for as long as you keep it.

ScenarioServerless function, billed while runningOn-demand RTX 4090 machine
One 5-minute job, machine kept up for an hourabout $0.15about $0.44
One 30-minute job, machine kept up for an hourabout $0.89about $0.44
2 hours of jobs spread over an 8-hour workdayabout $3.54about $3.52
8 hours of steady trainingabout $14.16about $3.52

The break-even sits near 25% utilization: $0.44 divided by $1.77. If the GPU would be busy less than a quarter of the time the machine is up, the serverless function is cheaper, and you never have to remember to stop anything. For long training runs that keep the GPU busy, an on-demand machine is the better deal.

CPU and memory are billed separately on serverless, and at a higher rate when attached to a GPU. A function that asks for 8 cores it never uses pays for them anyway, so request what the job needs. Dropping the example above to 1 core and 8 GiB brings it to about $1.23 an hour.

Choosing a way to run Python on a cloud GPU

The options split into three groups. Serverless function platforms (Beam, Modal) run a decorated function and stop billing when it returns. GPU pods (RunPod Pods, on-demand machines anywhere) give you a whole machine that bills until you stop it. Notebooks (Google Colab) are the fastest way to try something interactively, with runtime limits that make them a poor fit for long or scheduled jobs.

Beam functionModal functionRunPod PodRunPod ServerlessGoogle Colab
How code gets thereDecorator; working directory synced on each runDecorator; modal run app.pySSH, Jupyter, or your own container imageHandler function in a Docker image you build and pushNotebook cells
Mid-size GPU, listed rateRTX 4090 at $0.69/hr, plus CPU and RAML4 at $0.80/hr or A10 at $1.10/hr, plus CPU and RAMRTX 4090 at $0.74/hr, with 6 vCPU and 41 GB RAM included4090 PRO worker at $1.10/hrDepends on plan and compute units
Billing stops whenThe function returnsThe function returnsYou stop the podThe worker scales downThe runtime disconnects or times out
Longest runConfigurable; -1 for no limit24 hoursAs long as the pod runsSet per endpointUp to 12 hours; up to 24 on Pro+
Keeps running after you disconnectWith headless=TrueWith modal run --detachYesYesOnly with background execution on paid plans
Monthly fee or free creditDeveloper plan $0, Team $89/monthStarter $0 with $30/month of free computeNoneNoneFree tier, plus paid plans

Rates come from the Beam pricing page and docs, Modal's pricing and timeout docs, RunPod's pricing page, and the Colab FAQ.

Where each one fits:

  • Beam suits Python code that should run on a consumer or data-center GPU for minutes to hours, with no Dockerfile and no machine to shut down afterward. Its docs list four serverless GPU types, fewer than Modal offers; H100-class cards are listed as on-demand machines.
  • Modal has a similar decorator workflow and a wider range of serverless GPUs, from T4 up to B300, plus $30 a month of free compute on its Starter plan. It doesn't offer an RTX 4090, and its default function timeout is five minutes, so long jobs need an explicit timeout. The Modal pricing breakdown compares its per-second rates in detail.
  • RunPod Pods are cheap for all-day work and give you a full machine you can SSH into. You manage the environment yourself and pay until you stop the pod, which is the usual way a forgotten GPU runs up a bill.
  • RunPod Serverless is built for inference endpoints. Running a one-off function means writing a handler, building a Docker image, and deploying it first.
  • Google Colab is the quickest way to try a GPU from a browser, and the free tier costs nothing. Idle timeouts and the 12-hour cap make it a poor home for long training or anything that has to run unattended. See Google Colab alternatives for other notebook options.

For a wider view of the serverless platforms and how quickly each one starts a GPU container, see the top serverless GPU providers and the explainer on how serverless GPUs work.

FAQ

Can I run Python on a GPU without owning one?

Yes. Cloud GPU platforms rent GPU time by the second or the hour. With a serverless platform such as Beam or Modal, you add a decorator that names the GPU and call the function; the platform provides the machine, installs your packages, and shuts it down when the function returns.

How much does it cost to run Python on a cloud GPU?

It depends on the GPU and how long the code runs. On Beam, an RTX 4090 with 2 CPU cores and 16 GiB of RAM costs about $1.77 an hour on serverless, billed only while the function runs, so a 5-minute job is about $0.15. A dedicated on-demand RTX 4090 machine is $0.44 an hour but bills for as long as you keep it.

Do I need Docker to run Python on a cloud GPU?

Not on Beam or Modal. Both build the container for you from a list of Python packages declared in code. You can still bring an existing Docker image if you have one. RunPod Serverless does require you to build and push your own image.

How do I run a background training job on a cloud GPU and disconnect safely?

Use a platform option that keeps the job running after the client exits: headless=True on a Beam function, or modal run --detach on Modal. Set a timeout long enough for the whole job, write checkpoints to persistent storage as you go, and check progress later with the platform's logs command.

Am I billed while the GPU starts?

On Beam, no. You aren't charged for waiting for a machine to start or for pulling your container image. Billing starts when your code starts running, which includes any time your code spends downloading or loading a model.

Is Google Colab enough for running Python on a GPU?

For interactive experiments, often yes. For long or unattended work, Colab's limits get in the way: runtimes time out when idle, free sessions run for at most 12 hours, and background execution needs a paid plan. A function-based platform suits jobs that should run to completion without a browser tab open.

Run your Python on a GPU with Beam

Add one decorator, pick a GPU, and run the script you already have. Billing stops when the function returns. See more patterns on the GPU training page, or start on Beam with the free Developer plan.

Eli Mernit
Eli Mernit
Published September 29, 2026
Pay as you gobilled by the millisecond

Start shipping on infra
you won’t outgrow.

Run sandboxes and GPU workloads on your cloud, and scale out to ours when you need to. No infra to manage.