beam-logo
← All posts
Engineering

GPU Cron Jobs: Run Scheduled Jobs on GPUs

Eli MernitEli Mernit
September 29, 20268 min read
GPU Cron Jobs: Run Scheduled Jobs on GPUs

To run a GPU job on a schedule, attach a cron trigger to a function on a serverless GPU platform. The GPU starts when the schedule fires and shuts off when the function returns, so you pay for the minutes the job runs rather than for a machine that waits all day. On Beam, that means one @schedule decorator with a cron expression and a gpu argument, followed by beam deploy.

Deploy it with beam deploy app.py:embed_new_docs. Every night at 03:00 UTC, Beam starts a container with an RTX 4090, runs the function against whatever arrived in the inbox, writes the vectors to the volume, and scales back to zero.

What a scheduled GPU job involves

A scheduled GPU job is ordinary batch work (retraining, re-embedding a corpus, scoring yesterday's records, rendering a report) that needs a GPU for a few minutes or hours on a fixed timetable. The hard part is not the cron syntax. It's avoiding paying for a GPU that sits idle between runs, while still handling runs that start slowly, run long, overlap, or fail.

The questions to answer before picking a setup:

  • Idle cost. Does anything keep billing between runs, such as a VM, a node pool, or a monthly plan fee?
  • Start-up time. How long does it take from the cron tick to your code running on a GPU, and are you billed for that time?
  • Run length. Is there a hard cap on how long a GPU task can run?
  • Overlap. If Tuesday's run is still going when Wednesday's fires, what happens?
  • Time zone. Is the schedule read in UTC or in a zone you choose, and what happens around daylight saving changes?

How to schedule a GPU job with cron

On Beam you define the job as a Python function, put the schedule and hardware in the @schedule decorator, and deploy it. There is no separate scheduler service, IAM role, or queue to wire up. The same decorator takes the arguments a normal Beam function does, including gpu, cpu, memory, image, volumes, secrets, timeout, and retries.

Writing the cron expression

The when argument takes a standard five-field cron expression (minute, hour, day of month, month, day of week) or a predefined macro. Beam always reads the schedule in UTC, so convert your local time before you write it.

ScheduleExpressionRuns at
Hourly@hourlyMinute 0 of every hour
Nightly@daily00:00 UTC every day
Weekly@weekly00:00 UTC every Sunday
Monthly@monthly00:00 UTC on the 1st
Every 6 hours0 /600:00, 06:00, 12:00, 18:00 UTC
Weekdays at 07:30 UTC30 7 1-5Monday to Friday
Every 15 minutes/15 *Four times an hour

A job meant to run at 03:00 in New York needs 0 7 * * * during daylight saving time and 0 8 * * * the rest of the year. If exact local wall-clock time matters, the platform comparison below lists schedulers that accept a time zone.

Deploying and checking the next runs

beam deploy builds the image, registers the schedule, and prints the next three run times in UTC and in your local zone, which makes it easy to catch a wrong hour before it costs anything.

Running it once by hand

You don't need to wait for the first tick to find out whether the job works. Call the function's .remote() method to run it once on the same GPU and image, with logs streamed to your terminal.

Changing or stopping a schedule

Deploying a new version replaces the old schedule, so editing when and redeploying is how you change it. To stop a schedule, find the deployment ID and stop that deployment.

What a scheduled GPU run costs compared with an always-on GPU

A scheduled run on serverless GPU is billed only while the container runs, so a job that takes 20 minutes a night costs about 20 minutes of GPU time. An always-on machine costs the same whether the job runs or not. For most recurring jobs the scheduled run is far cheaper; the flat-rate machine only wins once the job runs for roughly six hours a day or more.

The figures below use Beam's published rates. For serverless, that's an RTX 4090 with 2 physical cores and 16 GiB of RAM at the GPU-attached rates, which Beam's pricing page puts at $1.77 an hour ($0.69 GPU, $0.76 CPU, $0.32 RAM). The always-on figure is an on-demand RTX 4090 machine at $0.44 an hour, which includes 8 vCPU and 64 GB of RAM, left on for a 730-hour month.

Nightly job lengthScheduled serverless run, 30 runs a monthAlways-on on-demand RTX 4090
20 minutesabout $17.70about $321
1 hourabout $53about $321
3 hoursabout $159about $321
6 hoursabout $319about $321

Two details affect the serverless numbers. Beam doesn't bill for spinning up the server or loading the container image, but it does bill for the time your code spends loading, so a job that downloads a 10 GB model on every run pays for that download each time. Keep weights on a volume and the load step shrinks to a read from disk. The RTX 4090 price guide covers what the same card costs to buy or rent elsewhere.

How to handle overlapping runs, retries, and failures

Beam starts a new run on every tick, whether or not the previous run has finished. The simplest guard is to set timeout below the gap between runs, so a stuck run is stopped before the next one starts, and to make the job idempotent, so a repeat run for the same window does nothing. Retries and completion callbacks cover the remaining failure cases.

Idempotency usually means keying output on the scheduled window rather than on the current time. A retry that starts at 02:10 still belongs to the 00:00 window:

The other settings that matter for recurring jobs:

  • `timeout` defaults to 3,600 seconds. Raise it for long jobs, or set it to -1 to remove the limit, but keep it under the schedule interval if overlap would cause trouble.
  • `retries` defaults to 3 and applies when the container crashes. Lower it for jobs where a rerun is expensive and a human should look first.
  • `callback_url` receives a request when a task completes, times out, or is cancelled. Point it at your alerting webhook so a failed nightly run gets noticed the next morning, not the next week.

Kubernetes handles overlap at the scheduler level with concurrencyPolicy: Forbid, which skips a tick while the last job is still running. If you would rather have that than a guard in code, it's a real point in Kubernetes' favor, and the next section weighs it against running your own GPU nodes.

Choosing a platform for scheduled GPU jobs

The right platform depends on how long the job runs, which GPU it needs, and how much infrastructure you want to own. Beam and Modal put the schedule in the function decorator and scale to zero between runs. The hyperscaler options pair a scheduler with a separate compute service. Kubernetes gives the most control over scheduling behavior and the most nodes to look after.

BeamModalCloud Run jobs + Cloud SchedulerEventBridge Scheduler + AWS BatchKubernetes CronJob
Where the schedule lives@schedule(when=...) in codemodal.Cron or modal.Period in codeCloud Scheduler jobEventBridge scheduleCronJob manifest
Time zoneUTC onlyAny zone on modal.CronConfigurableAny IANA zoneAny IANA zone (timeZone, stable since 1.27)
GPUsRTX 4090, RTX 5090, H100 and moreH100, A100, L40S, L4, A10 and moreL4, RTX PRO 6000 BlackwellAny EC2 GPU instance you configureWhatever your node pools have
Longest GPU runConfigurable, or no limit24 hours1 hourSet by your job timeoutNo limit by default
Overlap controlGuard in codeNot documentedNot documentedNot documentedconcurrencyPolicy
Pieces to set upDecorator and beam deployDecorator and modal deployJob, scheduler, service accountSchedule, IAM role, compute environment, job queue, job definitionCluster, GPU node pool, manifest
Cost between runsNothing on the free Developer planNothing on Starter, $250/month on TeamNothing for the jobNothing if the compute environment scales to zeroGPU nodes unless the pool scales to zero

Figures come from the Beam pricing page and SDK, Modal's pricing, scheduling, and timeout docs, Google Cloud's Cloud Run GPU and task-timeout docs, AWS EventBridge Scheduler docs, and the Kubernetes CronJob docs.

Where each one fits:

  • Beam suits recurring jobs that need a consumer or data-center GPU, might run for hours, and shouldn't need a separate scheduler service. The trade-off is UTC-only schedules and overlap guarded in code rather than by the scheduler.
  • Modal has the closest developer experience and better time zone support. Its Starter plan allows 5 deployed crons; unlimited scheduled functions need the $250/month Team plan. The Modal pricing breakdown covers how its per-second GPU rates compare.
  • Cloud Run jobs are a reasonable pick if you're already on Google Cloud and the job fits in an hour on an L4 or RTX PRO 6000. The one-hour cap on GPU tasks rules it out for retraining and large backfills.
  • EventBridge with AWS Batch gives you flexible time windows, time zones, and any GPU instance type, at the price of five or so resources to create and keep in sync.
  • Kubernetes CronJob makes sense when you already run a GPU cluster. Starting one just for a nightly job means paying for, or autoscaling, GPU nodes you would otherwise not need. We wrote about why we don't use Kubernetes to scale our GPU workloads.

For jobs too large for one GPU, schedule a function that splits the work and fans it out across containers; the batch inference on serverless GPU guide shows the fan-out pattern with .map().

FAQ

Can you run a cron job on a GPU?

Yes. Any platform that can start a GPU container from a trigger can run a cron job on a GPU. Serverless GPU platforms such as Beam and Modal accept the cron expression directly in code, while Google Cloud and AWS pair a scheduler service with a GPU-capable compute service.

Which providers support scheduled GPU training?

Beam, Modal, Google Cloud (Cloud Scheduler with Cloud Run jobs), AWS (EventBridge Scheduler with Batch), and any Kubernetes cluster with GPU nodes all support it. For training runs longer than an hour, rule out Cloud Run jobs, whose GPU tasks are capped at one hour, and check that the platform's maximum task length covers your longest run.

What is the cheapest serverless option for scheduled GPU jobs?

The cheapest option is a platform that bills per second, charges nothing between runs, and has no plan fee for the number of schedules you need. Beam's Developer plan and Modal's Starter plan both cost nothing when idle; Modal's Starter caps you at 5 deployed crons. After that, compare the GPU-hour rate for the specific card you need, plus CPU and memory, because those are billed separately on serverless platforms.

What time zone do Beam scheduled jobs use?

UTC, always. Convert your local target time to UTC when writing the expression, and remember that the UTC hour for a local time shifts by one when daylight saving starts or ends. beam deploy prints the next three runs in both UTC and your local zone so you can check.

Can a scheduled GPU job run longer than an hour?

On Beam, yes. The default timeout is one hour; raise it to whatever the job needs, or set it to -1 for no limit. Modal allows up to 24 hours. Cloud Run jobs cap GPU tasks at one hour.

How do I stop a scheduled job on Beam?

Run beam deployment list to find the deployment ID, then beam deployment stop <deployment-id>. Deploying a new version of the same job also replaces the previous schedule.

Run your recurring GPU jobs on Beam

Put a cron expression and a GPU in one decorator, deploy it, and pay only for the minutes the job runs. See more patterns on the batch processing page, or start on Beam with the free Developer plan.

Eli Mernit
Eli Mernit
Published September 29, 2026
Pay as you gobilled by the millisecond

Start shipping on infra
you won’t outgrow.

Run sandboxes and GPU workloads on your cloud, and scale out to ours when you need to. No infra to manage.