beam-logo
← All posts
Tutorials

Batch Embeddings at Scale on Serverless GPU

Eli MernitEli Mernit
September 29, 20268 min read
Batch Embeddings at Scale on Serverless GPU

You can embed a large dataset by wrapping an open-source embedding model in a Beam @function with a GPU, then calling .map() over your chunks so the work fans out across many containers at once. Each container loads the model once and encodes its shard in GPU-sized batches; Beam scales the fleet to zero when the queue is finished and bills by the second. A ten-million-document job that takes days on a CPU finishes in hours, and you never send a token to a paid API.

Run it with python embed.py. embed.map(shards) sends each shard to its own container and sends results back as they finish, so the same code embeds four shards or four thousand without a cluster to manage.

What embedding a large dataset at scale actually involves

Embedding a big corpus is three steps — chunk the text, run it through a model, write the vectors somewhere — and the model call is the easy one. The work that decides your wall-clock time and bill is the plumbing around it: splitting the corpus into shards, getting a GPU for each shard, keeping every GPU busy with full batches, retrying the shards that die on a bad row, and shutting the whole fleet off the moment the last shard lands.

A single GPU helps but does not finish the job. An NVIDIA A10G or A100 can push thousands of text chunks per minute through a small model, so the difference between a corpus that embeds in three hours and one that embeds in three days is how many GPUs you can run in parallel and how cheaply you can turn them off in between. That is a throughput and scheduling problem, not a modeling one.

  • Chunking. Documents get split into passages that fit the model's context window (often 256–512 tokens) before encoding, because you retrieve on passages, not whole files.
  • GPU batching. Encoding one chunk at a time wastes the card. You feed the model a batch_size of a few hundred chunks so the GPU stays saturated — this is where most of the speedup lives.
  • Fan-out. One GPU is a bottleneck at ten million rows. The job needs to run many GPUs at once over disjoint shards, then release them.
  • Storage that outlives the run. Vectors have to land in a database or object store, because the container disappears when the shard finishes.

This is the offline, throughput-first cousin of a live endpoint, and it shares the mechanics of any batch inference job on a serverless GPU — the embedding model is just what sits inside the fan-out.

How to batch-embed a large dataset on a serverless GPU

The hero snippet is the whole path; the five moves below are what take it from a toy to a ten-million-row run. The container is your entire cluster, so almost all of the work is in the model call and the data movement, not the infrastructure.

  1. Load the model once with `on_start`. Pass a loader to on_start and read context.on_start_value inside the function. The loader runs a single time when a container boots, so the weights are already in VRAM before the first shard arrives instead of reloading per shard.
  2. Tune `batch_size` to the GPU. Start around 128–256 and raise it until you are near the card's memory ceiling. A larger batch keeps the GPU busy and is the single biggest lever on throughput; too large and you hit an out-of-memory error, so leave headroom.
  3. Fan out with `.map()`. embed.map(shards) runs each shard on its own container concurrently and streams results back as they complete. Beam owns the queue and the autoscaling, so you scale from four to four thousand shards by changing the input, not the code.
  4. Write vectors as they land. Upsert each returned batch straight into your vector store — pgvector, Qdrant, or a plain object store for later indexing. Stream results in as .map() yields them so you never hold the whole corpus in memory.
  5. Re-embed only new rows on a schedule. Most datasets grow. Deploy the same function behind @schedule with a cron expression and embed just the rows added since the last run, so the nightly job is small and the GPUs fire only when there is work.

For inputs that trickle in over time rather than arriving as one list, swap .map() for a task_queue with retries set — the same pattern the batch inference walkthrough covers in depth, applied here to embeddings.

Choosing an embedding model for a batch job

For a large offline job, favor a small open-source model that batches well on a GPU over a giant one. Models like bge-small-en, e5-small, and all-MiniLM-L6-v2 produce 384-dimensional vectors, run several times faster per chunk than a 7B model, and score close enough on retrieval benchmarks that the throughput win usually pays off. Step up to a bge-large or a multilingual model only when your evaluation shows the smaller one missing.

Two knobs matter at scale: dimension and speed. Smaller vectors cost less to store and search across millions of rows, and a faster model is fewer GPU-hours for the same corpus. The full tradeoff — accuracy, dimensionality, context length, and where each model wins — is in our guide to choosing an embedding model; this page assumes you have picked one and want to run it over everything.

Self-hosting vs a per-token embedding API: the cost crossover

A hosted API is the right call until the dataset gets big, and then it flips. OpenAI's text-embedding-3-small is $0.02 per million tokens and text-embedding-3-large is $0.13 per million tokens, with the batch API roughly half that. Embedding ten million passages of about 256 tokens each is ~2.56 billion tokens: about $51 on the small model, or about $333 on the large one, per full pass — and you pay it again every time you re-embed after a model or chunking change.

Running an open-source model yourself replaces that per-token meter with a fixed amount of GPU time. You pay for the GPU only while the job runs and nothing between runs, there is no per-token charge no matter how many passes you make, your documents never leave your own infrastructure, and there is no API rate limit throttling a ten-million-row backfill. The API stays simpler for a few thousand documents; self-hosted batch on a serverless GPU wins on cost, data residency, and control once you are embedding at scale or re-embedding often.

What to look for in a platform for batch embeddings

The platform's only job here is to hand you GPUs while the job runs and charge you nothing when it does not. Judge the options on how much infrastructure you operate and how you pay for idle accelerators.

  • Per-second billing and scale to zero. Embedding is bursty — a few hours of heavy GPU work, then nothing until the next batch. A reserved GPU box bills around the clock; Beam bills GPU time by the second and drops to zero when the queue drains, so an overnight backfill costs the hours it ran, not the day it sat idle.
  • Real fan-out, not a single card. .map() turns one input list into hundreds of concurrent GPU containers with no queue or cluster to run, which is the difference between hours and days at ten million rows.
  • Throughput-friendly GPUs. A mid-range card like an A10G is often the sweet spot for small embedding models — enough VRAM for a large batch_size without paying for an H100 you cannot saturate.
  • Data stays in your infrastructure. Self-hosting the model means the corpus and the vectors never transit a third-party API, which matters for private or regulated data.
  • Reproducible environments. The image is defined in code and cached, so the job that embedded last month rebuilds identically today — the same serverless GPU model used for inference, pointed at an embedding backfill.

Batch embedding approaches compared

ApproachInfra to runCost modelData residencyRate limitsBest for
Hosted embedding APINonePer token, every passLeaves your infraYesA few thousand docs; no ops
Self-hosted batch on serverless GPUNone — .map() in a functionPer GPU-second, scales to zeroStays in your infraNoMillions of docs; frequent re-embeds
Spark / Ray GPU clusterYou run the clusterReserved nodes while upStays in your infraNoTeams already operating a cluster

A hosted API is the fastest way to embed a small collection and skip all operations. A Spark or Ray cluster makes sense when you already run one for other work. For the middle case most teams are in — a large corpus, embedded on a serverless GPU without standing infrastructure, re-run whenever the data or model changes — the fan-out job is the cheapest path that keeps the data yours.

FAQ

How long does it take to embed millions of documents? It depends on the model and how many GPUs you run in parallel, not on any single card. A small model like MiniLM or bge-small pushes thousands of chunks per minute per GPU, so a ten-million-passage corpus that would take days on one CPU finishes in a few hours when .map() spreads it across dozens of GPU containers at once. The more shards you fan out to, the shorter the wall-clock time.

What batch size should I use for embedding on a GPU? Start at 128–256 and raise it until VRAM is nearly full. A bigger batch keeps the GPU saturated and is the largest single lever on throughput, but too big triggers an out-of-memory error, so leave a little headroom. Small models on a mid-range card like an A10G comfortably handle a few hundred chunks per batch.

Is it cheaper to self-host an embedding model or use an API? For a few thousand documents, a hosted API is simpler and the cost is trivial. Once you are embedding millions of rows — or re-embedding often after model or chunking changes — self-hosting an open-source model on a serverless GPU is usually cheaper, because you pay for GPU time only while the job runs instead of a per-token fee on every pass, and the data never leaves your infrastructure.

Which embedding model should I use for a large batch job? Favor a small, fast open-source model such as bge-small, e5-small, or all-MiniLM-L6-v2 — they batch efficiently, produce compact 384-dimensional vectors that are cheap to store and search, and score close to much larger models on retrieval. Move up to a large or multilingual model only if your evaluation shows the small one falling short. Our embedding model guide covers the tradeoffs.

How do I embed only new rows instead of the whole dataset each time? Deploy the embedding function behind @schedule with a cron expression and query for rows added since the last run, so each scheduled job embeds just the new data. Beam spins up the GPUs for that smaller batch and shuts them off when it finishes, so an incremental nightly re-embed costs a fraction of a full backfill.

Embed your whole corpus in an afternoon, not a week. Get started on Beam — pay only for the compute you use.

Eli Mernit
Eli Mernit
Published September 29, 2026
Pay as you gobilled by the millisecond

Start shipping on infra
you won’t outgrow.

Run sandboxes and GPU workloads on your cloud, and scale out to ours when you need to. No infra to manage.