beam-logo
← All posts
Engineering

Using OpenAI-Compatible API for Open-Source LLMs

Eli MernitEli Mernit
September 29, 20269 min read
Using OpenAI-Compatible API for Open-Source LLMs

To host an OpenAI-compatible API for an open-source LLM, run vLLM's OpenAI server on a GPU behind an authenticated HTTPS URL, then set base_url and api_key in the OpenAI SDK. On Beam, that's one VLLM object in our SDK and beam deploy. You get a /v1 endpoint that autoscales and shuts off when traffic stops.

Deploy it:

The deploy prints a URL like https://qwen25-7b-1a2b3c4-v1.app.beam.cloud. Add /v1 and use it like the OpenAI API:

Qwen2.5-7B-Instruct is Apache-2.0 licensed, so it downloads without a Hugging Face token. The weights take about 15 GB in 16-bit precision, which leaves room on the RTX 4090's 24 GB for the KV cache.

What an OpenAI-compatible LLM API involves

An OpenAI-compatible API is a server that accepts the same requests as OpenAI's API (/v1/chat/completions, /v1/completions, /v1/models) and returns the same response objects. Any client built for OpenAI then works against it once you change the base URL and key. Most of the effort goes into running that server where it stays up, stays private, and doesn't bill you while idle.

To host one yourself, you need:

  • An inference server. vLLM is the common choice. It ships an OpenAI-compatible server and batches concurrent requests on the GPU. SGLang and Hugging Face TGI are alternatives.
  • A GPU sized to the model. The weights, plus KV cache for your context length, have to fit in GPU memory.
  • A public URL with TLS and auth. vLLM on its own is a bare HTTP server on a port.
  • Scaling. More replicas when traffic spikes, and ideally none when nobody is calling it.
  • Cost control. A GPU that sits idle overnight is the biggest line item for low-traffic APIs.

Most guides stop at vllm serve on a single machine, or jump to Kubernetes with KEDA for autoscaling. A serverless GPU platform covers the last three items for you.

How to host an OpenAI-compatible API on a serverless GPU

Our SDK's VLLM class wraps vLLM's OpenAI server as a deployable app. You pass the hardware (gpu, cpu, memory), scaling settings, and a VLLMArgs object whose fields match vllm serve flags. We build the container, mount a volume to cache model weights, put the server behind an authenticated URL, and scale it with traffic.

Install the CLI and deploy

Create an account at platform.beam.cloud, copy an API key, and install the client:

Save the hero snippet as app.py and run beam deploy app.py:qwen. The first container downloads the model into a volume named vllm_cache, so later cold starts read the weights from the volume instead of pulling them from Hugging Face again.

Point existing OpenAI code at the endpoint

The OpenAI Python SDK reads OPENAI_BASE_URL and OPENAI_API_KEY from the environment. If your code already calls OpenAI() with no arguments, you can switch it to your own model without editing it:

The one change you can't avoid is the model string. It has to match served_model_name, and client.models.list() returns it if you're unsure. The server also exposes /v1/completions for base models and /v1/embeddings for embedding models.

Endpoints are authenticated by default: requests need your Beam token as a Bearer token, which is exactly what the OpenAI SDK sends as api_key. Set authorized=False only if you put your own auth in front of it.

Let vLLM batch requests: raise concurrent_requests

vLLM's throughput comes from continuous batching. It runs many sequences through the GPU together, so 32 concurrent requests take far less than 32 times as long as one. The VLLM class defaults to concurrent_requests=1, which sends one request at a time to each container and leaves that batching unused.

Set concurrent_requests to the number of requests you want each GPU to handle at once. For a 7B model on a 24 GB card with 8K context, 16 to 64 is a reasonable starting range. Raise it until time to first token gets worse than you're willing to accept. Keep tasks_per_container in the autoscaler equal to it, so we only add a container when the current ones are full.

Scale out and scale to zero

The autoscaler adds containers as the request queue grows, up to max_containers. Its default is one container, so set it explicitly. When requests stop, each container stays up for keep_warm_seconds (60 by default for VLLM) and then shuts down. With no containers running, you pay nothing.

The trade-off is the cold start. The next request after a quiet period waits for a GPU container to boot and load the weights. We don't bill for machine startup or image pulls, but model loading does count as billed time. If users can't wait for that, keep one container running:

That container bills around the clock, so check the cost section below before you set it.

Serve gated or private models

Llama and some Mistral checkpoints on Hugging Face are gated. Store a Hugging Face token as a Beam secret and pass it to the app:

The same pattern works for your own fine-tuned weights pushed to a private Hugging Face repo. The Llama 3 fine-tuning guide walks through training one and serving it with vLLM.

Turn on tool calling and streaming

OpenAI-style function calling needs two vLLM flags: enable_auto_tool_choice=True and a tool_call_parser that matches the model's format. Qwen2.5 uses hermes, Mistral models use mistral, and Llama 3.1 uses llama3_json. With those set, tools=[...] in chat.completions.create returns tool_calls the same way OpenAI does.

Streaming works as it does with OpenAI: pass stream=True and iterate over the chunks. Our vLLM chat example in the beam-cloud/examples repo streams responses this way.

Pin the vLLM version

The VLLM class installs vLLM 0.8.4 by default. That release supports Qwen2.5, Llama 3.x, and Mistral, but models released after it may need a newer vLLM. Set vllm_version="..." to change it, and match huggingface_hub_version if the new release needs it. Newer vLLM releases add flags that VLLMArgs doesn't list; the SDK docs say you may need to subclass VLLMArgs to pass them.

Pinning also protects you in the other direction. A deploy that worked last month keeps working, because the image doesn't pick up a new vLLM release on its own.

Which GPU to pick for an open-source LLM

Size the GPU by the model's weights plus KV cache. In 16-bit precision, weights take about 2 bytes per parameter; 4-bit quantization (AWQ or GPTQ) cuts that to roughly 0.5 to 0.6 bytes. Leave several GB on top for the KV cache, which grows with context length and concurrent requests.

Model sizeWeights in 16-bitWeights in 4-bitGPU that fits
7B to 8Babout 15 to 16 GBabout 5 GBRTX 4090 or A10G (24 GB)
14Babout 28 to 30 GBabout 9 GB24 GB card with 4-bit, or RTX 5090 (32 GB) with short context in 16-bit
32Babout 64 GBabout 18 to 20 GB24 GB card with 4-bit and short context, or an 80 GB card
70Babout 140 GBabout 40 GBTwo or more 80 GB GPUs

Beam offers four serverless GPU types: T4 (16 GB), A10G (24 GB), RTX 4090 (24 GB), and RTX 5090 (32 GB). H100, H200, A100 80GB, and L40S are available as dedicated on-demand machines. Multi-GPU containers (gpu_count above one) need to be enabled on your account, which matters for 70B-class models. If you run a 4-bit checkpoint, set quantization="awq" (or the matching method) in VLLMArgs. The Qwen 2.5 overview covers which sizes of that family are worth serving.

What it costs to host your own LLM API

On Beam's serverless pricing, the hero config (RTX 4090, 2 CPU cores, 16 GiB of RAM) costs about $0.00049 per second, or $1.76 per hour while a container runs. That breaks down as $0.000192/s for the GPU, $0.000105 per core per second for CPU attached to a GPU, and $0.0000055 per GiB per second for RAM.

What you pay depends on how often a container is running:

Traffic patternContainer timeApproximate cost
A burst of requests, then quietRequest time plus 300 s keep-warmAbout $0.15 of idle time per burst
Office hours, 8 h a day, 22 daysAbout 176 h a monthAbout $310 a month
Always on with min_containers=1About 730 h a monthAbout $1,290 a month

Always-on is where serverless stops making sense. We also rent a dedicated on-demand RTX 4090 machine from $0.44 an hour, about $320 for a full month. Serverless is cheaper while the endpoint runs less than about a quarter of the time. Past that, a dedicated machine wins.

Hosted per-token APIs are the other comparison. For a popular model at low or spiky volume, paying per token is often cheaper than any GPU you rent yourself. Self-hosting pays off when you need a fine-tuned or niche model, want requests to stay in infrastructure you control, or have enough steady traffic to keep a GPU busy. The Modal pricing breakdown goes deeper on where per-second GPU billing gets expensive.

Choosing where to host an OpenAI-compatible endpoint

The main options are a serverless GPU platform that runs vLLM for you, a managed endpoint service, or vLLM on a GPU machine you manage. They differ in setup effort, idle cost, and who handles auth and scaling.

Beam VLLMModalRunPod Serverless vLLMHugging Face Inference EndpointsvLLM on a rented GPU VM
SetupOne Python object, beam deployPython class that starts vllm serve as a subprocessPick the vLLM worker in the console, enter a model namePick a model and instance in the consoleInstall vLLM, run vllm serve, add TLS and a proxy
Auth by defaultBeam token requiredThe official example is unauthenticatedRunPod API keyConfigurable per endpointNone until you add it
24 GB GPU, listed rateRTX 4090 at $0.69/hr plus CPU and RAML4 at $0.80/hr or A10 at $1.10/hr plus CPU and RAML4, A5000, or 3090 at $0.69/hr, RTX 4090 at $1.10/hrL4 at $0.80/hr, A10G at $1.00/hrRunPod RTX 4090 pod at $0.74/hr
Scales to zeroYes, by defaultYesYesYesNo
Largest single GPUServerless up to RTX 5090; H100 class as dedicated machinesUp to B300Up to B300Up to H200Whatever you rent
vLLM versionPinned to 0.8.4 unless you set itYou chooseSet by the worker releaseManaged by Hugging FaceYou choose

Rates come from each provider's published pricing page and docs. Hugging Face bills by the minute.

Where each one fits:

  • Beam is built for teams that want the endpoint defined in Python next to the rest of their code, authenticated by default, with scale-to-zero and a cheap 24 GB serverless GPU. Our serverless GPU list is short, so models that need 80 GB cards run on our dedicated machines.
  • Modal has a wider serverless GPU range, up to B300, and fine control over the vLLM process because you launch it yourself. You write more code, and the official example leaves auth to you. Its $30 a month of free compute on the Starter plan covers light testing.
  • RunPod Serverless is the fastest way to get a vLLM endpoint from a browser: pick the worker, type a model name, and deploy. Its base URL has a longer /v2/<endpoint-id>/openai/v1 path, and tuning happens through environment variables in the console.
  • Hugging Face Inference Endpoints fits teams already on the Hub who want a managed endpoint with per-minute billing. Rates for the same 24 GB class are higher than ours or RunPod's. The Hugging Face Inference Endpoints alternatives guide compares it with other hosts.
  • A GPU VM is cheapest for steady, all-day traffic and gives you full control. You own TLS, auth, restarts, upgrades, and scaling, and you pay while it sits idle.

For how serverless platforms compare on cold starts, see the top serverless GPU providers. Our earlier post on serving vLLM on Beam covers why we built the integration.

FAQ

What does OpenAI-compatible mean for an LLM API?

It means the server accepts the same request format and returns the same response objects as OpenAI's API, on the same routes such as /v1/chat/completions. Clients built for OpenAI, including the official SDKs and frameworks that use them, work after you change the base URL, API key, and model name.

Can I use the OpenAI Python SDK with an open-source model?

Yes. Create the client with base_url set to your server's /v1 URL and api_key set to that server's key, then pass the served model name as model. You can also set OPENAI_BASE_URL and OPENAI_API_KEY as environment variables and leave the code unchanged.

Which open-source models can I serve this way?

Any model vLLM supports, which covers most popular open-weight families on Hugging Face, including Qwen, Llama, Mistral, Gemma, and DeepSeek distills. Gated models need a Hugging Face token passed as a secret. Very new architectures may need a newer vLLM than the default, which you set with vllm_version.

How much does it cost to host your own LLM API?

On Beam's serverless pricing, an RTX 4090 container with 2 CPU cores and 16 GiB of RAM costs about $1.76 an hour while it runs, and nothing when it has scaled to zero. An endpoint busy 8 hours a day on weekdays runs about $310 a month. Past roughly 25% utilization, a dedicated machine is cheaper.

Does a self-hosted vLLM endpoint support streaming and tool calling?

Yes. Streaming works with stream=True as it does with OpenAI. Tool calling needs enable_auto_tool_choice=True and a tool_call_parser that matches the model, such as hermes for Qwen2.5 or mistral for Mistral models.

How do I avoid cold starts on a self-hosted LLM API?

Raise keep_warm_seconds so containers stay up between bursts, or set min_containers=1 so one container always runs. Both bill while the container is idle. Caching the weights on a volume, which our VLLM class does by default, shortens cold starts but doesn't remove the time to load weights onto the GPU.

Host your open-source LLM behind an OpenAI API

Define the model in Python, deploy it, and point your OpenAI client at the URL. The endpoint scales with traffic and shuts off when it's idle. Start on Beam with the free Developer plan.

Eli Mernit
Eli Mernit
Published September 29, 2026
Pay as you gobilled by the millisecond

Start shipping on infra
you won’t outgrow.

Run sandboxes and GPU workloads on your cloud, and scale out to ours when you need to. No infra to manage.