Using OpenAI-Compatible API for Open-Source LLMs
Eli Mernit
To host an OpenAI-compatible API for an open-source LLM, run vLLM's OpenAI server on a GPU behind an authenticated HTTPS URL, then set base_url and api_key in the OpenAI SDK. On Beam, that's one VLLM object in our SDK and beam deploy. You get a /v1 endpoint that autoscales and shuts off when traffic stops.
Deploy it:
The deploy prints a URL like https://qwen25-7b-1a2b3c4-v1.app.beam.cloud. Add /v1 and use it like the OpenAI API:
Qwen2.5-7B-Instruct is Apache-2.0 licensed, so it downloads without a Hugging Face token. The weights take about 15 GB in 16-bit precision, which leaves room on the RTX 4090's 24 GB for the KV cache.
What an OpenAI-compatible LLM API involves
An OpenAI-compatible API is a server that accepts the same requests as OpenAI's API (/v1/chat/completions, /v1/completions, /v1/models) and returns the same response objects. Any client built for OpenAI then works against it once you change the base URL and key. Most of the effort goes into running that server where it stays up, stays private, and doesn't bill you while idle.
To host one yourself, you need:
- An inference server. vLLM is the common choice. It ships an OpenAI-compatible server and batches concurrent requests on the GPU. SGLang and Hugging Face TGI are alternatives.
- A GPU sized to the model. The weights, plus KV cache for your context length, have to fit in GPU memory.
- A public URL with TLS and auth. vLLM on its own is a bare HTTP server on a port.
- Scaling. More replicas when traffic spikes, and ideally none when nobody is calling it.
- Cost control. A GPU that sits idle overnight is the biggest line item for low-traffic APIs.
Most guides stop at vllm serve on a single machine, or jump to Kubernetes with KEDA for autoscaling. A serverless GPU platform covers the last three items for you.
How to host an OpenAI-compatible API on a serverless GPU
Our SDK's VLLM class wraps vLLM's OpenAI server as a deployable app. You pass the hardware (gpu, cpu, memory), scaling settings, and a VLLMArgs object whose fields match vllm serve flags. We build the container, mount a volume to cache model weights, put the server behind an authenticated URL, and scale it with traffic.
Install the CLI and deploy
Create an account at platform.beam.cloud, copy an API key, and install the client:
Save the hero snippet as app.py and run beam deploy app.py:qwen. The first container downloads the model into a volume named vllm_cache, so later cold starts read the weights from the volume instead of pulling them from Hugging Face again.
Point existing OpenAI code at the endpoint
The OpenAI Python SDK reads OPENAI_BASE_URL and OPENAI_API_KEY from the environment. If your code already calls OpenAI() with no arguments, you can switch it to your own model without editing it:
The one change you can't avoid is the model string. It has to match served_model_name, and client.models.list() returns it if you're unsure. The server also exposes /v1/completions for base models and /v1/embeddings for embedding models.
Endpoints are authenticated by default: requests need your Beam token as a Bearer token, which is exactly what the OpenAI SDK sends as api_key. Set authorized=False only if you put your own auth in front of it.
Let vLLM batch requests: raise concurrent_requests
vLLM's throughput comes from continuous batching. It runs many sequences through the GPU together, so 32 concurrent requests take far less than 32 times as long as one. The VLLM class defaults to concurrent_requests=1, which sends one request at a time to each container and leaves that batching unused.
Set concurrent_requests to the number of requests you want each GPU to handle at once. For a 7B model on a 24 GB card with 8K context, 16 to 64 is a reasonable starting range. Raise it until time to first token gets worse than you're willing to accept. Keep tasks_per_container in the autoscaler equal to it, so we only add a container when the current ones are full.
Scale out and scale to zero
The autoscaler adds containers as the request queue grows, up to max_containers. Its default is one container, so set it explicitly. When requests stop, each container stays up for keep_warm_seconds (60 by default for VLLM) and then shuts down. With no containers running, you pay nothing.
The trade-off is the cold start. The next request after a quiet period waits for a GPU container to boot and load the weights. We don't bill for machine startup or image pulls, but model loading does count as billed time. If users can't wait for that, keep one container running:
That container bills around the clock, so check the cost section below before you set it.
Serve gated or private models
Llama and some Mistral checkpoints on Hugging Face are gated. Store a Hugging Face token as a Beam secret and pass it to the app:
The same pattern works for your own fine-tuned weights pushed to a private Hugging Face repo. The Llama 3 fine-tuning guide walks through training one and serving it with vLLM.
Turn on tool calling and streaming
OpenAI-style function calling needs two vLLM flags: enable_auto_tool_choice=True and a tool_call_parser that matches the model's format. Qwen2.5 uses hermes, Mistral models use mistral, and Llama 3.1 uses llama3_json. With those set, tools=[...] in chat.completions.create returns tool_calls the same way OpenAI does.
Streaming works as it does with OpenAI: pass stream=True and iterate over the chunks. Our vLLM chat example in the beam-cloud/examples repo streams responses this way.
Pin the vLLM version
The VLLM class installs vLLM 0.8.4 by default. That release supports Qwen2.5, Llama 3.x, and Mistral, but models released after it may need a newer vLLM. Set vllm_version="..." to change it, and match huggingface_hub_version if the new release needs it. Newer vLLM releases add flags that VLLMArgs doesn't list; the SDK docs say you may need to subclass VLLMArgs to pass them.
Pinning also protects you in the other direction. A deploy that worked last month keeps working, because the image doesn't pick up a new vLLM release on its own.
Which GPU to pick for an open-source LLM
Size the GPU by the model's weights plus KV cache. In 16-bit precision, weights take about 2 bytes per parameter; 4-bit quantization (AWQ or GPTQ) cuts that to roughly 0.5 to 0.6 bytes. Leave several GB on top for the KV cache, which grows with context length and concurrent requests.
| Model size | Weights in 16-bit | Weights in 4-bit | GPU that fits |
|---|---|---|---|
| 7B to 8B | about 15 to 16 GB | about 5 GB | RTX 4090 or A10G (24 GB) |
| 14B | about 28 to 30 GB | about 9 GB | 24 GB card with 4-bit, or RTX 5090 (32 GB) with short context in 16-bit |
| 32B | about 64 GB | about 18 to 20 GB | 24 GB card with 4-bit and short context, or an 80 GB card |
| 70B | about 140 GB | about 40 GB | Two or more 80 GB GPUs |
Beam offers four serverless GPU types: T4 (16 GB), A10G (24 GB), RTX 4090 (24 GB), and RTX 5090 (32 GB). H100, H200, A100 80GB, and L40S are available as dedicated on-demand machines. Multi-GPU containers (gpu_count above one) need to be enabled on your account, which matters for 70B-class models. If you run a 4-bit checkpoint, set quantization="awq" (or the matching method) in VLLMArgs. The Qwen 2.5 overview covers which sizes of that family are worth serving.
What it costs to host your own LLM API
On Beam's serverless pricing, the hero config (RTX 4090, 2 CPU cores, 16 GiB of RAM) costs about $0.00049 per second, or $1.76 per hour while a container runs. That breaks down as $0.000192/s for the GPU, $0.000105 per core per second for CPU attached to a GPU, and $0.0000055 per GiB per second for RAM.
What you pay depends on how often a container is running:
| Traffic pattern | Container time | Approximate cost |
|---|---|---|
| A burst of requests, then quiet | Request time plus 300 s keep-warm | About $0.15 of idle time per burst |
| Office hours, 8 h a day, 22 days | About 176 h a month | About $310 a month |
| Always on with min_containers=1 | About 730 h a month | About $1,290 a month |
Always-on is where serverless stops making sense. We also rent a dedicated on-demand RTX 4090 machine from $0.44 an hour, about $320 for a full month. Serverless is cheaper while the endpoint runs less than about a quarter of the time. Past that, a dedicated machine wins.
Hosted per-token APIs are the other comparison. For a popular model at low or spiky volume, paying per token is often cheaper than any GPU you rent yourself. Self-hosting pays off when you need a fine-tuned or niche model, want requests to stay in infrastructure you control, or have enough steady traffic to keep a GPU busy. The Modal pricing breakdown goes deeper on where per-second GPU billing gets expensive.
Choosing where to host an OpenAI-compatible endpoint
The main options are a serverless GPU platform that runs vLLM for you, a managed endpoint service, or vLLM on a GPU machine you manage. They differ in setup effort, idle cost, and who handles auth and scaling.
| Beam VLLM | Modal | RunPod Serverless vLLM | Hugging Face Inference Endpoints | vLLM on a rented GPU VM | |
|---|---|---|---|---|---|
| Setup | One Python object, beam deploy | Python class that starts vllm serve as a subprocess | Pick the vLLM worker in the console, enter a model name | Pick a model and instance in the console | Install vLLM, run vllm serve, add TLS and a proxy |
| Auth by default | Beam token required | The official example is unauthenticated | RunPod API key | Configurable per endpoint | None until you add it |
| 24 GB GPU, listed rate | RTX 4090 at $0.69/hr plus CPU and RAM | L4 at $0.80/hr or A10 at $1.10/hr plus CPU and RAM | L4, A5000, or 3090 at $0.69/hr, RTX 4090 at $1.10/hr | L4 at $0.80/hr, A10G at $1.00/hr | RunPod RTX 4090 pod at $0.74/hr |
| Scales to zero | Yes, by default | Yes | Yes | Yes | No |
| Largest single GPU | Serverless up to RTX 5090; H100 class as dedicated machines | Up to B300 | Up to B300 | Up to H200 | Whatever you rent |
| vLLM version | Pinned to 0.8.4 unless you set it | You choose | Set by the worker release | Managed by Hugging Face | You choose |
Rates come from each provider's published pricing page and docs. Hugging Face bills by the minute.
Where each one fits:
- Beam is built for teams that want the endpoint defined in Python next to the rest of their code, authenticated by default, with scale-to-zero and a cheap 24 GB serverless GPU. Our serverless GPU list is short, so models that need 80 GB cards run on our dedicated machines.
- Modal has a wider serverless GPU range, up to B300, and fine control over the vLLM process because you launch it yourself. You write more code, and the official example leaves auth to you. Its $30 a month of free compute on the Starter plan covers light testing.
- RunPod Serverless is the fastest way to get a vLLM endpoint from a browser: pick the worker, type a model name, and deploy. Its base URL has a longer
/v2/<endpoint-id>/openai/v1path, and tuning happens through environment variables in the console. - Hugging Face Inference Endpoints fits teams already on the Hub who want a managed endpoint with per-minute billing. Rates for the same 24 GB class are higher than ours or RunPod's. The Hugging Face Inference Endpoints alternatives guide compares it with other hosts.
- A GPU VM is cheapest for steady, all-day traffic and gives you full control. You own TLS, auth, restarts, upgrades, and scaling, and you pay while it sits idle.
For how serverless platforms compare on cold starts, see the top serverless GPU providers. Our earlier post on serving vLLM on Beam covers why we built the integration.
FAQ
What does OpenAI-compatible mean for an LLM API?
It means the server accepts the same request format and returns the same response objects as OpenAI's API, on the same routes such as /v1/chat/completions. Clients built for OpenAI, including the official SDKs and frameworks that use them, work after you change the base URL, API key, and model name.
Can I use the OpenAI Python SDK with an open-source model?
Yes. Create the client with base_url set to your server's /v1 URL and api_key set to that server's key, then pass the served model name as model. You can also set OPENAI_BASE_URL and OPENAI_API_KEY as environment variables and leave the code unchanged.
Which open-source models can I serve this way?
Any model vLLM supports, which covers most popular open-weight families on Hugging Face, including Qwen, Llama, Mistral, Gemma, and DeepSeek distills. Gated models need a Hugging Face token passed as a secret. Very new architectures may need a newer vLLM than the default, which you set with vllm_version.
How much does it cost to host your own LLM API?
On Beam's serverless pricing, an RTX 4090 container with 2 CPU cores and 16 GiB of RAM costs about $1.76 an hour while it runs, and nothing when it has scaled to zero. An endpoint busy 8 hours a day on weekdays runs about $310 a month. Past roughly 25% utilization, a dedicated machine is cheaper.
Does a self-hosted vLLM endpoint support streaming and tool calling?
Yes. Streaming works with stream=True as it does with OpenAI. Tool calling needs enable_auto_tool_choice=True and a tool_call_parser that matches the model, such as hermes for Qwen2.5 or mistral for Mistral models.
How do I avoid cold starts on a self-hosted LLM API?
Raise keep_warm_seconds so containers stay up between bursts, or set min_containers=1 so one container always runs. Both bill while the container is idle. Caching the weights on a volume, which our VLLM class does by default, shortens cold starts but doesn't remove the time to load weights onto the GPU.
Host your open-source LLM behind an OpenAI API
Define the model in Python, deploy it, and point your OpenAI client at the URL. The endpoint scales with traffic and shuts off when it's idle. Start on Beam with the free Developer plan.


