An OpenAI-compatible inference server for open-weight models, run on GPUs in Switzerland, the EU or your own datacentre, so prompts and answers never reach a model vendor. Pilae runs it on your own servers, or in Zurich, Switzerland, and eleven other Pilae Cloud regions.
- Licence
- Apache-2.0
- Runs on
- Your own hardware, or any of twelve Pilae regions — six of them in Switzerland and the EU
- Upgrades
- Pinned, tested against your configuration, applied in your window
- Upstream
- vllm.ai
Running vLLM in production: what it takes
Size, then pin and deploy
A vLLM image we have run, the model from the local mirror, and the tensor parallel size, context length and memory share in one config file in your repository. Usage statistics reporting to the project is switched off.
Put the API behind the gate
vLLM answers only on the private network, with a key per application. The gate in front passes the /v1 routes your clients need and nothing else.
Watch the queue and the KV cache
Probes check /health every 60 seconds. The metrics endpoint reports KV cache usage, waiting requests, time to first token and preemptions, and an engineer is alerted when the queue grows faster than it drains.
Back up the config and the weights
vLLM keeps no database. The config, compose files, mirrored weights and any fine-tuned adapters go offsite daily, encrypted, and the weights only change when the model does. Once a month we rebuild the server from them in a scratch environment and send it a test prompt.
Upgrade against your config
vLLM ships a minor release about every two weeks, and a deprecated flag can be gone two minor releases later. The Pilae Agent tests each upgrade against your config and model on a copy, waits for your approval, applies it in your window and keeps the previous image ready.
What vLLM is, and who runs it
Self-hosted vLLM for serving open-weight models
vLLM is an inference and serving engine for large language models. It started at UC Berkeley’s Sky Computing Lab and is now a hosted project of the PyTorch Foundation. PagedAttention keeps the KV cache in small blocks instead of one reservation per request, and continuous batching adds requests to the running batch as others finish. Tensor parallelism splits one model across the GPUs in a machine. Applications reach it through an OpenAI-compatible API.
It is built for many people or applications calling the same model at once. For one GPU machine serving a team, or several smaller models on modest hardware, Ollama is simpler to run. In a private AI deployment, vLLM usually sits behind a chat front end such as Open WebUI or LibreChat. Your own applications can also call it directly in place of the OpenAI API, and Langfuse beside it records each prompt and answer, so a change of model is judged on your own traffic.
vLLM in production: model weights, API access and GPU memory
vLLM runs on GPU machines dedicated to it, yours or ours in a
Pilae Cloud region: Zurich, Frankfurt, Falkenstein, Gravelines, Amsterdam or
Helsinki inside Switzerland and the EU, or one of six further afield. The weights are mirrored
locally with the revision recorded, so a restart never waits on a download. The server answers
only on your private network. Its API key covers /v1 and a few other
prefixes but leaves many routes open, including /invocations, which runs inference, so the gate
in front passes only the routes your clients call.
The failures that matter are about memory. vLLM refuses to start when the context length asked for does not fit in the KV cache, so the context length is set from the sizing before the first start. Under a burst it queues or preempts requests instead of failing them, so an overload shows up as slow answers rather than errors. Probes check the server every 60 seconds, and a queue that grows faster than it drains alerts an engineer. The config, compose files, weights and adapters are backed up daily, encrypted, to an offsite location in your chosen country, and once a month a scratch server is rebuilt from them. With a release every couple of weeks, the Pilae Agent runs each one against your config and model on a copy before you approve it for your window, and records the change in the console.
vLLM licence and model licences
vLLM is Apache-2.0. There is no paid edition and no enterprise directory, so nothing is kept back from the open-source code. Two other licences sit next to it. The model weights carry their own: some are Apache-2.0 or MIT, while others, such as Meta’s Llama models, come with their own licence and acceptable use policy. We check the one you choose before it is deployed. The CUDA libraries in the image come from NVIDIA’s own base image and are under NVIDIA’s licence terms. Pricing for our operation is on request. Talk to us about the model you want to serve.
vLLM system requirements
Before anything is deployed, this is what has to exist. We size it with you in the first session, and we say so when your own hardware is already enough.
- GPU
- NVIDIA, compute capability 7.5+With memory for the weights plus the KV cache at the context length and concurrency you need. AMD and Intel GPUs are supported through their own builds.
- Driver and runtime
- NVIDIA driver + Container ToolkitThe official image carries its own CUDA libraries, and the host driver has to be new enough for them. On datacentre GPUs the compatibility libraries in the image can bridge an older driver; otherwise, moving to a newer image can mean upgrading the driver first.
- Model storage
- Local NVMe, weights mirroredWeights are read from disk at every start, and upstream documents the default loader as most efficient on local storage. They are copied once from Hugging Face with the revision pinned.
- Shared memory
- ipc: hostvLLM uses PyTorch, which passes data between processes through shared memory, above all for tensor parallelism. Upstream says to give the container the host IPC namespace or a --shm-size, and we use the host namespace.
- Access
- API key + /v1 allowlistThe API key covers /v1 and a few other prefixes. Routes such as /invocations, /tokenize and /pause answer without it, so the gate passes only the routes clients use.
Migrating from OpenAI API to vLLM
Code written against the OpenAI API usually moves by changing three settings: the base URL, the key and the model name. vLLM serves the same Chat Completions, Completions, Embeddings and Responses endpoints, and tool calls work once the right parser is set for the model. What does not come across is the model. OpenAI's hosted models cannot be downloaded, so the work is choosing an open-weight model and testing it on your own prompts, which were tuned for another model and will need changes. Stored files, vector stores, fine-tuned models and Azure's content filters stay behind. Embeddings come from a separate embedding model, and vectors from two models cannot be mixed, so every vector index is rebuilt. Changing the code is the short part; that evaluation is the long one.
Read the usage before choosing a model
Which endpoints, models and features your applications call, with token volumes and per-minute request peaks from the OpenAI or Azure usage records. Concurrency and context length decide the GPUs, not the number of users.
Shortlist and test open-weight models
Two or three candidates whose licences fit, run against a set of your real prompts in a sandbox and scored beside the answers you get today. Prompts are adjusted here, not after the switch.
Mirror the weights and size the server
The chosen weights are downloaded once, with the revision and checksums recorded. Gated models need access granted on Hugging Face to one of your users first, and some authors approve requests by hand. Tensor parallel size, context length and memory share are set from the measured load.
Change the base URL, keep a fallback
Applications switch by configuration, one at a time, with the old key valid for one cycle. Anything that stores embeddings is re-indexed with the new embedding model before it switches.
A vLLM server, configured from one file in your repository
# /etc/vllm/qwen3-32b.yaml -- acme-inference-01, zur1, 2 x 80 GB GPUs # Started as: vllm serve --config /etc/vllm/qwen3-32b.yaml # VLLM_API_KEY comes from the secret store; VLLM_NO_USAGE_STATS=1 is set. model: /models/Qwen3-32B # local mirror, never pulled at start served-model-name: qwen3-32b # the name clients send as "model" tensor-parallel-size: 2 # one model split across both GPUs gpu-memory-utilization: 0.90 # vLLM's share; what weights leave is KV cache max-model-len: 32768 # longest prompt plus answer accepted max-num-seqs: 64 # sequences scheduled together enable-auto-tool-choice: true tool-call-parser: hermes reasoning-parser: qwen3 host: 10.20.0.14 # private network address only port: 8000 uvicorn-log-level: warning
What Pilae is responsible for
A pinned version
A version we have run, not whatever latest resolves to that day.
A runbook
What it depends on, how it fails, what to do about it. In your repository.
A restore drill
Backups restored on a schedule. A backup nobody has restored is a file.
A patch window
Security updates in a window you agreed, with a rollback ready.
Someone watching
Every endpoint probed on the minute. An alert reaches a person, not a dashboard nobody opens.
- Where it runs
- zur1, fra1, fal1, gra1, ams1, hel1, lon1, ash1, hil1, sin1, tok1, syd1, on-premZurich, Frankfurt, Falkenstein, Gravelines, Amsterdam, Helsinki, London, Ashburn, Hillsboro, Singapore, Tokyo, Sydney, Your own hardware
- Who holds the credentials
- You do. Ours are separate, named, logged and revocable with one command. We ask before anything changes outside an agreed window.
- If you leave
- The machine, the data, the compose files and the runbook are already yours. Nothing stops when our access does.
What drives the price of running vLLM
Pricing is on request: a fixed price for onboarding, then a monthly price for vLLM, quoted in writing within five business days. The plans set what every deployment includes; these are the inputs the quote is built from.
- Instance size
- The CPU, memory and, where a model runs, the GPUs the app needs for your users and your data.
- High availability
- One machine with tested restores, or a replicated setup that keeps serving when a node fails.
- Storage and backups
- How much data it holds, how long backups are kept, and point-in-time recovery for its database.
- Plan and support
- Essential, Business or Enterprise: support hours, response times in the contract and how often we review the service with you.
- Region
- Your own hardware, where the infrastructure is already yours, or a Pilae Cloud region, where it is passed through at cost plus a fixed margin.
- Sign-on and integrations
- Single sign-on, directory sync, mail relays and the other systems the app has to reach.
vLLM: common questions
Is vLLM open source?
Where do prompts and answers go?
Can our applications keep using the OpenAI SDK?
Should we run vLLM or Ollama?
What hardware does vLLM need?
Also in ai and llm
Open WebUI
A chat interface over models you host, so prompts and the documents people paste into them never leave your network.
Replaces ChatGPT Team, Microsoft Copilot
Ollama
A server that loads open-weight models on demand behind an OpenAI-compatible API, run on GPUs in Switzerland, the EU or your own datacentre so prompts never reach a model vendor.
Replaces OpenAI API, Azure OpenAI
LibreChat
One chat interface over every model your organisation allows, local or hosted, with the keys, the history and the choice of provider held on your side.
Replaces ChatGPT Team, Poe
Bring us your vLLM. We will tell you what it takes.
Thirty minutes on the deployment you already have, or the one you are about to start.