Managed vLLM hosting

AI and LLMOn-prem or sovereign site

An OpenAI-compatible inference server for open-weight models, run on GPUs in Switzerland, the EU or your own datacentre, so prompts and answers never reach a model vendor. Pilae runs it on your own servers, or in Zurich, Switzerland, and eleven other Pilae Cloud regions.

Talk to us about vLLM

Licence
Apache-2.0
Runs on
Your own hardware, or any of twelve Pilae regions — six of them in Switzerland and the EU
Upgrades
Pinned, tested against your configuration, applied in your window
Upstream
vllm.ai

Running vLLM in production: what it takes

  1. Size, then pin and deploy

    A vLLM image we have run, the model from the local mirror, and the tensor parallel size, context length and memory share in one config file in your repository. Usage statistics reporting to the project is switched off.

  2. Put the API behind the gate

    vLLM answers only on the private network, with a key per application. The gate in front passes the /v1 routes your clients need and nothing else.

  3. Watch the queue and the KV cache

    Probes check /health every 60 seconds. The metrics endpoint reports KV cache usage, waiting requests, time to first token and preemptions, and an engineer is alerted when the queue grows faster than it drains.

  4. Back up the config and the weights

    vLLM keeps no database. The config, compose files, mirrored weights and any fine-tuned adapters go offsite daily, encrypted, and the weights only change when the model does. Once a month we rebuild the server from them in a scratch environment and send it a test prompt.

  5. Upgrade against your config

    vLLM ships a minor release about every two weeks, and a deprecated flag can be gone two minor releases later. The Pilae Agent tests each upgrade against your config and model on a copy, waits for your approval, applies it in your window and keeps the previous image ready.

What vLLM is, and who runs it

Self-hosted vLLM for serving open-weight models

vLLM is an inference and serving engine for large language models. It started at UC Berkeley’s Sky Computing Lab and is now a hosted project of the PyTorch Foundation. PagedAttention keeps the KV cache in small blocks instead of one reservation per request, and continuous batching adds requests to the running batch as others finish. Tensor parallelism splits one model across the GPUs in a machine. Applications reach it through an OpenAI-compatible API.

It is built for many people or applications calling the same model at once. For one GPU machine serving a team, or several smaller models on modest hardware, Ollama is simpler to run. In a private AI deployment, vLLM usually sits behind a chat front end such as Open WebUI or LibreChat. Your own applications can also call it directly in place of the OpenAI API, and Langfuse beside it records each prompt and answer, so a change of model is judged on your own traffic.

vLLM in production: model weights, API access and GPU memory

vLLM runs on GPU machines dedicated to it, yours or ours in a Pilae Cloud region: Zurich, Frankfurt, Falkenstein, Gravelines, Amsterdam or Helsinki inside Switzerland and the EU, or one of six further afield. The weights are mirrored locally with the revision recorded, so a restart never waits on a download. The server answers only on your private network. Its API key covers /v1 and a few other prefixes but leaves many routes open, including /invocations, which runs inference, so the gate in front passes only the routes your clients call.

The failures that matter are about memory. vLLM refuses to start when the context length asked for does not fit in the KV cache, so the context length is set from the sizing before the first start. Under a burst it queues or preempts requests instead of failing them, so an overload shows up as slow answers rather than errors. Probes check the server every 60 seconds, and a queue that grows faster than it drains alerts an engineer. The config, compose files, weights and adapters are backed up daily, encrypted, to an offsite location in your chosen country, and once a month a scratch server is rebuilt from them. With a release every couple of weeks, the Pilae Agent runs each one against your config and model on a copy before you approve it for your window, and records the change in the console.

vLLM licence and model licences

vLLM is Apache-2.0. There is no paid edition and no enterprise directory, so nothing is kept back from the open-source code. Two other licences sit next to it. The model weights carry their own: some are Apache-2.0 or MIT, while others, such as Meta’s Llama models, come with their own licence and acceptable use policy. We check the one you choose before it is deployed. The CUDA libraries in the image come from NVIDIA’s own base image and are under NVIDIA’s licence terms. Pricing for our operation is on request. Talk to us about the model you want to serve.

vLLM system requirements

Before anything is deployed, this is what has to exist. We size it with you in the first session, and we say so when your own hardware is already enough.

GPU
NVIDIA, compute capability 7.5+With memory for the weights plus the KV cache at the context length and concurrency you need. AMD and Intel GPUs are supported through their own builds.
Driver and runtime
NVIDIA driver + Container ToolkitThe official image carries its own CUDA libraries, and the host driver has to be new enough for them. On datacentre GPUs the compatibility libraries in the image can bridge an older driver; otherwise, moving to a newer image can mean upgrading the driver first.
Model storage
Local NVMe, weights mirroredWeights are read from disk at every start, and upstream documents the default loader as most efficient on local storage. They are copied once from Hugging Face with the revision pinned.
Shared memory
ipc: hostvLLM uses PyTorch, which passes data between processes through shared memory, above all for tensor parallelism. Upstream says to give the container the host IPC namespace or a --shm-size, and we use the host namespace.
Access
API key + /v1 allowlistThe API key covers /v1 and a few other prefixes. Routes such as /invocations, /tokenize and /pause answer without it, so the gate passes only the routes clients use.

Migrating from OpenAI API to vLLM

Code written against the OpenAI API usually moves by changing three settings: the base URL, the key and the model name. vLLM serves the same Chat Completions, Completions, Embeddings and Responses endpoints, and tool calls work once the right parser is set for the model. What does not come across is the model. OpenAI's hosted models cannot be downloaded, so the work is choosing an open-weight model and testing it on your own prompts, which were tuned for another model and will need changes. Stored files, vector stores, fine-tuned models and Azure's content filters stay behind. Embeddings come from a separate embedding model, and vectors from two models cannot be mixed, so every vector index is rebuilt. Changing the code is the short part; that evaluation is the long one.

  1. Read the usage before choosing a model

    Which endpoints, models and features your applications call, with token volumes and per-minute request peaks from the OpenAI or Azure usage records. Concurrency and context length decide the GPUs, not the number of users.

  2. Shortlist and test open-weight models

    Two or three candidates whose licences fit, run against a set of your real prompts in a sandbox and scored beside the answers you get today. Prompts are adjusted here, not after the switch.

  3. Mirror the weights and size the server

    The chosen weights are downloaded once, with the revision and checksums recorded. Gated models need access granted on Hugging Face to one of your users first, and some authors approve requests by hand. Tensor parallel size, context length and memory share are set from the measured load.

  4. Change the base URL, keep a fallback

    Applications switch by configuration, one at a time, with the old key valid for one cycle. Anything that stores embeddings is re-indexed with the new embedding model before it switches.

A vLLM server, configured from one file in your repository

# /etc/vllm/qwen3-32b.yaml -- acme-inference-01, zur1, 2 x 80 GB GPUs
# Started as: vllm serve --config /etc/vllm/qwen3-32b.yaml
# VLLM_API_KEY comes from the secret store; VLLM_NO_USAGE_STATS=1 is set.

model: /models/Qwen3-32B          # local mirror, never pulled at start
served-model-name: qwen3-32b      # the name clients send as "model"
tensor-parallel-size: 2           # one model split across both GPUs
gpu-memory-utilization: 0.90      # vLLM's share; what weights leave is KV cache
max-model-len: 32768              # longest prompt plus answer accepted
max-num-seqs: 64                  # sequences scheduled together
enable-auto-tool-choice: true
tool-call-parser: hermes
reasoning-parser: qwen3
host: 10.20.0.14                  # private network address only
port: 8000
uvicorn-log-level: warning
An example config for acme, read at start by vllm serve --config. The values that decide whether a model fits on the GPUs sit in one reviewed file, so a rebuilt server comes back with the same memory budget, context length and model name the applications call.

What Pilae is responsible for

A pinned version

A version we have run, not whatever latest resolves to that day.

A runbook

What it depends on, how it fails, what to do about it. In your repository.

A restore drill

Backups restored on a schedule. A backup nobody has restored is a file.

A patch window

Security updates in a window you agreed, with a rollback ready.

Someone watching

Every endpoint probed on the minute. An alert reaches a person, not a dashboard nobody opens.

Where it runs
zur1, fra1, fal1, gra1, ams1, hel1, lon1, ash1, hil1, sin1, tok1, syd1, on-premZurich, Frankfurt, Falkenstein, Gravelines, Amsterdam, Helsinki, London, Ashburn, Hillsboro, Singapore, Tokyo, Sydney, Your own hardware
Who holds the credentials
You do. Ours are separate, named, logged and revocable with one command. We ask before anything changes outside an agreed window.
If you leave
The machine, the data, the compose files and the runbook are already yours. Nothing stops when our access does.

What drives the price of running vLLM

Pricing is on request: a fixed price for onboarding, then a monthly price for vLLM, quoted in writing within five business days. The plans set what every deployment includes; these are the inputs the quote is built from.

Instance size
The CPU, memory and, where a model runs, the GPUs the app needs for your users and your data.
High availability
One machine with tested restores, or a replicated setup that keeps serving when a node fails.
Storage and backups
How much data it holds, how long backups are kept, and point-in-time recovery for its database.
Plan and support
Essential, Business or Enterprise: support hours, response times in the contract and how often we review the service with you.
Region
Your own hardware, where the infrastructure is already yours, or a Pilae Cloud region, where it is passed through at cost plus a fixed margin.
Sign-on and integrations
Single sign-on, directory sync, mail relays and the other systems the app has to reach.

vLLM: common questions

Is vLLM open source?

Yes. vLLM is Apache-2.0 and hosted by the PyTorch Foundation, with no paid edition and no features held back for one. The models you serve with it are a separate question: each set of weights carries its own licence, and some come with use restrictions.

Where do prompts and answers go?

To the GPU machine and no further: your own hardware, or a dedicated server in the Pilae region you picked, in ISO 27001-certified datacentres. vLLM keeps no conversation store by default, and we leave its Responses store and request logging off. Left at its defaults, it also sends anonymous usage statistics about hardware and configuration to the project; we switch that off.

Can our applications keep using the OpenAI SDK?

Yes, for Chat Completions, Completions, Embeddings and Responses. The OpenAI client libraries call vLLM once the base URL, key and model name are changed. Tool calling needs a parser set for the model family, and that sits in the server's config, not in your code. OpenAI's stored files, vector stores and fine-tuning have no equivalent in vLLM.

Should we run vLLM or Ollama?

vLLM for many concurrent requests to one model, Ollama for one GPU machine serving a team or several smaller models. Ollama is simpler to run: it loads models on demand and runs quantised models on modest hardware. vLLM reserves most of the GPU memory at start and holds one base model per server, so serving several models means several servers.

What hardware does vLLM need?

GPUs with enough memory for the weights plus the KV cache at your context length and concurrency. NVIDIA with compute capability 7.5 or higher is the main path, and AMD and Intel GPUs are supported through their own builds. We size it from your traffic before anything is installed, and say so when the model you want does not fit.

Also in ai and llm

Back to apps

Bring us your vLLM. We will tell you what it takes.

Thirty minutes on the deployment you already have, or the one you are about to start.