A server that loads open-weight models on demand behind an OpenAI-compatible API, run on GPUs in Switzerland, the EU or your own datacentre so prompts never reach a model vendor. Pilae runs it on your own servers, or in Zurich, Switzerland, and eleven other Pilae Cloud regions.
- Licence
- MIT
- Runs on
- Your own hardware, or any of twelve Pilae regions — six of them in Switzerland and the EU
- Upgrades
- Pinned, tested against your configuration, applied in your window
- Upstream
- ollama.com
Running Ollama in production: what it takes
Prepare the GPU host
Driver and container toolkit installed and checked with nvidia-smi inside a container, then a pinned Ollama build we have run, bound to a private address.
Pin the models
Each model is pulled once, in a change window, and its digest recorded in the runbook. Context length and parameters live in Modelfiles in your repository. Cloud features are off.
Size concurrency and memory
Parallel requests, loaded models, keep-alive and context length are set so the models you serve stay in GPU memory. Probes every 60 seconds check each one answers and is still on the GPU, and alert an engineer when one has fallen back to the CPU.
Back up the models and Modelfiles
Ollama holds no conversations. The models directory, Modelfiles and compose files go offsite daily, encrypted. Once a month we restore them into a scratch environment, check the digests and ask a question.
Upgrade with the driver checked
A new Ollama build can need a newer GPU driver, and falls back to the CPU when it does not get one. The Pilae Agent tests each upgrade on a copy with your models, waits for your approval, applies it in your window and keeps the previous build ready.
What Ollama is, and who runs it
Self-hosted Ollama: open models behind an OpenAI-compatible API
Ollama is a model server. It pulls open models from its library, keeps them on disk, loads them
into GPU memory when a request arrives and answers over an HTTP API, including an
OpenAI-compatible one under /v1. The library publishes models in quantised builds, so a useful
model fits on one card. Open WebUI, LibreChat and
Dify connect to it as a model provider, and so do scripts and internal tools that
already speak the OpenAI API.
Because it loads and unloads models on demand, it suits one GPU machine serving a team, a prototype, or several models on modest hardware. Each loaded model answers a set number of requests at a time and queues the rest, so it is not the engine for hundreds of people on one model at once. For that we run vLLM.
Ollama in production: GPU memory, keep-alive and access control
Ollama fails quietly. If the GPU driver or the container toolkit is wrong, it starts anyway and runs on the CPU at a fraction of the speed, so our probes check every 60 seconds that each model answers and is still on the GPU. At the default keep-alive a model unloads after five minutes idle, and the next person waits while it loads again. Every parallel request adds its own context to GPU memory. We set keep-alive, parallel requests and context length in the compose file in your repository instead of leaving them to defaults. The local API asks for no credentials, so Ollama answers only on your private network and people reach it through an application with sign-on.
Ollama gets a GPU machine of its own, either yours or one of ours in a Pilae Cloud region, where six of the 12 regions are in Switzerland and the EU. We pin the models as well as the server. Each model is pulled in a change window with its digest recorded, and Ollama’s cloud features are switched off. A daily, encrypted copy of the models directory, Modelfiles and compose files goes to an offsite location in your chosen country, and a monthly drill restores it into a scratch environment. Once the models are staged, Ollama answers without internet access, which is how it runs air-gapped. The Pilae Agent tries each upgrade on a copy with your models loaded, then applies it in your window once you approve it and records it in the console.
Ollama licence and model licences
Nothing in the software licence limits how you run Ollama inside your organisation. The conditions that matter come with the models. We price our operation on request. Talk to us about the models you want to run and the GPUs you have.
Ollama system requirements
Before anything is deployed, this is what has to exist. We size it with you in the first session, and we say so when your own hardware is already enough.
- GPU
- NVIDIA, driver 550+Compute capability 5.0 or newer. Cards from 5.0 to 6.2 need driver 570+. Supported AMD cards work through ROCm 7. GPU memory decides the model: one that does not fit is split onto the CPU and slows sharply.
- Container runtime
- NVIDIA Container ToolkitHow Docker hands the GPU to the container. When discovery fails, Ollama falls back to the CPU instead of refusing to start, so the probe checks where each model is loaded.
- Model storage
- Local NVMe, sized to modelsWeights are read from disk every time a model loads. Library models range from under a gigabyte to well over a hundred, and every pinned version is kept.
- API port
- 11434/tcp, private onlyThe local API asks for no credentials, and the port that answers prompts also accepts pull and delete requests. It is never published to the internet.
- Sign-on
- OIDC, in front of OllamaThe Ollama server has no user accounts. Identity lives in the application or gateway in front of it, through Keycloak or your own IdP such as Microsoft Entra ID.
Migrating from OpenAI API to Ollama
In the code, moving off the OpenAI API is mostly a change of base URL. Ollama answers chat completions, embeddings and the stateless Responses API on its own /v1 path, so the OpenAI SDK in your applications stays. The Batch and Files APIs, stateful Responses, tool_choice, logprobs and image URLs do not come across. Neither do the embeddings you already have: vectors from one model mean nothing to another, so each document index is rebuilt. The time goes on the model, not the code. An open model that fits your GPUs will be weaker than the hosted one at some tasks and as good at others, and we find out which on your own prompts before anyone repoints an application.
List the calls
Which applications call the API, with which models, and which features they rely on. Tools, JSON output and streaming carry over. Batches, stateful Responses and logprobs need rework, and we say where.
Choose the model on your prompts
Candidate open models are run against a set of your real prompts, and their answers are graded. The smallest model that passes wins, and its digest goes into the runbook.
Repoint the client
The OpenAI SDK stays and the base URL moves to the private endpoint. Context length is set in a Modelfile in your repository, because the OpenAI API has no field for it.
Re-embed, then switch
Document indexes are rebuilt with the new embedding model. Each application then runs against both for a cycle, and moves across once its owners have compared the answers.
An Ollama server, configured in your repository
# compose.yaml -- acme-llm-01, one GPU machine
services:
ollama:
image: ollama/ollama:${OLLAMA_TAG} # pinned in .env, tested on a copy first
restart: unless-stopped
ports:
- '10.20.4.11:11434:11434' # private address only
volumes:
- /srv/ollama:/root/.ollama # models/manifests and models/blobs
environment:
OLLAMA_NO_CLOUD: '1' # no hosted models, no web search
OLLAMA_NOHISTORY: '1' # no CLI prompt history in the volume
OLLAMA_CONTEXT_LENGTH: '16384'
OLLAMA_NUM_PARALLEL: '4' # memory grows with parallel x context
OLLAMA_MAX_LOADED_MODELS: '2'
OLLAMA_MAX_QUEUE: '128' # beyond this, callers get a 503
OLLAMA_KEEP_ALIVE: '12h' # resident through the working day
OLLAMA_FLASH_ATTENTION: '1'
OLLAMA_KV_CACHE_TYPE: 'q8_0' # about half the cache memory of f16
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
What Pilae is responsible for
A pinned version
A version we have run, not whatever latest resolves to that day.
A runbook
What it depends on, how it fails, what to do about it. In your repository.
A restore drill
Backups restored on a schedule. A backup nobody has restored is a file.
A patch window
Security updates in a window you agreed, with a rollback ready.
Someone watching
Every endpoint probed on the minute. An alert reaches a person, not a dashboard nobody opens.
- Where it runs
- zur1, fra1, fal1, gra1, ams1, hel1, lon1, ash1, hil1, sin1, tok1, syd1, on-premZurich, Frankfurt, Falkenstein, Gravelines, Amsterdam, Helsinki, London, Ashburn, Hillsboro, Singapore, Tokyo, Sydney, Your own hardware
- Who holds the credentials
- You do. Ours are separate, named, logged and revocable with one command. We ask before anything changes outside an agreed window.
- If you leave
- The machine, the data, the compose files and the runbook are already yours. Nothing stops when our access does.
What drives the price of running Ollama
Pricing is on request: a fixed price for onboarding, then a monthly price for Ollama, quoted in writing within five business days. The plans set what every deployment includes; these are the inputs the quote is built from.
- Instance size
- The CPU, memory and, where a model runs, the GPUs the app needs for your users and your data.
- High availability
- One machine with tested restores, or a replicated setup that keeps serving when a node fails.
- Storage and backups
- How much data it holds, how long backups are kept, and point-in-time recovery for its database.
- Plan and support
- Essential, Business or Enterprise: support hours, response times in the contract and how often we review the service with you.
- Region
- Your own hardware, where the infrastructure is already yours, or a Pilae Cloud region, where it is passed through at cost plus a fixed margin.
- Sign-on and integrations
- Single sign-on, directory sync, mail relays and the other systems the app has to reach.
Ollama: common questions
Is Ollama open source?
Where do prompts and models live?
Should we run Ollama or vLLM?
Can people sign in to Ollama?
Will an open model answer as well as the OpenAI API?
Also in ai and llm
Open WebUI
A chat interface over models you host, so prompts and the documents people paste into them never leave your network.
Replaces ChatGPT Team, Microsoft Copilot
vLLM
An OpenAI-compatible inference server for open-weight models, run on GPUs in Switzerland, the EU or your own datacentre, so prompts and answers never reach a model vendor.
Replaces OpenAI API, Azure OpenAI
LibreChat
One chat interface over every model your organisation allows, local or hosted, with the keys, the history and the choice of provider held on your side.
Replaces ChatGPT Team, Poe
Bring us your Ollama. We will tell you what it takes.
Thirty minutes on the deployment you already have, or the one you are about to start.