Managed Ollama hosting

AI and LLMOn-prem or sovereign site

A server that loads open-weight models on demand behind an OpenAI-compatible API, run on GPUs in Switzerland, the EU or your own datacentre so prompts never reach a model vendor. Pilae runs it on your own servers, or in Zurich, Switzerland, and eleven other Pilae Cloud regions.

Talk to us about Ollama

Licence
MIT
Runs on
Your own hardware, or any of twelve Pilae regions — six of them in Switzerland and the EU
Upgrades
Pinned, tested against your configuration, applied in your window
Upstream
ollama.com

Running Ollama in production: what it takes

  1. Prepare the GPU host

    Driver and container toolkit installed and checked with nvidia-smi inside a container, then a pinned Ollama build we have run, bound to a private address.

  2. Pin the models

    Each model is pulled once, in a change window, and its digest recorded in the runbook. Context length and parameters live in Modelfiles in your repository. Cloud features are off.

  3. Size concurrency and memory

    Parallel requests, loaded models, keep-alive and context length are set so the models you serve stay in GPU memory. Probes every 60 seconds check each one answers and is still on the GPU, and alert an engineer when one has fallen back to the CPU.

  4. Back up the models and Modelfiles

    Ollama holds no conversations. The models directory, Modelfiles and compose files go offsite daily, encrypted. Once a month we restore them into a scratch environment, check the digests and ask a question.

  5. Upgrade with the driver checked

    A new Ollama build can need a newer GPU driver, and falls back to the CPU when it does not get one. The Pilae Agent tests each upgrade on a copy with your models, waits for your approval, applies it in your window and keeps the previous build ready.

What Ollama is, and who runs it

Self-hosted Ollama: open models behind an OpenAI-compatible API

Ollama is a model server. It pulls open models from its library, keeps them on disk, loads them into GPU memory when a request arrives and answers over an HTTP API, including an OpenAI-compatible one under /v1. The library publishes models in quantised builds, so a useful model fits on one card. Open WebUI, LibreChat and Dify connect to it as a model provider, and so do scripts and internal tools that already speak the OpenAI API.

Because it loads and unloads models on demand, it suits one GPU machine serving a team, a prototype, or several models on modest hardware. Each loaded model answers a set number of requests at a time and queues the rest, so it is not the engine for hundreds of people on one model at once. For that we run vLLM.

Ollama in production: GPU memory, keep-alive and access control

Ollama fails quietly. If the GPU driver or the container toolkit is wrong, it starts anyway and runs on the CPU at a fraction of the speed, so our probes check every 60 seconds that each model answers and is still on the GPU. At the default keep-alive a model unloads after five minutes idle, and the next person waits while it loads again. Every parallel request adds its own context to GPU memory. We set keep-alive, parallel requests and context length in the compose file in your repository instead of leaving them to defaults. The local API asks for no credentials, so Ollama answers only on your private network and people reach it through an application with sign-on.

Ollama gets a GPU machine of its own, either yours or one of ours in a Pilae Cloud region, where six of the 12 regions are in Switzerland and the EU. We pin the models as well as the server. Each model is pulled in a change window with its digest recorded, and Ollama’s cloud features are switched off. A daily, encrypted copy of the models directory, Modelfiles and compose files goes to an offsite location in your chosen country, and a monthly drill restores it into a scratch environment. Once the models are staged, Ollama answers without internet access, which is how it runs air-gapped. The Pilae Agent tries each upgrade on a copy with your models loaded, then applies it in your window once you approve it and records it in the console.

Ollama licence and model licences

Nothing in the software licence limits how you run Ollama inside your organisation. The conditions that matter come with the models. We price our operation on request. Talk to us about the models you want to run and the GPUs you have.

Ollama is MIT-licensed and there is no paid edition of the server: every feature of the software is in the open-source build. Ollama's paid plans buy usage of its hosted cloud models. We switch those features off with OLLAMA_NO_CLOUD, so no request can leave your network that way. The models are licensed separately from the software, each by its publisher. Llama 3.1 is under the Llama 3.1 Community License, with attribution and acceptable-use terms, while Mistral 7B and Qwen3 are Apache-2.0. We record the licence of every model we pin and flag any whose terms restrict how you may use it.

Ollama system requirements

Before anything is deployed, this is what has to exist. We size it with you in the first session, and we say so when your own hardware is already enough.

GPU
NVIDIA, driver 550+Compute capability 5.0 or newer. Cards from 5.0 to 6.2 need driver 570+. Supported AMD cards work through ROCm 7. GPU memory decides the model: one that does not fit is split onto the CPU and slows sharply.
Container runtime
NVIDIA Container ToolkitHow Docker hands the GPU to the container. When discovery fails, Ollama falls back to the CPU instead of refusing to start, so the probe checks where each model is loaded.
Model storage
Local NVMe, sized to modelsWeights are read from disk every time a model loads. Library models range from under a gigabyte to well over a hundred, and every pinned version is kept.
API port
11434/tcp, private onlyThe local API asks for no credentials, and the port that answers prompts also accepts pull and delete requests. It is never published to the internet.
Sign-on
OIDC, in front of OllamaThe Ollama server has no user accounts. Identity lives in the application or gateway in front of it, through Keycloak or your own IdP such as Microsoft Entra ID.

Migrating from OpenAI API to Ollama

In the code, moving off the OpenAI API is mostly a change of base URL. Ollama answers chat completions, embeddings and the stateless Responses API on its own /v1 path, so the OpenAI SDK in your applications stays. The Batch and Files APIs, stateful Responses, tool_choice, logprobs and image URLs do not come across. Neither do the embeddings you already have: vectors from one model mean nothing to another, so each document index is rebuilt. The time goes on the model, not the code. An open model that fits your GPUs will be weaker than the hosted one at some tasks and as good at others, and we find out which on your own prompts before anyone repoints an application.

  1. List the calls

    Which applications call the API, with which models, and which features they rely on. Tools, JSON output and streaming carry over. Batches, stateful Responses and logprobs need rework, and we say where.

  2. Choose the model on your prompts

    Candidate open models are run against a set of your real prompts, and their answers are graded. The smallest model that passes wins, and its digest goes into the runbook.

  3. Repoint the client

    The OpenAI SDK stays and the base URL moves to the private endpoint. Context length is set in a Modelfile in your repository, because the OpenAI API has no field for it.

  4. Re-embed, then switch

    Document indexes are rebuilt with the new embedding model. Each application then runs against both for a cycle, and moves across once its owners have compared the answers.

An Ollama server, configured in your repository

# compose.yaml -- acme-llm-01, one GPU machine
services:
  ollama:
    image: ollama/ollama:${OLLAMA_TAG} # pinned in .env, tested on a copy first
    restart: unless-stopped
    ports:
      - '10.20.4.11:11434:11434' # private address only
    volumes:
      - /srv/ollama:/root/.ollama # models/manifests and models/blobs
    environment:
      OLLAMA_NO_CLOUD: '1' # no hosted models, no web search
      OLLAMA_NOHISTORY: '1' # no CLI prompt history in the volume
      OLLAMA_CONTEXT_LENGTH: '16384'
      OLLAMA_NUM_PARALLEL: '4' # memory grows with parallel x context
      OLLAMA_MAX_LOADED_MODELS: '2'
      OLLAMA_MAX_QUEUE: '128' # beyond this, callers get a 503
      OLLAMA_KEEP_ALIVE: '12h' # resident through the working day
      OLLAMA_FLASH_ATTENTION: '1'
      OLLAMA_KV_CACHE_TYPE: 'q8_0' # about half the cache memory of f16
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: all
              capabilities: [gpu]
An example compose file for one GPU machine at acme. Every setting that decides memory use, and who can reach the port, is written down and versioned instead of left to a default that changes with the GPU. The port is bound to a private address because Ollama asks callers for no credentials.

What Pilae is responsible for

A pinned version

A version we have run, not whatever latest resolves to that day.

A runbook

What it depends on, how it fails, what to do about it. In your repository.

A restore drill

Backups restored on a schedule. A backup nobody has restored is a file.

A patch window

Security updates in a window you agreed, with a rollback ready.

Someone watching

Every endpoint probed on the minute. An alert reaches a person, not a dashboard nobody opens.

Where it runs
zur1, fra1, fal1, gra1, ams1, hel1, lon1, ash1, hil1, sin1, tok1, syd1, on-premZurich, Frankfurt, Falkenstein, Gravelines, Amsterdam, Helsinki, London, Ashburn, Hillsboro, Singapore, Tokyo, Sydney, Your own hardware
Who holds the credentials
You do. Ours are separate, named, logged and revocable with one command. We ask before anything changes outside an agreed window.
If you leave
The machine, the data, the compose files and the runbook are already yours. Nothing stops when our access does.

What drives the price of running Ollama

Pricing is on request: a fixed price for onboarding, then a monthly price for Ollama, quoted in writing within five business days. The plans set what every deployment includes; these are the inputs the quote is built from.

Instance size
The CPU, memory and, where a model runs, the GPUs the app needs for your users and your data.
High availability
One machine with tested restores, or a replicated setup that keeps serving when a node fails.
Storage and backups
How much data it holds, how long backups are kept, and point-in-time recovery for its database.
Plan and support
Essential, Business or Enterprise: support hours, response times in the contract and how often we review the service with you.
Region
Your own hardware, where the infrastructure is already yours, or a Pilae Cloud region, where it is passed through at cost plus a fixed margin.
Sign-on and integrations
Single sign-on, directory sync, mail relays and the other systems the app has to reach.

Ollama: common questions

Is Ollama open source?

Yes. The server is MIT-licensed and has no paid edition. Ollama sells plans for its hosted cloud models, which we switch off. The models you run carry their own licences, set by their publishers, and we record each one before it is pinned.

Where do prompts and models live?

On the GPU machine and nowhere else: your own hardware, or a dedicated machine in a Pilae region such as Zurich or Frankfurt, in ISO 27001-certified datacentres. The Ollama server keeps no conversation history and we leave its request logging off. With cloud features off, a request cannot be forwarded to hosted models.

Should we run Ollama or vLLM?

Ollama for one GPU machine serving a team, a prototype, or several models on modest hardware, since it loads and unloads them on demand. vLLM for many people using one model at once, where throughput decides. Both speak the OpenAI API, so moving a busy model to vLLM changes the base URL, key and model name in each application, not its code.

Can people sign in to Ollama?

No. The Ollama server has no user accounts and its local API asks for no credentials. People reach it through an application with sign-on, such as Open WebUI or LibreChat, and services through a gateway that checks a token. The port itself answers only on the private network.

Will an open model answer as well as the OpenAI API?

At some tasks, yes, and at others it is weaker. The gap depends on the model your GPUs can hold. We measure it on your own prompts before any application is repointed, and we tell you when a task should stay where it is.

Also in ai and llm

Back to apps

Bring us your Ollama. We will tell you what it takes.

Thirty minutes on the deployment you already have, or the one you are about to start.