A vector database for semantic search and RAG, run on your own hardware or in Switzerland and the EU, so the embeddings of your documents stay where the documents are. Pilae runs it on your own servers, or in Zurich, Switzerland, and eleven other Pilae Cloud regions.
- Licence
- Apache-2.0
- Runs on
- Your own hardware, or any of twelve Pilae regions — six of them in Switzerland and the EU
- Upgrades
- Pinned, tested against your configuration, applied in your window
- Upstream
- qdrant.tech
Running Qdrant in production: what it takes
Size it from the vector count
We count points, dimensions and replicas, add the index and payload on top, and decide what sits in RAM, what is read from disk and whether quantisation pays. Recall is tested on a sample of your data before a machine is ordered.
Lock it down before the first collection
We set an admin key, a read-only key and TLS before any collection exists, turn off the anonymous usage statistics Qdrant sends by default, and keep every port on the private network.
Keep each application to its own collections
With JWT access control on, an application that writes to collections we create for it gets a token scoped to those collections, so it cannot read another application's documents. Dify, and Open WebUI in its default mode, create a collection for each knowledge base, which a scoped token is not allowed to do, so each of those gets a Qdrant instance of its own. The tokens are signed with the admin key, so rotating it is a planned change.
Snapshot every collection, restore it monthly
A snapshot of every collection, taken on every node in a cluster, goes offsite daily, encrypted. Once a month we restore them into a scratch instance on the same version and run the baseline queries against it.
Upgrade one minor version at a time
Qdrant cannot skip minor versions, even on one node, because each step applies its own data migrations. Each step is tested on a copy by the Pilae Agent and applied in your window once you approve it. With two replicas, nodes restart in turn without downtime.
What Qdrant is, and who runs it
Self-hosted Qdrant for semantic search and RAG
Qdrant is a vector database written in Rust. It stores each embedding with a JSON payload beside it and answers nearest-neighbour queries filtered on that payload, over HTTP or gRPC. Dense, sparse and multi-vectors can sit in one collection, so a semantic search and a keyword-style search can run in the same query. In a RAG system it is the retrieval tier: Dify, Open WebUI and n8n can each use it as their vector store.
It is not always needed. When your documents already live in PostgreSQL and the index fits in that machine’s memory, pgvector in the same database is our default, because it adds nothing new to back up or secure. We move retrieval to Qdrant when the vectors outgrow the database machine, when quantisation is what keeps memory affordable, when queries filter heavily or combine dense and sparse search, or when several applications in a private AI deployment share one retrieval tier.
Qdrant in production: memory, shards and upgrades
Memory is the budget. Float32 vectors cost points × dimensions × 4 bytes per replica, and the HNSW graph, the payload indexes and the ID tracker come on top. Whether the search runs on the full vectors or on an int8 or binary copy is checked against recall on your own queries. The shard count is fixed when a collection is created, so we set it for the cluster you will have, not the one you start with.
It runs on dedicated machines next to the documents it indexes: on your own hardware, or in the Pilae Cloud region that holds the rest of the deployment. It answers only on your private network, with keys and TLS in place before any data arrives. Probes check every node every 60 seconds, and an engineer is alerted when one drops out.
Qdrant cannot skip a minor version when it upgrades, and a snapshot restores only into the same or the next minor version, so a deployment left to fall behind has to be walked forward one step at a time. We keep it current through the Pilae Agent: each step tested on a copy, approved by you, applied in your window and recorded in the console.
Qdrant licence and editions
Qdrant Solutions GmbH, which develops it, is in Berlin. The database you self-host has no fee per node or per user. Pricing for our operation is on request. Talk to us about the index you want to move off Pinecone.
Qdrant system requirements
Before anything is deployed, this is what has to exist. We size it with you in the first session, and we say so when your own hardware is already enough.
- Memory
- points × dims × 4 B × replicasThe float32 floor for the vectors alone, before the HNSW graph, payload indexes and 20% headroom. Quantisation brings it down: an int8 copy in RAM, with the originals on disk, needs a quarter of it.
- CPU
- 2+ vCPU · 64-bitIndexing and segment optimisation are CPU work, and they run while queries are served. Only 64-bit x86 and ARM are supported.
- Disk
- Local SSD or NVMe, POSIXQdrant needs block storage with a POSIX filesystem and does not work on NFS or object storage. Budget a fifth more for the write-ahead log, snapshots and optimiser segments.
- Access
- API key · read-only key · TLSSelf-hosted Qdrant starts with no authentication on every interface, so none of these is optional.
- Ports
- 6333 HTTP · 6334 gRPC · 6335 p2pAll on the private network. 6335 carries traffic between nodes, is open only when Qdrant runs as a cluster and never checks an API key, so only the other nodes can reach it.
Migrating from Pinecone to Qdrant
The vectors come across; the embedding model does not. Qdrant's own migration tool copies a serverless Pinecone index into a collection: vectors, sparse values and metadata. Pod-based indexes cannot be listed that way, so for those we re-embed from the source documents. Pinecone IDs are strings and Qdrant wants integers or UUIDs, so the original ID travels in a payload field. Namespaces have no direct equivalent: they become a tenant field or separate collections, decided before the copy. If Pinecone embedded the text for you, that model now has to run somewhere, and changing it means re-embedding everything. Moving the bytes is quick; proving that search results still match is the slow part.
Record a baseline
Vector counts, a sample of metadata and the top results for a set of real queries, captured from Pinecone before anything moves. The new collection is judged against this, not against impressions.
Create the collection deliberately
Dimensions, distance, shard count, quantisation and the tenant field are set before the copy. The shard count cannot be changed later without creating the collection again.
Copy with the migration tool
The tool reads the serverless index in batches and, if the copy is interrupted, picks up where it stopped. The application then looks records up by the payload field that holds the Pinecone ID.
Compare, then switch the client
The baseline queries run against both. Once recall matches and any score threshold in the application has been checked again, the client library changes and Pinecone stays read-only for one cycle.
A Qdrant collection, sized before it is created
# acme-docs · 8M chunks · 1024-dim embeddings · 3 nodes
# originals, float32: 8M × 1024 × 4 B × 2 replicas = 65.5 GB, on disk
# int8 copies in RAM: 8M × 1024 × 1 B × 2 replicas = 16.4 GB
# HNSW graph 2.5 GB · ID tracker 0.8 GB · plus 20% headroom
PUT /collections/acme-docs
{
"vectors": { "size": 1024, "distance": "Cosine", "memory": "cold" },
"quantization_config": {
"scalar": { "type": "int8", "memory": "pinned" }
},
"hnsw_config": { "memory": "cached" },
"shard_number": 6,
"replication_factor": 2,
"write_consistency_factor": 1
}
# config/production.yaml, identical on each node
telemetry_disabled: true
cluster:
enabled: true
p2p:
enable_tls: true
service:
enable_tls: true
jwt_rbac: true
tls:
cert: ./tls/cert.pem
key: ./tls/key.pem
ca_cert: ./tls/cacert.pem
# api_key and read_only_api_key arrive from the secret store as
# QDRANT__SERVICE__API_KEY and QDRANT__SERVICE__READ_ONLY_API_KEY
What Pilae is responsible for
A pinned version
A version we have run, not whatever latest resolves to that day.
A runbook
What it depends on, how it fails, what to do about it. In your repository.
A restore drill
Backups restored on a schedule. A backup nobody has restored is a file.
A patch window
Security updates in a window you agreed, with a rollback ready.
Someone watching
Every endpoint probed on the minute. An alert reaches a person, not a dashboard nobody opens.
- Where it runs
- zur1, fra1, fal1, gra1, ams1, hel1, lon1, ash1, hil1, sin1, tok1, syd1, on-premZurich, Frankfurt, Falkenstein, Gravelines, Amsterdam, Helsinki, London, Ashburn, Hillsboro, Singapore, Tokyo, Sydney, Your own hardware
- Who holds the credentials
- You do. Ours are separate, named, logged and revocable with one command. We ask before anything changes outside an agreed window.
- If you leave
- The machine, the data, the compose files and the runbook are already yours. Nothing stops when our access does.
What drives the price of running Qdrant
Pricing is on request: a fixed price for onboarding, then a monthly price for Qdrant, quoted in writing within five business days. The plans set what every deployment includes; these are the inputs the quote is built from.
- Instance size
- The CPU, memory and, where a model runs, the GPUs the app needs for your users and your data.
- High availability
- One machine with tested restores, or a replicated setup that keeps serving when a node fails.
- Storage and backups
- How much data it holds, how long backups are kept, and point-in-time recovery for its database.
- Plan and support
- Essential, Business or Enterprise: support hours, response times in the contract and how often we review the service with you.
- Region
- Your own hardware, where the infrastructure is already yours, or a Pilae Cloud region, where it is passed through at cost plus a fixed margin.
- Sign-on and integrations
- Single sign-on, directory sync, mail relays and the other systems the app has to reach.
Qdrant: common questions
Is Qdrant open source?
Do we need Qdrant, or is pgvector enough?
Where do the vectors and the text live?
Can Qdrant sign in with our directory?
What happens if a node fails?
Also in ai and llm
Open WebUI
A chat interface over models you host, so prompts and the documents people paste into them never leave your network.
Replaces ChatGPT Team, Microsoft Copilot
Ollama
A server that loads open-weight models on demand behind an OpenAI-compatible API, run on GPUs in Switzerland, the EU or your own datacentre so prompts never reach a model vendor.
Replaces OpenAI API, Azure OpenAI
vLLM
An OpenAI-compatible inference server for open-weight models, run on GPUs in Switzerland, the EU or your own datacentre, so prompts and answers never reach a model vendor.
Replaces OpenAI API, Azure OpenAI
Bring us your Qdrant. We will tell you what it takes.
Thirty minutes on the deployment you already have, or the one you are about to start.