Tracing, evaluation and prompt management for LLM applications, run in Switzerland, the EU or your own datacentre so the prompts and answers it records stay with you. Pilae runs it on your own servers, or in Zurich, Switzerland, and eleven other Pilae Cloud regions.
- Licence
- MIT
- Runs on
- Your own hardware, or any of twelve Pilae regions — six of them in Switzerland and the EU
- Upgrades
- Pinned, tested against your configuration, applied in your window
- Upstream
- langfuse.com
Running Langfuse in production: what it takes
Deploy the six services
Web and worker containers, PostgreSQL, ClickHouse, Valkey and object storage with two buckets. Each service is pinned to a version we have run and runs in UTC. Nothing is published unless applications outside your network send traces, and then only the SDK endpoints and the media bucket, through the gate on port 443.
Sign in through your directory
Keycloak or Microsoft Entra ID over OpenID Connect, with password sign-in off and sessions capped at twelve hours. New users join your organisation with the role you chose.
Decide retention before the disk does
Self-hosted Langfuse keeps every trace until it is deleted, and automatic retention is in the Enterprise edition. On the MIT build a lifecycle rule expires raw events in the event bucket, the ClickHouse system log tables are trimmed, and we alert on ClickHouse disk well before it fills. If traces must go after a fixed period, we say before deployment that you need the Enterprise edition.
Back up traces, buckets and secrets as one set
PostgreSQL, a native ClickHouse backup and both buckets go offsite daily, encrypted, with the secrets file alongside. Once a month we restore the set into a scratch environment and open a trace there.
Plan each major release
Minor releases migrate the databases on start. Major releases are manual. One of them moved traces into new ClickHouse tables, with an optional backfill that wants roughly three times the disk, and stopped accepting the oldest SDKs. Each upgrade is rehearsed on a copy by the Pilae Agent and waits for your approval before it reaches your window.
What Langfuse is, and who runs it
Self-hosted Langfuse for LLM tracing and evaluation
Langfuse records what an LLM application actually did: each call, the prompt that went in, the context retrieval added, the answer that came back, and the tokens and cost. On top of those traces it keeps versioned prompts, datasets, human annotation and LLM-as-a-judge evaluations, so a change to a prompt or a model is tested against real cases before it ships.
It covers what teams use LangSmith for, and the tracing side of Helicone, which has been in maintenance mode since Mintlify bought it in March 2026. It does not need to sit in front of your model provider: applications send traces through its SDKs or OpenTelemetry, and Dify sends them natively. Langfuse itself has been part of ClickHouse, Inc. since January 2026, and the project says it has no planned changes to its licence.
Langfuse in production: ClickHouse, retention and where traces live
Langfuse looks like one app and runs as six services. The web container takes each batch of traces, writes it to a bucket and queues a reference in Valkey. The worker reads the batch and writes it to ClickHouse, while PostgreSQL holds users, projects and keys. A spike queues work instead of overloading the database, and a full restore brings back PostgreSQL, ClickHouse, both buckets and the secrets as one set.
Disk is what needs watching: self-hosted Langfuse keeps every trace until someone deletes it, and ClickHouse grows with every call your applications make. We settle retention on the first day and alert on ClickHouse disk well before it fills. Probes check the instance every 60 seconds, and an alert reaches an engineer.
Where Langfuse runs matters as much as where the model runs, because every trace carries the prompt and the answer. So it goes where the models are, beside vLLM or Ollama in a private AI deployment: on your own hardware, or in whichever of our twelve Pilae Cloud regions holds the GPUs. It answers only on your private network. Backups go offsite daily, encrypted, to your chosen country, and are restored into a scratch environment every month.
Langfuse licence and editions
Tracing, prompt management, datasets, evaluations and the playground are all in the MIT build, with no limit on usage and no per-user fee, and single sign-on through Keycloak or Entra ID works on it. We tell you before deployment whether your requirements need the Enterprise edition. Minor and major upgrades alike go through the Pilae Agent, tested on a copy and applied in your window. Our operation is priced on request. Talk to us about the applications you want to trace.
Langfuse system requirements
Before anything is deployed, this is what has to exist. We size it with you in the first session, and we say so when your own hardware is already enough.
- CPU and memory
- 9 vCPU · 22 GBThe sum of upstream's minimums for the web and worker containers, PostgreSQL, Valkey and ClickHouse. ClickHouse is the part that grows, and upstream suggests 16 GB for it alone once volume picks up.
- Trace store
- ClickHouse 25.12+Traces, observations and scores. It must run in UTC: upstream warns that any other timezone returns wrong or empty results. Upstream recommends three replicas in production. A single node, backed up and restore-tested, fits where a restore rather than a failover is acceptable.
- Database
- PostgreSQL 15+Users, organisations, projects, API keys, datasets and settings. Usually small next to ClickHouse, and also kept in UTC.
- Queue and cache
- Valkey 8+ · noevictionQueues a reference to every incoming batch and caches API keys and prompts. Redis 7+ works too. Under any other eviction policy, queued jobs can be dropped when memory runs short.
- Object storage
- S3-compatible, 2 bucketsEvery incoming batch is written here before the worker reads it. Media, if you trace images, audio or files, goes to a second bucket. SDKs and browsers upload and fetch media there directly through presigned URLs, so your applications and your users must be able to reach it, not only Langfuse.
Migrating from LangSmith to Langfuse
Traces, datasets and prompts come across from LangSmith. Experiment scores and feedback do not, and evaluators are rebuilt rather than imported. Live traffic moves first and goes to both backends until the cutover. Datasets and Prompt Hub prompts are copied with the two SDKs, and prompt variables change from single to double braces. Past runs can be re-ingested over OpenTelemetry with their original timestamps, which is worth doing for the traces you still test against and rarely for the rest. The part that needs care is threads, users and tags, which LangSmith exports under its own attribute names.
Send traces to both
LangSmith's OpenTelemetry mode in Python, or the Langfuse callback handler in LangChain apps, sends each trace to both backends. JavaScript services move one process at a time. Nothing switches until the two views agree.
Copy datasets and prompts
A script reads examples and Prompt Hub commits with the LangSmith SDK and writes them with the Langfuse SDK, reusing example IDs so a rerun updates dataset items instead of duplicating them. Prompts are copied once, because every import adds a version.
Rebuild evaluators, rerun experiments
Evaluators become Langfuse LLM-as-a-judge evaluators or scores your own pipeline writes through the API. Experiments run again on the copied datasets rather than being imported.
Map attributes, then cut over
Threads, users and tags are mapped from LangSmith's attribute names to Langfuse sessions, users and tags. LangSmith export is switched off once a full release cycle has been traced in Langfuse.
The environment file behind a Langfuse deployment
# /srv/langfuse/langfuse.env (acme, zur1) # Read by langfuse-web and langfuse-worker. Passwords, SALT, # ENCRYPTION_KEY, NEXTAUTH_SECRET and client secrets are in # langfuse.secrets.env. NEXTAUTH_URL=https://langfuse.acme.internal DATABASE_HOST=pg-01.internal DATABASE_USERNAME=langfuse DATABASE_NAME=langfuse CLICKHOUSE_URL=http://ch-01.internal:8123 CLICKHOUSE_MIGRATION_URL=clickhouse://ch-01.internal:9000 CLICKHOUSE_USER=langfuse CLICKHOUSE_CLUSTER_ENABLED=false REDIS_HOST=valkey-01.internal REDIS_PORT=6379 # Raw event batches. Lifecycle rule on the bucket: expire after 30 days. LANGFUSE_S3_EVENT_UPLOAD_BUCKET=acme-langfuse-events LANGFUSE_S3_EVENT_UPLOAD_ENDPOINT=https://s3.zur1.acme.internal LANGFUSE_S3_EVENT_UPLOAD_FORCE_PATH_STYLE=true # Traces reference these objects, so no lifecycle rule here. LANGFUSE_S3_MEDIA_UPLOAD_BUCKET=acme-langfuse-media LANGFUSE_S3_MEDIA_UPLOAD_ENDPOINT=https://s3.zur1.acme.internal LANGFUSE_S3_MEDIA_UPLOAD_FORCE_PATH_STYLE=true # Keycloak only. First sign-in joins acme as a viewer; sessions last 12 h. AUTH_KEYCLOAK_CLIENT_ID=langfuse AUTH_KEYCLOAK_ISSUER=https://sso.acme.internal/realms/acme AUTH_DISABLE_USERNAME_PASSWORD=true AUTH_SESSION_MAX_AGE=720 LANGFUSE_DEFAULT_ORG_ID=acme LANGFUSE_DEFAULT_ORG_ROLE=VIEWER TELEMETRY_ENABLED=false
What Pilae is responsible for
A pinned version
A version we have run, not whatever latest resolves to that day.
A runbook
What it depends on, how it fails, what to do about it. In your repository.
A restore drill
Backups restored on a schedule. A backup nobody has restored is a file.
A patch window
Security updates in a window you agreed, with a rollback ready.
Someone watching
Every endpoint probed on the minute. An alert reaches a person, not a dashboard nobody opens.
- Where it runs
- zur1, fra1, fal1, gra1, ams1, hel1, lon1, ash1, hil1, sin1, tok1, syd1, on-premZurich, Frankfurt, Falkenstein, Gravelines, Amsterdam, Helsinki, London, Ashburn, Hillsboro, Singapore, Tokyo, Sydney, Your own hardware
- Who holds the credentials
- You do. Ours are separate, named, logged and revocable with one command. We ask before anything changes outside an agreed window.
- If you leave
- The machine, the data, the compose files and the runbook are already yours. Nothing stops when our access does.
What drives the price of running Langfuse
Pricing is on request: a fixed price for onboarding, then a monthly price for Langfuse, quoted in writing within five business days. The plans set what every deployment includes; these are the inputs the quote is built from.
- Instance size
- The CPU, memory and, where a model runs, the GPUs the app needs for your users and your data.
- High availability
- One machine with tested restores, or a replicated setup that keeps serving when a node fails.
- Storage and backups
- How much data it holds, how long backups are kept, and point-in-time recovery for its database.
- Plan and support
- Essential, Business or Enterprise: support hours, response times in the contract and how often we review the service with you.
- Region
- Your own hardware, where the infrastructure is already yours, or a Pilae Cloud region, where it is passed through at cost plus a fixed margin.
- Sign-on and integrations
- Single sign-on, directory sync, mail relays and the other systems the app has to reach.
Langfuse: common questions
Is Langfuse open source?
Where do our traces live?
What does the open-source edition leave out?
Can Langfuse replace LangSmith?
Can evaluations run on our own models?
Also in ai and llm
Open WebUI
A chat interface over models you host, so prompts and the documents people paste into them never leave your network.
Replaces ChatGPT Team, Microsoft Copilot
Ollama
A server that loads open-weight models on demand behind an OpenAI-compatible API, run on GPUs in Switzerland, the EU or your own datacentre so prompts never reach a model vendor.
Replaces OpenAI API, Azure OpenAI
vLLM
An OpenAI-compatible inference server for open-weight models, run on GPUs in Switzerland, the EU or your own datacentre, so prompts and answers never reach a model vendor.
Replaces OpenAI API, Azure OpenAI
Bring us your Langfuse. We will tell you what it takes.
Thirty minutes on the deployment you already have, or the one you are about to start.