Managed Langfuse hosting

AI and LLMOn-prem or sovereign site

Tracing, evaluation and prompt management for LLM applications, run in Switzerland, the EU or your own datacentre so the prompts and answers it records stay with you. Pilae runs it on your own servers, or in Zurich, Switzerland, and eleven other Pilae Cloud regions.

Talk to us about Langfuse

Licence
MIT
Runs on
Your own hardware, or any of twelve Pilae regions — six of them in Switzerland and the EU
Upgrades
Pinned, tested against your configuration, applied in your window
Upstream
langfuse.com

Running Langfuse in production: what it takes

  1. Deploy the six services

    Web and worker containers, PostgreSQL, ClickHouse, Valkey and object storage with two buckets. Each service is pinned to a version we have run and runs in UTC. Nothing is published unless applications outside your network send traces, and then only the SDK endpoints and the media bucket, through the gate on port 443.

  2. Sign in through your directory

    Keycloak or Microsoft Entra ID over OpenID Connect, with password sign-in off and sessions capped at twelve hours. New users join your organisation with the role you chose.

  3. Decide retention before the disk does

    Self-hosted Langfuse keeps every trace until it is deleted, and automatic retention is in the Enterprise edition. On the MIT build a lifecycle rule expires raw events in the event bucket, the ClickHouse system log tables are trimmed, and we alert on ClickHouse disk well before it fills. If traces must go after a fixed period, we say before deployment that you need the Enterprise edition.

  4. Back up traces, buckets and secrets as one set

    PostgreSQL, a native ClickHouse backup and both buckets go offsite daily, encrypted, with the secrets file alongside. Once a month we restore the set into a scratch environment and open a trace there.

  5. Plan each major release

    Minor releases migrate the databases on start. Major releases are manual. One of them moved traces into new ClickHouse tables, with an optional backfill that wants roughly three times the disk, and stopped accepting the oldest SDKs. Each upgrade is rehearsed on a copy by the Pilae Agent and waits for your approval before it reaches your window.

What Langfuse is, and who runs it

Self-hosted Langfuse for LLM tracing and evaluation

Langfuse records what an LLM application actually did: each call, the prompt that went in, the context retrieval added, the answer that came back, and the tokens and cost. On top of those traces it keeps versioned prompts, datasets, human annotation and LLM-as-a-judge evaluations, so a change to a prompt or a model is tested against real cases before it ships.

It covers what teams use LangSmith for, and the tracing side of Helicone, which has been in maintenance mode since Mintlify bought it in March 2026. It does not need to sit in front of your model provider: applications send traces through its SDKs or OpenTelemetry, and Dify sends them natively. Langfuse itself has been part of ClickHouse, Inc. since January 2026, and the project says it has no planned changes to its licence.

Langfuse in production: ClickHouse, retention and where traces live

Langfuse looks like one app and runs as six services. The web container takes each batch of traces, writes it to a bucket and queues a reference in Valkey. The worker reads the batch and writes it to ClickHouse, while PostgreSQL holds users, projects and keys. A spike queues work instead of overloading the database, and a full restore brings back PostgreSQL, ClickHouse, both buckets and the secrets as one set.

Disk is what needs watching: self-hosted Langfuse keeps every trace until someone deletes it, and ClickHouse grows with every call your applications make. We settle retention on the first day and alert on ClickHouse disk well before it fills. Probes check the instance every 60 seconds, and an alert reaches an engineer.

Where Langfuse runs matters as much as where the model runs, because every trace carries the prompt and the answer. So it goes where the models are, beside vLLM or Ollama in a private AI deployment: on your own hardware, or in whichever of our twelve Pilae Cloud regions holds the GPUs. It answers only on your private network. Backups go offsite daily, encrypted, to your chosen country, and are restored into a scratch environment every month.

Langfuse licence and editions

Tracing, prompt management, datasets, evaluations and the playground are all in the MIT build, with no limit on usage and no per-user fee, and single sign-on through Keycloak or Entra ID works on it. We tell you before deployment whether your requirements need the Enterprise edition. Minor and major upgrades alike go through the Pilae Agent, tested on a copy and applied in your window. Our operation is priced on request. Talk to us about the applications you want to trace.

Langfuse is MIT-licensed except for its ee directories, which are under the Langfuse Enterprise License: free for development and testing, and for production only with a subscription. Project-level roles, SCIM, audit logs, data retention policies, server-side data masking, UI customisation and the instance management API are in that edition; single sign-on and organisation roles are in the MIT build. Langfuse has been part of ClickHouse, Inc. since January 2026 and sells that edition bundled with a ClickHouse commercial plan, including one that runs on your own machines. Under that edition, its usage telemetry cannot be switched off. We deploy the MIT build unless you need those features. If you do, the licence is held in your name and we say so in the proposal.

Langfuse system requirements

Before anything is deployed, this is what has to exist. We size it with you in the first session, and we say so when your own hardware is already enough.

CPU and memory
9 vCPU · 22 GBThe sum of upstream's minimums for the web and worker containers, PostgreSQL, Valkey and ClickHouse. ClickHouse is the part that grows, and upstream suggests 16 GB for it alone once volume picks up.
Trace store
ClickHouse 25.12+Traces, observations and scores. It must run in UTC: upstream warns that any other timezone returns wrong or empty results. Upstream recommends three replicas in production. A single node, backed up and restore-tested, fits where a restore rather than a failover is acceptable.
Database
PostgreSQL 15+Users, organisations, projects, API keys, datasets and settings. Usually small next to ClickHouse, and also kept in UTC.
Queue and cache
Valkey 8+ · noevictionQueues a reference to every incoming batch and caches API keys and prompts. Redis 7+ works too. Under any other eviction policy, queued jobs can be dropped when memory runs short.
Object storage
S3-compatible, 2 bucketsEvery incoming batch is written here before the worker reads it. Media, if you trace images, audio or files, goes to a second bucket. SDKs and browsers upload and fetch media there directly through presigned URLs, so your applications and your users must be able to reach it, not only Langfuse.

Migrating from LangSmith to Langfuse

Traces, datasets and prompts come across from LangSmith. Experiment scores and feedback do not, and evaluators are rebuilt rather than imported. Live traffic moves first and goes to both backends until the cutover. Datasets and Prompt Hub prompts are copied with the two SDKs, and prompt variables change from single to double braces. Past runs can be re-ingested over OpenTelemetry with their original timestamps, which is worth doing for the traces you still test against and rarely for the rest. The part that needs care is threads, users and tags, which LangSmith exports under its own attribute names.

  1. Send traces to both

    LangSmith's OpenTelemetry mode in Python, or the Langfuse callback handler in LangChain apps, sends each trace to both backends. JavaScript services move one process at a time. Nothing switches until the two views agree.

  2. Copy datasets and prompts

    A script reads examples and Prompt Hub commits with the LangSmith SDK and writes them with the Langfuse SDK, reusing example IDs so a rerun updates dataset items instead of duplicating them. Prompts are copied once, because every import adds a version.

  3. Rebuild evaluators, rerun experiments

    Evaluators become Langfuse LLM-as-a-judge evaluators or scores your own pipeline writes through the API. Experiments run again on the copied datasets rather than being imported.

  4. Map attributes, then cut over

    Threads, users and tags are mapped from LangSmith's attribute names to Langfuse sessions, users and tags. LangSmith export is switched off once a full release cycle has been traced in Langfuse.

The environment file behind a Langfuse deployment

# /srv/langfuse/langfuse.env  (acme, zur1)
# Read by langfuse-web and langfuse-worker. Passwords, SALT,
# ENCRYPTION_KEY, NEXTAUTH_SECRET and client secrets are in
# langfuse.secrets.env.
NEXTAUTH_URL=https://langfuse.acme.internal
DATABASE_HOST=pg-01.internal
DATABASE_USERNAME=langfuse
DATABASE_NAME=langfuse
CLICKHOUSE_URL=http://ch-01.internal:8123
CLICKHOUSE_MIGRATION_URL=clickhouse://ch-01.internal:9000
CLICKHOUSE_USER=langfuse
CLICKHOUSE_CLUSTER_ENABLED=false
REDIS_HOST=valkey-01.internal
REDIS_PORT=6379

# Raw event batches. Lifecycle rule on the bucket: expire after 30 days.
LANGFUSE_S3_EVENT_UPLOAD_BUCKET=acme-langfuse-events
LANGFUSE_S3_EVENT_UPLOAD_ENDPOINT=https://s3.zur1.acme.internal
LANGFUSE_S3_EVENT_UPLOAD_FORCE_PATH_STYLE=true
# Traces reference these objects, so no lifecycle rule here.
LANGFUSE_S3_MEDIA_UPLOAD_BUCKET=acme-langfuse-media
LANGFUSE_S3_MEDIA_UPLOAD_ENDPOINT=https://s3.zur1.acme.internal
LANGFUSE_S3_MEDIA_UPLOAD_FORCE_PATH_STYLE=true

# Keycloak only. First sign-in joins acme as a viewer; sessions last 12 h.
AUTH_KEYCLOAK_CLIENT_ID=langfuse
AUTH_KEYCLOAK_ISSUER=https://sso.acme.internal/realms/acme
AUTH_DISABLE_USERNAME_PASSWORD=true
AUTH_SESSION_MAX_AGE=720
LANGFUSE_DEFAULT_ORG_ID=acme
LANGFUSE_DEFAULT_ORG_ROLE=VIEWER
TELEMETRY_ENABLED=false
An example environment file for the web and worker containers. Every endpoint is on the private network. The two buckets get different retention, because traces point at media. Password sign-in and upstream telemetry are off. The secrets file is backed up with the data: the LLM credentials Langfuse stores for evaluations and the playground are encrypted with ENCRYPTION_KEY and cannot be read without it.

What Pilae is responsible for

A pinned version

A version we have run, not whatever latest resolves to that day.

A runbook

What it depends on, how it fails, what to do about it. In your repository.

A restore drill

Backups restored on a schedule. A backup nobody has restored is a file.

A patch window

Security updates in a window you agreed, with a rollback ready.

Someone watching

Every endpoint probed on the minute. An alert reaches a person, not a dashboard nobody opens.

Where it runs
zur1, fra1, fal1, gra1, ams1, hel1, lon1, ash1, hil1, sin1, tok1, syd1, on-premZurich, Frankfurt, Falkenstein, Gravelines, Amsterdam, Helsinki, London, Ashburn, Hillsboro, Singapore, Tokyo, Sydney, Your own hardware
Who holds the credentials
You do. Ours are separate, named, logged and revocable with one command. We ask before anything changes outside an agreed window.
If you leave
The machine, the data, the compose files and the runbook are already yours. Nothing stops when our access does.

What drives the price of running Langfuse

Pricing is on request: a fixed price for onboarding, then a monthly price for Langfuse, quoted in writing within five business days. The plans set what every deployment includes; these are the inputs the quote is built from.

Instance size
The CPU, memory and, where a model runs, the GPUs the app needs for your users and your data.
High availability
One machine with tested restores, or a replicated setup that keeps serving when a node fails.
Storage and backups
How much data it holds, how long backups are kept, and point-in-time recovery for its database.
Plan and support
Essential, Business or Enterprise: support hours, response times in the contract and how often we review the service with you.
Region
Your own hardware, where the infrastructure is already yours, or a Pilae Cloud region, where it is passed through at cost plus a fixed margin.
Sign-on and integrations
Single sign-on, directory sync, mail relays and the other systems the app has to reach.

Langfuse: common questions

Is Langfuse open source?

Yes, apart from its ee directories. Everything else is MIT, including tracing, prompt management, datasets, LLM-as-a-judge evaluations, annotation queues, the playground and single sign-on. The ee directories are under the Langfuse Enterprise License, which needs a subscription for production use.

Where do our traces live?

In ClickHouse and in a bucket, in the same place as the models whose calls they record: your own hardware, or dedicated machines in that Pilae region, in ISO 27001-certified datacentres. A trace holds the prompt, the retrieved context and the answer, so it is as sensitive as the documents behind it. Self-hosted Langfuse needs no internet access, and we switch off its usage telemetry.

What does the open-source edition leave out?

Data retention policies, project-level roles, SCIM, audit logs, server-side data masking, UI customisation and the instance management API. Retention is the one that matters in practice, because without it every trace is kept until someone deletes it. On the MIT build we expire raw events in the event bucket and size ClickHouse for the history you keep. If you need traces deleted on a schedule, per-project roles or an audit trail for a regulator, that is the Enterprise edition, which Langfuse sells with a ClickHouse plan.

Can Langfuse replace LangSmith?

Yes, for tracing, datasets, prompt management and evaluations. It does not assume LangChain: it takes OpenTelemetry and has its own SDKs. What does not come across is LangSmith's history of experiment scores and feedback, and its tuned evaluators do not port. We list those during the inventory, before anything switches.

Can evaluations run on our own models?

Yes. LLM-as-a-judge and the playground call any OpenAI-compatible endpoint, so they can use a model served with vLLM or Ollama on your GPUs instead of sending traces to a hosted provider for scoring. The judge model has to support structured output, and because Langfuse blocks private addresses for LLM connections by default, we add your model host to its allow-list.

Also in ai and llm

Back to apps

Bring us your Langfuse. We will tell you what it takes.

Thirty minutes on the deployment you already have, or the one you are about to start.