Business intelligence from the Apache Software Foundation for data teams that write SQL, run next to your warehouse in Switzerland, the EU or your own datacentre. Pilae runs it on your own servers, or in Zurich, Switzerland, and eleven other Pilae Cloud regions.
- Licence
- Apache-2.0
- Runs on
- Your own hardware, or any of twelve Pilae regions — six of them in Switzerland and the EU
- Upgrades
- Pinned, tested against your configuration, applied in your window
- Upstream
- superset.apache.org
Running Apache Superset in production: what it takes
Build the image, then deploy app, workers and beat
A pinned upstream release, extended with the drivers for your databases and Playwright with a headless Chromium for reports. The web app, the Celery workers and exactly one beat run on PostgreSQL and Redis, from the same configuration file.
Map directory groups to roles
OpenID Connect to Keycloak or Entra ID, with Keycloak groups or Entra ID app roles mapped to Superset roles at every login. Row-level security filters are assigned to the same roles, so a leaver loses the dashboards and the rows together.
Give each database a read-only account
Every connection reaches its database over the private network with an account that can read only the schemas people need. SQL Lab is granted per role and per database, and long queries run asynchronously on the workers instead of timing out in the browser.
Back up the database and the key together
The metadata database goes offsite daily, encrypted, to your chosen country, and the secret key is escrowed beside it. Once a month we restore both into a scratch environment and run a chart through a restored connection, which proves the key matches.
Upgrade with a snapshot ready
Each release runs superset db upgrade and superset init. Some migrations do not reverse cleanly, so the rollback kept ready is a snapshot taken just before. The Pilae Agent tests the upgrade on a copy, waits for your approval, applies it in your window and records it in the console.
What Apache Superset is, and who runs it
Self-hosted Apache Superset for SQL analytics and dashboards
Apache Superset is a business intelligence platform from the Apache Software Foundation. Analysts write queries in SQL Lab, define metrics once on a dataset, and build charts and dashboards on top, geospatial maps included. It reads from almost any SQL database with a Python driver and a SQLAlchemy dialect: PostgreSQL, ClickHouse, Trino, Snowflake, BigQuery and others.
It is built for teams that already write SQL and publish to the rest of the organisation. If most of your readers will ask simple questions without writing any, Metabase is easier to adopt and lighter to run, and we say so. Superset belongs next to the warehouse it reads, so that is where we put it: on your own hardware, or in the Pilae Cloud region that already holds the data, one of 12, six of them in Switzerland and the EU.
Apache Superset in production: workers, beat and the secret key
Superset is several processes, not one. The web app answers people, Celery workers run long queries and take the screenshots for reports, and a single beat decides when. Redis sits between them. Their failures are quiet: a second beat can send every report twice, and a worker that meets the sign-in page fails the report. The default SQLite file does not work once workers share the metadata database, so we start on PostgreSQL. We run one beat and point the worker at the internal address. Probes check the web app every 60 seconds, and an engineer is alerted when a report schedule stops completing.
The metadata database is backed up daily, encrypted, to an offsite location in your chosen country. The secret key is escrowed beside it, because without it the stored connection passwords cannot be decrypted. Sign-on goes through Keycloak or your own IdP, and the app answers only on your private network. Upgrades run database migrations and reset the built-in roles to their defaults, so custom permissions live in custom roles. The Pilae Agent runs each upgrade on a copy first and applies it in your window after you approve it.
Apache Superset licence and editions
With nothing held back for a paid edition, there is no seat count to track and no licence to hold in your name. We operate the Apache release, on infrastructure you choose. Our operation is priced on request. Talk to us about the workbooks you want to bring across from Tableau.
Apache Superset system requirements
Before anything is deployed, this is what has to exist. We size it with you in the first session, and we say so when your own hardware is already enough.
- CPU and memory
- 4 vCPU · 8 GBOur starting point for the web app, workers and beat on one machine. Screenshot reports run a headless Chromium inside the worker, and that is what wants the memory.
- Metadata database
- PostgreSQL 14–17Holds charts, dashboards, users, query history and the encrypted connection passwords. Upstream supports PostgreSQL and MySQL; the default SQLite file is not for production.
- Cache and queue
- RedisCaches chart results and carries the Celery queue for asynchronous queries, alerts and reports. Out of the box the broker is a SQLite file, which upstream says to replace in production.
- Secret key
- SECRET_KEY, 42 random bytesSuperset refuses to start on the default key. It signs sessions and encrypts the stored database passwords, so a restore without the same key cannot decrypt them.
- Data sources
- Read-only accounts + driverEach engine needs its Python driver in the image, and each connection an account that can only read the schemas people need.
Migrating from Tableau to Apache Superset
Tableau workbooks do not open in Superset and there is no built-in importer, so dashboards are rebuilt. The charts are the quick part; the data is the slow one. Superset has no extract engine: it queries the database live, with a result cache in front, so anything that only exists in a Tableau extract, or in a blend of two sources, has to land in a SQL database first. Calculated fields become metrics and calculated columns on Superset datasets, written in SQL and shared by every chart built on them. A dashboard laid out to the pixel will not look the same, and a few Tableau chart types have no Superset counterpart. Those go on a list the owners see before any rebuild is scheduled.
Start from the workbooks people open
Tableau Server and Tableau Cloud record who opened what. The last three months of it pick the workbooks worth rebuilding, and the others lapse with the licence instead of being carried across.
Land the extracts in a database
Data that only lives in Tableau extracts or cross-source blends is loaded into a SQL database Superset can reach, often a PostgreSQL we already run for you, with a scheduled load in place of the extract refresh.
Rebuild calculations as datasets
Calculated fields and LOD expressions become SQL metrics, calculated columns or virtual datasets, checked figure by figure against the Tableau views they replace.
Move the subscriptions last
Email subscriptions become Superset reports once owners have checked a full reporting cycle in both tools. Tableau is kept read-only until the first month-end has closed on Superset.
One configuration file for the app, the workers and the beat
# superset_config.py (excerpt), mounted read-only into app, worker and beat
import os
from celery.schedules import crontab
from flask_appbuilder.security.manager import AUTH_OAUTH
from superset.config import CeleryConfig
SECRET_KEY = os.environ["SUPERSET_SECRET_KEY"] # escrowed with the backups
SQLALCHEMY_DATABASE_URI = os.environ["SUPERSET_METADATA_URI"]
REDIS = "redis://redis.internal:6379"
DATA_CACHE_CONFIG = {
"CACHE_TYPE": "RedisCache",
"CACHE_REDIS_URL": f"{REDIS}/1",
"CACHE_KEY_PREFIX": "acme_data",
"CACHE_DEFAULT_TIMEOUT": 3600,
}
class AcmeCeleryConfig(CeleryConfig):
broker_url = f"{REDIS}/0"
result_backend = f"{REDIS}/0"
beat_schedule = {
**CeleryConfig.beat_schedule,
"prune_query": {
"task": "prune_query",
"schedule": crontab(minute=0, hour=0, day_of_month=1),
"kwargs": {"retention_period_days": 180},
},
}
CELERY_CONFIG = AcmeCeleryConfig
FEATURE_FLAGS = {"ALERT_REPORTS": True, "PLAYWRIGHT_REPORTS_AND_THUMBNAILS": True}
WEBDRIVER_BASEURL = "http://superset-app:8088/"
WEBDRIVER_BASEURL_USER_FRIENDLY = "https://bi.acme.internal/"
AUTH_TYPE = AUTH_OAUTH
OAUTH_PROVIDERS = [{
"name": "keycloak",
"token_key": "access_token",
"remote_app": {
"client_id": "superset",
"client_secret": os.environ["OIDC_CLIENT_SECRET"],
"api_base_url": "https://sso.acme.internal/realms/acme/protocol/openid-connect",
"access_token_url": "https://sso.acme.internal/realms/acme/protocol/openid-connect/token",
"authorize_url": "https://sso.acme.internal/realms/acme/protocol/openid-connect/auth",
"client_kwargs": {"scope": "openid email profile"},
},
}]
AUTH_USER_REGISTRATION = True
AUTH_USER_REGISTRATION_ROLE = "Gamma"
# group paths from the "groups" claim, set by a Keycloak group-membership mapper
AUTH_ROLES_MAPPING = {"/bi-analysts": ["Alpha", "sql_lab"], "/bi-admins": ["Admin"]}
AUTH_ROLES_SYNC_AT_LOGIN = True
ENABLE_PROXY_FIX = True
What Pilae is responsible for
A pinned version
A version we have run, not whatever latest resolves to that day.
A runbook
What it depends on, how it fails, what to do about it. In your repository.
A restore drill
Backups restored on a schedule. A backup nobody has restored is a file.
A patch window
Security updates in a window you agreed, with a rollback ready.
Someone watching
Every endpoint probed on the minute. An alert reaches a person, not a dashboard nobody opens.
- Where it runs
- zur1, fra1, fal1, gra1, ams1, hel1, lon1, ash1, hil1, sin1, tok1, syd1, on-premZurich, Frankfurt, Falkenstein, Gravelines, Amsterdam, Helsinki, London, Ashburn, Hillsboro, Singapore, Tokyo, Sydney, Your own hardware
- Who holds the credentials
- You do. Ours are separate, named, logged and revocable with one command. We ask before anything changes outside an agreed window.
- If you leave
- The machine, the data, the compose files and the runbook are already yours. Nothing stops when our access does.
What drives the price of running Apache Superset
Pricing is on request: a fixed price for onboarding, then a monthly price for Apache Superset, quoted in writing within five business days. The plans set what every deployment includes; these are the inputs the quote is built from.
- Instance size
- The CPU, memory and, where a model runs, the GPUs the app needs for your users and your data.
- High availability
- One machine with tested restores, or a replicated setup that keeps serving when a node fails.
- Storage and backups
- How much data it holds, how long backups are kept, and point-in-time recovery for its database.
- Plan and support
- Essential, Business or Enterprise: support hours, response times in the contract and how often we review the service with you.
- Region
- Your own hardware, where the infrastructure is already yours, or a Pilae Cloud region, where it is passed through at cost plus a fixed margin.
- Sign-on and integrations
- Single sign-on, directory sync, mail relays and the other systems the app has to reach.
Apache Superset: common questions
Is Apache Superset open source?
Superset or Metabase: which should we choose?
Where do our data and dashboards live?
Can people run any SQL they like against our databases?
Why would a scheduled report not arrive?
Also in workflow and data
n8n
Workflow automation your own team writes, running next to the systems it talks to instead of reaching them across the internet.
Replaces Zapier, Make
NocoDB
A spreadsheet interface over a real database, so a department can build the tool it needs without anyone provisioning a new system.
Replaces Airtable
Directus
A data platform and headless CMS over your own SQL database. Editors get an admin app, and every table gets an API.
Replaces Contentful, Airtable
Bring us your Apache Superset. We will tell you what it takes.
Thirty minutes on the deployment you already have, or the one you are about to start.