Managed Apache Superset hosting

Workflow and dataOn-prem or sovereign site

Business intelligence from the Apache Software Foundation for data teams that write SQL, run next to your warehouse in Switzerland, the EU or your own datacentre. Pilae runs it on your own servers, or in Zurich, Switzerland, and eleven other Pilae Cloud regions.

Talk to us about Apache Superset

Licence
Apache-2.0
Runs on
Your own hardware, or any of twelve Pilae regions — six of them in Switzerland and the EU
Upgrades
Pinned, tested against your configuration, applied in your window
Upstream
superset.apache.org

Running Apache Superset in production: what it takes

  1. Build the image, then deploy app, workers and beat

    A pinned upstream release, extended with the drivers for your databases and Playwright with a headless Chromium for reports. The web app, the Celery workers and exactly one beat run on PostgreSQL and Redis, from the same configuration file.

  2. Map directory groups to roles

    OpenID Connect to Keycloak or Entra ID, with Keycloak groups or Entra ID app roles mapped to Superset roles at every login. Row-level security filters are assigned to the same roles, so a leaver loses the dashboards and the rows together.

  3. Give each database a read-only account

    Every connection reaches its database over the private network with an account that can read only the schemas people need. SQL Lab is granted per role and per database, and long queries run asynchronously on the workers instead of timing out in the browser.

  4. Back up the database and the key together

    The metadata database goes offsite daily, encrypted, to your chosen country, and the secret key is escrowed beside it. Once a month we restore both into a scratch environment and run a chart through a restored connection, which proves the key matches.

  5. Upgrade with a snapshot ready

    Each release runs superset db upgrade and superset init. Some migrations do not reverse cleanly, so the rollback kept ready is a snapshot taken just before. The Pilae Agent tests the upgrade on a copy, waits for your approval, applies it in your window and records it in the console.

What Apache Superset is, and who runs it

Self-hosted Apache Superset for SQL analytics and dashboards

Apache Superset is a business intelligence platform from the Apache Software Foundation. Analysts write queries in SQL Lab, define metrics once on a dataset, and build charts and dashboards on top, geospatial maps included. It reads from almost any SQL database with a Python driver and a SQLAlchemy dialect: PostgreSQL, ClickHouse, Trino, Snowflake, BigQuery and others.

It is built for teams that already write SQL and publish to the rest of the organisation. If most of your readers will ask simple questions without writing any, Metabase is easier to adopt and lighter to run, and we say so. Superset belongs next to the warehouse it reads, so that is where we put it: on your own hardware, or in the Pilae Cloud region that already holds the data, one of 12, six of them in Switzerland and the EU.

Apache Superset in production: workers, beat and the secret key

Superset is several processes, not one. The web app answers people, Celery workers run long queries and take the screenshots for reports, and a single beat decides when. Redis sits between them. Their failures are quiet: a second beat can send every report twice, and a worker that meets the sign-in page fails the report. The default SQLite file does not work once workers share the metadata database, so we start on PostgreSQL. We run one beat and point the worker at the internal address. Probes check the web app every 60 seconds, and an engineer is alerted when a report schedule stops completing.

The metadata database is backed up daily, encrypted, to an offsite location in your chosen country. The secret key is escrowed beside it, because without it the stored connection passwords cannot be decrypted. Sign-on goes through Keycloak or your own IdP, and the app answers only on your private network. Upgrades run database migrations and reset the built-in roles to their defaults, so custom permissions live in custom roles. The Pilae Agent runs each upgrade on a copy first and applies it in your window after you approve it.

Apache Superset licence and editions

With nothing held back for a paid edition, there is no seat count to track and no licence to hold in your name. We operate the Apache release, on infrastructure you choose. Our operation is priced on request. Talk to us about the workbooks you want to bring across from Tableau.

The code in the apache/superset repository is Apache-2.0; the two fonts it bundles, Inter and Fira Code, are under the SIL Open Font License. There is no enterprise directory and no paid build: single sign-on, row-level security, SQL Lab, alerts and reports are all in the release we deploy. Preset, the company founded by the creator of Superset, sells hosted and self-managed editions of its own under a separate contract with Preset. You do not need them for anything on this page. Apache Superset and its logo are trademarks of the Apache Software Foundation, which does not endorse this service.

Apache Superset system requirements

Before anything is deployed, this is what has to exist. We size it with you in the first session, and we say so when your own hardware is already enough.

CPU and memory
4 vCPU · 8 GBOur starting point for the web app, workers and beat on one machine. Screenshot reports run a headless Chromium inside the worker, and that is what wants the memory.
Metadata database
PostgreSQL 14–17Holds charts, dashboards, users, query history and the encrypted connection passwords. Upstream supports PostgreSQL and MySQL; the default SQLite file is not for production.
Cache and queue
RedisCaches chart results and carries the Celery queue for asynchronous queries, alerts and reports. Out of the box the broker is a SQLite file, which upstream says to replace in production.
Secret key
SECRET_KEY, 42 random bytesSuperset refuses to start on the default key. It signs sessions and encrypts the stored database passwords, so a restore without the same key cannot decrypt them.
Data sources
Read-only accounts + driverEach engine needs its Python driver in the image, and each connection an account that can only read the schemas people need.

Migrating from Tableau to Apache Superset

Tableau workbooks do not open in Superset and there is no built-in importer, so dashboards are rebuilt. The charts are the quick part; the data is the slow one. Superset has no extract engine: it queries the database live, with a result cache in front, so anything that only exists in a Tableau extract, or in a blend of two sources, has to land in a SQL database first. Calculated fields become metrics and calculated columns on Superset datasets, written in SQL and shared by every chart built on them. A dashboard laid out to the pixel will not look the same, and a few Tableau chart types have no Superset counterpart. Those go on a list the owners see before any rebuild is scheduled.

  1. Start from the workbooks people open

    Tableau Server and Tableau Cloud record who opened what. The last three months of it pick the workbooks worth rebuilding, and the others lapse with the licence instead of being carried across.

  2. Land the extracts in a database

    Data that only lives in Tableau extracts or cross-source blends is loaded into a SQL database Superset can reach, often a PostgreSQL we already run for you, with a scheduled load in place of the extract refresh.

  3. Rebuild calculations as datasets

    Calculated fields and LOD expressions become SQL metrics, calculated columns or virtual datasets, checked figure by figure against the Tableau views they replace.

  4. Move the subscriptions last

    Email subscriptions become Superset reports once owners have checked a full reporting cycle in both tools. Tableau is kept read-only until the first month-end has closed on Superset.

One configuration file for the app, the workers and the beat

# superset_config.py (excerpt), mounted read-only into app, worker and beat
import os
from celery.schedules import crontab
from flask_appbuilder.security.manager import AUTH_OAUTH
from superset.config import CeleryConfig

SECRET_KEY = os.environ["SUPERSET_SECRET_KEY"]  # escrowed with the backups
SQLALCHEMY_DATABASE_URI = os.environ["SUPERSET_METADATA_URI"]

REDIS = "redis://redis.internal:6379"
DATA_CACHE_CONFIG = {
    "CACHE_TYPE": "RedisCache",
    "CACHE_REDIS_URL": f"{REDIS}/1",
    "CACHE_KEY_PREFIX": "acme_data",
    "CACHE_DEFAULT_TIMEOUT": 3600,
}

class AcmeCeleryConfig(CeleryConfig):
    broker_url = f"{REDIS}/0"
    result_backend = f"{REDIS}/0"
    beat_schedule = {
        **CeleryConfig.beat_schedule,
        "prune_query": {
            "task": "prune_query",
            "schedule": crontab(minute=0, hour=0, day_of_month=1),
            "kwargs": {"retention_period_days": 180},
        },
    }

CELERY_CONFIG = AcmeCeleryConfig
FEATURE_FLAGS = {"ALERT_REPORTS": True, "PLAYWRIGHT_REPORTS_AND_THUMBNAILS": True}
WEBDRIVER_BASEURL = "http://superset-app:8088/"
WEBDRIVER_BASEURL_USER_FRIENDLY = "https://bi.acme.internal/"

AUTH_TYPE = AUTH_OAUTH
OAUTH_PROVIDERS = [{
    "name": "keycloak",
    "token_key": "access_token",
    "remote_app": {
        "client_id": "superset",
        "client_secret": os.environ["OIDC_CLIENT_SECRET"],
        "api_base_url": "https://sso.acme.internal/realms/acme/protocol/openid-connect",
        "access_token_url": "https://sso.acme.internal/realms/acme/protocol/openid-connect/token",
        "authorize_url": "https://sso.acme.internal/realms/acme/protocol/openid-connect/auth",
        "client_kwargs": {"scope": "openid email profile"},
    },
}]
AUTH_USER_REGISTRATION = True
AUTH_USER_REGISTRATION_ROLE = "Gamma"
# group paths from the "groups" claim, set by a Keycloak group-membership mapper
AUTH_ROLES_MAPPING = {"/bi-analysts": ["Alpha", "sql_lab"], "/bi-admins": ["Admin"]}
AUTH_ROLES_SYNC_AT_LOGIN = True
ENABLE_PROXY_FIX = True
An example superset_config.py for acme, trimmed. The same file is mounted into the web app, every worker and the single beat. Secrets come from the environment, reports render against the internal address so the worker never meets the sign-in page, and the extra beat entry prunes SQL Lab history, which Superset otherwise keeps indefinitely.

What Pilae is responsible for

A pinned version

A version we have run, not whatever latest resolves to that day.

A runbook

What it depends on, how it fails, what to do about it. In your repository.

A restore drill

Backups restored on a schedule. A backup nobody has restored is a file.

A patch window

Security updates in a window you agreed, with a rollback ready.

Someone watching

Every endpoint probed on the minute. An alert reaches a person, not a dashboard nobody opens.

Where it runs
zur1, fra1, fal1, gra1, ams1, hel1, lon1, ash1, hil1, sin1, tok1, syd1, on-premZurich, Frankfurt, Falkenstein, Gravelines, Amsterdam, Helsinki, London, Ashburn, Hillsboro, Singapore, Tokyo, Sydney, Your own hardware
Who holds the credentials
You do. Ours are separate, named, logged and revocable with one command. We ask before anything changes outside an agreed window.
If you leave
The machine, the data, the compose files and the runbook are already yours. Nothing stops when our access does.

What drives the price of running Apache Superset

Pricing is on request: a fixed price for onboarding, then a monthly price for Apache Superset, quoted in writing within five business days. The plans set what every deployment includes; these are the inputs the quote is built from.

Instance size
The CPU, memory and, where a model runs, the GPUs the app needs for your users and your data.
High availability
One machine with tested restores, or a replicated setup that keeps serving when a node fails.
Storage and backups
How much data it holds, how long backups are kept, and point-in-time recovery for its database.
Plan and support
Essential, Business or Enterprise: support hours, response times in the contract and how often we review the service with you.
Region
Your own hardware, where the infrastructure is already yours, or a Pilae Cloud region, where it is passed through at cost plus a fixed margin.
Sign-on and integrations
Single sign-on, directory sync, mail relays and the other systems the app has to reach.

Apache Superset: common questions

Is Apache Superset open source?

Yes. It is Apache-2.0 and a top-level project of the Apache Software Foundation. There is no enterprise edition of the software itself: sign-on, row-level security, alerts and reports are in the release we deploy. Preset sells hosted and self-managed editions of its own, which are separate products you do not need.

Superset or Metabase: which should we choose?

Metabase, if most of the people using it will ask simple questions without writing SQL: it is easier to adopt and easier to run, as one application and a database. Superset, if a data team writes SQL, wants a wider range of charts and a shared semantic layer, and needs row-level security without a paid licence. It costs more to operate: workers, a beat, Redis and a headless browser.

Where do our data and dashboards live?

Your data stays in your databases, which Superset queries live. Chart and dashboard definitions, users and query history sit in its PostgreSQL metadata database, and cached results in Redis, all on the machine beside your warehouse: your own hardware, or a dedicated machine in the Pilae region that holds the data, in ISO 27001-certified datacentres.

Can people run any SQL they like against our databases?

Only people whose role includes SQL Lab and access to that database. Upstream is clear that Superset is not a database firewall, so we do not rely on it as one: every connection uses an account that can only read the schemas it needs, and heavy analysis points at a replica. Row-level security filters apply to charts and dashboards but not, by default, to SQL Lab, so people limited to some rows do not get SQL Lab on that database.

Why would a scheduled report not arrive?

Because the worker could not take the screenshot. Reports are screenshots taken by a headless browser in the worker, so they fail when the worker cannot reach Superset, lands on a sign-in page, or has no browser installed. A second running beat causes the opposite fault: each report can go out twice. We configure the worker to render against the internal address, run one beat, and alert when a report schedule stops completing.

Also in workflow and data

Back to apps

Bring us your Apache Superset. We will tell you what it takes.

Thirty minutes on the deployment you already have, or the one you are about to start.