Managed Paperless-ngx hosting

BusinessOn-prem or sovereign site

OCR, filing and full-text search for scanned and emailed documents, run in Switzerland, the EU or your own datacentre so invoices, contracts and letters stay on storage you control. Pilae runs it on your own servers, or in Zurich, Switzerland, and eleven other Pilae Cloud regions.

Talk to us about Paperless-ngx

Licence
GPL-3.0
Runs on
Your own hardware, or any of twelve Pilae regions — six of them in Switzerland and the EU
Upgrades
Pinned, tested against your configuration, applied in your window
Upstream
docs.paperless-ngx.com

Running Paperless-ngx in production: what it takes

  1. Size OCR to your post, then deploy

    A version we have run, on a dedicated machine with encrypted disks: PostgreSQL rather than the default SQLite file, a Redis-compatible broker, and Tika and Gotenberg for Office files and email. OCR languages and worker counts are set to match your own post.

  2. Give every intake path an owner

    The scanner share, each mailbox and the API get a workflow that assigns the owner, group permissions and document type on arrival. Without one, a file consumed from the folder has no owner and every user can open it.

  3. Turn off password login in the web app

    OpenID Connect to Keycloak or your own IdP such as Microsoft Entra ID, with groups synced from the token and password login off in the web app. The API and the Django admin still accept local credentials, so the instance answers only on your private network.

  4. Back up files and database together

    The media volume and PostgreSQL go offsite at the same moment, daily, encrypted, while nothing is being consumed. Once a month we restore both into a scratch instance on the same pinned version and run the sanity checker, which compares every original and archive file with its stored checksum.

  5. Upgrade on a copy of your archive

    Each release runs its database migrations when the container starts, and the search index rebuilds itself when its format changes. The last major release replaced OCR, consumer and database settings and could only be reached from the final release of the line before. The Pilae Agent tests each upgrade on a copy of your archive, waits for your approval and applies it in your window, with the backup taken just before as the rollback.

What Paperless-ngx is, and who runs it

Self-hosted Paperless-ngx for scanned and emailed documents

Paperless-ngx turns paper and PDFs into an archive you can search. Documents arrive from a watched folder, a mailbox, the web upload or the API. It runs OCR with Tesseract through OCRmyPDF, keeps the original file and, for scans, stores a searchable PDF/A copy beside it. Matching rules, or a classifier trained on what you have already filed, assign the correspondent, document type and tags, and full-text search covers the lot.

It grew out of personal and small-office use, and the design still reflects that. The project is community-supported, the successor to Paperless and Paperless-ng, and its own README calls a local server in your own home the safest place to run it. Three things matter at organisation scale: permissions are set per document owner, user and group rather than per cabinet; there are no approval workflows, check-out or retention locks; and files sit unencrypted on an ordinary filesystem. It is what teams move to from DocuWare or M-Files when what they use every day is capture, filing and search rather than workflows.

Paperless-ngx in production: scanners, mailboxes and permissions

It runs on a dedicated machine with encrypted disks, in your own building or in a Pilae Cloud region, where six of our 12 regions are in Switzerland and the EU, and it answers only on your private network. Share links are public URLs that need no login, so they reach nobody outside until you choose to publish them through the gate on port 443.

Scanners drop into a share with one folder per department, and mail rules read the shared mailboxes. A workflow on each intake path sets the owner and group permissions, because a file consumed from the folder otherwise has none. Sign-in goes through Keycloak or Entra ID, with groups taken from the token.

Probes check the instance every 60 seconds, and a failed consumption task alerts an engineer. Without that alert, a bad attachment is skipped in silence: the message is left where it was but recorded as processed, so Paperless-ngx does not try it again. Backups of the database and the media volume go daily, encrypted, to an offsite location in your chosen country, and the monthly restore drill ends with the sanity checker. Upgrades go through the Pilae Agent: tested on a copy, approved by you, applied in your window.

The OCR text is available through the REST API, so it can feed retrieval in Open WebUI, provided the document permissions travel with it. Paperless-ngx has its own opt-in AI features, and we point them at Ollama on your network rather than at a hosted API.

Paperless-ngx licence and support

Every feature on this page is in the build we deploy, unmodified. Pricing for our operation is on request. Talk to us about the post and the archive you want to move.

Paperless-ngx is GPL-3.0, with no paid edition, no enterprise directory and no feature behind a licence key. The project calls itself community-supported and points users to GitHub Discussions and a Matrix room for help, so there is no support contract to buy from it, and the licence disclaims any warranty. Running it for your own organisation places no obligation on you: the GPL attaches to distributing the software, not to using it, and unlike the AGPL it has no network clause. The operating commitment, patching and restores included, is ours and is written into the contract with you.

Paperless-ngx system requirements

Before anything is deployed, this is what has to exist. We size it with you in the first session, and we say so when your own hardware is already enough.

CPU and memory
4 vCPU · 8 GBOur starting size with Tika and Gotenberg beside it; upstream publishes no minimum. OCR decides the size: task workers times threads per worker should not exceed the core count.
Database
PostgreSQL 14+Recommended upstream for new installations. The default is a SQLite file in the data volume, and MariaDB is supported with caveats about case sensitivity.
Message broker
Redis or ValkeyCarries the task queue, scheduled tasks included. While it is down, nothing is consumed and no mail is fetched.
Office files and email
Tika + GotenbergOptional upstream, and needed for Word, Excel, PowerPoint and .eml files. Gotenberg renders email with JavaScript off and only local files allowed, so a message cannot load a tracking pixel.
Sign-in
OIDC via django-allauthOpenID Connect to Keycloak or Entra ID, with groups synced from a claim onto Paperless-ngx groups of the same name. Password login can be turned off for the web app but not for the API or the Django admin, so those stay on the private network.

Migrating from DocuWare to Paperless-ngx

Paperless-ngx has no DocuWare importer, and its own importer reads only what its exporter wrote. The route out is docXporter, a DocuWare add-on licensed separately, which writes each document in its original format with its index data. Each file then goes in through the Paperless-ngx API with its date and index values already set. Paperless-ngx stores every file as it arrives and runs OCR on the ones without a text layer. Stamps and annotations come across only if docXporter burns them into a PDF copy. DocuWare workflows and the old audit trail stay behind: Paperless-ngx workflows act on intake, on a change or on a date, and have no approval steps. Mapping file cabinets and index fields takes longer than the copy.

  1. Map the cabinets and index fields

    Which file cabinets are still filed into, which are only kept for their retention period, and what each index field becomes in Paperless-ngx: a document type, a correspondent, a tag or a custom field. This mapping decides the rest.

  2. Export with docXporter

    Cabinet by cabinet, with each export counted against DocuWare before anything is loaded. docXporter needs its own DocuWare licence, which we raise at the start rather than at the end.

  3. Load through the API

    Each file is posted with its title, date, archive serial number and mapped values. A workflow on API intake sets the owner and group permissions, and the classifier learns from the migrated archive, so new post is filed the way the old archive was.

  4. Switch the scanners and mailboxes last

    Scanners move to the consumption share and mail rules take over the invoice mailbox. DocuWare stays read-only until one month-end has closed on Paperless-ngx, and its licence lapses after that.

A Paperless-ngx environment file for acme

# docker-compose.env (excerpt) · acme-paperless · zur1
PAPERLESS_URL=https://paperless.acme.internal
PAPERLESS_SECRET_KEY_FILE=/run/secrets/paperless_secret_key
PAPERLESS_TIME_ZONE=Europe/Zurich

# PostgreSQL on the private network, TLS required
PAPERLESS_DBENGINE=postgresql
PAPERLESS_DBHOST=db-01.acme.internal
PAPERLESS_DBPASS_FILE=/run/secrets/paperless_db
PAPERLESS_DB_OPTIONS=sslmode=require
PAPERLESS_REDIS=redis://broker:6379

# French, German and English ship with the image
PAPERLESS_OCR_LANGUAGE=fra+deu+eng
PAPERLESS_TASK_WORKERS=2
PAPERLESS_THREADS_PER_WORKER=2

# Word, Excel and .eml through Tika and Gotenberg
PAPERLESS_TIKA_ENABLED=true
PAPERLESS_TIKA_ENDPOINT=http://tika:9998
PAPERLESS_TIKA_GOTENBERG_ENDPOINT=http://gotenberg:3000

# consume/ is the scanner share, one folder per department
PAPERLESS_CONSUMER_RECURSIVE=true
PAPERLESS_CONSUMER_SUBDIRS_AS_TAGS=true

# Sign-in through Keycloak only, groups from the token
PAPERLESS_APPS=allauth.socialaccount.providers.openid_connect
PAPERLESS_SOCIALACCOUNT_PROVIDERS_FILE=/run/secrets/paperless_oidc
PAPERLESS_SOCIAL_AUTO_SIGNUP=true
PAPERLESS_SOCIAL_ACCOUNT_SYNC_GROUPS=true
PAPERLESS_DISABLE_REGULAR_LOGIN=true
PAPERLESS_REDIRECT_LOGIN_TO_SSO=true

# Originals survive an emptied trash; AI stays off until asked for
PAPERLESS_EMPTY_TRASH_DIR=../media/trash
PAPERLESS_AI_ENABLED=false
An example docker-compose.env excerpt. Secrets come from files. OCR runs only in the languages acme's post arrives in, because each extra language costs Tesseract CPU time. Each scanner folder becomes a tag, and sign-in goes through Keycloak only. Owners and group permissions are not in this file: workflows set them, and workflows live in the database and are restored with it.

What Pilae is responsible for

A pinned version

A version we have run, not whatever latest resolves to that day.

A runbook

What it depends on, how it fails, what to do about it. In your repository.

A restore drill

Backups restored on a schedule. A backup nobody has restored is a file.

A patch window

Security updates in a window you agreed, with a rollback ready.

Someone watching

Every endpoint probed on the minute. An alert reaches a person, not a dashboard nobody opens.

Where it runs
zur1, fra1, fal1, gra1, ams1, hel1, lon1, ash1, hil1, sin1, tok1, syd1, on-premZurich, Frankfurt, Falkenstein, Gravelines, Amsterdam, Helsinki, London, Ashburn, Hillsboro, Singapore, Tokyo, Sydney, Your own hardware
Who holds the credentials
You do. Ours are separate, named, logged and revocable with one command. We ask before anything changes outside an agreed window.
If you leave
The machine, the data, the compose files and the runbook are already yours. Nothing stops when our access does.

What drives the price of running Paperless-ngx

Pricing is on request: a fixed price for onboarding, then a monthly price for Paperless-ngx, quoted in writing within five business days. The plans set what every deployment includes; these are the inputs the quote is built from.

Instance size
The CPU, memory and, where a model runs, the GPUs the app needs for your users and your data.
High availability
One machine with tested restores, or a replicated setup that keeps serving when a node fails.
Storage and backups
How much data it holds, how long backups are kept, and point-in-time recovery for its database.
Plan and support
Essential, Business or Enterprise: support hours, response times in the contract and how often we review the service with you.
Region
Your own hardware, where the infrastructure is already yours, or a Pilae Cloud region, where it is passed through at cost plus a fixed margin.
Sign-on and integrations
Single sign-on, directory sync, mail relays and the other systems the app has to reach.

Paperless-ngx: common questions

Is Paperless-ngx open source?

Yes. It is GPL-3.0, with no paid edition and no features held back. It is the official successor to Paperless and Paperless-ng, maintained by a team of community contributors. The GPL places no obligation on an organisation that runs it for itself; its source terms apply when the software is distributed.

Where do our documents live?

On one dedicated machine, in Zurich or another Pilae region you choose, in ISO 27001-certified datacentres, or on your own hardware. Originals and PDF/A copies are plain files on disk and the metadata is in PostgreSQL. Upstream stores documents unencrypted and warns against running it on a host you do not trust, so the disks are encrypted, and for records that must never leave your building we run it on your hardware.

Can Paperless-ngx replace DocuWare?

For capture, OCR, filing and search, yes. It has no approval workflows, no check-out and no retention locks, and permissions are set per document owner, user and group rather than per file cabinet. If invoices pass through a multi-step approval in DocuWare today, that step moves to your ERP or to a workflow tool that calls the Paperless-ngx API, and we say which before the migration starts.

How do permissions work across departments?

Per document. Each one can have an owner, and view or edit rights can be granted to users and groups. Files consumed from the scanner folder have no owner by default, which makes them visible to everyone, so a workflow on each intake path assigns the owner and the right group. Superusers see every document, so administrators use an ordinary account day to day.

Does Paperless-ngx send our documents to an AI service?

Not unless you turn it on. The built-in suggestions come from a local classifier trained on the documents you have already filed. The LLM features — suggestions, similar documents and a document chat — are off by default. If you want them, we point them at Ollama on your network. Remote OCR, which sends documents to Microsoft Azure, is also off by default, and we leave it off.

Also in business

Back to apps

Bring us your Paperless-ngx. We will tell you what it takes.

Thirty minutes on the deployment you already have, or the one you are about to start.