Monitoring for self-hosted apps, with engineers on the alerts

Every app probed every 60 seconds, every machine measured continuously, and every alert sent to a Pilae engineer, day and night.

Start with one workloadBook a 30-minute briefing

Probes
Every 60 seconds, per app
Metrics
CPU, memory, disk and network, per machine
Alerts
To Pilae engineers, around the clock
Report
Monthly, per app and per machine

We know before your users do

Most teams learn about an outage from a colleague. With Pilae, a probe fails and an engineer is alerted before the first ticket arrives.

Probes every minute

Each app is checked every 60 seconds: it answers, it answers correctly, it answers fast, and its certificate is valid.

Machine metrics

CPU, memory, disk and network for every machine, kept as history so a disk filling up is caught weeks ahead.

Alerts to engineers

A failed probe or a crossed threshold alerts a Pilae engineer, at any hour, on every plan.

Status you can see

The health of every app and machine is shown in the console, with incidents and their updates as they happen.

Tied to the agent

A failing probe opens a task for the Pilae Agent, which proposes a fix with a backup and a rollback.

A monthly report

Availability against your commitment, incidents, changes applied, backups and restore tests, in one document.

Monitoring that someone actually reads

Plenty of self-hosted estates have monitoring. Fewer have someone whose job is to read the alert. With Pilae, every alert reaches a Pilae engineer, on every plan. Response times follow your plan, and around-the-clock incident response is part of Enterprise.

Probes run every 60 seconds against each app. They check that it answers, that it returns what it should, that it answers in time and that its certificate is valid. Machine metrics track CPU, memory, disk and network, so a disk filling up or a memory leak shows weeks before it causes an outage.

Built on open-source tools

The probes run on Gatus and machine metrics on Beszel, both open-source projects. They run in your environment, on your premises or in Pilae Cloud, and report into the console. Your team sees the same health, history and incidents as our engineers.

From alert to fix

A failing probe opens a task for the Pilae Agent. It checks recent changes, proposes a fix or a rollback, and takes a backup before anything runs. An engineer decides with you, and the whole run is recorded. Our managed operations service sets out who does what during an incident.

Reports your management can read

Each month you receive a report per app: availability against your 99.9% commitment, incidents and their causes, changes applied, and the backups and restore tests of the month. Every major incident gets its own report. The service level agreement defines availability, service credits and response times. To agree what we watch, book a scoping call.

What happens when something breaks

  1. Detect

    A probe fails or a metric crosses its threshold. Checks repeat to rule out a blip.

  2. Alert

    A Pilae engineer is alerted at once, with the failing check, the machine and recent changes attached.

  3. Respond

    The engineer investigates within the response times of your plan, and the Pilae Agent proposes a fix or a rollback for approval.

  4. Inform

    You are told what happened, what was done and whether any data was affected, within the times your plan sets.

  5. Review

    The incident, its cause and its fix appear in the console and in the monthly report.

What your contract includes

Probe interval
60 seconds for every app we operate.
Alerting
Monitoring and alerts around the clock on every plan. Around-the-clock incident response on Enterprise.
Response times
Set per plan and per severity in your contract.
Availability
A 99.9% monthly availability commitment with service credits, on every plan.
Reporting
A monthly report per app, and a post-incident report for every major incident.

The apps behind it

Questions

How often does Pilae check my apps?

Every 60 seconds. Each probe checks that the app answers, returns the expected response, answers within a set time and has a valid certificate.

Who receives the alerts?

Pilae engineers, around the clock, on every plan. Around-the-clock incident response is part of the Enterprise plan. On Essential, support is in business hours. Your team can also receive alerts if you wish.

Do I get a report on availability?

Yes. Each month you receive a report with availability against your 99.9% commitment, incidents, changes applied, backups taken and restore tests run.

Which tools does Pilae use for monitoring?

Open-source tools: Gatus for the 60-second probes and Beszel for machine metrics. Both run in your environment, and their results appear in the Pilae console.

How is monitoring done for an air-gapped deployment?

The probes and metrics run inside your perimeter. Alerts go to your on-site team, and to Pilae engineers only through a channel your security policy allows, if any.

Related

Put engineers on your alerts.

Thirty minutes with an engineer to agree what we probe, what we alert on and who hears about it.