Choose what to monitor
Which checks exist, how thresholds work, and how to pick which disks are watched.
A freshly enrolled server reports numbers but pages nobody. Alerting is something you turn on per check, deliberately, because a default threshold on somebody else's server is a guess.
The checks
| Check | What it measures | Default threshold | Opens an incident |
|---|---|---|---|
| CPU usage | Percentage of CPU time in use, averaged over the reporting window | at or above 90% | Yes |
| Load average per core | One-minute load average divided by the core count | at or above 2 | Yes |
| Disk usage | Percentage of a filesystem in use | at or above 85% | Yes |
| Inode usage | Percentage of a filesystem's inodes in use | at or above 85% | Yes |
| Uptime | Seconds since boot | — | No |
| Memory usage | Percentage of RAM in use, excluding buffers and cache | at or above 90% | Yes |
| Swap usage | Percentage of swap in use | at or above 50% | Yes |
A check appears on the server's page only once the agent has said it can perform it. That is why the list is shorter on an older agent, and why the page says "Needs agent x.y.z" rather than silently omitting something you were expecting.
Thresholds
Open the server, choose a check, and set a threshold. Three things to decide:
The value. Every check that alerts today does so when a number goes above one — 90% disk, 85% memory. The direction is configurable because later checks will not all be that shape, but nothing shipping now alerts on a value being too low.
Uptime is the one check in the list that never alerts, whatever you set. A reboot shows up as the value resetting, and a server that has gone away is already covered by it going quiet, so alerting here as well would page you twice for one event.
How long it has to hold. A threshold has to be breached on two consecutive reports before an incident opens. One minute of 95% CPU is a backup running. Ten minutes of it is a problem. This is not configurable yet and the confirmation is always two reports.
Severity, which decides what happens:
| Severity | Effect |
|---|---|
| Critical | Opens an incident and notifies you, like a down website |
| Warning | Opens an incident and notifies you, marked lower |
| Security | Recorded on the server page as a finding; no incident, no notification |
| Recommendation | Recorded as a finding; no incident, no notification |
| Info | Recorded for context only |
Security and Recommendation exist for the configuration and hardening checks that are coming. Nothing in the current list uses them.
Critical and Warning use the notification channels the team already has, so a full disk reaches you the same way a down site does. There is no separate place to configure that.
Disks: which ones get watched
Most servers have more filesystems than you would ever want alerts about. A typical Ubuntu box reports twenty-odd, most of them kernel bookkeeping and snap images. So you choose how that list is handled.
Automatic (the default)
We watch the real filesystems and leave out the ones that cannot run out of space or are not yours to manage. Specifically:
These filesystem types are excluded outright:
autofs, binfmt_misc, bpf, cgroup*, configfs, debugfs, devpts, devtmpfs, efivarfs, fuse.gvfsd-fuse, fuse.snapfuse, fusectl, hugetlbfs, mqueue, nsfs, overlay, overlayfs, proc, procfs, pstore, ramfs, securityfs, squashfs, sysfs, tmpfs, tracefs
Three of those are worth explaining, because they are the ones that would otherwise generate the false alarms:
squashfsis a snap package image. Every one is exactly 100% full by design, and a stock Ubuntu install has dozens. Alerting on these is the single most common false positive in host monitoring.tmpfsandramfs—/run,/dev/shm— live in memory, not on a disk. If they fill, the memory check is what should tell you.overlayis a container layer. A busy Docker host has one per container, appearing and vanishing as containers come and go.
The rest are kernel bookkeeping and hold none of your data.
Also excluded:
- Anything mounted read-only, whatever its type. It cannot grow.
- Duplicates. A bind mount and its origin are the same disk; we watch it once, under the shortest path.
What is left is watched automatically — including a volume you attach next month, which is the point: a new disk gets monitored without anyone remembering to come back here.
You can narrow it further with include and exclude patterns, using globs like /mnt/* or /var/lib/docker/*. Exclude wins where both match.
Manual
Pick specific mounts from the list and nothing else is watched. A new disk will appear in the list as Available but stays unmonitored until you say otherwise.
Choose this when you have a lot of mounts you genuinely do not care about, or when you want the monitored set to be a decision rather than a consequence.
Either way, you can see the whole list
The server's disk page shows every filesystem the machine reported, each marked Monitored, Available, or excluded with the reason — read-only, virtual filesystem, snap image, duplicate of /, your exclude rule. If you disagree with a decision we made, you can see that we made it, and why.
A check that stops reporting
If the agent stops being able to perform a check — a monitored disk gets unmounted, say — the check is marked stale rather than treated as healthy. An incident that was already open stays open: an unmounted disk is not a fixed disk, and closing the incident because the number went away is how a real problem gets quietly dropped.
Last updated 2 October 2026
Still stuck?
If this did not answer your question, tell us and we will fix the page as well as answer you.