Skip to content

product / infrastructure

Infrastructure

Know the state of the machines before the users do: the servers, the addresses they answer on and the jobs that should run by themselves.

The fleet at a glance

A small agent on each server reports processor, memory, disk and network, plus containers, system services and databases. A disk that is filling up warns you before it is full.

What is included
  • CPU, memory, disk, network, Docker, systemd, GPU and sensors
  • Databases and backups of each server
  • Predictive disk alerts
  • One year of history
Servers

web-0110 s ago

CPU34%

Memory58%

Disk48%

db-primary10 s ago

CPU67%

Memory67%

Disk86%

workers10 s ago

CPU25%

Memory44%

Disk39%

What is down, and since when

HTTP monitors per project and environment check your addresses and keep an honest history of incidents, with SLA over 24 hours, 7 days, 30 days and one year.

What is included
  • Incident timeline with start and resolution
  • Public status page, opt-in per monitor
  • Monitored addresses are checked against internal-network abuse
Uptime
drogheria.exampleUp

last 24 hours99.98%

api.drogheria.exampleSlow

last 24 hours99.71%

Payments · webhookDown for 3 h

last 24 hours86.11%

The job that did not run

A scheduled job tells CloseYourIt each time it runs. When the check-in does not arrive, or arrives as failed, you get the alert that silence never gives you.

What is included
  • Expected check-ins per job
  • Alert on a missed or failed run
  • One HTTP call from any scheduler
Scheduled jobs

nightly-backup

0 3 * * *

On time

send-invoices

*/15 * * * *

On time

sync-catalog

0 * * * *

Missed

An alert when a number crosses the line

A rule watches one number against a threshold, for the whole organization. When it fires, it reaches each person the way they chose, and stays silent during their quiet hours.

What is included
  • Threshold rules at organization level
  • In-app notification center
  • Email delivery with per-user preferences
  • Quiet hours per user
Alert rules

CPU above 85% for 5 minutes

db-primary

Firing

Disk full within 6 days

db-primary

Soon

Error rate above 2%

production

OK

p95 above 800 ms

production

OK

Quiet hours 22:00–07:00

The other areas