Skip to content

gatus health checks

gatus runs every health check from the swarm, with its UI on port 8085. it replaced uptime kuma because its config lives in git.

compose.yml, 123 lines, 6 notes

each in the code opens a note on that line. download compose.yml

services:
  git-sync:
    image: registry.k8s.io/git-sync/git-sync:v4.7.1
    user: "1000:1000"
    environment:
      - TZ=America/Los_Angeles
      - GITSYNC_REPO=ssh://git@ssh.github.com:443/<you>/<repo>.git
      - GITSYNC_REF=deploy/swarm/gatus
      - GITSYNC_PERIOD=60s
      - GITSYNC_ROOT=/git
      - GITSYNC_LINK=repo
      - GITSYNC_SSH_KEY_FILE=/dev/shm/key
      - GITSYNC_SSH_KNOWN_HOSTS_FILE=/known_hosts
      - GITSYNC_ADD_USER=true
      - GITSYNC_FILTER=blob:none
      - GITSYNC_SPARSE_CHECKOUT_FILE=/sparse-checkout
      - GITSYNC_HTTP_BIND=:8080
    entrypoint:
      - /bin/sh
      - -c
      - |
        umask 077
        printf '%s\n' "$$(cat /run/secrets/gitsync_ssh_key_v1)" > /dev/shm/key
        [ -s /dev/shm/key ] || { echo "git-sync: could not write the SSH key to /dev/shm" >&2; exit 1; }
        mkdir -p /tmp/.ssh && cp /ssh_config /tmp/.ssh/config || { echo "git-sync: could not install the ssh config" >&2; exit 1; }
        exec /git-sync
    volumes:
      - type: bind  # (1)!
        source: /usr/share/zoneinfo
        target: /usr/share/zoneinfo
        read_only: true
      - type: volume
        source: gitrepo_v2
        target: /git
        volume:
          nocopy: true
    configs:
      - source: gatus_known_hosts_v2
        target: /known_hosts
      - source: gatus_ssh_config_v2
        target: /ssh_config
      - source: gatus_sparse_checkout_v1
        target: /sparse-checkout
    secrets:
      - gitsync_ssh_key_v1
    networks:
      discovery:
        aliases:
          - gatus-git-sync
    deploy:
      mode: replicated
      replicas: 1

  gatus:
    image: ghcr.io/scyto/gatus:v5.36.0-12-secrets-backup@sha256:61a8988f286fed144c91cdfd1acb7d429d2ba0fecc47007c95df2e0a0c840518
    environment:  # (2)!
      - TZ=America/Los_Angeles
      - GATUS_CONFIG_PATH=/git/repo/stacks/swarm/gatus/config
      - GATUS_LOG_LEVEL=INFO
      - PROXMOX_TOKEN_FILE=/run/secrets/proxmox_api_token_v1
      - GATUS_SQLITE_BACKUP_PATH=/data/backup/gatus.db
    volumes:
      - type: bind
        source: /usr/share/zoneinfo
        target: /usr/share/zoneinfo
        read_only: true
      - gitrepo_v2:/git:ro
      - data:/data
    secrets:
      - proxmox_api_token_v1
    networks:
      - discovery  # (3)!
      - traefik-api
    extra_hosts:
      - mydomain1.com:192.168.1.45
    ports:
      - 8085:8080
    deploy:
      mode: replicated
      replicas: 1
      labels:  # (4)!
        - homepage.group=Monitoring
        - homepage.name=Gatus
        - homepage.icon=gatus.png
        - homepage.href=https://gatus.mydomain.com
        - homepage.description=Health checks, config in git
        - homepage.widget.type=gatus
        - homepage.widget.url=http://192.168.1.45:8085

volumes:  # (5)!
  gitrepo_v2:  # (6)!
    driver: local
    driver_opts:
      type: none
      device: "/mnt/docker-cephFS/gatus_git_v2"
      o: bind
  data:
    driver: local
    driver_opts:
      type: none
      device: "/mnt/docker-cephFS/gatus_data"
      o: bind

networks:
  discovery:
    external: true
  traefik-api:
    external: true

configs:
  gatus_known_hosts_v2:
    file: ./known_hosts
  gatus_ssh_config_v2:
    file: ./ssh_config
  gatus_sparse_checkout_v1:
    file: ./sparse-checkout

secrets:
  proxmox_api_token_v1:
    external: true

  gitsync_ssh_key_v1:
    external: true
  1. neither the git-sync image nor the gatus image has time zone data, so TZ alone changes nothing. both services mount the host's /usr/share/zoneinfo read-only.
  2. turns on the fork's hourly VACUUM INTO copy of the database, for the cephfs backup. the fork says why it isn't a sidecar.
  3. the discovery overlay, where gatus reaches the swarm's read-only docker socket proxy as dockerproxy:2375 and both git-sync sidecars by their aliases. none of them has a port on the lan.
  4. the tile on the homepage dashboard, under deploy.labels because homepage reads labels from the swarm's service specs. the widget reads gatus through the swarm VIP, 192.168.1.45, so it works on whichever node gatus runs.
  5. gatus_git_v2 and gatus_data must exist on the cephfs mount before the first deploy, with gatus_git_v2 owned by uid 1000, which git-sync runs as. a missing folder fails the task.
  6. the name carries _v2 because docker never updates an existing volume's options, so a node that already has the volume keeps mounting the old folder. a new name makes every node create it again.

before you deploy

  1. create both folders on the cephfs mount, and give gatus_git_v2 to uid 1000, which git-sync runs as:

    sudo mkdir -p /mnt/docker-cephFS/gatus_git_v2 /mnt/docker-cephFS/gatus_data
    sudo chown 1000:1000 /mnt/docker-cephFS/gatus_git_v2
    
    • a missing folder fails the task, because each volume binds its folder by path
  2. create the discovery overlay on a manager. the stack joins it and does not own it:

    docker network create -d overlay --attachable --scope swarm discovery
    
    • discovery belongs to no stack, so removing gatus never takes it away from homepage or the docker socket proxy

    gatus also joins traefik-api, where it reads traefik's API. create it as in traefik's steps

  3. create the two docker secrets. gitsync_ssh_key_v1 is the private half of a read-only deploy key on the repo, and proxmox_api_token_v1 is the secret of the proxmox api token pve-auditor@pam!homepage

state considerations

both volumes are named binds on cephfs, see stack conventions:

  • data, from /mnt/docker-cephFS/gatus_data, holds the check history, a SQLite database, and the hourly copy the fork makes of it in /data/backup/
  • gitrepo_v2, from /mnt/docker-cephFS/gatus_git_v2, holds git-sync's checkout. gatus mounts it read-only and reads its config from it

on cephfs, git-sync and gatus don't have to share a node, so gatus can move when a node goes down.

network considerations

  • gatus publishes its UI's 8080 as 8085 through the ingress mesh, so every swarm node and the keepalived VIP answer on it, 192.168.1.45:8085 included. its dashboard tile links to it by name, https://gatus.mydomain.com
  • it joins discovery to reach the swarm's read-only docker socket proxy as dockerproxy:2375, which has no port on the lan. git-sync joins it too, as gatus-git-sync, so gatus's check reaches the sidecar's health endpoint by name
  • it joins traefik-api to read traefik's API at http://traefik_traefik:8080, since traefik.mydomain.com is behind oauth

config in git

gatus reads its config straight from a checkout of the repo, which a git-sync sidecar in the same stack keeps current:

  • gatus reads GATUS_CONFIG_PATH=/git/repo/stacks/swarm/gatus/config and reloads when a file changes, so a merged check is live within a minute or two and nothing restarts
  • skip-invalid-config-update: true keeps the last good config running if a change doesn't parse

the checks are split by what they cover, one file each:

config/00-global.yaml: storage, concurrency, and keeping the last good config, 13 lines

download 00-global.yaml

metrics: true

storage:
  type: sqlite
  path: /data/gatus.db

ui:
  title: mydomain homelab health
  header: mydomain homelab

concurrency: 10

skip-invalid-config-update: true
config/10-network.yaml: the gateway and both adguards, 85 lines

download 10-network.yaml

endpoints:
  - name: gateway
    group: network
    url: "icmp://192.168.1.1"
    interval: 60s
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[CONNECTED] == true"

  - name: adguard1
    group: network
    url: "192.168.1.5"
    interval: 60s
    dns:
      query-name: "mydomain.com"
      query-type: "A"
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[DNS_RCODE] == NOERROR"

  - name: adguard2
    group: network
    url: "192.168.1.6"
    interval: 60s
    dns:
      query-name: "mydomain.com"
      query-type: "A"
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[DNS_RCODE] == NOERROR"

  - name: proxy name via winserver01
    group: network
    url: "192.168.1.35"
    interval: 60s
    dns:
      query-name: "traefik.mydomain.com"
      query-type: "A"
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[DNS_RCODE] == NOERROR"
      - "[BODY] == 192.168.1.45"

  - name: proxy name via winserver02
    group: network
    url: "192.168.1.36"
    interval: 60s
    dns:
      query-name: "traefik.mydomain.com"
      query-type: "A"
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[DNS_RCODE] == NOERROR"
      - "[BODY] == 192.168.1.45"

  - name: proxy name via adguard1
    group: network
    url: "192.168.1.5"
    interval: 60s
    dns:
      query-name: "traefik.mydomain.com"
      query-type: "A"
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[DNS_RCODE] == NOERROR"
      - "[BODY] == 192.168.1.45"

  - name: proxy name via adguard2
    group: network
    url: "192.168.1.6"
    interval: 60s
    dns:
      query-name: "traefik.mydomain.com"
      query-type: "A"
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[DNS_RCODE] == NOERROR"
      - "[BODY] == 192.168.1.45"
config/20-proxmox.yaml: the nodes, cluster quorum, and glances on each node, 72 lines

download 20-proxmox.yaml

endpoints:
  - name: pve1
    group: proxmox
    url: "icmp://pve1.mydomain.com"
    interval: 60s
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[CONNECTED] == true"

  - name: pve2
    group: proxmox
    url: "icmp://pve2.mydomain.com"
    interval: 60s
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[CONNECTED] == true"

  - name: pve3
    group: proxmox
    url: "icmp://pve3.mydomain.com"
    interval: 60s
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[CONNECTED] == true"

  - name: glances pve1
    group: proxmox
    url: "http://192.168.1.81:61208/api/4/quicklook"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[STATUS] == 200"
      - "[BODY].mem > 0"

  - name: glances pve2
    group: proxmox
    url: "http://192.168.1.82:61208/api/4/quicklook"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[STATUS] == 200"
      - "[BODY].mem > 0"

  - name: glances pve3
    group: proxmox
    url: "http://192.168.1.83:61208/api/4/quicklook"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[STATUS] == 200"
      - "[BODY].mem > 0"

  - name: pve cluster quorum
    group: proxmox
    url: "https://pve.mydomain.com:8006/api2/json/cluster/status"
    interval: 120s
    headers:
      Authorization: "PVEAPIToken=pve-auditor@pam!homepage=${PROXMOX_TOKEN}"
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[STATUS] == 200"
      - "[BODY].data[0].type == cluster"
      - "[BODY].data[0].quorate == 1"
      - "len([BODY].data) == 4"
      - "[CERTIFICATE_EXPIRATION] > 240h"
config/30-swarm.yaml: the swarm nodes, the VIP, and the swarm's web services, 147 lines

download 30-swarm.yaml

endpoints:
  - name: docker01
    group: swarm
    url: "icmp://192.168.1.41"
    interval: 60s
    ui:
      resolve-successful-conditions: true
    conditions: ["[CONNECTED] == true"]

  - name: docker02
    group: swarm
    url: "icmp://192.168.1.42"
    interval: 60s
    ui:
      resolve-successful-conditions: true
    conditions: ["[CONNECTED] == true"]

  - name: docker03
    group: swarm
    url: "icmp://192.168.1.43"
    interval: 60s
    ui:
      resolve-successful-conditions: true
    conditions: ["[CONNECTED] == true"]

  - name: swarm vip
    group: swarm
    url: "icmp://192.168.1.45"
    interval: 60s
    ui:
      resolve-successful-conditions: true
    conditions: ["[CONNECTED] == true"]

  - name: nginx proxy manager
    group: swarm
    url: "http://192.168.1.45:181"
    interval: 60s
    ui:
      resolve-successful-conditions: true
    conditions: ["[STATUS] == 200"]

  - name: traefik
    group: swarm
    url: "http://traefik_traefik:8080/api/version"
    interval: 60s
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[STATUS] == 200"
      - "has([BODY].Version) == true"

  - name: traefik by name
    group: swarm
    url: "https://auth.mydomain.com/ping"
    interval: 60s
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[STATUS] == 200"
      - "[BODY] == OK"

  - name: bentopdf
    group: swarm
    url: "http://192.168.1.45:8091/"
    interval: 60s
    ui:
      resolve-successful-conditions: true
    conditions: ["[STATUS] == 200"]

  - name: omni-tools
    group: swarm
    url: "http://192.168.1.45:8090/"
    interval: 60s
    ui:
      resolve-successful-conditions: true
    conditions: ["[STATUS] == 200"]

  - name: apprise
    group: swarm
    url: "http://192.168.1.45:8050/status"
    interval: 60s
    ui:
      resolve-successful-conditions: true
    conditions: ["[STATUS] == 200", "[BODY] == OK"]

  - name: unifiapibrowser
    group: swarm
    url: "http://192.168.1.45:8010/"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions: ["[STATUS] == 200", "[BODY] == pat(*UniFi API Browser*)"]

  - name: portainer
    group: swarm
    url: "http://192.168.1.45:9000/api/system/status"
    interval: 60s
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[STATUS] == 200"

  - name: homepage
    group: swarm
    url: "http://192.168.1.45:3000/api/healthcheck"
    interval: 60s
    ui:
      resolve-successful-conditions: true
    conditions: ["[STATUS] == 200"]

  - name: dozzle
    group: swarm
    url: "http://192.168.1.41:8888/healthcheck"
    interval: 60s
    ui:
      resolve-successful-conditions: true
    conditions: ["[STATUS] == 200"]

  - name: glances docker01
    group: swarm
    url: "http://192.168.1.41:61208/api/4/quicklook"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[STATUS] == 200"
      - "[BODY].mem > 0"

  - name: glances docker02
    group: swarm
    url: "http://192.168.1.42:61208/api/4/quicklook"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[STATUS] == 200"
      - "[BODY].mem > 0"

  - name: glances docker03
    group: swarm
    url: "http://192.168.1.43:61208/api/4/quicklook"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[STATUS] == 200"
      - "[BODY].mem > 0"
config/40-truenas.yaml: truenas1 and its apps, 134 lines

download 40-truenas.yaml

endpoints:
  - name: truenas1
    group: truenas
    url: "icmp://192.168.1.86"
    interval: 60s
    ui:
      resolve-successful-conditions: true
    conditions: ["[CONNECTED] == true"]

  - name: frigate cameras
    group: truenas
    url: "http://192.168.1.86:5000/api/stats"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[STATUS] == 200"
      - "[BODY].cameras.front_door.camera_fps > 0"
      - "[BODY].cameras.garage.camera_fps > 0"
      - "[BODY].cameras.kitchen.camera_fps > 0"
      - "[BODY].cameras.living_room.camera_fps > 0"
      - "[BODY].cameras.dining_room.camera_fps > 0"
      - "[BODY].cameras.cat_room.camera_fps > 0"
      - "[BODY].cameras.study.camera_fps > 0"
      - "[BODY].cameras.sarah_study.camera_fps > 0"

  - name: prowlarr
    group: truenas
    url: "http://192.168.1.86:9696/ping"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions: ["[STATUS] == 200", "[BODY].status == OK"]

  - name: radarr
    group: truenas
    url: "http://192.168.1.86:7878/ping"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions: ["[STATUS] == 200", "[BODY].status == OK"]

  - name: sonarr
    group: truenas
    url: "http://192.168.1.86:8989/ping"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions: ["[STATUS] == 200", "[BODY].status == OK"]

  - name: bazarr
    group: truenas
    url: "http://192.168.1.86:6767/api/system/ping"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions: ["[STATUS] == 200", "[BODY].status == OK"]

  - name: qbittorrent
    group: truenas
    url: "http://192.168.1.86:8080/"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions: ["[STATUS] == 200", "[BODY] == pat(*qBittorrent*)"]

  - name: sabnzbd
    group: truenas
    url: "http://192.168.1.86:8081/api?mode=version"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions: ["[STATUS] == 200", "has([BODY].version) == true"]

  - name: seerr
    group: truenas
    url: "http://192.168.1.86:5055/api/v1/status"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions: ["[STATUS] == 200", "has([BODY].version) == true"]

  - name: profilarr
    group: truenas
    url: "http://192.168.1.86:6868/"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions: ["[STATUS] == 200"]

  - name: jellyfin
    group: truenas
    url: "http://192.168.1.86:8096/health"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions: ["[STATUS] == 200"]

  - name: glances
    group: truenas
    url: "http://192.168.1.86:61208/api/4/quicklook"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[STATUS] == 200"
      - "[BODY].mem > 0"

  - name: prometheus
    group: truenas
    url: "http://192.168.1.86:30104/-/healthy"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions: ["[STATUS] == 200"]

  - name: unpoller
    group: truenas
    url: "http://192.168.1.86:30104/api/v1/query?query=unpoller_device_uptime_seconds%20%3E%200"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[STATUS] == 200"
      - "[BODY].status == success"
      - "len([BODY].data.result) > 0"

  - name: victorialogs
    group: truenas
    url: "http://192.168.1.86:9428/health"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions: ["[STATUS] == 200"]
config/50-plumbing.yaml: everything on the dashboard's plumbing tab, 256 lines, 1 note

each in the code opens a note on that line. download 50-plumbing.yaml

lines 1-221: 24 checks, mosquitto to cloudflare ddns
endpoints:
  - name: mosquitto
    group: plumbing
    url: "tcp://192.168.1.45:1883"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions: ["[CONNECTED] == true"]

  - name: oauth2 proxy
    group: plumbing
    url: "http://192.168.1.45:4180/ping"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[STATUS] == 200"
      - "[BODY] == OK"

  - name: docker proxy swarm
    group: plumbing
    url: "http://dockerproxy:2375/version"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[STATUS] == 200"
      - "has([BODY].ApiVersion) == true"

  - name: config sync gatus
    group: plumbing
    url: "http://gatus-git-sync:8080/"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[STATUS] == 200"

  - name: config sync homepage
    group: plumbing
    url: "http://homepage-git-sync:8080/"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[STATUS] == 200"

  - name: docker proxy truenas1
    group: plumbing
    url: "http://192.168.1.86:2375/version"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[STATUS] == 200"
      - "has([BODY].ApiVersion) == true"

  - name: portainer agent truenas1
    group: plumbing
    url: "tcp://192.168.1.86:9001"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions: ["[CONNECTED] == true"]

  - name: portainer agent docker01
    group: plumbing
    url: "tcp://192.168.1.41:9001"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions: ["[CONNECTED] == true"]

  - name: portainer agent docker02
    group: plumbing
    url: "tcp://192.168.1.42:9001"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions: ["[CONNECTED] == true"]

  - name: portainer agent docker03
    group: plumbing
    url: "tcp://192.168.1.43:9001"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions: ["[CONNECTED] == true"]

  - name: dozzle agent docker01
    group: plumbing
    url: "tcp://192.168.1.41:7007"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions: ["[CONNECTED] == true"]

  - name: dozzle agent docker02
    group: plumbing
    url: "tcp://192.168.1.42:7007"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions: ["[CONNECTED] == true"]

  - name: dozzle agent docker03
    group: plumbing
    url: "tcp://192.168.1.43:7007"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions: ["[CONNECTED] == true"]

  - name: dozzle agent truenas1
    group: plumbing
    url: "tcp://192.168.1.86:7007"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions: ["[CONNECTED] == true"]

  - name: flaresolverr
    group: plumbing
    url: "http://192.168.1.86:8191/"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[STATUS] == 200"
      - "[BODY].msg == pat(*FlareSolverr is ready*)"

  - name: ollama
    group: plumbing
    url: "http://192.168.1.86:30068/api/tags"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[STATUS] == 200"
      - "len([BODY].models) > 0"

  - name: dozzle agent syn02
    group: plumbing
    url: "tcp://192.168.1.31:7007"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions: ["[CONNECTED] == true"]

  - name: dozzle agent pi-zwave01
    group: plumbing
    url: "tcp://192.168.1.96:7007"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions: ["[CONNECTED] == true"]

  - name: portainer agent syn02
    group: plumbing
    url: "tcp://192.168.1.31:9001"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions: ["[CONNECTED] == true"]

  - name: portainer agent pi-zwave01
    group: plumbing
    url: "tcp://192.168.1.96:9001"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions: ["[CONNECTED] == true"]

  - name: npm database
    group: plumbing
    url: "http://192.168.1.45:181/api/"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[STATUS] == 200"
      - "[BODY].status == OK"
      - "[BODY].setup == true"

  - name: wordpress database
    group: plumbing
    url: "http://192.168.1.45:8180/wp-json/"
    headers:
      Host: mydomain.com
    client:
      ignore-redirect: true
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[STATUS] == 200"
      - "has([BODY].name) == true"

  - name: open webui redis
    group: plumbing
    url: "http://192.168.1.86:31028/ready"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[STATUS] == 200"
      - "[BODY].status == true"

  - name: cloudflare ddns
    group: plumbing
    url: "https://external.mydomain.com/ping"
    headers:
      Host: auth.mydomain.com
    client:
      dns-resolver: "tcp://1.1.1.1:53"
    interval: 300s
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[STATUS] == 200"
      - "[BODY] == OK"
  - name: acme synology  # (1)!
    group: plumbing
    url: "https://syn02.mydomain.com:5101/"
    interval: 1h
    ui:
      resolve-successful-conditions: true
    conditions: ["[CERTIFICATE_EXPIRATION] > 504h"]
  1. both acme checks, this one and acme asrock bmc, go by hostname with certificate verification on, so a device that fell back to its self-signed factory certificate fails the handshake. an expiry check alone would pass it, because that certificate is long-lived.
lines 231-256: 3 more checks, acme asrock bmc to docker autolabel
  - name: acme asrock bmc
    group: plumbing
    url: "https://asrock-bmc.mydomain.com/"
    interval: 1h
    ui:
      resolve-successful-conditions: true
    conditions: ["[CERTIFICATE_EXPIRATION] > 504h"]

  - name: acme traefik
    group: plumbing
    url: "https://auth.mydomain.com/ping"
    interval: 1h
    ui:
      resolve-successful-conditions: true
    conditions: ["[CERTIFICATE_EXPIRATION] > 504h"]

  - name: docker autolabel
    group: plumbing
    url: "http://dockerproxy:2375/nodes"
    interval: 300s
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[STATUS] == 200"
      - "[BODY] == pat(*\"running_adguard1\":\"1\"*)"
      - "[BODY] == pat(*\"running_adguard2\":\"1\"*)"
config/60-hosts.yaml: syn02 and pi-zwave01, 52 lines

download 60-hosts.yaml

endpoints:
  - name: syn02
    group: hosts
    url: "icmp://192.168.1.31"
    interval: 60s
    ui:
      resolve-successful-conditions: true
    conditions: ["[CONNECTED] == true"]

  - name: glances syn02
    group: hosts
    url: "http://192.168.1.31:61208/api/4/quicklook"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[STATUS] == 200"
      - "[BODY].mem > 0"

  - name: pi-zwave01
    group: hosts
    url: "icmp://192.168.1.96"
    interval: 60s
    ui:
      resolve-successful-conditions: true
    conditions: ["[CONNECTED] == true"]

  - name: glances pi-zwave01
    group: hosts
    url: "http://192.168.1.96:61208/api/4/quicklook"
    interval: 120s
    ui:
      resolve-successful-conditions: true
    conditions:
      - "[STATUS] == 200"
      - "[BODY].mem > 0"

  - name: zwave-js-ui
    group: hosts
    url: "http://192.168.1.96:80/health/zwave/"
    interval: 60s
    ui:
      resolve-successful-conditions: true
    conditions: ["[STATUS] == 200"]

  - name: zigbee2mqtt
    group: hosts
    url: "http://192.168.1.96:8080/"
    interval: 60s
    ui:
      resolve-successful-conditions: true
    conditions: ["[STATUS] == 200", "[BODY] == pat(*Zigbee2MQTT*)"]

the sidecar and gatus's sparse-checkout, ssh_config and known_hosts are on config from git.

the fork

i run a small fork, scyto/gatus, for two things upstream doesn't do.

a token from a file. the proxmox quorum check needs an api token. upstream gatus only substitutes environment variables into its config, and a token in an environment variable is readable by anything that can read the service. the fork also reads PROXMOX_TOKEN_FILE, a docker secret, and substitutes ${PROXMOX_TOKEN} from it.

a backup of gatus's database. the check history is SQLite in WAL mode, so a cephfs snapshot can catch it mid-write. nothing outside gatus can copy it safely: WAL needs every reader on the same host, swarm can't keep a sidecar on gatus's node, and the image has no shell. so the fork copies it with VACUUM INTO to GATUS_SQLITE_BACKUP_PATH, at start and hourly at :50.

what each check asserts

each check asserts something only a working service returns, beyond a 200 from its web server:

check asserts
oauth2-proxy /ping returns the body OK
frigate cameras each camera the check names reports camera_fps > 0 in /api/stats
glances, every host /api/4/quicklook returns collected memory, mem > 0
proxmox an authenticated cluster/status says the cluster is quorate

checking the proxy

gatus checks traefik in five ways, and never through an address outside my lan:

  • every route gets a check, generated with the routes. each goes through the VIP with the name as the Host header and no DNS, so one mis-wired route goes red on its own
  • traefik itself is checked by reading its API over traefik-api, with no name and no DNS
  • one check goes by name, to auth.mydomain.com/ping, the one name with no sign-in in front. it covers DNS, the VIP, traefik and the certificate
  • an app behind oauth also gets a check on its own address. its route check only reaches the sign-in, which answers with a redirect whether the app is up or not
  • a name outside my domain has a certificate of its own, which the route checks can't see. it gets a check by name, for the certificate's expiry, and gatus's extra_hosts points that name at the VIP, so the check stays on the lan

checking a job by its result

some services have no port: they wake, do a job and sleep. gatus checks those by what the job produces.

service checked by
cloudflare ddns the DDNS record, resolved through a public resolver because my lan has its own copy of that name, answers over TLS with a valid certificate only my proxy holds
acme.sh, BMC and synology the certificate each device serves verifies for its hostname and has more than 21 days left
auto-label nodes the node labels it maintains are present, read through the docker socket proxy

checking a database through the app that uses it

the databases are only on their own stack's network. gatus joins none of those networks and checks each database through its app:

database checked through
npm's mariadb NPM's /api/, which queries it on every call
wordpress's mysql /wp-json/, which reads the site's options table
open webui's redis open webui's /ready, which pings redis

checking per node

  • address nodes by their own IP, never the VIP. keepalived's VIP moves, so a check against it stays green while a node is down
  • the adguards are macvlan containers, which a container on the same host can't reach. a shim on every swarm node covers their two IPv4 addresses only, so gatus runs on any node and checks them over IPv4. the shim is in troubleshooting

small things

  • endpoint keys are <group>_<name> with spaces turned into hyphens and nothing else escaped, so names stay free of brackets. the dashboard reads results by key
  • resolve-successful-conditions: true under each endpoint's ui: shows the value a passing condition saw, as well as the condition
  • before relying on a new check, i break it and confirm it fails

checking it

this prints the key of every check whose latest result failed:

curl -s 'http://192.168.1.45:8085/api/v1/endpoints/statuses?page=1&pageSize=1' | jq -r '.[] | select(.results[0].success | not) | .key'

no output means every check passed.