services:git-sync:image:registry.k8s.io/git-sync/git-sync:v4.7.1user:"1000:1000"environment:-TZ=America/Los_Angeles-GITSYNC_REPO=ssh://git@ssh.github.com:443/<you>/<repo>.git-GITSYNC_REF=deploy/swarm/gatus-GITSYNC_PERIOD=60s-GITSYNC_ROOT=/git-GITSYNC_LINK=repo-GITSYNC_SSH_KEY_FILE=/dev/shm/key-GITSYNC_SSH_KNOWN_HOSTS_FILE=/known_hosts-GITSYNC_ADD_USER=true-GITSYNC_FILTER=blob:none-GITSYNC_SPARSE_CHECKOUT_FILE=/sparse-checkout-GITSYNC_HTTP_BIND=:8080entrypoint:-/bin/sh--c-|umask 077printf '%s\n' "$$(cat /run/secrets/gitsync_ssh_key_v1)" > /dev/shm/key[ -s /dev/shm/key ] || { echo "git-sync: could not write the SSH key to /dev/shm" >&2; exit 1; }mkdir -p /tmp/.ssh && cp /ssh_config /tmp/.ssh/config || { echo "git-sync: could not install the ssh config" >&2; exit 1; }exec /git-syncvolumes:-type:bind# (1)!source:/usr/share/zoneinfotarget:/usr/share/zoneinforead_only:true-type:volumesource:gitrepo_v2target:/gitvolume:nocopy:trueconfigs:-source:gatus_known_hosts_v2target:/known_hosts-source:gatus_ssh_config_v2target:/ssh_config-source:gatus_sparse_checkout_v1target:/sparse-checkoutsecrets:-gitsync_ssh_key_v1networks:discovery:aliases:-gatus-git-syncdeploy:mode:replicatedreplicas:1gatus:image:ghcr.io/scyto/gatus:v5.36.0-12-secrets-backup@sha256:61a8988f286fed144c91cdfd1acb7d429d2ba0fecc47007c95df2e0a0c840518environment:# (2)!-TZ=America/Los_Angeles-GATUS_CONFIG_PATH=/git/repo/stacks/swarm/gatus/config-GATUS_LOG_LEVEL=INFO-PROXMOX_TOKEN_FILE=/run/secrets/proxmox_api_token_v1-GATUS_SQLITE_BACKUP_PATH=/data/backup/gatus.dbvolumes:-type:bindsource:/usr/share/zoneinfotarget:/usr/share/zoneinforead_only:true-gitrepo_v2:/git:ro-data:/datasecrets:-proxmox_api_token_v1networks:-discovery# (3)!-traefik-apiextra_hosts:-mydomain1.com:192.168.1.45ports:-8085:8080deploy:mode:replicatedreplicas:1labels:# (4)!-homepage.group=Monitoring-homepage.name=Gatus-homepage.icon=gatus.png-homepage.href=https://gatus.mydomain.com-homepage.description=Health checks, config in git-homepage.widget.type=gatus-homepage.widget.url=http://192.168.1.45:8085volumes:# (5)!gitrepo_v2:# (6)!driver:localdriver_opts:type:nonedevice:"/mnt/docker-cephFS/gatus_git_v2"o:binddata:driver:localdriver_opts:type:nonedevice:"/mnt/docker-cephFS/gatus_data"o:bindnetworks:discovery:external:truetraefik-api:external:trueconfigs:gatus_known_hosts_v2:file:./known_hostsgatus_ssh_config_v2:file:./ssh_configgatus_sparse_checkout_v1:file:./sparse-checkoutsecrets:proxmox_api_token_v1:external:truegitsync_ssh_key_v1:external:true
neither the git-sync image nor the gatus image has time zone data, so TZ alone changes
nothing. both services mount the host's /usr/share/zoneinfo read-only.
turns on the fork's hourly VACUUM INTO copy of the database, for the
cephfs backup. the fork says why it isn't a sidecar.
the discovery overlay, where gatus reaches the swarm's read-only docker socket proxy as
dockerproxy:2375 and both git-sync sidecars by their aliases. none of them has a port on the
lan.
the tile on the homepage dashboard, under deploy.labels because homepage reads
labels from the swarm's service specs. the widget reads gatus through the swarm VIP,
192.168.1.45, so it works on whichever node gatus runs.
gatus_git_v2 and gatus_data must exist on the cephfs mount before the first deploy, with
gatus_git_v2 owned by uid 1000, which git-sync runs as. a missing folder fails the task.
the name carries _v2 because docker never updates an existing volume's options, so a node that
already has the volume keeps mounting the old folder. a new name makes every node create it again.
discovery belongs to no stack, so removing gatus never takes it away from homepage or the docker socket proxy
gatus also joins traefik-api, where it reads traefik's API. create it as in traefik's steps
create the two docker secrets. gitsync_ssh_key_v1 is the private half of a read-only deploy key on the repo, and proxmox_api_token_v1 is the secret of the proxmox api token pve-auditor@pam!homepage
gatus publishes its UI's 8080 as 8085 through the ingress mesh, so every swarm node and the keepalived VIP answer on it, 192.168.1.45:8085 included. its dashboard tile links to it by name, https://gatus.mydomain.com
it joins discovery to reach the swarm's read-only docker socket proxy as dockerproxy:2375, which has no port on the lan. git-sync joins it too, as gatus-git-sync, so gatus's check reaches the sidecar's health endpoint by name
it joins traefik-api to read traefik's API at http://traefik_traefik:8080, since traefik.mydomain.com is behind oauth
gatus reads its config straight from a checkout of the repo, which a git-sync sidecar in the same stack keeps current:
gatus reads GATUS_CONFIG_PATH=/git/repo/stacks/swarm/gatus/config and reloads when a file changes, so a merged check is live within a minute or two and nothing restarts
skip-invalid-config-update: true keeps the last good config running if a change doesn't parse
the checks are split by what they cover, one file each:
config/00-global.yaml: storage, concurrency, and keeping the last good config, 13 lines
endpoints:-name:gatewaygroup:networkurl:"icmp://192.168.1.1"interval:60sui:resolve-successful-conditions:trueconditions:-"[CONNECTED]==true"-name:adguard1group:networkurl:"192.168.1.5"interval:60sdns:query-name:"mydomain.com"query-type:"A"ui:resolve-successful-conditions:trueconditions:-"[DNS_RCODE]==NOERROR"-name:adguard2group:networkurl:"192.168.1.6"interval:60sdns:query-name:"mydomain.com"query-type:"A"ui:resolve-successful-conditions:trueconditions:-"[DNS_RCODE]==NOERROR"-name:proxy name via winserver01group:networkurl:"192.168.1.35"interval:60sdns:query-name:"traefik.mydomain.com"query-type:"A"ui:resolve-successful-conditions:trueconditions:-"[DNS_RCODE]==NOERROR"-"[BODY]==192.168.1.45"-name:proxy name via winserver02group:networkurl:"192.168.1.36"interval:60sdns:query-name:"traefik.mydomain.com"query-type:"A"ui:resolve-successful-conditions:trueconditions:-"[DNS_RCODE]==NOERROR"-"[BODY]==192.168.1.45"-name:proxy name via adguard1group:networkurl:"192.168.1.5"interval:60sdns:query-name:"traefik.mydomain.com"query-type:"A"ui:resolve-successful-conditions:trueconditions:-"[DNS_RCODE]==NOERROR"-"[BODY]==192.168.1.45"-name:proxy name via adguard2group:networkurl:"192.168.1.6"interval:60sdns:query-name:"traefik.mydomain.com"query-type:"A"ui:resolve-successful-conditions:trueconditions:-"[DNS_RCODE]==NOERROR"-"[BODY]==192.168.1.45"
config/20-proxmox.yaml: the nodes, cluster quorum, and glances on each node, 72 lines
both acme checks, this one and acme asrock bmc, go by hostname with certificate verification on,
so a device that fell back to its self-signed factory certificate fails the handshake. an expiry
check alone would pass it, because that certificate is long-lived.
lines 231-256: 3 more checks, acme asrock bmc to docker autolabel
i run a small fork, scyto/gatus, for two things upstream doesn't do.
a token from a file. the proxmox quorum check needs an api token. upstream gatus only substitutes environment variables into its config, and a token in an environment variable is readable by anything that can read the service. the fork also reads PROXMOX_TOKEN_FILE, a docker secret, and substitutes ${PROXMOX_TOKEN} from it.
a backup of gatus's database. the check history is SQLite in WAL mode, so a cephfs snapshot can catch it mid-write. nothing outside gatus can copy it safely: WAL needs every reader on the same host, swarm can't keep a sidecar on gatus's node, and the image has no shell. so the fork copies it with VACUUM INTO to GATUS_SQLITE_BACKUP_PATH, at start and hourly at :50.
gatus checks traefik in five ways, and never through an
address outside my lan:
every route gets a check, generated with the routes. each goes through the
VIP with the name as the Host header and no DNS, so one mis-wired route goes
red on its own
traefik itself is checked by reading its API over traefik-api, with no name
and no DNS
one check goes by name, to auth.mydomain.com/ping, the one name with no
sign-in in front. it covers DNS, the VIP, traefik and the certificate
an app behind oauth also gets a check on its own address. its route check
only reaches the sign-in, which answers with a redirect whether the app is
up or not
a name outside my domain has a certificate of its own, which the route
checks can't see. it gets a check by name, for the certificate's expiry,
and gatus's extra_hosts points that name at the VIP, so the check stays on
the lan
some services have no port: they wake, do a job and sleep. gatus checks those by what the job produces.
service
checked by
cloudflare ddns
the DDNS record, resolved through a public resolver because my lan has its own copy of that name, answers over TLS with a valid certificate only my proxy holds
acme.sh, BMC and synology
the certificate each device serves verifies for its hostname and has more than 21 days left
auto-label nodes
the node labels it maintains are present, read through the docker socket proxy
address nodes by their own IP, never the VIP. keepalived's VIP moves, so a check against it stays green while a node is down
the adguards are macvlan containers, which a container on the same host can't reach. a shim on every swarm node covers their two IPv4 addresses only, so gatus runs on any node and checks them over IPv4. the shim is in troubleshooting
endpoint keys are <group>_<name> with spaces turned into hyphens and nothing else escaped, so names stay free of brackets. the dashboard reads results by key
resolve-successful-conditions: true under each endpoint's ui: shows the value a passing condition saw, as well as the condition
before relying on a new check, i break it and confirm it fails