ups on the proxmox cluster¶
all three nodes are on the study UPS. each runs NUT against the UPS card, and when the battery gets low they shut the cluster down together. checked against Proxmox VE 9.2, NUT 2.8.1 and Ceph on the thunderbolt mesh.
what happens in a power cut¶
every node does the same thing, and no node talks to another. they all watch one UPS, so they act within a few seconds of each other.
- the UPS goes on battery
- when the charge is under 30% or the runtime estimate is under 5 minutes, the
driver on each node reports low battery, and
upsmonrunspowerfail-shutdown - each node sets four Ceph flags, writes down a time 150 seconds ahead, and starts a normal shutdown
- the node's guests stop. HA stops its own and does not try to start them on another node
- the node waits until the time it wrote down. this is the hold
- Ceph and corosync stop and the node powers off
- NUT tells the UPS to turn off. see the overview
- when mains returns the nodes boot, HA starts the guests, and once every OSD is up the flags are cleared
the hold is there because Ceph serves storage with two of three nodes up and stops with one. without it a node whose guests stop quickly would leave Ceph while another node's guest was still writing.
if mains returns before step 2, nothing happens.
one cluster setting¶
Datacenter > Options > HA Settings > Shutdown Policy: freeze.
or:
the default, conditional, leaves a powered off node's guests to be started on
another node about two minutes later. that is wrong when all three are going
down.
with freeze, powering one node off on purpose no longer moves its guests. for
planned work on one node, put it in maintenance mode first, which migrates them:
ha-manager crm-command node-maintenance enable pve1
# and when done
ha-manager crm-command node-maintenance disable pve1
NUT on each node¶
/etc/nut/ups.conf, the same on all three:
maxretry = 3
[ups-proxmox]
driver = snmp-ups
port = 192.168.1.73
desc = "Study shelf, PR750LCD"
mibs = cyberpower
snmp_version = v3
secLevel = authPriv
secName = nutups
authProtocol = SHA
privProtocol = AES
authPassword = <the auth key set on the card>
privPassword = <the privacy key set on the card>
ignorelb
override.battery.charge.low = 30
override.battery.runtime.low = 300
ignorelbmakes the driver ignore the UPS's own low battery flag, which comes too late, and use the twooverridevalues instead- this package starts a driver as soon as a UPS appears in
ups.conf. do not also start the driver by hand to test it: two drivers share one socket, and stopping the second one deletes it.systemctl restart nut-driver@ups-proxmoxputs it back
/etc/nut/nut.conf: MODE=netserver on pve1 and MODE=standalone on pve2 and
pve3.
/etc/nut/upsd.conf. pve1 is the one PeaNUT and Home Assistant read, so it
also listens on the LAN:
pve2 and pve3 have only the first line.
/etc/nut/upsd.users, with a different password on each node:
/etc/nut/upsmon.conf:
MONITOR ups-proxmox@127.0.0.1 1 upsmon <the upsmon password> primary
MINSUPPLIES 1
SHUTDOWNCMD "/usr/local/sbin/powerfail-shutdown upsmon"
NOTIFYCMD /usr/sbin/upssched
NOTIFYFLAG ONBATT SYSLOG+WALL+EXEC
NOTIFYFLAG ONLINE SYSLOG+WALL+EXEC
NOTIFYFLAG LOWBATT SYSLOG+WALL
NOTIFYFLAG COMMBAD SYSLOG
NOTIFYFLAG COMMOK SYSLOG
NOTIFYFLAG NOCOMM SYSLOG+WALL
NOTIFYFLAG REPLBATT SYSLOG
POLLFREQ 5
POLLFREQALERT 5
HOSTSYNC 15
DEADTIME 15
POWERDOWNFLAG /etc/killpower
RBWARNTIME 43200
NOCOMMWARNTIME 300
FINALDELAY 5
use 127.0.0.1, not localhost. upsd listens on IPv4 only, and localhost
can resolve to ::1 first.
systemctl restart nut-driver-enumerator.service
systemctl enable --now nut-server.service nut-monitor.service
upsc ups-proxmox@127.0.0.1
the shutdown scripts¶
three scripts in /usr/local/sbin, mode 755, and two units in
/etc/systemd/system.
/etc/default/powerfail:
# 1: log what would be done and change nothing. Leave at 1 until the dry run in
# the test plan has been seen to work on this node. 0 arms it.
DRY_RUN=1
# Seconds this node stays up, counted from the moment the shutdown is ordered,
# after its own guests have stopped. Must be the same on every node, and longer
# than the slowest guest in the cluster takes to shut down.
HOLD_SECONDS=150
/usr/local/sbin/powerfail-shutdown, which upsmon runs as root:
#!/bin/bash
# Shut this node down because the UPS feeding the cluster is on battery.
#
# upsmon runs this as root, as its SHUTDOWNCMD, on every node at about the same
# moment: all three watch the same UPS. Nothing here talks to another node. The
# clock is what coordinates them.
#
# What it does, in order:
# 1. writes the time until which this node must stay up, for powerfail-hold
# 2. sets the Ceph flags that stop the cluster reacting to nodes going away
# 3. starts a normal shutdown
#
# Every step is safe to run three times, because every node runs it.
set -u
STATE_DIR=/var/lib/powerfail
FLAGS="noout norebalance norecover nobackfill"
# How long after this moment the node stays up once its own guests have stopped,
# so that no node leaves Ceph while another node's guest is still writing.
HOLD_SECONDS=150
# 1 means log what would be done and change nothing. Installs start this way.
DRY_RUN=1
# shellcheck disable=SC1091
[ -r /etc/default/powerfail ] && . /etc/default/powerfail
log() { logger -t powerfail -- "$*"; echo "powerfail: $*" >&2; }
reason="${1:-unspecified}"
log "shutdown requested (reason: ${reason}, hold ${HOLD_SECONDS}s, dry run ${DRY_RUN})"
if [ "$DRY_RUN" = "1" ]; then
log "DRY RUN: would write ${STATE_DIR}/hold-until, set Ceph flags [${FLAGS}], and shut down"
exit 0
fi
mkdir -p "$STATE_DIR"
# First, before anything that can hang: if Ceph does not answer, the hold must
# still be in place when the shutdown reaches it.
deadline=$(( $(date +%s) + HOLD_SECONDS ))
echo "$deadline" > "${STATE_DIR}/hold-until"
# The marker powerfail-resume looks for on the next boot. Written before the
# flags are set, so a flag is never left set with nothing recording why.
date -Is > "${STATE_DIR}/flags-set"
for flag in $FLAGS; do
if timeout 20 ceph osd set "$flag" >/dev/null 2>&1; then
log "ceph flag set: ${flag}"
else
log "could not set ceph flag ${flag}; continuing"
fi
done
log "starting shutdown; this node holds until $(date -d "@${deadline}" +%H:%M:%S) after its guests stop"
/sbin/shutdown -h +0 "UPS on battery: cluster shutdown"
/usr/local/sbin/powerfail-hold:
#!/bin/bash
# Keep this node up until the hold deadline. Run by powerfail-hold.service as it
# stops, which systemd orders after the guests and HA have stopped and before
# Ceph, the cluster filesystem and the network stop.
#
# On an ordinary shutdown or reboot there is no deadline file, and this exits at
# once.
set -u
DEADLINE_FILE=/var/lib/powerfail/hold-until
# A deadline further away than this is not believed: the file is stale or wrong,
# and a node must never refuse to shut down because of it.
MAX_WAIT_SECONDS=900
log() { logger -t powerfail -- "$*"; echo "powerfail: $*" >&2; }
[ -r "$DEADLINE_FILE" ] || exit 0
deadline=$(tr -cd '0-9' < "$DEADLINE_FILE")
now=$(date +%s)
if [ -z "$deadline" ]; then
log "hold: ${DEADLINE_FILE} holds no time; not holding"
exit 0
fi
remaining=$(( deadline - now ))
if [ "$remaining" -le 0 ]; then
log "hold: deadline already passed ${remaining#-}s ago; not holding"
exit 0
fi
if [ "$remaining" -gt "$MAX_WAIT_SECONDS" ]; then
log "hold: deadline is ${remaining}s away, more than ${MAX_WAIT_SECONDS}s; not believed, not holding"
exit 0
fi
log "hold: guests on this node have stopped; keeping Ceph up for ${remaining}s more"
sleep "$remaining"
log "hold: released"
exit 0
/etc/systemd/system/powerfail-hold.service. stop order is the reverse of start
order, so starting after Ceph and before the guests means it stops after the
guests and before Ceph:
[Unit]
Description=Hold this node up until every node's guests have stopped (power failure only)
# Stop order is the reverse of start order. Starting after Ceph, the cluster
# filesystem and the network means this stops before they do. Starting before
# the guests and HA means this stops after they have.
After=ceph.target ceph-osd.target ceph-mon.target ceph-mgr.target ceph-mds.target
After=pve-cluster.service corosync.service network-online.target
Before=pve-guests.service pve-ha-lrm.service pve-ha-crm.service
Wants=network-online.target
[Service]
Type=oneshot
RemainAfterExit=yes
ExecStart=/bin/true
ExecStop=/usr/local/sbin/powerfail-hold
# Longer than the longest hold powerfail-hold will honour.
TimeoutStopSec=1000
[Install]
WantedBy=multi-user.target
/usr/local/sbin/powerfail-resume:
#!/bin/bash
# After a power-failure shutdown, clear the Ceph flags once the cluster is whole.
#
# powerfail-resume.service runs this at boot, only when the marker left by
# powerfail-shutdown exists. Every node runs it. Clearing a flag that is already
# clear does nothing.
#
# It waits for every OSD to be up before clearing anything. If that does not
# happen in time, a node or a disk has not come back, and the flags are left set
# on purpose: with them set, Ceph does not start moving data around to cover for
# something that may only be slow to boot. That case needs a person.
#
# This procedure owns the four flags below. If one of them had been set by hand
# for other maintenance before the power failed, it is cleared here too.
set -u
STATE_DIR=/var/lib/powerfail
FLAGS="nobackfill norecover norebalance noout"
WAIT_SECONDS=1800
POLL_SECONDS=15
log() { logger -t powerfail -- "$*"; echo "powerfail: $*" >&2; }
[ -e "${STATE_DIR}/flags-set" ] || exit 0
# Prints "<up> <total>", or nothing if Ceph does not answer. Plain grep, so the
# script needs nothing that is not on a bare node.
osd_counts() {
local json up total
json=$(timeout 20 ceph osd stat --format json 2>/dev/null) || return 0
up=$(printf '%s' "$json" | grep -o '"num_up_osds": *[0-9]*' | grep -o '[0-9]*$' | head -1)
total=$(printf '%s' "$json" | grep -o '"num_osds": *[0-9]*' | grep -o '[0-9]*$' | head -1)
[ -n "$up" ] && [ -n "$total" ] && echo "$up $total"
}
log "resume: waiting up to ${WAIT_SECONDS}s for every OSD to be up"
waited=0
while [ "$waited" -lt "$WAIT_SECONDS" ]; do
read -r up total <<< "$(osd_counts)"
if [ -n "${up:-}" ] && [ -n "${total:-}" ] && [ "$total" -gt 0 ] && [ "$up" -eq "$total" ]; then
for flag in $FLAGS; do
if timeout 20 ceph osd unset "$flag" >/dev/null 2>&1; then
log "resume: ceph flag cleared: ${flag}"
else
log "resume: could not clear ceph flag ${flag}"
fi
done
rm -f "${STATE_DIR}/flags-set" "${STATE_DIR}/hold-until"
log "resume: all ${total} OSDs are up; flags cleared"
exit 0
fi
sleep "$POLL_SECONDS"
waited=$(( waited + POLL_SECONDS ))
done
log "resume: only ${up:-?} of ${total:-?} OSDs are up after ${WAIT_SECONDS}s; LEAVING THE CEPH FLAGS SET. Check the cluster, then: for f in ${FLAGS}; do ceph osd unset \$f; done; rm ${STATE_DIR}/flags-set"
exit 1
/etc/systemd/system/powerfail-resume.service:
[Unit]
Description=Clear the Ceph flags set by a power-failure shutdown, once the cluster is whole
After=ceph.target pve-cluster.service network-online.target
Wants=network-online.target
ConditionPathExists=/var/lib/powerfail/flags-set
[Service]
Type=oneshot
ExecStart=/usr/local/sbin/powerfail-resume
# Longer than the wait inside the script.
TimeoutStartSec=2100
[Install]
WantedBy=multi-user.target
systemctl daemon-reload
systemctl enable --now powerfail-hold.service
systemctl enable powerfail-resume.service
on an ordinary shutdown or reboot there is no deadline file, so the hold exits at once, and no marker file, so the resume unit does not run.
if the OSDs are not all up 30 minutes after boot, powerfail-resume leaves the
flags set and fails. a node or a disk has not come back, and that needs a
person. the log says what to run.
testing it¶
read the log after each step with journalctl -t powerfail.
-
with
DRY_RUN=1,upsmon -c fsdshould logDRY RUN: would write ...and change nothing. it leaves two things behind: the UPS showsFSDto every reader, and/etc/killpowerexists. clean up with: -
check the hold is in the right place with one planned reboot:
then in
journalctl -b -1, in this order:pve-guests,pve-ha-lrmandpve-ha-crmstopped, thehold:lines, then the Ceph daemons and corosync stopping. remove the file afterwards -
check the flags are cleared:
-
set
DRY_RUN=0on every node and pull the mains
the full test, 2026-10-01¶
with the 5 minute timer this page used before the thresholds:
| time | |
|---|---|
| 21:37:31 | UPS on battery |
| 21:42:33 | all three nodes ordered the shutdown, within 5 seconds of each other |
| 21:44:00 | last guest stopped |
| 21:45:07 to 21:45:12 | the three holds released |
| 21:45:20 | all three nodes off |
| 21:49:16 to 21:49:25 | all three booted after mains returned |
| 21:51:13 | every OSD up, flags cleared |
it ordered the shutdown with 93% of the battery left, which is why it now waits for 30% or 5 minutes of runtime. the thresholds have not had a full test yet.
how long each guest took to stop:
| guest | seconds |
|---|---|
| homeassistant | 62 |
| winserver01, winserver02 | 29, 26 |
| docker01, docker02, docker03 | 16 to 21 |
| the two containers | 8 to 10 |