frigate¶
frigate is my NVR, with eight cameras and four inference accelerators. it runs on the NAS because that is where the accelerators are installed.
it isn't a TrueNAS app. portainer deploys it from git as a compose stack, see stacks.
compose.yml, 84 lines, 7 notes
each in the code opens a note on that line. download compose.yml
- labels for the tile on the homepage dashboard. the widget reads
frigate's plain api on port 5000 by address, which needs no key, and the tile's link opens the
authenticated UI on 8971 through traefik, as
nvr. - go2rtc reaches the cameras directly on the lan, and frigate listens on five ports (5000, 8971, 8554, 1984 and 8555), so host networking leaves no port list to keep in step.
- frigate's detector processes pass frames through
/dev/shm, and docker's default of 64 MB is too small for eight cameras. - the numbers are truenas1's video (44), render (107) and apps (568) groups.
- the first bind is the Frigate+ key, which frigate reads from
/run/secrets/PLUS_API_KEYwhen thePLUS_API_KEYenvironment variable is unset. create the file before the first deploy, holding only the key, owned by root with mode 600, because docker makes a directory at a missing bind source. - an anonymous volume, because frigate needs a writable
/tmpseparate from/tmp/cache. each recreate leaves the old one behind, anddocker volume pruneclears them. - passes only when every camera is streaming and the four named detectors are all inferring, and
prints the reason when it fails. docker never restarts a standalone container on health, so
unhealthy shows on the dashboard and in
docker inspectand nothing else happens.
before you deploy¶
-
check the hardware is there:
- docker refuses to start a container whose device node is missing
-
create
/mnt/fast/configs/frigate/frigate_plus_api_keyholding only the Frigate+ key, owned by root with mode 600. the compose binds it read-only at/run/secrets/PLUS_API_KEY- docker creates a directory at a missing bind source, so the file has to exist before the first deploy
why it isn't a catalog app¶
the MemryX MX3 detector needs the container to run privileged. the device node
alone is not enough, because the runtime talks to the mxa-manager daemon, see
sysexts.
the catalog app has a Devices list but no privileged toggle. editing the app's
questions.yaml and its compose template adds one, and i offered that upstream
as truenas/apps#5164. it was
declined.
an app update overwrites both files, and nothing in the UI says the setting has gone. the container still starts, passes its healthcheck and records, but with three detectors instead of four, because the one that needs privileged can't initialise. so i own the compose file.
what portainer gives it¶
- the compose file is in git, reviewed, and deployed from a branch
privileged: trueis in that file and nothing else rewrites it- the same place manages the swarm, so it is one list of stacks
detectors¶
all four run at once, each with its own model:
| detector | device | accelerator |
|---|---|---|
hailo8l |
PCIe |
Hailo-8 |
coral1 |
pci:0 |
Coral Edge TPU |
coral2 |
pci:1 |
Coral Edge TPU |
memryx |
PCIe:0 |
MemryX MX3 |
frigate round-robins detection across them, so the number that matters is their combined throughput, not any one card's latency.
the compose, and why each part is there¶
devices and the daemon socket¶
the compose passes in all four device nodes and mounts the MX3's socket directory:
devices:
- /dev/apex_0:/dev/apex_0
- /dev/apex_1:/dev/apex_1
- /dev/hailo0:/dev/hailo0
- /dev/memx0:/dev/memx0
volumes:
- /run/mxa_manager:/run/mxa_manager
it mounts the whole directory, not one socket, because the directory holds several.
privileged, with a capability list¶
the compose sets privileged and a capability list:
privileged: true
cap_drop: [ALL]
cap_add: [CHOWN, FOWNER, DAC_OVERRIDE, SETGID, SETUID, PERFMON, KILL]
under privileged the capability list is advisory. i keep it so that if the
MX3 ever stops needing privileged, removing that one line leaves a correct set,
not an empty one.
one MIG slice for ffmpeg¶
ffmpeg decodes on one MIG slice of the GPU, not the whole GPU:
environment:
NVIDIA_DRIVER_CAPABILITIES: all
NVIDIA_VISIBLE_DEVICES: MIG-<uuid>
CUDA_VISIBLE_DEVICES: MIG-<uuid>
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids: [MIG-<uuid>]
capabilities: [gpu]
the host's docker default runtime is already nvidia, so the compose needs no
runtime: line. get the UUID from nvidia-smi -L. it goes in both variables
and in the reservation. If you are not using MIG you can specify the card's normal
UUID, provided by the command above.
storage¶
| path | dataset | holds |
|---|---|---|
/config |
fast/configs/frigate |
config.yaml, the sqlite database, model cache |
/media |
fast/frigate/media |
recordings, clips, exports |
/tmp/cache |
fast/frigate/cache |
in-progress recording segments |
- config is the only irreplaceable one, and the only one snapshotted. media and cache must not be, see storage
- all three are host paths on datasets i made, not ixVolumes
- cache churns hard. it is scratch, and frigate rebuilds it
certificates¶
frigate serves TLS itself, so the certificate is bound in from the TrueNAS store:
volumes:
- /etc/certificates/frigate_cert.crt:/etc/letsencrypt/live/frigate/fullchain.pem:ro
- /etc/certificates/frigate_cert.key:/etc/letsencrypt/live/frigate/privkey.pem:ro
leaving the app catalog made one thing worse. TrueNAS renews the certificate in place, and as an app frigate would be restarted to pick it up. TrueNAS doesn't manage this stack, so frigate keeps serving the old certificate until something restarts it, which i now have to remember.
one portainer quirk¶
portainer deploys standalone stacks with the compose it bundles, which can be older than the daemon underneath. on a host running engine 29.x, mine refused the stack with:
so the compose file leaves that field out. without it the healthcheck runs at its normal interval during startup, and the container reports healthy slightly later. nothing else changes.
checking it¶
the healthcheck prints why it failed:
to count the detectors by hand:
the inference times show that the cards are in use, not only present:
detectors lists each detector with an inference speed. cameras lists camera
and detection fps per camera.