Skip to content

frigate

frigate is my NVR, with eight cameras and four inference accelerators. it runs on the NAS because that is where the accelerators are installed.

it isn't a TrueNAS app. portainer deploys it from git as a compose stack, see stacks.

compose.yml, 84 lines, 7 notes

each in the code opens a note on that line. download compose.yml

services:
  frigate:
    image: ghcr.io/blakeblackshear/frigate:0.17.1@sha256:1724960349dad0bd2ae8ec884171a6fd5755a4dc242a0e66cadbda9c0e85c99b
    container_name: frigate
    labels:  # (1)!
      - homepage.group=Media
      - homepage.name=Frigate
      - homepage.icon=frigate.png
      - homepage.href=https://nvr.mydomain.com
      - homepage.description=NVR, 8 cameras, 4 detectors
      - homepage.widget.type=frigate
      - homepage.widget.url=http://192.168.1.86:5000
    platform: linux/amd64
    restart: unless-stopped

    privileged: true

    network_mode: host  # (2)!

    shm_size: 512M  # (3)!

    user: "0:0"
    cap_drop:
      - ALL
    cap_add:
      - CHOWN
      - FOWNER
      - DAC_OVERRIDE
      - SETGID
      - SETUID
      - PERFMON
      - KILL
    security_opt:
      - no-new-privileges=true
    group_add:  # (4)!
      - "44"
      - "107"
      - "568"

    devices:
      - /dev/apex_0:/dev/apex_0
      - /dev/apex_1:/dev/apex_1
      - /dev/hailo0:/dev/hailo0
      - /dev/memx0:/dev/memx0

    environment:
      TZ: America/Los_Angeles
      UMASK: "002"
      UMASK_SET: "002"
      NVIDIA_DRIVER_CAPABILITIES: all
      NVIDIA_VISIBLE_DEVICES: MIG-87144e2e-17bb-5983-95c6-b84fa3d63613
      CUDA_VISIBLE_DEVICES: MIG-87144e2e-17bb-5983-95c6-b84fa3d63613

    volumes:  # (5)!
      - /mnt/fast/configs/frigate/frigate_plus_api_key:/run/secrets/PLUS_API_KEY:ro
      - /mnt/fast/configs/frigate:/config
      - /mnt/fast/frigate/media:/media
      - /run/mxa_manager:/run/mxa_manager
      - /mnt/fast/frigate/cache:/tmp/cache
      - /etc/certificates/frigate_cert.crt:/etc/letsencrypt/live/frigate/fullchain.pem:ro
      - /etc/certificates/frigate_cert.key:/etc/letsencrypt/live/frigate/privkey.pem:ro
      - /tmp  # (6)!

    deploy:
      resources:
        limits:
          cpus: "2"
          memory: 8192M
        reservations:
          devices:
            - driver: nvidia
              device_ids:
                - MIG-87144e2e-17bb-5983-95c6-b84fa3d63613
              capabilities:
                - gpu

    healthcheck:  # (7)!
      test:
        - CMD-SHELL
        - 'body=$$(curl -fsS --max-time 10 http://127.0.0.1:5000/api/stats) || { echo "frigate api unreachable"; exit 1; }; out=$$(printf "%s" "$$body" | jq -r ''[ (if (.cameras | length) == 0 then "no cameras configured" else empty end), (.cameras | to_entries[] | select(.value.camera_fps <= 0) | "camera \(.key) not streaming"), (if (.detectors | keys | sort) != ["coral1","coral2","hailo8l","memryx"] then "detector set is \(.detectors | keys | sort) expected [coral1,coral2,hailo8l,memryx]" else empty end), (.detectors | to_entries[] | select(.value.inference_speed <= 0) | "detector \(.key) stalled") ] as $$p | if ($$p | length) == 0 then "ok" else ($$p | join("; ")) end'') || { echo "frigate stats unparseable"; exit 1; }; [ "$$out" = ok ] || { echo "$$out"; exit 1; }'
      interval: 60s
      timeout: 15s
      retries: 3
      start_period: 360s
  1. labels for the tile on the homepage dashboard. the widget reads frigate's plain api on port 5000 by address, which needs no key, and the tile's link opens the authenticated UI on 8971 through traefik, as nvr.
  2. go2rtc reaches the cameras directly on the lan, and frigate listens on five ports (5000, 8971, 8554, 1984 and 8555), so host networking leaves no port list to keep in step.
  3. frigate's detector processes pass frames through /dev/shm, and docker's default of 64 MB is too small for eight cameras.
  4. the numbers are truenas1's video (44), render (107) and apps (568) groups.
  5. the first bind is the Frigate+ key, which frigate reads from /run/secrets/PLUS_API_KEY when the PLUS_API_KEY environment variable is unset. create the file before the first deploy, holding only the key, owned by root with mode 600, because docker makes a directory at a missing bind source.
  6. an anonymous volume, because frigate needs a writable /tmp separate from /tmp/cache. each recreate leaves the old one behind, and docker volume prune clears them.
  7. passes only when every camera is streaming and the four named detectors are all inferring, and prints the reason when it fails. docker never restarts a standalone container on health, so unhealthy shows on the dashboard and in docker inspect and nothing else happens.

before you deploy

  1. check the hardware is there:

    ls -l /dev/memx0 /dev/hailo0 /dev/apex_*
    
    • docker refuses to start a container whose device node is missing
  2. create /mnt/fast/configs/frigate/frigate_plus_api_key holding only the Frigate+ key, owned by root with mode 600. the compose binds it read-only at /run/secrets/PLUS_API_KEY

    • docker creates a directory at a missing bind source, so the file has to exist before the first deploy

why it isn't a catalog app

the MemryX MX3 detector needs the container to run privileged. the device node alone is not enough, because the runtime talks to the mxa-manager daemon, see sysexts.

the catalog app has a Devices list but no privileged toggle. editing the app's questions.yaml and its compose template adds one, and i offered that upstream as truenas/apps#5164. it was declined.

an app update overwrites both files, and nothing in the UI says the setting has gone. the container still starts, passes its healthcheck and records, but with three detectors instead of four, because the one that needs privileged can't initialise. so i own the compose file.

what portainer gives it

  • the compose file is in git, reviewed, and deployed from a branch
  • privileged: true is in that file and nothing else rewrites it
  • the same place manages the swarm, so it is one list of stacks

detectors

all four run at once, each with its own model:

detector device accelerator
hailo8l PCIe Hailo-8
coral1 pci:0 Coral Edge TPU
coral2 pci:1 Coral Edge TPU
memryx PCIe:0 MemryX MX3

frigate round-robins detection across them, so the number that matters is their combined throughput, not any one card's latency.

the compose, and why each part is there

devices and the daemon socket

the compose passes in all four device nodes and mounts the MX3's socket directory:

compose.yml
    devices:
      - /dev/apex_0:/dev/apex_0
      - /dev/apex_1:/dev/apex_1
      - /dev/hailo0:/dev/hailo0
      - /dev/memx0:/dev/memx0
    volumes:
      - /run/mxa_manager:/run/mxa_manager

it mounts the whole directory, not one socket, because the directory holds several.

privileged, with a capability list

the compose sets privileged and a capability list:

compose.yml
    privileged: true
    cap_drop: [ALL]
    cap_add: [CHOWN, FOWNER, DAC_OVERRIDE, SETGID, SETUID, PERFMON, KILL]

under privileged the capability list is advisory. i keep it so that if the MX3 ever stops needing privileged, removing that one line leaves a correct set, not an empty one.

one MIG slice for ffmpeg

ffmpeg decodes on one MIG slice of the GPU, not the whole GPU:

compose.yml
    environment:
      NVIDIA_DRIVER_CAPABILITIES: all
      NVIDIA_VISIBLE_DEVICES: MIG-<uuid>
      CUDA_VISIBLE_DEVICES: MIG-<uuid>
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              device_ids: [MIG-<uuid>]
              capabilities: [gpu]

the host's docker default runtime is already nvidia, so the compose needs no runtime: line. get the UUID from nvidia-smi -L. it goes in both variables and in the reservation. If you are not using MIG you can specify the card's normal UUID, provided by the command above.

storage

path dataset holds
/config fast/configs/frigate config.yaml, the sqlite database, model cache
/media fast/frigate/media recordings, clips, exports
/tmp/cache fast/frigate/cache in-progress recording segments
  • config is the only irreplaceable one, and the only one snapshotted. media and cache must not be, see storage
  • all three are host paths on datasets i made, not ixVolumes
  • cache churns hard. it is scratch, and frigate rebuilds it

certificates

frigate serves TLS itself, so the certificate is bound in from the TrueNAS store:

compose.yml
    volumes:
      - /etc/certificates/frigate_cert.crt:/etc/letsencrypt/live/frigate/fullchain.pem:ro
      - /etc/certificates/frigate_cert.key:/etc/letsencrypt/live/frigate/privkey.pem:ro

leaving the app catalog made one thing worse. TrueNAS renews the certificate in place, and as an app frigate would be restarted to pick it up. TrueNAS doesn't manage this stack, so frigate keeps serving the old certificate until something restarts it, which i now have to remember.

one portainer quirk

portainer deploys standalone stacks with the compose it bundles, which can be older than the daemon underneath. on a host running engine 29.x, mine refused the stack with:

can't set healthcheck.start_interval as feature require Docker Engine v25 or later

so the compose file leaves that field out. without it the healthcheck runs at its normal interval during startup, and the container reports healthy slightly later. nothing else changes.

checking it

the healthcheck prints why it failed:

docker inspect --format '{{json .State.Health.Log}}' frigate | python3 -m json.tool | tail -8

to count the detectors by hand:

docker exec frigate ps -eo args | grep -c 'frigate.detector:'     # expect 4

the inference times show that the cards are in use, not only present:

docker exec frigate curl -s http://127.0.0.1:5000/api/stats | python3 -m json.tool | head -40

detectors lists each detector with an inference speed. cameras lists camera and detection fps per camera.