hardware and base install¶
what the NAS is built from, and the parts of the base configuration the rest of these pages assume. it runs the TrueNAS 26.0 beta on bare metal.
it was a VM on a single proxmox host for about a year before this, with every drive passed through over PCIe. that setup is kept for reference in virtualized on proxmox.
hardware¶
| part | what |
|---|---|
| board | ASRock Rack GENOAD8UD-2T/X550 |
| cpu | AMD EPYC 9115, 16 cores / 32 threads |
| memory | 192 GB ECC |
| network | 25GbE Mellanox (ConnectX, fibre) in use; onboard dual 10GbE (X550) idle |
| gpu | NVIDIA RTX PRO 6000 Blackwell Workstation Edition |
| accelerators | Hailo-8, MemryX MX3, 2 × Coral Edge TPU |
the board is single socket. the unusual part is the accelerators: frigate can run a detector on each of three different inference accelerators at once, which is why this box has all three, see frigate.
they all need a driver the base OS does not ship, which is what sysexts are for.
the gpu is partitioned¶
the RTX PRO 6000 is split with MIG into three instances rather than handed to one consumer whole:
| instance | size |
|---|---|
| 0 | 2g.48gb |
| 1 | 1g.24gb |
| 2 | 1g.24gb |
- MIG gives each consumer a hard slice of memory and compute, so one app cannot starve another
- each instance has its own UUID, and an app is given that UUID rather than the whole GPU. frigate gets one 1g.24gb instance, ollama and open-webui use the others
- the partitioning survives reboots via a sysext, see sysexts
the gpu is power capped¶
the card may draw 600 W by default. i cap it at 350 W, so the box cannot exceed the UPS's power budget:
| what | power |
|---|---|
| the UPS | CyberPower OR1500LCDRTXL2U, 1500 VA / 1125 W |
| other equipment on it | about 33% of that, roughly 370 W |
| left for this box | roughly 750 W |
- with the server and the card both at full load, and the card uncapped, the UPS went over its limit in watts
-
the cap takes 250 W off the most the card can draw
-
the limit is not saved, the card is back at 600 W after every boot. so the line runs as a POSTINIT command, see tasks and scripts
- POSTINIT, not PREINIT: the driver is loaded by a PREINIT script, see
sysexts, and
nvidia-smineeds it -pm 1is persistence mode. it keeps the driver loaded when nothing is using the card, so the limit stays set- the limit is for the whole card, not per MIG instance, and the card accepts 150 W to 600 W
to check it:
disks and pools¶
18 drives, three pools:
| pool | topology | size | holds |
|---|---|---|---|
boot-pool |
mirror, 2 × 960 GB NVMe | 888 GB | the OS and boot environments |
fast |
2 × mirror, 4 × 4 TB NVMe | 7.25 TB | apps, app config, container storage |
rust |
raidz2, 6 × 24 TB HDD + SLOG + L2ARC | 131 TB | media, backups, S3 buckets |
the drives:
| drive | count | where |
|---|---|---|
| Seagate IronWolf Pro 24 TB | 6 | rust, raidz2 |
| Seagate FireCuda 530 4 TB NVMe | 4 | fast, two mirrors |
| Intel Optane 905P 960 GB | 1 | rust SLOG |
| ADATA XPG SX8200 Pro 2 TB NVMe | 1 | rust L2ARC |
| Kingston DC2000B 960 GB NVMe | 2 | boot-pool mirror |
| Intel Optane 905P (960 GB, 2 × 1.5 TB), Micron 7400 Pro 3.84 TB | 4 | not in a pool |
fastis two mirrors rather than one raidz because it holds app databases and container root filesystems, so IOPS matter more than usable capacityrustis raidz2 because 24 TB drives take a long time to resilver, and raidz2 survives a second failure during that window- the SLOG is Optane, for low latency sync writes. the L2ARC is an ordinary NVMe drive
- every dataset uses lz4 compression, none are encrypted, and the ones i made
use the default 128K record size except the PBS datastore,
rust/local-backups/pbs, which uses 1M, see PBS - both pools are scrubbed on a schedule, see tasks and scripts
the dataset layout and which datasets to snapshot are on their own page, see storage.
networking¶
one 25GbE link carries everything:
| setting | value |
|---|---|
| interface | one 25GbE fibre port (mlx5), static IPv4 |
| MTU | 9000 |
| other ports | onboard dual 10GbE (ixgbe), down |
- MTU 9000 here matches the rest of the LAN. use jumbo frames end to end or not at all: a path where one hop is 1500 gives you connections that open and then stall on the first large transfer
- no LAGG. one 25GbE link is well past what the pools can serve, and the onboard 10GbE ports stay down rather than add paths nothing needs
- the pbs container gets its own MAC on this same interface rather than a NAT, so it appears on the LAN as its own host
identity and access¶
- the box is joined to an Active Directory domain, so SMB shares authenticate
against it and domain accounts resolve as users on the box. domain users
appear as
MYDOMAIN\userwith high uids, distinct from local accounts - the join has the account cache and dynamic DNS updates on
- a script also registers the box's IPv4 and IPv6 addresses in DNS with
nsupdate, at boot and nightly, see tasks and scripts - apps do not run as you. most catalog apps run as uid/gid 568 (
apps), so a host path you created as root gives permission denied until it is chowned, see the app gotchas - containers are id-mapped and isolated, so root inside the container is an unprivileged uid on the host, see containers
- certificates come from the TrueNAS certificate store, issued by Let's Encrypt over a DNS challenge. apps pick one in their config and TrueNAS restarts them when it renews
use the full hostname the certificate was issued for. the short name fails verification and clients like portainer refuse it.