portainer¶
one portainer runs all six nodes from one UI: the swarm's three and the three standalone hosts.
i run the business edition on the paid home & student licence. it costs US$155 a year and covers up to 15 nodes with every business feature, for personal use only. i pay for it because the free business licence covers 3 nodes and i have 6.
the business feature this setup depends on is role based access. the homepage dashboard's api token belongs to a user who has read access through the helpdesk role on every environment.
on the swarm it is two stacks: the server, and an agent on every node. together they replace portainer's stock swarm file.
compose.yml: the server, 31 lines, 1 note
each in the code opens a note on that line. download compose.yml
- the
-H tcp://tasks.agent:9001on thecommand:line is left from the stock file, and the server cannot resolve that name, since the agents are on another overlay. portainer reads it only when it has no environment yet, and each environment here points at an address on 9001.
compose.yml: the agent, on every node, 38 lines, 3 notes
each in the code opens a note on that line. download compose.yml
- the host's
/, read-only. portainer only reads/hostto show a node's devices and storage and to browse its files, and write access would add nothing but uploading, deleting and renaming files there. - the agent image has no time zone data, so
TZalone changes nothing. the host's/usr/share/zoneinfois mounted read-only. - host mode, so the agent on the node that holds an address answers on it, the keepalived VIP included. through swarm's ingress, any node's agent could answer.
apart from names, the server's file differs from the stock one in three ways: it has no agent service, the image is pinned to a version instead of lts, and the data volume is bound to the replicated storage. that lets the server start on any manager and find its database.
install on the swarm¶
on one manager:
-
create the data folder on the cephfs mount:
- a missing folder fails the task, because the data volume binds its folder by path
-
delete the
command:line (-H tcp://tasks.agent:9001 --tlsskipverify) from the downloaded file, then deploy the server from it:- the line is left from the stock file. it names an agent the server can't resolve, and the server only reads it on a first start, before it has an environment. without the line the server starts with no environment, and step 5 adds one
- the UI is on
https://<node>:9443.9000serves it over plain http, and8000is the tunnel for edge agents, which i don't use
-
browse to it, create the admin user and enter the licence key
-
start a temporary agent on a spare port:
docker run -d --name temp-agent -p 9002:9001 \ -v /var/run/docker.sock:/var/run/docker.sock \ -v /var/lib/docker/volumes:/var/lib/docker/volumes \ portainer/agent:2.45.1- portainer creates a stack through an environment's agent, so the agent stack needs an agent before it exists
9002leaves9001free for the agent stack, which publishes it on every node
-
in portainer, add a docker swarm environment using the agent, at
<manager>:9002 - create the
agentstack from git, as in cut a stack over: referencerefs/heads/deploy/swarm/agent, compose pathstacks/swarm/agent/compose.yml -
point the environment at the keepalived VIP,
192.168.1.45:9001, and check the agent runs on every node:REPLICASreads3/3, one task on each of the three nodes
-
remove the temporary agent:
-
add the standalone hosts, see add a host
agents on every endpoint¶
| environment | agent | how it is installed |
|---|---|---|
| swarm | a global service, on 9001 on each node |
the agent stack above, from git |
| truenas1 | truenas custom app | truenas apps |
| syn02, pi-zwave01 | a container, restart: always |
add a host |
each environment points at an address on 9001. the swarm's is the keepalived VIP, and each standalone host's is the host itself. portainer reaches the swarm through that address, and removing the agent stack cuts it off.
signing in with entra¶
users sign in to portainer with their entra ID accounts, through portainer's own oauth setting. in settings → authentication → oauth, choose the microsoft provider with a custom configuration.
-
register an app in entra with a web redirect URI of portainer's own address,
https://portainer.mydomain.com, and a client secret- set assignment required on its enterprise application and assign who may sign in
-
fill in the form:
field value client ID, client secret the app's authorization URL https://login.microsoftonline.com/<tenant-id>/oauth2/v2.0/authorizeaccess token URL https://login.microsoftonline.com/<tenant-id>/oauth2/v2.0/tokenresource URL https://graph.microsoft.com/v1.0/meredirect URL portainer's own address, the same as the registered redirect URI logout URL https://login.microsoftonline.com/<tenant-id>/oauth2/v2.0/logoutuser identifier userPrincipalNamescopes openid profile- the tenant's own userinfo endpoint,
https://graph.microsoft.com/oidc/userinfo, has nouserPrincipalName, and signing in then fails with "failed to extract username from oauth resource"
- the tenant's own userinfo endpoint,
-
turn automatic user provisioning off, and create each user in portainer, named with their user principal name, before you log out
- with it off, only users made in portainer get in. without the user, your first sign-in with entra has nowhere to land
the initial admin can always log in with a password, whatever the setting,
and API keys skip the login altogether. so from outside,
traefik keeps entra in front of /api/auth and of any
request that carries an API key.
what deploys from git¶
portainer itself is the one stack not deployed from git. it would be applying changes to itself, and a bad commit leaves no UI to fix it with.
the agent deploys from git like any other stack. a change to it is an ordinary service update, which the swarm manager finishes even if portainer's connection drops while the agents restart.
a broken agent compose reaches every node a few minutes after it merges, and portainer loses its view of the swarm. to recover, point the environment at the temporary agent from install step 4 while a fixed or reverted compose deploys.
upgrades¶
- upgrade the server first, then the agents, and keep them on the same version. a newer server can talk to older agents, but the reverse is not guaranteed
- take a cold copy of the server's data before upgrading it
- renovate proposes bumps for both only through its dependency dashboard, so neither restarts until i approve it
checking it¶
on a manager:
REPLICAS reads 1/1 for the server and 3/3 for the agent.
This page started as a gist: the original, with its comments