Image Update Rollout Execution Bundle
Provide one approval-ready execution path for the seven WUD-reported image updates, with DEV-validated manifest changes, per-step rollback, and validation commands.
Purpose
Provide one approval-ready execution path for the seven container image updates reported by WUD on 2026-08-14, staged into independent windows so each can be rolled back alone.
Two of the seven WUD suggestions are unsafe as reported and must not be applied verbatim. See Rejected Suggestions.
Scope
- Stack area:
swarm/stacks/ - Manifests touched:
swarm/stacks/services/gitea/gitea.ymlswarm/stacks/services/vaultwarden/vaultwarden.ymlswarm/stacks/services/rustfs/rustfs.ymlswarm/stacks/services/zitadel/zitadel.ymlswarm/stacks/services/zitadel/zitadel-login.ymlswarm/stacks/panels/wud.ymlswarm/stacks/platform/redis/redis.yml
- Launchers used:
swarm/scripts/services/service.sh,swarm/scripts/services/zitadel.sh,swarm/scripts/panels/panels.sh,swarm/scripts/platform/platform.sh
Approval Gate
- Required approver: owner
- Approval rule: explicit owner approval per window; approval of one window does not carry to the next
- Every step below is a VPS / PROD mutation under
AGENTS.md§5 and §15
Rejected Suggestions
Do not apply these two. Both were verified against the upstream registries.
Redis 8-alpine → 32bit-stretch — REJECT
redis:32bit-stretch was published 2019-07-04 and is a 32-bit Debian Stretch build of the Redis 5.x line. Applying it is a three-major-version downgrade; a Redis 8 AOF/RDB will not load under Redis 5, so the outcome is data loss rather than a downgrade.
Root cause is a missing label on the running container, not a bad manifest. swarm/stacks/platform/redis/redis.yml already carries the correct guardrail:
- wud.watch=true
- wud.tag.include=^\d+(\.\d+(\.\d+)?)?-alpine$$32bit-stretch does not match that pattern, so the running container is not carrying the label. Reconcile the running container with the manifest (Window 1) rather than applying the update.
The genuine Redis update is a digest bump on the same tag and is covered in Window 5.
Gitea 1.22 → 1.27-rootless — REJECT the tag, ACCEPT the upgrade
Production runs the non-rootless image. The two variants are not interchangeable — image configs pulled from Docker Hub:
1.27.2 | 1.27.2-rootless | |
|---|---|---|
| Cmd | /usr/bin/s6-svscan /etc/s6 | (none) |
| Entrypoint | /usr/bin/entrypoint | dumb-init -- docker-entrypoint.sh |
| Volumes | /data | /etc/gitea, /var/lib/gitea |
GITEA_CUSTOM | /data/gitea | /var/lib/gitea/custom |
| User | root → git | 1000:1000 |
The rootless image would ignore the scm_gitea_data volume mounted at /data and would not resolve the menu.tmpl bind at /data/gitea/templates/. Correct target is gitea/gitea:1.27.2 (non-rootless), covered in Window 4.
DEV Validation Already Performed
Executed 2026-08-14, DEV only, using a Compose rehearsal that mirrors the production manifest env and command wrapper exactly (postgis/postgis:18-3.6 + Gitea, same secret-injection wrapper).
Finding — latent bug in the Gitea command wrapper. s6-svscan moved between releases:
| Image | Path present |
|---|---|
gitea/gitea:1.22 | /bin/s6-svscan only |
gitea/gitea:1.27.2 | /usr/bin/s6-svscan only |
The manifest hardcoded /bin/s6-svscan, so the first 1.27.2 boot failed with exec: /bin/s6-svscan: not found, exit code 127 — a crash loop into max_attempts: 3. Fixed by letting PATH resolve the binary; verified to resolve correctly in both images, so the change is safe on the currently deployed 1.22.
This fix is already applied to swarm/stacks/services/gitea/gitea.yml and is a prerequisite for Window 4.
Rehearsal results after the fix, 1.22 → 1.27.2 as a single direct jump:
- Container reached
healthy; no[E],fatal, orpaniclines - Schema migrated 299 → 343 with no manual intervention
- Data preserved: 2 users, 1 package, 1 package version — unchanged across the upgrade
- Generic package downloaded successfully with byte-identical content
/api/healthzreturned{"status":"pass"};gitea --versionreported 1.27.2/v2/correctly returned401(Docker registry auth challenge intact)- Repo-unit config persisted correctly in
app.ini, including the load-bearingDEFAULT_FORK_REPO_UNITS = repo.code - Custom
org/menu.tmplrendered against 1.27.2: org page200, Packages tab present, no template errors
Finding — WUD wrapper needs no change. Upstream changed Cmd from node dist/index.js to node dist/index between 8.2.2 and 8.3.1, but dist/index.js still exists in 8.3.1 (Node resolves the extensionless path to it). A DEV run of 8.3.1 with the exact production command wrapper booted cleanly and registered the custom pv registry from the injected secret. Window 2 is a pure image bump.
Preconditions
- Owner approval captured in a change record instance.
- Window 0 completed — running stack names reconciled against the repo.
- Free disk verified for volume snapshots (
df -h /var/lib/docker). - Tier-1 backup timer state known (
systemctl status tier1-db-backup.timer).
Window 0 — Reconcile running stack names (read-only, blocking)
WUD reports containers named infra-platform_redis, infra-gitea_gitea, infra-operations_wud, infra-object-storage_rustfs, infra-vaultwarden_vaultwarden, infra-zitadel_zitadel. The repository's authoritative # Swarm stack: headers declare platform, service, panel, and service-zitadel. swarm/ is clean in Git, so production appears to still run the pre-rename infra-* stacks.
This is blocking: every later window uses docker service update, whose service names depend on the answer, and running a launcher against a renamed stack would create a second stack rather than updating the existing one.
# VPS / PROD — read-only
docker stack ls
docker service ls --format '{{.Name}}\t{{.Image}}'
docker service inspect infra-platform_redis \
--format '{{json .Spec.Labels}}' 2>/dev/null | tr ',' '\n' | grep -i wudOutcomes:
- Names match the repo (
platform_redis,service_gitea, …) — WUD's store is stale. Force a rescan (POST /api/containers/watch) and use the repo names in all later windows. - Names are still
infra-*— production predates the rename rollout. Use theinfra-*names in all later windows, and treat the rename as a separate migration; do not fold it into an image update.
Record the answer before proceeding. Do not run any stack launcher until this is settled.
Window 1 — Redis label reconcile (no image change)
Restores the wud.tag.include guardrail on the running container so the 32bit-stretch suggestion stops recurring. No image change; Redis restarts.
# VPS / PROD — requires approval. `<SVC>` from Window 0.
docker service update \
--label-add 'wud.watch=true' \
--label-add 'wud.tag.include=^\d+(\.\d+(\.\d+)?)?-alpine$' \
<SVC_redis>Validation:
docker service ls --filter name=<SVC_redis> # expect 1/1
docker exec $(docker ps -qf name=<SVC_redis>) \
redis-cli -a "$(cat /run/secrets/redis_password)" ping # expect PONGRollback: docker service update --label-rm wud.tag.include --label-rm wud.watch <SVC_redis>
Redis holds cache/session data under allkeys-lru; a restart is tolerable but will drop the cache. Confirm no Kener/app workload is mid-critical first.
Window 2 — WUD 8.2.2 → 8.3.1
Lowest risk: monitoring only, no persistent application data. Do this first among image bumps so the updated detector reports the rest.
# swarm/stacks/panels/wud.yml
- image: getwud/wud:8.2.2@sha256:2df7cafa0fc84cdd66ad976163abc5390fa5995e0b5788a90e6a5303706bc21a
+ image: getwud/wud:8.3.1@sha256:42554d2cd50d5de7a47563170bd0629193dd1e086a1fd71d1ec10fc4c2409f5dThe command: wrapper stays unchanged — DEV-verified above.
# VPS / PROD — requires approval
docker service update --image \
getwud/wud:8.3.1@sha256:42554d2cd50d5de7a47563170bd0629193dd1e086a1fd71d1ec10fc4c2409f5d \
<SVC_wud>Validation:
docker service ls --filter name=<SVC_wud> # expect 1/1
docker service logs --tail 40 <SVC_wud> | grep -i 'version = 8.3.1'
docker service logs --tail 80 <SVC_wud> | grep -i 'registry.custom.pv' # custom registry registeredFrom a NetBird client, confirm wud.perspective-v.com loads and the container inventory is populated.
Rollback: docker service update --rollback <SVC_wud>
Window 3 — Vaultwarden 1.36.0 → 1.37.1
Entrypoint (/start.sh), Cmd and volume (/data) are byte-identical between the two images, so the ADMIN_TOKEN wrapper stays valid.
Pre-flight — Vaultwarden is backed by vaultwarden_db in PostgreSQL and may carry migrations:
# VPS / PROD — requires approval
docker exec $(docker ps -qf name=postgres) \
pg_dump -U postgres -Fc vaultwarden_db > /var/backups/vaultwarden_db-$(date -u +%Y%m%dT%H%M%SZ).dump# swarm/stacks/services/vaultwarden/vaultwarden.yml
- image: vaultwarden/server:1.36.0@sha256:d626d04934cd1192ad8ced1adb975099fca78cec33ab467d2d3c923cde7f3b0c
+ image: vaultwarden/server:1.37.1@sha256:ebdfe70701c60ac0c28c697e787cea767d7972940b786037b29fe0d507f821e8docker service update --image \
vaultwarden/server:1.37.1@sha256:ebdfe70701c60ac0c28c697e787cea767d7972940b786037b29fe0d507f821e8 \
<SVC_vaultwarden>Validation: service 1/1; curl -sI https://${VAULT_HOST}/alive returns 200; one interactive vault unlock from a browser; admin page still gated by the injected ADMIN_TOKEN.
Rollback: docker service update --rollback <SVC_vaultwarden>, then restore the dump only if schema migration is confirmed to have run and failed.
Window 4 — Gitea 1.22 → 1.27.2
Five minor versions, 44 schema migrations. DEV-rehearsed end to end (see above). Requires its own window — this service backs every image pull on the node.
Pre-flight, in order:
# VPS / PROD — requires approval
# 1. Database dump
docker exec $(docker ps -qf name=postgres) \
pg_dump -U postgres -Fc gitea_db > /var/backups/gitea_db-$(date -u +%Y%m%dT%H%M%SZ).dump
# 2. Volume snapshot — scm_gitea_data holds every package blob and is NOT
# covered by operations/backups/. This is the only copy.
docker run --rm -v scm_gitea_data:/src:ro -v /var/backups:/dst alpine \
tar czf /dst/scm_gitea_data-$(date -u +%Y%m%dT%H%M%SZ).tar.gz -C /src .
# 3. Record current digest for rollback
docker service inspect <SVC_gitea> --format '{{.Spec.TaskTemplate.ContainerSpec.Image}}'Manifest changes — the wrapper fix is already applied; the image bump is not:
# swarm/stacks/services/gitea/gitea.yml
- image: gitea/gitea:1.22@sha256:538658de667c5d098a274f2f63aa6ec891d88f670cdd5282cf27221ba747dda4
+ image: gitea/gitea:1.27.2@sha256:d20286ca2b2e170fdf628e7231b8a31a3220ade39ff462b55041d43d1fc757dd
command:
- /bin/sh
- -c
- - export GITEA__database__PASSWD="$$(cat /run/secrets/gitea_db_password)"; exec /bin/s6-svscan /etc/s6
+ - export GITEA__database__PASSWD="$$(cat /run/secrets/gitea_db_password)"; exec s6-svscan /etc/s6Apply both together — the image bump without the wrapper fix produces exit 127.
docker service update \
--image gitea/gitea:1.27.2@sha256:d20286ca2b2e170fdf628e7231b8a31a3220ade39ff462b55041d43d1fc757dd \
--args '/bin/sh,-c,export GITEA__database__PASSWD="$(cat /run/secrets/gitea_db_password)"; exec s6-svscan /etc/s6' \
<SVC_gitea>Validation, in order — stop and roll back on any failure:
docker service ls --filter name=<SVC_gitea> # 1/1
docker service logs --tail 100 <SVC_gitea> | grep -iE 'fatal|panic|\[E\]|no default' # expect none
curl -s https://gitea.perspective-v.com/api/healthz # {"status":"pass"}
curl -so /dev/null -w '%{http_code}\n' https://gitea.perspective-v.com/v2/ # 401
docker exec $(docker ps -qf name=<SVC_gitea>) gitea --version # 1.27.2Then confirm data and the package path:
# schema should have advanced past 299
docker exec $(docker ps -qf name=postgres) psql -U gitea_db_usr -d gitea_db \
-tAc "select version from version; select count(*) from package_version;"
# authenticated pull of a real service image
docker pull gitea.perspective-v.com/perspective-v/dbskc-web:latestFinally, load https://gitea.perspective-v.com/perspective-v in a browser and confirm the custom org menu still shows only Packages / Members / Teams / Settings.
Rollback:
docker service update --rollback <SVC_gitea>The --rollback restores the 1.22 image and previous args. The schema is not rolled back — 1.22 cannot run against schema 343. If the upgrade must be reverted after migrations have run, restore gitea_db from the dump and the scm_gitea_data volume from the tarball, then redeploy 1.22.
Known constraint from prior work: env-to-ini writes app.ini onto the volume and never removes keys deleted from the environment. If a bad setting needs clearing, the volume must be wiped and the schema reset — which is why the volume snapshot above is mandatory, not optional.
Window 5 — Zitadel core + login → v4.17.1
WUD's two Zitadel entries are the same upgrade. The reported "running digest" cf6c2e88… is the amd64 manifest inside the v4.15.2 index the repo pins, and the suggested d38f9217… is the amd64 manifest of v4.17.0 — so there is no repo/production drift, only WUD reporting platform digests instead of index digests.
Update both services in lockstep. Version skew between core and login UI breaks the sign-in flow, and the /ui/v2/login routing is hand-built across both manifests.
v4.17.1 (2026-08-14) supersedes WUD's v4.17.0 suggestion and is the recommended target.
Pre-flight:
# VPS / PROD — requires approval
docker exec $(docker ps -qf name=postgres) \
pg_dump -U postgres -Fc zitadel > /var/backups/zitadel-$(date -u +%Y%m%dT%H%M%SZ).dump# swarm/stacks/services/zitadel/zitadel.yml
- image: ghcr.io/zitadel/zitadel:v4.15.2@sha256:1e10033861f04fd2d5371f838ff97b44e3ca86789401d2f9b9fbd560eafba321
+ image: ghcr.io/zitadel/zitadel:v4.17.1@sha256:3ac6910685d48f32481f01f45e3e6215efe5a9df2c069591b481e9a101712db5
# swarm/stacks/services/zitadel/zitadel-login.yml
- image: ghcr.io/zitadel/zitadel-login:v4.15.2@sha256:06832e5dcbd98d4ff9db55087f5c4b3e2ca69a6958ad68fa3c335fb8d4f3abe6
+ image: ghcr.io/zitadel/zitadel-login:v4.17.1@sha256:8035df2409afb35a3999482ee98e453261715f98d47e4b62e948e4a1ddf4345fUpdate core first, let migrations settle, then the login UI:
docker service update --image \
ghcr.io/zitadel/zitadel:v4.17.1@sha256:3ac6910685d48f32481f01f45e3e6215efe5a9df2c069591b481e9a101712db5 \
<SVC_zitadel>
# wait for 1/1 and migration completion in logs, then:
docker service update --image \
ghcr.io/zitadel/zitadel-login:v4.17.1@sha256:8035df2409afb35a3999482ee98e453261715f98d47e4b62e948e4a1ddf4345f \
<SVC_zitadel_login>Validation: both services 1/1; curl -sI https://${ZITADEL_EXTERNALDOMAIN}/debug/healthz returns 200; https://${ZITADEL_EXTERNALDOMAIN}/ redirects to /ui/v2/login/; one full interactive login completes.
Rollback: docker service update --rollback on both, newest first. As with Gitea, schema migrations are not reversed — restore the dump if v4.15.2 refuses the migrated schema.
Window 6 — RustFS beta.6 → rc.1
Highest data risk; schedule last and alone.
The jump crosses beta.7 through beta.12 and then rc.1 on a pre-1.0 product. RustFS is also the Tier-1 backup destination, so a failure here removes the ability to restore anything else — never run this in the same window as another change.
Additional pre-flight beyond the others:
- Read the RustFS changelog for beta.7 → rc.1 for on-disk format or API changes.
- Confirm the Tier-1 backup timer has a recent successful run, and that at least one backup set exists off RustFS.
swarm/stacks/services/rustfs/rustfs.ymloverridesentrypoint: ["/bin/sh","-c"]with a custom command — verify the rc.1 image still exposes the same binary path and flags before applying.
# swarm/stacks/services/rustfs/rustfs.yml
- image: rustfs/rustfs:1.0.0-beta.6-glibc@sha256:c6c65aa5010f923452d8057cb4084b2238245b2ce740b075f082e2db2ee39cd6
+ image: rustfs/rustfs:1.0.0-rc.1-glibc@sha256:b97262df87e0d9d4b4e925fbaf892c8083e54732f6388e8036b9992fb91930baValidation: service 1/1; console reachable over NetBird at ${RUSTFS_CONSOLE_HOST}; API returns an auth gate at ${RUSTFS_API_HOST}; list an existing bucket and download one known object; confirm non-NetBird access still returns 403.
Rollback: docker service update --rollback <SVC_rustfs>. If the on-disk format was migrated, rollback requires restoring /var/lib/rustfs/data from a pre-change snapshot — take one before starting.
Out Of Scope
WUD reported 7 updates; a direct registry sweep on 2026-08-14 found roughly 12 further digest-level updates it did not report — traefik, mysql, rabbitmq, postgis, mongo, kong, portainer-ee, pgadmin4, redisinsight, both NetBird images, and mssql. The likely cause is those containers lacking wud.watch=true. Reconcile WUD's watch coverage after Window 0 resolves the naming question, then handle those separately.
Three images are current only because upstream stopped publishing: pantsel/konga (last push 2020-05-16), containrrr/watchtower (2023-11-11), and mongo-express (2024-05-22). Replacement is a separate decision, not an update.
Post-Execution
- Record each executed window in a change record instance from
docs/operations/templates/production-change-record-template.mdx. - Update
docs/state/already-implemented.mdxwith the versions actually deployed. - Update
docs/state/next-steps.mdxto remove completed windows and keep the deferred ones accurate. - Remove
/var/backupsartifacts only after the rollback window closes; never commit them.
Execution And Evidence
Execution bundles and verification records that support controlled changes and restore confidence.
Monitoring Evaluation Execution Bundle
Approval-ready execution path for installing Cockpit, Beszel and Netdata on the host, evaluating Beszel against Netdata on real workload, and retiring the loser.