Change records
Swarm Role Layout Production Rollout
Production migration to role-organized split stacks and canonical operations tooling.
Change metadata
- Change id:
CR-2026-08-14-swarm-role-layout-rollout - Date and time: 2026-08-14 18:44–19:06 UTC
- Environment: VPS / production
- Requester and approver: repository owner, explicit approval in the active session
- Executor: Codex
- Status: completed with one pre-existing registry credential follow-up
Scope
- Split the shared platform stack into
infra-redisandinfra-rabbitmq. - Rename the updater stack to
infra-watchtower. - Reclassify RustFS, Kong, and NetBird as
svc-rustfs,svc-kong, andsvc-netbird. - Replace the administration-stack identity with
panelswhile preserving the five-suspended/three-running baseline. - Redeploy
infra-edgefrom separate Traefik and Kener fragments with service identitiesedge-traefikandedge-kener. - Install both backup timers from canonical
operations/systemd/paths and retire compatibility executables.
Preconditions and backup
- Captured stacks, replicas, images, volumes, networks, secrets, and timer state.
- Confirmed 97 GiB free disk space.
- Detected that the legacy Tier-1 unit had skipped runs since July because
its
ConditionPathExistsreferenced a nonexistent path. - Created fresh local Tier-1 backup
tier1-20260814T184617Zfor PostgreSQL, MySQL, and MongoDB; all 14 artifact checksums passed. - Kept MSSQL intentionally at
0/0and outside the scheduled engine list.
Execution and validation
- Removed
infra-platform, then deployedinfra-redisandinfra-rabbitmq; both reached1/1using their original external volumes. - Replaced
infra-operationswithinfra-watchtower; both updater scopes are healthy at1/1. Initial default-overlay allocation retries and duplicate Watchtower cleanup converged without intervention. - Migrated RustFS, Kong, and NetBird one at a time. Public behavior matched
baseline: RustFS routes
403, Kong unmatched root404, NetBird instance API200. - Corrected Kong's one-shot migration task to load its existing Docker secret; the final task completed successfully and Kong remained healthy.
- Replaced the administration stack with
panels. The role filter returns exactly eight services; pgAdmin completed a controlled0 -> 1 -> 0test and returned the expected non-NetBird403after Traefik reconciliation. - Redeployed
infra-edge; Kener returned200, while Kong and NetBird routes continued returning their expected statuses. - Installed and enabled both canonical backup timers. The persistent Tier-1
timer immediately ran the missed job successfully as
tier1-20260814T185509Z; next schedules are Sunday 03:00 UTC for databases and the last day of the month at 03:00 UTC for WordPress.
Preservation checks
- Production Docker volume inventory is unchanged.
- Production Docker secret inventory is unchanged.
- The only network inventory change is the expected updater default-overlay
rename to
infra-watchtower_default. - Migrated floating-tag workloads retained their captured image digests.
- Redis returned
PONG; RabbitMQ diagnostics returnedPing succeeded. - Vaultwarden health
200, Zitadel readiness200and root308, Gitea health200and registry challenge401.
Follow-up
- An authenticated pull from the Gitea Docker registry still returns
unauthorizedbecause the VPS Docker client's dedicatedsvc-pullercredential is absent or expired. Restore that credential without reusing another service token, then repeat the pull.
Rollback
Each migrated stack can be removed and its prior identity redeployed from the pre-change revision against the same external volumes, networks, and secrets. No rollback was required.