Change records
Header-Defined Swarm Stack Production Rollout
Production migration to the stack namespaces declared by each Swarm manifest header.
Change metadata
- Change id:
CR-2026-08-14-swarm-header-stack-rollout - Date and time: 2026-08-14 19:29–20:03 UTC
- Environment: VPS / production
- Requester and approver: repository owner, explicit approval in the active session
- Executor: Codex
- Status: completed; known pre-existing application and external follow-ups remain
Scope
Each fragment's first-line # Swarm stack: value became its authoritative
deployed namespace. Fragments with the same value are operated together:
infra-{postgres,mysql,mongodb,mssql}becamedatabase.infra-{redis,rabbitmq,watchtower}becameplatform.infra-edgebecameedge;panelsbecamepanel.svc-{gitea,rustfs,vaultwarden}becameservice.svc-kong,svc-netbird, andsvc-zitadelbecameservice-kong,service-netbird, andservice-zitadel.- The six deployed website groups moved from
svc-*towebsites-*. - Syassociates remained undeployed; only its source and launcher identity changed.
Canonical grouped launchers now always deploy the complete fragment set for
database, platform, and service. Database and WordPress backup defaults
were updated to resolve the new database_* task names.
Preconditions and rollback evidence
- Captured stacks, services, replicas, exact service image specs, volumes, networks, secrets, and timer state.
- Confirmed the latest scheduled Tier-1 job completed successfully.
- Revalidated all 14 artifacts in
tier1-20260814T185509Zby SHA-256. - Kept MSSQL explicitly at
0/0in the new manifest. - Stored sanitized before/after evidence under the ignored local directory
tmp/backups/swarm-role-layout-20260814T192920Z-header-rollout/.
Execution and validation
- Migrated each namespace only after its prior tasks stopped, preventing two identities from mounting the same external state or advertising duplicate routes.
databasereached PostgreSQL/MySQL/MongoDB1/1and MSSQL0/0; all three running task containers reported healthy.platformreached4/4; Redis returnedPONGand RabbitMQ diagnostics returnedPing succeeded.edgereached2/2; Kener returned200.panelcontains exactly eight role-labeled services and preserved its five-suspended/three-running baseline. A controlled pgAdmin0 -> 1 -> 0cycle passed and the route returned the expected non-NetBird403.- The service stacks reached their expected replicas. Kong's migration task
completed, NetBird's instance API returned
200, Zitadel returned ready200and root308, Vaultwarden and Gitea health returned200, the Gitea registry challenged with401, and RustFS remained NetBird-restricted. - All six active website stacks reached their expected replicas. A drifted
Nishat web image and its separate environment source were restored. Both
services are
1/1;nishatcolony.pkreturns200, while the intended WordPress route atnew.nishatcolony.pkretains its pre-existing500. - Both production backup dry-runs passed against
database_*; both systemd timers remain enabled and active.
Preservation checks
- Exact service image specs and desired replica values match the pre-rollout inventory.
- Named volume and Docker secret inventories are unchanged. Anonymous task volumes churned as expected when old containers were replaced.
- The only named network change is
infra-watchtower_default->platform_default. - All tested public routes and NetBird restrictions returned their expected status codes.
Additional corrections
- Pinned all four database manifests to their captured production image digests so the namespace change could not trigger an image upgrade.
- Added the explicit
kong-adminrouter-to-service binding required when Kong exposes both proxy and admin ports through Traefik. - Pinned
nishatcolony-webback to its deployed Gitea image after the live env rendered an unintended generic WordPress image during the rename. Its launcher now loads both intentional Nishat environment sources so the two hostnames do not collapse onto one router.
Follow-up
- Restore the VPS Docker client's dedicated
svc-pullerlogin and repeat an authenticated Gitea image pull; do not substitute another service token. data.arnexglobal.comstill resolves away from this VPS, so its external TLS behavior cannot validate this deployment until DNS is corrected.- Diagnose the existing
new.nishatcolony.pkWordPress HTTP 500 separately; its task is healthy and the naming rollout preserved its state and route.
Rollback
Remove the replacement namespace, wait for all of its tasks to stop, then use the pre-change revision to deploy the prior namespace against the same external volumes, networks, and secrets. Roll back one namespace group at a time and validate it before continuing. No rollback was required.