Perspective V Docs

next-steps

**Documentation Alignment Complete:** All references in state files (next-steps.md, already-implemented.md, requirements.md, accepted-suggestions.md) have been updated to reflect the new folder structure with infrastructure stack separation (`runtime/stacks/infrastructure/*` and `runtime/environments/{dev,vps}/infrastructure/*`). Test environment references have been removed from operational docs (dev-only for templates). Legacy path patterns have been eliminated.

Next Steps

Current Snapshot (2026-04-28)

Documentation Alignment Complete: All references in state files (next-steps.md, already-implemented.md, requirements.md, accepted-suggestions.md) have been updated to reflect the new folder structure with infrastructure stack separation (runtime/stacks/infrastructure/* and runtime/environments/{dev,vps}/infrastructure/*). Test environment references have been removed from operational docs (dev-only for templates). Legacy path patterns have been eliminated.

  • Runtime operator guide set added under docs/runtime for master setup flow, per-stack rollout, environments, CI, service deployments, and wrapper-script usage.

  • Cross-platform launcher scripts added under runtime/scripts for active infrastructure and service compose targets.

  • Edge stack operational behind Traefik with NetBird-only admin access model.

  • NetBird core deployed and integrated with Existing Traefik routing.

  • Database engines and DB admin UIs are deployed from runtime/stacks/infrastructure/databases.

  • Databases runbook includes explicit stop/remove commands for all DB and DB UI containers.

  • Git ignore policy tracks named env files under runtime/environments while keeping plain .env files ignored.

  • Tracked named env files under runtime/environments/vps now keep example-only password/secret placeholders while preserving host/domain values.

  • Runtime structure migration started (docs/docker -> runtime/stacks, docs/ci -> runtime/ci, non-production templates -> runtime/environments/dev, and active VPS env values -> runtime/environments/vps).

  • Watchtower split to runtime/stacks/infrastructure/operations is deployed on server with Docker API compatibility fix and active schedule.

  • Shared service stacks are deployed in enforced order (platform -> operations -> registry -> feeds).

  • Registry basic-auth cutover (Traefik middleware + Registry Admin basic mode + NetBird-only UI) deployed and validated.

  • Object storage baseline is RustFS from runtime/stacks/infrastructure/object-storage/rustfs; MinIO has been decommissioned.

  • Database and Redis cutover to strict NetBird-only still pending final validation.

  • Redis developer hostname access model (redis.perspective-v.com over NetBird-only firewall) applied on server; final NetBird/non-NetBird client validation still pending.

  • Split admin-panel auth policy applied in runtime/stacks/infrastructure/edge compose (dashboard keeps Traefik basic auth; built-in-login panels use app auth).

  • Edge base stack artifacts are split into docker-compose.traefik.yml and docker-compose.portainer.yml with per-service env folders under runtime/environments/vps/infrastructure/edge/{traefik,portainer,kener}.

  • Edge dev env templates are split into runtime/environments/dev/infrastructure/edge/{traefik,portainer,kener}.

  • Recreate Traefik and Portainer from split edge artifacts (docker-compose.traefik.yml + docker-compose.portainer.yml) in the approved change window.

  • Kener parallel-pilot repository artifacts prepared in runtime/stacks/infrastructure/edge (docker-compose.kener.yml + README updates), pending server deployment.

  • Owner approved Kener pilot rollout; execution is deferred to direct VPS compose session after docs update pass.

  • Top-level folder rename completed: planning -> docs and deployed -> prod, with docs path references updated.

  • Set Traefik ACME runtime path to host storage (/var/lib/traefik/letsencrypt/acme.json) with file mode 600.

  • Recreate Traefik from runtime/environments/vps/infrastructure/edge/traefik/.env and verify ACME entries/routes persist after restart.

  • Retire prod/ folder after runtime sync verification.

  • From a NetBird-connected client, validate Redis access:
nc -vz redis.perspective-v.com 6379
redis-cli -h redis.perspective-v.com -p 6379 -a "<REDIS_PASSWORD>" ping
  • From a non-NetBird client, validate Redis denial:
nc -vz redis.perspective-v.com 6379
  • Confirm production services still use internal Docker hostname redis in runtime env/config.
  • Validate the new runtime/scripts/*.sh launchers from a Linux shell on the VPS before using them in maintenance workflows.
  • Add tracked placeholder and dev-template env files for wp.dbskc.com and new.nishatcolony.pk if those variants need reproducible non-VPS deployments.
  • Complete final encrypted backup activation: four-site dry-run/real workflow, complete active-engine logical set, first encrypted PostgreSQL physical snapshot, isolated physical restore, scratch cleanup, offsite flags, and all three timers.
  • Complete encrypted canary upload/cryptcheck/download validation, one WordPress scratch restore, and one Tier-1 PostgreSQL scratch restore before setting DB_OFFSITE_ENABLED=true or changing production timers.
  • Monitor the first naturally scheduled WordPress and PostgreSQL physical runs and retain their sanitized evidence. Review the preserved failed physical staging only after the owner decides whether it should be removed.
  • Fix store.gorsistudio.com's service compose env_file (currently points at the parent gorsistudio.com env file) so the running container and backup registry agree on gorsistudio_store_db.
  • After Redis validation, run remaining registry/feed CI-path checks in Phase 2.5.
  • Recreate service stacks in runtime/stacks/services after Watchtower label policy update so running containers include com.centurylinklabs.watchtower.scope=fast.
  • Recreate Traefik-exposed service stacks after label hardening so running containers include traefik.docker.network=proxy and current TLS router settings.
  • Validate public service routes after recreate (dbskc.com, nishatcolony.pk, console.perspective-v.com) return HTTPS responses with valid Let's Encrypt certificates.
  • Validate container health status is healthy after recreate for stacks updated with new healthchecks.
  • Confirm no remaining service stack healthchecks rely on bash for Alpine/nginx-based images.
  • Recreate service containers that should report com.docker.compose.project=perspective-v if they are still running with legacy project labels.
  • Recreate only the NetBird dashboard container so updated embedded-IdP env is loaded:
cd /opt/docs/runtime/stacks/infrastructure/netbird
docker compose -f docker-compose.yml up -d --force-recreate dashboard
  • Validate NetBird dashboard login no longer returns Error: Unauthenticated.

  • Update CI build/push jobs in runtime/ci to emit Docker schema v2 compatible manifests (oci-mediatypes=false, disable provenance/SBOM attestations when needed for compatibility).

  • Re-push active service tags in Docker schema v2 compatible format so Registry Admin catalog remains visible.

  • For registry basic-auth cutover, deploy registry-admin first, then registry.

  • Validate API challenge header:

curl -Ik https://registry.perspective-v.com/v2/
  • Expected challenge: WWW-Authenticate: Basic realm="traefik".

Phase 0: Fresh VPS Baseline

  • Place this docs repository on VPS at /opt/docs (or finalized PLAN_DIR).
  • Verify Docker and Docker Compose versions.
  • Install host prerequisites (curl, jq, ufw, netcat-openbsd, dnsutils).
  • Confirm DNS A and AAAA for netbird.perspective-v.com to VPS.
  • Apply base UFW policy (public: 80/tcp, 443/tcp, 3478/udp only).

Phase 1: NetBird Core Bring-up

  • Deploy edge stack from runtime/stacks/infrastructure/edge and verify Traefik is active on 80/443.
  • Install NetBird via quickstart script in /opt/netbird.
  • During script, select Existing Traefik option [1].
  • Keep NetBird proxy service disabled.
  • Create first admin user from /setup.
  • Recreate NetBird dashboard service after embedded-IdP env correction (AUTH_CLIENT_SECRET=).
  • Validate NetBird dashboard interactive login succeeds without Error: Unauthenticated.
  • Join all required developer/admin machines to NetBird.
  • Detect and record active NetBird interface name on VPS (expected wt0).

Phase 2: Private Admin Panels (NetBird-Only)

  • Apply split admin-route policy in edge stack (Traefik dashboard keeps admin-auth; Portainer uses NetBird allowlist + app login).
  • Validate from NetBird client: dashboard returns 401 before credentials, while Portainer shows app login without Traefik basic-auth challenge (server-side NetBird-address path validation recorded in docs/operations/validations/2026-04-15-netbird-route-validation.mdx).
  • Validate admin route denial from non-NetBird host after split policy rollout (recorded in docs/operations/validations/2026-04-15-netbird-route-validation.mdx).

Phase 2.1: Kener Public Cutover (Built-in Auth)

  • Add Kener overlay compose artifact at runtime/stacks/infrastructure/edge/docker-compose.kener.yml.
  • Update edge runbook with Kener pilot launch flow and required Kener env variable documentation.
  • Populate server-local runtime/environments/vps/infrastructure/edge/kener/.env with KENER_HOST, KENER_ORIGIN, KENER_SECRET_KEY, and KENER_REDIS_URL.
  • Populate server-local runtime/environments/vps/infrastructure/edge/kener/.env with SMTP_HOST, SMTP_PORT, SMTP_USER, SMTP_PASSWORD, SMTP_FROM_EMAIL, and SMTP_SECURE.
  • Ask owner approval for Kener rollout model change.
  • Execute approved Kener deployment on VPS using edge stack compose artifacts.
  • Validate from public source that kener.perspective-v.com returns Kener app route (HTTP 200).
  • Validate Kener health endpoint on kener.perspective-v.com/healthcheck returns ok.
  • Retire Uptime Kuma runtime artifacts (service/container/volume/image) and remove Kuma service from edge stack artifacts.
  • Keep fallback Kener NetBird middleware configuration commented in runtime/stacks/infrastructure/edge/docker-compose.kener.yml.
  • Resolve Kener startup warning by deciding Redis eviction policy (allkeys-lru observed; Kener recommends noeviction) within current Platform resource constraints.
  • Create Kener owner admin account and baseline pages for monitor separation.
  • Manually recreate critical Uptime Kuma monitors in Kener (phase-1 subset).
  • Run 3-7 day post-cutover Kener alert fidelity observation window.

Phase 2.5: Shared Services Baseline (Platform -> Operations -> Registry -> Feeds)

  • Prepare platform.env from runtime/environments/vps/infrastructure/platform/platform.env.
  • Deploy runtime/stacks/infrastructure/platform stack (Redis + RabbitMQ).
  • Prepare operations.env from runtime/environments/vps/infrastructure/operations/operations.env.
  • Deploy runtime/stacks/infrastructure/operations Watchtower stack and verify scheduled label-based execution.
  • Add dual Watchtower runtime definitions in operations stack (watchtower baseline scope none + watchtower-fast scope fast, 1-hour poll interval).
  • Add WUD compose service in runtime/stacks/infrastructure/operations with pinned image, Discord trigger variables, and NetBird-only Traefik labels for wud.perspective-v.com.
  • Add WUD non-production placeholders in runtime/environments/dev/infrastructure/operations/operations.dev.env.
  • Add live WUD values in server-local runtime/environments/vps/infrastructure/operations/.env (set real Discord webhook, keep tracked env placeholders secret-free).
  • Deploy/recreate operations stack to start WUD container alongside Watchtower.
  • Recreate operations stack on VPS to start watchtower-fast and apply baseline/fast scope split.
  • Validate scope separation on VPS: watchtower-fast updates only com.centurylinklabs.watchtower.scope=fast, while baseline watchtower handles default scope none.
  • Validate WUD non-NetBird denial from server-source check (HTTP 403).
  • Validate WUD route from a separate NetBird client at wud.perspective-v.com.
  • Validate Discord notification delivery using a controlled WUD trigger test.
  • Configure WUD custom private registry provider for registry.perspective-v.com with authenticated access.
  • Validate WUD now lists container inventory and shows custom.pv on the registries page/API.
  • Create dedicated wud-monitor registry basic-auth credential and wire WUD to use it instead of admin.
  • Recreate registry and operations stacks after credential update and validate WUD still reports custom.pv successfully.
  • Keep WUD in discovery mode on VPS (WUD_WATCHER_LOCAL_WATCHBYDEFAULT=true) and retain selected wud.tag.include guardrails for mysql/rabbitmq/redis/verdaccio.
  • Trigger immediate WUD full scan (POST /api/containers/watch) after mode switch and validate full inventory is visible (4 -> 25 containers, including dbskc-web).
  • Decide long-term WUD scope policy after one week of alert-noise observation (stay discovery-default or move back to opt-in labels).
  • Execute staged major upgrades with rollback checkpoints and pre-change volume backups for Verdaccio (5 -> 6), Redis (7-alpine -> 8-alpine), RabbitMQ (3.13-management-alpine -> 4.2-management-alpine), and MySQL (8.4 -> 9.6-oraclelinux9).
  • Validate post-upgrade service health and core runtime checks (Verdaccio HTTP ping 200, Redis authenticated PING, RabbitMQ broker status, MySQL 9.6 healthy with working root/app auth).
  • Run dependent application smoke tests against upgraded Redis/RabbitMQ/MySQL/Verdaccio clients before closing the maintenance window.
  • Prepare registry.env from runtime/environments/vps/infrastructure/registry/registry.env.
  • Initialize shared registry auth file /var/lib/traefik/registry-auth/registry.htpasswd on VPS and verify file permissions are restricted.
  • Validate Traefik registry auth now uses file-provider middleware (registry-basic-auth@file) from the shared htpasswd source.
  • Deploy basic-auth registry cutover (registry-admin + registry) from runtime/stacks/infrastructure/registry.
  • Validate registry API returns HTTP Basic challenge.
  • Validate docker login, push, and pull on registry API with admin credential.
  • Validate CI service credentials can login/push/pull after cutover (initial scoped-user validation completed).
  • Validate creating/updating a user in Registry Admin updates shared htpasswd and enables docker login without registry stack recreate (validated for nishatcolony-ci after Traefik auth-state refresh via docker restart traefik).
  • Finalize post-user-change auth refresh runbook as manual Traefik restart (Portainer restart accepted); no helper automation required in this phase.
  • Retire legacy docker-registry-ui service from prod registry stack and remove registry-ui host route.
  • Validate registry API access from non-NetBird CI runner using authenticated docker login.
  • Validate registry API rate-limit settings (average/burst) allow normal CI push throughput.
  • Validate registry-admin UI is NetBird-only and denied from non-NetBird source (HTTP 403 expected).
  • Validate retired registry-ui hostname is no longer routed after service removal (HTTP 404 observed).
  • Update CI templates in runtime/ci to push Docker schema v2 compatible manifests for Registry Admin (oci-mediatypes=false).
  • Pin registry runtime image target to registry:3 in runtime/stacks/infrastructure/registry/docker-compose.registry.yml for v3 migration implementation start.
  • Add operator-safe registry cleanup helper script (runtime/stacks/infrastructure/registry/scripts/registry-cleanup.sh) with dry-run default, tag delete, repository purge, and GC guidance.
  • Add registry cleanup policy template (runtime/stacks/infrastructure/registry/cleanup-policy.json) with protected repo/tag controls and monthly include/exclude targeting.
  • Document registry cleanup and garbage-collection workflow in runtime/stacks/infrastructure/registry/README.md.
  • Add registry staging rehearsal helper script (runtime/stacks/infrastructure/registry/scripts/staging-rehearsal.sh) for standardized v3 pre-cutover checks.
  • Add registry test env template at runtime/environments/dev/infrastructure/registry/registry.dev.env for non-production rehearsal values.
  • Capture pre-cutover rollback snapshots for registry volumes and image references at tmp/backups/registry-v3-cutover-20260413-225028.
  • Execute owner-approved direct VPS cutover for registry service to registry:3 from runtime/stacks/infrastructure/registry.
  • Validate immediate post-cutover checks on VPS (/v2/ challenge header, registry container running registry:3, and registry log activity on existing repositories).
  • Populate registry monthly cleanup policy include/exclude rules for real repositories before first apply-mode run.
  • Validate monthly registry cleanup job behavior on VPS in dry-run mode first, including protected repository/tag exclusions.
  • Re-push required service tags in Docker schema v2 compatible format and confirm Registry Admin catalog visibility.
  • Validate RabbitMQ UI from NetBird reaches its built-in login without Traefik basic-auth challenge (recorded in docs/operations/validations/2026-04-15-netbird-route-validation.mdx).
  • Prepare feeds.env from runtime/environments/vps/infrastructure/feeds/feeds.env.
  • Deploy runtime/stacks/infrastructure/feeds stack (BaGet + Verdaccio).
  • Validate NuGet restore from non-NetBird source at nuget.perspective-v.com (HTTP 200).
  • Retire legacy feed bridge service/routes from runtime/stacks/infrastructure/feeds/docker-compose.feeds.yml and keep canonical feeds only.
  • Simplify feed env contracts in runtime/environments/dev/infrastructure/feeds/feeds.dev.env and runtime/environments/vps/infrastructure/feeds/feeds.env to canonical keys only.
  • Add NPM_FEED_API_KEY to feed env artifacts (runtime/environments/dev/infrastructure/feeds/feeds.dev.env, runtime/environments/vps/infrastructure/feeds/feeds.env, and live runtime/environments/vps/infrastructure/feeds/.env).
  • Remove legacy feed bridge bootstrap helper from runtime/stacks/infrastructure/feeds/scripts.
  • Normalize feed runbook to canonical endpoint operations in runtime/stacks/infrastructure/feeds/README.md.
  • Apply Verdaccio auth hardening in runtime/stacks/infrastructure/feeds/verdaccio/conf/config.yaml (max_users=-1, access=$all, publish/unpublish=$authenticated) on VPS runtime.
  • Validate unknown-user npm adduser self-registration is rejected on https://npm.perspective-v.com.
  • Validate NuGet restore from CI runner using https://nuget.perspective-v.com/v3/index.json (owner-confirmed; see docs/operations/validations/2026-04-16-ci-nuget-feed-token-validation.mdx).
  • Validate NuGet publish on BaGet using API-key model.
  • Validate npm install/publish from CI using https://npm.perspective-v.com/ with NPM_FEED_API_KEY (owner-confirmed; see docs/operations/validations/2026-04-16-ci-npm-feed-api-key-validation.mdx).
  • Validate canonical endpoints remain available for CI and developer workflows (nuget.perspective-v.com, npm.perspective-v.com).
  • Apply npm CI variable/secret cutover in service repositories to NPM_FEED_URL/NPM_FEED_API_KEY using the same value as runtime/environments/vps/infrastructure/feeds/.env.
  • Apply NuGet CI variable cutover in service repositories to NUGET_FEED_URL/NUGET_FEED_TOKEN.
  • Retire legacy feed bridge after Verdaccio auth validation, CI cutover completion, and package inventory checks.
  • Recreate VPS feeds stack from canonical compose using live runtime/environments/vps/infrastructure/feeds/.env and --remove-orphans to remove running legacy proget container.
  • Capture rollback snapshot/checksum for legacy ProGet volume before cleanup at tmp/backups/feeds-proget-retirement-20260416T180458Z.
  • Remove residual legacy ProGet VPS artifacts (feeds_proget_packages volume and proget.inedo.com/productimages/inedo/proget:25.0.25 image).
  • Enforce stack order for baseline rollout: platform -> operations -> registry -> feeds.
  • Validate only stateless services keep com.centurylinklabs.watchtower.enable=true.
  • Audit Redis consumers and confirm platform Redis is cache/session/ephemeral only.
  • If any critical Redis workload exists, plan migration to a dedicated stateful Redis stack with backup/restore tests.
  • Create DNS record for redis.perspective-v.com to VPS.
  • Apply NetBird-only Redis firewall rule (6379 on NetBird interface, no broad public allow).
  • Validate Redis connectivity from NetBird client at redis.perspective-v.com:6379.
  • Validate Redis denial from non-NetBird source at redis.perspective-v.com:6379.
  • Confirm production service connection strings keep internal Redis host redis.
  • Add Redis Insight env values and deploy redis-insight from runtime/stacks/infrastructure/platform (recorded in docs/operations/validations/2026-04-15-netbird-route-validation.mdx).
  • Validate Redis Insight UI from NetBird client at redis-insight.perspective-v.com (recorded in docs/operations/validations/2026-04-15-netbird-route-validation.mdx).
  • Validate Redis Insight UI denial from non-NetBird source (HTTP 403 expected) (recorded in docs/operations/validations/2026-04-15-netbird-route-validation.mdx).

Phase 2.6: Gateway Baseline (Dev-First)

  • Provision kong_db and kong_user in existing PostgreSQL for Kong metadata storage (recorded in docs/operations/validations/2026-04-15-netbird-route-validation.mdx).
  • Add runtime/stacks/infrastructure/gateway/docker-compose.kong.yml with Kong + Konga only (no dedicated Kong DB service).
  • Add dev gateway env template at runtime/environments/dev/infrastructure/gateway/kong.dev.env with placeholder secrets.
  • Add VPS gateway env artifacts at runtime/environments/vps/infrastructure/gateway/.env and runtime/environments/vps/infrastructure/gateway/kong.env.
  • Add gateway runbook at runtime/stacks/infrastructure/gateway/README.md aligned with stack conventions.
  • Deploy gateway stack on VPS from runtime/stacks/infrastructure/gateway using runtime/environments/vps/infrastructure/gateway/.env.
  • Validate Kong migrations complete successfully against shared PostgreSQL (postgres on postgres-network).
  • Validate api.perspective-v.com route reaches Kong proxy through Traefik (HTTP 404 from Kong before route provisioning is expected).
  • Validate konga.perspective-v.com non-NetBird access is denied (HTTP 403).
  • Validate konga.perspective-v.com shows Konga built-in login from a NetBird-connected client (server-side NetBird-address path validation recorded in docs/operations/validations/2026-04-15-netbird-route-validation.mdx).
  • Validate Kong Admin API is not host-exposed and is reachable from Konga over Docker network (http://kong:8001).
  • Convert Ocelot routes to Kong services/routes/plugins using runtime/stacks/infrastructure/gateway/kong-bootstrap-ocelot.sh and confirm inventory parity in Kong/Konga (5 services, 8 routes, expected plugin bindings).
  • Populate real gateway secrets in runtime/environments/vps/infrastructure/gateway/.env before VPS deployment.
  • Add standalone service stack artifacts for identity, graph, and console under runtime/stacks/services with matching VPS tracked env templates under runtime/environments/vps/services.
  • Align Kong and service upstream connectivity to the existing external proxy network (no dedicated gateway-only network).
  • In Kong/Konga, backend upstream services and routes are provisioned for identity (identity:62258) and graph (graph:5138) under host api.perspective-v.com.
  • Keep Console out of Kong routes: no console route is provisioned in Kong migration set and console remains direct Traefik host console.perspective-v.com.
  • Run end-to-end route smoke tests for migrated public/protected gateway paths (JWT allow/deny and rewrite behavior) and capture evidence.
  • Capture and store dedicated gateway parity validation record under docs/operations/validations/2026-04-16-gateway-ocelot-parity-validation.mdx with route-level evidence.
  • Add and apply Kong compatibility routes for client URL stability (/identity/docs/*, /identity/swagger/v1/swagger.json, /swagger/v1/swagger.json, and /resume as UI passthrough).
  • Validate public compatibility outcomes: https://api.perspective-v.com/identity/docs/index.html HTTP 200, https://api.perspective-v.com/identity/swagger/v1/swagger.json HTTP 200, https://api.perspective-v.com/swagger/v1/swagger.json HTTP 200, and https://api.perspective-v.com/resume upstream HTTP 400 (no Kong route-miss 404).
  • Add and apply Kong gopher compatibility routes (/graph/gopher/{credential,serviceprovider,serviceprovideremail,platform} as explicit public API routes, plus /gopher/* UI passthrough) with priorities above graph-protected.
  • Validate gopher compatibility outcomes: graph gopher API paths now return upstream responses (GET 400 / POST 415 where applicable) instead of pre-fix route/auth failures (GET 404 / POST 401), and UI paths /resume plus /gopher/{credential,serviceprovider,serviceprovideremail,platform} resolve through Kong as passthrough UI routes.
  • Extend Kong global CORS allow-headers with Apollo client headers (apollographql-client-name, apollographql-client-version) to satisfy browser preflight for console-origin graph API calls.
  • Update live Kong CORS allowed origins to include https://console.perspective-v.com and https://hassan.taj.contact, then validate graph endpoint preflight returns HTTP 200 for both origins.
  • Add https://hassantaj.github.io to live Kong CORS allowed origins and enable OPTIONS on /graph/resume, then validate preflight returns HTTP 200 with matching allow-origin.
  • Create a complete Kong route/service mapping guide covering service/route additions, rewrite/auth/plugin model, and validation workflow.
  • Add repository-side graph parity extension in gateway bootstrap for /graph/docs/* (public docs rewrite) and /graph/v1/* (identity-style JWT-protected rewrite to /api/v1/*) while keeping existing graph compatibility routes unchanged.
  • Add rollout artifacts for graph v1/docs under docs/operations/validations/2026-04-21-graph-v1-kong-access-validation.mdx and docs/operations/change-records/cr-2026-04-21-graph-v1-kong-route-rollout.mdx.
  • Add Hoppscotch-importable graph REST template at runtime/stacks/infrastructure/gateway/client-imports/graph-v1-hoppscotch-openapi.json and link import steps in gateway README.
  • Execute owner-approved VPS apply for graph v1/docs bootstrap changes and capture runtime route/plugin inventory evidence.
  • Run post-apply route smoke validation for /graph/docs (public) and /graph/v1/{endpoint} (JWT allow/deny) and record observed HTTP outcomes.
  • Add declarative Kong state file runtime/stacks/infrastructure/gateway/kong.yml for decK gateway sync/diff automation and document sync commands in gateway README.
  • Install decK natively on VPS host (non-Docker), validate runtime/stacks/infrastructure/gateway/kong.yml locally with deck file validate, and verify online Kong Admin connectivity with deck gateway ping.
  • Remove obsolete legacy graph service/routes from runtime/stacks/infrastructure/gateway/kong.yml so decK dry-run no longer proposes unintended entity creation.
  • Decide and execute JWT secret alignment in live Kong via host-native decK sync (envsubst rendered state): credential ocelot-jwt-issuer now matches VPS KONG_JWT_HMAC_SECRET and post-sync diff is zero.
  • Resolve upstream graph docs outage by refreshing graph:latest runtime image and recreating the graph container; https://api.perspective-v.com/graph/docs/ now serves Scalar UI.
  • Investigate upstream graph application responses on /api/v1/health if non-auth errors persist under JWT-authenticated probes.
  • Apply decK sync after removing JWT from graph-resume-get-by-token-public so /graph/v1/resume/GetByAccessToken is public.

Phase 3: Object Storage Baseline (RustFS)

  • Add RustFS compose artifact at runtime/stacks/infrastructure/object-storage/rustfs/docker-compose.rustfs.yml with pinned image, NetBird-only routes, and explicit Traefik router-service bindings.
  • Add RustFS dev env template at runtime/environments/dev/infrastructure/object-storage/rustfs/rustfs.dev.env.
  • Prepare server-local RustFS runtime values in runtime/environments/vps/infrastructure/object-storage/rustfs/.env (no tracked secrets).
  • Set explicit non-default RustFS API credentials in server-local env (RUSTFS_ACCESS_KEY and RUSTFS_SECRET_KEY).
  • Create /var/lib/rustfs/data and /var/lib/rustfs/logs on VPS with ownership 10001:10001.
  • Deploy RustFS from runtime/stacks/infrastructure/object-storage/rustfs and confirm rustfs container health.
  • Validate RustFS console/API through NetBird-address path (SNI --resolve to server NetBird IP) showing backend responses (console index HTTP 200, API auth gate as XML AccessDenied).
  • Validate RustFS console/API from NetBird client machines (successful key-login path on rustfs.perspective-v.com with server URL rustfs-api.perspective-v.com).
  • Validate RustFS route denial from non-NetBird sources (HTTP 403 expected).
  • Decommission MinIO runtime completely (container, data volume, images, and runtime stack/env artifacts removed).
  • Update object-storage runbook and tracking docs to RustFS-only operating model with decommission rationale.
  • Sweep runtime app env/service-stack artifacts and remove residual MinIO host or MINIO_* references (RustFS-only endpoint standard).
  • Validate app-container S3 access to RustFS over object-storage Docker network.
  • Validate signed URL delivery through app layer using RustFS backend.
  • If first certificate issuance fails under strict SNI, temporarily set sniStrict=false, issue certs for RustFS hosts, then restore sniStrict=true.

Phase 4: Databases (MSSQL + PostgreSQL + MySQL + MongoDB)

  • Prepare mssql.env, postgres.env, mysql.env, and mongodb.env from runtime/environments/dev/infrastructure/databases/<engine>/<engine>.dev.env templates.
  • Create DNS records for mssql/pgsql/mysql/mongo and pgadmin/phpmyadmin/mongo-express hostnames.
  • Keep DB/UI DNS hostnames as public records; do not create a separate NetBird-only DNS zone for this phase.
  • Complete live DB credential rotation validation for all engines after VPS .env secret hardening (MSSQL/PostgreSQL/MySQL/MongoDB validated in running containers).
  • Deploy MSSQL, PostgreSQL, MySQL, and MongoDB from runtime/stacks/infrastructure/databases.
  • Deploy pgAdmin, phpMyAdmin, and mongo-express admin profiles.
  • Add pgAdmin-specific Traefik middleware override to keep NetBird-only access while allowing SAMEORIGIN frame behavior for Query Tool.
  • Recreate PostgreSQL from runtime/environments/vps/infrastructure/databases/postgres/.env and remediate persisted-role password drift so non-local auth uses .env password.
  • Recreate PostgreSQL admin profile to apply updated pgAdmin Traefik labels and verify running container label update.
  • Validate pgAdmin Query Tool opens without "refused to connect" from a NetBird client.
  • Redeploy/validate Mongo Express built-in auth credentials from VPS .env and confirm DB UI routes stay on NetBird-only + security headers (no Traefik admin-auth).
  • Add compose teardown commands in databases README for full DB stack stop/remove workflow.
  • In NetBird dashboard, verify VPS peer and required developer/admin peers are connected before connectivity tests.
  • Record active NetBird interface and server NetBird IP for cutover commands.
  • Validate friendly DB hostnames over NetBird (mssql/pgsql/mysql/mongo) from dedicated NetBird client machines.
  • Validate DB management UI hostnames over NetBird (pgadmin/phpmyadmin/mongo-express).
  • Validate DB UI hosts from NetBird return app login responses (typically 200/302) without Traefik 401 challenges.
  • Validate DB UI route denial from non-NetBird source (HTTP 403).
  • Run operations/migrations/netbird-db-port-cutover.sh with the detected NetBird interface.
  • Confirm public DB/Redis access is closed and NetBird-only DB/Redis access works from separate NetBird/non-NetBird clients.

Phase 5: Hardening and Resource Optimization

  • Install and tune Fail2ban.
  • Finalize interim backup destination as RustFS on the same VPS for immediate baseline coverage.
  • Finalize critical-service recovery targets at RTO 4 hours and RPO 1 hour.
  • Add Tier-1 database backup helper at operations/backups/db-backup.sh with strict rolling keep-last-backups retention and manifest output.
  • Add systemd scheduling artifacts under operations/systemd/.
  • Add approval-ready execution bundle for first backup rollout at docs/operations/execution/phase5-backup-rollout-execution-bundle.mdx.
  • Add prefilled change record draft at docs/operations/change-records/cr-2026-04-15-tier1-backup-timer-rollout.mdx.
  • Add first-run evidence helper at operations/diagnostics/collect-tier1-backup-evidence.sh.
  • Implement automated backups for Tier 1 stateful services and keep retention documented (first successful run captured at /home/repo/contabo-server-setup/tmp/backups/tier1-20260415T205728Z).
  • Install approved production scheduling for backup automation and keep first-run evidence (tier1-db-backup.timer active; evidence file /home/repo/contabo-server-setup/tmp/backups/evidence/tier1-backup-evidence-20260415T205803Z.txt).
  • Update repository Tier-1 backup policy defaults to weekly Sunday 03:00 UTC and strict keep-last-3 retention across scripts, installer, systemd templates, and runbook references.
  • Ask owner approval, then apply updated weekly keep-last-3 timer/service policy on VPS and capture refreshed timer/evidence outputs (completed; evidence /home/repo/contabo-server-setup/tmp/backups/evidence/tier1-backup-evidence-20260416T222414Z.txt).
  • Run and document one end-to-end restore validation against RTO/RPO targets (docs/operations/restore-validations/rv-20260415t210554z-postgres-tier1.mdx; evidence /home/repo/contabo-server-setup/tmp/backups/evidence/restore-validation-20260415T210554Z.txt).
  • Add restore drill template at docs/operations/templates/restore-validation-template.mdx.
  • Finish the encrypted Google Drive rollout for four WordPress sites and active PostgreSQL/MySQL/MongoDB engines, including physical PostgreSQL recovery validation, offsite flags, sliding retention, diagnostics, and all three timers.
  • Keep monitoring lightweight for 4 vCPU / 8 GB (no heavy monitoring stack yet) (docs/operations/baselines/lightweight-monitoring-baseline.mdx).
  • Schedule controlled monthly DB maintenance updates (docs/operations/baselines/monthly-db-maintenance-schedule.mdx).

Phase 5.1: Governance and Runbooks

  • Add disaster recovery policy at docs/operations/governance/disaster-recovery-plan.mdx.
  • Add incident response playbook at docs/operations/governance/incident-response-playbook.mdx.
  • Add post-incident host integrity checklist at docs/operations/governance/host-integrity-checklist.mdx.
  • Add production change control policy at docs/operations/governance/change-control-policy.mdx.
  • Add production change record template at docs/operations/templates/production-change-record-template.mdx.
  • Add incident closure template at docs/operations/templates/incident-closure-template.mdx.
  • Add NetBird access validation matrix at docs/operations/governance/netbird-access-validation-matrix.mdx.
  • Execute one tabletop drill using the new incident-response and host-integrity runbooks (docs/operations/tabletop-drills/tt-20260415t220000z-ssh-egress-policy.mdx).
  • Execute one approved production change using the change-control policy template and record outcomes (docs/operations/change-records/cr-2026-04-15-tier1-backup-timer-rollout.mdx).

Phase 6: App and Feed Migration (After Secure Baseline)

Superseded (2026-07-01): the package layer described here (registry.perspective-v.com + BaGet/Verdaccio/ProGet feeds and per-service REGISTRY_USERNAME/REGISTRY_PASSWORD basic-auth) has been replaced by the Gitea package registry and those 5 services were retired — see the "Gitea Unified Package Registry Migration (2026-07-01)" entry above. Open ([ ]) items below that still reference registry.perspective-v.com or old feed hosts are obsolete; the equivalent work is done against gitea.perspective-v.com/perspective-v/* with token auth.

  • Deploy production services from runtime/stacks/services/perspective-v.
  • Confirm service images resolve from registry.perspective-v.com with valid docker login.
  • Validate services requiring object storage are attached to object-storage network.
  • Prepare reusable CI templates under runtime/ci for GitHub Actions and Azure Pipelines.
  • Add dbskc CI templates under runtime/ci for GitHub Actions and Azure Pipelines with branch trigger deploy/dbskc.
  • Add nishatcolony CI templates under runtime/ci for GitHub Actions and Azure Pipelines with branch trigger deploy/nishatcolony.
  • Align identity/graph/console CI templates under runtime/ci to deploy branches (deploy/identity, deploy/graph, deploy/console) for both GitHub Actions and Azure Pipelines.
  • Enable Watchtower labels for stateless application service compose templates under runtime/stacks/services.
  • Copy service-specific CI templates (identity/graph/console) into matching repositories and set REGISTRY_USERNAME/REGISTRY_PASSWORD secrets.
  • Copy dbskc CI templates into the dbskc application repo(s) and set REGISTRY_USERNAME/REGISTRY_PASSWORD secrets.
  • Copy nishatcolony CI templates into the nishatcolony application repo(s) and set REGISTRY_USERNAME/REGISTRY_PASSWORD secrets.
  • Add a dedicated dbskc-ci registry basic-auth user on VPS and set matching plaintext REGISTRY_USERNAME/REGISTRY_PASSWORD in Azure pipeline variables.
  • Add a dedicated nishatcolony-ci registry basic-auth user on VPS and set matching plaintext REGISTRY_USERNAME/REGISTRY_PASSWORD in Azure pipeline variables.
  • Run dbskc CI from deploy/dbskc and confirm registry.perspective-v.com/dbskc-web receives both latest and v1.0.<run> tags.
  • Run nishatcolony CI from deploy/nishatcolony and confirm registry.perspective-v.com/nishatcolony-web receives both latest and v1.0.<run> tags.
  • Validate dbskc Azure pipeline registry login returns Login Succeeded after credential sync.
  • Prepare dbskc VPS runtime env at runtime/environments/vps/services/perspective-v/dbskc/.env with tracked placeholder runtime/environments/vps/services/perspective-v/dbskc/dbskc.env.
  • Prepare nishatcolony VPS runtime env at runtime/environments/vps/services/perspective-v/nishatcolony/.env with tracked placeholder runtime/environments/vps/services/perspective-v/nishatcolony/nishatcolony.env.
  • Mirror Perspective-V service env layout by service under runtime/environments/{dev,vps}/services/perspective-v/{identity,graph,console,dbskc,nishatcolony}.
  • Populate server-local service-local .env files for identity/graph/console at runtime/environments/vps/services/perspective-v/{identity,graph,console}/.env.
  • Ask owner approval, then run per-service deployment/recreate from runtime/stacks/services/perspective-v/{identity,graph,console} using service-local env files.
  • Validate DNS for dbskc.com resolves to VPS before first dbskc deploy.
  • Finalize nishatcolony public host as nishatcolony.pk (instead of app.nishatcolony.pk) and validate DNS resolves to VPS before first nishatcolony deploy.
  • Deploy dbskc stack from runtime/stacks/services/perspective-v/dbskc using VPS env source and verify container healthy.
  • Deploy nishatcolony stack from runtime/stacks/services/perspective-v/nishatcolony using VPS env source and verify container healthy.
  • Validate dbskc HTTPS route response on https://dbskc.com and confirm Traefik router/certificate behavior.
  • Validate nishatcolony HTTPS route response on https://nishatcolony.pk and confirm Traefik router/certificate behavior.
  • Validate one controlled dbskc update path: push a new latest image and confirm watchtower-fast applies rollout within configured fast interval.
  • Create the versioned Gitea auth Swarm secret and deploy the updated watchtower-fast stack.
  • Confirm the next watchtower-fast poll scans dev-dbskc-web without unauthorized.
  • Merge dbskc into deploy/dev.dbskc.com, build dev-dbskc-web, bootstrap websites-dev-dbskc, and validate https://dev.dbskc.com plus fast Watchtower rollout.
  • Validate one controlled nishatcolony update path: push a new latest image and confirm watchtower-fast applies rollout within configured fast interval.
  • Run CI pipeline once per service and confirm both latest and v1.0.<run> tags are pushed to registry.perspective-v.com.
  • Bootstrap Perspective-V services on VPS once, then keep rolling updates on Watchtower-managed latest tags.
  • Validate one controlled baseline Watchtower rollout (push new latest for one non-fast service and confirm Sunday schedule is applied).
  • Execute phased cutover per domain plan.
  • Add syassociates compose artifacts under runtime/stacks/services/syassociates.pk for syassociates.pk and service.syassociates.pk.
  • Deploy and validate syassociates frontend HTTPS and Kong-backed API routing after owner approval.

Phase 9: Docs Taxonomy Normalization

  • Keep decision/state source-of-truth files under docs/state/.
  • Keep operational artifacts and records under docs/operations/.
  • Keep setup and operator guides under docs/runtime/.
  • Normalize state-document references from legacy docs/* paths to docs/operations/* where applicable.
  • Update feed migration guidance in docs/runtime/legacy-setup-guide.mdx to canonical feed routing model.
  • Run a final docs path/read-through pass; active Swarm role paths and stack identities no longer use the superseded layout.

Security Incident Follow-up (2026-04-11)

  • Archive incident package under Incidents/2026-04-12-0015-ssh-outbound-spike/ with markdown reports and raw evidence files.
  • Draft provider-ready Contabo response and attach supporting evidence references.
  • Run same-day post-containment revalidation and archive evidence files 23 through 30 (live SSH egress snapshot, process attribution, scheduler scan, Uptime Kuma monitor audit, and control snapshots).
  • Verify whether accepted SSH login from 86.121.67.247 at 2026-04-09 12:00 UTC+0 was authorized; if not authorized, treat as compromise indicator.
  • Verify that established inbound SSH source IPs observed during revalidation are authorized (202.66.181.251 and 36.132.36.134 in evidence file 25).
  • Rotate root password and regenerate/rotate SSH keys in authorized_keys; remove any unknown key material.
  • Disable provider VNC/remote console access for normal operations.
  • Reset VS Code SSH and Tabby SSH access credentials/keys and keep trusted principals only.
  • Finalize permanent outbound SSH policy: dual-stack deny-by-default with logging-enabled TCP/22 block, GitHub access via ssh.github.com:443, owner-approved one-hour exception workflow, and Discord alert path.
  • Implement outbound SSH policy tooling (runtime/stacks/infrastructure/operations/ssh-egress-policy.sh) and enable recurring maintenance timer.
  • Investigate 2026-04-17 SSH egress Discord alert flood and identify root cause from VPS evidence (detector false positives caused by inbound UFW BLOCK matching plus DPT=22 prefix matching 22xx ports).
  • Patch runtime/stacks/infrastructure/operations/ssh-egress-policy.sh blocked-scan filter to detect only outbound exact DPT=22 events (IN= OUT=<iface> + DPT=22 exact) and exclude 22xx matches.
  • Fix blocked-detection Discord formatter in runtime/stacks/infrastructure/operations/ssh-egress-policy.sh by escaping Markdown code-fence backticks so systemd logs no longer show text: command not found / [UFW: command not found during alert rendering.
  • After owner approval, apply patched SSH egress policy script on VPS and validate with one synthetic outbound TCP/22 probe plus one inbound 22xx check to confirm only true outbound SSH events alert Discord (completed 2026-04-17: false-positive window replay returned 0 matches, synthetic outbound probe produced only exact DPT=22 outbound matches).
  • Verify Discord render quality for embed-card SSH egress alerts using underlined heading plus ordered-list body during one blocked-detection event and one exception open/close cycle.
  • Run one controlled exception drill (open and close) using full audit fields to validate runbook quality under change pressure.
  • Verify effective SSH auth matrix (sshd -T) and confirm no password fallback/auth-method drift remains after access reset.
  • Send Contabo the prepared incident summary (source process, destination pattern, containment actions, current status) and keep provider confirmation in the same incident folder.
  • Add incident response playbook and host integrity checklist in docs/operations/ for repeatable post-incident operations.
  • Complete deep host integrity checks (package/service audit, startup persistence sweep, credential review) and decide on rebuild-vs-recovery posture.

Docker Swarm Migration (2026-06-30)

  • Swarm migration plan finalized with Docker Secrets, pre-built WordPress images, separate swarm/ directory
  • swarm/stacks/ — active Swarm services are maintained as one YAML fragment per service under six role directories.
  • swarm/configs/traefik/ — Swarm provider config + file-based middlewares (@file instead of @docker)
  • swarm/secrets/ — create-all-secrets.sh + secrets-map.md for 30+ Docker secrets
  • swarm/scripts/ — role-organized stack launchers plus guarded panel management.
  • swarm/README.md — deployment docs, dependency graph, rollback procedure
  • docs/runtime/swarm-migration.mdx — migration guide and rationale
  • WordPress images built and pushed to registry.perspective-v.com (dbskc-web, nishatcolony-pk, gorsistudio-web, gorsistudio-store, wcblahore-pk)
  • VPS pre-flight: full DB backup via tier1 scripts, volume snapshot (tmp/backups/swarm-preflight-volumes-*.txt)
  • VPS cutover: swarm init, overlay networks, secrets, stacks deployed
  • Post-migration validation: services 1/1, HTTP checks, DB connectivity, NetBird access, ACME certs

Post-cutover fixes (2026-07-01)

  • Data preservation: all stacks map data volumes to the original compose volumes via external: true (e.g. netbird_netbird_data, databases_postgres-data); pre-migration data intact. Orphan infra-*/svc-* volumes from the first (pre-mapping) deploy are unused.
  • Traefik ports switched to mode: host (now swarm/stacks/edge/traefik.yml) — swarm ingress SNAT hid the client IP (10.0.0.2), which broke the netbird-only (100.64.0.0/10) allowlist on all admin UIs. Admin UIs now reachable over the NetBird VPN.
  • create-all-secrets.sh quote-strip fix — .env values like MYSQL_PWD="…" were stored with literal quotes, breaking auth; WP + kener_secret_key + vaultwarden_admin_token secrets recreated cleanly. WordPress sites (dbskc, gorsistudio, store, wcblahore) now 200.
  • Split-stack env sourcing: edge.sh/.bat preloads Kener's env and the panels launcher loads each panel's owning subsystem env; website launchers preserve per-service database variables. .sh/.bat pairs are reconciled.
  • _FILE-less services fed from secrets via command wrappers: kener (KENER_SECRET_KEY, SMTP_PASSWORD), vaultwarden (ADMIN_TOKEN).
  • arnexglobal healthcheck fixed (wget→bash TCP probe); zitadel /ui/v2/login router added; Traefik dashboard router fixed (dummy service port); dbskc-web-coming-soon duplicate removed.
  • Traefik dashboard basic auth reset (admin / htpasswd at /var/lib/traefik/registry-auth/registry.htpasswd).
  • phpMyAdmin extra themes (darkwolf, boodark) bind-mounted from /var/lib/phpmyadmin/themes/* (swarm/stacks/panels/phpmyadmin.yml).

Remaining follow-ups (blocked / optional)

  • Migrate the eight administration UIs to panel, preserving image digests, persistent data, access rules, and the prior replica baseline; controlled pgAdmin 0 → 1 → 0 validation passed.
  • Complete the interim Vaultwarden, Zitadel, and Gitea svc-* rename; this identity was later superseded by the header-defined namespace rollout below.
  • svc-puller login restored (2026-08-15). Root cause was not a bad or expired token: /root/.docker/config.json had an empty auths object, so the node held no credential for gitea.perspective-v.com at all. The account itself was healthy (active, 1 token) and the registry challenged correctly (GET /v2/ -> 401 with Bearer + Basic; note HEAD /v2/ returns 405, so probe with GET). Existing token values are unrecoverable (Gitea stores hashes), so a fresh read:package token was minted with gitea admin user generate-access-token -u svc-puller --scopes read:package --raw (run as the git user — the CLI refuses to run as root) and piped straight into docker login --password-stdin so the value was never printed. Verified: all 11 gitea.perspective-v.com/perspective-v/* images pull successfully.
    • The previous unusable token is still on the account; revoke it in the Gitea UI when convenient.
    • docker login stores credentials base64-encoded (not encrypted) in /root/.docker/config.json — standard Docker behavior, worth knowing for host-integrity reviews.
  • data.arnexglobal.com TLS — BLOCKED: DNS A record points to 66.165.248.146, not the server 161.97.83.142; ACME tlsChallenge fails until the record is repointed. App itself runs fine.
  • Diagnose the pre-existing HTTP 500 from new.nishatcolony.pk; the websites-nishatcolony_nishatcolony-pk task is healthy at 1/1 and its data volume was preserved.
  • perspective-v.com apex and syassociates.pk not routed/deployed (syassociates deploys via template-env fallback when needed).

Gitea Unified Package Registry Migration (2026-07-01)

A root-level working log /home/repo/contabo-server-setup/next-steps.md exists from this migration and should be consolidated into this file; docs/state/next-steps.mdx is canonical.

  • Deploy the package-only registry at gitea.perspective-v.com (git disabled), now running as service at 1/1 with valid LE cert; org owner perspective-v, admin pvadmin.
  • Migrate all package data into Gitea: NuGet 45 versions (BaGet), Docker 20 tags/11 repos (old registry), npm @pv/core 14 versions (Azure DevOps pv-ng); Verdaccio empty.
  • Repoint CI/CD templates to Gitea (runtime/ci/github-actions/*, runtime/ci/azure-pipelines/*, runtime/ci/README.md).
  • Reconfigure + prepare WUD to watch Gitea (dedicated wud-monitor read:package token and recreated wud_registry_password secret); WUD is now sourced from swarm/stacks/panels/wud.yml.
  • Cut over all 11 running services to gitea.perspective-v.com/perspective-v/*; all are 1/1. The later header-defined rollout restored the drifted Nishat web source to its Gitea image and restored the separate nishatcolony.pk router; the intended WordPress route at new.nishatcolony.pk still returns its pre-existing 500.
  • Retire legacy 5 services: removed infra-registry (registry + registry-admin) and infra-feeds (proget + baget + verdaccio); old hosts now 404. Removed secrets registry_*, baget_api_key, npm_feed_api_key.
  • Publish guide docs/runtime/stacks/gitea.mdx and retirement runbook docs/runtime/gitea-retirement-runbook.mdx.

Remaining follow-ups (optional cleanup — needs go)

  • Legacy data volumes removed by the owner via Portainer on 2026-08-15 (32 -> 24 volumes). Note: the documented precondition — verifying ProGet held no unmigrated NuGet packages before pruning its volume — was not performed first, so that check is now moot. Only BaGet's 45 versions were ever migrated to Gitea.
  • Retired registry.perspective-v.com/* image tags removed from the node (2026-08-15). They were byte-identical duplicates: each shared the same image ID as its gitea.perspective-v.com/perspective-v/* counterpart, so removing them only dropped the obsolete tag and freed no image data.
  • Remove old DNS records: registry., registry-admin., nuget., npm., proget.perspective-v.com.
  • Remove retired Registry/Feeds Swarm manifests, launchers, and teardown helper.
  • Remove retained Registry/Feeds Compose environment artifacts and remaining stale secret/documentation references after their rollback window.
  • Fix .claude/settings.local.json health check(s) still pointing at proget.perspective-v.com.
  • Validate WUD in panel and both Watchtower services in platform; all three are healthy after migration. The separate svc-puller registry-login follow-up remains open below.

Container image update rollout (2026-08-14)

Execution bundle: docs/operations/execution/image-update-rollout-execution-bundle.mdx. WUD reported 7 available updates; 2 were rejected as unsafe after verification.

  • Verify all 7 WUD-reported updates against upstream registries (Docker Hub, GHCR, MCR).
  • Reject Redis 8-alpine -> 32bit-stretch: that tag is a 2019 32-bit Redis 5.x build; applying it would be a 3-major downgrade and a Redis 8 AOF/RDB will not load under it. Cause is a missing wud.tag.include label on the running container, not a manifest defect.
  • Reject Gitea 1.27-rootless: production runs the non-rootless image; rootless uses /var/lib/gitea + /etc/gitea and a different entrypoint, so the scm_gitea_data volume at /data and the menu.tmpl bind would both be ignored. Correct target is gitea/gitea:1.27.2.
  • DEV-rehearse Gitea 1.22 -> 1.27.2 against a mirrored Compose stack: schema migrated 299 -> 343, users/packages preserved, package download byte-identical, /v2/ 401 intact, custom org template renders.
  • Fix latent Gitea wrapper bug found in rehearsal: s6-svscan lives at /bin/ in 1.22 but /usr/bin/ in 1.27, so the hardcoded /bin/s6-svscan exits 127 on upgrade. swarm/stacks/services/gitea/gitea.yml now resolves it via PATH (verified in both images).
  • DEV-verify WUD 8.3.1 boots with the existing command wrapper (dist/index.js still present); no manifest change needed beyond the image bump.
  • Confirm the two Zitadel WUD entries are one upgrade and that cf6c2e88 is the amd64 manifest of the pinned v4.15.2 index (no repo/production drift).

Executed 2026-08-15 (all validated, backups in /var/backups/preupgrade-20260814/)

  • Window 0: stack names resolved — production is on the renamed stacks (platform, service, panel, service-zitadel). The infra-* names in WUD's report were stale store entries for containers that no longer exist.
  • Window 1 not required: platform_redis already carried the wud.* labels; the real defect was label placement (see below).
  • WUD 8.2.2 -> 8.3.1.
  • Vaultwarden 1.36.0 -> 1.37.1 (Rocket launched, no errors).
  • Gitea 1.22 -> 1.27.2: schema migrated 299 -> 343; 21 packages / 113 versions / 5 users preserved; /v2/ 401 intact; custom org template renders.
  • Zitadel core + login -> v4.17.1 in lockstep; migrations verified; / -> /ui/v2/login returns 200.
  • Kener, Redis (8.10.0), RabbitMQ (4.3.4), MySQL (9.7.2), Kong (3.9.3) updated; all WordPress sites, API gateway routes, and queue/DB checks pass.
  • RustFS rc.1 attempted and ROLLED BACK to 1.0.0-beta.6-glibc. rc.1 requires --console-enable (beta.6 starts the console by default), moves the console from / to /rustfs/console/, and rejects the stored public bucket policy (unknown field ID, expected Id), which broke anonymous public-object reads with HTTP 500. Public object serving verified restored (HTTP 200) after rollback.

WUD reporting defects found and fixed (2026-08-15)

  • wud.* labels were under deploy.labels, which Swarm applies to the service; WUD reads container labels, so wud.tag.include had never taken effect. Moved to service-level labels: on redis and added guardrails to traefik, mysql, postgres, mongodb, rabbitmq, gitea, vaultwarden, rustfs. Before the fix WUD proposed traefik -> v3.7-windowsservercore-ltsc2025, mysql -> 26.7-oraclelinux9, postgres -> 19beta1-master, rabbitmq -> 4.3-rc-management-alpine, and redis -> 32bit-stretch; all are now filtered.
  • Docker Hub digest watching was off. hub/Hub.js overrides shouldWatchDigest() and returns false unless WUD_REGISTRY_HUB_PUBLIC_WATCHDIGEST=true or a per-container wud.watch.digest label is set — other registries default to true. That is why GHCR (Zitadel) reported digest updates but every Hub :latest image was skipped with "not a semver and digest watching is disabled". The watcher-level WUD_WATCHER_LOCAL_WATCHDIGEST is deprecated in 8.x and does not control this.
  • Discord notifications were failing entirely with HTTP 400. MODE=batch renders one message and the verbose upstream body exceeded Discord's limit, so nothing was delivered. Replaced with fixed-width container | current | new | type rows.
  • Batch message reformatted as a single fenced code block. Discord does not render markdown tables (pipes stay literal), and upstream renderBatchBody hardcodes a - markdown bullet per container with no configuration to disable it. Overrode that one method via a read-only bind mount at swarm/stacks/panels/custom/wud/Trigger.js (same pattern already used for Gitea's menu.tmpl), emitting one ```vb fence containing a header row plus one aligned row per update. New env knobs: WUD_TRIGGER_BATCH_LANG, WUD_TRIGGER_BATCH_HEADER, WUD_TRIGGER_BATCH_MAXCHARS.
    • Maintenance: the mounted Trigger.js was extracted from image 8.3.1. Re-extract and re-apply the patch on every WUD upgrade, otherwise a stale copy of this file is mounted over the new image's version.
    • Body templates are JS template literals eval'd as eval('`'+template+'`') and must not contain a literal backtick — use String.fromCharCode(96) if one is ever needed.
    • WUD_TRIGGER_BATCH_MAXCHARS guards Discord's 1024-character embed field-value limit and appends ... +N more rather than failing the whole webhook.

Second pass — all remaining stable updates applied 2026-08-15

  • Digest refreshes applied and validated: postgres (PostgreSQL 18.4 / PostGIS 3.6, 9 DBs intact), mongodb, traefik v3.7 (routes + ACME cert preserved), portainer, netbird-server, netbird-dashboard, and the scaled-to-zero pgadmin (9.16 -> 9.17), redis-insight, mssql. Full pg_dumpall taken first (19 MB, 22 databases).
  • RustFS beta.6 -> 1.0.0-rc.2-glibc — the rc.1 rollback blockers were fixed first, so both prior regressions are resolved: anonymous public reads return HTTP 200 with the correct byte count and there are zero bucket_metadata_parse_failed entries.
    • The stored public bucket policy contained a legacy "ID":"" field that rc.x rejects. Corrected in place via PutBucketPolicy while still on beta.6 (dropping the empty field only, no semantic change); original preserved at /var/backups/preupgrade-20260814/rustfs-public-policy.orig.json.
    • --console-enable added to the command (rc.x makes the console opt-in).
    • rc.x serves the console under /rustfs/console/ rather than /, so a rustfs-console-root router + redirectregex middleware (priority 300) redirects the bare host to it, mirroring zitadel-root. /public/ (priority 200) and the API host are unaffected.
    • RustFS has no stable release published — only alpha/beta/rc — so rc.2 is the closest-to-stable option available. It is the only pre-release image in the fleet.
  • Full sweep re-run: zero digest drift across all 26 public images.
  • Confirm from a NetBird client that https://rustfs.perspective-v.com/ redirects to /rustfs/console/ and the console loads. The router carries netbird-only, so this cannot be verified from the node itself.
  • websites-perspective-v_graph has a newer image in Gitea — intentionally not applied, since pulling it deploys new application code (a release decision, not a maintenance update).

Tooling note — pin digests that are actually addressable

MCR content-negotiates on Accept, and the two answers are not equivalent:

  • Index-only Accept returns docker-content-digest: sha256:0730f368… for tag 2022-latest, but fetching that digest returns 404 — it is not independently addressable.
  • Broad Accept (including manifest.v2/image.manifest.v1) returns sha256:ba4c8329…, which does resolve by digest and is what Docker can pull.

A pin taken from the index-only header therefore produced an unpullable database_mssql manifest, which would only have surfaced the next time that service was scaled up. Corrected to ba4c8329… and verified. Docker Hub and GHCR return the same digest either way; only MCR differs.

Rule: after changing any pinned digest, verify it resolves by digest (docker manifest inspect <repo>@<digest>), not just that the tag reports it. A sweep that only compares tag headers can both raise false drift and hide a broken pin.

  • Decide replacements for images whose upstreams are dead: pantsel/konga (last push 2020-05-16), containrrr/watchtower (2023-11-11), mongo-express (2024-05-22).
  • Prune /var/backups/preupgrade-20260814/ once the rollback window closes (never commit it).

Host monitoring rollout — Cockpit + Beszel + Netdata (2026-08-15)

Execution bundle: docs/operations/execution/monitoring-evaluation-bundle.mdx. Capacity review recorded in docs/operations/baselines/lightweight-monitoring-baseline.mdx. Repository artifacts are complete and validated; nothing is deployed yet.

  • All three tools designed as host installs, none in Swarm, for two reasons: docker stack deploy silently ignores cap_add/pid/security_opt (so a Swarm Netdata would report healthy while missing apps.plugin — the capability being evaluated), and monitoring must not depend on the thing it monitors (a containerised dashboard is down exactly when Docker breaks).
  • Artifacts added under operations/monitoring/{cockpit,netdata,beszel,loadtest}, operations/systemd/{beszel,netdata}, operations/diagnostics/monitoring-diagnostics.sh, and swarm/configs/traefik/dynamic/host-services.yml.
  • Config-drift detection built in from the start (install-monitoring-config.sh --check): a Swarm manifest is the deployed artifact, but a host config is only a copy, and nothing otherwise detects a live edit to /etc diverging from the repo.
  • swarm/stacks/platform/rabbitmq/rabbitmq.yml gains a loopback-only 15672 publish — RabbitMQ published no host ports at all, so a host-installed Netdata could not reach its management API. AMQP (5672) stays unpublished; overlay and Traefik paths unchanged.
  • Hostnames fixed as cockpit./netdata./beszel.perspective-v.com.
  • Create DNS A records for all three → 161.97.83.142 before Phase A, or ACME issuance fails the same way data.arnexglobal.com did.
  • Phase A — Cockpit EXECUTED 2026-08-15. Account hassan created (sudo, no SSH key); root refused via /etc/cockpit/disallowed-users; route returns 403 from non-NetBird sources with a valid LE certificate; Traefik reaches the backend (200 from the proxy overlay); config drift clean. Cost: 12 MB RSS, 7.4 MB disk.
    • cockpit-networkmanager deliberately excluded — it pulls in network-manager, and this host runs systemd-networkd via netplan with no out-of-band console (provider VNC disabled after the 2026-04 incident). A second network manager touching eth0/wt0 on a remote-only box is not worth a configuration GUI. Only the Networking config page is lost; metrics come from Netdata/Beszel.
    • cockpit.socket is bound to both 172.27.0.1 (Traefik) and the NetBird address. The bridge address only exists while Docker runs, so a bridge-only bind would make Cockpit unreachable exactly when Docker is broken.
    • Owner action: run passwd hassan. The account is created --disabled-password (status L) because a generated password must not appear in a transcript. Cockpit login does not work until this is set.
    • Owner validates from a NetBird client: hassan logs in, root is refused, and the systemd/journal/storage/terminal pages load.
  • cockpit-pcp added 2026-08-15 (owner installed it via the Cockpit UI for the Metrics history page). Recorded in operations/monitoring/README.md so a rebuild reproduces it — a package installed through a UI is exactly the drift the --check mode exists to catch. It pulls in Performance Co-Pilot: pmcd + pmlogger, ~14 MB RSS, archives in /var/log/pcp (watch growth). pmcd listens on 0.0.0.0:44321 — not publicly reachable (INPUT policy DROP, no ACCEPT rule), but the host's -A INPUT -i wt0 -j ACCEPT means any NetBird peer can query it unauthenticated. PCP is now a fourth metrics store; reconsider it if Netdata wins.
  • Netdata auth redesigned before install. It was going to use admin-auth, which shares its htpasswd with registry-basic-auth — a file holding five CI credentials distributed to build pipelines. Any CI token would have unlocked host metrics. Added a dedicated monitoring-auth middleware backed by a separate /etc/traefik/registry-auth/monitoring.htpasswd, and changed the Traefik mount from a single file to the directory so further credential sets need no edge redeploy.
  • host-services.yml staged: only the cockpit router is live. Routers for services that are not yet installed return 502, and the netdata one logged a repeating htpasswd error. Netdata and Beszel blocks are appended during their phases.
  • Phase B — Beszel EXECUTED 2026-08-15. Hub 8 MB and agent 6 MB (both capped 128M), bound to 172.27.0.1 only; agent detected eth0/wt0 confirming host-level visibility; route 403s from non-NetBird with a valid LE cert. The agent key was derived from the hub's own keypair (ssh-keygen -y -f /var/lib/beszel/beszel_data/id_ed25519), so the UI was not needed to install it.
    • Owner: create the hub admin at https://beszel.perspective-v.com/_/, then Add System → 172.27.0.1:45876, then set the Discord webhook.
    • Note: one benign HUB_URL not set warning at agent start — Beszel 0.18 tries WebSocket mode before falling back to the SSH listener used here. Not a loop.
  • Phase C — Netdata EXECUTED 2026-08-15. v2.11.0, 175 MB (ceiling 400M), bound to 127.0.0.1 + 172.27.0.1, routed at https://netdata.perspective-v.com behind netbird-only + monitoring-auth. Collectors verified across two restarts: postgres 1977 charts, mysql 45, redis 22, plus docker 72, apps.plugin 1480 and cgroups 948.
    • monitoring.htpasswd holds ONLY the admin entry copied from registry.htpasswd — same password the owner already uses, but the five CI credentials in that file do not unlock host metrics. Traefik's mount changed from a single file to the directory.
    • RabbitMQ collector dropped — cannot be done safely. Every option was tested: short-syntax 127.0.0.1:15672:15672 is ignored by swarm ingress and installs a DOCKER-INGRESS DNAT that bypasses UFW (briefly exposed publicly, reverted within minutes); long-syntax host_ip: is rejected by docker stack deploy; mode: host binds 0.0.0.0. The host also cannot reach overlay addresses. Documented in the manifest so nobody retries it.

Netdata gotchas found the hard way (all now fixed and documented)

  • Inline comments silently invert settings. Netdata does not strip trailing # comments — ebpf = no # heavy becomes the literal value no # heavy, which is not no. Result: ebpf stayed ENABLED and go.d stayed DISABLED, so every database collector was missing while the config read as correct. All comments moved to their own lines.
  • [db] mode renamed to [db] db in v2. The old name works via a migration shim (the effective config literally prints migrated from: [db].mode); now set explicitly.
  • Unquoted DSNs are not parsed. go.d silently refused to register the mysql job until dsn: was quoted. Postgres/redis quoted too for consistency.
  • go.d gives up on first failure. A collector that cannot connect during a restart never appears and logs nothing. autodetection_retry: 30 added to all three jobs; verified across two consecutive restarts.
  • Collection rate raised 2s → 5s after measurement: at 2s the docker/cgroups collectors held dockerd+containerd at ~24% CPU on a 4-core box. Idle recovered from ~50% to ~78%.
  • Netdata Cloud is not claimed (/var/lib/netdata/cloud.d empty). v2.11 exposes no documented switch to hide the dashboard's Sign-in control, so it remains but is inert — it only links to app.netdata.cloud. The agent has no local login of its own; the monitoring-auth basic-auth prompt is the self-hosted login.

Netdata release channel — was nightly, corrected to stable

  • The kickstart in operations/systemd/netdata/install.sh did not pass --stable-channel, so Netdata installed from the edge (nightly) repository as 2.11.0-14-nightly. This surfaced as a dashboard TypeError on the Appearance page — a nightly UI regression, not a configuration problem. Tracking nightly builds on a production host is wrong regardless of that symptom. Fixed: apt source switched repos/edge -> repos/stable, all 18 netdata packages downgraded to 2.11.0 (including netdata-dashboard, which carries the UI), and the installer now passes --stable-channel. Note the downgrade must move every netdata package together — netdata-user conflicts otherwise. Old source file backed up. Re-verified after the change: postgres 1977 / mysql 45 / redis 22 / docker 72 charts, apps.plugin 1480, all tuning intact, no auto-updater installed.

Monitoring decision closed + Traefik dashboard replaced (2026-08-15)

  • Beszel retained, Netdata retired. Owner ended the evaluation early. Netdata cost ~175–206 MB vs Beszel's ~21 MB, held dockerd+containerd at ~24% CPU at 2s collection, and needed materially more care to configure correctly — every misconfiguration failed silently. Full rationale in docs/operations/baselines/lightweight-monitoring-baseline.mdx. Retirement verified: 19 packages purged, apt repo removed, /etc/netdata, /var/lib/netdata, /var/cache/netdata, /var/log/netdata deleted, the three least-privilege DB monitoring users dropped from PostgreSQL/MySQL/Redis, Traefik route and UFW rule removed, all repo artifacts deleted, no listener on 19999. Final architecture: Cockpit + Portainer + Kener + Beszel (~48 MB total).

  • Built-in Traefik dashboard disabled; Traefik Manager deployed at https://traefik-manager.perspective-v.com (netbird-only, LE cert, 1/1). It replaces a dashboard whose only protection was admin-auth — the shared htpasswd that also holds five CI credentials. Traefik Manager has bcrypt cost-12 auth with optional TOTP 2FA.

    • api.dashboard: false, api.insecure: true. The API binds :8080 inside the container only (no ports: entry), so it is reachable from the proxy overlay and nowhere else. Accepted trade-off: containers on that overlay can read routing topology, but no secrets — basicAuth appears as a usersFile path and no TLS keys are served. Routing api@internal through Traefik cannot work, as the manager is a container and netbird-only would reject it.
    • Config ownership is split by file in /var/lib/traefik/dynamic, because the file provider watches only one directory: security.yml and host-services.yml are installed from Git by operations/traefik/install-dynamic-config.sh; managed.yml is the only file mounted into the manager. Verified by write test — /data/traefik.yml returns Read-only file system, /data/dynamic.yml writes through.
    • Scope note: 36 of 39 routers come from swarm deploy.labels and cannot be edited in the UI. The manager's editing value is routes for services with no built-in login.
    • Old dashboard host traefik.perspective-v.com now returns 404; its DNS record can be retired.

Two Traefik failure modes discovered — both 404 every site

  • traefik.enable=true without a loadbalancer port. The swarm provider fails with service "edge-edge-traefik" error: port is missing, and that error is not scoped to that service — it aborts the entire provider config, dropping all routers. Hit while removing the dashboard labels (the dummy port went with them). Fixed by removing traefik.enable entirely, with a warning comment in the manifest.

  • An invalid file anywhere in the dynamic directory. A managed.yml seeded with empty maps was rejected by Traefik, and one bad file aborts the whole directory load — every @file middleware vanished and every router referencing one was disabled. The seed is now comments-only. Both failures present identically: Traefik healthy and listening, every site 404; diagnose via api/http/routers showing disabled.

  • operations/traefik/README.md documents the ownership split, the API trade-off and both failure modes.

Traefik Manager data sources wired up (2026-08-15)

Certificates, Logs and Plugins were all reporting "not mounted" / "not configured". All three are now connected, read-only:

  • Plugins — STATIC_CONFIG_PATH=/data/traefik.yml, pointing at the static config that was already mounted read-only for the editor.

  • Certificates — /var/lib/traefik/letsencrypt/acme.json:/app/acme.json:ro. Accepted risk: 568 KB embedding the private key of every one of 45 certificates plus the ACME account key. Read-only prevents modification but not disclosure, so compromising this container yields every TLS private key on the host. Accepted because it is NetBird-only, has its own authentication and no Docker socket — revisit if any of those three change.

  • Logs — accessLog.filePath: /var/log/traefik/access.log added to the static config (it was previously stdout-only), mounted read-only into the manager. fields.headers.defaultMode: drop prevents Authorization and Cookie from being written verbatim. Traefik cannot redact query strings, and this estate puts JWTs there (/notifications/hub?access_token=eyJ...), so log lines remain credential-bearing. Capped at 7 compressed days by operations/traefik/logrotate-traefik, root-only, with copytruncate — mandatory, because Traefik holds the file open and does not reopen on SIGHUP, so a plain rotate would leave the new file empty.

  • TLS tab stays empty by design — no TLS options block is defined and the defaults are appropriate. Add one to managed.yml only if a minimum TLS version needs pinning.

  • State persistence proven: after the redeploy, setup_complete: true, must_change_password: false and zero bootstrap events — the owner's password survived, confirming the /app/config volume fix.

  • Recorded as CR-2026-08-15-traefik-dashboard-retirement.

  • docs/runtime/stacks/edge.mdx corrected — it still documented the dashboard at traefik.perspective-v.com as NetBird + basic auth.

  • Retire the DNS A record for traefik.perspective-v.com (now 404).

Failed systemd units resolved (2026-08-15)

Surfaced by Cockpit's failed-units view — its first practical value.

  • ssh-egress-policy-maintenance.service — dead since the repo reorganisation. Exit 203/EXEC: the unit's ExecStart pointed at runtime/stacks/operations/ssh-egress-policy.sh, but the role-layout rollout moved the script to runtime/stacks/infrastructure/operations/. The unit had been failing on every timer tick, meaning the SSH egress enforcement sweep and blocked-attempt Discord alerting from the April incident had not run since the reorg. Path corrected; the timer now runs and the service exits 0. Old unit saved to /var/backups/preupgrade-20260814/ssh-egress-unit.bak.
  • Unit brought into operations/systemd/ssh-egress-policy/ with the path templated as @REPO_ROOT@ and substituted at install time, plus a --check mode that fails if ExecStart does not point at an existing script. A repo move can no longer silently disable it. Installed and verified: timer active, service exits 0, no drift.
  • systemd-networkd-wait-online.service — stale failure from a slow boot on 2026-06-10. The netplan drop-in already scopes it to eth0, eth0 is routable, and the unit only runs at boot. Cleared with reset-failed; no configuration change needed.

Log tampering found on /var/log/wtmp (2026-04-09) — evidence preserved, logging restored

  • systemd-update-utmp.service was failing with Failed to write utmp record: Is a directory. Root cause: /var/log/wtmp had been replaced by a directory with mode d--------- and the immutable attribute set.

    Birth : 2026-04-09 10:02:09 UTC   directory created in place of the file
    Change: 2026-04-09 10:03:46 UTC   chattr +i applied, 97 seconds later

    Why this was escalated rather than quietly repaired:

    • No wtmp, utmp or chattr command appears in root's shell history for that window; the history there is w, top, ls -a, top, bash.
    • Later the same day at 12:39 / 12:41 UTC the admin ran chattr -i -a /root/.ssh/authorized_keys and chattr -i -a /root/.ssh — removing immutable flags that something else had set. The .ssh instance was found and cleaned; the wtmp instance never was.
    • It sits ~2h before the SSH login from 86.121.67.247 that remains unverified in the Security Incident Follow-up section below.
    • Only wtmp (successful logins) was affected. btmp, lastlog, auth.log and syslog were intact — the asymmetry an intruder would want.
    • It was the only immutable file on the entire system.

    Effect: no successful login was recorded on this host between 2026-04-09 and 2026-08-15 (~4 months). last was non-functional throughout.

  • Evidence preserved at /var/backups/incident-20260409-wtmp/ before any change: full stat/lsattr record, 4,456 journal lines covering 2026-04-09 09:30–13:00 UTC, and the root shell history for the window.

  • Logging restored: immutable flag cleared, directory removed, /var/log/wtmp recreated as a regular file 0664 root:utmp. systemd-update-utmp is active and last works.

  • Owner confirmed both keys in /root/.ssh/authorized_keys are theirs (2026-08-15). That file was not modified.

  • Work this against docs/operations/governance/host-integrity-checklist.mdx and close the open 86.121.67.247 verification item — a four-month gap in login history is a material finding for that review.

Measured cost (all three tools, vs the 2026-08-15 baseline)

Memory
Netdata175 MB
Beszel hub15 MB
Beszel agent6 MB
Cockpit (socket-activated)~12 MB
PCP (from cockpit-pcp)~15 MB
Total~220 MB, under the ≤300 MB budget

Available RAM 3263 MB vs 2904 MB baseline (higher, from page-cache reclaim). Load is elevated versus the 0.30 baseline and should be re-measured once the install session has quiesced; CPU was ~78% idle with wa=0 at the end of the work.

  • Phase C — Netdata + four least-privilege DB monitoring users; verify the config actually applied by diffing against curl -s http://localhost:19999/netdata.conf, since several [db] keys were renamed between v1 and v2 and unknown keys are ignored silently.
  • Phase D/E — measure after each phase; evaluate 1–2 weeks, using operations/monitoring/loadtest/synthetic-load.sh if no real incident occurs.
  • Phase F — retire the loser and record the outcome in the baseline doc.

Accepted limitation — container-to-host access is not NetBird-gated

Owner decision 2026-08-15: keep Traefik routing and accept this.

Traefik proxies from a container to a host port, and at the network layer the host cannot distinguish it from the other 30 containers on docker_gwbridge. Any container on an overlay network can reach 172.27.0.1:{9090,19999,8090} directly, bypassing the netbird-only middleware. Containers on the default bridge (docker0) cannot — the UFW rules are scoped to docker_gwbridge, which narrows the exposure to swarm-attached containers rather than all of them. No UFW or ipAllowList rule closes this — bridge addresses are dynamic and indistinguishable.

External access remains properly gated (UFW default-deny with interface-scoped additions only, services bound to 172.27.0.1 never 0.0.0.0, netbird-only on every router). Cockpit and Beszel require their own login, so a container reaches only a login page. Netdata has no built-in authentication, so admin-auth was layered onto its router — that protects the browser path but not direct container access. Closing it fully would mean dropping Traefik and binding to the NetBird interface, losing TLS and the hostnames.

Committed production secrets (2026-08-15) — OPEN, owner deferred

Owner decision on 2026-08-15: note and defer; no rotation or history rewrite performed.

  • 40 real production secret values are committed to Git, across 29 tracked .env files under runtime/environments/vps/**, pushed to github.com/HassanTaj/contabo-server-setup (origin/main) since 2026-05-03 across 15 commits. The repository is private — verified anonymously via the GitHub API (HTTP 404) — so the values are not world-readable, but they are readable by anyone with repository access and by any leaked credential with that access. Highest-risk values, in rough order: ZITADEL_MASTERKEY (encrypts the whole identity provider), KONG_JWT_HMAC_SECRET (permits minting valid JWTs, i.e. auth bypass on api.perspective-v.com), all four database root passwords (MYSQL_ROOT_PASSWORD, POSTGRES_PASSWORD, MONGO_INITDB_ROOT_PASSWORD, MSSQL_SA_PASSWORD), REDIS_PASSWORD, RABBITMQ_DEFAULT_PASS, the RustFS root/access/secret keys, the Vaultwarden ADMIN_TOKEN, the registry admin bootstrap password, SMTP passwords and the Discord webhook. AGENTS.md §16 requires committed secrets be treated as compromised; rotation and history remediation remain outstanding.

  • Root cause — .gitignore negation. Line 436 !runtime/environments/** un-ignores the whole tree, and the blanket guard runtime/environments/**/.env on line 439 is commented out. Only backups/.env was ever explicitly re-ignored, so every other live .env is tracked by default.

  • The documented verification method is misleading here. AGENTS.md §3 says to verify ignore behaviour before populating live env files, but git check-ignore -v on these paths exits 0 while printing the negation rule, which reads as "ignored". Only git status --porcelain <path> (showing ??) reveals the file is tracked. Use git status, not check-ignore, for this check.

  • New monitoring env paths (.../monitoring/netdata/.env, .../monitoring/beszel/.env) explicitly re-ignored in .gitignore and verified with git status, so the monitoring rollout adds nothing to the exposure. A warning comment now documents the misleading check-ignore behaviour inline.

Backup coverage gap (2026-08-14)

  • scm_gitea_data has no backup. operations/backups/db-backup.sh captures gitea_db incidentally (it dumps every non-template Postgres DB), but the volume holding the actual package blobs — Docker layers, .nupkg files, npm tarballs — is covered by nothing. Losing it means re-pushing every image and package by hand. Highest-value open data-loss risk.

Backup implementation finding (2026-09-08)

  • Logical Tier-1 local retention does not distinguish successful sets from failed staging. operations/backups/db-backup.sh prunes every tier1-* directory once the count exceeds DB_KEEP_LOCAL; unlike the PostgreSQL physical helper, it has no success marker. A later successful run or explicit prune can therefore delete preserved failed staging, contrary to the requirement that failed staging remain available for investigation.

Platform evaluation decisions

  • Evaluated OneDev as a Gitea replacement (2026-08-14): rejected, keep Gitea. Gitea here is a package registry (git/actions/SSH disabled), and OneDev serves the same Docker/npm/NuGet formats, so the swap is lateral on capability while costing a JVM runtime with a 2 GB documented minimum against Gitea's ~112 MiB on a 4-core/8 GB node running 39 services. It would also re-point 47 files referencing gitea.perspective-v.com six weeks after the last cutover. Self-hosted CI on this node was rejected for the same resource reasons; hosted runners pushing to the VPS registry remain the model.

Swarm role-layout rollout

  • Split active Swarm stacks into one YAML fragment per service and reorganize manifests/launchers into the six role directories.
  • Relocate host-wide operational tooling into operations/ and validate backup tests plus systemd installer dry-runs.
  • Reinstall both backup systemd units from operations/systemd/, verify their schedules and a successful persistent Tier-1 run, then remove the retired compatibility paths.
  • Migrate Redis/RabbitMQ, Watchtower, RustFS, Kong, NetBird, Panels, and Edge identities one workload at a time with preserved state and health checks.
  • Make every manifest's # Swarm stack: header authoritative and migrate production to database, platform, edge, panel, service, the three dedicated service-* stacks, and six active websites-* stacks.
  • Update database and WordPress backup defaults to database_*; both production dry-runs pass and both systemd timers remain enabled/active.
  • Preserve the prior service image and replica sets, named external volumes, and secrets; validate all expected routes plus a panel 0 → 1 → 0 cycle.

Homelab Pi-hole administration

  • TCP 8053 is present in NetBird policy contabo-homelab-server-link; no user or peer group was broadened.
  • Add private NetBird DNS record pihole.home.perspective-v.com pointing to Contabo's NetBird IP 100.83.72.162.
  • Install the repo-managed homelab-pihole Traefik router and backend, validate trusted TLS, NetBird access, public denial, and configuration drift. The rollback archive is /var/backups/traefik/dynamic-pre-pihole-host-rewrite2-20260914T231325Z.tar.gz.

Home Assistant route

  • Install the repo-managed assistant.home.perspective-v.com router and backend at 100.83.117.37:8123, protected by netbird-only and security headers. Live dynamic-config drift check is clean. Rollback archive: /var/backups/traefik/dynamic-pre-assistant-rename-20260929T204500Z.tar.gz.
  • Point private NetBird DNS and the public Namecheap A record at Contabo (100.83.72.162 privately; 161.97.83.142 publicly); preserve existing wildcard and sibling records. Trusted TLS is issued and public denial is verified.
  • Keep the old Traefik Yui host rule retired; no service router matches it.
  • After the user creates the Home Assistant owner account, confirm the imported HTTP settings in Settings > System > Network: Trust X-Forwarded-For and trusted proxy 100.83.72.162. Then verify the onboarding page through the canonical HTTPS URL.
  • Remove the old private NetBird yui.home record after the user confirmed the irreversible deletion. The assistant.home record remains, and public wildcard and sibling records are unchanged.

Homelab LAN certificate export

  • Add the restricted exporter for home, next, images, media, pihole, and portainer certificates.
  • Keep the exporter under review during the first local-edge renewal cycle; if the homelab sync timer reports a failure, retain the last known-good local files and investigate before changing the Contabo ACME design.
  • Publish the Contabo-side homelab edge guide, DNS boundary, certificate handoff, and operational validation notes in the Mintlify documentation.

On this page

Next StepsCurrent Snapshot (2026-04-28)Recommended Actions (Immediate)Phase 0: Fresh VPS BaselinePhase 1: NetBird Core Bring-upPhase 2: Private Admin Panels (NetBird-Only)Phase 2.1: Kener Public Cutover (Built-in Auth)Phase 2.5: Shared Services Baseline (Platform -> Operations -> Registry -> Feeds)Phase 2.6: Gateway Baseline (Dev-First)Phase 3: Object Storage Baseline (RustFS)Phase 4: Databases (MSSQL + PostgreSQL + MySQL + MongoDB)Phase 5: Hardening and Resource OptimizationPhase 5.1: Governance and RunbooksPhase 6: App and Feed Migration (After Secure Baseline)Phase 9: Docs Taxonomy NormalizationSecurity Incident Follow-up (2026-04-11)Docker Swarm Migration (2026-06-30)Post-cutover fixes (2026-07-01)Remaining follow-ups (blocked / optional)Gitea Unified Package Registry Migration (2026-07-01)Remaining follow-ups (optional cleanup — needs go)Container image update rollout (2026-08-14)Executed 2026-08-15 (all validated, backups in /var/backups/preupgrade-20260814/)WUD reporting defects found and fixed (2026-08-15)Second pass — all remaining stable updates applied 2026-08-15Tooling note — pin digests that are actually addressableHost monitoring rollout — Cockpit + Beszel + Netdata (2026-08-15)Netdata gotchas found the hard way (all now fixed and documented)Netdata release channel — was nightly, corrected to stableMonitoring decision closed + Traefik dashboard replaced (2026-08-15)Two Traefik failure modes discovered — both 404 every siteTraefik Manager data sources wired up (2026-08-15)Failed systemd units resolved (2026-08-15)Log tampering found on /var/log/wtmp (2026-04-09) — evidence preserved, logging restoredMeasured cost (all three tools, vs the 2026-08-15 baseline)Accepted limitation — container-to-host access is not NetBird-gatedCommitted production secrets (2026-08-15) — OPEN, owner deferredBackup coverage gap (2026-08-14)Backup implementation finding (2026-09-08)Platform evaluation decisionsSwarm role-layout rolloutHomelab Pi-hole administrationHome Assistant routeHomelab LAN certificate export