next-steps
**Documentation Alignment Complete:** All references in state files (next-steps.md, already-implemented.md, requirements.md, accepted-suggestions.md) have been updated to reflect the new folder structure with infrastructure stack separation (`runtime/stacks/infrastructure/*` and `runtime/environments/{dev,vps}/infrastructure/*`). Test environment references have been removed from operational docs (dev-only for templates). Legacy path patterns have been eliminated.
Next Steps
Current Snapshot (2026-04-28)
Documentation Alignment Complete: All references in state files (next-steps.md, already-implemented.md, requirements.md, accepted-suggestions.md) have been updated to reflect the new folder structure with infrastructure stack separation (runtime/stacks/infrastructure/* and runtime/environments/{dev,vps}/infrastructure/*). Test environment references have been removed from operational docs (dev-only for templates). Legacy path patterns have been eliminated.
-
Runtime operator guide set added under
docs/runtimefor master setup flow, per-stack rollout, environments, CI, service deployments, and wrapper-script usage. -
Cross-platform launcher scripts added under
runtime/scriptsfor active infrastructure and service compose targets. -
Edge stack operational behind Traefik with NetBird-only admin access model.
-
NetBird core deployed and integrated with Existing Traefik routing.
-
Database engines and DB admin UIs are deployed from runtime/stacks/infrastructure/databases.
-
Databases runbook includes explicit stop/remove commands for all DB and DB UI containers.
-
Git ignore policy tracks named env files under runtime/environments while keeping plain
.envfiles ignored. -
Tracked named env files under runtime/environments/vps now keep example-only password/secret placeholders while preserving host/domain values.
-
Runtime structure migration started (docs/docker -> runtime/stacks, docs/ci -> runtime/ci, non-production templates -> runtime/environments/dev, and active VPS env values -> runtime/environments/vps).
-
Watchtower split to runtime/stacks/infrastructure/operations is deployed on server with Docker API compatibility fix and active schedule.
-
Shared service stacks are deployed in enforced order (platform -> operations -> registry -> feeds).
-
Registry basic-auth cutover (Traefik middleware + Registry Admin basic mode + NetBird-only UI) deployed and validated.
-
Object storage baseline is RustFS from
runtime/stacks/infrastructure/object-storage/rustfs; MinIO has been decommissioned. -
Database and Redis cutover to strict NetBird-only still pending final validation.
-
Redis developer hostname access model (
redis.perspective-v.comover NetBird-only firewall) applied on server; final NetBird/non-NetBird client validation still pending. -
Split admin-panel auth policy applied in runtime/stacks/infrastructure/edge compose (dashboard keeps Traefik basic auth; built-in-login panels use app auth).
-
Edge base stack artifacts are split into
docker-compose.traefik.ymlanddocker-compose.portainer.ymlwith per-service env folders underruntime/environments/vps/infrastructure/edge/{traefik,portainer,kener}. -
Edge dev env templates are split into
runtime/environments/dev/infrastructure/edge/{traefik,portainer,kener}. -
Recreate Traefik and Portainer from split edge artifacts (
docker-compose.traefik.yml+docker-compose.portainer.yml) in the approved change window. -
Kener parallel-pilot repository artifacts prepared in runtime/stacks/infrastructure/edge (
docker-compose.kener.yml+ README updates), pending server deployment. -
Owner approved Kener pilot rollout; execution is deferred to direct VPS compose session after docs update pass.
-
Top-level folder rename completed:
planning->docsanddeployed->prod, with docs path references updated. -
Set Traefik ACME runtime path to host storage (
/var/lib/traefik/letsencrypt/acme.json) with file mode 600. -
Recreate Traefik from
runtime/environments/vps/infrastructure/edge/traefik/.envand verify ACME entries/routes persist after restart. -
Retire
prod/folder after runtime sync verification.
Recommended Actions (Immediate)
- From a NetBird-connected client, validate Redis access:
nc -vz redis.perspective-v.com 6379
redis-cli -h redis.perspective-v.com -p 6379 -a "<REDIS_PASSWORD>" ping- From a non-NetBird client, validate Redis denial:
nc -vz redis.perspective-v.com 6379- Confirm production services still use internal Docker hostname
redisin runtime env/config. - Validate the new
runtime/scripts/*.shlaunchers from a Linux shell on the VPS before using them in maintenance workflows. - Add tracked placeholder and dev-template env files for
wp.dbskc.comandnew.nishatcolony.pkif those variants need reproducible non-VPS deployments. - Complete final encrypted backup activation: four-site dry-run/real workflow, complete active-engine logical set, first encrypted PostgreSQL physical snapshot, isolated physical restore, scratch cleanup, offsite flags, and all three timers.
- Complete encrypted canary upload/cryptcheck/download validation, one WordPress
scratch restore, and one Tier-1 PostgreSQL scratch restore before setting
DB_OFFSITE_ENABLED=trueor changing production timers. - Monitor the first naturally scheduled WordPress and PostgreSQL physical runs and retain their sanitized evidence. Review the preserved failed physical staging only after the owner decides whether it should be removed.
- Fix
store.gorsistudio.com's service composeenv_file(currently points at the parentgorsistudio.comenv file) so the running container and backup registry agree ongorsistudio_store_db. - After Redis validation, run remaining registry/feed CI-path checks in Phase 2.5.
- Recreate service stacks in
runtime/stacks/servicesafter Watchtower label policy update so running containers includecom.centurylinklabs.watchtower.scope=fast. - Recreate Traefik-exposed service stacks after label hardening so running containers include
traefik.docker.network=proxyand current TLS router settings. - Validate public service routes after recreate (
dbskc.com,nishatcolony.pk,console.perspective-v.com) return HTTPS responses with valid Let's Encrypt certificates. - Validate container health status is
healthyafter recreate for stacks updated with new healthchecks. - Confirm no remaining service stack healthchecks rely on
bashfor Alpine/nginx-based images. - Recreate service containers that should report
com.docker.compose.project=perspective-vif they are still running with legacy project labels. - Recreate only the NetBird dashboard container so updated embedded-IdP env is loaded:
cd /opt/docs/runtime/stacks/infrastructure/netbird
docker compose -f docker-compose.yml up -d --force-recreate dashboard-
Validate NetBird dashboard login no longer returns
Error: Unauthenticated. -
Update CI build/push jobs in
runtime/cito emit Docker schema v2 compatible manifests (oci-mediatypes=false, disable provenance/SBOM attestations when needed for compatibility). -
Re-push active service tags in Docker schema v2 compatible format so Registry Admin catalog remains visible.
-
For registry basic-auth cutover, deploy
registry-adminfirst, thenregistry. -
Validate API challenge header:
curl -Ik https://registry.perspective-v.com/v2/- Expected challenge:
WWW-Authenticate: Basic realm="traefik".
Phase 0: Fresh VPS Baseline
- Place this docs repository on VPS at /opt/docs (or finalized PLAN_DIR).
- Verify Docker and Docker Compose versions.
- Install host prerequisites (curl, jq, ufw, netcat-openbsd, dnsutils).
- Confirm DNS A and AAAA for netbird.perspective-v.com to VPS.
- Apply base UFW policy (public: 80/tcp, 443/tcp, 3478/udp only).
Phase 1: NetBird Core Bring-up
- Deploy edge stack from runtime/stacks/infrastructure/edge and verify Traefik is active on 80/443.
- Install NetBird via quickstart script in /opt/netbird.
- During script, select Existing Traefik option [1].
- Keep NetBird proxy service disabled.
- Create first admin user from /setup.
- Recreate NetBird
dashboardservice after embedded-IdP env correction (AUTH_CLIENT_SECRET=). - Validate NetBird dashboard interactive login succeeds without
Error: Unauthenticated. - Join all required developer/admin machines to NetBird.
- Detect and record active NetBird interface name on VPS (expected wt0).
Phase 2: Private Admin Panels (NetBird-Only)
- Apply split admin-route policy in edge stack (Traefik dashboard keeps
admin-auth; Portainer uses NetBird allowlist + app login). - Validate from NetBird client: dashboard returns 401 before credentials, while Portainer shows app login without Traefik basic-auth challenge (server-side NetBird-address path validation recorded in
docs/operations/validations/2026-04-15-netbird-route-validation.mdx). - Validate admin route denial from non-NetBird host after split policy rollout (recorded in
docs/operations/validations/2026-04-15-netbird-route-validation.mdx).
Phase 2.1: Kener Public Cutover (Built-in Auth)
- Add Kener overlay compose artifact at
runtime/stacks/infrastructure/edge/docker-compose.kener.yml. - Update edge runbook with Kener pilot launch flow and required Kener env variable documentation.
- Populate server-local
runtime/environments/vps/infrastructure/edge/kener/.envwithKENER_HOST,KENER_ORIGIN,KENER_SECRET_KEY, andKENER_REDIS_URL. - Populate server-local
runtime/environments/vps/infrastructure/edge/kener/.envwithSMTP_HOST,SMTP_PORT,SMTP_USER,SMTP_PASSWORD,SMTP_FROM_EMAIL, andSMTP_SECURE. - Ask owner approval for Kener rollout model change.
- Execute approved Kener deployment on VPS using edge stack compose artifacts.
- Validate from public source that
kener.perspective-v.comreturns Kener app route (HTTP 200). - Validate Kener health endpoint on
kener.perspective-v.com/healthcheckreturnsok. - Retire Uptime Kuma runtime artifacts (service/container/volume/image) and remove Kuma service from edge stack artifacts.
- Keep fallback Kener NetBird middleware configuration commented in
runtime/stacks/infrastructure/edge/docker-compose.kener.yml. - Resolve Kener startup warning by deciding Redis eviction policy (
allkeys-lruobserved; Kener recommendsnoeviction) within current Platform resource constraints. - Create Kener owner admin account and baseline pages for monitor separation.
- Manually recreate critical Uptime Kuma monitors in Kener (phase-1 subset).
- Run 3-7 day post-cutover Kener alert fidelity observation window.
Phase 2.5: Shared Services Baseline (Platform -> Operations -> Registry -> Feeds)
- Prepare platform.env from runtime/environments/vps/infrastructure/platform/platform.env.
- Deploy runtime/stacks/infrastructure/platform stack (Redis + RabbitMQ).
- Prepare operations.env from runtime/environments/vps/infrastructure/operations/operations.env.
- Deploy runtime/stacks/infrastructure/operations Watchtower stack and verify scheduled label-based execution.
- Add dual Watchtower runtime definitions in operations stack (
watchtowerbaseline scopenone+watchtower-fastscopefast, 1-hour poll interval). - Add WUD compose service in runtime/stacks/infrastructure/operations with pinned image, Discord trigger variables, and NetBird-only Traefik labels for
wud.perspective-v.com. - Add WUD non-production placeholders in runtime/environments/dev/infrastructure/operations/operations.dev.env.
- Add live WUD values in server-local runtime/environments/vps/infrastructure/operations/.env (set real Discord webhook, keep tracked env placeholders secret-free).
- Deploy/recreate operations stack to start WUD container alongside Watchtower.
- Recreate operations stack on VPS to start
watchtower-fastand apply baseline/fast scope split. - Validate scope separation on VPS:
watchtower-fastupdates onlycom.centurylinklabs.watchtower.scope=fast, while baselinewatchtowerhandles default scopenone. - Validate WUD non-NetBird denial from server-source check (HTTP 403).
- Validate WUD route from a separate NetBird client at wud.perspective-v.com.
- Validate Discord notification delivery using a controlled WUD trigger test.
- Configure WUD custom private registry provider for
registry.perspective-v.comwith authenticated access. - Validate WUD now lists container inventory and shows
custom.pvon the registries page/API. - Create dedicated
wud-monitorregistry basic-auth credential and wire WUD to use it instead of admin. - Recreate registry and operations stacks after credential update and validate WUD still reports
custom.pvsuccessfully. - Keep WUD in discovery mode on VPS (
WUD_WATCHER_LOCAL_WATCHBYDEFAULT=true) and retain selectedwud.tag.includeguardrails for mysql/rabbitmq/redis/verdaccio. - Trigger immediate WUD full scan (
POST /api/containers/watch) after mode switch and validate full inventory is visible (4 -> 25 containers, including dbskc-web). - Decide long-term WUD scope policy after one week of alert-noise observation (stay discovery-default or move back to opt-in labels).
- Execute staged major upgrades with rollback checkpoints and pre-change volume backups for Verdaccio (5 -> 6), Redis (7-alpine -> 8-alpine), RabbitMQ (3.13-management-alpine -> 4.2-management-alpine), and MySQL (8.4 -> 9.6-oraclelinux9).
- Validate post-upgrade service health and core runtime checks (Verdaccio HTTP ping 200, Redis authenticated PING, RabbitMQ broker status, MySQL 9.6 healthy with working root/app auth).
- Run dependent application smoke tests against upgraded Redis/RabbitMQ/MySQL/Verdaccio clients before closing the maintenance window.
- Prepare registry.env from runtime/environments/vps/infrastructure/registry/registry.env.
- Initialize shared registry auth file
/var/lib/traefik/registry-auth/registry.htpasswdon VPS and verify file permissions are restricted. - Validate Traefik registry auth now uses file-provider middleware (
registry-basic-auth@file) from the shared htpasswd source. - Deploy basic-auth registry cutover (registry-admin + registry) from runtime/stacks/infrastructure/registry.
- Validate registry API returns HTTP Basic challenge.
- Validate docker login, push, and pull on registry API with admin credential.
- Validate CI service credentials can login/push/pull after cutover (initial scoped-user validation completed).
- Validate creating/updating a user in Registry Admin updates shared htpasswd and enables
docker loginwithout registry stack recreate (validated fornishatcolony-ciafter Traefik auth-state refresh viadocker restart traefik). - Finalize post-user-change auth refresh runbook as manual Traefik restart (Portainer restart accepted); no helper automation required in this phase.
- Retire legacy
docker-registry-uiservice from prod registry stack and removeregistry-uihost route. - Validate registry API access from non-NetBird CI runner using authenticated docker login.
- Validate registry API rate-limit settings (average/burst) allow normal CI push throughput.
- Validate registry-admin UI is NetBird-only and denied from non-NetBird source (HTTP 403 expected).
- Validate retired
registry-uihostname is no longer routed after service removal (HTTP 404 observed). - Update CI templates in
runtime/cito push Docker schema v2 compatible manifests for Registry Admin (oci-mediatypes=false). - Pin registry runtime image target to
registry:3inruntime/stacks/infrastructure/registry/docker-compose.registry.ymlfor v3 migration implementation start. - Add operator-safe registry cleanup helper script (
runtime/stacks/infrastructure/registry/scripts/registry-cleanup.sh) with dry-run default, tag delete, repository purge, and GC guidance. - Add registry cleanup policy template (
runtime/stacks/infrastructure/registry/cleanup-policy.json) with protected repo/tag controls and monthly include/exclude targeting. - Document registry cleanup and garbage-collection workflow in
runtime/stacks/infrastructure/registry/README.md. - Add registry staging rehearsal helper script (
runtime/stacks/infrastructure/registry/scripts/staging-rehearsal.sh) for standardized v3 pre-cutover checks. - Add registry test env template at
runtime/environments/dev/infrastructure/registry/registry.dev.envfor non-production rehearsal values. - Capture pre-cutover rollback snapshots for registry volumes and image references at
tmp/backups/registry-v3-cutover-20260413-225028. - Execute owner-approved direct VPS cutover for registry service to
registry:3fromruntime/stacks/infrastructure/registry. - Validate immediate post-cutover checks on VPS (
/v2/challenge header, registry container runningregistry:3, and registry log activity on existing repositories). - Populate registry monthly cleanup policy include/exclude rules for real repositories before first apply-mode run.
- Validate monthly registry cleanup job behavior on VPS in dry-run mode first, including protected repository/tag exclusions.
- Re-push required service tags in Docker schema v2 compatible format and confirm Registry Admin catalog visibility.
- Validate RabbitMQ UI from NetBird reaches its built-in login without Traefik basic-auth challenge (recorded in
docs/operations/validations/2026-04-15-netbird-route-validation.mdx). - Prepare feeds.env from runtime/environments/vps/infrastructure/feeds/feeds.env.
- Deploy runtime/stacks/infrastructure/feeds stack (BaGet + Verdaccio).
- Validate NuGet restore from non-NetBird source at nuget.perspective-v.com (HTTP 200).
- Retire legacy feed bridge service/routes from
runtime/stacks/infrastructure/feeds/docker-compose.feeds.ymland keep canonical feeds only. - Simplify feed env contracts in
runtime/environments/dev/infrastructure/feeds/feeds.dev.envandruntime/environments/vps/infrastructure/feeds/feeds.envto canonical keys only. - Add
NPM_FEED_API_KEYto feed env artifacts (runtime/environments/dev/infrastructure/feeds/feeds.dev.env,runtime/environments/vps/infrastructure/feeds/feeds.env, and liveruntime/environments/vps/infrastructure/feeds/.env). - Remove legacy feed bridge bootstrap helper from
runtime/stacks/infrastructure/feeds/scripts. - Normalize feed runbook to canonical endpoint operations in
runtime/stacks/infrastructure/feeds/README.md. - Apply Verdaccio auth hardening in
runtime/stacks/infrastructure/feeds/verdaccio/conf/config.yaml(max_users=-1,access=$all,publish/unpublish=$authenticated) on VPS runtime. - Validate unknown-user
npm adduserself-registration is rejected onhttps://npm.perspective-v.com. - Validate NuGet restore from CI runner using
https://nuget.perspective-v.com/v3/index.json(owner-confirmed; seedocs/operations/validations/2026-04-16-ci-nuget-feed-token-validation.mdx). - Validate NuGet publish on BaGet using API-key model.
- Validate npm install/publish from CI using
https://npm.perspective-v.com/withNPM_FEED_API_KEY(owner-confirmed; seedocs/operations/validations/2026-04-16-ci-npm-feed-api-key-validation.mdx). - Validate canonical endpoints remain available for CI and developer workflows (
nuget.perspective-v.com,npm.perspective-v.com). - Apply npm CI variable/secret cutover in service repositories to
NPM_FEED_URL/NPM_FEED_API_KEYusing the same value asruntime/environments/vps/infrastructure/feeds/.env. - Apply NuGet CI variable cutover in service repositories to
NUGET_FEED_URL/NUGET_FEED_TOKEN. - Retire legacy feed bridge after Verdaccio auth validation, CI cutover completion, and package inventory checks.
- Recreate VPS feeds stack from canonical compose using live
runtime/environments/vps/infrastructure/feeds/.envand--remove-orphansto remove running legacyprogetcontainer. - Capture rollback snapshot/checksum for legacy ProGet volume before cleanup at
tmp/backups/feeds-proget-retirement-20260416T180458Z. - Remove residual legacy ProGet VPS artifacts (
feeds_proget_packagesvolume andproget.inedo.com/productimages/inedo/proget:25.0.25image). - Enforce stack order for baseline rollout: platform -> operations -> registry -> feeds.
- Validate only stateless services keep
com.centurylinklabs.watchtower.enable=true. - Audit Redis consumers and confirm platform Redis is cache/session/ephemeral only.
- If any critical Redis workload exists, plan migration to a dedicated stateful Redis stack with backup/restore tests.
- Create DNS record for redis.perspective-v.com to VPS.
- Apply NetBird-only Redis firewall rule (6379 on NetBird interface, no broad public allow).
- Validate Redis connectivity from NetBird client at redis.perspective-v.com:6379.
- Validate Redis denial from non-NetBird source at redis.perspective-v.com:6379.
- Confirm production service connection strings keep internal Redis host
redis. - Add Redis Insight env values and deploy
redis-insightfrom runtime/stacks/infrastructure/platform (recorded indocs/operations/validations/2026-04-15-netbird-route-validation.mdx). - Validate Redis Insight UI from NetBird client at redis-insight.perspective-v.com (recorded in
docs/operations/validations/2026-04-15-netbird-route-validation.mdx). - Validate Redis Insight UI denial from non-NetBird source (HTTP 403 expected) (recorded in
docs/operations/validations/2026-04-15-netbird-route-validation.mdx).
Phase 2.6: Gateway Baseline (Dev-First)
- Provision
kong_dbandkong_userin existing PostgreSQL for Kong metadata storage (recorded indocs/operations/validations/2026-04-15-netbird-route-validation.mdx). - Add
runtime/stacks/infrastructure/gateway/docker-compose.kong.ymlwith Kong + Konga only (no dedicated Kong DB service). - Add dev gateway env template at
runtime/environments/dev/infrastructure/gateway/kong.dev.envwith placeholder secrets. - Add VPS gateway env artifacts at
runtime/environments/vps/infrastructure/gateway/.envandruntime/environments/vps/infrastructure/gateway/kong.env. - Add gateway runbook at
runtime/stacks/infrastructure/gateway/README.mdaligned with stack conventions. - Deploy gateway stack on VPS from
runtime/stacks/infrastructure/gatewayusingruntime/environments/vps/infrastructure/gateway/.env. - Validate Kong migrations complete successfully against shared PostgreSQL (
postgresonpostgres-network). - Validate
api.perspective-v.comroute reaches Kong proxy through Traefik (HTTP 404 from Kong before route provisioning is expected). - Validate
konga.perspective-v.comnon-NetBird access is denied (HTTP 403). - Validate
konga.perspective-v.comshows Konga built-in login from a NetBird-connected client (server-side NetBird-address path validation recorded indocs/operations/validations/2026-04-15-netbird-route-validation.mdx). - Validate Kong Admin API is not host-exposed and is reachable from Konga over Docker network (
http://kong:8001). - Convert Ocelot routes to Kong services/routes/plugins using
runtime/stacks/infrastructure/gateway/kong-bootstrap-ocelot.shand confirm inventory parity in Kong/Konga (5 services, 8 routes, expected plugin bindings). - Populate real gateway secrets in
runtime/environments/vps/infrastructure/gateway/.envbefore VPS deployment. - Add standalone service stack artifacts for
identity,graph, andconsoleunderruntime/stacks/serviceswith matching VPS tracked env templates underruntime/environments/vps/services. - Align Kong and service upstream connectivity to the existing external
proxynetwork (no dedicated gateway-only network). - In Kong/Konga, backend upstream services and routes are provisioned for
identity(identity:62258) andgraph(graph:5138) under hostapi.perspective-v.com. - Keep Console out of Kong routes: no console route is provisioned in Kong migration set and console remains direct Traefik host
console.perspective-v.com. - Run end-to-end route smoke tests for migrated public/protected gateway paths (JWT allow/deny and rewrite behavior) and capture evidence.
- Capture and store dedicated gateway parity validation record under
docs/operations/validations/2026-04-16-gateway-ocelot-parity-validation.mdxwith route-level evidence. - Add and apply Kong compatibility routes for client URL stability (
/identity/docs/*,/identity/swagger/v1/swagger.json,/swagger/v1/swagger.json, and/resumeas UI passthrough). - Validate public compatibility outcomes:
https://api.perspective-v.com/identity/docs/index.htmlHTTP 200,https://api.perspective-v.com/identity/swagger/v1/swagger.jsonHTTP 200,https://api.perspective-v.com/swagger/v1/swagger.jsonHTTP 200, andhttps://api.perspective-v.com/resumeupstream HTTP 400 (no Kong route-miss 404). - Add and apply Kong gopher compatibility routes (
/graph/gopher/{credential,serviceprovider,serviceprovideremail,platform}as explicit public API routes, plus/gopher/*UI passthrough) with priorities abovegraph-protected. - Validate gopher compatibility outcomes: graph gopher API paths now return upstream responses (
GET400 /POST415 where applicable) instead of pre-fix route/auth failures (GET404 /POST401), and UI paths/resumeplus/gopher/{credential,serviceprovider,serviceprovideremail,platform}resolve through Kong as passthrough UI routes. - Extend Kong global CORS allow-headers with Apollo client headers (
apollographql-client-name,apollographql-client-version) to satisfy browser preflight for console-origin graph API calls. - Update live Kong CORS allowed origins to include
https://console.perspective-v.comandhttps://hassan.taj.contact, then validate graph endpoint preflight returns HTTP 200 for both origins. - Add
https://hassantaj.github.ioto live Kong CORS allowed origins and enableOPTIONSon/graph/resume, then validate preflight returns HTTP 200 with matching allow-origin. - Create a complete Kong route/service mapping guide covering service/route additions, rewrite/auth/plugin model, and validation workflow.
- Add repository-side graph parity extension in gateway bootstrap for
/graph/docs/*(public docs rewrite) and/graph/v1/*(identity-style JWT-protected rewrite to/api/v1/*) while keeping existing graph compatibility routes unchanged. - Add rollout artifacts for graph v1/docs under
docs/operations/validations/2026-04-21-graph-v1-kong-access-validation.mdxanddocs/operations/change-records/cr-2026-04-21-graph-v1-kong-route-rollout.mdx. - Add Hoppscotch-importable graph REST template at
runtime/stacks/infrastructure/gateway/client-imports/graph-v1-hoppscotch-openapi.jsonand link import steps in gateway README. - Execute owner-approved VPS apply for graph v1/docs bootstrap changes and capture runtime route/plugin inventory evidence.
- Run post-apply route smoke validation for
/graph/docs(public) and/graph/v1/{endpoint}(JWT allow/deny) and record observed HTTP outcomes. - Add declarative Kong state file
runtime/stacks/infrastructure/gateway/kong.ymlfor decK gateway sync/diff automation and document sync commands in gateway README. - Install decK natively on VPS host (non-Docker), validate
runtime/stacks/infrastructure/gateway/kong.ymllocally withdeck file validate, and verify online Kong Admin connectivity withdeck gateway ping. - Remove obsolete legacy
graphservice/routes fromruntime/stacks/infrastructure/gateway/kong.ymlso decK dry-run no longer proposes unintended entity creation. - Decide and execute JWT secret alignment in live Kong via host-native decK sync (
envsubstrendered state): credentialocelot-jwt-issuernow matches VPSKONG_JWT_HMAC_SECRETand post-sync diff is zero. - Resolve upstream graph docs outage by refreshing
graph:latestruntime image and recreating the graph container;https://api.perspective-v.com/graph/docs/now serves Scalar UI. - Investigate upstream graph application responses on
/api/v1/healthif non-auth errors persist under JWT-authenticated probes. - Apply decK sync after removing JWT from
graph-resume-get-by-token-publicso/graph/v1/resume/GetByAccessTokenis public.
Phase 3: Object Storage Baseline (RustFS)
- Add RustFS compose artifact at
runtime/stacks/infrastructure/object-storage/rustfs/docker-compose.rustfs.ymlwith pinned image, NetBird-only routes, and explicit Traefik router-service bindings. - Add RustFS dev env template at
runtime/environments/dev/infrastructure/object-storage/rustfs/rustfs.dev.env. - Prepare server-local RustFS runtime values in
runtime/environments/vps/infrastructure/object-storage/rustfs/.env(no tracked secrets). - Set explicit non-default RustFS API credentials in server-local env (
RUSTFS_ACCESS_KEYandRUSTFS_SECRET_KEY). - Create
/var/lib/rustfs/dataand/var/lib/rustfs/logson VPS with ownership10001:10001. - Deploy RustFS from
runtime/stacks/infrastructure/object-storage/rustfsand confirmrustfscontainer health. - Validate RustFS console/API through NetBird-address path (SNI
--resolveto server NetBird IP) showing backend responses (console indexHTTP 200, API auth gate as XMLAccessDenied). - Validate RustFS console/API from NetBird client machines (successful key-login path on
rustfs.perspective-v.comwith server URLrustfs-api.perspective-v.com). - Validate RustFS route denial from non-NetBird sources (HTTP 403 expected).
- Decommission MinIO runtime completely (container, data volume, images, and runtime stack/env artifacts removed).
- Update object-storage runbook and tracking docs to RustFS-only operating model with decommission rationale.
- Sweep runtime app env/service-stack artifacts and remove residual MinIO host or
MINIO_*references (RustFS-only endpoint standard). - Validate app-container S3 access to RustFS over object-storage Docker network.
- Validate signed URL delivery through app layer using RustFS backend.
- If first certificate issuance fails under strict SNI, temporarily set
sniStrict=false, issue certs for RustFS hosts, then restoresniStrict=true.
Phase 4: Databases (MSSQL + PostgreSQL + MySQL + MongoDB)
- Prepare mssql.env, postgres.env, mysql.env, and mongodb.env from
runtime/environments/dev/infrastructure/databases/<engine>/<engine>.dev.envtemplates. - Create DNS records for mssql/pgsql/mysql/mongo and pgadmin/phpmyadmin/mongo-express hostnames.
- Keep DB/UI DNS hostnames as public records; do not create a separate NetBird-only DNS zone for this phase.
- Complete live DB credential rotation validation for all engines after VPS
.envsecret hardening (MSSQL/PostgreSQL/MySQL/MongoDB validated in running containers). - Deploy MSSQL, PostgreSQL, MySQL, and MongoDB from runtime/stacks/infrastructure/databases.
- Deploy pgAdmin, phpMyAdmin, and mongo-express admin profiles.
- Add pgAdmin-specific Traefik middleware override to keep NetBird-only access while allowing SAMEORIGIN frame behavior for Query Tool.
- Recreate PostgreSQL from
runtime/environments/vps/infrastructure/databases/postgres/.envand remediate persisted-role password drift so non-local auth uses.envpassword. - Recreate PostgreSQL admin profile to apply updated pgAdmin Traefik labels and verify running container label update.
- Validate pgAdmin Query Tool opens without "refused to connect" from a NetBird client.
- Redeploy/validate Mongo Express built-in auth credentials from VPS
.envand confirm DB UI routes stay on NetBird-only + security headers (no Traefik admin-auth). - Add compose teardown commands in databases README for full DB stack stop/remove workflow.
- In NetBird dashboard, verify VPS peer and required developer/admin peers are connected before connectivity tests.
- Record active NetBird interface and server NetBird IP for cutover commands.
- Validate friendly DB hostnames over NetBird (mssql/pgsql/mysql/mongo) from dedicated NetBird client machines.
- Validate DB management UI hostnames over NetBird (pgadmin/phpmyadmin/mongo-express).
- Validate DB UI hosts from NetBird return app login responses (typically 200/302) without Traefik 401 challenges.
- Validate DB UI route denial from non-NetBird source (HTTP 403).
- Run
operations/migrations/netbird-db-port-cutover.shwith the detected NetBird interface. - Confirm public DB/Redis access is closed and NetBird-only DB/Redis access works from separate NetBird/non-NetBird clients.
Phase 5: Hardening and Resource Optimization
- Install and tune Fail2ban.
- Finalize interim backup destination as RustFS on the same VPS for immediate baseline coverage.
- Finalize critical-service recovery targets at RTO 4 hours and RPO 1 hour.
- Add Tier-1 database backup helper at
operations/backups/db-backup.shwith strict rolling keep-last-backups retention and manifest output. - Add systemd scheduling artifacts under
operations/systemd/. - Add approval-ready execution bundle for first backup rollout at
docs/operations/execution/phase5-backup-rollout-execution-bundle.mdx. - Add prefilled change record draft at
docs/operations/change-records/cr-2026-04-15-tier1-backup-timer-rollout.mdx. - Add first-run evidence helper at
operations/diagnostics/collect-tier1-backup-evidence.sh. - Implement automated backups for Tier 1 stateful services and keep retention documented (first successful run captured at
/home/repo/contabo-server-setup/tmp/backups/tier1-20260415T205728Z). - Install approved production scheduling for backup automation and keep first-run evidence (
tier1-db-backup.timeractive; evidence file/home/repo/contabo-server-setup/tmp/backups/evidence/tier1-backup-evidence-20260415T205803Z.txt). - Update repository Tier-1 backup policy defaults to weekly Sunday 03:00 UTC and strict keep-last-3 retention across scripts, installer, systemd templates, and runbook references.
- Ask owner approval, then apply updated weekly keep-last-3 timer/service policy on VPS and capture refreshed timer/evidence outputs (completed; evidence
/home/repo/contabo-server-setup/tmp/backups/evidence/tier1-backup-evidence-20260416T222414Z.txt). - Run and document one end-to-end restore validation against RTO/RPO targets (
docs/operations/restore-validations/rv-20260415t210554z-postgres-tier1.mdx; evidence/home/repo/contabo-server-setup/tmp/backups/evidence/restore-validation-20260415T210554Z.txt). - Add restore drill template at
docs/operations/templates/restore-validation-template.mdx. - Finish the encrypted Google Drive rollout for four WordPress sites and active PostgreSQL/MySQL/MongoDB engines, including physical PostgreSQL recovery validation, offsite flags, sliding retention, diagnostics, and all three timers.
- Keep monitoring lightweight for 4 vCPU / 8 GB (no heavy monitoring stack yet) (
docs/operations/baselines/lightweight-monitoring-baseline.mdx). - Schedule controlled monthly DB maintenance updates (
docs/operations/baselines/monthly-db-maintenance-schedule.mdx).
Phase 5.1: Governance and Runbooks
- Add disaster recovery policy at
docs/operations/governance/disaster-recovery-plan.mdx. - Add incident response playbook at
docs/operations/governance/incident-response-playbook.mdx. - Add post-incident host integrity checklist at
docs/operations/governance/host-integrity-checklist.mdx. - Add production change control policy at
docs/operations/governance/change-control-policy.mdx. - Add production change record template at
docs/operations/templates/production-change-record-template.mdx. - Add incident closure template at
docs/operations/templates/incident-closure-template.mdx. - Add NetBird access validation matrix at
docs/operations/governance/netbird-access-validation-matrix.mdx. - Execute one tabletop drill using the new incident-response and host-integrity runbooks (
docs/operations/tabletop-drills/tt-20260415t220000z-ssh-egress-policy.mdx). - Execute one approved production change using the change-control policy template and record outcomes (
docs/operations/change-records/cr-2026-04-15-tier1-backup-timer-rollout.mdx).
Phase 6: App and Feed Migration (After Secure Baseline)
Superseded (2026-07-01): the package layer described here (registry.perspective-v.com + BaGet/Verdaccio/ProGet feeds and per-service
REGISTRY_USERNAME/REGISTRY_PASSWORDbasic-auth) has been replaced by the Gitea package registry and those 5 services were retired — see the "Gitea Unified Package Registry Migration (2026-07-01)" entry above. Open ([ ]) items below that still referenceregistry.perspective-v.comor old feed hosts are obsolete; the equivalent work is done againstgitea.perspective-v.com/perspective-v/*with token auth.
- Deploy production services from runtime/stacks/services/perspective-v.
- Confirm service images resolve from registry.perspective-v.com with valid docker login.
- Validate services requiring object storage are attached to object-storage network.
- Prepare reusable CI templates under runtime/ci for GitHub Actions and Azure Pipelines.
- Add dbskc CI templates under runtime/ci for GitHub Actions and Azure Pipelines with branch trigger
deploy/dbskc. - Add nishatcolony CI templates under runtime/ci for GitHub Actions and Azure Pipelines with branch trigger
deploy/nishatcolony. - Align identity/graph/console CI templates under runtime/ci to deploy branches (
deploy/identity,deploy/graph,deploy/console) for both GitHub Actions and Azure Pipelines. - Enable Watchtower labels for stateless application service compose templates under
runtime/stacks/services. - Copy service-specific CI templates (identity/graph/console) into matching repositories and set REGISTRY_USERNAME/REGISTRY_PASSWORD secrets.
- Copy dbskc CI templates into the dbskc application repo(s) and set REGISTRY_USERNAME/REGISTRY_PASSWORD secrets.
- Copy nishatcolony CI templates into the nishatcolony application repo(s) and set REGISTRY_USERNAME/REGISTRY_PASSWORD secrets.
- Add a dedicated
dbskc-ciregistry basic-auth user on VPS and set matching plaintextREGISTRY_USERNAME/REGISTRY_PASSWORDin Azure pipeline variables. - Add a dedicated
nishatcolony-ciregistry basic-auth user on VPS and set matching plaintextREGISTRY_USERNAME/REGISTRY_PASSWORDin Azure pipeline variables. - Run dbskc CI from
deploy/dbskcand confirmregistry.perspective-v.com/dbskc-webreceives bothlatestandv1.0.<run>tags. - Run nishatcolony CI from
deploy/nishatcolonyand confirmregistry.perspective-v.com/nishatcolony-webreceives bothlatestandv1.0.<run>tags. - Validate dbskc Azure pipeline registry login returns
Login Succeededafter credential sync. - Prepare dbskc VPS runtime env at
runtime/environments/vps/services/perspective-v/dbskc/.envwith tracked placeholderruntime/environments/vps/services/perspective-v/dbskc/dbskc.env. - Prepare nishatcolony VPS runtime env at
runtime/environments/vps/services/perspective-v/nishatcolony/.envwith tracked placeholderruntime/environments/vps/services/perspective-v/nishatcolony/nishatcolony.env. - Mirror Perspective-V service env layout by service under
runtime/environments/{dev,vps}/services/perspective-v/{identity,graph,console,dbskc,nishatcolony}. - Populate server-local service-local
.envfiles for identity/graph/console atruntime/environments/vps/services/perspective-v/{identity,graph,console}/.env. - Ask owner approval, then run per-service deployment/recreate from
runtime/stacks/services/perspective-v/{identity,graph,console}using service-local env files. - Validate DNS for
dbskc.comresolves to VPS before first dbskc deploy. - Finalize nishatcolony public host as
nishatcolony.pk(instead ofapp.nishatcolony.pk) and validate DNS resolves to VPS before first nishatcolony deploy. - Deploy dbskc stack from
runtime/stacks/services/perspective-v/dbskcusing VPS env source and verify container healthy. - Deploy nishatcolony stack from
runtime/stacks/services/perspective-v/nishatcolonyusing VPS env source and verify container healthy. - Validate dbskc HTTPS route response on
https://dbskc.comand confirm Traefik router/certificate behavior. - Validate nishatcolony HTTPS route response on
https://nishatcolony.pkand confirm Traefik router/certificate behavior. - Validate one controlled dbskc update path: push a new
latestimage and confirmwatchtower-fastapplies rollout within configured fast interval. - Create the versioned Gitea auth Swarm secret and deploy the updated
watchtower-faststack. - Confirm the next
watchtower-fastpoll scansdev-dbskc-webwithoutunauthorized. - Merge
dbskcintodeploy/dev.dbskc.com, builddev-dbskc-web, bootstrapwebsites-dev-dbskc, and validatehttps://dev.dbskc.complus fast Watchtower rollout. - Validate one controlled nishatcolony update path: push a new
latestimage and confirmwatchtower-fastapplies rollout within configured fast interval. - Run CI pipeline once per service and confirm both latest and
v1.0.<run>tags are pushed to registry.perspective-v.com. - Bootstrap Perspective-V services on VPS once, then keep rolling updates on Watchtower-managed
latesttags. - Validate one controlled baseline Watchtower rollout (push new
latestfor one non-fast service and confirm Sunday schedule is applied). - Execute phased cutover per domain plan.
- Add syassociates compose artifacts under
runtime/stacks/services/syassociates.pkforsyassociates.pkandservice.syassociates.pk. - Deploy and validate syassociates frontend HTTPS and Kong-backed API routing after owner approval.
Phase 9: Docs Taxonomy Normalization
- Keep decision/state source-of-truth files under
docs/state/. - Keep operational artifacts and records under
docs/operations/. - Keep setup and operator guides under
docs/runtime/. - Normalize state-document references from legacy
docs/*paths todocs/operations/*where applicable. - Update feed migration guidance in
docs/runtime/legacy-setup-guide.mdxto canonical feed routing model. - Run a final docs path/read-through pass; active Swarm role paths and stack identities no longer use the superseded layout.
Security Incident Follow-up (2026-04-11)
- Archive incident package under
Incidents/2026-04-12-0015-ssh-outbound-spike/with markdown reports and raw evidence files. - Draft provider-ready Contabo response and attach supporting evidence references.
- Run same-day post-containment revalidation and archive evidence files
23through30(live SSH egress snapshot, process attribution, scheduler scan, Uptime Kuma monitor audit, and control snapshots). - Verify whether accepted SSH login from
86.121.67.247at 2026-04-09 12:00 UTC+0 was authorized; if not authorized, treat as compromise indicator. - Verify that established inbound SSH source IPs observed during revalidation are authorized (
202.66.181.251and36.132.36.134in evidence file25). - Rotate root password and regenerate/rotate SSH keys in
authorized_keys; remove any unknown key material. - Disable provider VNC/remote console access for normal operations.
- Reset VS Code SSH and Tabby SSH access credentials/keys and keep trusted principals only.
- Finalize permanent outbound SSH policy: dual-stack deny-by-default with logging-enabled TCP/22 block, GitHub access via
ssh.github.com:443, owner-approved one-hour exception workflow, and Discord alert path. - Implement outbound SSH policy tooling (
runtime/stacks/infrastructure/operations/ssh-egress-policy.sh) and enable recurring maintenance timer. - Investigate 2026-04-17 SSH egress Discord alert flood and identify root cause from VPS evidence (detector false positives caused by inbound
UFW BLOCKmatching plusDPT=22prefix matching22xxports). - Patch
runtime/stacks/infrastructure/operations/ssh-egress-policy.shblocked-scan filter to detect only outbound exactDPT=22events (IN= OUT=<iface>+DPT=22exact) and exclude22xxmatches. - Fix blocked-detection Discord formatter in
runtime/stacks/infrastructure/operations/ssh-egress-policy.shby escaping Markdown code-fence backticks so systemd logs no longer showtext: command not found/[UFW: command not foundduring alert rendering. - After owner approval, apply patched SSH egress policy script on VPS and validate with one synthetic outbound TCP/22 probe plus one inbound
22xxcheck to confirm only true outbound SSH events alert Discord (completed 2026-04-17: false-positive window replay returned 0 matches, synthetic outbound probe produced only exactDPT=22outbound matches). - Verify Discord render quality for embed-card SSH egress alerts using underlined heading plus ordered-list body during one blocked-detection event and one exception open/close cycle.
- Run one controlled exception drill (open and close) using full audit fields to validate runbook quality under change pressure.
- Verify effective SSH auth matrix (
sshd -T) and confirm no password fallback/auth-method drift remains after access reset. - Send Contabo the prepared incident summary (source process, destination pattern, containment actions, current status) and keep provider confirmation in the same incident folder.
- Add incident response playbook and host integrity checklist in
docs/operations/for repeatable post-incident operations. - Complete deep host integrity checks (package/service audit, startup persistence sweep, credential review) and decide on rebuild-vs-recovery posture.
Docker Swarm Migration (2026-06-30)
- Swarm migration plan finalized with Docker Secrets, pre-built WordPress images, separate
swarm/directory -
swarm/stacks/— active Swarm services are maintained as one YAML fragment per service under six role directories. -
swarm/configs/traefik/— Swarm provider config + file-based middlewares (@fileinstead of@docker) -
swarm/secrets/—create-all-secrets.sh+secrets-map.mdfor 30+ Docker secrets -
swarm/scripts/— role-organized stack launchers plus guarded panel management. -
swarm/README.md— deployment docs, dependency graph, rollback procedure -
docs/runtime/swarm-migration.mdx— migration guide and rationale - WordPress images built and pushed to
registry.perspective-v.com(dbskc-web, nishatcolony-pk, gorsistudio-web, gorsistudio-store, wcblahore-pk) - VPS pre-flight: full DB backup via tier1 scripts, volume snapshot (
tmp/backups/swarm-preflight-volumes-*.txt) - VPS cutover: swarm init, overlay networks, secrets, stacks deployed
- Post-migration validation: services 1/1, HTTP checks, DB connectivity, NetBird access, ACME certs
Post-cutover fixes (2026-07-01)
- Data preservation: all stacks map data volumes to the original compose volumes via
external: true(e.g.netbird_netbird_data,databases_postgres-data); pre-migration data intact. Orphaninfra-*/svc-*volumes from the first (pre-mapping) deploy are unused. - Traefik ports switched to
mode: host(nowswarm/stacks/edge/traefik.yml) — swarm ingress SNAT hid the client IP (10.0.0.2), which broke thenetbird-only(100.64.0.0/10) allowlist on all admin UIs. Admin UIs now reachable over the NetBird VPN. -
create-all-secrets.shquote-strip fix —.envvalues likeMYSQL_PWD="…"were stored with literal quotes, breaking auth; WP +kener_secret_key+vaultwarden_admin_tokensecrets recreated cleanly. WordPress sites (dbskc, gorsistudio, store, wcblahore) now 200. - Split-stack env sourcing:
edge.sh/.batpreloads Kener's env and the panels launcher loads each panel's owning subsystem env; website launchers preserve per-service database variables..sh/.batpairs are reconciled. -
_FILE-less services fed from secrets via command wrappers: kener (KENER_SECRET_KEY,SMTP_PASSWORD), vaultwarden (ADMIN_TOKEN). - arnexglobal healthcheck fixed (
wget→bash TCP probe); zitadel/ui/v2/loginrouter added; Traefik dashboard router fixed (dummy service port);dbskc-web-coming-soonduplicate removed. - Traefik dashboard basic auth reset (
admin/ htpasswd at/var/lib/traefik/registry-auth/registry.htpasswd). - phpMyAdmin extra themes (darkwolf, boodark) bind-mounted from
/var/lib/phpmyadmin/themes/*(swarm/stacks/panels/phpmyadmin.yml).
Remaining follow-ups (blocked / optional)
- Migrate the eight administration UIs to
panel, preserving image digests, persistent data, access rules, and the prior replica baseline; controlled pgAdmin0 → 1 → 0validation passed. - Complete the interim Vaultwarden, Zitadel, and Gitea
svc-*rename; this identity was later superseded by the header-defined namespace rollout below. -
svc-pullerlogin restored (2026-08-15). Root cause was not a bad or expired token:/root/.docker/config.jsonhad an emptyauthsobject, so the node held no credential forgitea.perspective-v.comat all. The account itself was healthy (active, 1 token) and the registry challenged correctly (GET /v2/-> 401 with Bearer + Basic; noteHEAD /v2/returns 405, so probe with GET). Existing token values are unrecoverable (Gitea stores hashes), so a freshread:packagetoken was minted withgitea admin user generate-access-token -u svc-puller --scopes read:package --raw(run as thegituser — the CLI refuses to run as root) and piped straight intodocker login --password-stdinso the value was never printed. Verified: all 11gitea.perspective-v.com/perspective-v/*images pull successfully.- The previous unusable token is still on the account; revoke it in the Gitea UI when convenient.
docker loginstores credentials base64-encoded (not encrypted) in/root/.docker/config.json— standard Docker behavior, worth knowing for host-integrity reviews.
-
data.arnexglobal.comTLS — BLOCKED: DNS A record points to66.165.248.146, not the server161.97.83.142; ACMEtlsChallengefails until the record is repointed. App itself runs fine. - Diagnose the pre-existing HTTP 500 from
new.nishatcolony.pk; thewebsites-nishatcolony_nishatcolony-pktask is healthy at 1/1 and its data volume was preserved. -
perspective-v.comapex andsyassociates.pknot routed/deployed (syassociates deploys via template-env fallback when needed).
Gitea Unified Package Registry Migration (2026-07-01)
A root-level working log
/home/repo/contabo-server-setup/next-steps.mdexists from this migration and should be consolidated into this file;docs/state/next-steps.mdxis canonical.
- Deploy the package-only registry at
gitea.perspective-v.com(git disabled), now running asserviceat 1/1 with valid LE cert; org ownerperspective-v, adminpvadmin. - Migrate all package data into Gitea: NuGet 45 versions (BaGet), Docker 20 tags/11 repos (old registry), npm
@pv/core14 versions (Azure DevOpspv-ng); Verdaccio empty. - Repoint CI/CD templates to Gitea (
runtime/ci/github-actions/*,runtime/ci/azure-pipelines/*,runtime/ci/README.md). - Reconfigure + prepare WUD to watch Gitea (dedicated
wud-monitorread:packagetoken and recreatedwud_registry_passwordsecret); WUD is now sourced fromswarm/stacks/panels/wud.yml. - Cut over all 11 running services to
gitea.perspective-v.com/perspective-v/*; all are 1/1. The later header-defined rollout restored the drifted Nishat web source to its Gitea image and restored the separatenishatcolony.pkrouter; the intended WordPress route atnew.nishatcolony.pkstill returns its pre-existing 500. - Retire legacy 5 services: removed
infra-registry(registry + registry-admin) andinfra-feeds(proget + baget + verdaccio); old hosts now 404. Removed secretsregistry_*,baget_api_key,npm_feed_api_key. - Publish guide
docs/runtime/stacks/gitea.mdxand retirement runbookdocs/runtime/gitea-retirement-runbook.mdx.
Remaining follow-ups (optional cleanup — needs go)
- Legacy data volumes removed by the owner via Portainer on 2026-08-15 (32 -> 24 volumes). Note: the documented precondition — verifying ProGet held no unmigrated NuGet packages before pruning its volume — was not performed first, so that check is now moot. Only BaGet's 45 versions were ever migrated to Gitea.
- Retired
registry.perspective-v.com/*image tags removed from the node (2026-08-15). They were byte-identical duplicates: each shared the same image ID as itsgitea.perspective-v.com/perspective-v/*counterpart, so removing them only dropped the obsolete tag and freed no image data. - Remove old DNS records:
registry.,registry-admin.,nuget.,npm.,proget.perspective-v.com. - Remove retired Registry/Feeds Swarm manifests, launchers, and teardown helper.
- Remove retained Registry/Feeds Compose environment artifacts and remaining stale secret/documentation references after their rollback window.
- Fix
.claude/settings.local.jsonhealth check(s) still pointing atproget.perspective-v.com. - Validate WUD in
paneland both Watchtower services inplatform; all three are healthy after migration. The separatesvc-pullerregistry-login follow-up remains open below.
Container image update rollout (2026-08-14)
Execution bundle: docs/operations/execution/image-update-rollout-execution-bundle.mdx.
WUD reported 7 available updates; 2 were rejected as unsafe after verification.
- Verify all 7 WUD-reported updates against upstream registries (Docker Hub, GHCR, MCR).
- Reject Redis
8-alpine->32bit-stretch: that tag is a 2019 32-bit Redis 5.x build; applying it would be a 3-major downgrade and a Redis 8 AOF/RDB will not load under it. Cause is a missingwud.tag.includelabel on the running container, not a manifest defect. - Reject Gitea
1.27-rootless: production runs the non-rootless image; rootless uses/var/lib/gitea+/etc/giteaand a different entrypoint, so thescm_gitea_datavolume at/dataand themenu.tmplbind would both be ignored. Correct target isgitea/gitea:1.27.2. - DEV-rehearse Gitea 1.22 -> 1.27.2 against a mirrored Compose stack: schema migrated 299 -> 343, users/packages preserved, package download byte-identical,
/v2/401 intact, custom org template renders. - Fix latent Gitea wrapper bug found in rehearsal:
s6-svscanlives at/bin/in 1.22 but/usr/bin/in 1.27, so the hardcoded/bin/s6-svscanexits 127 on upgrade.swarm/stacks/services/gitea/gitea.ymlnow resolves it viaPATH(verified in both images). - DEV-verify WUD 8.3.1 boots with the existing command wrapper (
dist/index.jsstill present); no manifest change needed beyond the image bump. - Confirm the two Zitadel WUD entries are one upgrade and that
cf6c2e88is the amd64 manifest of the pinnedv4.15.2index (no repo/production drift).
Executed 2026-08-15 (all validated, backups in /var/backups/preupgrade-20260814/)
- Window 0: stack names resolved — production is on the renamed stacks (
platform,service,panel,service-zitadel). Theinfra-*names in WUD's report were stale store entries for containers that no longer exist. - Window 1 not required:
platform_redisalready carried thewud.*labels; the real defect was label placement (see below). - WUD 8.2.2 -> 8.3.1.
- Vaultwarden 1.36.0 -> 1.37.1 (Rocket launched, no errors).
- Gitea 1.22 -> 1.27.2: schema migrated 299 -> 343; 21 packages / 113 versions / 5 users preserved;
/v2/401 intact; custom org template renders. - Zitadel core + login -> v4.17.1 in lockstep; migrations verified;
/->/ui/v2/loginreturns 200. - Kener, Redis (8.10.0), RabbitMQ (4.3.4), MySQL (9.7.2), Kong (3.9.3) updated; all WordPress sites, API gateway routes, and queue/DB checks pass.
- RustFS
rc.1attempted and ROLLED BACK to1.0.0-beta.6-glibc. rc.1 requires--console-enable(beta.6 starts the console by default), moves the console from/to/rustfs/console/, and rejects the storedpublicbucket policy (unknown field ID, expected Id), which broke anonymous public-object reads with HTTP 500. Public object serving verified restored (HTTP 200) after rollback.
WUD reporting defects found and fixed (2026-08-15)
-
wud.*labels were underdeploy.labels, which Swarm applies to the service; WUD reads container labels, sowud.tag.includehad never taken effect. Moved to service-levellabels:on redis and added guardrails to traefik, mysql, postgres, mongodb, rabbitmq, gitea, vaultwarden, rustfs. Before the fix WUD proposedtraefik -> v3.7-windowsservercore-ltsc2025,mysql -> 26.7-oraclelinux9,postgres -> 19beta1-master,rabbitmq -> 4.3-rc-management-alpine, andredis -> 32bit-stretch; all are now filtered. - Docker Hub digest watching was off.
hub/Hub.jsoverridesshouldWatchDigest()and returns false unlessWUD_REGISTRY_HUB_PUBLIC_WATCHDIGEST=trueor a per-containerwud.watch.digestlabel is set — other registries default to true. That is why GHCR (Zitadel) reported digest updates but every Hub:latestimage was skipped with "not a semver and digest watching is disabled". The watcher-levelWUD_WATCHER_LOCAL_WATCHDIGESTis deprecated in 8.x and does not control this. - Discord notifications were failing entirely with HTTP 400.
MODE=batchrenders one message and the verbose upstream body exceeded Discord's limit, so nothing was delivered. Replaced with fixed-widthcontainer | current | new | typerows. - Batch message reformatted as a single fenced code block. Discord does not render markdown tables (pipes stay literal), and upstream
renderBatchBodyhardcodes a-markdown bullet per container with no configuration to disable it. Overrode that one method via a read-only bind mount atswarm/stacks/panels/custom/wud/Trigger.js(same pattern already used for Gitea'smenu.tmpl), emitting one ```vb fence containing a header row plus one aligned row per update. New env knobs:WUD_TRIGGER_BATCH_LANG,WUD_TRIGGER_BATCH_HEADER,WUD_TRIGGER_BATCH_MAXCHARS.- Maintenance: the mounted
Trigger.jswas extracted from image8.3.1. Re-extract and re-apply the patch on every WUD upgrade, otherwise a stale copy of this file is mounted over the new image's version. - Body templates are JS template literals eval'd as
eval('`'+template+'`')and must not contain a literal backtick — useString.fromCharCode(96)if one is ever needed. WUD_TRIGGER_BATCH_MAXCHARSguards Discord's 1024-character embed field-value limit and appends... +N morerather than failing the whole webhook.
- Maintenance: the mounted
Second pass — all remaining stable updates applied 2026-08-15
- Digest refreshes applied and validated:
postgres(PostgreSQL 18.4 / PostGIS 3.6, 9 DBs intact),mongodb,traefikv3.7 (routes + ACME cert preserved),portainer,netbird-server,netbird-dashboard, and the scaled-to-zeropgadmin(9.16 -> 9.17),redis-insight,mssql. Fullpg_dumpalltaken first (19 MB, 22 databases). - RustFS beta.6 ->
1.0.0-rc.2-glibc— the rc.1 rollback blockers were fixed first, so both prior regressions are resolved: anonymous public reads return HTTP 200 with the correct byte count and there are zerobucket_metadata_parse_failedentries.- The stored
publicbucket policy contained a legacy"ID":""field that rc.x rejects. Corrected in place viaPutBucketPolicywhile still on beta.6 (dropping the empty field only, no semantic change); original preserved at/var/backups/preupgrade-20260814/rustfs-public-policy.orig.json. --console-enableadded to the command (rc.x makes the console opt-in).- rc.x serves the console under
/rustfs/console/rather than/, so arustfs-console-rootrouter +redirectregexmiddleware (priority 300) redirects the bare host to it, mirroringzitadel-root./public/(priority 200) and the API host are unaffected. - RustFS has no stable release published — only alpha/beta/rc — so rc.2 is the closest-to-stable option available. It is the only pre-release image in the fleet.
- The stored
- Full sweep re-run: zero digest drift across all 26 public images.
- Confirm from a NetBird client that
https://rustfs.perspective-v.com/redirects to/rustfs/console/and the console loads. The router carriesnetbird-only, so this cannot be verified from the node itself. -
websites-perspective-v_graphhas a newer image in Gitea — intentionally not applied, since pulling it deploys new application code (a release decision, not a maintenance update).
Tooling note — pin digests that are actually addressable
MCR content-negotiates on Accept, and the two answers are not equivalent:
- Index-only
Acceptreturnsdocker-content-digest: sha256:0730f368…for tag2022-latest, but fetching that digest returns 404 — it is not independently addressable. - Broad
Accept(includingmanifest.v2/image.manifest.v1) returnssha256:ba4c8329…, which does resolve by digest and is what Docker can pull.
A pin taken from the index-only header therefore produced an unpullable database_mssql manifest, which would only have surfaced the next time that service was scaled up. Corrected to ba4c8329… and verified. Docker Hub and GHCR return the same digest either way; only MCR differs.
Rule: after changing any pinned digest, verify it resolves by digest (docker manifest inspect <repo>@<digest>), not just that the tag reports it. A sweep that only compares tag headers can both raise false drift and hide a broken pin.
- Decide replacements for images whose upstreams are dead:
pantsel/konga(last push 2020-05-16),containrrr/watchtower(2023-11-11),mongo-express(2024-05-22). - Prune
/var/backups/preupgrade-20260814/once the rollback window closes (never commit it).
Host monitoring rollout — Cockpit + Beszel + Netdata (2026-08-15)
Execution bundle: docs/operations/execution/monitoring-evaluation-bundle.mdx.
Capacity review recorded in docs/operations/baselines/lightweight-monitoring-baseline.mdx.
Repository artifacts are complete and validated; nothing is deployed yet.
- All three tools designed as host installs, none in Swarm, for two reasons:
docker stack deploysilently ignorescap_add/pid/security_opt(so a Swarm Netdata would report healthy while missingapps.plugin— the capability being evaluated), and monitoring must not depend on the thing it monitors (a containerised dashboard is down exactly when Docker breaks). - Artifacts added under
operations/monitoring/{cockpit,netdata,beszel,loadtest},operations/systemd/{beszel,netdata},operations/diagnostics/monitoring-diagnostics.sh, andswarm/configs/traefik/dynamic/host-services.yml. - Config-drift detection built in from the start (
install-monitoring-config.sh --check): a Swarm manifest is the deployed artifact, but a host config is only a copy, and nothing otherwise detects a live edit to/etcdiverging from the repo. -
swarm/stacks/platform/rabbitmq/rabbitmq.ymlgains a loopback-only15672publish — RabbitMQ published no host ports at all, so a host-installed Netdata could not reach its management API. AMQP (5672) stays unpublished; overlay and Traefik paths unchanged. - Hostnames fixed as
cockpit./netdata./beszel.perspective-v.com. - Create DNS A records for all three →
161.97.83.142before Phase A, or ACME issuance fails the same waydata.arnexglobal.comdid. - Phase A — Cockpit EXECUTED 2026-08-15. Account
hassancreated (sudo, no SSH key);rootrefused via/etc/cockpit/disallowed-users; route returns 403 from non-NetBird sources with a valid LE certificate; Traefik reaches the backend (200 from theproxyoverlay); config drift clean. Cost: 12 MB RSS, 7.4 MB disk.cockpit-networkmanagerdeliberately excluded — it pulls innetwork-manager, and this host runssystemd-networkdvia netplan with no out-of-band console (provider VNC disabled after the 2026-04 incident). A second network manager touchingeth0/wt0on a remote-only box is not worth a configuration GUI. Only the Networking config page is lost; metrics come from Netdata/Beszel.cockpit.socketis bound to both172.27.0.1(Traefik) and the NetBird address. The bridge address only exists while Docker runs, so a bridge-only bind would make Cockpit unreachable exactly when Docker is broken.- Owner action: run
passwd hassan. The account is created--disabled-password(statusL) because a generated password must not appear in a transcript. Cockpit login does not work until this is set. - Owner validates from a NetBird client:
hassanlogs in,rootis refused, and the systemd/journal/storage/terminal pages load.
-
cockpit-pcpadded 2026-08-15 (owner installed it via the Cockpit UI for the Metrics history page). Recorded inoperations/monitoring/README.mdso a rebuild reproduces it — a package installed through a UI is exactly the drift the--checkmode exists to catch. It pulls in Performance Co-Pilot:pmcd+pmlogger, ~14 MB RSS, archives in/var/log/pcp(watch growth).pmcdlistens on0.0.0.0:44321— not publicly reachable (INPUT policy DROP, no ACCEPT rule), but the host's-A INPUT -i wt0 -j ACCEPTmeans any NetBird peer can query it unauthenticated. PCP is now a fourth metrics store; reconsider it if Netdata wins. - Netdata auth redesigned before install. It was going to use
admin-auth, which shares its htpasswd withregistry-basic-auth— a file holding five CI credentials distributed to build pipelines. Any CI token would have unlocked host metrics. Added a dedicatedmonitoring-authmiddleware backed by a separate/etc/traefik/registry-auth/monitoring.htpasswd, and changed the Traefik mount from a single file to the directory so further credential sets need no edge redeploy. -
host-services.ymlstaged: only the cockpit router is live. Routers for services that are not yet installed return 502, and the netdata one logged a repeating htpasswd error. Netdata and Beszel blocks are appended during their phases. - Phase B — Beszel EXECUTED 2026-08-15. Hub 8 MB and agent 6 MB (both capped 128M),
bound to
172.27.0.1only; agent detectedeth0/wt0confirming host-level visibility; route 403s from non-NetBird with a valid LE cert. The agent key was derived from the hub's own keypair (ssh-keygen -y -f /var/lib/beszel/beszel_data/id_ed25519), so the UI was not needed to install it.- Owner: create the hub admin at
https://beszel.perspective-v.com/_/, then Add System →172.27.0.1:45876, then set the Discord webhook. - Note: one benign
HUB_URL not setwarning at agent start — Beszel 0.18 tries WebSocket mode before falling back to the SSH listener used here. Not a loop.
- Owner: create the hub admin at
- Phase C — Netdata EXECUTED 2026-08-15. v2.11.0, 175 MB (ceiling 400M), bound to
127.0.0.1+172.27.0.1, routed athttps://netdata.perspective-v.combehindnetbird-only+monitoring-auth. Collectors verified across two restarts: postgres 1977 charts, mysql 45, redis 22, plus docker 72,apps.plugin1480 andcgroups948.monitoring.htpasswdholds ONLY theadminentry copied fromregistry.htpasswd— same password the owner already uses, but the five CI credentials in that file do not unlock host metrics. Traefik's mount changed from a single file to the directory.- RabbitMQ collector dropped — cannot be done safely. Every option was tested:
short-syntax
127.0.0.1:15672:15672is ignored by swarm ingress and installs aDOCKER-INGRESSDNAT that bypasses UFW (briefly exposed publicly, reverted within minutes); long-syntaxhost_ip:is rejected bydocker stack deploy;mode: hostbinds0.0.0.0. The host also cannot reach overlay addresses. Documented in the manifest so nobody retries it.
Netdata gotchas found the hard way (all now fixed and documented)
- Inline comments silently invert settings. Netdata does not strip trailing
#comments —ebpf = no # heavybecomes the literal valueno # heavy, which is notno. Result: ebpf stayed ENABLED andgo.dstayed DISABLED, so every database collector was missing while the config read as correct. All comments moved to their own lines. -
[db] moderenamed to[db] dbin v2. The old name works via a migration shim (the effective config literally printsmigrated from: [db].mode); now set explicitly. - Unquoted DSNs are not parsed. go.d silently refused to register the mysql job
until
dsn:was quoted. Postgres/redis quoted too for consistency. - go.d gives up on first failure. A collector that cannot connect during a restart
never appears and logs nothing.
autodetection_retry: 30added to all three jobs; verified across two consecutive restarts. - Collection rate raised 2s → 5s after measurement: at 2s the docker/cgroups
collectors held
dockerd+containerdat ~24% CPU on a 4-core box. Idle recovered from ~50% to ~78%. - Netdata Cloud is not claimed (
/var/lib/netdata/cloud.dempty). v2.11 exposes no documented switch to hide the dashboard's Sign-in control, so it remains but is inert — it only links to app.netdata.cloud. The agent has no local login of its own; themonitoring-authbasic-auth prompt is the self-hosted login.
Netdata release channel — was nightly, corrected to stable
- The kickstart in
operations/systemd/netdata/install.shdid not pass--stable-channel, so Netdata installed from the edge (nightly) repository as2.11.0-14-nightly. This surfaced as a dashboardTypeErroron the Appearance page — a nightly UI regression, not a configuration problem. Tracking nightly builds on a production host is wrong regardless of that symptom. Fixed: apt source switchedrepos/edge->repos/stable, all 18 netdata packages downgraded to 2.11.0 (includingnetdata-dashboard, which carries the UI), and the installer now passes--stable-channel. Note the downgrade must move every netdata package together —netdata-userconflicts otherwise. Old source file backed up. Re-verified after the change: postgres 1977 / mysql 45 / redis 22 / docker 72 charts,apps.plugin1480, all tuning intact, no auto-updater installed.
Monitoring decision closed + Traefik dashboard replaced (2026-08-15)
-
Beszel retained, Netdata retired. Owner ended the evaluation early. Netdata cost ~175–206 MB vs Beszel's ~21 MB, held
dockerd+containerdat ~24% CPU at 2s collection, and needed materially more care to configure correctly — every misconfiguration failed silently. Full rationale indocs/operations/baselines/lightweight-monitoring-baseline.mdx. Retirement verified: 19 packages purged, apt repo removed,/etc/netdata,/var/lib/netdata,/var/cache/netdata,/var/log/netdatadeleted, the three least-privilege DB monitoring users dropped from PostgreSQL/MySQL/Redis, Traefik route and UFW rule removed, all repo artifacts deleted, no listener on 19999. Final architecture:Cockpit + Portainer + Kener + Beszel(~48 MB total). -
Built-in Traefik dashboard disabled; Traefik Manager deployed at
https://traefik-manager.perspective-v.com(netbird-only, LE cert, 1/1). It replaces a dashboard whose only protection wasadmin-auth— the shared htpasswd that also holds five CI credentials. Traefik Manager has bcrypt cost-12 auth with optional TOTP 2FA.api.dashboard: false,api.insecure: true. The API binds:8080inside the container only (noports:entry), so it is reachable from theproxyoverlay and nowhere else. Accepted trade-off: containers on that overlay can read routing topology, but no secrets — basicAuth appears as a usersFile path and no TLS keys are served. Routingapi@internalthrough Traefik cannot work, as the manager is a container andnetbird-onlywould reject it.- Config ownership is split by file in
/var/lib/traefik/dynamic, because the file provider watches only one directory:security.ymlandhost-services.ymlare installed from Git byoperations/traefik/install-dynamic-config.sh;managed.ymlis the only file mounted into the manager. Verified by write test —/data/traefik.ymlreturnsRead-only file system,/data/dynamic.ymlwrites through. - Scope note: 36 of 39 routers come from swarm
deploy.labelsand cannot be edited in the UI. The manager's editing value is routes for services with no built-in login. - Old dashboard host
traefik.perspective-v.comnow returns 404; its DNS record can be retired.
Two Traefik failure modes discovered — both 404 every site
-
traefik.enable=truewithout a loadbalancer port. The swarm provider fails withservice "edge-edge-traefik" error: port is missing, and that error is not scoped to that service — it aborts the entire provider config, dropping all routers. Hit while removing the dashboard labels (the dummy port went with them). Fixed by removingtraefik.enableentirely, with a warning comment in the manifest. -
An invalid file anywhere in the dynamic directory. A
managed.ymlseeded with empty maps was rejected by Traefik, and one bad file aborts the whole directory load — every@filemiddleware vanished and every router referencing one was disabled. The seed is now comments-only. Both failures present identically: Traefik healthy and listening, every site 404; diagnose viaapi/http/routersshowingdisabled. -
operations/traefik/README.mddocuments the ownership split, the API trade-off and both failure modes.
Traefik Manager data sources wired up (2026-08-15)
Certificates, Logs and Plugins were all reporting "not mounted" / "not configured". All three are now connected, read-only:
-
Plugins —
STATIC_CONFIG_PATH=/data/traefik.yml, pointing at the static config that was already mounted read-only for the editor. -
Certificates —
/var/lib/traefik/letsencrypt/acme.json:/app/acme.json:ro. Accepted risk: 568 KB embedding the private key of every one of 45 certificates plus the ACME account key. Read-only prevents modification but not disclosure, so compromising this container yields every TLS private key on the host. Accepted because it is NetBird-only, has its own authentication and no Docker socket — revisit if any of those three change. -
Logs —
accessLog.filePath: /var/log/traefik/access.logadded to the static config (it was previously stdout-only), mounted read-only into the manager.fields.headers.defaultMode: droppreventsAuthorizationandCookiefrom being written verbatim. Traefik cannot redact query strings, and this estate puts JWTs there (/notifications/hub?access_token=eyJ...), so log lines remain credential-bearing. Capped at 7 compressed days byoperations/traefik/logrotate-traefik, root-only, withcopytruncate— mandatory, because Traefik holds the file open and does not reopen on SIGHUP, so a plain rotate would leave the new file empty. -
TLS tab stays empty by design — no TLS
optionsblock is defined and the defaults are appropriate. Add one tomanaged.ymlonly if a minimum TLS version needs pinning. -
State persistence proven: after the redeploy,
setup_complete: true,must_change_password: falseand zero bootstrap events — the owner's password survived, confirming the/app/configvolume fix. -
Recorded as
CR-2026-08-15-traefik-dashboard-retirement. -
docs/runtime/stacks/edge.mdxcorrected — it still documented the dashboard attraefik.perspective-v.comas NetBird + basic auth. -
Retire the DNS A record for
traefik.perspective-v.com(now 404).
Failed systemd units resolved (2026-08-15)
Surfaced by Cockpit's failed-units view — its first practical value.
-
ssh-egress-policy-maintenance.service— dead since the repo reorganisation. Exit203/EXEC: the unit'sExecStartpointed atruntime/stacks/operations/ssh-egress-policy.sh, but the role-layout rollout moved the script toruntime/stacks/infrastructure/operations/. The unit had been failing on every timer tick, meaning the SSH egress enforcement sweep and blocked-attempt Discord alerting from the April incident had not run since the reorg. Path corrected; the timer now runs and the service exits 0. Old unit saved to/var/backups/preupgrade-20260814/ssh-egress-unit.bak. - Unit brought into
operations/systemd/ssh-egress-policy/with the path templated as@REPO_ROOT@and substituted at install time, plus a--checkmode that fails ifExecStartdoes not point at an existing script. A repo move can no longer silently disable it. Installed and verified: timer active, service exits 0, no drift. -
systemd-networkd-wait-online.service— stale failure from a slow boot on 2026-06-10. The netplan drop-in already scopes it toeth0,eth0isroutable, and the unit only runs at boot. Cleared withreset-failed; no configuration change needed.
Log tampering found on /var/log/wtmp (2026-04-09) — evidence preserved, logging restored
-
systemd-update-utmp.servicewas failing withFailed to write utmp record: Is a directory. Root cause:/var/log/wtmphad been replaced by a directory with moded---------and the immutable attribute set.Birth : 2026-04-09 10:02:09 UTC directory created in place of the file Change: 2026-04-09 10:03:46 UTC chattr +i applied, 97 seconds laterWhy this was escalated rather than quietly repaired:
- No
wtmp,utmporchattrcommand appears in root's shell history for that window; the history there isw,top,ls -a,top,bash. - Later the same day at 12:39 / 12:41 UTC the admin ran
chattr -i -a /root/.ssh/authorized_keysandchattr -i -a /root/.ssh— removing immutable flags that something else had set. The.sshinstance was found and cleaned; thewtmpinstance never was. - It sits ~2h before the SSH login from
86.121.67.247that remains unverified in the Security Incident Follow-up section below. - Only
wtmp(successful logins) was affected.btmp,lastlog,auth.logandsyslogwere intact — the asymmetry an intruder would want. - It was the only immutable file on the entire system.
Effect: no successful login was recorded on this host between 2026-04-09 and 2026-08-15 (~4 months).
lastwas non-functional throughout. - No
-
Evidence preserved at
/var/backups/incident-20260409-wtmp/before any change: fullstat/lsattrrecord, 4,456 journal lines covering 2026-04-09 09:30–13:00 UTC, and the root shell history for the window. -
Logging restored: immutable flag cleared, directory removed,
/var/log/wtmprecreated as a regular file0664 root:utmp.systemd-update-utmpis active andlastworks. -
Owner confirmed both keys in
/root/.ssh/authorized_keysare theirs (2026-08-15). That file was not modified. -
Work this against
docs/operations/governance/host-integrity-checklist.mdxand close the open86.121.67.247verification item — a four-month gap in login history is a material finding for that review.
Measured cost (all three tools, vs the 2026-08-15 baseline)
| Memory | |
|---|---|
| Netdata | 175 MB |
| Beszel hub | 15 MB |
| Beszel agent | 6 MB |
| Cockpit (socket-activated) | ~12 MB |
PCP (from cockpit-pcp) | ~15 MB |
| Total | ~220 MB, under the ≤300 MB budget |
Available RAM 3263 MB vs 2904 MB baseline (higher, from page-cache reclaim). Load is
elevated versus the 0.30 baseline and should be re-measured once the install session has
quiesced; CPU was ~78% idle with wa=0 at the end of the work.
- Phase C — Netdata + four least-privilege DB monitoring users; verify the config
actually applied by diffing against
curl -s http://localhost:19999/netdata.conf, since several[db]keys were renamed between v1 and v2 and unknown keys are ignored silently. - Phase D/E — measure after each phase; evaluate 1–2 weeks, using
operations/monitoring/loadtest/synthetic-load.shif no real incident occurs. - Phase F — retire the loser and record the outcome in the baseline doc.
Accepted limitation — container-to-host access is not NetBird-gated
Owner decision 2026-08-15: keep Traefik routing and accept this.
Traefik proxies from a container to a host port, and at the network layer the host cannot
distinguish it from the other 30 containers on docker_gwbridge. Any container on an overlay network can reach
172.27.0.1:{9090,19999,8090} directly, bypassing the netbird-only middleware. Containers
on the default bridge (docker0) cannot — the UFW rules are scoped to docker_gwbridge,
which narrows the exposure to swarm-attached containers rather than all of them. No UFW or
ipAllowList rule closes this — bridge addresses are dynamic and indistinguishable.
External access remains properly gated (UFW default-deny with interface-scoped additions
only, services bound to 172.27.0.1 never 0.0.0.0, netbird-only on every router).
Cockpit and Beszel require their own login, so a container reaches only a login page.
Netdata has no built-in authentication, so admin-auth was layered onto its router —
that protects the browser path but not direct container access. Closing it fully would mean
dropping Traefik and binding to the NetBird interface, losing TLS and the hostnames.
Committed production secrets (2026-08-15) — OPEN, owner deferred
Owner decision on 2026-08-15: note and defer; no rotation or history rewrite performed.
-
40 real production secret values are committed to Git, across 29 tracked
.envfiles underruntime/environments/vps/**, pushed togithub.com/HassanTaj/contabo-server-setup(origin/main) since 2026-05-03 across 15 commits. The repository is private — verified anonymously via the GitHub API (HTTP 404) — so the values are not world-readable, but they are readable by anyone with repository access and by any leaked credential with that access. Highest-risk values, in rough order:ZITADEL_MASTERKEY(encrypts the whole identity provider),KONG_JWT_HMAC_SECRET(permits minting valid JWTs, i.e. auth bypass onapi.perspective-v.com), all four database root passwords (MYSQL_ROOT_PASSWORD,POSTGRES_PASSWORD,MONGO_INITDB_ROOT_PASSWORD,MSSQL_SA_PASSWORD),REDIS_PASSWORD,RABBITMQ_DEFAULT_PASS, the RustFS root/access/secret keys, the VaultwardenADMIN_TOKEN, the registry admin bootstrap password, SMTP passwords and the Discord webhook. AGENTS.md §16 requires committed secrets be treated as compromised; rotation and history remediation remain outstanding. -
Root cause —
.gitignorenegation. Line 436!runtime/environments/**un-ignores the whole tree, and the blanket guardruntime/environments/**/.envon line 439 is commented out. Onlybackups/.envwas ever explicitly re-ignored, so every other live.envis tracked by default. -
The documented verification method is misleading here. AGENTS.md §3 says to verify ignore behaviour before populating live env files, but
git check-ignore -von these paths exits 0 while printing the negation rule, which reads as "ignored". Onlygit status --porcelain <path>(showing??) reveals the file is tracked. Usegit status, notcheck-ignore, for this check. -
New monitoring env paths (
.../monitoring/netdata/.env,.../monitoring/beszel/.env) explicitly re-ignored in.gitignoreand verified withgit status, so the monitoring rollout adds nothing to the exposure. A warning comment now documents the misleadingcheck-ignorebehaviour inline.
Backup coverage gap (2026-08-14)
-
scm_gitea_datahas no backup.operations/backups/db-backup.shcapturesgitea_dbincidentally (it dumps every non-template Postgres DB), but the volume holding the actual package blobs — Docker layers,.nupkgfiles, npm tarballs — is covered by nothing. Losing it means re-pushing every image and package by hand. Highest-value open data-loss risk.
Backup implementation finding (2026-09-08)
- Logical Tier-1 local retention does not distinguish successful sets from failed staging.
operations/backups/db-backup.shprunes everytier1-*directory once the count exceedsDB_KEEP_LOCAL; unlike the PostgreSQL physical helper, it has no success marker. A later successful run or explicitprunecan therefore delete preserved failed staging, contrary to the requirement that failed staging remain available for investigation.
Platform evaluation decisions
- Evaluated OneDev as a Gitea replacement (2026-08-14): rejected, keep Gitea. Gitea here is a package registry (git/actions/SSH disabled), and OneDev serves the same Docker/npm/NuGet formats, so the swap is lateral on capability while costing a JVM runtime with a 2 GB documented minimum against Gitea's ~112 MiB on a 4-core/8 GB node running 39 services. It would also re-point 47 files referencing
gitea.perspective-v.comsix weeks after the last cutover. Self-hosted CI on this node was rejected for the same resource reasons; hosted runners pushing to the VPS registry remain the model.
Swarm role-layout rollout
- Split active Swarm stacks into one YAML fragment per service and reorganize manifests/launchers into the six role directories.
- Relocate host-wide operational tooling into
operations/and validate backup tests plus systemd installer dry-runs. - Reinstall both backup systemd units from
operations/systemd/, verify their schedules and a successful persistent Tier-1 run, then remove the retired compatibility paths. - Migrate Redis/RabbitMQ, Watchtower, RustFS, Kong, NetBird, Panels, and Edge identities one workload at a time with preserved state and health checks.
- Make every manifest's
# Swarm stack:header authoritative and migrate production todatabase,platform,edge,panel,service, the three dedicatedservice-*stacks, and six activewebsites-*stacks. - Update database and WordPress backup defaults to
database_*; both production dry-runs pass and both systemd timers remain enabled/active. - Preserve the prior service image and replica sets, named external volumes,
and secrets; validate all expected routes plus a panel
0 → 1 → 0cycle.
Homelab Pi-hole administration
- TCP 8053 is present in NetBird policy
contabo-homelab-server-link; no user or peer group was broadened. - Add private NetBird DNS record
pihole.home.perspective-v.compointing to Contabo's NetBird IP100.83.72.162. - Install the repo-managed
homelab-piholeTraefik router and backend, validate trusted TLS, NetBird access, public denial, and configuration drift. The rollback archive is/var/backups/traefik/dynamic-pre-pihole-host-rewrite2-20260914T231325Z.tar.gz.
Home Assistant route
- Install the repo-managed
assistant.home.perspective-v.comrouter and backend at100.83.117.37:8123, protected bynetbird-onlyand security headers. Live dynamic-config drift check is clean. Rollback archive:/var/backups/traefik/dynamic-pre-assistant-rename-20260929T204500Z.tar.gz. - Point private NetBird DNS and the public Namecheap A record at Contabo
(
100.83.72.162privately;161.97.83.142publicly); preserve existing wildcard and sibling records. Trusted TLS is issued and public denial is verified. - Keep the old Traefik Yui host rule retired; no service router matches it.
- After the user creates the Home Assistant owner account, confirm the
imported HTTP settings in Settings > System > Network: Trust X-Forwarded-For
and trusted proxy
100.83.72.162. Then verify the onboarding page through the canonical HTTPS URL. - Remove the old private NetBird
yui.homerecord after the user confirmed the irreversible deletion. Theassistant.homerecord remains, and public wildcard and sibling records are unchanged.
Homelab LAN certificate export
- Add the restricted exporter for
home,next,images,media,pihole, andportainercertificates. - Keep the exporter under review during the first local-edge renewal cycle; if the homelab sync timer reports a failure, retain the last known-good local files and investigate before changing the Contabo ACME design.
- Publish the Contabo-side homelab edge guide, DNS boundary, certificate handoff, and operational validation notes in the Mintlify documentation.