Accepted Suggestions
Migrated from the repository documentation set.
- Keep apex (@) on GitHub Pages.
- Keep www and learn on GitHub Pages.
- Use wildcard DNS to route app subdomains to VPS by default.
- Enable IPv6 now.
- Use GitHub Actions as the first CI/CD target.
- Use hybrid CI/CD: GitHub Actions for GitHub repos and Azure Pipelines for Azure DevOps repos.
- Keep reusable CI templates under runtime/ci for both GitHub Actions and Azure Pipelines.
- Use service-specific CI files for split repositories: build-and-push-identity, build-and-push-graph, and build-and-push-console.
- Keep build-and-push-single-image as the fallback starter template for new services.
- Use
docker login --password-stdinin Azure Pipelines image workflows and fail fast when registry secrets are missing. - Add dbskc-specific CI templates under runtime/ci for both GitHub Actions and Azure Pipelines.
- Add nishatcolony-specific CI templates under runtime/ci for both GitHub Actions and Azure Pipelines.
- Trigger dbskc CI image builds from branch
deploy/dbskcand publish toregistry.perspective-v.com/dbskc-web. - Trigger nishatcolony CI image builds from branch
deploy/nishatcolonyand publish toregistry.perspective-v.com/nishatcolony-web. - Trigger identity CI image builds from branch
deploy/identityand publish toregistry.perspective-v.com/identity. - Trigger graph CI image builds from branch
deploy/graphand publish toregistry.perspective-v.com/graph. - Trigger console CI image builds from branch
deploy/consoleand publish toregistry.perspective-v.com/console. - Deploy dbskc website from
runtime/stacks/services/dbskc.com/dbskc.comusingdbskc.comand imageregistry.perspective-v.com/dbskc-web:latest. - Deploy nishatcolony service from
runtime/stacks/services/nishatcolony.pk/nishatcolony.pkusing imageregistry.perspective-v.com/nishatcolony-web:latest. - Finalize nishatcolony production host as
nishatcolony.pkfor VPS route mapping. - Use
APP_INTERNAL_PORT=80for nishatcolony nginx-based image routing behind Traefik. - Keep dbskc rollout model as bootstrap manual deploy plus fast-scope Watchtower updates for new
latestdigests. - Keep nishatcolony rollout model as bootstrap manual deploy plus fast-scope Watchtower updates for new
latestdigests. - Use private registry login with CI secrets (REGISTRY_USERNAME and REGISTRY_PASSWORD) in all image pipelines.
- Push moving latest tags plus immutable auto semantic-style version tags
v1.0.<run>from CI. - Execute migration in phases, starting with api.perspective-v.com, identity.perspective-v.com, and graph.perspective-v.com.
- Use Let's Encrypt in Traefik with automatic certificate renewal.
- Use TLS challenge on port 443 for certificate issuance.
- Use notifications at perspective-v.com as Let's Encrypt contact email.
- Install Portainer for host operations.
- Retire Uptime Kuma completely from edge runtime (service, container, volume, and image).
- Promote Kener from pilot to active monitoring endpoint at kener.perspective-v.com.
- Keep Kener publicly reachable and rely on Kener built-in authentication.
- Keep Kener NetBird middleware configuration commented in
runtime/stacks/infrastructure/edge/docker-compose.kener.ymlfor rollback fallback. - Enable Kener SMTP settings from edge env values for invitation and password-recovery workflows.
- Split edge base stack artifacts into dedicated compose files:
runtime/stacks/infrastructure/edge/docker-compose.traefik.ymlandruntime/stacks/infrastructure/edge/docker-compose.portainer.yml. - Split edge VPS env artifacts into per-service folders under
runtime/environments/vps/infrastructure/edge/{traefik,portainer,kener}. - Split edge dev env templates into per-service folders under
runtime/environments/dev/infrastructure/edge/{traefik,portainer,kener}using{service}.dev.envnaming. - Keep Kener runtime env keys in
runtime/environments/vps/infrastructure/edge/kener/.env(live runtime) andruntime/environments/vps/infrastructure/edge/kener/kener.env(tracked placeholder template). - Execute edge rollout with per-service compose files (
docker-compose.traefik.yml,docker-compose.portainer.yml,docker-compose.kener.yml). - Keep Portainer on Business Edition image (
portainer/portainer-ee:2.39.1) because the existing Portainer data volume is BE-format and CE downgrades fail. - Continue manual monitor recreation in Kener (critical monitors first).
- Skip Dozzle for now.
- Add Fail2ban for host intrusion protection.
- Use Watchtower in label-based mode instead of global auto-updates.
- Schedule Watchtower for weekly Sunday 03:00 updates.
- Add a second scoped Watchtower instance with 1-hour polling for explicit fast-update stateless services.
- Run Watchtower from a dedicated operations stack (
runtime/stacks/infrastructure/operations) instead of coupling it to Platform or Registry stacks. - For stateless application service stacks under
runtime/stacks/services, enable Watchtower labels and use scope labels to control cadence (nonebaseline,fastexplicit immediate-target). - Configure Redis with password, 512 MB cap, and persistence while keeping production container traffic on internal host
redis. - Publish Redis host port for developer access only through NetBird-scoped UFW rules.
- Use
redis.perspective-v.comas the developer-friendly Redis hostname. - Treat Platform Redis as cache/session/ephemeral only; critical source-of-truth data is not allowed on this shared Redis instance.
- If Redis durable critical state is required, move it to a dedicated stateful Redis stack with backup/restore controls.
- Configure RabbitMQ with persistence and expose management UI at rmq.perspective-v.com.
- Keep Platform stack scope as Redis + RabbitMQ for this phase.
- Use shared stack deployment order:
runtime/stacks/infrastructure/platform->runtime/stacks/infrastructure/operations->runtime/stacks/infrastructure/registry->runtime/stacks/infrastructure/feeds. - Track already-implemented.md as server-only completed items.
- Use phased checklist format in next-steps.md for execution tracking.
- Add a current operator guide set under
docs/runtimefor runtime rebuild, per-stack rollout, environments, CI, services, and wrapper-script usage. - Retire the legacy
docs/folder after its content is migrated into the Mintlify project underdocs/. - Add cross-platform
.shand.batlauncher scripts underruntime/scriptsfor each active infrastructure and service compose target, defaulting to live VPS.envwith tracked-placeholder fallback anddevtemplate support. - Protect developer DB access with self-hosted NetBird and close public DB ports after validation.
- Retire legacy feed bridge service/routes from feeds runtime after canonical npm and NuGet validation gates close.
- Use canonical post-retirement feed endpoints:
https://nuget.perspective-v.com(BaGet) andhttps://npm.perspective-v.com(Verdaccio). - Enforce Verdaccio no-signup policy with
auth.htpasswd.max_users=-1. - Enforce Verdaccio package policy: anonymous read/install (
access: $all) and authenticated publish/unpublish ($authenticated). - Restrict npm publish credentials to manually provisioned CI-only service account(s).
- Standardize npm CI publish secret naming as
NPM_FEED_API_KEY, mapped to npm auth token usage. - Keep feed env artifacts explicit for npm CI token handoff by defining
NPM_FEED_API_KEYinruntime/environments/dev/infrastructure/feeds/feeds.dev.env,runtime/environments/vps/infrastructure/feeds/feeds.env, and liveruntime/environments/vps/infrastructure/feeds/.env. - Keep BaGet publish model on API key with public restore/download behavior.
- For live feed runtime env values containing dollar signs, escape as
$$inruntime/environments/vps/infrastructure/feeds/.envso Docker Compose interpolation does not corrupt secrets. - Keep feed persistence on local Docker volumes for this phase.
- For approved feed bridge retirement execution on VPS, recreate feeds with canonical compose using
docker compose up -d --remove-orphansso legacy containers are removed from runtime. - For approved legacy feed artifact cleanup on VPS, capture rollback snapshot/checksum under
tmp/backups/feeds-proget-retirement-<timestamp>before deleting legacy volume/image remnants. - Use netbird.perspective-v.com as the NetBird host.
- Use NetBird embedded IdP quickstart mode initially.
- For NetBird embedded IdP quickstart, keep
AUTH_CLIENT_SECRETempty in dashboard env files (runtime/environments/vps/infrastructure/netbird/.envandruntime/environments/vps/infrastructure/netbird/netbird.env) to prevent dashboard loginUnauthenticatederrors after updates. - Use runtime/stacks/infrastructure/netbird/TABBY-EXECUTION-SEQUENCE.md as the primary operator command sequence.
- Use dedicated production DB compose files with env-based secrets and localhost-first bindings.
- Keep DB server containers as long-lived stateful services with controlled monthly maintenance updates.
- Run DB engines with NetBird-ready publish binds and enforce access using interface-scoped UFW rules on NetBird interface.
- Keep production DB compose/env files under runtime/stacks/infrastructure/databases (moved out of local setup).
- Use prebuilt registry images for production services deployment.
- Include api-gateway, identity, graph, and console in current production service scope.
- Exclude pvwebsite from current production compose rollout.
- Keep production service artifacts under
runtime/stacks/services/<domain>(current:runtime/stacks/services/perspective-v.com). - Mirror Perspective-V service env layout by service under
runtime/environments/{dev,vps}/services/perspective-v.com/<service-domain>/with tracked placeholders as<service-domain>.env(or<service-domain>.dev.env) and live VPS runtime values in service-local.envfiles. - Roll out remaining Perspective-V services (
identity,graph,console) together fromruntime/stacks/services/perspective-v.comusing one compose command with the VPS service-local env file. - Keep console startup dependency scoped to
identityandgraphso these three services can bootstrap without forcing api-gateway startup. - Keep Kong plus service upstream connectivity on the existing external
proxynetwork; do not introduce a dedicated gateway-only network. - Keep
identityandgraphas Kong upstream backends (internal service targets), and disable direct Traefik exposure for those services. - Keep
consoleas the public Angular app directly exposed by Traefik onconsole.perspective-v.com(not routed through Kong). - Use service hosts: api.perspective-v.com, identity.perspective-v.com, graph.perspective-v.com, and console.perspective-v.com.
- Use initial image placeholders: api-gateway:latest, identity:latest, graph:latest, console:latest.
- Use compose project names for grouping: Services="Perspective-V" and Databases="Databases".
- Use compose project names for shared stacks: Platform="Platform", Operations="Operations", Feeds="Feeds", NetBird="NetBird".
- Create per-domain services starter folders under runtime/stacks/services for arnexglobal, nishatcolony, and dbskc.
- Keep this naming style for future domain groupings.
- Reset deployment baseline to a clean VPS (Ubuntu + Docker + Docker Compose + zsh) and rebuild from scratch.
- Use dedicated edge Traefik from runtime/stacks/infrastructure/edge as day-one reverse proxy baseline.
- Integrate NetBird using Existing Traefik mode (option 1) from day one.
- Keep NetBird proxy service disabled for now.
- Keep Traefik dashboard NetBird-only behind Traefik source-range allowlist and basic auth.
- Keep Portainer NetBird-only behind Traefik source-range allowlist and use built-in panel authentication (no shared Traefik admin-auth).
- Keep Kener publicly reachable with built-in panel authentication and keep NetBird middleware commented for fallback in overlay compose.
- Keep 80 and 443 publicly available for internet-facing apps and APIs routed by Traefik.
- Repository database-backup scope includes MSSQL, PostgreSQL, MySQL, and MongoDB. For
the current Google Drive activation, keep MSSQL excluded while its production service
intentionally remains
0/0; restore its default backup coverage when it is re-enabled. - Keep Watchtower as the only automatic updater for this rollout phase.
- Add WUD as alerting-only update monitoring in runtime/stacks/infrastructure/operations (no automatic update actions).
- Expose WUD at wud.perspective-v.com behind Traefik with NetBird-only access policy.
- Use Discord webhook notifications as the initial WUD alert channel.
- Register private registry.perspective-v.com in WUD as a custom registry with authenticated access.
- Run WUD in discovery mode (
WUD_WATCHER_LOCAL_WATCHBYDEFAULT=true) so all running containers are inventoried by default; use explicit label opt-outs only when needed. - Use WUD manual full-scan endpoint (
POST /api/containers/watch) after watcher-mode changes to refresh inventory immediately. - Add per-service
wud.tag.includeguardrails (for example Redis alpine, RabbitMQ management-alpine, MySQL oraclelinux channel, Verdaccio numeric stable tags) to reduce noisy major-alert channels. - Use a dedicated
wud-monitorregistry basic-auth account for WUD private-registry checks while keeping admin credentials separate. - Execute approved major upgrades in controlled, one-service-at-a-time maintenance windows with pre-change volume backups and explicit rollback checkpoints.
- Adopt upgraded stable targets for this phase: Verdaccio
6, Redis8-alpine, RabbitMQ4.2-management-alpine, and MySQL9.6-oraclelinux9. - If MySQL persisted-volume credential drift is detected during upgrade windows, run controlled recovery-mode alignment so live
root/app credentials match VPS.envsource-of-truth values. - Use Portainer stack update actions for quick manual one-click redeploys when needed.
- Add a private Docker Registry (distribution) plus Registry Admin under runtime/stacks/infrastructure/registry.
- Keep Registry API public behind Traefik TLS with Traefik HTTP basic auth to support Docker login from CI and developer machines.
- Apply Traefik rate limiting on the public Registry API route.
- Keep Registry Admin UI NetBird-only behind Traefik allowlist without Traefik basic auth middleware.
- Remove public token-realm dependency (
/api/v1/registry/auth) from the registry API auth path. - Retire
registry-auth-proxyfrom runtime model. - Retire legacy docker-registry-ui and keep registry-admin as the sole registry UI endpoint.
- Use one shared host htpasswd file (
/var/lib/traefik/registry-auth/registry.htpasswd) as the single source for registry API credentials. - Keep Traefik registry auth on file-provider middleware (
registry-basic-auth@file) so credential updates apply without registry stack recreate. - Point Registry Admin htpasswd path to the same shared host file so users created in Registry Admin can immediately use
docker login. - If a newly created or updated Registry Admin user still gets
401 Unauthorized, refresh Traefik auth state by restarting only the Traefik container (Portainer restart is acceptable; no registry stack recreate). - Add Redis Insight to runtime/stacks/infrastructure/platform at redis-insight.perspective-v.com with NetBird-only Traefik route and no shared Traefik basic auth.
- Keep all runbook/checklist documentation under docs only; do not maintain markdown docs under prod.
- Retire MinIO from runtime because required self-hosted admin workflows are no longer available in the current community-mode behavior.
- Adopt RustFS as the sole object-storage baseline at
runtime/stacks/infrastructure/object-storage/rustfs/docker-compose.rustfs.yml. - Keep RustFS NetBird-only on
rustfs.perspective-v.com(console) andrustfs-api.perspective-v.com(API) with RustFS built-in authentication. - Enforce explicit non-default RustFS API credentials through
RUSTFS_ACCESS_KEYandRUSTFS_SECRET_KEY. - Keep RustFS persistence on host bind mounts
/var/lib/rustfs/dataand/var/lib/rustfs/logswith UID/GID10001:10001. - Keep object-storage artifacts only under RustFS paths:
runtime/stacks/infrastructure/object-storage/rustfsandruntime/environments/{dev,vps}/infrastructure/object-storage/rustfs. - Remove residual MinIO host/env references from runtime app env and service-stack artifacts; keep RustFS endpoint standard only.
- If first certificate issuance for RustFS hosts fails under strict SNI, temporarily set
sniStrict=false, issue certs, then restoresniStrict=true. - Use developer-friendly DB hostnames over NetBird: mssql/pgsql/mysql/mongo under perspective-v.com.
- Keep MSSQL on
MSSQL_PID=Developerfor the current VPS environment by explicit decision, and reassess edition only if compliance/commercial constraints change. - Expose pgAdmin, phpMyAdmin, and mongo-express via Traefik with NetBird-only plus security headers and built-in UI authentication.
- For pgAdmin route hardening, keep NetBird-only access and security headers but allow SAMEORIGIN frame behavior so Query Tool panels load correctly.
- Keep Traefik + NetBird as the network access gate for DB UI routes and keep Mongo Express built-in auth enabled.
- Enforce strict go-live gates for DB/object-storage rollout: unique strong secrets, pinned stateful images, healthchecks, and NetBird-vs-non-NetBird access validation before release.
- Add explicit compose down commands in databases README to stop and remove all DB/UI containers without deleting volumes by default.
- Use
{service}.envfor production runtime env files underruntime/environments/vps, and{service}.dev.env/{service}.test.envfor non-production templates where needed. - Use
docker-compose.{service}.ymlnaming across runtime/stacks (no.prodcompose suffix). - Remove
.exampleenv template style and keep tracked non-production templates under runtime/environments/dev (*.dev.env). - For database templates, keep per-engine folders under runtime/environments/dev/infrastructure/databases with
<engine>.dev.envfiles. - Track named environment files under runtime/environments (
{service}.env,*.dev.env,*.test.env) and keep plain.envfiles ignored. - For live VPS deployments, execute compose with service-local
runtime/environments/vps/**/.envfiles; keep trackedruntime/environments/vps/**/{service}.envfiles as placeholders only. - For edge stack operations, use service-local env files under
runtime/environments/vps/infrastructure/edge/{traefik,portainer,kener}/.env; keep tracked placeholders as{service}.envin the same folders. - For persisted PostgreSQL volumes, after recreate with
.env, align the livepostgresrole password to.envwhen drift is detected (recreate alone does not backfill existing role credentials). - Treat prod/docker stack artifacts as source-of-truth snapshots for active VPS runtime and sync them into runtime/stacks + runtime/environments/vps before retiring prod.
- Keep Traefik ACME runtime file host-managed at /var/lib/traefik/letsencrypt/acme.json with file mode 600.
- Retire transitional prod folder after syncing source-of-truth VPS stacks and env files into runtime.
- Keep deployed service connection strings on internal Docker DB names (mssql/postgres/mysql/mongodb).
- Keep DB DNS records public and enforce access at NetBird/UFW and Traefik middleware; do not add a separate NetBird-only DNS zone in this phase.
- Require NetBird dashboard peer-readiness checks and VPS interface/IP detection before running DB firewall cutover.
- Continue immediately with Phase 4 server execution (env prep, DB engines, DB UIs, firewall cutover, and validation sequence).
- Run a consistency sweep so Setup Guide and NetBird runbooks use the same NetBird-side DB configuration language as runtime/stacks/infrastructure/databases docs.
- Point domain service image examples to private registry paths (registry.perspective-v.com/*).
- Use one private registry repository per service name (for example, registry.perspective-v.com/api-gateway).
- Keep Registry Admin pointed to internal Docker registry host (
http://registry:5000) in compose runtime. - Enforce Docker schema v2 media types for CI image pushes (
oci-mediatypes=false) to keep Registry Admin sync/catalog compatibility. - Start registry major-version migration by pinning registry runtime image to
registry:3in stack artifacts. - Keep registry manifest deletion enabled for cleanup workflows (
REGISTRY_STORAGE_DELETE_ENABLED=true). - Add an operator-safe registry cleanup helper workflow (dry-run by default) for tag delete and repository purge operations.
- Treat full repository cleanup as manifest deletion workflow plus garbage collection, not as a single API hard-delete call.
- Run registry cleanup on a monthly cadence with dry-run preview before apply.
- Keep repository cleanup approval model as single-owner in this rollout phase.
- Use a tracked registry cleanup policy file to protect critical repositories/tags and drive monthly include/exclude targeting.
- Require explicit monthly include targets before running cleanup apply mode.
- Add a reusable registry staging rehearsal helper script to standardize v3 pre-cutover validation checks.
- Track a registry dev env template under
runtime/environments/dev/infrastructure/registry/registry.dev.envfor non-production rehearsal. - Allow owner-approved direct VPS cutover for registry v3 without staging rehearsal when pre-cutover volume backups and rollback metadata are captured first.
- Attach domain service templates to external object-storage network for RustFS-ready connectivity.
- Deploy MongoDB from runtime/stacks/infrastructure/databases as part of the active database rollout scope.
- Pin edge Traefik to v3.6 and set DOCKER_API_VERSION=1.40 for Docker Engine 29 API compatibility.
- Define Let's Encrypt cert resolver in static Traefik config (runtime/stacks/infrastructure/edge/traefik/traefik.yml).
- Retire Uptime Kuma image and compose service from active edge runtime artifacts.
- Pin Kener pilot image to explicit version tag
rajnandan1/kener:v4.0.16in repository compose overlay. - Track active NetBird runtime artifacts in runtime/stacks/infrastructure/netbird (docker-compose.yml and config.yaml) and keep active VPS env values in
runtime/environments/vps/infrastructure/netbird/netbird.env. - Keep NetBird dashboard image compatibility by loading live VPS env via compose (
runtime/stacks/infrastructure/netbird/docker-compose.ymldashboardenv_file->../../environments/vps/infrastructure/netbird/.env) before recreate/pull operations. - Complete transitional top-level rename to docs and retire the temporary prod mirror after runtime sync.
- Move reusable CI templates under runtime/ci and keep docs focused on markdown documentation.
- Keep passwords/secrets in tracked runtime/environments named env files as example placeholders only, and set real secret values only in local server-side copies.
- Add gateway stack artifacts under
runtime/stacks/infrastructure/gatewayusingdocker-compose.kong.yml. - Use Kong + Konga as the API gateway management path for Ocelot replacement in this phase.
- Reuse existing PostgreSQL stack for Kong metadata (
kong_db+kong_user) onpostgres-network; do not deploy a dedicated Kong DB container. - Keep Kong Admin API internal-only on Docker network (no host port exposure).
- Expose Kong proxy through Traefik host
api.perspective-v.com. - Expose Konga through Traefik host
konga.perspective-v.comwithnetbird-only+security-headersmiddleware and Konga built-in login. - Keep dev gateway env template at
runtime/environments/dev/infrastructure/gateway/kong.dev.envand add VPS gateway env artifacts underruntime/environments/vps/infrastructure/gateway(.envlive runtime +kong.envtracked placeholder). - Keep Ocelot bearer parity on Kong by applying JWT only to protected routes and keeping public routes without JWT plugin.
- Use Kong JWT settings with key claim
iss, exp verification, and HS256 credential key mapped to issuer URL. - Use internal Kong upstream targets for migrated identity and graph routes:
identity:62258andgraph:5138. - Approximate Ocelot period limits in Kong using second-based rate limiting (
10/10s -> second=1,10/5s -> second=2). - Keep global CORS allowed headers including
AppCode,appcode, andAPPCODEfor client compatibility. - Keep all migrated Kong entities tagged as
ocelot-migrationfor Konga audit visibility.
Security Incident Decisions (2026-04-11)
- Use full host plus Docker forensics when diagnosing outbound SSH/port 22 anomalies.
- Include immediate emergency triage plus longer-term hardening for SSH anomaly incidents.
- Produce support-ready evidence for provider communication (process attribution, destination IPs, timestamps, mitigation).
- Allow temporary outbound SSH containment (
deny out 22/tcp) when unknown high-rate traffic is detected. - Rotate root password and regenerate trusted SSH key material (
authorized_keys) after host-level compromise indicators. - Disable provider VNC/remote console access by default and enable only for time-boxed break-glass use.
- Reset VS Code SSH and Tabby SSH access credentials/keys after SSH security incidents.
- Finalize outbound SSH policy as permanent deny-by-default or explicit destination allowlist with documented exceptions.
- Perform same-day post-containment revalidation sweeps (live SSH egress snapshot, process attribution, scheduler/persistence check, and Uptime Kuma monitor audit) and archive the outputs in the same incident evidence bundle.
- Set permanent outbound SSH policy to deny-by-default on both IPv4 and IPv6.
- Use logging-enabled outbound deny rules for TCP/22 so blocked-attempt alerts have reliable source events.
- Keep GitHub VPS workflows on SSH over 443 (
Host github.com -> ssh.github.com:443) instead of broad outbound TCP/22 allow rules. - Use script-managed temporary outbound SSH exceptions with required audit fields and one-hour default maximum duration.
- Restrict outbound SSH exception approvals to single-owner authorization.
- Send immediate Discord alerts when outbound SSH attempts are blocked by policy.
- Format SSH egress Discord notifications using Markdown with an underlined heading plus ordered-list structure for readability (Discord webhook content, not HTML).
- Keep this runbook constraint during incident handling: do not run docker compose commands without explicit approval.
- Store each incident package in-repo under
Incidents/<date-time>-<incident-name>/with both markdown reports and raw evidence files. - Enforce Watchtower fast-scope labeling for all service stacks under
runtime/stacks/servicesby setting bothcom.centurylinklabs.watchtower.enable=trueandcom.centurylinklabs.watchtower.scope=fast. - Enforce public Traefik service labeling baseline in
runtime/stacks/services:traefik.enable=true,traefik.docker.network=proxy,websecureentrypoint, TLS enabled, Let's Encrypt certresolver,security-headers@file, and explicit load balancer server port. - Keep healthchecks enabled on all service stack containers (
runtime/stacks/services) using internal container ports from env defaults where applicable. - Use shell-compatible healthchecks for Alpine/nginx-based images (for example
wgetprobes) and avoidbash-dependent checks on images that ship only/bin/sh. - Use compose project name
perspective-vfor all service compose files underruntime/stacks/services.
Governance and DR Decisions (2026-04-15)
- Start implementation with governance and runbooks first before broader operational rollout tasks.
- Keep production high-risk change approvals under single-owner authority, with explicit approval or direct manual execution by owner.
- Use RustFS bucket storage on the same VPS as the temporary backup destination baseline.
- Finalize critical-service recovery targets at RTO 4 hours and RPO 1 hour.
- Add and maintain runbooks under docs for disaster recovery, incident response, host integrity checks, and production change control.
- Add and maintain Tier-1 database backup helper automation at
operations/backups/db-backup.shwith checksum manifest and strict rolling keep-last-backups retention. - Use weekly Tier-1 backup scheduling at Sunday 03:00 UTC and keep its templates
and installer under
operations/systemd/tier1-db-backup/. - After owner approval, apply the updated weekly keep-last-3 Tier-1 timer/service policy on VPS and capture refreshed evidence output.
- Add and maintain operational evidence templates for production change records, restore validation, and incident closure under
docs/operations/templates/. - Add and maintain a NetBird access validation matrix at
docs/operations/governance/netbird-access-validation-matrix.mdxfor source-allow and source-deny proof capture. - Prepare and maintain an approval-ready execution bundle for first backup timer rollout at
docs/operations/execution/phase5-backup-rollout-execution-bundle.mdx. - Prepare and maintain an instantiated production change record draft under
docs/operations/change-records/before server execution. - Add and maintain one-command first-run evidence capture helper at
operations/diagnostics/collect-tier1-backup-evidence.sh. - Keep the monitoring baseline lightweight on this VPS: Kener + WUD + Watchtower/Watchtower-fast only, with no heavy observability stack in this phase.
- Schedule controlled monthly DB maintenance on the first Saturday at 02:00 UTC with owner-approved execution windows.
- Use separately authorized published OAuth Desktop applications for the central and Gorsi accounts, root-only writable rclone configuration, separate client crypt keys, and non-secret per-site destination mappings for multi-Drive backup delivery.
- Use a hybrid PostgreSQL recovery model: weekly globals and collision-safe custom dumps
for every connectable non-template database, plus a monthly online
pg_basebackupsnapshot with included WAL. Never tar the live PostgreSQL volume. - Keep sliding retention at four WordPress archives per site, twelve weekly database sets remotely/three locally, and four PostgreSQL physical snapshots remotely/two locally; prune only after upload verification.
- Limit WordPress backup scope to
dbskc.com,gorsistudio.com,store.gorsistudio.com, andwcblahore.pk; exclude both Nishat variants. - Execute the encrypted Google Drive rollout in separately approved production phases. OAuth/crypt configuration, two-Drive canaries, and WordPress/logical PostgreSQL scratch restores preceded the approved full workflows, physical backup/restore, offsite flags, and final three-timer activation completed on 2026-08-14. Never transmit backup secrets through chat or Git.
- Keep Kong compatibility routes for client URL stability:
/identity/docs/*,/identity/swagger/v1/swagger.json,/swagger/v1/swagger.json, and/resumeas a UI passthrough route so existing callers avoid route-miss regressions. - Keep Kong graph routing aligned with .NET GraphQL middleware mapping: UI routes (
/resume,/gopher/*) must remain UI passthrough paths, while API routes remain under/graph/*(including/graph/gopher/{credential,serviceprovider,serviceprovideremail,platform}). - Keep Kong global CORS allow-headers compatible with Apollo GraphQL clients from
https://console.perspective-v.com, includingapollographql-client-nameandapollographql-client-version. - Keep Kong global CORS allowed origins explicitly including both
https://console.perspective-v.comandhttps://hassan.taj.contactfor browser graph API preflight. - Keep Kong global CORS allowed origins explicitly including
https://hassantaj.github.ioand ensure/graph/resumeacceptsOPTIONSfor browser preflight. - Keep additive Kong graph parity extension with identity-style versioned API and docs paths:
/graph/v1/{endpoint}(JWT-protected, rewrite to/api/v1/{endpoint}) and/graph/docs/*(public, rewrite to/docs*) while retaining existing/graph/resume,/graph/gopher/*,/resume, and/gopher/*routes. - Add and maintain declarative Kong state at
runtime/stacks/infrastructure/gateway/kong.ymlfor decK-based gateway automation and drift-safe sync workflows.
Gitea Unified Package Registry Decisions (2026-07-01)
- Consolidate the five separate package services (docker-registry, registry-admin, ProGet, BaGet, Verdaccio) into one Gitea instance serving Docker images, npm, and NuGet.
- Run Gitea as a package-only registry with git features disabled (SSH off, registration disabled, no repos), host
gitea.perspective-v.combehind edge Traefik TLS. - Use organization
perspective-vas the single owner namespace for all Docker/npm/NuGet packages. - Make the instance public with sign-in required (registration disabled); admin user is
pvadmin(Gitea reservesadmin). - Use Gitea user + access-token auth:
write:packageto push (CI, token),read:packageto pull (runtime pull usersvc-puller, and WUD monitor userwud-monitor). - Migrate all existing package data into Gitea (NuGet from BaGet, Docker from the old registry, npm
@pv/corefrom the Azure DevOpspv-ngfeed), repoint CI/CD templates and runtime service image refs to Gitea, and reconfigure WUD to watch Gitea. - Parallel-run the old five services during migration, then retire them: remove
infra-registry+infra-feedsstacks and free the old subdomains (registry.,registry-admin.,nuget.,npm.,proget.). - Publish a Gitea usage/token guide (
docs/runtime/stacks/gitea.mdx) and a legacy-service retirement runbook (docs/runtime/gitea-retirement-runbook.mdx).
Swarm Service Naming Decisions (2026-08-14)
Superseded later the same day by the header-defined stack decisions below.
- Group independently scalable administration UIs in one
panelsstack and label themcom.perspective-v.role=panel. - Limit panel scaling to zero or one replica because multiple panels use single-instance persistent state.
- Keep Kener in
infra-edge; it is the public status component, not part of administration-panel suspension controls. - Classify Vaultwarden, Zitadel, and Gitea as application services with
svc-vaultwarden,svc-zitadel, andsvc-giteastack identities while retaining their existing repository and environment paths.
Swarm Role Layout and Operations Decisions (2026-08-14)
Directory and split-file decisions remain active; deployed namespaces were superseded later the same day by the header-defined stack decisions below.
- Organize active Swarm manifests and launchers into
databases,platform,edge,services,panels, andwebsites. - Maintain one YAML fragment per active service and require launchers to deploy the complete ordered fragment set for each split stack.
- Split Redis and RabbitMQ into
infra-redisandinfra-rabbitmq; rename Watchtower toinfra-watchtower. - Classify Kong, NetBird, and RustFS as
svc-kong,svc-netbird, andsvc-rustfs. - Keep Traefik and Kener together in
infra-edgeas service keysedge-traefikandedge-kener. - Use one administration stack named
panels, producingpanels_*services. - Maintain host-wide backup, restore, diagnostic, migration, systemd, and test
tooling under top-level
operations/.
Header-Defined Swarm Stack Decisions (2026-08-14)
- Treat each fragment's first-line
# Swarm stack:value as the authoritative deployed stack namespace; fragments with the same value deploy together. - Use shared stacks
database,platform,edge,panel, andservice. - Keep Kong, NetBird, and Zitadel in dedicated
service-kong,service-netbird, andservice-zitadelstacks. - Use one
websites-<site>namespace per public/client-facing application group, while leaving undeployed website groups inactive until separately approved. - Preserve one YAML fragment per service, external state names, image identity, routes, secrets, access policy, replica counts, and backup integration during namespace changes.
Homelab Pi-hole Administration (2026-09-15)
- Expose
pihole.home.perspective-v.comonly through Contabo Traefik with thenetbird-onlyandsecurity-headersmiddlewares. - Proxy the route over NetBird to
http://100.83.117.37:8053; do not add a public or LAN dashboard listener on the homelab. - Add only TCP 8053 to the existing
contabo-homelab-server-linkpolicy and point the private NetBird DNS record at Contabo's NetBird IP100.83.72.162.
Home Assistant access (2026-09-30)
- Route
assistant.home.perspective-v.comthrough Contabo withnetbird-onlyandsecurity-headers-allow-sameorigin. - Forward only to the homelab NetBird address at
100.83.117.37:8123. - Add only TCP 8123 to the existing
contabo-homelab-server-linkpolicy; point private NetBird DNS at100.83.72.162and public DNS at Contabo's public IP161.97.83.142. Keep the route behindnetbird-only.
Homelab LAN certificate handoff (2026-09-15)
- Keep Contabo's ACME store and TLS-ALPN-01 resolver authoritative for the existing public hostnames.
- Export only the six homelab certificate/key pairs through a root-only repository script over the already-established homelab-to-Contabo SSH path.
- Do not expose the ACME store, add DNS provider credentials, or change public Traefik routes as part of the LAN split-DNS rollout.
Homelab split-DNS documentation (2026-09-15)
- Keep Namecheap public records and Contabo Traefik authoritative for remote access; implement the LAN optimization only with Pi-hole overrides to 192.168.1.135.
- Reuse the same hostnames and HTTPS certificates on both paths so application callbacks and remote access remain unchanged.
- Keep pihole.home.perspective-v.com and other administration routes NetBird-only; do not expose admin services through public DNS or LAN port forwards.