Perspective V Docs

Lightweight Monitoring Baseline

No heavy metrics/log stack (Prometheus/Loki/Grafana/Elasticsearch) is deployed in this phase.

Status

  • Effective date: 2026-04-15
  • VPS capacity reference: 4 CPU, 8 GB RAM
  • Objective: keep monitoring and update-observability lightweight for current host limits

Active Monitoring/Update Components

  • kener (rajnandan1/kener:v4.0.16)
  • wud (getwud/wud:8.2.2)
  • watchtower (containrrr/watchtower:latest)
  • watchtower-fast (containrrr/watchtower:latest)

No heavy metrics/log stack (Prometheus/Loki/Grafana/Elasticsearch) is deployed in this phase.

Resource Snapshot (UTC 2026-04-15)

Host capacity (Docker):

  • cpus=4
  • mem_bytes=8326971392

Container snapshot (docker stats --no-stream):

  • kener cpu=1.49% mem=149.7MiB / 7.755GiB
  • wud cpu=2.63% mem=87.99MiB / 7.755GiB
  • watchtower cpu=0.00% mem=5.18MiB / 7.755GiB
  • watchtower-fast cpu=0.00% mem=6.871MiB / 7.755GiB

Combined memory footprint in this snapshot is under 260 MiB.

Guardrails

  • Keep Kener + WUD + Watchtower model as baseline for current VPS size.
  • Do not add heavy telemetry stack components without explicit capacity review and owner approval.
  • Keep Discord as primary alert channel for this phase.
  • Re-check monitoring component footprint monthly during DB maintenance window.

Capacity review — host monitoring rollout (2026-08-15)

This baseline previously forbade heavier telemetry without an explicit capacity review. That review is recorded here rather than leaving the guardrail silently contradicted.

Decision: add Cockpit (host administration, permanent) and run Beszel and Netdata concurrently for 1–2 weeks to select one permanently. Owner-approved after being shown the resource cost, including the deliberate choice to keep Netdata despite its footprint.

Capacity at time of decision: 2,904 MB RAM available, load 0.30/0.63/0.78 on 4 cores, 106 GB disk free, 31 containers.

Budget: Cockpit ~0 idle (socket-activated), Beszel hub+agent ~40 MB, Netdata tuned ≤250 MB → ≤300 MB total, roughly 10% of remaining headroom.

Controls that keep this bounded:

  • Hard systemd ceilings: Netdata MemoryMax=400M / CPUQuota=50%; Beszel hub and agent MemoryMax=128M / CPUQuota=25% each. These are kernel-enforced backstops, not tuning — a misconfigured collector cannot starve production.
  • Netdata tuned down: machine learning disabled (its largest RAM/CPU consumer), 2 s collection instead of 1 s, two storage tiers, and ebpf/python.d/charts.d/slabinfo/ perf plugins off.
  • Abort criteria: stop and reduce scope if available RAM drops below ~2 GB or load average sustains above 2.0.
  • operations/diagnostics/monitoring-diagnostics.sh reports actual cost against the baseline above, plus metrics-database growth, after each phase.

Still excluded: Prometheus, Loki, Grafana, Elasticsearch, and any distributed telemetry stack. Unchanged for this VPS size.

OUTCOME (2026-08-15): Beszel retained, Netdata retired

The owner ended the evaluation early and chose Beszel. Netdata was fully removed the same day rather than left running.

Final architecture: Cockpit + Portainer + Kener + Beszel.

Rationale, and what the short evaluation still established:

  • Cost. Netdata settled at ~175–206 MB against Beszel's ~21 MB (hub 15 + agent 6) — roughly 9x for a single host. At 2s collection it also held dockerd + containerd at ~24% CPU on a 4-core box shared with 30 production services; raising it to 5s recovered idle from ~50% to ~78%, but the polling cost was structural.
  • Operational burden. Netdata needed materially more care to run correctly: inline comments silently inverted settings, [db] mode was renamed in v2, unquoted DSNs were rejected without logging, collectors gave up permanently on a single startup failure, and the default install channel was nightly. Each was fixable, but the cumulative maintenance load was the deciding factor for one host.
  • RabbitMQ could not be collected safely at all — see the execution bundle.
  • Depth was never in doubt: Netdata produced ~2,300 go.d charts plus per-process attribution. It was more capable and more expensive; for this host the capability was not worth the cost.

Retired on 2026-08-15: all 19 netdata packages purged, the apt repository removed, /etc/netdata, /var/lib/netdata, /var/cache/netdata and /var/log/netdata deleted, the three least-privilege DB monitoring users dropped from PostgreSQL/MySQL/Redis, the Traefik route and UFW rule removed, and all repository artifacts deleted.

Current monitoring footprint: Beszel ~21 MB + Cockpit ~12 MB + PCP ~15 MB ≈ 48 MB, comfortably back inside the original lightweight guardrail.

Note that the component list above is from the 2026-04-15 snapshot and is now stale in detail (Kener and WUD have since been updated; WUD is 8.3.1). Footprint figures should be re-measured with the diagnostics script rather than read from that snapshot.

On this page