Lightweight Monitoring Baseline
No heavy metrics/log stack (Prometheus/Loki/Grafana/Elasticsearch) is deployed in this phase.
Status
- Effective date: 2026-04-15
- VPS capacity reference: 4 CPU, 8 GB RAM
- Objective: keep monitoring and update-observability lightweight for current host limits
Active Monitoring/Update Components
kener(rajnandan1/kener:v4.0.16)wud(getwud/wud:8.2.2)watchtower(containrrr/watchtower:latest)watchtower-fast(containrrr/watchtower:latest)
No heavy metrics/log stack (Prometheus/Loki/Grafana/Elasticsearch) is deployed in this phase.
Resource Snapshot (UTC 2026-04-15)
Host capacity (Docker):
cpus=4mem_bytes=8326971392
Container snapshot (docker stats --no-stream):
kener cpu=1.49% mem=149.7MiB / 7.755GiBwud cpu=2.63% mem=87.99MiB / 7.755GiBwatchtower cpu=0.00% mem=5.18MiB / 7.755GiBwatchtower-fast cpu=0.00% mem=6.871MiB / 7.755GiB
Combined memory footprint in this snapshot is under 260 MiB.
Guardrails
- Keep Kener + WUD + Watchtower model as baseline for current VPS size.
- Do not add heavy telemetry stack components without explicit capacity review and owner approval.
- Keep Discord as primary alert channel for this phase.
- Re-check monitoring component footprint monthly during DB maintenance window.
Capacity review — host monitoring rollout (2026-08-15)
This baseline previously forbade heavier telemetry without an explicit capacity review. That review is recorded here rather than leaving the guardrail silently contradicted.
Decision: add Cockpit (host administration, permanent) and run Beszel and Netdata concurrently for 1–2 weeks to select one permanently. Owner-approved after being shown the resource cost, including the deliberate choice to keep Netdata despite its footprint.
Capacity at time of decision: 2,904 MB RAM available, load 0.30/0.63/0.78 on 4 cores,
106 GB disk free, 31 containers.
Budget: Cockpit ~0 idle (socket-activated), Beszel hub+agent ~40 MB, Netdata tuned ≤250 MB → ≤300 MB total, roughly 10% of remaining headroom.
Controls that keep this bounded:
- Hard systemd ceilings: Netdata
MemoryMax=400M/CPUQuota=50%; Beszel hub and agentMemoryMax=128M/CPUQuota=25%each. These are kernel-enforced backstops, not tuning — a misconfigured collector cannot starve production. - Netdata tuned down: machine learning disabled (its largest RAM/CPU consumer), 2 s
collection instead of 1 s, two storage tiers, and
ebpf/python.d/charts.d/slabinfo/perfplugins off. - Abort criteria: stop and reduce scope if available RAM drops below ~2 GB or load average sustains above 2.0.
operations/diagnostics/monitoring-diagnostics.shreports actual cost against the baseline above, plus metrics-database growth, after each phase.
Still excluded: Prometheus, Loki, Grafana, Elasticsearch, and any distributed telemetry stack. Unchanged for this VPS size.
OUTCOME (2026-08-15): Beszel retained, Netdata retired
The owner ended the evaluation early and chose Beszel. Netdata was fully removed the same day rather than left running.
Final architecture: Cockpit + Portainer + Kener + Beszel.
Rationale, and what the short evaluation still established:
- Cost. Netdata settled at ~175–206 MB against Beszel's ~21 MB (hub 15 + agent 6) —
roughly 9x for a single host. At 2s collection it also held
dockerd+containerdat ~24% CPU on a 4-core box shared with 30 production services; raising it to 5s recovered idle from ~50% to ~78%, but the polling cost was structural. - Operational burden. Netdata needed materially more care to run correctly: inline
comments silently inverted settings,
[db] modewas renamed in v2, unquoted DSNs were rejected without logging, collectors gave up permanently on a single startup failure, and the default install channel was nightly. Each was fixable, but the cumulative maintenance load was the deciding factor for one host. - RabbitMQ could not be collected safely at all — see the execution bundle.
- Depth was never in doubt: Netdata produced ~2,300 go.d charts plus per-process attribution. It was more capable and more expensive; for this host the capability was not worth the cost.
Retired on 2026-08-15: all 19 netdata packages purged, the apt repository removed,
/etc/netdata, /var/lib/netdata, /var/cache/netdata and /var/log/netdata deleted,
the three least-privilege DB monitoring users dropped from PostgreSQL/MySQL/Redis, the
Traefik route and UFW rule removed, and all repository artifacts deleted.
Current monitoring footprint: Beszel ~21 MB + Cockpit ~12 MB + PCP ~15 MB ≈ 48 MB, comfortably back inside the original lightweight guardrail.
Note that the component list above is from the 2026-04-15 snapshot and is now stale in
detail (Kener and WUD have since been updated; WUD is 8.3.1). Footprint figures should be
re-measured with the diagnostics script rather than read from that snapshot.