Monitoring Evaluation Execution Bundle
Approval-ready execution path for installing Cockpit, Beszel and Netdata on the host, evaluating Beszel against Netdata on real workload, and retiring the loser.
Purpose
Install host administration (Cockpit, permanent) and run Beszel and Netdata side by side to decide which single observability tool earns a permanent place.
OUTCOME — CLOSED 2026-08-15: Beszel retained, Netdata retired
The owner ended the evaluation early and chose Beszel. Netdata was removed the same day.
Final architecture: Cockpit + Portainer + Kener + Beszel.
Netdata reached ~175–206 MB against Beszel's ~21 MB, and at 2s collection held
dockerd+containerd at ~24% CPU on a 4-core box shared with 30 production services. It
also proved materially more demanding to configure correctly — see the gotchas recorded in
Phase C, every one of which produced a silent failure rather than an error. Its depth was
never in question (~2,300 go.d charts plus per-process attribution); the capability simply
did not justify the cost and maintenance load on one host.
Retirement performed: 19 packages purged, apt repository removed, /etc/netdata,
/var/lib/netdata, /var/cache/netdata and /var/log/netdata deleted, the three
least-privilege DB monitoring users dropped, Traefik route and UFW rule removed, and all
repository artifacts deleted. Verified: no listener on 19999, binary absent, no drift.
The Phase C record below is retained deliberately — it documents Netdata behaviours that would otherwise have to be rediscovered if it is ever reconsidered.
Responsibilities — do not let these overlap
| Tool | Answers | Status |
|---|---|---|
| Kener | "Is the service alive and reachable?" | existing |
| Portainer | "Manage Docker / Swarm." | existing |
| Cockpit | "Manage the Ubuntu host." | new, permanent |
| Beszel or Netdata | "What has the infrastructure been doing, and why did it break?" | new, one survives |
Cockpit has no Docker plugin (only cockpit-podman, unused here). Containers remain
Portainer's job.
Scope
- Host tooling:
operations/monitoring/,operations/systemd/{beszel,netdata}/ - Diagnostics:
operations/diagnostics/monitoring-diagnostics.sh - Routing:
swarm/configs/traefik/dynamic/host-services.yml - One production manifest change:
swarm/stacks/platform/rabbitmq/rabbitmq.ymlgains a loopback-only15672publish so the host-installed Netdata can reach its RabbitMQ collector. AMQP (5672) stays unpublished; the overlay and Traefik paths are unchanged.
Approval gate
Owner approval required per phase. Package installs, system-account creation, firewall changes, database-user creation and the RabbitMQ redeploy are all VPS/PROD mutations under AGENTS.md §5. Execute phases one at a time with a measurement pass between each.
Why nothing here runs in Swarm
Two independent reasons, both load-bearing:
- Swarm silently drops
cap_add,pid,security_opt,privileged,devicesandtmpfs. No warning; the service reports healthy. Netdata needsSYS_PTRACE/SYS_ADMINandpid: hostforapps.pluginand full cgroup access, so a Swarm-deployed Netdata would come up "working" while missing exactly the capability being evaluated — producing a rigged comparison against a healthy Beszel. - Monitoring must not depend on what it monitors. A containerised dashboard is unavailable precisely when Docker or Swarm breaks. So the Beszel hub is host-installed too, though it would run fine in a container.
Access model
| URL | Backend | Auth |
|---|---|---|
https://cockpit.perspective-v.com | host 172.27.0.1:9090 | NetBird + PAM (dedicated account) |
https://netdata.perspective-v.com | host 172.27.0.1:19999 | NetBird + admin-auth |
https://beszel.perspective-v.com | host 172.27.0.1:8090 | NetBird + Beszel login |
All three need DNS A records → 161.97.83.142 before first deploy, or ACME issuance
fails the same way data.arnexglobal.com did.
Known limitation — the container path is not gated
Traefik runs in a container and reaches these services at the docker_gwbridge gateway. At
the network layer the host cannot distinguish Traefik from the other 30 containers on that
bridge, so any of them can reach 172.27.0.1:{9090,19999,8090} directly and bypass the
netbird-only middleware. No UFW or ipAllowList rule fixes this — bridge addresses are
dynamic and indistinguishable. It is inherent to proxying from a container to a host port.
Consequences, and why this was accepted (owner decision, 2026-08-15):
- Cockpit — PAM login required; a container reaches a login page only.
rootis refused via/etc/cockpit/disallowed-users. - Beszel — own login required; a container reaches a login page only.
- Netdata — has no built-in authentication; its only native control is the bind
address.
admin-authwas added to its router to protect the browser path, but a container could still read host metrics directly.
Closing it entirely requires removing Traefik from the path and binding to the NetBird interface, trading away TLS and the hostnames. Revisit if the threat model changes.
Pre-flight
Baseline already recorded — Phase 1 of the original plan is complete:
2026-08-15 available RAM 2904 MB load 0.30/0.63/0.78 disk free 106 GB 31 containersBudget: Cockpit ~0 idle (socket-activated), Beszel hub+agent ~40 MB, Netdata tuned ≤250 MB → ≤300 MB total, roughly 10% of remaining headroom.
Abort criteria: stop and reduce scope if available RAM drops below ~2 GB or load average sustains above 2.0.
Phase A — Cockpit ✅ EXECUTED 2026-08-15
# VPS / PROD
apt install --no-install-recommends cockpit cockpit-system cockpit-storaged cockpit-packagekit
adduser --disabled-password --gecos "Cockpit administration" hassan
usermod -aG sudo hassan
operations/systemd/cockpit/install.sh # bind override
operations/monitoring/install-monitoring-config.sh --tools cockpit \
--cockpit-host cockpit.perspective-v.com
operations/monitoring/ufw-monitoring.sh --ports 9090cockpit-networkmanager deliberately EXCLUDED
The dry run showed it pulls in network-manager as a dependency. This host runs
systemd-networkd via netplan (50-cloud-init.yaml), NetworkManager is not installed, and
the interfaces at stake are eth0 (the only public path), wt0 (the NetBird admin path)
and the Docker bridges. Provider VNC console is disabled per the 2026-04 incident
follow-up, so there is no out-of-band recovery if networking breaks.
Introducing a second network manager on a remote-only production host, for a configuration GUI, is not a worthwhile trade. What is lost is Cockpit's Networking configuration page only; network metrics come from Netdata/Beszel, and interface configuration on this box belongs in netplan regardless.
Also excluded: cockpit-machines, cockpit-389-ds, cockpit-sosreport, cockpit-tests.
Installed set is 19 packages, no NetworkManager.
Bind addresses — two, on purpose
operations/systemd/cockpit/install.sh narrows cockpit.socket from all interfaces to:
172.27.0.1:9090— how Traefik reaches it for the HTTPS hostname<netbird-ip>:9090— direct fallback
The second is not redundant. docker_gwbridge only exists while Docker is running, so
binding solely to it would make Cockpit unreachable exactly when Docker is broken — the
scenario it exists for. FreeBind=true removes any boot-ordering dependency.
Edge case: without Origins in /etc/cockpit/cockpit.conf, Cockpit rejects every
proxied request with a websocket origin error. This is the guaranteed first-attempt failure.
Results
| Check | Result |
|---|---|
| NetworkManager installed | no (0 packages) |
eth0 / wt0 after install | both UP; dbskc.com still 200 |
cockpit.socket bind | 172.27.0.1:9090 + <netbird-ip>:9090 only |
| UFW rule | 9090/tcp on docker_gwbridge — interface-scoped, not public |
https://cockpit.perspective-v.com from the node | 403 (non-NetBird correctly denied) |
| TLS certificate | issued by Let's Encrypt, valid to 2026-11-13 |
| Traefik → backend | 200 from a container on the proxy overlay |
| Config drift | none (install-monitoring-config.sh --check exit 0) |
| Memory cost | 12 MB RSS (cockpit-ws 6 MB + cockpit-tls 5 MB), 7.4 MB disk |
root login | refused via /etc/cockpit/disallowed-users |
Containers on the default bridge (docker0) cannot reach 9090 — the UFW rule is scoped
to docker_gwbridge, so the container-path caveat applies only to containers on overlay
networks, not every container on the host.
Outstanding
- Owner must set the account password:
passwd hassan. The account is created--disabled-password(statusL) because a generated password must not be written to a transcript or log. Cockpit login will not work until this is done. - Owner validates from a NetBird client: login as
hassansucceeds,rootis refused, and the systemd / journal / storage / terminal pages load.
Phase B — Beszel ✅ EXECUTED 2026-08-15
Hub and agent both installed as host binaries under systemd, checksum-verified against
beszel_0.18.7_checksums.txt (the asset is versioned — checksums.txt does not exist and
returns 404).
| Check | Result |
|---|---|
beszel-hub | active, 8 MB (ceiling 128M), listening 172.27.0.1:8090 only |
beszel-agent | active, 6 MB (ceiling 128M), listening 172.27.0.1:45876 only |
| Agent host visibility | detected eth0 and wt0 — confirms real host metrics, not container-scoped |
| Route | https://beszel.perspective-v.com → 403 from non-NetBird (correct) |
| TLS | Let's Encrypt, valid to 2026-11-13 |
| UFW | 8090/tcp on docker_gwbridge — interface-scoped |
The agent logs one HUB_URL environment variable not set warning at startup. This is
benign and not a loop: Beszel 0.18 attempts WebSocket mode first, then falls back to the
SSH listener this deployment uses (hub → agent on 45876).
The hub's agent key did not require the UI — it is derivable from the hub's own keypair:
ssh-keygen -y -f /var/lib/beszel/beszel_data/id_ed25519Outstanding (owner)
- Open
https://beszel.perspective-v.com/_/over NetBird and create the hub admin account. Not scriptable without writing a live password into a transcript. - Add System → host
172.27.0.1, port45876. The agent is already configured with the matching key. - Configure the Discord webhook under Settings → Notifications (shoutrrr, same library Watchtower uses, so the existing webhook value works unchanged).
Phase B — reference
# VPS / PROD — requires approval
operations/systemd/beszel/install.sh --components hub
operations/monitoring/install-monitoring-config.sh --tools beszel \
--beszel-host beszel.perspective-v.com
operations/monitoring/ufw-monitoring.sh --ports 8090
# then, over NetBird: create the admin account, Add System, copy the public key
operations/systemd/beszel/install.sh --components agent --hub-key '<key>'Binaries are checksum-verified against beszel_<version>_checksums.txt (note: the asset is
versioned, not checksums.txt); installation aborts if verification fails.
Configure the Discord webhook in the hub UI. Beszel uses shoutrrr — the same library Watchtower uses — so the existing webhook value works unchanged.
Phase C — Netdata
# VPS / PROD — requires approval
operations/systemd/netdata/install.sh
operations/monitoring/netdata/create-monitoring-users.sh # --dry-run first
# RabbitMQ needs its loopback management port before its collector can work:
swarm/scripts/platform/platform.sh deploy vps # restarts rabbitmq
operations/monitoring/install-monitoring-config.sh --tools netdata
operations/monitoring/ufw-monitoring.sh --ports 19999
systemctl restart netdataVerify the config actually applied. Several [db] keys were renamed between Netdata v1
and v2 and unknown keys are ignored silently:
curl -s http://localhost:19999/netdata.conf > /tmp/effective.conf
diff /tmp/effective.conf /etc/netdata/netdata.confAny key present in the repo file but absent from that output is doing nothing. Correct it in
operations/monitoring/netdata/netdata.conf and reinstall rather than assuming it took.
Collector privileges (least privilege, created by create-monitoring-users.sh):
| Engine | Grant | Cannot |
|---|---|---|
| PostgreSQL | pg_monitor | read table data |
| MySQL | PROCESS, REPLICATION CLIENT, SELECT on performance_schema | read application schemas |
| Redis | ACL -@all plus info/ping/slowlog/latency | read keys |
| RabbitMQ | monitoring tag, no vhost permissions | publish or consume |
Accepted trade-off: these credentials live in /etc/netdata/go.d/*.conf (0640
root:netdata) as host plaintext, outside the Docker-secrets model. They are rendered from
the git-ignored env file rather than typed into /etc, so every live secret stays in one
known location.
Phase D — Measure
Run after each phase, so cost is attributable rather than aggregate:
operations/diagnostics/monitoring-diagnostics.shReports host capacity against the recorded baseline, per-unit memory against its ceiling, metrics-database growth, config drift, and whether anything became publicly reachable.
Phase E — Evaluation window (1–2 weeks)
Diagnose every real incident in both tools. The host is stable (73 days uptime, load 0.3 at planning time), so if no incident occurs, force a fair comparison:
operations/monitoring/loadtest/synthetic-load.sh --dry-run # review the plan first
operations/monitoring/loadtest/synthetic-load.shBounded by design: never saturates all cores, refuses to run below 1.5 GB free RAM or 5 GB free disk, and touches no Docker, database or application data. Run it in an agreed low-traffic window.
Evaluation matrix
Score each on whether it answered the question, and how fast.
| Area | Beszel | Netdata |
|---|---|---|
| Current CPU/RAM | ||
| Historical CPU/RAM | ||
| Container monitoring | ||
| Disk usage | ||
| Disk I/O diagnosis | ||
| Network monitoring | ||
| Alerts | ||
| Historical incident investigation | ||
| PostgreSQL visibility | ||
| Redis visibility | ||
| RabbitMQ visibility | ||
| systemd visibility | ||
| Root-cause analysis | ||
| Ease of use | ||
| Dashboard clarity | ||
| Resource consumption | ||
| Storage consumption | ||
| Maintenance burden | ||
| Security complexity | ||
| Actual frequency of use |
Decision rule
Do not pick Netdata merely because it has more features. Decide on operational value.
Keep Beszel if the real questions are mostly "is CPU high", "which container is eating RAM", "how much disk is left", "what happened a few hours ago" — and Beszel answers them quickly. Netdata's complexity is then unjustified.
Keep Netdata if questions like "why is iowait high", "which disk has latency", "was memory pressure involved", "did PostgreSQL activity correlate with the API slowdown" recur and Netdata materially answers them.
Keep both only if they provide genuinely distinct recurring value. Duplicate dashboards cost maintenance, attack surface, upgrades and storage.
Phase F — Teardown of the loser
Nothing here is a Swarm stack, so teardown touches no running container:
# Beszel
systemctl disable --now beszel-hub beszel-agent
rm -f /etc/systemd/system/beszel-{hub,agent}.service /usr/local/bin/beszel{,-agent}
rm -rf /var/lib/beszel /etc/beszel && userdel beszel
operations/monitoring/ufw-monitoring.sh --remove --ports 8090
# Netdata — also drop the four monitoring users
apt purge netdata && rm -rf /etc/netdata /var/lib/netdata /var/cache/netdata
operations/monitoring/ufw-monitoring.sh --remove --ports 19999Then remove the router/service block from host-services.yml (file provider has
watch: true, so no Traefik restart) and, if Netdata loses, revert the RabbitMQ loopback
port publish.
Update docs/operations/baselines/lightweight-monitoring-baseline.mdx and
docs/state/next-steps.mdx with the outcome and rationale.
Rollback
Every phase is independently reversible and none modifies an existing production service,
volume or secret — the only production manifest change is the additive RabbitMQ loopback
port. Traefik routing is a new file; deleting it restores prior behaviour with no restart.
UFW rules are additive and interface-scoped; remove with --remove.
Image Update Rollout Execution Bundle
Provide one approval-ready execution path for the seven WUD-reported image updates, with DEV-validated manifest changes, per-step rollback, and validation commands.
Phase 5 Backup Rollout Execution Bundle
Provide one approval-ready execution path for installing the Tier-1 backup timer and capturing first-run evidence.