Perspective V Docs

Monitoring Evaluation Execution Bundle

Approval-ready execution path for installing Cockpit, Beszel and Netdata on the host, evaluating Beszel against Netdata on real workload, and retiring the loser.

Purpose

Install host administration (Cockpit, permanent) and run Beszel and Netdata side by side to decide which single observability tool earns a permanent place.

OUTCOME — CLOSED 2026-08-15: Beszel retained, Netdata retired

The owner ended the evaluation early and chose Beszel. Netdata was removed the same day.

Final architecture: Cockpit + Portainer + Kener + Beszel.

Netdata reached ~175–206 MB against Beszel's ~21 MB, and at 2s collection held dockerd+containerd at ~24% CPU on a 4-core box shared with 30 production services. It also proved materially more demanding to configure correctly — see the gotchas recorded in Phase C, every one of which produced a silent failure rather than an error. Its depth was never in question (~2,300 go.d charts plus per-process attribution); the capability simply did not justify the cost and maintenance load on one host.

Retirement performed: 19 packages purged, apt repository removed, /etc/netdata, /var/lib/netdata, /var/cache/netdata and /var/log/netdata deleted, the three least-privilege DB monitoring users dropped, Traefik route and UFW rule removed, and all repository artifacts deleted. Verified: no listener on 19999, binary absent, no drift.

The Phase C record below is retained deliberately — it documents Netdata behaviours that would otherwise have to be rediscovered if it is ever reconsidered.

Responsibilities — do not let these overlap

ToolAnswersStatus
Kener"Is the service alive and reachable?"existing
Portainer"Manage Docker / Swarm."existing
Cockpit"Manage the Ubuntu host."new, permanent
Beszel or Netdata"What has the infrastructure been doing, and why did it break?"new, one survives

Cockpit has no Docker plugin (only cockpit-podman, unused here). Containers remain Portainer's job.

Scope

  • Host tooling: operations/monitoring/, operations/systemd/{beszel,netdata}/
  • Diagnostics: operations/diagnostics/monitoring-diagnostics.sh
  • Routing: swarm/configs/traefik/dynamic/host-services.yml
  • One production manifest change: swarm/stacks/platform/rabbitmq/rabbitmq.yml gains a loopback-only 15672 publish so the host-installed Netdata can reach its RabbitMQ collector. AMQP (5672) stays unpublished; the overlay and Traefik paths are unchanged.

Approval gate

Owner approval required per phase. Package installs, system-account creation, firewall changes, database-user creation and the RabbitMQ redeploy are all VPS/PROD mutations under AGENTS.md §5. Execute phases one at a time with a measurement pass between each.

Why nothing here runs in Swarm

Two independent reasons, both load-bearing:

  1. Swarm silently drops cap_add, pid, security_opt, privileged, devices and tmpfs. No warning; the service reports healthy. Netdata needs SYS_PTRACE/ SYS_ADMIN and pid: host for apps.plugin and full cgroup access, so a Swarm-deployed Netdata would come up "working" while missing exactly the capability being evaluated — producing a rigged comparison against a healthy Beszel.
  2. Monitoring must not depend on what it monitors. A containerised dashboard is unavailable precisely when Docker or Swarm breaks. So the Beszel hub is host-installed too, though it would run fine in a container.

Access model

URLBackendAuth
https://cockpit.perspective-v.comhost 172.27.0.1:9090NetBird + PAM (dedicated account)
https://netdata.perspective-v.comhost 172.27.0.1:19999NetBird + admin-auth
https://beszel.perspective-v.comhost 172.27.0.1:8090NetBird + Beszel login

All three need DNS A records → 161.97.83.142 before first deploy, or ACME issuance fails the same way data.arnexglobal.com did.

Known limitation — the container path is not gated

Traefik runs in a container and reaches these services at the docker_gwbridge gateway. At the network layer the host cannot distinguish Traefik from the other 30 containers on that bridge, so any of them can reach 172.27.0.1:{9090,19999,8090} directly and bypass the netbird-only middleware. No UFW or ipAllowList rule fixes this — bridge addresses are dynamic and indistinguishable. It is inherent to proxying from a container to a host port.

Consequences, and why this was accepted (owner decision, 2026-08-15):

  • Cockpit — PAM login required; a container reaches a login page only. root is refused via /etc/cockpit/disallowed-users.
  • Beszel — own login required; a container reaches a login page only.
  • Netdata — has no built-in authentication; its only native control is the bind address. admin-auth was added to its router to protect the browser path, but a container could still read host metrics directly.

Closing it entirely requires removing Traefik from the path and binding to the NetBird interface, trading away TLS and the hostnames. Revisit if the threat model changes.

Pre-flight

Baseline already recorded — Phase 1 of the original plan is complete:

2026-08-15   available RAM 2904 MB   load 0.30/0.63/0.78   disk free 106 GB   31 containers

Budget: Cockpit ~0 idle (socket-activated), Beszel hub+agent ~40 MB, Netdata tuned ≤250 MB → ≤300 MB total, roughly 10% of remaining headroom.

Abort criteria: stop and reduce scope if available RAM drops below ~2 GB or load average sustains above 2.0.

Phase A — Cockpit ✅ EXECUTED 2026-08-15

# VPS / PROD
apt install --no-install-recommends cockpit cockpit-system cockpit-storaged cockpit-packagekit
adduser --disabled-password --gecos "Cockpit administration" hassan
usermod -aG sudo hassan
operations/systemd/cockpit/install.sh                     # bind override
operations/monitoring/install-monitoring-config.sh --tools cockpit \
    --cockpit-host cockpit.perspective-v.com
operations/monitoring/ufw-monitoring.sh --ports 9090

cockpit-networkmanager deliberately EXCLUDED

The dry run showed it pulls in network-manager as a dependency. This host runs systemd-networkd via netplan (50-cloud-init.yaml), NetworkManager is not installed, and the interfaces at stake are eth0 (the only public path), wt0 (the NetBird admin path) and the Docker bridges. Provider VNC console is disabled per the 2026-04 incident follow-up, so there is no out-of-band recovery if networking breaks.

Introducing a second network manager on a remote-only production host, for a configuration GUI, is not a worthwhile trade. What is lost is Cockpit's Networking configuration page only; network metrics come from Netdata/Beszel, and interface configuration on this box belongs in netplan regardless.

Also excluded: cockpit-machines, cockpit-389-ds, cockpit-sosreport, cockpit-tests. Installed set is 19 packages, no NetworkManager.

Bind addresses — two, on purpose

operations/systemd/cockpit/install.sh narrows cockpit.socket from all interfaces to:

  • 172.27.0.1:9090 — how Traefik reaches it for the HTTPS hostname
  • <netbird-ip>:9090 — direct fallback

The second is not redundant. docker_gwbridge only exists while Docker is running, so binding solely to it would make Cockpit unreachable exactly when Docker is broken — the scenario it exists for. FreeBind=true removes any boot-ordering dependency.

Edge case: without Origins in /etc/cockpit/cockpit.conf, Cockpit rejects every proxied request with a websocket origin error. This is the guaranteed first-attempt failure.

Results

CheckResult
NetworkManager installedno (0 packages)
eth0 / wt0 after installboth UP; dbskc.com still 200
cockpit.socket bind172.27.0.1:9090 + <netbird-ip>:9090 only
UFW rule9090/tcp on docker_gwbridge — interface-scoped, not public
https://cockpit.perspective-v.com from the node403 (non-NetBird correctly denied)
TLS certificateissued by Let's Encrypt, valid to 2026-11-13
Traefik → backend200 from a container on the proxy overlay
Config driftnone (install-monitoring-config.sh --check exit 0)
Memory cost12 MB RSS (cockpit-ws 6 MB + cockpit-tls 5 MB), 7.4 MB disk
root loginrefused via /etc/cockpit/disallowed-users

Containers on the default bridge (docker0) cannot reach 9090 — the UFW rule is scoped to docker_gwbridge, so the container-path caveat applies only to containers on overlay networks, not every container on the host.

Outstanding

  • Owner must set the account password: passwd hassan. The account is created --disabled-password (status L) because a generated password must not be written to a transcript or log. Cockpit login will not work until this is done.
  • Owner validates from a NetBird client: login as hassan succeeds, root is refused, and the systemd / journal / storage / terminal pages load.

Phase B — Beszel ✅ EXECUTED 2026-08-15

Hub and agent both installed as host binaries under systemd, checksum-verified against beszel_0.18.7_checksums.txt (the asset is versioned — checksums.txt does not exist and returns 404).

CheckResult
beszel-hubactive, 8 MB (ceiling 128M), listening 172.27.0.1:8090 only
beszel-agentactive, 6 MB (ceiling 128M), listening 172.27.0.1:45876 only
Agent host visibilitydetected eth0 and wt0 — confirms real host metrics, not container-scoped
Routehttps://beszel.perspective-v.com → 403 from non-NetBird (correct)
TLSLet's Encrypt, valid to 2026-11-13
UFW8090/tcp on docker_gwbridge — interface-scoped

The agent logs one HUB_URL environment variable not set warning at startup. This is benign and not a loop: Beszel 0.18 attempts WebSocket mode first, then falls back to the SSH listener this deployment uses (hub → agent on 45876).

The hub's agent key did not require the UI — it is derivable from the hub's own keypair:

ssh-keygen -y -f /var/lib/beszel/beszel_data/id_ed25519

Outstanding (owner)

  1. Open https://beszel.perspective-v.com/_/ over NetBird and create the hub admin account. Not scriptable without writing a live password into a transcript.
  2. Add System → host 172.27.0.1, port 45876. The agent is already configured with the matching key.
  3. Configure the Discord webhook under Settings → Notifications (shoutrrr, same library Watchtower uses, so the existing webhook value works unchanged).

Phase B — reference

# VPS / PROD — requires approval
operations/systemd/beszel/install.sh --components hub
operations/monitoring/install-monitoring-config.sh --tools beszel \
    --beszel-host beszel.perspective-v.com
operations/monitoring/ufw-monitoring.sh --ports 8090
# then, over NetBird: create the admin account, Add System, copy the public key
operations/systemd/beszel/install.sh --components agent --hub-key '<key>'

Binaries are checksum-verified against beszel_<version>_checksums.txt (note: the asset is versioned, not checksums.txt); installation aborts if verification fails.

Configure the Discord webhook in the hub UI. Beszel uses shoutrrr — the same library Watchtower uses — so the existing webhook value works unchanged.

Phase C — Netdata

# VPS / PROD — requires approval
operations/systemd/netdata/install.sh
operations/monitoring/netdata/create-monitoring-users.sh          # --dry-run first
# RabbitMQ needs its loopback management port before its collector can work:
swarm/scripts/platform/platform.sh deploy vps                     # restarts rabbitmq
operations/monitoring/install-monitoring-config.sh --tools netdata
operations/monitoring/ufw-monitoring.sh --ports 19999
systemctl restart netdata

Verify the config actually applied. Several [db] keys were renamed between Netdata v1 and v2 and unknown keys are ignored silently:

curl -s http://localhost:19999/netdata.conf > /tmp/effective.conf
diff /tmp/effective.conf /etc/netdata/netdata.conf

Any key present in the repo file but absent from that output is doing nothing. Correct it in operations/monitoring/netdata/netdata.conf and reinstall rather than assuming it took.

Collector privileges (least privilege, created by create-monitoring-users.sh):

EngineGrantCannot
PostgreSQLpg_monitorread table data
MySQLPROCESS, REPLICATION CLIENT, SELECT on performance_schemaread application schemas
RedisACL -@all plus info/ping/slowlog/latencyread keys
RabbitMQmonitoring tag, no vhost permissionspublish or consume

Accepted trade-off: these credentials live in /etc/netdata/go.d/*.conf (0640 root:netdata) as host plaintext, outside the Docker-secrets model. They are rendered from the git-ignored env file rather than typed into /etc, so every live secret stays in one known location.

Phase D — Measure

Run after each phase, so cost is attributable rather than aggregate:

operations/diagnostics/monitoring-diagnostics.sh

Reports host capacity against the recorded baseline, per-unit memory against its ceiling, metrics-database growth, config drift, and whether anything became publicly reachable.

Phase E — Evaluation window (1–2 weeks)

Diagnose every real incident in both tools. The host is stable (73 days uptime, load 0.3 at planning time), so if no incident occurs, force a fair comparison:

operations/monitoring/loadtest/synthetic-load.sh --dry-run     # review the plan first
operations/monitoring/loadtest/synthetic-load.sh

Bounded by design: never saturates all cores, refuses to run below 1.5 GB free RAM or 5 GB free disk, and touches no Docker, database or application data. Run it in an agreed low-traffic window.

Evaluation matrix

Score each on whether it answered the question, and how fast.

AreaBeszelNetdata
Current CPU/RAM
Historical CPU/RAM
Container monitoring
Disk usage
Disk I/O diagnosis
Network monitoring
Alerts
Historical incident investigation
PostgreSQL visibility
Redis visibility
RabbitMQ visibility
systemd visibility
Root-cause analysis
Ease of use
Dashboard clarity
Resource consumption
Storage consumption
Maintenance burden
Security complexity
Actual frequency of use

Decision rule

Do not pick Netdata merely because it has more features. Decide on operational value.

Keep Beszel if the real questions are mostly "is CPU high", "which container is eating RAM", "how much disk is left", "what happened a few hours ago" — and Beszel answers them quickly. Netdata's complexity is then unjustified.

Keep Netdata if questions like "why is iowait high", "which disk has latency", "was memory pressure involved", "did PostgreSQL activity correlate with the API slowdown" recur and Netdata materially answers them.

Keep both only if they provide genuinely distinct recurring value. Duplicate dashboards cost maintenance, attack surface, upgrades and storage.

Phase F — Teardown of the loser

Nothing here is a Swarm stack, so teardown touches no running container:

# Beszel
systemctl disable --now beszel-hub beszel-agent
rm -f /etc/systemd/system/beszel-{hub,agent}.service /usr/local/bin/beszel{,-agent}
rm -rf /var/lib/beszel /etc/beszel && userdel beszel
operations/monitoring/ufw-monitoring.sh --remove --ports 8090

# Netdata — also drop the four monitoring users
apt purge netdata && rm -rf /etc/netdata /var/lib/netdata /var/cache/netdata
operations/monitoring/ufw-monitoring.sh --remove --ports 19999

Then remove the router/service block from host-services.yml (file provider has watch: true, so no Traefik restart) and, if Netdata loses, revert the RabbitMQ loopback port publish.

Update docs/operations/baselines/lightweight-monitoring-baseline.mdx and docs/state/next-steps.mdx with the outcome and rationale.

Rollback

Every phase is independently reversible and none modifies an existing production service, volume or secret — the only production manifest change is the additive RabbitMQ loopback port. Traefik routing is a new file; deleting it restores prior behaviour with no restart. UFW rules are additive and interface-scoped; remove with --remove.

On this page