From c40e692c404ab9237e6ea823762719701632dd4e Mon Sep 17 00:00:00 2001 From: Mark Date: Sun, 2 Aug 2026 11:50:05 +0800 Subject: [PATCH 1/3] Record live-host verification of the no-mesh TRMM installer MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Full interactive run: pinned script verified, both patches applied, compile succeeded, registration completed and the tacticalagent service is running on 2.11.0 with no mesh agent. Remaining checks are UI-side only. Also notes the two benign messages seen during the run — the getcwd warnings (operator's shell was in a deleted directory before the script started; reproduced in a container) and rmmagent's board_serial inventory warning. --- Progress.md | 27 +++++++++++++++++++++------ 1 file changed, 21 insertions(+), 6 deletions(-) diff --git a/Progress.md b/Progress.md index 9f1a002..a1c2559 100644 --- a/Progress.md +++ b/Progress.md @@ -74,14 +74,29 @@ carry the same pin and must be re-pinned together. - Guard unit tests (8/8) against the shipped `patch_community_script`: applies at exactly 1 match, aborts at 0 matches (re-patch) and at 2 matches (ambiguous). -**Still outstanding — needs the real Debian host** (nothing below can be done off-host): +**VERIFIED ON A LIVE HOST (2026-08-02)** — full interactive run against `rmm-api.resort-manager.com`, +registered to client Komune / site Hotel as a server: -1. Registration against the live TRMM server and the systemd unit (`systemctl status tacticalagent` active). -2. Agent appears in TRMM under the right client/site/type. -3. Web terminal connects from the TRMM UI — if not, check the **Use Terminal** role permission first. -4. `update-tacticalrmm-agent-linux.sh` rebuilds and restarts cleanly on the same host with the pinned script. +- Pinned build script downloaded, **SHA-256 matched**, both dispatcher patches reported applied. +- Source pre-fetch used (920K), 2 community download lines neutralised, compile succeeded. +- `Installation was successful!` — registration completed against the live server. +- `tacticalagent` service running; **agent version 2.11.0**. +- `No mesh agent present — registering without a mesh node id` — the no-mesh path works as designed. -Then update this section to fully verified with the version it landed on. +**Still outstanding (UI-side only):** + +1. Confirm the agent shows in the TRMM UI under Komune / Hotel with type server. +2. Web terminal connects from the TRMM UI — if not, check the **Use Terminal** role permission first. +3. `update-tacticalrmm-agent-linux.sh` rebuilds and restarts cleanly on the same host with the pinned script. + +**Two benign messages seen during the live run — neither is a script fault:** + +- `shell-init / job-working-directory: error retrieving current directory: getcwd: ...` on every fork. The + operator's shell was sitting in a deleted directory *before* the script started (the first one appears ahead + of the script banner). Reproduced in a container: a deleted cwd emits this on every subshell while the script + itself runs fine. Our scripts never `cd`, so they cannot cause it. Fix: `cd ~` or open a new shell. +- `WARNING: Unable to read board_serial: permission denied` — emitted by rmmagent itself while collecting + inventory on a host with restricted `/sys/class/dmi/id`. One inventory field is missing; nothing else. ### Bugs caught while building this (both fixed) From 8ae91baf8c3731d7888d9be39e5e4b0caada498e Mon Sep 17 00:00:00 2001 From: Mark Date: Sun, 2 Aug 2026 13:05:54 +0800 Subject: [PATCH 2/3] Detect orphaned Zabbix plugin configs by binary, not by package name MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A live Debian 13 host got stuck in a dpkg loop that this script was written to fix but would have missed. zabbix-agent2 could never configure because ExecStartPre runs `zabbix_agent2 -T`, which died with: plugin "EmberPlus": fork/exec /usr/libexec/zabbix/zabbix-agent2-plugin-ember-plus: no such file or directory The plugin package had been removed but its plugins.d/ember.conf survived, so the agent kept trying to launch a binary that no longer existed. The failure is self-perpetuating: `dpkg --configure -a` re-runs the same config test. Detection derived the package name from the config filename, but the config, package and binary names differ per plugin: ember.conf -> zabbix-agent2-plugin-ember-plus nvidia.conf -> zabbix-agent2-plugin-nvidia-gpu so it missed precisely the two plugins that break hosts in practice. It now checks whether the binary named in System.Path= is executable, which is name-independent and also catches a config dpkg no longer tracks — the case on the live host, which had no rc-state packages at all, only a leftover file. Repair additionally purges rc-state plugin packages, guarded against an empty package list so it cannot produce a bare `dpkg --purge` invocation. Verified 6/6 in a Debian 13 container staged in the host's state: report names both plugins by missing binary, repair clears them, and the config test passes afterwards. The orphan failure itself was reproduced byte-for-byte first. --- Progress.md | 28 ++++++++++++++++++++ repair-rmm-zabbix-linux.sh | 53 ++++++++++++++++++++++++++------------ 2 files changed, 65 insertions(+), 16 deletions(-) diff --git a/Progress.md b/Progress.md index a1c2559..cfda95f 100644 --- a/Progress.md +++ b/Progress.md @@ -157,6 +157,34 @@ check is present. Non-issues: `PROXY_MODE`/`DB_TYPE` in `install-zabbix-proxy.sh documenting hardcoded choices (the config writes `ProxyMode=0` literally) — cosmetic only; SC2076 in `migrate-ufw-to-iptables.sh` is a literal substring match, which is the intended behaviour. +### Zabbix orphaned plugin config — real-world case that broke my detection (2026-08-02) + +A live host (KomuneProxy, Debian 13) got stuck in a dpkg loop: `zabbix-agent2` could never configure because +`ExecStartPre` runs `zabbix_agent2 -T`, which died with + + plugin "EmberPlus": fork/exec /usr/libexec/zabbix/zabbix-agent2-plugin-ember-plus: no such file or directory + +The plugin package had been removed but its `plugins.d/ember.conf` survived, so the agent kept trying to launch a +binary that was gone. Reproduced byte-for-byte in a container (`dpkg --remove` leaves the conffile; purging it +restores `Validation successful`). + +**This exposed a bug in `repair-rmm-zabbix-linux.sh`**: it derived the package name from the config filename, but +the three names differ per plugin — + +| config | package | binary | +| --- | --- | --- | +| `ember.conf` | `zabbix-agent2-plugin-ember-plus` | `/usr/libexec/zabbix/zabbix-agent2-plugin-ember-plus` | +| `nvidia.conf` | `zabbix-agent2-plugin-nvidia-gpu` | `/usr/libexec/zabbix/zabbix-agent2-plugin-nvidia-gpu` | + +so it missed exactly the two plugins that break hosts in practice. Detection now checks whether the `System.Path=` +binary is executable — name-independent, and it also catches a config dpkg no longer tracks. Repair additionally +purges `rc`-state plugin packages (guarded against an empty list). Verified 6/6 against a container staged in the +host's state. + +Worth remembering: the failure is self-perpetuating — `dpkg --configure -a` re-runs the same config test, so the +host cannot recover on its own; and the file-level fix is required because on the live host there were **no** +`rc`-state packages at all, only an untracked leftover config. + ## Completed - 2026-06-29: VERIFIED WORKING end-to-end on a live agent: 2.10.0 -> 2.11.0, service running. Committed the TRMM bootstrap as `trmm-self-update-bootstrap.sh`, updated README (manual vs TRMM-bootstrap usage, with the self-restart/cgroup explanation) and CLAUDE.md structure. diff --git a/repair-rmm-zabbix-linux.sh b/repair-rmm-zabbix-linux.sh index eff6163..7a22b92 100755 --- a/repair-rmm-zabbix-linux.sh +++ b/repair-rmm-zabbix-linux.sh @@ -346,22 +346,27 @@ diagnose_zabbix() { [[ $has_pkg -eq 1 ]] && finding "zabbix" "broken" "package installed but $ZBX_CONF is missing" fi - # --- Plugin configs pointing at plugins that are not installed ---------- + # --- Plugin configs pointing at a binary that is not there -------------- # This is the classic crash loop: the agent loads plugins.d/*.conf and dies - # when the referenced loadable plugin binary is absent. + # when the referenced loadable plugin binary is absent, e.g. + # plugin "EmberPlus": fork/exec /usr/libexec/zabbix/zabbix-agent2-plugin-ember-plus: + # no such file or directory + # Check the System.Path binary rather than deriving a package name from the + # config filename: the two do not match for every plugin (ember.conf ships + # in zabbix-agent2-plugin-ember-plus, nvidia.conf in ...-nvidia-gpu), and a + # name-based guess misses exactly the plugins that break hosts in practice. + # It also catches a config left behind by an incomplete purge, where the + # package is gone from dpkg but its conffile survived. if [[ -d "$ZBX_PLUGINS_D" ]]; then local orphans=0 placeholders=0 while IFS= read -r conf; do [[ -e "$conf" ]] || continue - local base pkg + local base plugin_bin base=$(basename "$conf" .conf) - pkg="zabbix-agent2-plugin-${base}" - if ! pkg_installed "$pkg"; then - # Only flag plugins that ship as separate packages. - if apt-cache show "$pkg" >/dev/null 2>&1; then - log_bad "Plugin config ${base}.conf present but $pkg is not installed" - orphans=$((orphans+1)) - fi + plugin_bin=$(sed -n 's/^[[:space:]]*Plugins\..*\.System\.Path=//p' "$conf" 2>/dev/null | head -1) + if [[ -n "$plugin_bin" && ! -x "$plugin_bin" ]]; then + log_bad "Plugin config ${base}.conf points at a missing binary: ${plugin_bin}" + orphans=$((orphans+1)) fi if grep -q 'CHANGE_ME' "$conf" 2>/dev/null; then log_warn "Plugin config ${base}.conf still has placeholder credentials" @@ -519,23 +524,39 @@ do_repair() { [[ $removed -gt 0 ]] && { repaired "cleared ${removed} stale build leftover(s) from /tmp"; acted=1; } # --- Zabbix -------------------------------------------------------------- - # Disable plugin configs whose plugin package is missing — this is what puts - # the agent into a crash loop after a partial install. + # Disable plugin configs whose binary is missing — this is what puts the + # agent into a crash loop, and it survives an interrupted purge (the package + # is gone from dpkg but its conffile is left behind). if [[ -d "$ZBX_PLUGINS_D" ]]; then local disabled=0 while IFS= read -r conf; do [[ -e "$conf" ]] || continue - local base pkg + local base plugin_bin base=$(basename "$conf" .conf) - pkg="zabbix-agent2-plugin-${base}" - if ! pkg_installed "$pkg" && apt-cache show "$pkg" >/dev/null 2>&1; then + plugin_bin=$(sed -n 's/^[[:space:]]*Plugins\..*\.System\.Path=//p' "$conf" 2>/dev/null | head -1) + if [[ -n "$plugin_bin" && ! -x "$plugin_bin" ]]; then mv "$conf" "${conf}.disabled" 2>/dev/null \ - && { log_info "Disabled ${base}.conf (plugin $pkg not installed)"; disabled=$((disabled+1)); } + && { log_info "Disabled ${base}.conf (missing binary ${plugin_bin})"; disabled=$((disabled+1)); } fi done < <(find "$ZBX_PLUGINS_D" -maxdepth 1 -name '*.conf' 2>/dev/null) [[ $disabled -gt 0 ]] && { repaired "disabled ${disabled} orphaned plugin config(s)"; acted=1; } fi + # Config files left behind by packages that were removed but never purged. + # dpkg leaves these in "rc" state; the agent still reads them and dies on the + # missing binary, and `dpkg --configure -a` can never succeed while it does. + local rc_plugins + rc_plugins=$(dpkg -l 2>/dev/null | awk '/^rc[[:space:]]+zabbix-agent2-plugin/{print $2}') + if [[ -n "$rc_plugins" ]]; then + log_info "Purging leftover config from removed plugin packages: $(echo "$rc_plugins" | tr '\n' ' ')" + # shellcheck disable=SC2086 # deliberate word splitting: list of package names + if dpkg --purge $rc_plugins >/dev/null 2>&1; then + repaired "purged leftover config for $(echo "$rc_plugins" | wc -l) removed plugin package(s)"; acted=1 + else + log_warn "Could not purge leftover plugin config — remove the stale files under ${ZBX_PLUGINS_D} manually" + fi + fi + # Missing Include for plugins.d if [[ -f "$ZBX_CONF" ]] && [[ -d "$ZBX_PLUGINS_D" ]] \ && ! grep -qE "^Include=${ZBX_PLUGINS_D}/\*\.conf" "$ZBX_CONF" 2>/dev/null; then From 647503601fcbba20a17e229349e98aff619117f7 Mon Sep 17 00:00:00 2001 From: Mark Date: Sun, 2 Aug 2026 13:35:48 +0800 Subject: [PATCH 3/3] Stop clean-zabbix destroying a co-installed proxy or server config MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit clean-zabbix is documented as removing "all Zabbix Agent 2 components", but it ran `rm -rf /etc/zabbix` — a directory shared with zabbix-proxy and zabbix-server. On a host running both, it deleted the proxy's configuration for a service it was never asked to touch, leaving a running proxy that would not survive its next restart. The backup taken beforehand only covered the agent's own paths, so the deleted proxy config was not recoverable from it either. Both halves of that are fixed: - Remove only zabbix_agent2.conf and zabbix_agent2.d. The directory itself is removed only when nothing else lives in it, and anything preserved is listed in the output so it is obvious what was left behind. - Back up the whole of /etc/zabbix when it exists, not just the agent's files, so a mistake in that directory can be undone. Verified 5/5 in a Debian 13 container with the proxy and agent co-installed: the proxy config survives clean-zabbix byte-identical, the agent config and conf.d are removed as intended, and the proxy config is present in the backup archive. --- repair-rmm-zabbix-linux.sh | 23 +++++++++++++++++++++-- 1 file changed, 21 insertions(+), 2 deletions(-) diff --git a/repair-rmm-zabbix-linux.sh b/repair-rmm-zabbix-linux.sh index 7a22b92..cbe9f18 100755 --- a/repair-rmm-zabbix-linux.sh +++ b/repair-rmm-zabbix-linux.sh @@ -428,6 +428,10 @@ backup_configs() { for p in "$RMM_CONF" "$RMM_UNIT" "$ZBX_CONF" "$ZBX_CONF_D"; do [[ -e "$p" ]] && paths+=("$p") done + # Back up the whole of /etc/zabbix when it exists, not just the agent's own + # files: a co-installed proxy or server keeps its config in the same + # directory, and a backup that omits it cannot undo a mistake made there. + [[ -d /etc/zabbix ]] && paths+=(/etc/zabbix) [[ -e "/etc/systemd/system/${ZBX_SVC}.service" ]] && paths+=("/etc/systemd/system/${ZBX_SVC}.service") if [[ ${#paths[@]} -eq 0 ]]; then @@ -648,10 +652,25 @@ clean_zabbix() { log_info "No Zabbix agent packages installed" fi - rm -rf "$ZBX_CONF" "$ZBX_CONF_D" /etc/zabbix + # Remove ONLY agent-owned paths. /etc/zabbix is shared with zabbix-proxy and + # zabbix-server: an earlier version of this script did `rm -rf /etc/zabbix` + # here and destroyed a running proxy's configuration on a host where the + # proxy and the agent are co-installed. Never touch the directory itself. + rm -rf "$ZBX_CONF" "$ZBX_CONF_D" rm -f "/etc/systemd/system/${ZBX_SVC}.service" systemctl daemon-reload >/dev/null 2>&1 || true - log_ok "Removed Zabbix configuration" + log_ok "Removed Zabbix agent configuration" + + # Report anything else living in /etc/zabbix so it is obvious we left it. + local other_conf + other_conf=$(find /etc/zabbix -maxdepth 1 -name 'zabbix_*' ! -name 'zabbix_agent2*' 2>/dev/null) + if [[ -n "$other_conf" ]]; then + log_info "Left other Zabbix components untouched:" + while IFS= read -r f; do [[ -n "$f" ]] && echo " $f"; done <<< "$other_conf" + else + # Only remove the directory when nothing else owns anything in it. + rmdir /etc/zabbix 2>/dev/null && log_info "Removed the now-empty /etc/zabbix" + fi # The apt repo is left in place deliberately — a reinstall needs it, and # removing it would force the installer to re-add and re-key the repo.