/home/dev/lightwalletd is simultaneously a git checkout and the runtime directory, and monitor_lwd.sh -- the primary's live supervisor -- was a tracked file inside it. So a routine `git reset --hard` during an unrelated cherry-pick reverted the running supervisor to an older committed revision. That happened on 2026-08-26. Nothing noticed, because the running monitor was executing an already-deleted inode: the file on disk was broken while the live process was fine. It was recovered from /proc/<pid>/fd/255. What the reverted copy would have reintroduced, had it ever restarted: the loss of `-cache-size 5000`, so every relaunch warms the block cache from tip-400000 instead of tip-5000; and `wait "$LWD_PID" || true; EXIT_CODE=$?`, which reads the exit status of `|| true` and is therefore always 0, so the monitor logs "exited cleanly. Not restarting." and breaks its loop on every exit including crashes -- the bug behind an 11h48m outage on 2026-08-21. Move it to deploy/, from where it is copied to /home/dev/ and run. The runtime directory now holds only the binary and its logs, so no checkout, reset, rebase or branch switch can reach a running supervisor. The script no longer derives its paths from its own location: SCRIPT_DIR became an explicit LWD_DIR, because the script and the runtime directory are deliberately no longer the same place. lwd_watchdog.sh launches "$MONITOR" by absolute path and refuses to run if it is missing. Verified: run from /tmp, the relocated monitor resolves its binary and log through LWD_DIR rather than its own directory, and launches with -cache-size 5000 intact. The running monitor was not restarted. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FU87LdsJZiZkfq1eXubpeo
41 lines
1.9 KiB
Markdown
41 lines
1.9 KiB
Markdown
# deploy/
|
|
|
|
Operational scripts, versioned here but **executed from outside this repository**.
|
|
|
|
A node's runtime directory (`/home/dev/lightwalletd`) holds only the binary and
|
|
its logs. The supervisor scripts live in `/home/dev/`:
|
|
|
|
| repo (source of truth) | deployed to | invoked by |
|
|
|------------------------|--------------------|---------------------------|
|
|
| `deploy/monitor_lwd.sh`| `/home/dev/monitor_lwd.sh` | `/home/dev/lwd_watchdog.sh` (cron, primary) |
|
|
|
|
## Why they are not run from here
|
|
|
|
This repository's working tree used to *be* the runtime directory, and
|
|
`monitor_lwd.sh` — the primary's live supervisor — was a tracked file inside it.
|
|
On 2026-08-26 a routine `git reset --hard` during an unrelated cherry-pick
|
|
reverted it to an older committed revision that (a) dropped `-cache-size 5000`,
|
|
making every relaunch warm the block cache from tip-400000, and (b) reintroduced
|
|
`wait "$LWD_PID" || true; EXIT_CODE=$?`, which reads the exit status of `|| true`
|
|
and is therefore always 0, so the monitor logged "exited cleanly. Not restarting."
|
|
and broke its loop on every exit including crashes — the bug behind an 11h48m
|
|
outage on 2026-08-21.
|
|
|
|
Nothing noticed at the time because the running monitor was executing an
|
|
already-deleted inode: the working copy was broken while the live process was
|
|
fine. It was recovered from `/proc/<pid>/fd/255`.
|
|
|
|
Deploying these from outside the working tree means no checkout, reset, rebase or
|
|
branch switch can reach a running supervisor.
|
|
|
|
## Changing one
|
|
|
|
Edit it here, commit, then copy to the node and let the next relaunch pick it up:
|
|
|
|
cp deploy/monitor_lwd.sh /home/dev/monitor_lwd.sh.stage
|
|
chmod 755 /home/dev/monitor_lwd.sh.stage
|
|
mv -f /home/dev/monitor_lwd.sh.stage /home/dev/monitor_lwd.sh
|
|
|
|
`mv`, not `cp`: a rename cannot disturb a running process, and the currently
|
|
running monitor keeps its own inode until it next restarts.
|