Files
ObsidianDragon/docs/daemon-startup-hardening.md
DanS d188a08db7 test(daemon): add F1/F2 process-lifecycle integration tests; fix clobbered start error
Links the real EmbeddedDaemon into the ObsidianDragonTests target (its deps were
already present) and adds two POSIX integration tests that exercise the actual
fork/exec/waitpid fixes headlessly:

- testExecFailureReported (F2): start() against a non-executable file must fail with a
  precise "not executable or wrong architecture" reason.
- testDaemonCrashDetected (F1): a short-lived child that exits abnormally is still
  detected (crash_count_ increments) while isRunning() is hammered from the test
  thread — a regression test for the reap race.

Writing the F2 test surfaced a real bug: start()'s failure branch called
setState(State::Error, "Failed to start dragonxd process"), and setState stores the
Error message into last_error_ — clobbering the precise message startProcess() had
just set, so getLastError()/the UI only ever saw the generic string. Fixed to pass the
preserved detail to setState, so the precise reason survives and now also reaches the
state callback (crash panel / status).

ctest 1/1, green including the two new integration tests.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-02 13:11:15 -05:00

550 lines
34 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Daemon Startup Hardening — Implementation Plan
Eight verified edge-case defects in how ObsidianDragon brings up (and watches) the
`dragonxd` daemon at launch. Each entry is a buildable fix: the defect (with exact
line references), the chosen approach, the call sites, a representative change, and how
to verify it.
- **Scope:** full-node startup path (`--lite` excludes the embedded daemon entirely).
- **Source:** line references are exact against branch `dev` @ `45b652f`.
- **Provenance:** findings verified by direct source read; each fix designed by an
independent agent grounded in the cited files, with a sequencing pass for ordering,
shared helpers, and merge conflicts.
**Severity:** 2 High, 6 Medium · **Effort:** ≈ 2535 engineering-hours · **7 landing steps.**
Status legend: ☐ not started · ◐ in progress · ☑ landed & verified
**Status: all 8 landed & verified** (build-clean, `ctest` green after each) across four commits on
`dev` — lifecycle cluster (F1/F2/F4), filesystem+params cluster (F7/F6/F5), F3, and F8. Six new
pure-helper unit tests added.
**Wrap-up done:** release notes added (`CHANGELOG.md`, F8 breaking change front and center); i18n
back-fill applied additively to `res/lang/*.json` (42 keys — all 6 for es/de/fr/pt/ru; 6 zh/ja/ko
entries whose glyphs aren't in the current `NotoSansCJK-Subset.ttf` were left on English fallback
rather than render as tofu).
**Still owed before release:** a **CJK subset-font rebuild** (`scripts/build_cjk_subset.py`, needs
the Noto CJK source font) to cover the 6 deferred zh/ja/ko strings. *(F1 and F2 now have headless
integration-test coverage — see the progress log — so their GUI repros are optional, not blocking.)*
---
## Recommended rollout sequence
A real dependency order, not a checklist. The daemon-lifecycle cluster lands first
because it makes the `State::Error` / `crash_count_` contract trustworthy — which the
connect-stall panel and the lock gate both build on. The filesystem cluster lands
around a single shared helper. The connectivity-breaking security flip lands last.
| Step | Finding(s) | Site | Why here | Status |
|------|-----------|------|----------|--------|
| 1 | **F1** | `embedded_daemon.cpp` · `isRunning()` | Smallest/highest-severity; establishes the reliable Error/crash-count transition steps 3 & 6 depend on. | ☑ |
| 2 | **F2** | `embedded_daemon.cpp` · `startProcess()` | Same file family, different function; test the F1+F2 pair together with `kill -SEGV` / bad-binary repros. | ☑ |
| 3 | **F4** | `embedded_daemon.cpp` · `start()` | After F1/F2 so crash-count semantics are settled; its bail deliberately stays out of the crash path. | ☑ |
| 4 | **F7** | `util/platform` · `connection.cpp` | Structural owner of the fs-error idiom + `ConnectionConfig` that F5/F6/F8 reuse. | ☑ |
| 5 | **F6 + F5** | `app.cpp` · `verifySaplingParams()` | Same `startEmbeddedDaemon` / `verifySaplingParams` block; land together. | ☑ |
| 6 | **F3** | `app.cpp` · `renderLoadingOverlay()` | After F1 — panel is guarded off during `State::Error` (owned by the crash-count hint). | ☑ |
| 7 | **F8** | `connection.cpp` · `tryConnect()` | Largest; only connectivity-breaking default flip — land last, with release notes. | ☑ |
---
## F1 — Double-`waitpid` race can swallow a daemon crash
**Severity:** High · **Effort:** S (~12h) · **Status:** ☑ landed & verified
### The defect
`EmbeddedDaemon::isRunning()` (`embedded_daemon.cpp:1136`, POSIX branch) calls
`waitpid(WNOHANG)` — from the **UI thread, nearly every frame** — racing
`monitorProcess()`'s own reap at `:1244`. `waitpid` is one-shot: if the UI thread wins,
the monitor never decodes the exit, so `crash_count_` never increments, `State::Error`
never fires, and the 3-strike auto-restart cap (`app_network.cpp:479`) is defeated. The
sibling `XmrigManager::isRunning()` (`xmrig_manager.cpp:512`) already fixed exactly this
with an atomic read.
### The fix
Make `isRunning()` read the existing `std::atomic<State> state_` (member at
`embedded_daemon.h:253`) instead of calling `waitpid`, leaving `monitorProcess()` as the
sole reaper. Predicate is `Running || Stopping``Stopping` must stay "alive" because
`stop()`'s graceful/SIGTERM wait loops poll `isRunning()` before the process has exited.
### Files touched
- `src/daemon/embedded_daemon.cpp``isRunning()`, POSIX branch (~1136)
### Core change
```cpp
bool EmbeddedDaemon::isRunning() const // POSIX branch
{
// Read the atomic state_ instead of waitpid() — monitorProcess() is the
// sole reaper. Previously both threads reaped; if the UI thread won, the
// monitor never saw the exit (crash_count_ / exit code / Error all lost).
if (process_pid_ <= 0) return false;
State s = state_.load(std::memory_order_relaxed);
// Stopping stays "alive": stop()'s wait loops poll isRunning() while
// state_ == Stopping, before the process has actually terminated.
return (s == State::Running || s == State::Stopping);
}
```
### Verification
- Manual: `kill -SEGV` the daemon 1020×; the monitor must report the exit and increment `crash_count_` every time (previously intermittent).
- Regression: a normal Settings-driven stop still escalates SIGTERM→SIGKILL (the `Stopping` predicate).
- Not unit-testable (real fork/exec/waitpid) — consistent with the no-process-spawn harness.
### Dependencies
Mirrors `XmrigManager::isRunning()`. Flags a separate latent hazard (out of scope):
`stop()`'s final blocking `waitpid` (`:1220`) can still race a mid-sleep monitor
iteration — file as its own ticket.
---
## F2 — exec-after-fork silent failure: "Running" for a daemon that never started
**Severity:** High · **Effort:** S (~23h) · **Status:** ☑ landed & verified
### The defect
In `startProcess()` (`embedded_daemon.cpp:9571061`, POSIX) the parent runs
`process_pid_ = pid; return true;` **unconditionally** after `fork()` — with no
exec-status handshake. On a non-executable / wrong-arch / corrupt binary the child's
`execv` fails and it `_exit(127)`s, but `start()` has already set `State::Running`
(`:565`). The real cause never reaches `last_error_`; it surfaces later, generically,
as "exited unexpectedly (exit code 127)".
### The fix
Add a **close-on-exec self-pipe** handshake — `pipe() + fcntl(FD_CLOEXEC)`, deliberately
**not** `pipe2()` (macOS lacks it; the POSIX branch is shared). The child writes `errno`
only on `execv` failure; a successful exec closes the write end for free. Parent reads:
EOF ⇒ success; 4 bytes ⇒ reap the zombie, set a precise `last_error_` ("not executable
or wrong architecture"), and return `false` so `start()` never reports Running. EINTR-safe
on both ends. Also comments the unchecked parent-side `setpgid` at `:1053`.
### Files touched
- `src/daemon/embedded_daemon.cpp``startProcess()` parent read path
- `src/daemon/embedded_daemon.cpp` — child `execv`-failure write (~1043)
- `src/daemon/embedded_daemon.cpp``setpgid` best-effort comment (~1053)
### Core change
```cpp
// Self-pipe exec handshake (pipe()+FD_CLOEXEC; NOT pipe2 — macOS lacks it).
int execpipe[2]; pipe(execpipe);
fcntl(execpipe[0], F_SETFD, FD_CLOEXEC);
fcntl(execpipe[1], F_SETFD, FD_CLOEXEC);
pid_t pid = fork();
if (pid == 0) { // child
close(execpipe[0]);
/* setpgid / chdir / dup2 / argv … */
execv(binary_path.c_str(), argv.data());
int e = errno; // execv failed
while (write(execpipe[1], &e, sizeof e) < 0 && errno == EINTR) {}
_exit(127);
}
close(execpipe[1]); // parent: must close or read() never EOFs
int child_errno = 0, total = 0;
for (;;) { // EOF ⇒ exec ok; 4 bytes ⇒ exec failed
ssize_t n = read(execpipe[0], (char*)&child_errno + total, sizeof(int) - total);
if (n == 0) break;
if (n < 0) { if (errno == EINTR) continue; break; }
if ((total += n) >= (int)sizeof(int)) break;
}
close(execpipe[0]);
if (total >= (int)sizeof(int)) { // exec never happened
waitpid(pid, nullptr, 0); // reap the zombie
last_error_ = "dragonxd could not be executed: " +
std::string(strerror(child_errno)) +
" — not executable or wrong architecture";
return false; // start() no longer reports Running
}
```
### Verification
- Point at a `chmod -x` / wrong-arch file → `start()` returns false immediately, precise message, no leftover zombie.
- Success path: real binary still starts with no perceptible added latency.
- Optional pure `formatExecFailureError(errno)` helper for a `test_phase4.cpp` unit test.
### Dependencies
F1 (same function family; sequence F1→F2). **Highest-risk mistake:** forgetting
`FD_CLOEXEC` makes every successful start hang the parent read forever.
---
## F4 — Stale datadir-lock start → restart storm that wedges the UI
**Severity:** Medium · **Effort:** S (~35h) · **Status:** ☑ landed & verified
### The defect
`start()` (`embedded_daemon.cpp:466`) gates only on the RPC port (`:482`), never on
`isDaemonProcessRunning()` (`:1292`). A graceful shutdown frees the port but keeps the
datadir `.lock` for up to ~90s. A rapid stop→start spawns a daemon that dies "Cannot
obtain a lock on data directory" — routed to the generic crash path. With a ~4s retry
cadence, **three lock races in ~12s exhaust the 3-strike budget** and wedge the UI long
before the lock actually clears.
### The fix
Fail-fast with a **short bounded local wait (~300ms), not a 90s block**. After the port
bail, consult `isDaemonProcessRunning()` — gated by `!skip_port_check_` and exempt when
`override_datadir_` is set, so the isolated migrate-to-seed daemon still works. A pure
`evaluateDatadirLockGate()` returns a **distinct non-crash Error** that never increments
`crash_count_`. The connect loop's own retry then absorbs the transient.
### Files touched
- `src/daemon/embedded_daemon.h` — decision struct, helper decl, poll constants
- `src/daemon/embedded_daemon.cpp``start()` gate + `evaluateDatadirLockGate()`
### Core change
```cpp
static StartLockGateDecision evaluateDatadirLockGate(
bool skipPortCheck, bool isolatedOverride, bool stillRunningAfterWait) {
if (skipPortCheck || isolatedOverride) return {true, ""}; // migrate-to-seed exempt
if (!stillRunningAfterWait) return {true, ""};
return {false, "A previous dragonxd is still shutting down and holding the "
"data directory lock. Retrying shortly…"};
}
// start() — after the isPortInUse() bail, before setState(Starting):
if (!skip_port_check_ && override_datadir_.empty()) {
bool stillLocked = false; // ~300ms bounded wait, NOT ~90s
for (int i = 0; i < kDatadirLockWaitMaxPolls; ++i) {
if (!isDaemonProcessRunning()) { stillLocked = false; break; }
stillLocked = true;
std::this_thread::sleep_for(std::chrono::milliseconds(kDatadirLockWaitPollMs));
}
auto gate = evaluateDatadirLockGate(false, false, stillLocked);
if (!gate.proceed) { setState(State::Error, gate.errorMessage); return false; }
}
```
### Verification
- Unit: `evaluateDatadirLockGate()` across the skip / isolated / still-running matrix.
- Manual: rapid restart into a lingering lock → distinct message, no crash-cap wedge.
- Migrate-to-seed second daemon still starts (isolated exemption).
### Dependencies
F1/F2 (must not touch `crash_count_`; wording must not collide with the monitor's
"exited unexpectedly"). Same TU, different function.
---
## F5 — Extraction / copy write-failures never surfaced up front
**Severity:** Medium · **Effort:** S (~23h) · **Status:** ☑ landed & verified
### The defect
`startEmbeddedDaemon()` discards `extractEmbeddedResources()`'s `bool` return
(`app.cpp:4152`) and the second copy-fallback loop drops `copy_file`'s `error_code`
entirely (`:4236`). Only Sapling params **existence** is re-checked — never the daemon
binary/CLI/tx/asmap. A disk-full or truncated `dragonxd` write falls straight through to
spawn and fails opaquely. The innermost write already returns `false`
(`embedded_resources.cpp:307`) — the signal is simply thrown away.
### The fix
Minimal, surgical wiring — no new abstraction. Capture the extraction return and, on
failure, set `daemon_status_ = TR("sb_daemon_extract_failed")` and `return false` before
spawning. In the second copy loop, check `ec` after each `copy_file`, track `copyFailed`,
and abort with a dir-parameterized `sb_daemon_files_failed`. An **absent source** stays
fine (optional files); only an actual `error_code` counts. Written so F6/F7 slot in later
without re-touching this control flow.
### Files touched
- `src/app.cpp``startEmbeddedDaemon()` extraction check (~4152)
- `src/app.cpp` — second copy-fallback loop (~42104242)
- `src/util/i18n.cpp` + `res/lang/*.json` — 2 additive keys
### Core change
```cpp
// stop discarding the extraction result (~4152)
if (!resources::extractEmbeddedResources()) {
daemon_status_ = TR("sb_daemon_extract_failed"); // disk full / permission denied
return false; // abort before spawning
}
// second copy-fallback loop — was dropping ec entirely (~4236)
bool copyFailed = false;
for (const char* name : { "asmap.dat", "dragonxd", "dragonx-cli", "dragonx-tx" }) {
fs::path dst = fs::path(daemon_dir) / name;
if (fs::exists(dst)) continue; // already present — skip
for (const auto& dir : searchDirs) {
fs::path src = fs::path(dir) / name;
if (!fs::exists(src)) continue; // absent source is OK, not a failure
fs::copy_file(src, dst, ec);
if (ec) { copyFailed = true; ec.clear(); }
break;
}
}
if (copyFailed) {
char buf[512];
snprintf(buf, sizeof buf, TR("sb_daemon_files_failed"), daemon_dir.c_str());
daemon_status_ = buf;
return false; // don't fall through to spawn
}
```
### Verification
- Unit: `extractEmbeddedResources()` returns false without embedded resources.
- Extract the copy loop into a testable helper; force one dst write to fail (dst is an existing directory).
- Manual: near-full tmpfs / read-only dir → clear status, daemon controller never constructed.
### Dependencies
Shares the `daemon_status_` surfacing convention with F6; its early-return pattern is the
template F7 matches. Open item: remove truncated dst files so a retry re-copies.
---
## F6 — Sapling params validated by existence/size only, never hashed
**Severity:** Medium · **Effort:** S (~35h) · **Status:** ☑ landed & verified
> **As-built note.** `verifySaplingParams()` now delegates to a public, injectable
> `verifySaplingParamsIn(dir, digests)` so the integrity + marker-cache logic is unit-testable
> with synthetic small files (the real 48 MB params aren't in the repo). i18n keys for F5 were
> added to `i18n.cpp` (English source of truth); the `res/lang/*.json` back-fill via
> `scripts/add_missing_translations.py` is deferred to a single run at the end of the batch,
> per the cross-cutting note. Non-English locales fall back to English until then.
### The defect
`verifySaplingParams()` (`connection.cpp:123`) only calls `fs::exists()`;
`resourceNeedsUpdate()` (`embedded_resources.cpp:250`) is size-only. On Linux (no
embedded resources) a **truncated-but-present** param passes and is handed to the daemon,
which then fails to build shielded proofs mid-operation — far from the real cause.
### The fix
Add a pinned `{ filename → size, sha256 }` table (one source of truth, cross-referenced
to `scripts/build-lite-backend-artifact.sh`) and hash-check each param after the
existence check, reusing the existing `util::sha256Hex` (no second implementation). Since
these are ~48 MB, **cache the result** via a `.sapling_verified` marker keyed on
`size:mtime` — re-hash only when the stat line changes, so startup isn't slowed.
### Files touched
- `src/rpc/connection.h``verifySaplingParams` decl
- `src/rpc/connection.cpp` — digest table, marker helpers, rewrite
### Core change
```cpp
// connection.cpp — pinned known-good digests
// (source of truth: scripts/build-lite-backend-artifact.sh ensure_sapling_params)
constexpr SaplingParamDigest kSaplingParamDigests[] = {
{ "sapling-spend.params", 47958396, "8e48ffd2…efc13" },
{ "sapling-output.params", 3592860, "2f0ebbcb…fb0e4" },
};
bool Connection::verifySaplingParams() {
// existence check (unchanged) …
// cache: skip re-hashing a ~48 MB file unless size:mtime changed
if (readMarkerMatches(marker, statLines)) return true;
for (auto& d : kSaplingParamDigests)
if (util::sha256Hex(bytes) != d.sha256) return false; // reuse existing helper
writeMarker(marker, statLines);
return true;
}
```
### Verification
- Unit: good params pass; truncated / wrong-bytes rejected; marker cache short-circuits re-hash unless size/mtime changed. Real temp-file fixtures (matches existing `sha256Hex` tests).
### Dependencies
F7 (reuse fs-error idiom; shares the `startEmbeddedDaemon`/`verifySaplingParams` block).
Third caller of the existing `util::sha256Hex`.
---
## F7 — Directory-create errors universally ignored on the daemon-env path
**Severity:** Medium · **Effort:** S (~34h) · **Status:** ☑ landed & verified
> **As-built notes.** Two deviations from the original design, both confirmed against the code:
> (1) `embedded_resources.cpp:270` already checks its `error_code` and returns `false` on failure — it was **not** a bug, so it is left untouched.
> (2) Of the four `autoDetectConfig` callers, only the primary connect path (`app_network.cpp:243`) was wired to check `dir_error`; the other three degrade gracefully on their own — `app.cpp:4306` and `app_wizard.cpp:912` are stop paths that already gate on empty creds, and `settings_page.cpp:434` is read-only display. `dir_error` is set by `autoDetectConfig`, so they can be wired later if desired.
### The defect
Five startup directory-create sites either drop the `error_code` or use the throwing
overload with no `catch`: `main.cpp:730`, `connection.cpp:216` (can throw **uncaught**
through its callers), `embedded_resources.cpp:270`, `app.cpp:4172`/`4218`. A read-only
home or permission-denied yields a confusing "conf missing" / "binary not found"
downstream — or an uncaught `filesystem_error` — instead of a clear cause.
### The fix
One shared, non-throwing `Platform::ensureDirectory(dir, outError)` in
`util/platform.{h,cpp}` that produces a single consistent message. Replace all five
sites; `autoDetectConfig()` moves off the throwing overload and sets a new
`ConnectionConfig::dir_error` that its four callers check and bail on. This is the
**structural owner** of the fs-error idiom that F5 and F6 reuse.
### Files touched
- `src/util/platform.h` / `.cpp``ensureDirectory()`
- `src/rpc/connection.h` / `.cpp``dir_error` + `autoDetectConfig`
- `main.cpp`, `app.cpp`, `app_network.cpp`, `app_wizard.cpp`, `settings_page.cpp`, `embedded_resources.cpp` — 5 sites + 4 callers
- `tests/test_phase4.cpp``TestPlatformEnsureDirectory`
### Core change
```cpp
// util/platform.cpp — one shared, non-throwing helper
bool Platform::ensureDirectory(const std::string& dir, std::string* outError) {
std::error_code ec;
if (std::filesystem::is_directory(dir, ec)) return true;
ec.clear();
std::filesystem::create_directories(dir, ec);
if (ec) {
if (outError)
*outError = "Cannot create " + dir + ": " + ec.message() +
". Check permissions / free space.";
return false;
}
return true;
}
// Replaces 5 ad-hoc sites; autoDetectConfig() now sets ConnectionConfig::dir_error,
// and its 4 callers bail on it.
```
### Verification
- Unit `TestPlatformEnsureDirectory`: existing dir → true; fresh nested → created; POSIX unwritable → false + message.
- All four `autoDetectConfig` callers tolerate `dir_error`. Pre-App-init site (main.cpp) reports via stderr / MessageBox.
### Dependencies
**Owns** `Platform::ensureDirectory` (used by F5, F6) and the `ConnectionConfig`
extension (coordinated with F8). Land before F5/F6/F8.
---
## F8 — Plaintext-remote RPC credential transmission is warn-only
**Severity:** Medium · **Effort:** M (~69h) · **Status:** ☑ landed & verified
> **⚠️ RELEASE NOTES REQUIRED — breaking default flip.** A wallet configured to talk to a
> **remote** `rpchost` over **plain HTTP** (no `rpctls=1`) will now be **refused** at connect
> time instead of warned. Affected users must add **`rpcallowplaintext=1`** to `DRAGONX.conf`
> (or switch to `rpctls=1`) to reconnect. Local/embedded daemons (`127.0.0.0/8`, `localhost`,
> `::1`) are unaffected. Call this out prominently in the release notes.
>
> **As-built note.** Shipped the security-complete core: `isLocalHost` tightened to exact
> loopback (`isExactIPv4Loopback` — `127.evil.com` no longer passes), refuse-by-default in
> `tryConnect`, and the `rpcallowplaintext` conf-key opt-in. The **Settings toggle UI was
> deferred** — the RPC section of `settings_page.cpp` is read-only display and a security
> toggle there is riskier surface; the conf-key opt-in fully covers recovery, and the refusal
> status/notification tells the user exactly what to add. The toggle can be added later
> (persist a `Settings` flag and OR it into `allowsPlaintextRemote`).
### The defect
A remote `rpchost` without `rpctls=1` sends Basic-auth `rpcuser:rpcpassword` over
cleartext HTTP. `tryConnect()` (`app_network.cpp:314`) only shows a **dismissible
warning** then proceeds — a local-network MITM sees the credentials. Compounding it,
`isLocalHost()`'s naive `rfind("127.",0)==0` misclassifies `127.evil.com` as local,
suppressing even the warning.
### The fix
Change the policy to **refuse-by-default with an explicit, persisted opt-in** — a
`rpcallowplaintext=1` conf key (for hand-editors) and a Settings toggle. Block the
connect and show a **blocking modal** explaining the risk and how to enable TLS or opt
in; localhost is unaffected. Tighten `isLocalHost()` to exact `127.x.y.z` / `::1` /
`localhost` via `isExactIPv4Loopback()`. **Back-compat:** default off ⇒ existing remote
users hit a hard stop until they opt in — **ship with prominent release notes.**
### Files touched
- `src/rpc/connection.h` / `.cpp``isLocalHost`, `allow_plaintext_remote`, `parseConfFile`
- `src/config/settings.h` / `.cpp` — persisted opt-in
- `src/app_network.cpp`, `src/app.h` — refuse + modal dispatch
- `src/ui/windows/plaintext_remote_rpc_dialog.h` — new blocking modal
- `src/ui/pages/settings_page.cpp` — toggle UI
### Core change
```cpp
// Tightened loopback test — "127.evil.com" is NOT local
bool Connection::isLocalHost(const std::string& host) {
std::string h = stripBrackets(lowercase(host));
return h == "localhost" || h == "::1" || isExactIPv4Loopback(h); // exact 127.x.y.z
}
// Refuse-by-default with an explicit, persisted opt-in
const bool plaintextRemote = rpc::Connection::usesPlaintextRemote(config);
const bool plaintextAllowed = config.allow_plaintext_remote // rpcallowplaintext=1
|| settings_.getAllowPlaintextRemoteRpc(); // Settings toggle
if (plaintextRemote && !plaintextAllowed) {
connection_status_ = TR("sb_plaintext_remote_blocked");
showPlaintextRemoteRpcDialog(config.host + ":" + config.port); // blocking modal
return; // no creds sent
}
```
### Verification
- Unit: `isLocalHost``127.evil.com` false, `127.0.0.1`/`::1`/`localhost` true; `allowsPlaintextRemote` honors conf key + settings flag.
- Manual: remote plaintext blocked; modal fires; opt-in persists across restart.
### Dependencies
F7 (second extender of `ConnectionConfig`/`parseConfFile`; land after so the struct grows
once). Wire `renderPlaintextRemoteRpcDialog` into the app modal-dispatch list.
---
## Shared helpers & coordination points
| Helper | Purpose | Used by |
|--------|---------|---------|
| `Platform::ensureDirectory()` | Single non-throwing directory-create with one consistent message; replaces five ad-hoc sites. Owned by F7. | F7, F5, F6 |
| `ConnectionConfig` extension | Coordination point, not a function: F7 adds `dir_error`, F8 adds `allow_plaintext_remote`. Land F7→F8 so it grows once per step. | F7, F8 |
| `util::sha256Hex` *(existing)* | Already-compiled, curl-free SHA-256. F6 becomes its third caller — no second hash routine. | F6 |
| `connectHasStalled()` *(new, pure)* | Stall predicate split out of the ImGui/App code for unit testing, per the `*_updater_core.cpp` precedent. | F3 |
| `evaluateDatadirLockGate()` *(new, pure)* | Lock-gate decision as `{proceed, message}` from three booleans — unit-testable without real process/fs I/O. | F4 |
## F3 — Unbounded connect spinner (deferred to step 6)
**Severity:** Medium · **Effort:** S (~35h) · **Status:** ☑ landed & verified
> **As-built note.** `renderLoadingOverlay()` is a pure draw-list overlay with **no interactive
> widgets** (the existing crash case at ~5289 already communicates via guidance *text*, relying on
> the sidebar staying reachable). So rather than inject `ActionButton`s — which would fight the
> non-interactive overlay — the stall notice follows that same idiom: a "Taking longer than
> expected" title + a reassuring body (with elapsed seconds) + a full-node-gated hint ("Open
> Settings → Restart Daemon, or check the Console"). This let me drop the planned
> `WalletState::connect_stalled` flag too: the stalled state is computed locally in the overlay
> from `connect_stall_since_`, so the only new member is `App::connect_stall_since_`.
The connect loop retries forever while `!state_.connected` (`app.cpp:1239`);
`loading_timer_` only animates the spinner. Stamp `connect_stall_since_` when
"reachable but not ready" is first seen; a pure `connectHasStalled()` helper (new
`util/connect_stall.h`, default 45s from `ui.toml`) flips `state_.connect_stalled` at
threshold, and `renderLoadingOverlay()` shows a "Taking longer than expected" panel with
Retry / Restart daemon / Open console (full-node gated). The background retry keeps
firing — recovery clears the panel automatically. Guarded off while the daemon is in
`State::Error` (owned by F1's crash-count hint). Full detail lives in the sequencing/
design record; see the shared-helper table above.
---
## Cross-cutting notes
- **One TU, three functions.** `embedded_daemon.cpp` is edited by F1 (`isRunning`),
F2 (`startProcess`) and F4 (`start`) — no literal hunk overlap, but land in order to
keep "monitorProcess is the sole reaper" coherent.
- **Connection struct grows twice.** `connection.h/.cpp` is touched by F6, F7 and F8;
F7 and F8 both extend `ConnectionConfig` and `parseConfFile` — highest collision risk.
Sequence F7→F6→F8.
- **Testability split.** The three new pure predicates all get `tests/test_phase4.cpp`
coverage. F1/F2's fork/exec/waitpid changes are **not** unit-testable — they rely on
manual `kill` / non-executable-binary repros, consistent with the no-process-spawn harness.
- **i18n is additive-only.** Add each finding's English keys to `strings_`, then run
`scripts/add_missing_translations.py` **once at the very end**
(`json.dump indent=4, sort_keys=True, ensure_ascii=False`) — never bulk-regenerate a
`res/lang/*.json`.
- **F8 is a breaking default flip.** Refuse-plaintext-by-default stops existing
remote-RPC users cold until they opt in. Lands last, gated behind a persisted opt-in,
with release notes calling out the new `rpcallowplaintext` key and the Settings toggle.
- **Latent hazard, out of scope.** F1 surfaces (but doesn't fix) a second
double-`waitpid` window between `stop()`'s final blocking reap (`:1220`) and a
mid-sleep monitor iteration — file it as its own ticket.
---
## Progress log
- **F1/F2 integration tests** — ☑ added `testExecFailureReported` (F2) and `testDaemonCrashDetected` (F1) to `test_phase4.cpp`, driving the **real** `EmbeddedDaemon` fork/exec/waitpid code headlessly (POSIX; required linking `embedded_daemon.cpp` into the test target — its deps were already there). The F1 test hammers `isRunning()` from the test thread while the child exits, so it's a genuine regression test for the reap race. **The F2 test caught a real bug:** `start()`'s failure branch overwrote `startProcess()`'s precise `last_error_` ("…not executable or wrong architecture") with a generic "Failed to start dragonxd process" (because `setState(Error, …)` stores its message into `last_error_`), so the precise reason never reached `getLastError()`/the UI — **fixed** to preserve the detail (now also surfaced via the state callback / crash panel). Build-clean; `ctest` 1/1.
- **F1** — ☑ landed: `isRunning()` (POSIX) now reads the atomic `state_` (predicate `Running || Stopping`) instead of calling `waitpid`, leaving `monitorProcess()` the sole reaper. Clean build (all targets link); `ctest` 1/1 passing. Not unit-testable — needs the manual `kill -SEGV` repro before release.
- **F2** — ☑ landed: `startProcess()` (POSIX) now creates a `FD_CLOEXEC` self-pipe before `fork()`; the child writes `errno` to it on `execv` failure, the parent reads EOF-vs-errno and, on failure, reaps the zombie + sets a precise `last_error_` ("not executable or wrong architecture") + returns `false` (so `start()` no longer reports `Running` for a daemon that never started). Parent-side `setpgid` is now best-effort with a `DEBUG_LOGF` on failure. Clean build; `ctest` 1/1 passing. Not unit-testable — needs the manual non-executable / wrong-arch-binary repro before release.
- **F8** — ☑ landed: `isLocalHost()` tightened to exact loopback via `isExactIPv4Loopback` (a `127.`-prefixed *hostname* like `127.evil.com` is no longer misclassified as local). `tryConnect()` now **refuses** a plaintext connection to a remote host instead of warn-and-proceeding — a local-network MITM can no longer capture `rpcuser:rpcpassword` — unless the user opts in with `rpcallowplaintext=1` in `DRAGONX.conf` (new `ConnectionConfig::allow_plaintext_remote` + `allowsPlaintextRemote()` policy). The refusal surfaces via status line + a one-time notification. New `testIsLocalHost` (12 assertions) + `testAllowsPlaintextRemote` (5). Clean build; `ctest` 1/1 passing. **Breaking — needs release notes; Settings-toggle UI deferred (see as-built note).**
- **F3** — ☑ landed: the connect loop now stamps `connect_stall_since_ = ImGui::GetTime()` the moment the daemon first goes "reachable but not ready" (warmup branch + `applyDaemonInitStatus`), and clears it in `onConnected` / `onDisconnected` / warmup-complete — all in `app_network.cpp`. The pure `util::connectHasStalled(stallSince, now, threshold)` helper (new `util/connect_stall.h`, default 45 s from `ui.toml`) drives a draw-list "Taking longer than expected" notice in `renderLoadingOverlay()` (title + elapsed-seconds body + full-node hint), guarded off while the daemon is in `State::Error`. Background retry continues, so the notice self-clears on connect. New `testConnectHasStalled` unit test (7 assertions). Clean build; `ctest` 1/1 passing. (Draw-list text, not buttons — see as-built note above.)
- **F6** — ☑ landed: `verifySaplingParams()` now hash-verifies each Sapling param against its pinned canonical SHA-256 (from `build-lite-backend-artifact.sh`), replacing the existence-only check, so a truncated/corrupt-but-present param is rejected instead of failing later on a shielded op. A `<params_dir>/.sapling_verified` marker keyed on `size:mtime` skips re-hashing ~48 MB on every startup. Logic extracted to the injectable `verifySaplingParamsIn(dir, digests)`; new `testVerifySaplingParams` unit test (valid / marker fast-path / wrong-hash / truncated / missing). Clean build; `ctest` 1/1 passing.
- **F5** — ☑ landed: `startEmbeddedDaemon()` now checks `extractEmbeddedResources()`'s return (abort with `sb_daemon_extract_failed` on failure) and the previously-dropped `copy_file` `error_code` in the daemon-binary fallback loop (abort with `sb_daemon_files_failed` incl. the dir), so a disk-full / truncated `dragonxd` write is surfaced up front instead of failing opaquely at spawn. An absent source file stays non-fatal. Two i18n keys added to `i18n.cpp`. Clean build; `ctest` 1/1 passing.
- **F7** — ☑ landed: new non-throwing `Platform::ensureDirectory(dir, outError)` in `util/platform.{h,cpp}` with one consistent message. Replaces the unchecked/throwing directory-create sites at `main.cpp:730` (pre-init: now logs + `MessageBoxA` on Windows + `return 1`), `connection.cpp:216` (autoDetectConfig now uses the ec overload — **no more uncaught `filesystem_error`** — and sets the new `ConnectionConfig::dir_error`), and both `app.cpp` daemon-dir sites (surface via `daemon_status_` + `return false`). Primary connect path (`app_network.cpp:243`) checks `dir_error` and bails to the status line instead of mislabelling it "waiting for config". `embedded_resources.cpp:270` left as-is (already correct). New `testPlatformEnsureDirectory` unit test (existing-dir / fresh-nested / empty / parent-is-file). Clean build; `ctest` 1/1 passing.
- **F4** — ☑ landed: `start()` now gates on a lingering datadir lock after the port bail. When `!skip_port_check_ && override_datadir_.empty()`, it polls `isDaemonProcessRunning()` with a bounded ~300 ms wait (3 × 100 ms, breaks early), then a pure header-inline `evaluateDatadirLockGate()` decides: if a sibling `dragonxd` is still alive it bails with a distinct **non-crash** `State::Error` ("…holding the data directory lock. Retrying shortly…") that never touches `crash_count_`, so the 3-strike cap can't trip; the connect loop's retry resumes once the lock clears. Isolated migrate-to-seed starts are exempt. New `testDatadirLockGate` unit test (5 assertions, proceed/bail/2× exempt) added to `test_phase4.cpp`. Clean build; `ctest` 1/1 passing.