Two harness-side bugs that made a broken wn silently look like a working one:
1. whitenoise-skip-unprocessable-retry.patch's hunk header declared
`@@ -178,7 +178,23 @@` but the new side only has 22 lines. GNU patch
bails with "malformed patch at line 26" and leaves the target file
untouched. Fix the count to +178,22 so the patch actually applies.
2. setup.sh's patch-apply loop piped `patch` through `tee`, which
masked patch's exit code behind tee's. A miscounted hunk therefore
still touched the `.headless-patched-*` marker and the next run
skipped the "patch" step entirely, so the resulting wn ran the
stock 10-retry exponential backoff on every MlsMessageUnprocessable
— ~17min per event, which pegs the per-account serial event
processor and makes every later test timeout "just because".
Check patch's real exit status; fail preflight loudly instead.
Same class of bug was also hiding cargo build failures: the
`cargo build ... | tee` pipeline reported success even when rustup
couldn't fetch the toolchain manifest (503 from static.rust-lang.org)
or cargo couldn't hit the crates.io index (503 from fastly). setup.sh
then `info`-logged the expected binary paths and happily moved on —
until the daemon launch tripped over `nohup: No such file or directory`.
Now each cargo build runs in a 4-attempt retry loop that terminates
only when the binary actually exists on disk, and fails preflight
loudly if it never materialises. Matches the jitpack 503 retry
behaviour already in place for `:cli:installDist`.
https://claude.ai/code/session_018kSBco5VVfctkW6vAm7NzA
Narrowing the previous patch to only MlsMessageUnprocessable /
MlsMessagePreviouslyFailed still leaves wn stuck retrying
\`MdkCoreError("Failed to decrypt message with any exporter secret")\`
— mdk surfaces that as an Err variant rather than the Unprocessable
result enum, so it took the full 10-attempt exponential backoff (~17
minutes) before giving up. That was enough to block later
decryptable commits and regress test 11 (leave group) from pass back
to fail.
mdk-core doesn't retry internally, so ANY Err it returns is terminal
by construction. Treating the whole \`MdkCoreError\` variant as
one-shot is therefore equivalent to "trust mdk's verdict" rather
than "permissive skip" — if mdk decides the message is bad, it will
stay bad.
Net in the headless harness: 6/13 pass (up from 5/13 with the
narrower patch and 1/13 baseline).
https://claude.ai/code/session_016kAxdp6ubB5CnF9URhCEzP
mdk-core's `MessageProcessingResult::Unprocessable` means the received
kind:445 is genuinely outside this member's decrypt horizon — it's from
a pre-membership epoch, or an epoch we've retained past, and no amount
of retrying will make it decryptable. wn's account event processor was
running that down the same exponential retry ladder as a legitimate
transient failure (10 attempts, 1s..512s backoff, ~17 minutes to
exhaust), which in the interop harness meant every later decryptable
commit — add-member, rename, leave — had to wait behind a queue of
doomed retries and raced the 30s test timeout.
Add a third wn patch, `whitenoise-skip-unprocessable-retry.patch`,
that treats `MlsMessageUnprocessable` and `MlsMessagePreviouslyFailed`
as terminal: log the warning once, skip the reschedule. Applied in
preflight via `setup.sh`.
https://claude.ai/code/session_016kAxdp6ubB5CnF9URhCEzP