642aa6fe1fc292843eb60ea39dc99e22a0d9a708
8 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
fc38c4cb5d |
fix(marmot-interop): repair hunk count in skip-retry patch + harden preflight
Two harness-side bugs that made a broken wn silently look like a working one: 1. whitenoise-skip-unprocessable-retry.patch's hunk header declared `@@ -178,7 +178,23 @@` but the new side only has 22 lines. GNU patch bails with "malformed patch at line 26" and leaves the target file untouched. Fix the count to +178,22 so the patch actually applies. 2. setup.sh's patch-apply loop piped `patch` through `tee`, which masked patch's exit code behind tee's. A miscounted hunk therefore still touched the `.headless-patched-*` marker and the next run skipped the "patch" step entirely, so the resulting wn ran the stock 10-retry exponential backoff on every MlsMessageUnprocessable — ~17min per event, which pegs the per-account serial event processor and makes every later test timeout "just because". Check patch's real exit status; fail preflight loudly instead. Same class of bug was also hiding cargo build failures: the `cargo build ... | tee` pipeline reported success even when rustup couldn't fetch the toolchain manifest (503 from static.rust-lang.org) or cargo couldn't hit the crates.io index (503 from fastly). setup.sh then `info`-logged the expected binary paths and happily moved on — until the daemon launch tripped over `nohup: No such file or directory`. Now each cargo build runs in a 4-attempt retry loop that terminates only when the binary actually exists on disk, and fails preflight loudly if it never materialises. Matches the jitpack 503 retry behaviour already in place for `:cli:installDist`. https://claude.ai/code/session_018kSBco5VVfctkW6vAm7NzA |
||
|
|
d8d094db48 |
fix(marmot-interop): parse post-v0.2 wn --json shape in interactive harness
The `.result` wrapper fix landed in lib.sh's pollers and the headless tests
back in
|
||
|
|
e9df0155c1 |
feat(marmot): accept PrivateMessage commits (handshake ratchet + dispatch)
MDK-core (openmls) defaults group configuration to
`MIXED_CIPHERTEXT_WIRE_FORMAT_POLICY` — outgoing commits / proposals
ship as PrivateMessage. Quartz only handled PrivateMessage for
ContentType.APPLICATION; any handshake message came in via the
application ratchet, consumed the first application generation on
whatever sender sent an app message next, and then died with
`Generation 0 already consumed` the moment the receiver saw the real
PrivateMessage commit. That's why every A-side processing of B's
rename / promote / remove / leave commit timed out.
Fix:
1. SecretTree grows a parallel `handshakeKeyNonceForGeneration`
path, using the per-sender `handshakeSecret` / `handshakeGeneration`
the struct already tracked — with its own skipped-keys cache and
replay-detection set so handshake and application generations
never clash.
2. `MlsGroup.decrypt()` dispatches on the PrivateMessage's
`content_type` BEFORE consuming any ratchet — COMMIT and PROPOSAL
take the handshake path; APPLICATION stays on the existing
application path.
3. COMMIT bodies decode as `Commit || signature<V> || confirmation_tag<V>
|| padding` (RFC 9420 §6.3.1) and route into `processCommit` inline,
so the caller just sees the epoch advance — no new public API.
PROPOSAL still rejects (standalone proposals aren't used by
Marmot flows yet) with a clear message instead of a cryptic
generation-consumed error.
https://claude.ai/code/session_016kAxdp6ubB5CnF9URhCEzP
|
||
|
|
0b7b1353e8 |
fix(marmot-interop): broaden skip-retry patch to all MdkCoreError
Narrowing the previous patch to only MlsMessageUnprocessable /
MlsMessagePreviouslyFailed still leaves wn stuck retrying
\`MdkCoreError("Failed to decrypt message with any exporter secret")\`
— mdk surfaces that as an Err variant rather than the Unprocessable
result enum, so it took the full 10-attempt exponential backoff (~17
minutes) before giving up. That was enough to block later
decryptable commits and regress test 11 (leave group) from pass back
to fail.
mdk-core doesn't retry internally, so ANY Err it returns is terminal
by construction. Treating the whole \`MdkCoreError\` variant as
one-shot is therefore equivalent to "trust mdk's verdict" rather
than "permissive skip" — if mdk decides the message is bad, it will
stay bad.
Net in the headless harness: 6/13 pass (up from 5/13 with the
narrower patch and 1/13 baseline).
https://claude.ai/code/session_016kAxdp6ubB5CnF9URhCEzP
|
||
|
|
8ba424295c |
fix(marmot-interop): drop retries for provably-unprocessable MLS messages
mdk-core's `MessageProcessingResult::Unprocessable` means the received kind:445 is genuinely outside this member's decrypt horizon — it's from a pre-membership epoch, or an epoch we've retained past, and no amount of retrying will make it decryptable. wn's account event processor was running that down the same exponential retry ladder as a legitimate transient failure (10 attempts, 1s..512s backoff, ~17 minutes to exhaust), which in the interop harness meant every later decryptable commit — add-member, rename, leave — had to wait behind a queue of doomed retries and raced the 30s test timeout. Add a third wn patch, `whitenoise-skip-unprocessable-retry.patch`, that treats `MlsMessageUnprocessable` and `MlsMessagePreviouslyFailed` as terminal: log the warning once, skip the reschedule. Applied in preflight via `setup.sh`. https://claude.ai/code/session_016kAxdp6ubB5CnF9URhCEzP |
||
|
|
f08b010f50 |
fix(marmot,cli,interop): interop-compatible Marmot flows and harness correctness
Five protocol-level fixes and a batch of harness correctness fixes to get
the headless Marmot/Whitenoise interop harness from 1/13 to 5/13 passing
cleanly, with the remaining failures all rooted in wn's per-account
serial event-processor retry backoff (which drops undecryptable
pre-membership commits after several minutes) rather than amy behaviour.
quartz + commons
----------------
* MarmotGroupData: hold CURRENT_VERSION at 2. mdk-core (the Rust MLS
engine used by whitenoise-rs) strict-rejects v3 payloads with
`ExtensionFormatError("Trailing bytes in NostrGroupDataExtension")`
— our v3 welcomes and GCE commits never apply, so every cross-client
group flow dies at welcome processing. We still parse v3 happily on
the way in; we just don't emit it until mdk publishes the
forward-compat fix MIP-01 mandates.
* MarmotManager.updateGroupMetadata: MERGE extensions instead of
REPLACING. RFC 9420 §12.1.7 says GCE proposals blow away the old
extension list; callers that pass only [marmot_group_data] dropped
[required_capabilities], which peers then reject. Preserve every slot
except the one we're updating.
* MarmotManager.createGroup: new optional `initialMetadata` parameter
that bakes MarmotGroupData into epoch-0 GroupContext.extensions
directly. Without it, creators had to publish a pre-membership
"bootstrap" commit that no later joiner could decrypt — each such
peer then burned their retry budget on an undecryptable kind:445
before seeing the real state. Threaded through MlsGroup.create /
MlsGroupManager.createGroup.
* MarmotManager.mlsGroupIdHex: new translation helper so any code
juggling the MIP-01 nostr_group_id (what amy indexes on) and the
MLS GroupContext groupId (what mdk indexes on) can cross-reference
them without reaching into MlsGroupManager directly.
amethyst module
---------------
* Account.leaveMarmotGroup: self-demote before SelfRemove per MIP-01,
and promote a surviving member to admin first if the caller is the
sole admin (otherwise we'd throw "admin depletion"). Matches the
cli/GroupMembershipCommands.leave flow.
cli (amy)
---------
* Context.syncIncoming: don't advance `giftWrapSince` on empty polls
(so the first-ever sync doesn't bump the cursor past every
past-timestamped wrap we've ever been sent), subtract 2 days lookback
when filtering (NIP-59 randomWithTwoDays gift wraps can have any
createdAt in the last 48h), and only advance `groupSince` for groups
we actually received events for.
* Context.syncIncoming: after ingest, if any Welcome consumed a
KeyPackage, rotate and publish a fresh one immediately. MIP-00
requires this — a KP can only be welcomed once and leaving the
consumed one on relays just means later senders invite us with a
bundle we no longer have private keys for.
* Context.resolveGroupId: accept either nostr_group_id (amy's primary
key) or the MLS GroupContext groupId (what wn emits) on every verb
that takes a group id. Wired through GroupAdd/Remove/Leave/Metadata
/Read, Message send/list, and all the await* verbs so harness
scripts never have to juggle both forms for a single group.
* GroupCreateCommand: bake initial metadata into epoch 0 (see the
quartz change above). Dropped the now-redundant bootstrap commit
publish and tightened the JSON output to include `mls_group_id`.
* GroupMembershipCommands.leave: self-demote admin before SelfRemove;
promote an heir if we're the only admin, otherwise the GCE would
deplete admins and the leave aborts.
* MarmotIngest.ingestGiftWrap: unwrap the sealed-rumor layer. NIP-59
wrap is gift-wrap(kind:1059) → seal(kind:13) → rumor; the old code
only unwrapped once and then checked `inner.kind == 444`, which is
always false because inner is actually the seal. Unseal once more
before the Welcome check. This single fix is what unsticks every
amy-side Welcome ingestion.
wn harness patches + scripts
----------------------------
* whitenoise-defaults-env.patch: honour $WHITENOISE_DISCOVERY_RELAYS in
`Relay::defaults()` (release builds otherwise bake damus.io / primal
/ nos.lol into every new account's NIP-65 / Inbox / KeyPackage
lists, which breaks publishing and prevents the inbox subscription
plane from ever reaching an operational state in a sandbox).
* setup.sh: sleep 2s after amy's initial kind:30443 publish so
nostr-rs-relay has a chance to fsync before wn's first targeted
discovery query. Without it wn's `keys check` races the relay's
WAL flush and intermittently returns NotFound.
* lib.sh: peel wn's `{"result": …}` wrapper in `jq_group_id`,
`wait_for_invite`, `wait_for_message`, `wait_for_member`. Post-v0.2
wn `--json` output nests everything under `.result` (and
`groups invites[]` nests further under `.group.mls_group_id`) —
these helpers were still pattern-matching on the flat shape, so
they returned empty strings for a perfectly good response.
* tests-{create,manage,extras}.sh: track both group IDs per test
(amy's nostr + wn's MLS), pass each CLI the id it understands, and
bump the post-commit wait timeouts to 90–120s so wn's exponential
retry backoff has time to work through the pre-membership commits
it can't decrypt and get to the ones it can.
https://claude.ai/code/session_016kAxdp6ubB5CnF9URhCEzP
|
||
|
|
37515ec975 |
feat(marmot-interop): let wnd run in sandboxes without kernel keyring
Containers / CI often block the keyutils syscalls wnd uses by default to store secret keys (add_key returns EACCES, keyctl_search returns ENOSYS). The integration-tests feature ships a mock keyring store, but the stock wnd binary doesn't wire it up. Add a tiny opt-in patch. Changes: - headless/patches/whitenoise-mock-keyring.patch: if built with the integration-tests feature and \$WHITENOISE_MOCK_KEYRING is set, wnd calls Whitenoise::initialize_mock_keyring_store() before the normal init path. Harmless on production builds. - headless/setup.sh: apply both harness patches generically from a list; build wn/wnd with --features cli,integration-tests; export WHITENOISE_MOCK_KEYRING=1 alongside WHITENOISE_DISCOVERY_RELAYS when launching each daemon. - preflight swaps keyctl for patch in required tools — we no longer need a real session keyring. https://claude.ai/code/session_01M6dCKAF5Y1VyHGZPjzwDXq |
||
|
|
6213fb2197 |
feat(marmot-interop): patch wnd to honour \$WHITENOISE_DISCOVERY_RELAYS
Without this, wnd exits on startup with NoRelayConnections in any environment that can't reach wnd's baked-in public discovery relay set (nos.lol, relay.damus.io, …). The headless harness runs entirely on a loopback relay, so hitting the public internet is never OK. Changes: - headless/patches/whitenoise-discovery-env.patch: adds an env-var override at the top of DiscoveryPlaneConfig::curated_default_relays. When \$WHITENOISE_DISCOVERY_RELAYS is set (comma-separated) we use that list instead of the hardcoded public one. - headless/setup.sh: applies the patch in preflight (idempotent — checks for a .headless-discovery-patched marker), invalidates the previous wn/wnd build on first patch, rebuilds, and threads \$WHITENOISE_DISCOVERY_RELAYS=\$RELAY_URL into every start_daemon invocation. https://claude.ai/code/session_01M6dCKAF5Y1VyHGZPjzwDXq |