Step 1 of `2026-05-07-moq-relay-routing-investigation.md` ran a
5× sweep on `HangInteropTest` with per-test moq-relay trace
capture. Failure rate: 3/5, all in
`late_join_listener_still_decodes_tail`. Cross-referencing the
relay trace (`encoding self=Subscribe …catalog.json` on conn{id=0}
at T+2.07 s) with the speaker's `Log.d("NestTx")` lines (only
`ANNOUNCE inbound prefix=''` at T=0; NO `SUBSCRIBE inbound id=0`
in the failing window) shows the relay correctly forwards the
upstream SUBSCRIBE — the bidi just never reaches
`MoqLiteSession.handleInboundBidi`. So the prior plan's
"moq-relay 0.10.x per-broadcast subscribe-routing race"
hypothesis is wrong.
Re-scoped Priority 1 of the closure roadmap to point at
`:quic`'s peer-bidi surfacing path instead. The actual bug is
between QUIC's bidi-accept and `incomingBidiStreams()` flow.
Spotless-formatted Kotlin imports.
Trace artefacts preserved under
`nestsClient/plans/artefacts/2026-05-07-routing-race-disproven/`
so the next agent (QUIC owner) doesn't have to re-run the sweep.
8 plan files updated to reflect what's actually shipped vs. what
each doc said before:
- 2026-05-06-cross-stack-interop-test.md: 'Spec — ready to
implement' → 'Implemented and merged' with pointers to results,
gap matrix, and closure roadmap.
- 2026-05-06-cross-stack-interop-test-results.md: scenario
inventory updated to show all branches merged into
claude/cross-stack-interop-test-XAbYB; suite-flake caveats
flagged with cross-refs to the routing investigation; file
inventory now lists BrowserInteropTest + PlaywrightDriver +
the moved nestsClient/tests/browser-interop/ tree (no more 'in
sister branches not yet merged' section).
- 2026-05-06-cross-stack-interop-test-gap-matrix.md: T8/T10/T11/
T12/T13 all show hang ✅ + browser ✅; I13/I14 'NOT YET LANDED'
callouts removed; the 'in sister branch' qualifier on T13 / I7
/ I6 dropped; coverage-holes section replaced with caveat
pointers to the routing investigation.
- 2026-05-06-i4-stereo-cross-stack-scenario.md: 'Spec — ready for
implementation' → 'Landed' with PR #2755 + the three test
scenarios that ship the assertion.
- 2026-05-06-phase4-browser-harness.md: 'Spec — ready for
implementation' → 'Landed', with notes on how the implementation
diverged (used Container.Legacy.Consumer not @moq/watch; cert
pinning via serverCertificateHashes; I2/I3/I4/I13/I14 originally
deferred but later landed).
- 2026-05-06-phase4-browser-harness-results.md: status updated to
reflect 4.D CI removed per maintainer ask; layout note for
the nestsClient/tests/browser-interop/ relocation; full
scenario list reconciled with results doc.
- 2026-05-07-late-join-catalog-flake-investigation.md: status
updated to mention 00f6cba31 + 207057374 were reverted, and
link to 2026-05-07-moq-relay-routing-investigation.md as the
action plan. Adds note that the flake also affects four
browser-tier scenarios after Browser I7 landed.
- 2026-05-01-quic-stream-cliff-investigation.md: adds an HCgOY-
override banner — recommendation 2 (framesPerGroup = 5) was
superseded 4 days later by HCgOY field tests landing the prod
default at 50. Cross-refs the reconciliation + production-
rerun plans. Open follow-up #3 ('reset to 1') updated with
the same context.
No code or test changes.
The 600 ms speaker warmup and 250 ms hang-listen post-announced()
sleep were intended to give the relay more time to prime its
per-broadcast subscribe-pump. Sweep showed they actually made
things WORSE — combined 0/5 pass (down from 2/5 with the
single-subscribe fix alone) — because the cumulative ~850 ms
of pre-subscribe delay shrank the catalog-read window into the
speaker's tear-down region.
Lesson recorded: any mitigation that adds pre-subscribe delay
hurts more than it helps. Future attempts should target the
relay's subscribe-routing race directly (upstream fix, RUST_LOG
trace, or version bump) rather than smoothing it over with
delays.
Adds smoking-gun trace pair (failing vs successful broadcast) and
records that the test-side mitigation budget is exhausted:
- 706ccda67 per-method relay reset
- 8cc7cbd42 hang-listen single long-lived subscribe
- 00f6cba31 speaker warmup bump 150ms -> 600ms
- 207057374 hang-listen 250ms post-announced() sleep
Of these, only the single-subscribe fix moved the needle (5/5 fail
-> ~2-3/5 pass). The remaining flake is in moq-relay 0.10.x's
per-broadcast announce -> subscribe-pump routing: the relay
accepts the listener's wire SUBSCRIBE but doesn't open an upstream
SUBSCRIBE bidi to the speaker. Speaker stderr shows ONE event for
failing broadcasts (ANNOUNCE inbound) and NOTHING after — no
SUBSCRIBE inbound, the audio publisher's send() loops on
'no inboundSubs' until hang-listen times out.
Three next steps documented for upstream support:
1. Re-run with RUST_LOG=moq_relay=trace
2. File upstream bug at kixelated/moq
3. Try moq-relay > 0.10.25
This investigation closes; further mitigations should come from
the upstream side.
Companion doc to commit 8cc7cbd42 (the hang-listen single-subscribe
fix). Captures:
- Pre-fix root cause: moq-rs cancel cascade. Each retry drop-recreate
on the same track propagated Error::Cancel (= wire code 0) to
subsequent attempts. Sweep was 5/5 fail.
- Fix: hold ONE subscription open for the full 10 s catalog-read
budget. Inner timeouts on .next() poll for the first group; outer
timeout caps total wait. Eliminates the cancel-cascade. Sweep 2/5
pass post-fix.
- Residual root cause: unidentified. Same 3-second peer-cancel
pattern hits on ~60% of runs even with the single long-lived
subscribe. Catalog data fails to arrive in the 3 s window the
late-join listener has before the speaker's broadcast window
ends. Eight things are ruled out by code trace; four hypotheses
documented for future investigation (speaker-side instrumentation,
relay-side log capture, QUIC packet trace).
No production code change; test-only Rust patch already shipped.