Files
amethyst/nestsClient/plans/2026-05-07-moq-relay-routing-investigation.md
T
Claude 32065901c8 docs(nests): T16 closure roadmap + 4 follow-up plans
Plans the path from 'infra shipped' to 'full coverage with
correct behaviours'. Five new plan docs in nestsClient/plans/:

1. 2026-05-07-t16-closure-roadmap.md — index + priority order.
   Three sequential steps + one independent track.

2. 2026-05-07-moq-relay-routing-investigation.md (Priority 1)
   — root-cause the moq-relay 0.10.x per-broadcast subscribe-
   routing race. Step-by-step: capture relay-side trace, write
   minimum reproducer, file upstream OR bump moq-relay version.
   Smoking gun + hypotheses already in place from the
   late-join-flake investigation; this plan turns that into
   actionable next steps.

3. 2026-05-07-tighten-cross-stack-assertions.md (Priority 2)
   — once the routing race is closed, replace the five soft-
   pass scenarios in BrowserInteropTest with hard floors.
   Lists each scenario, its current soft-pass behavior, and
   the proposed hard threshold.

4. 2026-05-07-cross-stack-interop-ci-gating.md (Priority 3)
   — re-add the hang-interop + browser-interop GitHub Actions
   jobs that were dropped in 6829ab727 / b94737de7. Includes
   the exact YAML to restore and a 10/10 sweep stability bar
   before merge.

5. 2026-05-07-framespergroup-production-rerun.md (independent
   track) — re-run the HCgOY two-phone field tests against
   current nostrnests production at multiple framesPerGroup
   values to settle whether the test pin (5) and production
   default (50) can converge.

Estimated total: 2.5-3.5 days of focused work to fully close
T16. After these four: I7 post-reconnect cliff and I12 GOAWAY
remain open as genuinely upstream-territory items, tracked in
their existing investigation docs.
2026-05-07 14:36:21 +00:00

184 lines
7.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Plan: investigate moq-relay 0.10.x per-broadcast subscribe-routing race
**Status:** specced — pickup ready.
**Owns:** the residual flake that affects four T16 scenarios:
`late_join_listener_still_decodes_tail`,
`packet_loss_1pct_does_not_kill_audio`,
`long_broadcast_60s_tone_round_trips`, and the new
`chromium_publisher_*_kotlin_listener_recovers` tests in browser-tier.
**Blocks:** CI gating for `:nestsClient:jvmTest -DnestsHangInterop=true`
and `-DnestsBrowserInterop=true`. Re-evaluate the
`hang-interop` / `browser-interop` workflow jobs once this is closed.
**Cross-refs:**
- `nestsClient/plans/2026-05-07-late-join-catalog-flake-investigation.md`
(smoking-gun trace + 4 mitigation attempts, 2 of which were
net-negative and reverted).
- `nestsClient/plans/2026-05-07-i7-post-reconnect-cliff-investigation.md`
(same kind of routing issue surfacing across publisher cycles).
## What we know
For broadcasts that fail (sample suffixes from the trace:
`10d4b6f2…`, `c75e2648…`, `f1be27ef…`):
1. The Kotlin speaker side logs:
- `ANNOUNCE inbound prefix='' → emitted Active suffix='<broadcast>'`
- …then NOTHING for the entire test window.
- Audio publisher's `send()` repeats `no inboundSubs` at 50 fps
until the test times out.
2. The Rust hang-listen side logs:
- `connected, version=moq-lite-03`
- `broadcast announced path=<broadcast>`
- `subscribe started id=0 broadcast=<broadcast> track=catalog.json`
- …then `subscribe error err=remote error: code=0` exactly when
the speaker tears down at the broadcast-window end (= relay
forwarding `Cancel`).
The relay accepts the listener's wire SUBSCRIBE on its downstream
connection but **never opens an upstream SUBSCRIBE bidi to the
speaker** for the failing broadcast. The upstream subscribe-pump
that's supposed to forward downstream subscribes to the speaker
isn't wired up by the time the listener subscribes.
For broadcasts that succeed (same trace, same JVM, different test):
```
ANNOUNCE inbound prefix='' → emitted Active suffix=<broadcast>
SUBSCRIBE inbound id=0 broadcast=<broadcast> track='catalog.json'
SUBSCRIBE registered id=0 …
openGroupStream subId=0 seq=0
```
All log lines fire; the relay forwards the upstream subscribe
within ~1 ms of the downstream subscribe. Failure mode is binary:
the relay does or does not forward.
## Hypotheses, ranked by next step
### H1 — moq-rs 0.10.x bug in `Origin::announced()` → upstream-pump setup race
`Origin::announced().await` returns the broadcast as soon as the
speaker's announce lands in the relay's origin map. The relay's
upstream-subscribe pump for that broadcast is set up on a separate
async path. If a downstream listener subscribes before the pump is
fully wired, the SUBSCRIBE accepts on the listener's wire (the
relay has the broadcast in its origin) but never propagates
upstream.
**Status:** prime suspect; see "smoking gun" in
`2026-05-07-late-join-catalog-flake-investigation.md`.
### H2 — interaction with the `--auth-public ""` minimal config
The harness boots moq-relay with `--auth-public ""` to skip JWT
issuance. Production runs with full auth. It's possible the
auth-public path takes a different code path through the relay's
origin/subscribe wiring that's racier than the auth'd path.
**Status:** plausible; would explain why the flake isn't reported
against the production deployment.
### H3 — local-only timing race that resolves at higher latency
Loopback (127.0.0.1) has near-zero RTT. The relay's internal
async setup may rely on the natural RTT cushion of a real network
to sequence upstream-subscribe-pump setup vs. downstream-subscribe
acceptance. We bypass that cushion in the test.
**Status:** less likely (the cliff plan's evidence shows lossy
network actually makes things *worse* via the `serve_group` task
pool) but worth ruling out.
## Investigation plan
### Step 1 — capture relay-side traces
`NativeMoqRelayHarness.boot` currently launches `moq-relay` with
`RUST_LOG=info`. Bump to `RUST_LOG=moq_relay=trace,moq_lite=trace`
and capture stderr to a per-test tempfile. Cross-reference with
the failing test's hang-listen stdout AND the speaker-side
`Log.d("NestTx")` traces (already captured in
`<system-err>` per JUnit XML).
Concretely: in `NativeMoqRelayHarness.kt` add a `--log-stderr`
option that the @BeforeTest hook sets to a `<test-method>.log`
path under `nestsClient/build/relay-logs/`. The Kotlin side
already has the speaker-side traces; the Rust side is the gap.
What to look for in the failed-broadcast log:
- Was a SUBSCRIBE bidi opened to the speaker for the failing
broadcast suffix? (moq_lite span: `subscribe`).
- Did the relay's `Origin::publish_broadcast` call complete
before the listener's SUBSCRIBE arrived?
- Any `track.unused()` resolves on the publisher-side track that
would explain immediate cancellation?
### Step 2 — write a minimal reproducer
If Step 1 shows the bug is independent of our test framework,
extract a minimum reproducer:
```rust
// reproducer.rs
let mut cmd = std::process::Command::new("moq-relay")
.args(&["--server-bind", "127.0.0.1:0", "--auth-public", "",
"--tls-generate", "localhost"])
.spawn()?;
// Run a moq-lite SPEAKER on one client, a moq-lite LISTENER on
// another, both pointed at the relay. Listener subscribes immediately
// after the speaker announces. Repeat 100×; count how many succeed.
```
Then strip the SPEAKER's announce timing, the LISTENER's subscribe
timing, the relay's `--auth-public` flag — bisect to the smallest
form that still reproduces.
### Step 3 — file upstream
If Step 1 / 2 confirm a moq-rs bug, file a `kixelated/moq` issue
with:
- The reproducer.
- Smoking-gun trace pair from our test harness.
- Pin to moq-rs version `0.10.25` (per `nestsClient/tests/hang-interop/REV`).
- Cross-link to existing
`2026-05-01-quic-stream-cliff-investigation.md`'s open follow-up
#1 (the per-subscriber forward-queue cliff is a sister bug).
### Step 4 — try newer moq-relay version
Bump `MOQ_RELAY_VERSION` in `nestsClient/tests/hang-interop/REV`
and `nestsClient/build.gradle.kts` to the next minor release on
crates.io (whatever's current at the time of pickup). Run the 5×
sweep. If the flake disappears, the upstream may have already
fixed it; we can pin past 0.10.x.
**Risk:** newer moq-relay versions may have wire-format changes
that break our current `moq-lite-03` ALPN pin. The browser
harness's `@moq/lite` 0.2.x client offers `moq-lite-04` AND
`moq-lite-03`, so a newer relay that drops `03` would still
negotiate fine via 04.
## Acceptance criteria
- Sweep `for i in 1 2 3 4 5; do ./gradlew :nestsClient:jvmTest
--tests HangInteropTest -DnestsHangInterop=true --rerun-tasks; done`
passes 5/5.
- Browser-tier sweep similarly stable.
- Either:
(a) Upstream issue filed with reproducer (if the bug is in
moq-rs and we can't fix it locally), OR
(b) Local fix applied (e.g. version bump + REV update +
Cargo.lock regenerate).
## Out of scope
- The `:quic` module's `MAX_STREAMS_UNI` extension fix
(`d391ae1d`) — already shipped, separate concern.
- The production-side `framesPerGroup` reconciliation
(`2026-05-07-framespergroup-production-rerun.md`) — independent.