32065901c8
Plans the path from 'infra shipped' to 'full coverage with correct behaviours'. Five new plan docs in nestsClient/plans/: 1. 2026-05-07-t16-closure-roadmap.md — index + priority order. Three sequential steps + one independent track. 2. 2026-05-07-moq-relay-routing-investigation.md (Priority 1) — root-cause the moq-relay 0.10.x per-broadcast subscribe- routing race. Step-by-step: capture relay-side trace, write minimum reproducer, file upstream OR bump moq-relay version. Smoking gun + hypotheses already in place from the late-join-flake investigation; this plan turns that into actionable next steps. 3. 2026-05-07-tighten-cross-stack-assertions.md (Priority 2) — once the routing race is closed, replace the five soft- pass scenarios in BrowserInteropTest with hard floors. Lists each scenario, its current soft-pass behavior, and the proposed hard threshold. 4. 2026-05-07-cross-stack-interop-ci-gating.md (Priority 3) — re-add the hang-interop + browser-interop GitHub Actions jobs that were dropped in6829ab727/b94737de7. Includes the exact YAML to restore and a 10/10 sweep stability bar before merge. 5. 2026-05-07-framespergroup-production-rerun.md (independent track) — re-run the HCgOY two-phone field tests against current nostrnests production at multiple framesPerGroup values to settle whether the test pin (5) and production default (50) can converge. Estimated total: 2.5-3.5 days of focused work to fully close T16. After these four: I7 post-reconnect cliff and I12 GOAWAY remain open as genuinely upstream-territory items, tracked in their existing investigation docs.
184 lines
7.1 KiB
Markdown
184 lines
7.1 KiB
Markdown
# Plan: investigate moq-relay 0.10.x per-broadcast subscribe-routing race
|
||
|
||
**Status:** specced — pickup ready.
|
||
|
||
**Owns:** the residual flake that affects four T16 scenarios:
|
||
`late_join_listener_still_decodes_tail`,
|
||
`packet_loss_1pct_does_not_kill_audio`,
|
||
`long_broadcast_60s_tone_round_trips`, and the new
|
||
`chromium_publisher_*_kotlin_listener_recovers` tests in browser-tier.
|
||
|
||
**Blocks:** CI gating for `:nestsClient:jvmTest -DnestsHangInterop=true`
|
||
and `-DnestsBrowserInterop=true`. Re-evaluate the
|
||
`hang-interop` / `browser-interop` workflow jobs once this is closed.
|
||
|
||
**Cross-refs:**
|
||
- `nestsClient/plans/2026-05-07-late-join-catalog-flake-investigation.md`
|
||
(smoking-gun trace + 4 mitigation attempts, 2 of which were
|
||
net-negative and reverted).
|
||
- `nestsClient/plans/2026-05-07-i7-post-reconnect-cliff-investigation.md`
|
||
(same kind of routing issue surfacing across publisher cycles).
|
||
|
||
## What we know
|
||
|
||
For broadcasts that fail (sample suffixes from the trace:
|
||
`10d4b6f2…`, `c75e2648…`, `f1be27ef…`):
|
||
|
||
1. The Kotlin speaker side logs:
|
||
- `ANNOUNCE inbound prefix='' → emitted Active suffix='<broadcast>'`
|
||
- …then NOTHING for the entire test window.
|
||
- Audio publisher's `send()` repeats `no inboundSubs` at 50 fps
|
||
until the test times out.
|
||
|
||
2. The Rust hang-listen side logs:
|
||
- `connected, version=moq-lite-03`
|
||
- `broadcast announced path=<broadcast>`
|
||
- `subscribe started id=0 broadcast=<broadcast> track=catalog.json`
|
||
- …then `subscribe error err=remote error: code=0` exactly when
|
||
the speaker tears down at the broadcast-window end (= relay
|
||
forwarding `Cancel`).
|
||
|
||
The relay accepts the listener's wire SUBSCRIBE on its downstream
|
||
connection but **never opens an upstream SUBSCRIBE bidi to the
|
||
speaker** for the failing broadcast. The upstream subscribe-pump
|
||
that's supposed to forward downstream subscribes to the speaker
|
||
isn't wired up by the time the listener subscribes.
|
||
|
||
For broadcasts that succeed (same trace, same JVM, different test):
|
||
|
||
```
|
||
ANNOUNCE inbound prefix='' → emitted Active suffix=<broadcast>
|
||
SUBSCRIBE inbound id=0 broadcast=<broadcast> track='catalog.json'
|
||
SUBSCRIBE registered id=0 …
|
||
openGroupStream subId=0 seq=0
|
||
…
|
||
```
|
||
|
||
All log lines fire; the relay forwards the upstream subscribe
|
||
within ~1 ms of the downstream subscribe. Failure mode is binary:
|
||
the relay does or does not forward.
|
||
|
||
## Hypotheses, ranked by next step
|
||
|
||
### H1 — moq-rs 0.10.x bug in `Origin::announced()` → upstream-pump setup race
|
||
|
||
`Origin::announced().await` returns the broadcast as soon as the
|
||
speaker's announce lands in the relay's origin map. The relay's
|
||
upstream-subscribe pump for that broadcast is set up on a separate
|
||
async path. If a downstream listener subscribes before the pump is
|
||
fully wired, the SUBSCRIBE accepts on the listener's wire (the
|
||
relay has the broadcast in its origin) but never propagates
|
||
upstream.
|
||
|
||
**Status:** prime suspect; see "smoking gun" in
|
||
`2026-05-07-late-join-catalog-flake-investigation.md`.
|
||
|
||
### H2 — interaction with the `--auth-public ""` minimal config
|
||
|
||
The harness boots moq-relay with `--auth-public ""` to skip JWT
|
||
issuance. Production runs with full auth. It's possible the
|
||
auth-public path takes a different code path through the relay's
|
||
origin/subscribe wiring that's racier than the auth'd path.
|
||
|
||
**Status:** plausible; would explain why the flake isn't reported
|
||
against the production deployment.
|
||
|
||
### H3 — local-only timing race that resolves at higher latency
|
||
|
||
Loopback (127.0.0.1) has near-zero RTT. The relay's internal
|
||
async setup may rely on the natural RTT cushion of a real network
|
||
to sequence upstream-subscribe-pump setup vs. downstream-subscribe
|
||
acceptance. We bypass that cushion in the test.
|
||
|
||
**Status:** less likely (the cliff plan's evidence shows lossy
|
||
network actually makes things *worse* via the `serve_group` task
|
||
pool) but worth ruling out.
|
||
|
||
## Investigation plan
|
||
|
||
### Step 1 — capture relay-side traces
|
||
|
||
`NativeMoqRelayHarness.boot` currently launches `moq-relay` with
|
||
`RUST_LOG=info`. Bump to `RUST_LOG=moq_relay=trace,moq_lite=trace`
|
||
and capture stderr to a per-test tempfile. Cross-reference with
|
||
the failing test's hang-listen stdout AND the speaker-side
|
||
`Log.d("NestTx")` traces (already captured in
|
||
`<system-err>` per JUnit XML).
|
||
|
||
Concretely: in `NativeMoqRelayHarness.kt` add a `--log-stderr`
|
||
option that the @BeforeTest hook sets to a `<test-method>.log`
|
||
path under `nestsClient/build/relay-logs/`. The Kotlin side
|
||
already has the speaker-side traces; the Rust side is the gap.
|
||
|
||
What to look for in the failed-broadcast log:
|
||
- Was a SUBSCRIBE bidi opened to the speaker for the failing
|
||
broadcast suffix? (moq_lite span: `subscribe`).
|
||
- Did the relay's `Origin::publish_broadcast` call complete
|
||
before the listener's SUBSCRIBE arrived?
|
||
- Any `track.unused()` resolves on the publisher-side track that
|
||
would explain immediate cancellation?
|
||
|
||
### Step 2 — write a minimal reproducer
|
||
|
||
If Step 1 shows the bug is independent of our test framework,
|
||
extract a minimum reproducer:
|
||
|
||
```rust
|
||
// reproducer.rs
|
||
let mut cmd = std::process::Command::new("moq-relay")
|
||
.args(&["--server-bind", "127.0.0.1:0", "--auth-public", "",
|
||
"--tls-generate", "localhost"])
|
||
.spawn()?;
|
||
// Run a moq-lite SPEAKER on one client, a moq-lite LISTENER on
|
||
// another, both pointed at the relay. Listener subscribes immediately
|
||
// after the speaker announces. Repeat 100×; count how many succeed.
|
||
```
|
||
|
||
Then strip the SPEAKER's announce timing, the LISTENER's subscribe
|
||
timing, the relay's `--auth-public` flag — bisect to the smallest
|
||
form that still reproduces.
|
||
|
||
### Step 3 — file upstream
|
||
|
||
If Step 1 / 2 confirm a moq-rs bug, file a `kixelated/moq` issue
|
||
with:
|
||
- The reproducer.
|
||
- Smoking-gun trace pair from our test harness.
|
||
- Pin to moq-rs version `0.10.25` (per `nestsClient/tests/hang-interop/REV`).
|
||
- Cross-link to existing
|
||
`2026-05-01-quic-stream-cliff-investigation.md`'s open follow-up
|
||
#1 (the per-subscriber forward-queue cliff is a sister bug).
|
||
|
||
### Step 4 — try newer moq-relay version
|
||
|
||
Bump `MOQ_RELAY_VERSION` in `nestsClient/tests/hang-interop/REV`
|
||
and `nestsClient/build.gradle.kts` to the next minor release on
|
||
crates.io (whatever's current at the time of pickup). Run the 5×
|
||
sweep. If the flake disappears, the upstream may have already
|
||
fixed it; we can pin past 0.10.x.
|
||
|
||
**Risk:** newer moq-relay versions may have wire-format changes
|
||
that break our current `moq-lite-03` ALPN pin. The browser
|
||
harness's `@moq/lite` 0.2.x client offers `moq-lite-04` AND
|
||
`moq-lite-03`, so a newer relay that drops `03` would still
|
||
negotiate fine via 04.
|
||
|
||
## Acceptance criteria
|
||
|
||
- Sweep `for i in 1 2 3 4 5; do ./gradlew :nestsClient:jvmTest
|
||
--tests HangInteropTest -DnestsHangInterop=true --rerun-tasks; done`
|
||
passes 5/5.
|
||
- Browser-tier sweep similarly stable.
|
||
- Either:
|
||
(a) Upstream issue filed with reproducer (if the bug is in
|
||
moq-rs and we can't fix it locally), OR
|
||
(b) Local fix applied (e.g. version bump + REV update +
|
||
Cargo.lock regenerate).
|
||
|
||
## Out of scope
|
||
|
||
- The `:quic` module's `MAX_STREAMS_UNI` extension fix
|
||
(`d391ae1d`) — already shipped, separate concern.
|
||
- The production-side `framesPerGroup` reconciliation
|
||
(`2026-05-07-framespergroup-production-rerun.md`) — independent.
|