diff --git a/nestsClient/plans/2026-05-07-cross-stack-interop-ci-gating.md b/nestsClient/plans/2026-05-07-cross-stack-interop-ci-gating.md new file mode 100644 index 000000000..5f19f3168 --- /dev/null +++ b/nestsClient/plans/2026-05-07-cross-stack-interop-ci-gating.md @@ -0,0 +1,171 @@ +# Plan: wire CI gating for the cross-stack interop suite + +**Status:** specced — pickup ready. +**Depends on:** +- `2026-05-07-moq-relay-routing-investigation.md` closed +- `2026-05-07-tighten-cross-stack-assertions.md` closed +- 5/5 sweep stability verified + +This is the FINAL step of the T16 closure. With stable hard-pass +suites, CI gating becomes safe and meaningful. + +## What's needed + +### A) `.github/workflows/build.yml` — the hang-interop job + +The job was originally part of this branch but removed per +maintainer ask in commit `6829ab727` ("ci(nests): drop hang-interop +job from build.yml") because the suite was flaky. Resurrect the +exact same shape: + +```yaml +hang-interop: + needs: lint + runs-on: ubuntu-latest + timeout-minutes: 30 + steps: + - uses: actions/checkout@v6 + - uses: actions/setup-java@v5 + with: { distribution: 'zulu', java-version: 21 } + - uses: gradle/actions/setup-gradle@v4 + with: + cache-read-only: ${{ github.ref != 'refs/heads/main' }} + - uses: dtolnay/rust-toolchain@stable + - uses: actions/cache@v4 + with: + path: | + ~/.cargo/registry + ~/.cargo/git + nestsClient/tests/hang-interop/target + ~/.cache/amethyst-nests-interop/hang-interop-cargo + key: ${{ runner.os }}-cargo-${{ hashFiles('nestsClient/tests/hang-interop/Cargo.lock', 'nestsClient/tests/hang-interop/REV') }} + restore-keys: | + ${{ runner.os }}-cargo- + - name: Run cross-stack interop suite + run: ./gradlew :nestsClient:jvmTest -DnestsHangInterop=true + - uses: actions/upload-artifact@v7 + if: failure() + with: + name: Hang Interop Test Reports + path: nestsClient/build/reports/tests/jvmTest/ +``` + +The `git show 6829ab727 -- .github/workflows/build.yml` reverse +gives the exact diff to re-add. Linux-only is correct: the cargo +install of moq-relay 0.10.x has nontrivial native deps +(aws-lc-sys, ring) that take 5+ min cold; cached runs ~30 s. +macOS / Windows would double matrix cost without catching new +defects. + +### B) `.github/workflows/build.yml` — the browser-interop job + +Same shape as A, plus bun + Playwright caching: + +```yaml +browser-interop: + needs: lint + runs-on: ubuntu-latest + timeout-minutes: 30 + steps: + # ...same checkout + JDK + Gradle + Rust + cargo cache as hang-interop... + - uses: oven-sh/setup-bun@v2 + with: { bun-version: 1.3.11 } + - uses: actions/cache@v4 + with: + path: | + nestsClient-browser-interop/node_modules + nestsClient-browser-interop/dist + key: ${{ runner.os }}-bun-${{ hashFiles('nestsClient-browser-interop/package.json', 'nestsClient-browser-interop/bun.lock') }} + - uses: actions/cache@v4 + with: + path: ~/.cache/ms-playwright + key: ${{ runner.os }}-playwright-${{ hashFiles('nestsClient-browser-interop/package.json') }} + - name: Run browser cross-stack interop suite + run: | + ./gradlew :nestsClient:jvmTest \ + --tests "com.vitorpamplona.nestsclient.interop.native.BrowserInteropTest" \ + -DnestsHangInterop=true \ + -DnestsBrowserInterop=true + - uses: actions/upload-artifact@v7 + if: failure() + with: + name: Browser Interop Test Reports + path: | + nestsClient/build/reports/tests/jvmTest/ + nestsClient-browser-interop/test-results/ + nestsClient-browser-interop/playwright-report/ +``` + +Same `git show b94737de7 -- .github/workflows/build.yml` reverse +gives the exact diff (`feat/nests-browser-interop`'s removal +commit). + +### C) Cross-link with `:cli` interop tests + +The existing `nests-interop` opt-in pattern already lives in +`cli/tests/nests/nests-interop.sh`. Confirm both new jobs run +in parallel with that without resource contention. They use +different ports (NativeMoqRelayHarness reserves `ServerSocket(0)`) +so they're independent at the network level. + +## Stability bar + +Before flipping the CI switch, run: + +``` +for i in 1 2 3 4 5 6 7 8 9 10; do + echo "=== run $i ===" + ./gradlew :nestsClient:jvmTest \ + --tests HangInteropTest \ + --tests BrowserInteropTest \ + -DnestsHangInterop=true \ + -DnestsBrowserInterop=true \ + --rerun-tasks 2>&1 | grep -E "FAILED]|BUILD" +done +``` + +10/10 BUILD SUCCESSFUL. If even one fails, do NOT wire CI; loop +back to the routing investigation. + +## CI runtime budget + +- Hang-interop job: ~3-4 min on warm cache (one suite run, 60 s + long-broadcast scenario dominates), ~8 min cold (cargo install + moq-relay). +- Browser-interop job: ~5-7 min warm (Chromium boot × N + scenarios), ~10 min cold (Playwright install). +- Both run in parallel after `lint`. + +Total CI overhead: ~5-10 min on the critical path beyond the +existing build matrix. Acceptable. + +## Documentation updates + +After CI is green: + +1. `nestsClient/plans/2026-05-06-cross-stack-interop-test-results.md` + — replace the "CI integration: Not wired" section with "wired, + tracking flake-rate at 0/N runs". +2. `nestsClient/plans/2026-05-06-cross-stack-interop-test-gap-matrix.md` + — replace `#6 CI integration: ⏸ deferred` with `✅ live`. +3. Pick a maintainer to monitor the first 2 weeks of CI runs and + bisect any new flake immediately (don't let it accumulate as + "known flake" again). + +## Acceptance criteria + +- Both jobs added to `build.yml` and merge to main. +- 10/10 sweep before merge. +- First 2 weeks post-merge: ≥ 95% green rate. If lower, the + routing investigation isn't really done — pull the jobs again + until it is. + +## Optional follow-ups + +- **Add I-12 GOAWAY scenario IF an IETF moq-transport target lands.** + Currently N/A in moq-lite-03 per + `cross-stack-interop-test-results.md`'s I12 section. If an IETF + target ever ships, this is the cross-stack regression test. +- **Surface `framesPerGroup` as a per-deployment config** if the + framesPerGroup-rerun outcome shows the two rigs can't converge + (see `2026-05-07-framespergroup-production-rerun.md`). diff --git a/nestsClient/plans/2026-05-07-framespergroup-production-rerun.md b/nestsClient/plans/2026-05-07-framespergroup-production-rerun.md new file mode 100644 index 000000000..388313211 --- /dev/null +++ b/nestsClient/plans/2026-05-07-framespergroup-production-rerun.md @@ -0,0 +1,132 @@ +# Plan: re-run HCgOY field tests against current production + +**Status:** specced — pickup ready (needs prod-rig access). +**Cross-ref:** `nestsClient/plans/2026-05-07-framespergroup-reconciliation.md` +documents the conflict between the cliff plan's value (5) and +HCgOY's value (50). This plan settles which is current truth. + +## What we're trying to settle + +Production currently runs `NestMoqLiteBroadcaster.DEFAULT_FRAMES_PER_GROUP = 50` +based on HCgOY two-phone field tests at commit `6e4df4a` +(2026-05-05) which observed: + +| `framesPerGroup` | streams/sec @ 50 fps | observed cliff window | +|---|---|---| +| 1 | 50 | ~3 s | +| 5 | 10 | ~13 s | +| 10 | 5 | ~16 s | +| 50 | 1 | not reached | +| 100 | 0.5 | never observed | + +Interop tests pin `5` because the local `--auth-public ""` minimal +relay setup hits a *different* cliff (per-stream byte volume) at +`framesPerGroup = 50`. + +The production deployment may have changed since 2026-05-05: +- nostrnests may have updated their `moq-relay` version +- their resource limits may have shifted +- the upstream `kixelated/moq` may have addressed one or both + cliffs + +We don't know without re-running. **Cliff value at `framesPerGroup += 5` may now be unbounded — if so, both rigs can converge on 5 +and the test pin matches the prod default.** + +## Test setup + +### Rig A — interop env (already exists) + +`./gradlew :nestsClient:jvmTest --tests HangInteropTest -DnestsHangInterop=true` +runs against local `moq-relay 0.10.25 --auth-public "" --tls-generate localhost`. +Long-broadcast scenario `long_broadcast_60s_tone_round_trips` is +the existing 60-second sustained-stream test pinned at +`framesPerGroup = 5`. + +To probe other values, parameterize the helper: + +```kotlin +// HangInteropTest.kt — runSpeakerToHangListen helper +private suspend fun runSpeakerToHangListen( + speakerSeconds: Int, + framesPerGroup: Int = 5, // ← new parameter + // ...existing params +): HangListenOutput { ... } +``` + +Then add scenarios `long_broadcast_60s_framesPerGroup_50` etc. +that pin different values. Expected outcomes today (per the +2026-05-01 cliff plan): +- `framesPerGroup = 5` — passes (current pin) +- `framesPerGroup = 10` — passes +- `framesPerGroup = 50` — fails (per-stream byte volume cliff + at the local minimal relay) + +### Rig B — production deployment (needs maintainer access) + +This is the gap. Rerunning the HCgOY two-phone field test pattern +needs: +- Two physical Android devices +- A nostrnests room (production endpoint + `wss://nostrnests.com/v0/ws` per `NestsConnect.kt`) +- The diagnostic-build of Amethyst that emits the cliff-detector + trace logs (see commit `6e4df4a`'s logcat run from 18:37:43..18:38:08) + +The maintainer should run the same test pattern at: +- `framesPerGroup = 5` (current test value) +- `framesPerGroup = 10` +- `framesPerGroup = 25` (untested, midpoint) +- `framesPerGroup = 50` (current prod value) +- `framesPerGroup = 100` (full group; the cliff plan's + `fpg-all` reference) + +For each, broadcast for 120 s and observe: +- Total streams forwarded by the relay +- Time-to-cliff if any (when the listener-side flow-control + snapshot stops incrementing `peerInitiatedUni`) +- Audio dropouts (perceptual + sample-count) + +## Decision matrix after data lands + +| Rig A passes at | Rig B passes at | Decision | +|---|---|---| +| 5, 10 | 5, 10, 25, 50, 100 | Keep prod 50; test pins 5 (current state) | +| 5, 10, 50 | 5, 10, 25, 50, 100 | Both rigs converge → unify on 50, test pin matches prod | +| 5, 10 | 50, 100 only (5 still cliffs) | Current state is correct; document permanently as "two cliffs in one binary" | +| 5 only | 50, 100 only (5 still cliffs) | The two cliffs are real; consider per-environment config | +| 5, 10, 50 | 50, 100 only (5 still cliffs) | Test rig fixed; production cliff still hits at 5. Test pin doesn't catch prod regression — bigger problem. | + +The "decision" column drives the production-side change (or +non-change) to `DEFAULT_FRAMES_PER_GROUP`. + +## What lands as code + +After Rig B data is in: + +1. Update kdoc on `NestMoqLiteBroadcaster.DEFAULT_FRAMES_PER_GROUP` + citing the new run's logcat dates / commit. +2. If the values converge, change the default to match. Update + the test-side `framesPerGroup = 5` pin to match the new + default — keep both rigs aligned. +3. If values still diverge, document it explicitly as a known + environment-dependent value. Consider exposing + `framesPerGroup` as a per-deployment config (currently only + exposed as a constructor parameter — wire to a config knob if + product wants per-deployment tuning). +4. Update `nestsClient/plans/2026-05-07-framespergroup-reconciliation.md`'s + "Recommendation" section with the data-driven outcome. + +## Acceptance criteria + +- A logcat dump from Rig B with `framesPerGroup = 5` for ≥ 60 s + showing whether the cliff still hits at ~13 s. +- Decision logged in the framesPerGroup reconciliation doc. +- If a value change lands, the test-side pin and production + default agree. + +## Out of scope + +- The local interop env's per-stream byte cliff at + `framesPerGroup = 50`. That's a separate thread; addressing it + would require either a different relay configuration or + patching moq-relay itself. diff --git a/nestsClient/plans/2026-05-07-moq-relay-routing-investigation.md b/nestsClient/plans/2026-05-07-moq-relay-routing-investigation.md new file mode 100644 index 000000000..f9ace613d --- /dev/null +++ b/nestsClient/plans/2026-05-07-moq-relay-routing-investigation.md @@ -0,0 +1,183 @@ +# Plan: investigate moq-relay 0.10.x per-broadcast subscribe-routing race + +**Status:** specced — pickup ready. + +**Owns:** the residual flake that affects four T16 scenarios: +`late_join_listener_still_decodes_tail`, +`packet_loss_1pct_does_not_kill_audio`, +`long_broadcast_60s_tone_round_trips`, and the new +`chromium_publisher_*_kotlin_listener_recovers` tests in browser-tier. + +**Blocks:** CI gating for `:nestsClient:jvmTest -DnestsHangInterop=true` +and `-DnestsBrowserInterop=true`. Re-evaluate the +`hang-interop` / `browser-interop` workflow jobs once this is closed. + +**Cross-refs:** +- `nestsClient/plans/2026-05-07-late-join-catalog-flake-investigation.md` + (smoking-gun trace + 4 mitigation attempts, 2 of which were + net-negative and reverted). +- `nestsClient/plans/2026-05-07-i7-post-reconnect-cliff-investigation.md` + (same kind of routing issue surfacing across publisher cycles). + +## What we know + +For broadcasts that fail (sample suffixes from the trace: +`10d4b6f2…`, `c75e2648…`, `f1be27ef…`): + +1. The Kotlin speaker side logs: + - `ANNOUNCE inbound prefix='' → emitted Active suffix=''` + - …then NOTHING for the entire test window. + - Audio publisher's `send()` repeats `no inboundSubs` at 50 fps + until the test times out. + +2. The Rust hang-listen side logs: + - `connected, version=moq-lite-03` + - `broadcast announced path=` + - `subscribe started id=0 broadcast= track=catalog.json` + - …then `subscribe error err=remote error: code=0` exactly when + the speaker tears down at the broadcast-window end (= relay + forwarding `Cancel`). + +The relay accepts the listener's wire SUBSCRIBE on its downstream +connection but **never opens an upstream SUBSCRIBE bidi to the +speaker** for the failing broadcast. The upstream subscribe-pump +that's supposed to forward downstream subscribes to the speaker +isn't wired up by the time the listener subscribes. + +For broadcasts that succeed (same trace, same JVM, different test): + +``` +ANNOUNCE inbound prefix='' → emitted Active suffix= +SUBSCRIBE inbound id=0 broadcast= track='catalog.json' +SUBSCRIBE registered id=0 … +openGroupStream subId=0 seq=0 +… +``` + +All log lines fire; the relay forwards the upstream subscribe +within ~1 ms of the downstream subscribe. Failure mode is binary: +the relay does or does not forward. + +## Hypotheses, ranked by next step + +### H1 — moq-rs 0.10.x bug in `Origin::announced()` → upstream-pump setup race + +`Origin::announced().await` returns the broadcast as soon as the +speaker's announce lands in the relay's origin map. The relay's +upstream-subscribe pump for that broadcast is set up on a separate +async path. If a downstream listener subscribes before the pump is +fully wired, the SUBSCRIBE accepts on the listener's wire (the +relay has the broadcast in its origin) but never propagates +upstream. + +**Status:** prime suspect; see "smoking gun" in +`2026-05-07-late-join-catalog-flake-investigation.md`. + +### H2 — interaction with the `--auth-public ""` minimal config + +The harness boots moq-relay with `--auth-public ""` to skip JWT +issuance. Production runs with full auth. It's possible the +auth-public path takes a different code path through the relay's +origin/subscribe wiring that's racier than the auth'd path. + +**Status:** plausible; would explain why the flake isn't reported +against the production deployment. + +### H3 — local-only timing race that resolves at higher latency + +Loopback (127.0.0.1) has near-zero RTT. The relay's internal +async setup may rely on the natural RTT cushion of a real network +to sequence upstream-subscribe-pump setup vs. downstream-subscribe +acceptance. We bypass that cushion in the test. + +**Status:** less likely (the cliff plan's evidence shows lossy +network actually makes things *worse* via the `serve_group` task +pool) but worth ruling out. + +## Investigation plan + +### Step 1 — capture relay-side traces + +`NativeMoqRelayHarness.boot` currently launches `moq-relay` with +`RUST_LOG=info`. Bump to `RUST_LOG=moq_relay=trace,moq_lite=trace` +and capture stderr to a per-test tempfile. Cross-reference with +the failing test's hang-listen stdout AND the speaker-side +`Log.d("NestTx")` traces (already captured in +`` per JUnit XML). + +Concretely: in `NativeMoqRelayHarness.kt` add a `--log-stderr` +option that the @BeforeTest hook sets to a `.log` +path under `nestsClient/build/relay-logs/`. The Kotlin side +already has the speaker-side traces; the Rust side is the gap. + +What to look for in the failed-broadcast log: +- Was a SUBSCRIBE bidi opened to the speaker for the failing + broadcast suffix? (moq_lite span: `subscribe`). +- Did the relay's `Origin::publish_broadcast` call complete + before the listener's SUBSCRIBE arrived? +- Any `track.unused()` resolves on the publisher-side track that + would explain immediate cancellation? + +### Step 2 — write a minimal reproducer + +If Step 1 shows the bug is independent of our test framework, +extract a minimum reproducer: + +```rust +// reproducer.rs +let mut cmd = std::process::Command::new("moq-relay") + .args(&["--server-bind", "127.0.0.1:0", "--auth-public", "", + "--tls-generate", "localhost"]) + .spawn()?; +// Run a moq-lite SPEAKER on one client, a moq-lite LISTENER on +// another, both pointed at the relay. Listener subscribes immediately +// after the speaker announces. Repeat 100×; count how many succeed. +``` + +Then strip the SPEAKER's announce timing, the LISTENER's subscribe +timing, the relay's `--auth-public` flag — bisect to the smallest +form that still reproduces. + +### Step 3 — file upstream + +If Step 1 / 2 confirm a moq-rs bug, file a `kixelated/moq` issue +with: +- The reproducer. +- Smoking-gun trace pair from our test harness. +- Pin to moq-rs version `0.10.25` (per `nestsClient/tests/hang-interop/REV`). +- Cross-link to existing + `2026-05-01-quic-stream-cliff-investigation.md`'s open follow-up + #1 (the per-subscriber forward-queue cliff is a sister bug). + +### Step 4 — try newer moq-relay version + +Bump `MOQ_RELAY_VERSION` in `nestsClient/tests/hang-interop/REV` +and `nestsClient/build.gradle.kts` to the next minor release on +crates.io (whatever's current at the time of pickup). Run the 5× +sweep. If the flake disappears, the upstream may have already +fixed it; we can pin past 0.10.x. + +**Risk:** newer moq-relay versions may have wire-format changes +that break our current `moq-lite-03` ALPN pin. The browser +harness's `@moq/lite` 0.2.x client offers `moq-lite-04` AND +`moq-lite-03`, so a newer relay that drops `03` would still +negotiate fine via 04. + +## Acceptance criteria + +- Sweep `for i in 1 2 3 4 5; do ./gradlew :nestsClient:jvmTest + --tests HangInteropTest -DnestsHangInterop=true --rerun-tasks; done` + passes 5/5. +- Browser-tier sweep similarly stable. +- Either: + (a) Upstream issue filed with reproducer (if the bug is in + moq-rs and we can't fix it locally), OR + (b) Local fix applied (e.g. version bump + REV update + + Cargo.lock regenerate). + +## Out of scope + +- The `:quic` module's `MAX_STREAMS_UNI` extension fix + (`d391ae1d`) — already shipped, separate concern. +- The production-side `framesPerGroup` reconciliation + (`2026-05-07-framespergroup-production-rerun.md`) — independent. diff --git a/nestsClient/plans/2026-05-07-t16-closure-roadmap.md b/nestsClient/plans/2026-05-07-t16-closure-roadmap.md new file mode 100644 index 000000000..d962e0c3f --- /dev/null +++ b/nestsClient/plans/2026-05-07-t16-closure-roadmap.md @@ -0,0 +1,123 @@ +# T16 closure roadmap — full coverage with correct behaviours + +**Goal state.** Every spec'd cross-stack scenario green in suite-mode +sweeps, asserting its full design intent (no soft-passes, no vacuous +threshold loosening), with CI gating live and stable. + +**Where we are.** The merged `claude/cross-stack-interop-test-XAbYB` +branch ships 22 of 23 spec'd scenarios; each passes individually. +Suite-mode runs hit a residual moq-relay 0.10.x routing race on a +specific subset (~40-60% flake rate). Five scenarios soft-pass the +listener side as a known-flake mitigation. CI is intentionally +unwired pending stability. + +This roadmap takes the suite from "passes individually" to "passes +in suite + CI" through three sequential plans. None should be +parallelized — each unblocks the next. + +## Priority 1 — `2026-05-07-moq-relay-routing-investigation.md` + +**Why first.** The race is the root cause of every soft-pass and +the reason CI isn't wired. Without resolving it, downstream plans +mask flake rather than catch regressions. + +**What lands.** +- Either an upstream moq-relay version bump that closes the bug, + OR a documented relay configuration tweak that does. +- Or, if neither: a filed `kixelated/moq` issue with reproducer + + trace pair, plus a documented decision to keep CI unwired until + upstream resolves. + +**Acceptance bar.** 5/5 sweep BUILD SUCCESSFUL on the existing +HangInteropTest + BrowserInteropTest with their CURRENT soft-pass +assertions intact. (The next step tightens those.) + +## Priority 2 — `2026-05-07-tighten-cross-stack-assertions.md` + +**Why second.** Once the suite is stable, every soft-pass that +returned vacuous-pass on listener-side 0-frame outcomes is now +HIDING regressions instead of side-stepping flakes. Replace each +with a hard floor. + +**What lands.** +- Five BrowserInteropTest scenarios get hard sample-count + FFT + floors (or tightened existing ones). +- Gap matrix updated to reflect hard-pass coverage. + +**Acceptance bar.** 5/5 sweep AGAIN, this time with hard +assertions. If anything fail-flakes, the routing investigation +isn't really done — loop back. + +## Priority 3 — `2026-05-07-cross-stack-interop-ci-gating.md` + +**Why third.** Stability + hard-asserts in place → CI is now a +net positive (catches regressions, doesn't burn maintainer time +on false reds). + +**What lands.** +- Re-add `hang-interop` job (was at commit `6829ab727`'s parent; + `git show 6829ab727 -- .github/workflows/build.yml` reverse + gives the exact diff). +- Re-add `browser-interop` job (same pattern, plus bun + + Playwright caches). +- Documentation update across the results plan + gap matrix. + +**Acceptance bar.** 10/10 sweep before merge; ≥ 95% CI green +rate over the first 2 weeks. If lower, the upstream race isn't +fully closed — pull the jobs. + +## Independent track — `2026-05-07-framespergroup-production-rerun.md` + +This one **doesn't block the closure roadmap**. It can run any +time after Priority 1 is done; it settles whether the test pin +(5) and production default (50) can converge, or whether they +must remain different. Either outcome is shippable. + +**What lands.** +- Logcat data from a fresh two-phone field test against current + nostrnests production at multiple `framesPerGroup` values. +- A data-driven decision on whether to change the production + default, the test pin, or neither. + +**Why it's parallelizable.** Doesn't gate the test suite or CI; +it gates a one-line code change to `NestMoqLiteBroadcaster`'s +default constant. + +## After all four close — what remains + +Two open items, both genuinely upstream: + +1. **I7 post-reconnect listener cliff** — + `2026-05-07-i7-post-reconnect-cliff-investigation.md`. The I7 + reverse scenario passes its 2.5 s threshold but a regression + test of "all post-reconnect data arrives" would require the + moq-relay 0.10.x per-broadcast forward queue fix. Same upstream + class as the routing race. + +2. **I12 GOAWAY** — only re-emerges if an IETF moq-transport + target lands (currently moq-lite-03 only). Tracked in + `2026-05-06-cross-stack-interop-test-results.md`. + +Beyond those: T16 reaches "full coverage with correct behaviours" +when this roadmap closes. + +## Estimated wallclock + +- Priority 1: 1–2 days (depends on whether upstream version bump + fixes it, or we have to file + wait for upstream). +- Priority 2: 0.5 day (mechanical replacement of soft-passes + with floors, plus rerun verification). +- Priority 3: 0.5 day (re-add CI jobs, run the 10× sweep, merge). +- Independent track (framesPerGroup): 0.5 day (needs prod-rig + access). + +Total: 2.5–3.5 days of focused work to take T16 from "infra +shipped" to "fully closed". + +## Plan files + +- `2026-05-07-moq-relay-routing-investigation.md` +- `2026-05-07-tighten-cross-stack-assertions.md` +- `2026-05-07-cross-stack-interop-ci-gating.md` +- `2026-05-07-framespergroup-production-rerun.md` +- (this file) `2026-05-07-t16-closure-roadmap.md` diff --git a/nestsClient/plans/2026-05-07-tighten-cross-stack-assertions.md b/nestsClient/plans/2026-05-07-tighten-cross-stack-assertions.md new file mode 100644 index 000000000..2928c0a1e --- /dev/null +++ b/nestsClient/plans/2026-05-07-tighten-cross-stack-assertions.md @@ -0,0 +1,114 @@ +# Plan: tighten cross-stack interop assertions to hard-pass + +**Status:** specced — pickup ready. +**Depends on:** `2026-05-07-moq-relay-routing-investigation.md` +must be closed first (the soft-passes exist *because* of that flake; +removing them while the flake is unresolved produces fail-flakes, +not regression catches). + +## Why soft passes exist today + +Five scenarios currently soft-pass (vacuous-pass on listener-side +0-frame outcomes) to keep the test suite from fail-flaking on the +upstream relay-routing race documented in +`2026-05-07-late-join-catalog-flake-investigation.md`: + +| Scenario | File | Soft-pass behavior | +|---|---|---| +| `chromium_listener_late_join_still_decodes_tail` | `BrowserInteropTest.kt` | `if (pcm.size <= warmupSamples) return` | +| `chromium_listener_mid_broadcast_mute_shortens_pcm` | `BrowserInteropTest.kt` | same | +| `chromium_decoder_no_errors_through_warmup_window` (I14) | `BrowserInteropTest.kt` | no `decoderOutputs >= 4` floor | +| `chromium_publisher_baseline_kotlin_listener_decodes` | `BrowserInteropTest.kt` | hard-asserts publisher framesIn; soft-asserts listener | +| `chromium_publisher_reconnect_kotlin_listener_recovers` (Browser I7) | `BrowserInteropTest.kt` | same as baseline | + +All five have hard assertions on the `framesIn` / publisher-side +behavior; the listener-side is what's soft. None of these are +reaching their full design intent. + +## Soft-pass justification audit (per scenario) + +Re-read each scenario's kdoc. The soft-pass is honest right now +(captured 0 frames means harness flake, not regression). Once the +relay-routing race is fixed, the listener side becomes deterministic +and the soft-pass is no longer load-bearing — at that point, the +soft-pass HIDES regressions (a real T8/T11/T13 break could land in +a 0-frame outcome and pass vacuously). + +## Tighten plan + +### Step 1 — confirm sweep stability + +After the routing investigation lands, run: + +``` +for i in 1 2 3 4 5; do + echo "=== run $i ===" + ./gradlew :nestsClient:jvmTest \ + --tests HangInteropTest \ + --tests BrowserInteropTest \ + -DnestsHangInterop=true \ + -DnestsBrowserInterop=true \ + --rerun-tasks 2>&1 | grep -E "FAILED]|BUILD" +done +``` + +5/5 BUILD SUCCESSFUL with 0 `FAILED` lines = stability achieved. + +### Step 2 — replace each soft-pass with a hard floor + +For each scenario in the table above, remove the +`if (pcm.size <= warmupSamples) return` short-circuit and replace +with a meaningful sample-count floor. Tighten thresholds based on +observed steady-state numbers (see each scenario's kdoc for what +"steady-state" looks like — most run ≥ 1 s of audio in green-state). + +Concretely: + +- **Late-join**: `assertTrue(pcm.size >= ...)` floor. + Steady-state captures ~3 s on a 5 s broadcast with 2 s late-join, + minus warmup. Threshold: `≥ 1.5 s` — comfortably under the + steady-state but well over zero. +- **Mute-window**: tighten the upper bound (current 5.5 s) to + ~5.0 s. Add a lower bound asserting `≥ 2.5 s` — proves audio + arrived AND the muted segment shortened the total. +- **I14**: re-add `decoderOutputs >= 4` (3 warmup + ≥ 1 audio). + Current absence-only assertion is partial coverage. +- **Browser publisher baseline + reconnect**: remove the + vacuous-pass branches, add a `≥ 0.5 s of audio after warmup` + floor for baseline and `≥ 2.5 s` for reconnect (matches the + hang-tier I7 threshold). + +### Step 3 — reverify + +Re-run the 5× sweep. All scenarios must hard-pass 5/5. If a +scenario flakes after tightening, the relay-routing investigation +isn't fully done and we revert the tightening on that scenario +until it is. + +### Step 4 — update the gap matrix + +`nestsClient/plans/2026-05-06-cross-stack-interop-test-gap-matrix.md` +currently lists I14 with "browser ⏳" pending; flip to "✅" once +its hard floor is in. Same for any I-scenarios that now have +hard floors on both tiers. + +## Acceptance criteria + +- All BrowserInteropTest scenarios run with hard sample-count + AND FFT-peak assertions (no `return@runBlocking` short-circuits + on pcm.size). +- All HangInteropTest scenarios already hard-pass — no change + needed there. +- Gap matrix updated to reflect hard-pass coverage on each T#. +- Results plan updated to remove the "soft-pass on flake" + language. + +## Risk: post-tightening flake + +If any scenario fail-flakes after tightening, the routing +investigation isn't really done. Don't paper over with a wider +threshold; that's the same trap as the soft-passes. Either: +(a) revert the tightening on that scenario and keep +investigating, OR +(b) widen the threshold ONLY if the new value still excludes +the regression mode the test was designed to catch.