Files
amethyst/nestsClient/plans/2026-05-07-framespergroup-production-rerun.md
Claude 32065901c8 docs(nests): T16 closure roadmap + 4 follow-up plans
Plans the path from 'infra shipped' to 'full coverage with
correct behaviours'. Five new plan docs in nestsClient/plans/:

1. 2026-05-07-t16-closure-roadmap.md — index + priority order.
   Three sequential steps + one independent track.

2. 2026-05-07-moq-relay-routing-investigation.md (Priority 1)
   — root-cause the moq-relay 0.10.x per-broadcast subscribe-
   routing race. Step-by-step: capture relay-side trace, write
   minimum reproducer, file upstream OR bump moq-relay version.
   Smoking gun + hypotheses already in place from the
   late-join-flake investigation; this plan turns that into
   actionable next steps.

3. 2026-05-07-tighten-cross-stack-assertions.md (Priority 2)
   — once the routing race is closed, replace the five soft-
   pass scenarios in BrowserInteropTest with hard floors.
   Lists each scenario, its current soft-pass behavior, and
   the proposed hard threshold.

4. 2026-05-07-cross-stack-interop-ci-gating.md (Priority 3)
   — re-add the hang-interop + browser-interop GitHub Actions
   jobs that were dropped in 6829ab727 / b94737de7. Includes
   the exact YAML to restore and a 10/10 sweep stability bar
   before merge.

5. 2026-05-07-framespergroup-production-rerun.md (independent
   track) — re-run the HCgOY two-phone field tests against
   current nostrnests production at multiple framesPerGroup
   values to settle whether the test pin (5) and production
   default (50) can converge.

Estimated total: 2.5-3.5 days of focused work to fully close
T16. After these four: I7 post-reconnect cliff and I12 GOAWAY
remain open as genuinely upstream-territory items, tracked in
their existing investigation docs.
2026-05-07 14:36:21 +00:00

5.1 KiB

Plan: re-run HCgOY field tests against current production

Status: specced — pickup ready (needs prod-rig access). Cross-ref: nestsClient/plans/2026-05-07-framespergroup-reconciliation.md documents the conflict between the cliff plan's value (5) and HCgOY's value (50). This plan settles which is current truth.

What we're trying to settle

Production currently runs NestMoqLiteBroadcaster.DEFAULT_FRAMES_PER_GROUP = 50 based on HCgOY two-phone field tests at commit 6e4df4a (2026-05-05) which observed:

framesPerGroup streams/sec @ 50 fps observed cliff window
1 50 ~3 s
5 10 ~13 s
10 5 ~16 s
50 1 not reached
100 0.5 never observed

Interop tests pin 5 because the local --auth-public "" minimal relay setup hits a different cliff (per-stream byte volume) at framesPerGroup = 50.

The production deployment may have changed since 2026-05-05:

  • nostrnests may have updated their moq-relay version
  • their resource limits may have shifted
  • the upstream kixelated/moq may have addressed one or both cliffs

We don't know without re-running. Cliff value at framesPerGroup = 5 may now be unbounded — if so, both rigs can converge on 5 and the test pin matches the prod default.

Test setup

Rig A — interop env (already exists)

./gradlew :nestsClient:jvmTest --tests HangInteropTest -DnestsHangInterop=true runs against local moq-relay 0.10.25 --auth-public "" --tls-generate localhost. Long-broadcast scenario long_broadcast_60s_tone_round_trips is the existing 60-second sustained-stream test pinned at framesPerGroup = 5.

To probe other values, parameterize the helper:

// HangInteropTest.kt — runSpeakerToHangListen helper
private suspend fun runSpeakerToHangListen(
    speakerSeconds: Int,
    framesPerGroup: Int = 5,  // ← new parameter
    // ...existing params
): HangListenOutput { ... }

Then add scenarios long_broadcast_60s_framesPerGroup_50 etc. that pin different values. Expected outcomes today (per the 2026-05-01 cliff plan):

  • framesPerGroup = 5 — passes (current pin)
  • framesPerGroup = 10 — passes
  • framesPerGroup = 50 — fails (per-stream byte volume cliff at the local minimal relay)

Rig B — production deployment (needs maintainer access)

This is the gap. Rerunning the HCgOY two-phone field test pattern needs:

  • Two physical Android devices
  • A nostrnests room (production endpoint wss://nostrnests.com/v0/ws per NestsConnect.kt)
  • The diagnostic-build of Amethyst that emits the cliff-detector trace logs (see commit 6e4df4a's logcat run from 18:37:43..18:38:08)

The maintainer should run the same test pattern at:

  • framesPerGroup = 5 (current test value)
  • framesPerGroup = 10
  • framesPerGroup = 25 (untested, midpoint)
  • framesPerGroup = 50 (current prod value)
  • framesPerGroup = 100 (full group; the cliff plan's fpg-all reference)

For each, broadcast for 120 s and observe:

  • Total streams forwarded by the relay
  • Time-to-cliff if any (when the listener-side flow-control snapshot stops incrementing peerInitiatedUni)
  • Audio dropouts (perceptual + sample-count)

Decision matrix after data lands

Rig A passes at Rig B passes at Decision
5, 10 5, 10, 25, 50, 100 Keep prod 50; test pins 5 (current state)
5, 10, 50 5, 10, 25, 50, 100 Both rigs converge → unify on 50, test pin matches prod
5, 10 50, 100 only (5 still cliffs) Current state is correct; document permanently as "two cliffs in one binary"
5 only 50, 100 only (5 still cliffs) The two cliffs are real; consider per-environment config
5, 10, 50 50, 100 only (5 still cliffs) Test rig fixed; production cliff still hits at 5. Test pin doesn't catch prod regression — bigger problem.

The "decision" column drives the production-side change (or non-change) to DEFAULT_FRAMES_PER_GROUP.

What lands as code

After Rig B data is in:

  1. Update kdoc on NestMoqLiteBroadcaster.DEFAULT_FRAMES_PER_GROUP citing the new run's logcat dates / commit.
  2. If the values converge, change the default to match. Update the test-side framesPerGroup = 5 pin to match the new default — keep both rigs aligned.
  3. If values still diverge, document it explicitly as a known environment-dependent value. Consider exposing framesPerGroup as a per-deployment config (currently only exposed as a constructor parameter — wire to a config knob if product wants per-deployment tuning).
  4. Update nestsClient/plans/2026-05-07-framespergroup-reconciliation.md's "Recommendation" section with the data-driven outcome.

Acceptance criteria

  • A logcat dump from Rig B with framesPerGroup = 5 for ≥ 60 s showing whether the cliff still hits at ~13 s.
  • Decision logged in the framesPerGroup reconciliation doc.
  • If a value change lands, the test-side pin and production default agree.

Out of scope

  • The local interop env's per-stream byte cliff at framesPerGroup = 50. That's a separate thread; addressing it would require either a different relay configuration or patching moq-relay itself.