6f649e76d8c2c6d9c1433b9ecccb546d8d9f16d2
12890 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
6f649e76d8 |
refactor(nests): extract MoqLiteBroadcastHandle + HotSwappablePublisherSource (Audit-10)
`MoqLiteNestsSpeaker.kt` mixed three top-level concerns at 387 lines:
- The `MoqLiteNestsSpeaker` class itself — the speaker entry point
that builds a publisher + broadcaster pair on
`startBroadcasting`.
- `MoqLiteBroadcastHandle` — internal `BroadcastHandle`
implementation tying the broadcaster, audio publisher, and
catalog publisher together with a fixed-order shutdown.
- `HotSwappablePublisherSource` — internal interface that lets
`ReconnectingNestsSpeaker.runHotSwapIteration` retarget a
long-lived broadcaster onto fresh moq-lite session publishers
without restarting the AudioRecord / Opus encoder pipeline.
The handle and the interface are independently reachable from
`ReconnectingNestsSpeaker` and have no behavioural coupling to
`MoqLiteNestsSpeaker` beyond a `parent` reference (handle) or an
`as?` cast (interface). Move each to its own file:
- `MoqLiteBroadcastHandle.kt` (109 lines).
- `HotSwappablePublisherSource.kt` (62 lines).
`MoqLiteNestsSpeaker.kt` is now 276 lines focused on the speaker
class. Same package, same `internal` visibility — no call-site
changes needed elsewhere. Tests + spotless green.
https://claude.ai/code/session_014JfZJHSTvyYYWJbC9VbB47
|
||
|
|
fb7b6b7cd9 |
refactor(nests): split MoqLiteMessages.kt by concern (Audit-11)
The original `MoqLiteMessages.kt` mixed three categories of declaration:
- One `MoqLiteAlpn` object — the WebTransport sub-protocol
advertisement strings.
- Four enums — `MoqLiteControlType`, `MoqLiteDataType`,
`MoqLiteAnnounceStatus`, `MoqLiteSubscribeResponseType` — that
are wire-format discriminators the codec reads at message
boundaries to choose which body to decode.
- Eight data classes / objects — `MoqLiteAnnouncePlease`,
`MoqLiteAnnounce`, `MoqLiteSubscribe`, `MoqLiteSubscribeOk`,
`MoqLiteSubscribeDrop`, `MoqLiteSubscribeDropCode`,
`MoqLiteGroupHeader`, `MoqLiteProbe` — the body shapes themselves.
317 lines, three concerns, one file. Split by responsibility:
- `MoqLiteAlpn.kt` — the ALPN object alone.
- `MoqLiteControlCodes.kt` — the four discriminator enums.
- `MoqLiteMessages.kt` — body data classes only.
Same package, same visibility, identical wire behaviour. Imports
across the codebase resolve to the new locations automatically
since everything stayed in `com.vitorpamplona.nestsclient.moq.lite`.
Tests pass unchanged.
https://claude.ai/code/session_014JfZJHSTvyYYWJbC9VbB47
|
||
|
|
a96ec9db00 |
chore(nests): standardise Log import in audio-pipeline files
Audit-12 cleanup. The audio-pipeline sources mixed two import styles:
some files used `com.vitorpamplona.quartz.utils.Log.d/w/e` fully-
qualified (NestPlayer's 13 sites, AudioTrackPlayer's 8, two each in
the encoder/decoder, three in NestMoqLiteBroadcaster) while
AudioRecordCapture had already migrated to a top-level
`import com.vitorpamplona.quartz.utils.Log`. The lambda overload
(`Log.d(tag) { lazyMessage }`) is what the find-non-lambda-logs
skill enforces and is the only style the rest of the codebase uses;
the fully-qualified form is verbose noise on every call site.
Add the `import com.vitorpamplona.quartz.utils.Log` line to each
file and rewrite call sites to bare `Log.d` / `Log.w` / `Log.e`. No
behaviour change; logs still go through PlatformLog with the same
tag and lazy-message contract. Five files updated, ~30 call sites.
https://claude.ai/code/session_014JfZJHSTvyYYWJbC9VbB47
|
||
|
|
8a49486b37 |
fix(nests): G4+G5 plumb catalog sampleRate through decoder + AudioTrack
G4: Audit-9 (catalog-driven decoder reconfig) plumbed numberOfChannels
end-to-end but left sampleRate hardcoded at AudioFormat.SAMPLE_RATE_HZ
in MediaCodecOpusDecoder.buildOpusIdHeader / buildFormat and in
AudioTrackPlayer's MinBufferSize / 250 ms target / setSampleRate
calls. For Opus this is benign in practice — Codec2 always emits
48 kHz PCM regardless of OpusHead inputSampleRate — but hardcoding
the constant means a future codec or container variant whose
decoder DOES respect input sample rate would mis-clock playback.
And the OpusHead identification header should match what the
catalog declares either way.
Thread sampleRate alongside channelCount through every layer:
- MediaCodecOpusDecoder constructor takes
`sampleRate: Int = AudioFormat.SAMPLE_RATE_HZ`. buildFormat
and buildOpusIdHeader both take it as a parameter.
- AudioTrackPlayer constructor takes the same. Used in
AudioTrack.getMinBufferSize, the 250 ms-equivalent target
bytes calculation (`(sampleRate / 4) * BYTES_PER_SAMPLE *
channelCount`), AndroidAudioFormat.Builder.setSampleRate, and
the diagnostic log.
- decoderFactory and playerFactory in NestViewModel become
`(channelCount: Int, sampleRate: Int) -> ...`.
- awaitDecoderChannelCount → awaitAudioPipelineConfig, returning
a private `AudioPipelineConfig(channelCount, sampleRate)`
struct so openSubscription handles both fields uniformly.
`sampleRate` falls back to AudioFormat.SAMPLE_RATE_HZ on
timeout / non-positive declaration with a warning log.
- NestViewModelFactory + NestViewModelTest updated to the
two-arg factory shape.
G5: documented the SUBSCRIBE_BUFFER safety budget on
CATALOG_AWAIT_TIMEOUT_MS's kdoc. With the current 500 ms timeout +
SUBSCRIBE_BUFFER = 64-frame DROP_OLDEST flow + 50 fps Opus, at
most 25 frames buffer during the wait — leaving ≥ 39 frames of
margin (≈ 780 ms) before the oldest frame would be evicted. Even
at the production framesPerGroup = 50 (1 group/sec) cadence this
never trips during normal startup. No code change; just pinning
the rationale so a future timeout bump is checked against the
buffer size.
https://claude.ai/code/session_014JfZJHSTvyYYWJbC9VbB47
|
||
|
|
75f572ba3d |
test+fix(nests): pin stripLegacyTimestampPrefix; primaryAudio filters to legacy
Two audit-pass gaps:
G1 — `MoqLiteNestsListener.Companion.stripLegacyTimestampPrefix`
is the entire receive-side wire-format converter for audio frames
(every web speaker's Opus packet flows through it) and had zero
unit tests. A regression here is invisible to users — silent decode
failure of every web speaker — and to interop tests, because both
sides agree on the same bug. New StripLegacyTimestampPrefixTest pins:
- all four QUIC varint-length tiers (1 / 2 / 4 / 8 bytes), one
test per tier with a representative timestamp value plus a
tag-byte assertion so the test fails loudly if Varint.encode's
tier-selection changes.
- empty-payload and malformed-payload fast paths (returns
payload unchanged, same instance — no allocation).
- round-trip against `NestMoqLiteBroadcaster`'s exact wire
shape (`varint(timestamp_us) + raw_opus_packet`) across five
timestamp values spanning all four tiers.
G6 — `RoomSpeakerCatalog.primaryAudio()` previously returned
`audio.renditions.values.firstOrNull()` with no container filter.
A future publisher that emits CMAF before legacy in iteration
order would surface the CMAF rendition, and our decoder pipeline
would then feed CMAF MOOF/MDAT bytes to its legacy
`varint(ts)+opus` parser and decode garbage. Filter to
`container.kind == Container.LEGACY_KIND` so:
- mixed CMAF+legacy publishers always surface the legacy entry
regardless of map iteration order
- CMAF-only publishers correctly return null (caller falls
back to "unknown codec / no audio")
- publishers that omit `container` entirely also return null —
treating "no container declared" as "legacy by default" would
silently mis-decode a CMAF-only publisher that just forgot
to spell out `container.kind`
LEGACY_KIND const lives on `Container.Companion` so call sites
share one canonical spelling. Three new test cases pin the new
filter behaviour (mixed iteration order, CMAF-only, missing
container). The pre-existing `toleratesUnknownKeys` test is
updated to declare a legacy container — the unknown-key
tolerance we're checking is about unrecognised siblings, not
container-presence.
https://claude.ai/code/session_014JfZJHSTvyYYWJbC9VbB47
|
||
|
|
a94d12638f |
perf(nests): bound encoder CSD loop, single-alloc audio framing, cache catalog JSON
Three audit follow-ups, all in the audio publish hot path: Audit-3: MediaCodecOpusEncoder.encode could busy-loop on a buggy encoder that emits BUFFER_FLAG_CODEC_CONFIG on every dequeue. The existing format-change path has a one-shot `formatChangeAbsorbed` guard for exactly this reason; the new CSD-skip path didn't. Cap consecutive CSD skips per encode call at MAX_CSD_SKIPS_PER_CALL=4 (generous: Codec2 emits 1 OpusHead or 2 OpusHead+OpusTags in practice). On overshoot, log a warning and return ByteArray(0) — the broadcaster's `if (opus.isEmpty()) continue` already handles that contract. Audit-6: NestMoqLiteBroadcaster's per-frame send path allocated twice per Opus packet — once for the timestamp Varint and once for the concatenated `varint+opus` payload. At 50 fps × N speakers that's measurable young-gen pressure. Switch to the same single-allocation pattern PublisherStateImpl.send already uses: compute Varint.size(timestampUs), allocate the final payload buffer once, write the varint directly into it via Varint.writeTo, then copy the opus bytes. Drops audio-thread allocations from 3/frame to 1/frame (the third was the wrap in PublisherStateImpl.send, already optimised). Audit-7: MoqLiteNestsSpeaker.startBroadcasting and ReconnectingNestsSpeaker.runHotSwapIteration both encode the same fixed catalog JSON on every call — once per broadcast start and once per JWT-refresh hot-swap iteration. Cache the encoded bytes on MoqLiteHangCatalog.Companion.OPUS_MONO_48K_AUDIO_DATA_JSON_BYTES so the kotlinx.serialization run happens at class-init time and both call sites read a static reference. Trivial perf, but co-locates the "default Amethyst publish shape" in one place. https://claude.ai/code/session_014JfZJHSTvyYYWJbC9VbB47 |
||
|
|
ade8da3b5b |
fix(nests): NestPlayer keeps existing decoder when boundary factory throws
Audit-1: the publisher-boundary rebuild path released the old decoder
BEFORE asking the factory for a replacement (`runCatching {
decoder.release() }; decoder = factory()`). If `factory()` threw —
MediaCodec contention, audio policy denial mid-session, etc. — the
released decoder stayed assigned to the field and every subsequent
`decoder.decode(payload)` would throw `IllegalStateException` with
no recovery path. The room would silently fail for that subscription
until torn down.
Build the replacement first; only release the old one and swap on
success. On factory failure, log and keep the existing decoder
running. Cross-publisher Opus predictor state is wrong for one
group but at least audio keeps playing.
Audit-4+5: companion test cleanup —
- Drop the `byteArrayOf(0x0A) + it` / `0x0B` dead expressions in
the FakeOpusDecoder closures of the existing boundary-rebuild
test. They allocated a ByteArray and concatenated, then threw
the result away.
- Drop the local `IntBox` helper and use
`java.util.concurrent.atomic.AtomicInteger` like the rest of
the codebase does (`MoqLiteSession`, `MoqLiteNestsListener`).
Same semantics, no new vocabulary.
New test pins the dangling-field fix:
`publisher_boundary_keeps_old_decoder_when_factory_throws` flows
two MoqObjects with different trackAliases through a NestPlayer
whose factory always throws; asserts both frames decode through
the original decoder and the original decoder is released exactly
once on `stop()`.
https://claude.ai/code/session_014JfZJHSTvyYYWJbC9VbB47
|
||
|
|
7e76ab1139 |
fix(nests): T11 drop bestEffort=true on moq-lite group uni streams
`MoqLiteSession.openGroupStream` was opening each group's QUIC uni stream with `bestEffort = true`. `:quic`'s `SendBuffer.markLost` with that flag drops lost STREAM ranges WITHOUT retransmit AND WITHOUT `RESET_STREAM` (`quic/src/commonMain/kotlin/com/vitorpamplona/quic/ stream/SendBuffer.kt:300-309`). The peer's `ReceiveBuffer` ends up with chunks at `[0, P)` and `[Q, end)` and a permanent hole at `[P, Q)`; the application's `incoming` Flow parks on the hole forever. Web watchers (`@moq/hang` `Container.Consumer`) park their `Group.readFrame` until the relay's 30 s `MAX_GROUP_AGE` ages the broadcast queue out — manifesting as a 30 s silent dropout per lost packet on lossy networks (cellular, mobile WiFi, congested home routers). This is a real-world bug that's invisible on a clean LAN. The reference implementation (kixelated/moq-rs `Publisher::serve_group`, `rs/moq-lite/src/lite/publisher.rs:347-406`) writes to reliable QUIC streams with no `set_unreliable` call — `bestEffort` was Amethyst's private optimisation with no peer-side support. Drop it; let `:quic` retransmit lost ranges normally. A retransmit arriving 50–150 ms late still falls inside hang's default ~200 ms jitter buffer, so the cost is marginal extra bandwidth on retransmits and the win is no more silent dropouts on lossy networks. T11.2 — orthogonality check: stream-cliff fix is independent. Re-read `nestsClient/plans/2026-05-01-quic-stream-cliff-investigation.md` and grep'd both the plan and `NestMoqLiteBroadcaster.kt` for `bestEffort` / `best_effort` — neither references it. The cliff fix is `framesPerGroup = 5/50` (cadence reduction); load-bearing on stream-creation RATE, not loss handling. Drop is safe. T11.3 (stream priority — newer groups drain first under congestion) deferred to a follow-up commit; hooking it into `:quic`'s send-frame loop is bigger than a one-line change and warrants its own task. The kixelated/moq-rs publisher uses `stream.set_priority(priority.current())` to bias the writer; without it, our drain order under congestion is FIFO across streams rather than newest-first, which can mean the listener catches up on a stale group when a fresh one is more useful. Doesn't block today — production audio rarely hits transport congestion at 1 group/sec cadence. https://claude.ai/code/session_014JfZJHSTvyYYWJbC9VbB47 |
||
|
|
4714e3c723 |
fix(nests): T13 reset Opus decoder on publisher boundary in NestPlayer
The reissuing-subscribe wrapper splices fresh-publisher frames into the
same `SharedFlow` that NestPlayer consumes, so across a JWT-refresh
hot-swap (speaker side) or a cliff-detector recycle (listener side)
the decoder receives a discontinuous frame stream while keeping its
Opus predictor state. Result: an audible warble at every publisher
cycle. Single-stream tests don't catch it because they never present
two distinct publishers.
Detect the boundary via `MoqObject.trackAlias` — the underlying
`MoqLiteSession.subscribe` assigns a fresh subscribeId per SUBSCRIBE,
which the listener wrapper surfaces verbatim as `trackAlias` on every
emitted object. A change between consecutive objects = new publisher.
Add an optional `decoderFactory: (() -> OpusDecoder)?` to NestPlayer.
When non-null, NestPlayer tracks `lastTrackAlias` and on a change
releases the current decoder and rebuilds via the factory. The
default `null` preserves the legacy single-decoder behaviour so
existing tests / callers that don't care about boundaries stand
unchanged.
NestViewModel.openSubscription wires the factory: a closure capturing
the catalog-derived `channelCount` so a rebuild reuses the SAME
channel layout — without that capture, a rebuild after a stereo-
publisher cycle would default to mono and silently downmix.
Two new NestPlayerTest cases pin the behaviour:
- `publisher_boundary_rebuilds_decoder_when_factory_provided`:
factory invoked twice (initial + boundary), first decoder
released on boundary, second released on stop.
- `publisher_boundary_no_op_when_factory_is_null`: legacy path
holds onto the same decoder across trackAlias changes (unchanged
semantics).
NestPlayer's first constructor parameter is now `initialDecoder`
(was `decoder`); positional-arg call sites are unchanged but the
named-arg call sites in the test suite are updated accordingly.
https://claude.ai/code/session_014JfZJHSTvyYYWJbC9VbB47
|
||
|
|
be4e0b9f98 |
fix(nests): T12 carry audio group sequence across hot-swaps
Every JWT-refresh hot-swap reset the audio publisher's group sequence
counter to 0, because `PublisherStateImpl` always initialised
`nextSequence = 0L`. kixelated/hang's `Container.Consumer.#run`
discards any consumer with `sequence < this.#active`, so a watcher
whose `#active` had advanced to e.g. 50 from the previous session's
publisher would silently drop every group on the new session until
either `#active` rolls over or the watcher re-subscribes — audible as
"speaker goes silent for a minute every 9 minutes" (the proactive
JWT refresh window).
Carry the sequence forward end-to-end:
- MoqLiteSession.publish gains a `startSequence: Long = 0L`
parameter; PublisherStateImpl seeds `nextSequenceField` from it
instead of hard-coding 0.
- MoqLitePublisherHandle exposes a `nextSequence: Long` snapshot.
`@Volatile`-backed so the hot-swap caller can read it without
contending the publisher's gate; race-window between read and
a concurrent send is sub-millisecond and would only produce a
single duplicate sequence in the very unlikely overlap, which
listeners tolerate.
- HotSwappablePublisherSource.openPublisherForHotSwap takes
`startSequence: Long = 0L`. MoqLiteNestsSpeaker passes it
through to session.publish.
- NestMoqLiteBroadcaster exposes its current publisher via a
read-only `currentPublisher` so the hot-swap pump in
ReconnectingNestsSpeaker.runHotSwapIteration can read its
`nextSequence` before opening the replacement publisher and
pass it as the seed.
- Catalog publisher keeps `startSequence = 0L` (the default) —
catalog isn't subject to the same `#active` accumulation
because each new SUBSCRIBE triggers a fresh emit-on-subscribe
write that resets the watcher's catalog state.
Listener-side change is none — the watcher already reads whatever
sequence we send. A new MoqLiteSessionTest pins the contract:
publishing with startSequence=42 makes the first uni stream's
GroupHeader.sequence == 42 and advances to 43 after the first send.
https://claude.ai/code/session_014JfZJHSTvyYYWJbC9VbB47
|
||
|
|
c23da52795 |
fix(nests): T10 endGroup() on unmuted→muted transition
When the user mutes mid-broadcast, `NestMoqLiteBroadcaster`'s send loop just `continue`s past the mute check, leaving the current group's QUIC uni stream open with no FIN. The watcher's `Container.Consumer.#runGroup` parks on `await group.consumer.readFrame()` waiting for either another frame or a stream FIN — and gets neither. kixelated/hang's UI surfaces this as a "stalled" indicator on the speaker tile while we're muted, which is wrong: we're muted, not stalled. FIN the open uni stream once on the unmuted→muted edge via a `wasMuted` latch that the muted→unmuted edge clears. Doesn't change the timestamp counter — it still advances on muted frames so an unmute's first timestamp reflects the muted duration as a real wall- clock gap (not a collapse to zero). `runCatching` around endGroup because a transport-side close racing with the FIN is non-fatal — the broadcaster's terminal-failure path picks up real breakage on the next live frame. Also resets `framesInCurrentGroup` to 0 so the post-unmute send path opens a fresh group cleanly via the standard `currentGroup ?: openNextGroupLocked()` branch in `PublisherStateImpl.send`. https://claude.ai/code/session_014JfZJHSTvyYYWJbC9VbB47 |
||
|
|
96cfa1235a |
fix(nests): T8 skip BUFFER_FLAG_CODEC_CONFIG outputs in MediaCodecOpusEncoder
Android's `audio/opus` MediaCodec encoder emits `BUFFER_FLAG_CODEC_CONFIG` output buffers BEFORE any audio buffer — typically the 19-byte OpusHead identification header per RFC 7845, and (on some Codec2 stacks) the OpusTags comment header. These are decoder-config blobs, NOT audio frames. We weren't filtering them, so the first 1–2 wire frames every encoder lifetime were OpusHead bytes wrapped in our legacy-container varint(timestamp_us) prefix. The web watcher's WebCodecs `AudioDecoder` handles this by burning warmup slots (`@moq/watch/src/audio/decoder.ts`'s `warmed <= 3` policy) — so the listener mostly recovers — but on Android Codec2 stacks that emit BOTH OpusHead AND OpusTags as separate CSD buffers, two of the three warmup slots get absorbed by metadata and the listener hears a tiny click on the next group rollover. The hot-swap path (`ReconnectingNestsSpeaker`) repeats the warmup on every JWT refresh, so this fired once every 9 minutes in production. Filter via `bufferInfo.flags and MediaCodec.BUFFER_FLAG_CODEC_CONFIG` inside `encode`'s output-drain loop. CSD buffers are released and the loop continues to the next dequeue rather than returning the bytes. Logged once per encoder lifetime via `loggedCsdSkip` so we have proof the path fired without flooding logcat — a stack that emits CSD on every frame would otherwise be noisy. Returning `ByteArray(0)` for the first call (because the only output was a CSD + format-change pair) is unchanged behaviour and the broadcaster's `if (opus.isEmpty()) continue` already handles it. https://claude.ai/code/session_014JfZJHSTvyYYWJbC9VbB47 |
||
|
|
73722d2ad2 |
fix(nests): T14 recognise Goaway control type instead of silent FIN
moq-rs's `Publisher::run` accepts `ControlType::Goaway = 5` as a graceful-shutdown signal that asks the publisher to migrate to a different relay node (`rs/moq-lite/src/lite/publisher.rs`). Our enum only defined Session/Announce/Subscribe/Fetch/Probe; an inbound Goaway bidi fell through `MoqLiteControlType.fromCode` as `null`, hit the unknown-control branch in `handleInboundBidi`, and was FIN'd silently — losing the relay's shutdown notification entirely. Add `Goaway(5L)` to the enum and a dedicated arm in `handleInboundBidi` that logs the event and FINs cleanly. We don't act on the migration request today (no body decode, no preferred-relay failover); the `connectReconnecting*` wrappers' transport-loss reconnect path already recovers from the eventual hard disconnect, so all this arm needs is to surface the relay's intent in logcat instead of swallowing it. Wire body decoding is left as a follow-up. https://claude.ai/code/session_014JfZJHSTvyYYWJbC9VbB47 |
||
|
|
ea90685a86 |
revert(nests): drop moq-lite-04 ALPN advertisement (codec is wire-incompatible)
Re-investigation of `kixelated/moq` commit 45db108 ("moq-lite/moq-relay:
hop-based clustering") revealed that Lite-04 IS wire-incompatible with
our Lite-03 codec on Announce / AnnounceInterest / Probe — contradicting
the earlier reasoning that motivated
|
||
|
|
1c47fd2403 |
feat(nests): drive Opus decoder + AudioTrack channel count from catalog
Web speakers (kixelated/hang reference) emit stereo Opus when the
user's AudioContext picks a stereo input — and our catalog now lets
us know that ahead of decoder construction. Previously the decoder
+ AudioTrack were both hardcoded to mono via AudioFormat.CHANNELS, so
a stereo web publisher's frames either threw on the Opus TOC byte
mismatch, decoded with downmix artifacts, or output left-channel-only
depending on the device's MediaCodec implementation.
Wire the catalog-discovered numberOfChannels through:
- MediaCodecOpusDecoder accepts `channelCount: Int = 1`. Drives
MediaFormat.createAudioFormat's channel count, the OpusHead CSD-0
channel byte, and the per-frame output buffer size (stereo frame =
2× shorts vs mono). Validates 1..2 — RFC 7845 mapping family 0
covers both mono and stereo without an explicit channel mapping
table; multichannel needs family 1 which we don't support.
- AudioTrackPlayer accepts `channelCount: Int = 1`. Drives the
AudioTrack channel mask (CHANNEL_OUT_MONO vs CHANNEL_OUT_STEREO)
and the 250 ms wall-clock buffer-size target (stereo doubles bytes
per sample).
- NestViewModel.openSubscription now starts the catalog fetch
BEFORE waiting for it, then awaits `_speakerCatalogs[pubkey]` for
CATALOG_AWAIT_TIMEOUT_MS (500 ms) before constructing the decoder
+ player. With the speaker-side emit-on-subscribe hook in place
(
|
||
|
|
87b43d4d78 |
fix(nests): bump preroll to 200 ms + surface persistent decode failures
Two audit follow-ups for receive-path interop with kixelated/moq web publishers (NestsUI v2's reference encoder): 1. ROOM_PLAYER_PREROLL_FRAMES 5 → 10 (≈100 ms → ≈200 ms). The web reference encoder uses `groupDuration: 100ms` (5 Opus frames per group) and stacks another 100–200 ms of jitter on top via the relay's outbound forward queue under residential network conditions. 100 ms of preroll reliably underran on web speakers, producing audible clicks between adjacent groups; 200 ms covers the observed envelope while keeping perceived join latency well below the wire latency a listener already sees. 2. Per-subscription consecutive Opus decode-error counter in NestViewModel.openSubscription. Single-frame decode errors are normal noise — Opus PLC papers over a single bad frame and the prior code silently swallowed them — but a structural mismatch (wrong wire format, missing legacy-container varint strip, codec/channel-count mismatch, corrupted Opus) fails *every* frame in a row. The previous behaviour made exactly that class of bug invisible: no audio, no log, no UI signal. New logic increments a per-subscription AtomicLong on each DecoderError / EncoderError, resets on every successful onLevel tick. When the streak hits DECODE_ERROR_STREAK_LOG_THRESHOLD (50 frames ≈ 1 s of solid failures) we log ONE conspicuous warning tagged `NestRx` with the underlying throwable's stacktrace. Logged once per crossing (exact-equality, not modulus) so a stuck stream produces a single attention-grabbing line instead of a 50 Hz flood. Picks the non-lambda Log.w overload deliberately so the throwable stacktrace survives — fires exactly once per stuck-stream streak, so the unconditional message-string build is fine. Both changes are additive — no callers changed, no public API touched. The stale `ROOM_PLAYER_PREROLL_FRAMES = 5` reference in NestMoqLiteBroadcaster's framesPerGroup kdoc is updated to match. https://claude.ai/code/session_014JfZJHSTvyYYWJbC9VbB47 |
||
|
|
3417e44a6a |
refactor(nests): catalog emit-on-subscribe hook replaces 2s republish loop
The catalog publisher's previous shape was "send once at startup
(silently fails if no subs are attached yet) + a 2 s republish loop".
That worked but had two problems:
1. The startup send is a no-op (PublisherStateImpl.send returns false
when inboundSubs is empty), so the catalog isn't actually on the
wire until the first loop tick — up to 2 s of catalog dead time
for a watcher that subscribes immediately.
2. Re-emitting a static blob every 2 s for the broadcast's lifetime
wastes a fresh group on the relay's per-track cache; not a hot path
today but unnecessary.
Replace with an emit-on-subscribe hook. New API on MoqLitePublisherHandle:
fun setOnNewSubscriber(hook: (suspend () -> Unit)?)
PublisherStateImpl fires the hook once per accepted (track-matching)
inbound SUBSCRIBE, OUTSIDE its serialisation lock so the hook can
safely call send/endGroup without deadlock. Track-mismatched subs
(SubscribeDrop reply) do NOT fire the hook — covered by a new
"hook does not fire on track mismatch" test.
MoqLiteNestsSpeaker.startBroadcasting and ReconnectingNestsSpeaker's
hot-swap iteration both now register a hook that writes the catalog
JSON on every fresh subscribe instead of looping. The
catalogRepublishJob field on MoqLiteBroadcastHandle and the
hotSwapCatalogJob field on ReissuingBroadcastHandle (plus their close
paths) are dropped.
Net effect for hang.js / NestsUI-v2 watchers:
- Catalog reaches the watcher on the SAME group as the subscribe ack
(no 0–2 s startup gap).
- Late joiners hit the relay's track-latest cache for the catalog
blob — a fresh hook fire only happens when the relay opens a new
SUBSCRIBE bidi to us (cliff-detector recycle on listener side, or
JWT-refresh on our side).
- One catalog group per relay-side subscribe instead of one every
2 s for the broadcast lifetime.
Two tests pin the new behaviour: hook fires on inbound subscribe;
hook does NOT fire on a track-mismatched subscribe (which gets a
SubscribeDrop reply and never enters inboundSubs).
https://claude.ai/code/session_014JfZJHSTvyYYWJbC9VbB47
|
||
|
|
79c76d5831 |
fix(nests): reply SubscribeDrop on unsupported tracks instead of silent FIN
When a relay opens a SUBSCRIBE bidi for a track this session doesn't publish (e.g. a watcher subscribes to `video/data` against a broadcast that only declares `audio/data` + `catalog.json`), the session was silently FINing the bidi without writing any reply. The peer's wait on its receive side resolves to end-of-stream with no indication WHY — indistinguishable from "publisher disappeared mid-subscribe." Reply with [MoqLiteSubscribeDrop] carrying [MoqLiteSubscribeDropCode.TRACK_DOES_NOT_EXIST] (0x04, mirroring the IETF-MoQ `ErrorCode.TRACK_DOES_NOT_EXIST` convention) and an informational reason phrase listing the tracks we DO publish, then FIN. Matches what kixelated's `rs/moq-lite/src/lite/subscribe.rs` expects on this path. Doesn't change the happy path — every track NestsUI-v2 subscribes to (`audio/data` + `catalog.json`) is now published — but is the spec-correct behaviour for any future watcher that prospectively subscribes to a track our publisher hasn't declared. https://claude.ai/code/session_014JfZJHSTvyYYWJbC9VbB47 |
||
|
|
4b736263e0 |
feat(nests): advertise both moq-lite-03 and moq-lite-04 ALPNs
Production moq-relay (moq.nostrnests.com) and the kixelated browser client both ship Lite-04 now. We previously only advertised moq-lite-03 in the wt-available-protocols header, so a Lite-04-only relay deployment would refuse the WebTransport CONNECT (no mutual sub-protocol). Add moq-lite-04 to the advertised list. The on-the-wire framing for Subscribe / Group / Announce did NOT change between Lite-03 and Lite-04, so the existing codec handles either selection unchanged — this is purely a handshake-layer addition. moq-lite-03 stays listed first so a relay that supports both still picks the previously-tested path. Constants now live on MoqLiteAlpn alongside LITE_03 and LEGACY rather than being a string literal in QuicWebTransportFactory. https://claude.ai/code/session_014JfZJHSTvyYYWJbC9VbB47 |
||
|
|
361dbbe201 |
fix(nests): stamp Opus frames at frame-index × frame-duration, not wall clock
The legacy-container timestamp on each audio frame was previously
captured at send-time via `frameStartMark.elapsedNow()`. That mark
advances on wall clock, but `current.send(...)` is a suspending call
that backs off on transport backpressure — so a frame captured at T
gets stamped at T+δ once the previous send drains, and the watcher's
WebCodecs AudioDecoder schedules playback at the wrong instant. Audible
as a slowly-growing offset between speaker and listener over the life
of a session that's seen any meaningful send-side queueing.
Replace with a frame-index counter pinned to the codec grid:
`timestampUs = nextFrameIndex * AudioFormat.FRAME_DURATION_US`. The
counter advances once per captured PCM frame BEFORE the encode/mute/
send path, so:
- Encoder failures, empty-encoded frames, and muted frames each
consume one slot — preserving the wall-clock gap on the wire so
a 5 s mute shows up as a 5 s timestamp jump (matches the prior
wall-clock behaviour for mute and silence).
- Send-time stalls don't drift the timestamp: it's computed at the
top of the iteration, not after the suspending send.
- Survives publisher hot-swap (single-broadcast, single-encoder;
the new publisher inherits the running counter) — same guarantee
the prior `TimeMark` had.
Also adds [AudioFormat.FRAME_DURATION_US] = 20_000 (960 / 48 kHz × 1e6)
so the constant lives next to FRAME_SIZE_SAMPLES rather than being
implicit in every call site.
https://claude.ai/code/session_014JfZJHSTvyYYWJbC9VbB47
|
||
|
|
1b0740d4a8 |
feat(nests): emit catalog jitter hint matching Opus frame duration
hang.js's encoder publishes a `jitter` field (ms) in every audio rendition; the watcher uses it to size its jitter buffer. Without it the watcher falls back to a built-in default — works in practice today but means our publishes deviate from the canonical hang shape and the deviation could surface as audible buffer-sizing differences on a future watcher revision that takes the field as required. Add `jitter: 20` to MoqLiteHangCatalog.AudioRendition (matching the Opus frame cadence: 960 samples at 48 kHz = 20 ms; same value hang.js hard-codes at `js/publish/src/audio/encoder.ts:#runConfig`). Updates the byte-exact MoqLiteHangCatalogTest assertion and the parallel inline-literal in RoomSpeakerCatalogTest so both sides stay pinned. The commons-side `RoomSpeakerCatalog` parser doesn't surface the new field — `JsonMapper`'s `ignoreUnknownKeys = true` lets the watcher ingest it silently. Adding a typed `jitter` field there is a forward-compat task; not needed for interop with hang today. https://claude.ai/code/session_014JfZJHSTvyYYWJbC9VbB47 |
||
|
|
ff9bb25921 |
refactor(nests): replace hand-built catalog JSON with typed MoqLiteHangCatalog
The two SPEAKER_CATALOG_JSON call sites (MoqLiteNestsSpeaker for the non-reconnecting path, ReconnectingNestsSpeaker for the JWT-refresh hot- swap path) both built the manifest as a `\\` + `\"` + `+`-concatenated string literal. One missed escape and the watcher's parser fails silently — the broadcast becomes invisible to hang.js with no indication of why. The two literals had also drifted independently in review (same content today; nothing prevents drift tomorrow). Replace both with `MoqLiteHangCatalog.opusMono48k(track).encodeJsonBytes()`, a Serializable data class that encodes via kotlinx.serialization with `encodeDefaults = false` + `explicitNulls = false` pinned (so a future global default-flip can't surface `"bitrate": null` on hang.js's parser). The new type lives in :nestsClient because :nestsClient does not depend on :commons and the parser-side `RoomSpeakerCatalog` (in :commons) isn't reachable from the publish path. The two classes target the same wire shape independently; a byte-exact assertion in the new MoqLiteHangCatalogTest plus the existing RoomSpeakerCatalogTest's inline-literal round-trip together pin both sides — desync on either end fails one of the two tests immediately. https://claude.ai/code/session_014JfZJHSTvyYYWJbC9VbB47 |
||
|
|
aa37a931e1 |
test(okhttp): widen SurgeDns stale-refresh TTL to defeat double-expiry race
The 1ms positiveTtlMs let the second refreshed entry expire again between the post-refresh assertion and the cache-hit lookup that followed, triggering an extra background refresh that broke the calls=2 invariant on slow CI runners. 100ms TTL with sleep(150) keeps the expiry boundary clearly outside the assertion window. |
||
|
|
5310b3f574 | Merge remote-tracking branch 'origin/main' into claude/fix-nests-audio-receiver-HCgOY | ||
|
|
5b7e347ac0 |
feat(nests): align moq-lite catalog with kixelated/hang shape
Switch RoomSpeakerCatalog to the canonical hang catalog format (camelCase audio.renditions[<track>] with container.kind), wrap audio frames in the legacy varint(timestamp_us) prefix on publish, and strip it on subscribe. Also reply to inbound Probe/Fetch/Session bidis with a bitrate hint or clean FIN instead of hanging. |
||
|
|
5389b32e52 |
feat(nests): keep catalog.json published across JWT refresh
ReissuingBroadcastHandle's hot-swap iteration now opens its own catalog.json publisher on every fresh session, pushes the manifest once, and launches the same 2 s republish loop that MoqLiteNestsSpeaker.startBroadcasting uses for the legacy non-reconnecting path. Each iteration tears down the prior session's catalog publisher + republisher AFTER the new pair is installed, so watchers see at most a brief overlap rather than a gap when the 600 s moq-auth bearer triggers a session recycle. Without this, the catalog goes silent the moment the wrapper recycles the underlying session; any watcher that attaches AFTER the recycle finds nothing to subscribe to even though the audio path is hot- swapped seamlessly. With it, both audio and catalog survive every JWT refresh. Catalog teardown is also added to ReissuingBroadcastHandle.close() because — unlike the audio publisher which broadcaster.stop() handles internally — the catalog has no broadcaster wrapper. |
||
|
|
babf7623d8 |
feat(nests): publish catalog.json so browsers can discover broadcasts
Two-phone Amethyst sessions stream Opus frames just fine on the wire,
but the kixelated/moq browser watcher (the canonical web reference) sees
nothing — because we never published a catalog.json track. moq-lite
watchers discover a broadcast's tracks by subscribing to the
`catalog.json` track on the broadcast suffix and parsing the JSON
manifest. Without it the watcher has nothing to subscribe to.
Two parts:
1. MoqLiteSession.publish now supports multiple tracks per session,
sharing one broadcast suffix. The previous "one publisher per
session" invariant rejected the second publish() with
"already publishing", forcing a single track per session.
2. MoqLiteNestsSpeaker.startBroadcasting now opens both an audio
publisher AND a catalog publisher, pushes the manifest once, and
launches a 2s republish loop so late-attaching watchers don't wait
indefinitely for the next refresh.
The manifest shape mirrors RoomSpeakerCatalog (kotlinx.serialization
data class on the listener side):
{"version":1,"audio":[{"track":"audio/data","codec":"opus",
"sample_rate":48000,"channel_count":1}]}
Hard-coded for now since OpusEncoder is fixed at 48kHz mono.
Hot-swap path (ReconnectingNestsSpeaker session refresh) is NOT yet
wired to recycle the catalog publisher — acceptable trade-off for a
diagnostic landing; follow-up will move the catalog republisher into
HotSwappablePublisherSource so it survives JWT refreshes.
|
||
|
|
0874706b31 |
Merge pull request #2742 from vitorpamplona/claude/fix-eventstore-expiration-test-AbtSm
test(quartz): widen NIP-40 expiration buffer to outlive insert race |
||
|
|
d6be9e4574 |
test(quartz): drive nip40 expiration test through projection directly
Replace the wall-clock dance with a direct EventStoreProjection construction wired to a controllable nowProvider. Both events are inserted with expirations safely in the future so the SQL reject_expired_events trigger never fires; the test then bumps the fake clock and replays the StoreChange.DeleteExpired the observable would emit after a sweep. Tests the projection's reaction to the sweep signal — which is what this case was always about — without a 6s real-time delay or a flake window. Full SQL sweep behavior remains covered by ExpirationTest.testDeletingExpiredEvents. |
||
|
|
5f4d8ba747 |
test(quartz): widen NIP-40 expiration buffer to outlive insert race
The reject_expired_events trigger checks NEW.expiration <= unixepoch() at insert time, so signing/insert latency on a slow runner could push unixepoch() to time+1 before the row was committed, raising SQLiteException at runBlocking entry. Bumping the short event to time+5 with a matching delay(6000) mirrors ExpirationTest.testDeletingExpiredEvents and gives the trigger room to breathe without changing the behavior under test. |
||
|
|
44f3776e86 |
Merge pull request #2741 from vitorpamplona/claude/investigate-build-size-458NE
Optimize CI builds and reduce APK size with Gradle caching and font subsetting |
||
|
|
bedf21f57e |
chore(nests): auto-enable NestsTrace recorder in debug builds
Hook into Amethyst.onCreate behind the existing isDebug check so the JSONL session-trace recorder lights up the moment a debug APK launches — no per-room toggle to find, no rebuild needed to flip the flag, no extra step on top of the existing logcat capture flow. Release builds are unaffected: setRecording() is gated by `if (isDebug)`, so the volatile flag stays false and every emit site short-circuits at the field-load + branch. Capture from a connected receiver phone with: adb logcat -c adb logcat -s NestsTraceJsonl:D -v raw > nest-trace.jsonl The `-v raw` formatter strips the `D NestsTraceJsonl:` prefix so the file is valid JSONL ready for the (planned) TraceReplayingTransport to feed back through MoqLiteSession + NestViewModel for headless cliff-scenario regression tests. |
||
|
|
a86f19f069 |
feat(nests): NestsTrace recorder for replayable session captures
Adds an opt-in JSONL event recorder behind every receiver-side
moq-lite + NestViewModel decision point so a real two-phone
production session can be captured and (in a follow-up) replayed
through the unmodified pipeline as a unit test. Step 1 of the
"capture-then-replay" plan from the prior conversation.
Off by default — production release builds that never call
`NestsTrace.setRecording(true)` pay one volatile-load + branch per
emit site. The fields-builder lambda doesn't run on the disabled
path, so call sites can do non-trivial string concat freely.
Output goes to logcat under tag `NestsTraceJsonl`. Capture with:
adb logcat -c
adb logcat -s NestsTraceJsonl:D -v raw > nest-trace.jsonl
The `-v raw` formatter strips the `D NestsTraceJsonl:` prefix so
the captured file is valid JSONL ready for replay tooling.
Schema per line: `{"t_ms":N,"kind":"K", ...kind-specific fields}`,
where `t_ms` is milliseconds since `setRecording(true)` was first
called. Field names are lower_snake_case for stability across
clients.
Wired into 13 call sites this commit (matched 1:1 with existing
NestRx/NestTx human log lines so the trace and the log line stay
adjacent and the diff stays small):
`MoqLiteSession`:
- announce_bidi_opened (per session.announce call)
- announce_pump_emit (per Active/Ended received on a bidi)
- announce_bidi_ended_naturally / announce_bidi_threw
- subscribe_send / subscribe_ok / subscribe_drop
- subscribe_bidi_exited
- announce_watch_update / announce_watch_ended_closing_subs
- uni_pump_started
- group_header / group_fin / group_threw
`NestViewModel`:
- vm_observe_announce (per ann emission)
- cliff_tick (every CLIFF_DIAG_LOG_EVERY ticks: active + announced
sets + per-pubkey lastFrameAt elapsed-ms)
- cliff_recycle (when the detector forces a recycleSession())
Anonymisation: pubkeys + track names recorded verbatim — they're
already in the existing `NestRx`/`NestTx` log lines the user is
sharing. Frame payloads are NEVER recorded, only sizes — audio
content can't leak through a trace dump.
Tests: NestsTraceTest (9 cases) — exhaustive jsonStr/jsonArrStr
quoting + setRecording state-machine + emit-lambda-noop-when-
disabled coverage. The `emit` log-output side itself is untestable
in commonTest because `Log.d` writes to a platform actual; the
schema correctness we DO want to pin (a JSON syntax bug at one of
the 13 call sites would silently break replay) is covered by the
quoting helpers.
CliffDetectorTest 12/12 + MoqLiteSessionTest 11/11 still pass —
the trace wiring is purely additive next to existing log statements.
Follow-up: a `TraceReplayingTransport` reading these JSONL files
back through `WebTransportSession` to drive end-to-end regression
tests for the cliff-recovery scenarios captured from production.
|
||
|
|
19588be84b |
build(perf): trim build outputs for hooks, CI, and the shipped font
Headline numbers from a fresh `./gradlew assemble`:
- :commons material_symbols_outlined.ttf 11M -> 409K (subset to the 210
codepoints actually referenced from MaterialSymbols.kt)
- :amethyst per-ABI APKs no longer built in CI (-PdisableAbiSplits=true);
~600M of stripped_native_libs intermediates skipped per run
Changes:
- .git-hooks/pre-push runs only :amethyst:testPlayDebugUnitTest plus jvmTest
on the KMP modules instead of `./gradlew test`, which was compiling all six
amethyst variants (play/fdroid x debug/release/benchmark)
- amethyst/build.gradle adds three opt-in fast-build flags:
-PdisableAbiSplits=true skip per-ABI APK splits
-PdisableUniversalApk=true skip the universal APK output
-Pamethyst.skipMapping=true disable R8 (release+benchmark)
Defaults are unchanged; release pipelines must not set skipMapping.
- .github/workflows/build.yml uses gradle/actions/setup-gradle@v4 with
cache-read-only on PRs / cache-write on main, drops --no-daemon, collapses
the Android job to a single Gradle invocation (lint + focused debug unit
tests + assembleBenchmark with -PdisableAbiSplits=true), and globs both
APK naming patterns for the upload step.
- tools/material-symbols-subset/{subset.sh,README.md} regenerates the font
from upstream + MaterialSymbols.kt; run after adding/removing icons.
Verified: ./gradlew :amethyst:help with all three new -P flags parses cleanly,
and ./gradlew :commons:jvmJar --rerun-tasks succeeds with the subsetted font
(commons-jvm jar shrinks from ~13M to 2.9M).
https://claude.ai/code/session_01YSmkagXXN5AwGcY4upiUh1
|
||
|
|
a36ccb5692 |
fix(nests): close the relay-cliff at the source — framesPerGroup 10→50, cliff-threshold 4s→2.5s
Two-phone production logs at commit |
||
|
|
6e4df4aad4 |
fix(nests): cliff-detector — reset lastFrameAt on recycle + bump cooldown to 30s
Production logs at commit
|
||
|
|
ea08c43b71 |
fix(nests): don't teardown on wrapper Closed — let orchestrator reconnect
The cliff-detector recycle path emits Closed → Reconnecting → Connecting → Connected from the wrapper. The Closed branch was calling teardown(), which cancels cliffDetectorJob, announcesJob, and the wrapper itself, preventing the orchestrator from ever reopening the inner listener. Result pre-fix (visible in receiver log): cliff-detector fires after 4s of silence, teardown runs ~500ms later, wrapper never reopens, room permanently dead. User-initiated close goes through disconnect() / onCleared() which call teardown directly; the wrapper's subsequent Closed emission is redundant for those paths, so no-op'ing it is safe. UI still reflects Closed via state.toUiState(ui.connection). https://claude.ai/code/session_01UHN3fnXzWdj8UbSXnxSwwv |
||
|
|
457e0f5997 |
fix(nests): wrapper.announces() ALSO needs channelFlow (collectLatest emits cross-coroutine)
The channelFlow conversion in |
||
|
|
f17e7adfa7 |
fix(nests): announces flow MUST use channelFlow, not flow{}
The diagnostic logging in |
||
|
|
1fc8dbceee |
debug(nests): trace announce-flow chain to localise still-empty _announcedSpeakers
The replay=64 SharedFlow fix in
|
||
|
|
d6517cf346 |
fix(nests): announce SharedFlow late-attach race dropped Active updates
The cliff-detector diagnostic logs from the user's two-phone repro
(commit
|
||
|
|
46a07cfe43 |
Merge pull request #2739 from vitorpamplona/l10n_crowdin_translations
New Crowdin Translations |
||
|
|
84bd3560e5 |
Merge pull request #2736 from vitorpamplona/claude/disable-vlc-ci-checks-4BD0q
Cache vlc-setup downloads in desktop CI |
||
|
|
f2205dc97e |
test(nests): pure unit tests for cliff-detector predicate + diagnostics + Connected restart
The two-phone repro on commit
|
||
|
|
0887c65ce9 | New Crowdin translations by GitHub Action | ||
|
|
3915d98774 |
Merge pull request #2738 from vitorpamplona/claude/fix-session-swap-timeout-VB9NT
Fix race condition in ReconnectingNestsListenerTest frame delivery |
||
|
|
bd681a0a5f |
test(nests): close FRAME1 delivery race in subscribeSpeaker_survives_session_swap
The test emitted FRAME1 into first.frames and immediately called first.fail(...), trusting that FRAME1 would propagate through the pump to the wrapper's frames before the session swap. emit() is non-suspending — it only enqueues FRAME1 into the pump's slot. On slower hosts (observed on macOS CI) the orchestrator's reconnect path can flip activeListener to the second session before the pump's collect lambda runs, at which point collectLatest cancels pump-iteration-1 mid-resume and FRAME1 is dropped. The consumer's take(2) then only ever sees FRAME2 and the async's withTimeout fires after 5 s. Add a consumerProgress StateFlow that the consumer's collector bumps on each frame, and wait for it to reach 1 before failing the listener. Same shape as the existing consumerSubscribed gate that closed the "emit before consumer subscribes" race. |
||
|
|
cb61082bc4 |
fix(nests): close receiver-audio dropout cluster + auto-start broadcast on host create
Six changes that together close out the bug class the user reported (receiver "blinks green ring for ~5 s then stops") and the related host-side "circular spinner until first unmute" UX issue. Each is defensible independently; together they form a defense-in-depth against the moq-rs production-relay forward-queue cliff documented in `nestsClient/plans/2026-05-01-quic-stream-cliff-investigation.md`. #3 — `NestMoqLiteBroadcaster.onLevel` only fires when `current.send(opus)` returned `true` (frame reached an inbound subscriber). Two-phone logs showed the host's green/glow firing for the first ~7 s while no listener was attached and after a relay-side cliff while no audio was on the wire. The local "I'm broadcasting" ring now tracks the wire, not the runCatching outcome. #5 — `DEFAULT_FRAMES_PER_GROUP: 5 → 10` (200 ms of audio per group, 5 streams/sec). Two-phone production logs showed the relay still cliffing at ~135 streams under sustained load with the old 5-frame default — same bug class the plan thought it had closed, just shifted by half. Doubling the pack roughly doubles the worst-case cliff window. Kept the existing `framesPerGroup = 5` test scenarios intact since they're explicit on the call site. #4 — `SUBSCRIBE_RETRY_BACKOFF_INITIAL_MS: 250 → 100`. Logs of the join path showed 5 retries (250+500+1000+1000+1000 = 3.75 s) of "subscribe stream FIN before reply" before the relay observed the broadcast's announce. New ladder is 100+200+400+800+1000 = 2.5 s. The MAX cap stays at 1000 ms so a permanently-gone publisher still doesn't hammer the relay. #2 — `NestViewModel.openSubscription` now passes `maxLatencyMs = 500L` on every speaker SUBSCRIBE (new constant `ROOM_AUDIO_MAX_LATENCY_MS`). Tells moq-rs to evict groups older than 500 ms from the per-subscriber forward queue rather than queuing them up to the relay's `MAX_GROUP_AGE = 30 s` default — caps the queue growth that triggers the cliff. The wire support existed in the listener; only the call site was passing 0L. #1 — listener-side cliff detector. New `cliffDetectorJob` in `NestViewModel` periodically scans `activeSubscriptions` against `lastFrameAt` (per-pubkey monotonic TimeMark, updated in `onSpeakerActivity`). When a speaker has been announced Active AND we have an active subscription AND no frames have arrived for `ROOM_AUDIO_CLIFF_TIMEOUT_MS = 4 s`, force a `recycleSession()` so the orchestrator opens a fresh QUIC transport — a clean reset of the relay's per-subscriber forward queue. `ROOM_AUDIO_CLIFF_COOLDOWN_MS = 8 s` prevents back-to-back recycles while the new session is mid-handshake. Started from `launchConnect` after `observeAnnounces`, cancelled in `teardown`. Per-pubkey `lastFrameAt` entry dropped on `closeSubscription` so a removed speaker doesn't trigger a stale recycle. #6 — auto-start broadcast pre-muted on host create. `NestViewModel.startBroadcast` gets an `initialMuted: Boolean = false` parameter; when `true` the broadcaster opens the publisher session, calls `setMuted(true)` before the UI flips to `Broadcasting`, and the mic stays open + capture loop keeps running with `if (muted) continue` gating sends. New `LaunchedEffect(speakerPubkeyHex)` in `OnStageIdleControls` calls `startBroadcast(initialMuted = true)` automatically when mic permission is already granted, so a host who just created a room sees their avatar transition to "Live (silent)" and listeners can subscribe immediately without waiting for the host to tap "Talk" first. Tap-to-talk path also passes `initialMuted = true` now — unmute is a separate, explicit step on the mic toggle that appears once we're in `Broadcasting(isMuted = true)`. Permission- denied case still requires the manual Talk-button tap to launch the runtime permission prompt. Diagnostic logs from the previous two debug commits remain in place (throttled to ≤ 1 line/sec/speaker at 50 fps capture cadence) so the next two-phone repro can validate the fixes against logcat directly. |
||
|
|
0c8e1ef58e |
chore(desktop): drop skipVlcSetup escape hatch
The -PskipVlcSetup / AMETHYST_SKIP_VLC opt-out is no longer used: CI relies on the actions/cache step in build.yml to avoid hitting get.videolan.org, and the in-build retry budget covers cache misses. Remove the dead opt-out from desktopApp/build.gradle.kts and the env export from the session-start hook. |
||
|
|
c0262abe59 |
Merge pull request #2735 from vitorpamplona/claude/fix-keyboard-padding-messages-nIBGm
Add IME padding to MessagesTwoPane Scaffold |