Commit Graph

65 Commits

Author SHA1 Message Date
Claude 07dd572423 perf+test(quic): audit follow-ups for the multiplex lock fix
Audit of recent multiplex / PTO commits surfaced three concrete
improvements:

1. **Pre-encode requests outside streamsLock** in both prepareRequests
   impls. QPACK encoding (Http3GetClient) and string formatting
   (HqInteropGetClient) are non-trivial under multiplex load — 1999
   paths means we were holding streamsLock across all 32 chunks ×
   ~2KB of QPACK encoding per chunk = ~64 KB of CPU work serialised
   against the send loop. Now done outside the lock.

2. **Fix O(N²) byte merging in MultiplexingRoundTripTest.** Previous
   shape allocated a new array per STREAM frame; for real-size
   payloads this would dominate test runtime. Switched to a per-
   stream MutableList<ByteArray> joined once at the end.

3. **Pin coalescing end-to-end in MultiplexingRoundTripTest.**
   Pre-existing test verified per-stream content arrived but said
   nothing about the wire shape. Added totalDatagrams ≤ 12 assertion
   so the matrix's "one stream per datagram" failure mode would
   break the test instead of passing silently.

Audit also surfaced two non-actionable items, flagged for future
cleanup but not changed in this commit:

- LevelState.levelLock docstring claims writer/parser acquire it
  around sentPackets mutations; in practice neither does. The
  handlePtoFired call that takes levelLock currently serialises
  only against itself; SendBuffer's internal synchronized is what
  prevents the actual cryptoSend race. Kept the call (matches the
  documented design intent and future-proofs against the writer
  actually taking levelLock) but the docstring is stale.

- openBidiStreamLocked's check(streamsLock.isLocked) catches "no
  lock" and "wrong lock" callers but not "another coroutine holds
  streamsLock and I'm calling without holding it" — kotlinx Mutex
  doesn't expose owner-aware checks without an explicit owner arg
  we don't pass. Acceptable since the bug we just fixed and any
  future regression in the same shape are caught.

https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
2026-05-07 03:25:55 +00:00
Claude 991b1a1da3 fix(quic-interop): prepareRequests must hold streamsLock, not lifecycleLock
aioquic interop multiplexing 2026-05-06 qlog post-mortem:
  - 2898 packets sent in 60s, each carrying ONE STREAM frame
  - server log: streams created/discarded strictly serially, ~30-40ms apart
  - 1421/2000 files completed before the runner's 60s timeout
  - shape ≈ 1 RTT per stream — wire was emitting one stream per datagram

Cause: Http3GetClient.prepareRequests + HqInteropGetClient.prepareRequests
both did `conn.lock.withLock { ... openBidiStreamLocked() }`. Post the
lock-split refactor `conn.lock` is the deprecated alias for lifecycleLock.
The writer's drainOutbound takes streamsLock — not lifecycleLock — so the
send loop interleaved between every two openBidiStreamLocked calls,
draining one stream's data per pass.

Fix:
  1. Both prepareRequests impls now use conn.streamsLock.withLock.
  2. openBidiStreamLocked now `check`s streamsLock.isLocked at entry
     so this can never silently regress again — calling it without the
     lock (or with the wrong lock) throws IllegalStateException with
     a message naming streamsLock as the lock to acquire.
  3. New BatchedOpenLockContractTest pins the contract:
     - calling openBidiStreamLocked WITHOUT any lock throws
     - calling openBidiStreamLocked while holding lifecycleLock throws
       (the exact regression shape)
     - calling openBidiStreamLocked while holding streamsLock works
       (happy path)

The runtime check is the regression-proof part: future callers physically
cannot hold the wrong lock without the test (and prod) blowing up at the
first call site.

https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
2026-05-07 03:18:18 +00:00
Claude 46926a712b test(quic): in-process matrix-shape multiplexing round-trip
The runner's `multiplexing` testcase opens N parallel bidi streams,
each downloading one file, and asserts every file's content lands
on the right stream. Existing tests cover pieces of that:

- MultiplexingThroughputTest: opens 1000 streams in <2s — measures
  lock contention but never moves bytes server-side.
- MultiplexingCoalescingTest: pins that 64 streams coalesce into ≤6
  packets — encoder contract, no end-to-end.
- MultiStreamFinDeliveryTest: server pushes responses to 50 streams,
  client surfaces every FIN — but the CLIENT never sends a STREAM
  frame in that test, so any regression in the writer's request-
  side multiplex path is invisible.

This test runs the full request → response loop:
  1. Open 64 parallel bidi streams
  2. Each enqueues a tiny request + FIN
  3. Drain client outbound → decrypt → assert all 64 STREAM frames
     made it across, with their request bytes intact
  4. Server sends one response per stream + FIN
  5. Per-stream incoming.toList() must yield the expected response

Failures call out the specific stream id, so a regression points
at "stream X dropped its FIN" instead of a generic timeout.

64 streams keeps wall-clock under a second; the bug class the test
guards against (per-stream loss / mis-routing) fires identically at
64 and 1999.

https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
2026-05-07 02:55:18 +00:00
Claude b579c766a4 test(quic): exercise the actual driver PTO path, not a simulation
The existing PtoCryptoRetransmitTest simulated the driver inline:
it set pendingPing=true and called requeueAllInflightCrypto by hand,
then asserted the next drain emitted CRYPTO. That checked the helpers
worked but never noticed when the DRIVER stopped calling
requeueAllInflightCrypto — which is exactly the regression that bit
us in commits c0d7b6031 (qlog merge) and again in cf2303a38
(lock-split refactor).

Extract the PTO-fired logic from QuicConnectionDriver.sendLoop into
a top-level internal helper handlePtoFired(conn). Driver and test now
both call the same function — if anyone unwires the requeue from
that helper or rewrites sendLoop to do volatile-only updates, the
test breaks.

Verified the regression-catch by stripping the requeue from
handlePtoFired and re-running the test: it fails with an
AssertionError on the PTO retransmit packet check, as expected.
Restored the helper, test green again.

https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
2026-05-07 02:50:31 +00:00
Claude d920bf8fd0 Merge branch 'worktree-agent-acb67f8575e4086eb' into claude/research-quic-libraries-hH1Dc 2026-05-07 02:25:28 +00:00
Claude ef4bb99988 refactor(quic): split conn.lock into streamsLock + per-level lock + lifecycleLock
The single connection-wide `QuicConnection.lock` mutex serialised every
critical path: the read loop's `feedDatagram`, the send loop's
`drainOutbound`, and every public mutator (`openBidiStream`,
`streamById`, `flowControlSnapshot`, ...). The multiplexing testcase
opens hundreds of bidi streams in parallel and was capped at ~25
streams/sec by lock contention against the I/O loops.

Phase 1 of the lock split (see
`quic/plans/2026-05-08-lock-split-design.md`) introduces three
domain-specific mutexes:

  - `streamsLock` — streams registry, datagram queues, stream-id
    counters, connection-level flow-control bookkeeping, pending-
    retransmit maps for control frames
  - `LevelState.levelLock` (one per encryption level) — per-level
    pnSpace / sentPackets / ackTracker / CRYPTO buffers
  - `lifecycleLock` — status transitions, close reason/error code

Acquisition order: `lifecycleLock < streamsLock < levelLock`.
Per-stream `synchronized(this)` blocks inside SendBuffer/ReceiveBuffer
remain at the leaf — never acquire any QuicConnection mutex while
holding a per-stream lock.

The legacy `lock: Mutex` field is preserved as a deprecated alias of
`lifecycleLock` for source-compatibility with external test harnesses;
new code MUST use the appropriate domain lock.

Highlights:

  - `feedDatagram` / `drainOutbound` now require the caller to hold
    `streamsLock`; the driver wraps each call. Phase 1 keeps the whole
    feed/drain inside `streamsLock` for safety; phase 2 (deferred) will
    split frame-collection from encrypt + sentPackets-record so app
    coroutines can intersperse during the encrypt window.
  - `pendingPing`, `peerTransportParameters`, `status`,
    `handshakeComplete` are now @Volatile so observers read them
    without a lock.
  - `markClosedExternally` no longer needs any lock (status is
    @Volatile, signals are channel-thread-safe).
  - Driver's PTO bookkeeping uses the volatile fields directly — no
    lock needed.
  - Tests that manually acquired `conn.lock` to call
    `getOrCreatePeerStreamLocked` / `onTokensAcked` / `onTokensLost`
    now acquire `streamsLock` (the domain those routines mutate).
  - New `MultiplexingThroughputTest` locks in the contract: 1000
    parallel `openBidiStream` calls must complete in <2 s.

Test plan:

  - `:quic:jvmTest` — 294 tests pass (293 prior + 1 new throughput).
  - `MultiplexingThroughputTest`: 1000 bidi streams in 52 ms
    (~19,000 streams/sec on the in-memory pipe), well above the
    250+/sec target.
  - `:nestsClient:compileKotlinJvm` — clean, no API breaks.
  - `./gradlew :quic:spotlessApply` — clean.

https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
2026-05-07 02:24:51 +00:00
Claude 7ed3d55b31 test(quic): pin multiplexing coalescing contract
Two unit tests for the multiplexing throughput problem we just fixed
in the InteropClient (commit bc19e90c1):

1. `64 streams enqueued before drain coalesce into a small fixed
   number of packets` — opens 64 bidi streams, enqueues 50 bytes +
   FIN on each (no drain between), drains everything, asserts
   ≤6 packets and ≥10 streams/packet. Pre-fix shape (each enqueue
   followed by an immediate single-stream drain) would emit 64 packets.

2. `1000 streams enqueued in batches of 64 produce a tractable packet
   count` — stress version mirroring the runner's 1999-file
   multiplexing test. Asserts ≤150 packets total for 1000 streams.

Both tests run in <300ms on the in-memory pipe — fast iteration for
debugging the multiplexing-throughput regression cycle without
spinning up Docker.

These pin the contract that ZERO drains between enqueues = packets
batch many streams' frames. The runner-side fix (split GetClient.get
into prepareRequest + awaitResponse, batch prepareRequests serially,
wake once) ensures we hit this shape in real interop.

https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
2026-05-07 01:59:50 +00:00
Claude 17b80270d9 fix(quic): preserve Initial PN namespace across Retry (RFC 9001 §5.7)
aioquic's retry test result, surfaced via qlog:
  Check of downloaded files succeeded.
  Client reset the packet number. Check failed for PN 0

Our applyRetry called LevelState.resetForVersionNegotiation, which
creates a fresh PacketNumberSpaceState() — resetting PN to 0. The
qlog confirmed: PN=0 sent at t=388 (pre-Retry ClientHello), then PN=0
again at t=1468 (post-Retry retried ClientHello). Same PN reused
across the boundary.

RFC 9001 §5.7 + RFC 9000 §17.2.5: the Initial PN namespace CONTINUES
across the Retry boundary. The new Initial keys are derived from the
new DCID, but PN doesn't reset. Reusing a PN under different keys
makes the runner's pcap-decryption check fail (it's also a security
concern in the general case, hence the strict spec rule).

Fix: new LevelState.resetForRetry that's identical to
resetForVersionNegotiation EXCEPT it preserves pnSpace. applyRetry
calls resetForRetry. Two regression tests updated to assert the
post-Retry Initial uses PN=1 (continues from PN=0 of the pre-Retry
attempt) rather than PN=0 (the buggy reset behavior).

For Version Negotiation the original semantics still apply (RFC 9000
§6.2: client treats VN as if the original Initial was never sent;
PN reset to 0 is correct).

This should bring the retry testcase from ✕(S) to ✓(S) against
servers that exercise the Retry path. The handshake / transfer
already succeeded over the Retry per the qlog (the server's check
"Check of downloaded files succeeded." passed); only the PN-reuse
flag was failing the test.

https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
2026-05-07 00:56:59 +00:00
Claude 0107bbbac9 Merge branch 'worktree-agent-a5e247c7b4025a83c' into claude/research-quic-libraries-hH1Dc 2026-05-07 00:24:08 +00:00
Claude cfc305feb3 feat(quic): add qlog observer infrastructure for interop diagnostics
Hooks every QUIC protocol decision (packets sent / received / dropped,
TLS key updates, transport params, ALPN, loss detection, PTO, close)
into a [QlogObserver] interface. Production callers default to
[QlogObserver.NoOp] (zero allocation, single virtual call); the
:quic interop runner wires a [QlogWriter] writing JSON-NDJSON
(qlog 0.3 / JSON-SEQ format) to `<QLOGDIR>/client.sqlog`, consumable
by qvis (https://qvis.quictools.info/) and Wireshark. Goal: every
interop-runner test failure produces a qlog file the operator can
drop into qvis to see exactly what we did differently from the spec.

Hooked call sites:
  - QuicConnection.start                   → connection_started, parameters_set(local), version_information
  - QuicConnection.close                   → connection_closed(local)
  - QuicConnection.markClosedExternally    → connection_closed(remote)
  - QuicConnection.applyPeerTransportParameters → parameters_set(remote)
  - QuicConnection.tlsListener (handshake/app keys) → security:key_updated
  - QuicConnection.tlsListener (handshake done)     → alpn_information
  - QuicConnectionWriter.buildLongHeaderFromFrames  → packet_sent (initial / handshake)
  - QuicConnectionWriter.buildApplicationPacket     → packet_sent (1-rtt)
  - QuicConnectionWriter.buildBestLevelPacket       → packet_sent (close-path)
  - QuicConnectionParser.feedLongHeaderPacket       → packet_received / packet_dropped
  - QuicConnectionParser.feedShortHeaderPacket      → packet_received / packet_dropped
  - QuicConnectionParser AckFrame loss-detect       → recovery:packet_lost
  - QuicConnectionDriver.sendLoop PTO branch        → recovery:loss_timer_updated (pto)
2026-05-07 00:21:42 +00:00
Claude d0bc998cd2 Merge branch 'worktree-agent-a4e96f738ceb4bbd5' into claude/research-quic-libraries-hH1Dc 2026-05-07 00:21:38 +00:00
Claude aff2ee182b fix(quic): add explicit peer-uni-stream drainer to avoid H3 multiplex tear-down
Variant (B) from the three-way fix menu in the multiplexing-interop
investigation: keep `:quic` strict about per-stream backpressure (the
audit-4 #3 "INTERNAL_ERROR: stream … consumer overflowed" tear-down
stays the contract for app-data overflow on bidi streams) but expose
an explicit, opt-in helper for peer-initiated UNI streams that the
application has decided it does not need to interpret.

Root cause confirmed in QuicConnectionParser.kt:290: when the server
opens its three RFC 9114 §6.2.1 peer-uni streams (CONTROL +
QPACK_ENCODER + QPACK_DECODER) and the H3 client does not consume
them, the parser routes their bytes into each stream's bounded
incomingChannel (capacity 64). Once the QPACK encoder issues
dynamic-table inserts beyond 64 chunks the next chunk overflows
trySend, sets QuicStream.overflowed, and the parser maps that to
markClosedExternally — the entire connection dies.

Notes on scope:
  - The `Http3GetClient` and `:quic-interop` runner mentioned in the
    investigation prompt do NOT exist on the `main` worktree this
    branch starts from. The fix here is therefore `:quic`-only: the
    public `awaitIncomingPeerStream` API was already sufficient for
    an integrator to write the accept loop themselves; this commit
    wraps the common case in `drainPeerInitiatedUniStreamsIntoBlackHole`
    and updates the doc on `awaitIncomingPeerStream` so the next
    integrator landing the H3 GET client doesn't hit the same trap.
  - Variant (C) — silent default drain in `:quic` itself — was
    deliberately rejected: defaults that swallow application bytes
    are the misconfiguration we want type-system-or-API-explicit.
    The new helper requires the caller to pass a CoroutineScope, so
    opt-in is unmistakable in any callsite.

Regression test coverage in PeerUniStreamDrainTest:
  - pre_fix_no_consumer_overflows_and_tears_down_connection — pushes
    65 chunks (capacity + 1) on a SERVER_UNI stream with no consumer;
    asserts the connection transitions to CLOSED. Pins the existing
    backpressure contract.
  - drainPeerInitiatedUniStreamsIntoBlackHole_keeps_connection_alive
    — same setup but with the new helper running on a side scope;
    pushes 256 chunks (4× capacity) and asserts the connection stays
    CONNECTED. With the helper sabotaged, this test fails at
    line 119 with status=CLOSED, confirming it actually exercises
    the fix.

Full quic test suite: 295 tests, 0 failures, 0 errors.

https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
2026-05-07 00:19:40 +00:00
Claude 04e30465d9 Merge branch 'worktree-agent-a9d8336181fce16eb' into claude/research-quic-libraries-hH1Dc 2026-05-07 00:17:49 +00:00
Claude 350387f7e0 feat(quic): handle Version Negotiation packets per RFC 9000 §6
Adds the client-side VN flow needed for the interop runner's
`versionnegotiation` testcase:

- `QuicConnection` accepts an `initialVersion` constructor parameter
  (default `QuicVersion.V1`) and exposes a mutable `currentVersion`
  the writer stamps into outbound long-headers. `start()` now caches
  the ClientHello bytes for VN-driven re-emission.
- `applyVersionNegotiation(supportedVersions)` validates per §6.2
  (anti-downgrade: reject if list contains the offered version),
  picks v1 from the offered set, regenerates DCID, re-derives
  Initial keys against the new DCID, resets the Initial level via
  `LevelState.resetForVersionNegotiation`, re-enqueues the cached
  ClientHello, and latches `vnConsumed` so a second VN is dropped.
  Failure to find a mutually supported version closes the connection
  with `QuicVersionNegotiationException`.
- `QuicConnectionParser.feedDatagram` detects `version == 0` long
  headers BEFORE peekHeader (whose layout assumes v1) and dispatches
  to a new `feedVersionNegotiationPacket` that parses the §17.2.1
  shape and validates the echoed DCID.
- `QuicConnectionWriter` reads `conn.currentVersion` instead of the
  hardcoded `QuicVersion.V1`.
- `QuicVersion.FORCE_VERSION_NEGOTIATION = 0x1a2a3a4a` for the
  interop runner.
- `InteropRunner` honors `TESTCASE=versionnegotiation` (or
  `-DinteropTestcase=`) and offers the force-VN version.

Regression coverage in `VersionNegotiationTest`:

- happy path: VN switches `currentVersion` to v1, regenerates DCID,
  resets PN, and the next drain emits a v1 Initial on the wire.
- downgrade defense: VN listing the offered version is dropped.
- unsupported list: VN whose versions we can't speak fails the
  handshake and closes the connection.
- second VN: post-consumption VN is ignored.
- DCID mismatch: spoofed VN with wrong echoed DCID is dropped.
- backward compatibility: default `initialVersion` keeps v1 behavior
  for existing callers.

https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
2026-05-07 00:17:13 +00:00
Claude 2da5d42d70 Merge branch 'worktree-agent-a5a40cf58838c96dd' into claude/research-quic-libraries-hH1Dc 2026-05-06 23:48:39 +00:00
Claude 39f9ae2aab fix(quic): deliver FIN to per-stream Channel under concurrent multi-stream load
When QuicConnection tears down (CONNECTION_CLOSE, read-loop death,
INTERNAL_ERROR from a saturated stream channel, etc.) the per-stream
incomingChannel objects were left open, so any application coroutine
suspended on `stream.incoming.collect { … }` hung forever waiting for
a FIN that would never come. The connection-wide signal channels
(closedSignal, peerStreamSignal, incomingDatagramSignal) all closed
cleanly, but the per-stream Flows did not — surfacing in the
quic-interop-runner `multiplexing` case as 677 collectors stuck after
the connection died mid-response, so zero of the 1999 expected files
landed.

Fix: closeAllSignals() now also calls closeIncoming() on every stream
in streamsList. Channel.close() is idempotent, and consumeAsFlow drains
already-buffered chunks before honouring the close, so any bytes the
parser had already pushed are still surfaced to the collector before
the Flow terminates.

Adds MultiStreamFinDeliveryTest covering: (a) FIN delivery to N parallel
client-bidi streams, (b) connection-teardown unblocks every per-stream
Flow, (c) buffered bytes survive a teardown without an explicit FIN.
2026-05-06 23:47:28 +00:00
Claude 671f9c7050 Merge branch 'worktree-agent-ac6580a9453de5616' into claude/research-quic-libraries-hH1Dc 2026-05-06 23:07:21 +00:00
Claude fb35031b4e Merge branch 'worktree-agent-a4016e24b23c3e8ff' into claude/research-quic-libraries-hH1Dc 2026-05-06 23:07:13 +00:00
Claude d03e179816 feat(quic): wire Retry packet handling (RFC 9000 §17.2.5 + RFC 9001 §5.8)
The Retry parser + integrity-tag verifier already existed in
RetryPacket.kt, but feedDatagram dropped Retry packets on the floor.
Hook them up:

- QuicConnectionParser.feedLongHeaderPacket detects RETRY type before
  the standard parse-and-decrypt path, parses via RetryPacket, and
  dispatches to QuicConnection.applyRetry.

- QuicConnection.start() now caches the ClientHello bytes (TLS only
  emits ClientHello once; we need to re-queue the same bytes on the
  fresh Initial keys after Retry). New applyRetry method:
  verifies the integrity tag, swaps DCID to Retry's SCID, re-derives
  Initial keys, resets the Initial PN space + sentPackets +
  cryptoSend, re-enqueues the cached ClientHello, stores the Retry
  token, and latches retryConsumed so a second Retry is dropped.

- LevelState.restoreFromRetry / PacketNumberSpaceState.resetForRetry
  give applyRetry an in-place reset (the level reference is a `val`,
  so we mirror discardKeys' field-reset pattern).

- QuicConnectionWriter.buildLongHeaderFromFrames threads
  conn.retryToken through the Initial header's Token field on every
  Initial we emit after Retry.

Per RFC 9001 §5.8, a Retry with a bad integrity tag is silently
dropped; per RFC 9000 §17.2.5.2, only one Retry is honored per
connection. Both invariants are tested.

New test: RetryHandlingTest covers the happy path (DCID swap, PN
reset, token threading, ClientHello replay, ≥1200-byte padding),
the bad-tag path, and the second-retry path.
2026-05-06 23:02:15 +00:00
Claude c9e036f728 fix(quic): PTO retransmits unacked CRYPTO at Initial/Handshake (RFC 9002 §6.2.4)
Pre-handshake PTO previously only set `pendingPing`, which collapsed to
either nothing (the bug 86b6c609a fixed in a sibling branch) or a bare
PING. Against aioquic in the quic-interop-runner ns-3 sim, the first
ClientHello can be dropped (server not ready at t≈0.5s); a follow-up
PING with the same DCID is silently ignored because the server has no
state for that DCID. We need to retransmit the actual ClientHello bytes
so the server sees a full connection attempt.

Implementation:
  - SendBuffer.requeueAllInflight(): walks the inflight list and moves
    every sent-but-not-ACK'd range to the retransmit queue, preserving
    offset and FIN. Mirrors markLost's per-range path but applies to all
    inflight ranges in one shot. Idempotent + best-effort safe.
  - QuicConnection.requeueAllInflightCrypto(level): thin wrapper that
    drives the per-level cryptoSend buffer's new method.
  - QuicConnectionDriver.sendLoop: PTO branch now calls the new helper
    at the highest active pre-application level (Handshake > Initial)
    when 1-RTT keys aren't installed. The next drain naturally emits a
    CRYPTO frame at the original offset (takeChunk drains the
    retransmit queue first per existing semantics).
  - QuicConnectionWriter.collectHandshakeLevelFrames: honors pendingPing
    pre-handshake at the highest active level — but skips the PING when
    a CRYPTO frame is already in the same level's frame list (the
    CRYPTO retransmit covers the ack-eliciting requirement). Bare PING
    still goes out when there's nothing to retransmit.

Old RecoveryToken.Crypto entries in sentPackets for the original PNs
remain harmless: when loss detection eventually declares them lost,
markLost re-runs against ranges that have already moved on, which is
itself idempotent (clamps to flushedFloor / no-op on already-queued).

Test: PtoCryptoRetransmitTest reproduces the wire scenario — first drain
emits ClientHello in a ≥1200-byte Initial datagram; simulated PTO
calls requeueAllInflightCrypto + sets pendingPing; second drain must
contain a CRYPTO frame at the same offset with the same payload bytes,
not a bare PING. Datagram size still ≥1200 (RFC 9000 §14.1).

https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
2026-05-06 23:01:08 +00:00
Claude 9c86eee5e2 fix(quic): pad PING-only Initial datagrams to strict 1200 bytes (RFC 9000 §14.1)
The padding-rebuild branch in QuicConnectionWriter.drainOutbound computed
`padBytes = 1200 - natural`, but the QUIC long-header Length field is a
varint (RFC 9000 §16). When the natural-size payload was small enough for
Length to fit in 1 byte (body ≤ 63 bytes), the rebuild's larger body
crossed the 64-byte threshold and Length grew to 2 bytes — adding 1 wire
byte that wasn't in `natural`. PING-only PTO probe Initials therefore went
out at exactly 1199 bytes, one short of the §14.1 floor.

Fix: rebuild iteratively. After the first rebuild, measure the actual
datagram size; if still < 1200, bump padBytes by the residual and rebuild
once more. PADDING bytes inside the AEAD envelope add 1:1 to the wire
size and the Length varint grows monotonically, so the loop terminates
in ≤ 2 iterations for any reachable payload.

Same fix is applied to buildClosingDatagram so close-only Initial probes
on the boundary aren't tripped by future varint-growth changes.

Tightens the existing PTO-probe regression test to assert ≥ 1200 (was
relaxed to ≥ 1199 in 86b6c609a) and adds a new boundary test that builds
a single-byte-payload Initial and checks 1200 ≤ size ≤ 1203 — strict
floor with a tight ceiling so over-correction would also fail.

https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
2026-05-06 23:00:30 +00:00
Claude 72295915de fix(quic): tier-local round-robin so priority survives every drain
The previous priority-then-round-robin shape applied the rotating
start-index globally over the sorted list, so the cross-tier order
flipped on alternate drains: drain N had high-priority first, drain
N+1 advanced streamRoundRobinStart and had low-priority first. The
priority hint was silently defeated under any sustained traffic.

Replace it with strict priority across tiers + round-robin only
within each same-priority tier. Higher tiers always drain ahead of
lower ones; same-priority peers continue to take turns via the
existing rotating start. Default-priority callers see no behaviour
change (single tier, identical rotation semantics).

Tighten the test: drain twice and assert the higher-priority stream
emits first on BOTH drains — the regression case that the single-
drain version of the test missed. Add a regression guard for the
same-priority round-robin so a future refactor can't silently
serialise on the first stream.

https://claude.ai/code/session_01KWdr4RjVvyYZfEuPVaQfUa
2026-05-06 20:34:39 +00:00
Claude f1034b1f53 feat(quic): priority-aware stream scheduling for moq-lite groups
Bias the QUIC connection writer's drain loop toward higher-priority
streams so moq-lite group streams with newer (higher) sequence numbers
drain ahead of older ones under congestion. Implements the T11.3
follow-up flagged in nestsClient/plans/2026-05-06-stream-priority-
followup.md (now removed).

QuicStream gets a `@Volatile var priority: Int = 0`. The writer's
streamsView iteration is replaced by a stable sortedByDescending pass
so same-priority streams keep their existing rotating-start round
robin while higher-priority tiers always drain first.

WebTransportWriteStream gains a `setPriority(Int)` hook; the QUIC-
backed adapter forwards to the underlying QuicStream, while the
in-memory test fakes treat it as a no-op. MoqLiteSession.openGroupStream
calls `uni.setPriority(sequence)` (saturating to Int.MAX_VALUE) to
mirror moq-rs's `Publisher::serve_group`.

Tests: a new InMemoryQuicPipe.decryptClientApplicationFrames helper
walks past coalesced long-header packets to surface 1-RTT frames,
which lets QuicConnectionWriterTest assert that the higher-priority
stream's StreamFrame lands first inside a single drained packet.

https://claude.ai/code/session_01KWdr4RjVvyYZfEuPVaQfUa
2026-05-06 20:25:01 +00:00
Claude edd6eb5c10 fix(quic): pad short plaintext payloads for HP sample (RFC 9001 §5.4.2)
ShortHeaderPacket.build / LongHeaderPacket.build crashed with
`IllegalArgumentException: packet too short for HP sample` whenever the
plaintext payload was small enough that pnLen + payload < 4 — most
visibly on the 1-RTT path when buildApplicationPacket emitted a single
1-byte PING (PTO probe with no ACKs queued and no streams to drain) and
the packet number still fit in 1 byte. The crash tore down the writer
loop, which surfaced upstream as moq-lite "subscribe stream FIN before
reply" because in-flight bidi streams got FIN'd by the peer.

RFC 9001 §5.4.2 mandates the sender pad the plaintext so the encrypted
output (plaintext + 16-byte AEAD tag) has at least 20 bytes after the
packet-number offset for the 16-byte HP sample. The fix pads the
plaintext with trailing 0x00 bytes — those decode as PADDING frames per
RFC 9000 §19.1 and decodeFrames already absorbs them. For long-header
packets the padded size feeds back into the Length varint.
2026-05-05 13:37:41 +00:00
Claude fba0a5c952 feat(quic): bestEffort streams + park CC plan indefinitely
After drafting the congestion-control plan we concluded the audio-rooms
workload doesn't actually need CC — speakers push ~8 KB/sec, which
never fills any modern link's capacity. The one real concern that
surfaced — STREAM retransmit wasting bandwidth on stale Opus frames
on lossy uplinks — is much cheaper to fix directly than to bound via
a 14-test CC subsystem.

SendBuffer gains a `bestEffort: Boolean = false` constructor flag.
When true, markLost drops the lost ranges instead of moving them to
the retransmit queue and lets the underlying byte storage compact as
if the bytes had been ACK'd. The FIN flag (if covered) also stays
sent — best-effort skips FIN re-emission too. The peer may end up
with a truncated stream; moq-lite's per-stream timeouts handle that.

Plumbed through QuicStream → QuicConnection.openUniStream(bestEffort)
→ QuicWebTransportSessionState.openUniStream(bestEffort) →
WebTransportSession.openUniStream(bestEffort). Default is false
everywhere, so reliable streams (HTTP/3 control, moq-lite SUBSCRIBE
bidi, etc.) keep RFC 9000 §3.5 semantics.

MoqLiteSession.openGroupStream now passes `bestEffort = true` —
group streams carry a single Opus packet, are real-time, and don't
benefit from retransmit.

Internal cleanup: `removeOverlap`'s `ackedNotLost: Boolean` parameter
became `OverlapAction { ACK, RETRANSMIT, DROP }` so the third best-
effort disposition has a name. Same code paths, same tests, just
clearer at the call site.

CC plan (quic/plans/2026-05-05-congestion-control.md) is updated to
"parked indefinitely" with a note that this commit is the lighter-
weight alternative that addresses the only practical concern. The
plan is preserved as a reference if a future workload justifies CC.

New tests: SendBufferBestEffortTest (6 cases — reliable baseline,
best-effort drops, FIN drop in best-effort mode, partial overlap,
idempotent stale loss, ACK path still works).

https://claude.ai/code/session_01PYYez8a6sjiakyjAxsfCEQ
2026-05-05 02:06:08 +00:00
Claude 2053f50f35 fix(quic): discard Initial/Handshake keys per RFC 9001 §4.9
Pre-fix `:quic` held Initial AND Handshake encryption-level state
indefinitely once derived. AEAD cipher state, per-level CRYPTO
buffers, and the per-level sent-packet map all stayed alive for the
lifetime of the connection — a real memory leak for long sessions
(audio rooms run for hours).

LevelState.discardKeys() (idempotent):
- Nulls sendProtection / receiveProtection (frees AEAD state).
- Replaces cryptoSend / cryptoReceive with empty instances.
- Replaces ackTracker with an empty instance.
- Clears sentPackets and resets largestAckedPn /
  largestAckedSentTimeMs.
- Latches keysDiscarded = true.

Hook locations:
- Initial discard (RFC 9001 §4.9.1, client side): in
  QuicConnectionWriter.drainOutbound, after a Handshake-level packet
  is built into the outbound datagram. The next drainOutbound MUST
  NOT touch the Initial level; any retransmitted Initial from the
  peer is silently dropped (receiveProtection == null), which is
  correct per the same RFC since the server has also moved up
  encryption levels by then.
- Handshake discard (RFC 9001 §4.9.2 + §4.1.2, client side): in
  QuicConnectionParser, on receipt of a HANDSHAKE_DONE frame.

Once a level's protection is null, parser-side decrypt at that level
returns null silently (existing receiveProtection == null check) and
writer-side build skips it (existing sendProtection == null check),
so no further code paths needed updating.

New test: KeyDiscardTest (4 cases — Initial keys discarded after
first Handshake packet, Handshake keys still live until
HANDSHAKE_DONE, Handshake keys discarded on HANDSHAKE_DONE,
discardKeys is idempotent).

Listed in the audit-summary deferred-work as item 3
(`No Initial / Handshake key discard`).

https://claude.ai/code/session_01PYYez8a6sjiakyjAxsfCEQ
2026-05-05 01:41:28 +00:00
Claude e3a3ffd1d9 fix(quic): correct AckTracker.purgeBelow via ACK-of-ACK dispatch
Pre-fix, QuicConnectionParser purged the inbound AckTracker on every
inbound AckFrame using `frame.largestAcknowledged - frame.firstAckRange`
— but that value lives in OUR outbound PN space, while the tracker
holds inbound PNs we received from the peer. The two PN spaces are
unrelated; the bug mostly hid because they grow at similar rates,
but caused range-list bloat over long sessions where traffic is
asymmetric (e.g. listener receives ~50 audio frames/sec while sending
back ~1 ACK/sec).

The correct semantics: purge only when the peer has confirmed receipt
of OUR outbound ACK frame. Now driven by the ACK-of-ACK dispatch.

- RecoveryToken.Ack changed from data object to data class carrying
  (level, largestAcked) — the encryption level and the largest inbound
  PN our outbound ACK frame covered.
- QuicConnectionWriter populates these fields from the AckFrame at
  emit time.
- QuicConnection.onTokensAcked dispatches RecoveryToken.Ack to
  levelState(level).ackTracker.purgeBelow(largestAcked + 1).
- The wrong purge in QuicConnectionParser is removed (replaced with a
  comment pointing at the new dispatch path).

Listed in the audit-summary deferred-work as item 6
(`AckTracker.purgeBelow threshold semantics`).

New test: AckTrackerPurgeOnAckOfAckTest (4 cases — purge on
ack-of-ack, level routing, partial purge keeps higher PNs, out-of-order
ACKs are safe).

https://claude.ai/code/session_01PYYez8a6sjiakyjAxsfCEQ
2026-05-05 01:30:02 +00:00
Claude 086a9c75dc fix(quic): RESET_STREAM/STOP_SENDING first-call-wins + threading contract
Audit follow-up from the prior two commits.

1. resetStream() / stopSending() now no-op on the second call. RFC 9000
   §3.5 pins finalSize at first emission; replaying retransmits with a
   larger value (because the app enqueued more bytes between two
   resetStream calls) would trigger FINAL_SIZE_ERROR on the peer. The
   "idempotent: a second call overwrites with the newer error code"
   claim was simply wrong. Two new tests lock the contract:
   resetStream_secondCallIsNoOp_finalSizeFrozen and
   stopSending_secondCallIsNoOp.

2. resetEmitPending / resetAcked / stopSendingEmitPending /
   stopSendingAcked are now @Volatile. The public emit APIs are
   callable from any coroutine while the writer / loss / ACK
   dispatchers read the same fields under QuicConnection.lock; volatile
   gives the cross-thread happens-before, and the first-call-wins gate
   above eliminates the only multi-writer race (two app threads racing
   the writer's clear-after-emit).

3. SendBuffer's class-level KDoc still claimed range arithmetic was
   O(N) "swap to TreeMap if profiling flags it" — stale after the
   binary-search refactor in 303caa8. Updated to reflect the actual
   O(log N + k) cost.

4. The bulk-removal comment in removeOverlap overstated the win
   ("O(k) per call, single shift of trailing entries"). ArrayDeque
   removeAt(i) shifts on every call, so the actual cost is
   O(k * (size - end + k)). Toned down — it's still cheap because
   k is 1-2 in steady state.

https://claude.ai/code/session_01PYYez8a6sjiakyjAxsfCEQ
2026-05-05 01:04:32 +00:00
Claude 996ab39940 feat(quic): emit RESET_STREAM / STOP_SENDING + per-stream retransmit dispatch
Application code can now call QuicStream.resetStream(errorCode) and
stopSending(errorCode); the next writer drain emits the matching frame
with a RecoveryToken. Loss dispatch re-flags the per-stream emit-pending
bit; ACK dispatch latches resetAcked / stopSendingAcked so stale loss
notifications can't re-emit. NEW_CONNECTION_ID retransmit drains
QuicConnection.pendingNewConnectionId on next writer pass (no public
emit API since :quic doesn't rotate connection IDs, but the wiring is
in place for a future emit path).

Five tests in ResetStopSendingEmitTest mirror neqo's send-stream reset
coverage: emit-and-token, retransmit-on-loss, ack-then-stale-loss-drop,
stop-sending emission, and NEW_CONNECTION_ID drain.

https://claude.ai/code/session_01PYYez8a6sjiakyjAxsfCEQ
2026-05-05 00:23:07 +00:00
Claude 0c847b4f69 feat(quic): wire CRYPTO retransmit per encryption level
Closes commits D + E of the deferred-follow-ups pass. With
SendBuffer's retain-until-ACK semantics in commit B and the
Crypto / ResetStream / StopSending / NewConnectionId tokens added
in commit A, the writer now records a Crypto token per CRYPTO
frame at each encryption level (Initial, Handshake) so RFC 9002
retransmit recovers lost handshake bytes.

# Writer side

QuicConnectionWriter.collectHandshakeLevelFrames returns a new
HandshakeLevelContents(frames, tokens) pair instead of a bare
frame list. For each emitted CryptoFrame it appends a matching
RecoveryToken.Crypto(level, offset, length).

QuicConnectionWriter.buildLongHeaderFromFrames now also takes the
parallel tokens list, and after encryption records a SentPacket
in the matching LevelState.sentPackets map. Initial-level rebuilds
with padding (RFC 9000 §14.1) call pnSpace.rewindOutboundForRebuild
to reuse the same PN — the second build's map insert overwrites
the prior entry, so retention reflects the final padded packet.

drainOutbound's two callsites updated to pass tokens through.

# ACK / loss dispatch

Already wired in commit A. QuicConnection.onTokensAcked routes
Crypto tokens to LevelState.cryptoSend.markAcked at the matching
level, releasing buffer memory as the contiguous low end is ACK'd.
QuicConnection.onTokensLost routes them to markLost, re-queueing
the bytes for retransmit at the same level.

# RESET_STREAM / STOP_SENDING / NEW_CONNECTION_ID

Same dispatcher-only completion. The pendingResetStream /
pendingStopSending / pendingNewConnectionId maps on QuicConnection
are populated by the loss dispatcher when those token types are
seen. :quic doesn't currently emit any of those frames (no
application code triggers stream reset, connection-ID rotation
isn't wired), so the writer never drains the maps yet —
scaffolding for future emit support. The exhaustive when in
onTokensLost / onTokensAcked is now complete: any future addition
of a new RecoveryToken variant trips the compile-time exhaustive
check, mirroring the test in RecoveryTokenTest.

# Tests added (3, all pass)

CryptoRetransmitTest:
  - handshakePacket_carriesCryptoToken_inSentPacket: ClientHello
    emission produces an Initial-level SentPacket with a Crypto
    token at offset 0 with the expected level.
  - cryptoData_lostAndRetransmittedAtSameLevel: simulate loss via
    direct dispatch, observe re-queue in cryptoSend, verify next
    drain produces a fresh Initial packet replaying the same
    offset (RFC 9000 §13.3 idempotent).
  - cryptoAck_releasesBufferAtSameLevel: ACK via onTokensAcked,
    cryptoSend's readableBytes drops to 0.

# Net result

Lost handshake bytes (ClientHello, EncryptedExtensions, Certificate,
Finished, NewSessionTicket) are now recovered automatically. The
prior 1-second-fixed-PTO placeholder in QuicConnectionDriver
(commit step 7 of the prior plan) becomes meaningfully more useful
— PTO now wakes the writer to retransmit ACTUAL data, not just
emit empty PINGs.

Full :quic test suite + nestsClient moq-lite + amethyst Android
compile all pass.

https://claude.ai/code/session_01PYYez8a6sjiakyjAxsfCEQ
2026-05-05 00:10:57 +00:00
Claude f623e886c3 feat(quic): wire STREAM data retransmit — token emission + ACK/loss dispatch
Closes commit C of the deferred-follow-ups pass. With the
SendBuffer rewrite in commit B (retain-until-ACK with markAcked /
markLost), the connection now wires STREAM frames into the same
RFC 9002 retransmit path that already handles flow-control
extensions.

# Writer side

QuicConnectionWriter.buildApplicationPacket records a
RecoveryToken.Stream(streamId, offset, length, fin) for every
STREAM frame it emits. The token captures the on-wire byte range
plus the FIN bit so retransmit can reproduce the same StreamFrame
on next drain.

# ACK side

New QuicConnection.onTokensAcked() mirrors onTokensLost. The
parser's AckFrame handler iterates the drained packets and routes
each to onTokensAcked, which:
  - For Stream tokens: calls SendBuffer.markAcked(offset, length).
    The buffer removes the range from in-flight; if the contiguous
    low end is now fully ACK'd, flushedFloor advances and storage
    shifts forward.
  - For Crypto tokens: same shape, applied to the per-level
    cryptoSend buffer (commit E will exercise this path for
    handshake reliability — Crypto retransmit is wired now but the
    writer's CRYPTO emission path doesn't yet record Crypto tokens;
    that's commit E).
  - For control-frame and Ack tokens: ACK-no-op. The frame already
    did its job by reaching the peer; no per-buffer state to
    release.

# Loss side

onTokensLost (commit A) already routes Stream tokens to
SendBuffer.markLost. With commit B's real implementation (was a
no-op stub), this now actually re-queues the byte range for
retransmit. The next writer drain pulls from the retransmit queue
before any fresh sends, with the original offset preserved (RFC
9000 §13.3 idempotent retransmit).

# Tests added (3, all pass)

  - streamFrame_carriesStreamToken_inSentPacket: writer emits a
    Stream token whose fields match the StreamFrame on the wire
  - streamData_lostAndRetransmittedOnNextDrain: simulate loss via
    direct dispatch, observe re-emit at the same offset in a fresh
    SentPacket (different PN)
  - streamData_ackedReleasesBuffer: ACK via onTokensAcked,
    enqueue more bytes, observe the next send picks up at the
    post-ACK offset (proves bytes were released and floor advanced)

Full :quic test suite, nestsClient moq-lite tests, amethyst Android
compile all pass.

Net result: lost STREAM data (e.g. nestsClient bidi control-stream
bytes — moq-lite Subscribe/Announce control messages travel on
QUIC bidi streams) is now recovered automatically. Audio rooms
benefit indirectly: the relay's announce/subscribe path is more
resilient to packet loss.

https://claude.ai/code/session_01PYYez8a6sjiakyjAxsfCEQ
2026-05-04 23:24:26 +00:00
Claude 03cfb3188f feat(quic): rewrite SendBuffer for retain-until-ACK with markAcked/markLost
Foundation for STREAM and CRYPTO retransmit (commits C, D of the
deferred-follow-ups pass). Replaces the prior best-effort mode
where takeChunk released bytes immediately.

# Three logical regions

The buffer covers `[flushedFloor, nextOffset)`. Each byte is in one
of three states:

  - In-flight: sent but not yet ACK'd. Sorted-by-offset list.
  - Needs retransmit: declared lost; re-sent before any fresh bytes.
    FIFO queue.
  - Unsent: `[nextSendOffset, nextOffset)`.

# New API

  - markAcked(offset, length): removes the range from in-flight,
    advances flushedFloor through any contiguous low-end ACKs and
    shifts the underlying byte storage forward. Length=0 means
    FIN-only ACK; latches finAcked = true.
  - markLost(offset, length, fin): removes from in-flight, appends
    to retransmit queue. If fin, clears finSent so the FIN gets
    re-emitted. Idempotent: stale loss notifications below
    flushedFloor are absorbed.

takeChunk priority order:
  1. retransmit queue (preserves original offset; same byte data)
  2. fresh unsent bytes from [nextSendOffset, nextOffset)
  3. FIN-only zero-byte chunk if finPending and everything drained

# Range arithmetic

removeOverlap walks in-flight, computes the three-piece split for
each overlapping range (leftKept, covered, rightKept) and either
drops the covered portion (ACK) or pushes it onto retransmit
(loss). FIN belongs to the rightmost piece, so split at the end of
the original range correctly preserves it.

# Compaction

advanceFlushedFloorIfPossible bumps the floor whenever the lowest
in-flight / retransmit / unsent offset is above it, then
ByteArray.copyInto shifts the data window forward. Memory bounded
by `nextOffset - flushedFloor` rather than growing unboundedly.

# FIN handling

Treated as a virtual byte at offset = nextOffset. finPending arms
takeChunk to attach FIN to the final data chunk; finSent latches
true on emission; markLost(fin=true) clears it for re-emission;
finAcked latches true when the FIN-bearing range is ACK'd.

# Tests added (14, all pass)

  - takeChunk_releasesNothingUntilAcked
  - markAcked_full_releasesBytes_andDoesNotResend
  - markLost_movesBytesToRetransmitQueue_takeChunkReplaysSameOffset
  - markLost_partialRange_splitsInFlight
  - markAcked_partialRange_splitsInFlight
  - retransmitDrainsBeforeFreshBytes (priority order)
  - maxBytesSplits_acrossRetransmitAndFresh
  - fin_carriedOnFinalDataChunk_andRetransmittedOnLoss
  - fin_only_emittedAfterDataDrained
  - fin_only_lostAndRetransmits
  - markAcked_advancesFlushedFloor_releasesMemory
  - markAcked_outOfOrder_preventsFloorAdvance
  - markLost_belowFlushedFloor_isNoop (defensive)
  - finAcked_latchesTrueOnce
  - readableBytes_reflectsRetransmitPlusFresh
  - multipleSendsWithinSingleEnqueue_acksIndependently

Existing :quic tests (handshake, flow control, frame routing) +
nestsClient moq-lite + amethyst Android compile all pass.
SendBufferConcurrencyTest unchanged — the synchronized-on-this
discipline carries over from the prior implementation.

Wires nothing yet — commits C/D will route Stream/Crypto tokens
through markLost. The dispatcher in QuicConnection.onTokensLost
already calls markLost (added in commit A as a no-op stub); now
those calls actually do work.

https://claude.ai/code/session_01PYYez8a6sjiakyjAxsfCEQ
2026-05-04 23:20:39 +00:00
Claude 7f6d9085a4 feat(quic): extend RecoveryToken — Stream, Crypto, ResetStream, StopSending, NewConnectionId
Token-shape commit (commit A of the deferred-follow-ups pass that
extends RFC 9002 retransmit beyond the receive-side flow-control
extensions originally shipped at commit 9e6fa3d).

New variants on RecoveryToken sealed class:
  - Stream(streamId, offset, length, fin): STREAM data range we
    sent. On loss the dispatcher will re-queue the range on the
    stream's SendBuffer (commits C, D wire that).
  - Crypto(level, offset, length): CRYPTO bytes at a specific
    encryption level. RFC 9000 §17.2.5 + §13.3: handshake bytes
    are reliable; lost CRYPTO retransmits at the same encryption
    level. Commit E wires the per-level dispatch.
  - ResetStream(streamId, errorCode, finalSize),
    StopSending(streamId, errorCode),
    NewConnectionId(seq, retirePriorTo, cid, statelessResetToken):
    reliable per RFC 9000 §13.3, scaffolding-only since :quic
    doesn't currently emit any of these. The pendingResetStream /
    pendingStopSending / pendingNewConnectionId maps on
    QuicConnection are populated by the dispatcher but not yet
    drained by any writer code path.

QuicConnection.onTokensLost extended with the new variants:
  - Stream/Crypto: route to the matching SendBuffer.markLost
    (currently a no-op — see SendBuffer.markLost kdoc; commit C
    replaces the stub with the retain-until-ACK rewrite).
  - ResetStream/StopSending/NewConnectionId: write to the
    pending* maps; future emit code drains them.

NewConnectionId data-class needs explicit equals/hashCode because
its ByteArray fields would otherwise compare by identity (Kotlin
data-class auto-equals limitation). Tested.

Tests added (4):
  - whenDispatch_isExhaustive extended with all 10 variants;
    catches at compile time if a new variant is added without
    updating the dispatcher
  - newConnectionId_arrayEqualityIsByContent
  - stream_equalityByValue
  - crypto_equalityByValue

SendBuffer gains placeholder markLost(offset, length, fin) and
markAcked(offset, length) methods — both no-ops today. Their
signatures are stable so the dispatcher wiring lands once now and
doesn't need re-touching when commit C makes them functional.

Full :quic test suite passes.

https://claude.ai/code/session_01PYYez8a6sjiakyjAxsfCEQ
2026-05-04 23:14:42 +00:00
Claude c43c95184e feat(quic): steps 7, 8, 9 of RFC 9002 retransmit — PTO + integration test + revert workaround
Step 7: Probe Timeout (RFC 9002 §6.2).
  - QuicLossDetection.ptoBaseMs(maxAckDelayMs) computes
    `smoothed_rtt + max(4 * rttvar, 1ms) + max_ack_delay`.
  - QuicConnection gains pendingPing: Boolean and
    consecutivePtoCount: Int. The driver sets pendingPing = true
    when the PTO timer fires; the writer drains it as a PingFrame
    (smallest ack-eliciting frame). The peer ACKs the PING; that
    ACK feeds loss detection (steps 5–6) and triggers retransmit.
  - QuicConnectionDriver.sendLoop replaces the prior fixed 1-second
    placeholder with RFC 9002 PTO timing, doubling backoff per
    consecutivePtoCount per §6.2.2.
  - Parser resets consecutivePtoCount on any new ack-eliciting ACK.

Step 8: end-to-end integration test (RetransmitIntegrationTest, 2
tests). Drives the full chain (writer → SentPacket → loss detection
→ dispatch → re-emit) by simulating loss directly on
QuicConnection state, since the in-process pipe doesn't model loss:
  - maxStreamsUni_lostByPacketThreshold_isRetransmitted: emit a
    MAX_STREAMS_UNI, simulate ACK at PN+4 (above
    PACKET_THRESHOLD), verify retransmit lands in a NEW packet
    with a fresh PN.
  - lossDispatch_handlesSupersedeAcrossMultipleEmits: emit two
    successive MAX_STREAMS_UNI bumps (caps 6 and 10), declare the
    OLDER one lost. Supersede check drops the stale lost token —
    no retransmit because the newer cap covers it.

Step 9: revert the cap workaround.
  initialMaxStreamsUni: 1_000_000 → 10_000 (moq-rs's own default).
  The 1M value was a workaround for the moq-rs cliff that fired
  when our :quic emitted its first MAX_STREAMS_UNI extension. With
  retransmit now durable, a single dropped extension is recovered
  automatically — no need for the high-cap dodge. Lowering back
  exercises the rolling-extension path (and its retransmit) in
  production, which is what we want to validate the new code.

Tests added (4):
  - PtoTest x4: RFC 9002 §6.2.1 duration math
    (initial / with-ack-delay / after-rtt-sample / variance-floor)
  - RetransmitIntegrationTest x2: end-to-end retransmit cycle and
    supersede semantics

Full :quic test suite (~80 tests) + nestsClient moq-lite tests +
amethyst Android compile all green.

Closes the 9-step plan started at commit 9e6fa3d (step 1).

https://claude.ai/code/session_01PYYez8a6sjiakyjAxsfCEQ
2026-05-04 23:02:01 +00:00
Claude 15a6bfcc84 feat(quic): step 6 of RFC 9002 retransmit — dispatch lost tokens to pending*
QuicConnection.onTokensLost(tokens) closes the loop between loss
detection (step 5) and writer drain (step 4). For each lost token:

  - Ack: ignored (RFC 9000 §13.2.1: ACK frames not retransmittable)
  - MaxStreamsUni: pendingMaxStreamsUni := token.maxStreams
    iff token.maxStreams == advertisedMaxStreamsUni
  - MaxStreamsBidi / MaxData: same shape, against their advertised cap
  - MaxStreamData: pendingMaxStreamData[streamId] := maxData iff
    stream exists AND token.maxData == stream.receiveLimit

The supersede check (`lost == advertised`) mirrors neqo's
`fc.rs::frame_lost` line 322. If a higher extension has gone out
since, the older lost frame is irrelevant — the newer value
covers the receiver's grant. Without the check we'd resurrect
stale extensions and waste wire bandwidth re-emitting values the
peer already has.

Wired into the parser's AckFrame handler immediately after
detectAndRemoveLost: walk each lost packet's tokens and call
onTokensLost. Caller already holds the connection lock.

Tests added (8, all pass):
  - ackToken_doesNotPopulateAnyPending
  - lostMaxStreamsUni_matchingAdvertised_setsPending
  - lostMaxStreamsUni_supersededByHigherEmit_isDropped (the
    supersede-check invariant from neqo's fc.rs:322)
  - lostMaxStreamsBidi_matchingAdvertised_setsPending
  - lostMaxData_matchingAdvertised_setsPending
  - lostMaxData_supersededIsDropped
  - lostMaxStreamData_unknownStream_dropped (defensive)
  - multipleLostTokens_dispatchAll (one packet's worth of mixed
    tokens dispatched in one call)
  - lostTokensFromMultiplePackets_unionInPending (sequential
    dispatch across multiple lost packets — older stale, newer
    valid; the valid one survives)

Mirror of neqo's `streams.rs::lost` dispatch + the
`no_max_allowed_frame_after_old_loss` and
`set_max_active_equal_does_not_set_frame_pending` tests from
`fc.rs` deferred from step 4 per the plan.

Full :quic test suite + nestsClient moq-lite tests pass.

https://claude.ai/code/session_01PYYez8a6sjiakyjAxsfCEQ
2026-05-04 22:55:34 +00:00
Claude 1df6441639 feat(quic): step 5 of RFC 9002 retransmit — loss detection + RTT estimator
QuicLossDetection encapsulates the RFC 9002 §5–§6 algorithms:

  - RTT estimation (§5): smoothedRtt, rttVar, latestRtt, minRtt
    with first-sample bootstrap; ack-delay clamped against minRtt
    so a peer reporting a large delay can't push the estimate
    below its observed floor (§5.3 anti-exploit clamp).
  - Loss-delay (§6.1.2): max(latestRtt, smoothedRtt) * 9/8,
    clamped to GRANULARITY_MS (1 ms).
  - detectAndRemoveLost (§6.1): walks the in-flight set,
    removes entries that are either:
      - more than PACKET_THRESHOLD (3) PNs below largestAckedPn,
        OR
      - sent more than lossDelay ago.
    Returns the lost packets so step 6 can dispatch their tokens.

Wired into QuicConnectionParser's AckFrame handler:

  1. Snapshot largest-acked send time BEFORE drain
  2. Drain ACK'd packets (step 3)
  3. If largestAckedPn advanced AND any drained packet was
     ack-eliciting, update RTT (RFC 9002 §5.2 sample conditions)
  4. detectAndRemoveLost on the surviving set; lost list dropped
     for now — step 6 wires the dispatch

LevelState gains:
  - largestAckedPn: Long? (high-water mark for packet-threshold)
  - largestAckedSentTimeMs: Long? (RTT sample input)

QuicConnection gains:
  - lossDetection: QuicLossDetection (single instance, RTT is
    per-path; we model a single path)

Tests added (11, all pass):
  - firstRttSample_setsAllRttFieldsAtomically
  - secondRttSample_movesSmoothedRttTowardSample (math: 7/8 + 1/8)
  - ackDelay_clampedAgainstMinRtt (anti-exploit clamp)
  - negativeRttSample_isIgnored (clock skew defense)
  - lossDelay_floor (initial 333*9/8 = 374)
  - packetThresholdLost_removesPacketsBelowThreshold (PNs 0..6
    when largestAcked=10, threshold=3 → 7..9 survive)
  - timeThresholdLost_removesPacketsSentTooLongAgo (sentAt + delay
    < now → lost; recent → kept)
  - packetEqualToLargestAcked_notLost (edge case)
  - emptyMap_returnsEmpty
  - lostPackets_carryOriginalTokens (drain preserves token list)
  - packetThresholdAndTimeThresholdMatch_singleRemoval (no
    double-iteration)

Mirrors the subset of neqo's `recovery/mod.rs` tests in scope per
the plan: remove_acked, time_loss_detection_gap,
time_loss_detection_timeout, big_gap_loss,
duplicate_ack_does_not_update_largest_acked_sent_time. PTO tests
land in step 7.

Full :quic test suite passes.

https://claude.ai/code/session_01PYYez8a6sjiakyjAxsfCEQ
2026-05-04 22:52:49 +00:00
Claude 29282634e5 feat(quic): step 4 of RFC 9002 retransmit — pending* fields + writer drain
Adds the writer's drain side of the retransmit path. The QuicConnection
gains four `pending*` fields:

  - pendingMaxStreamsUni: Long?
  - pendingMaxStreamsBidi: Long?
  - pendingMaxData: Long?
  - pendingMaxStreamData: MutableMap<Long, Long>  (keyed by stream id)

Each non-null entry signals "the last extension we sent at this value
was lost; re-emit it". appendFlowControlUpdates now drains all four
ahead of the rolling-extension threshold check, emitting a fresh
frame + RecoveryToken for each pending entry and clearing it.

Step 4 only wires the consumer side; the setter side (loss
dispatcher) is step 6. Tests populate `pending*` directly to exercise
the drain in isolation.

Tests added (8, all pass):
  - pendingMaxStreamsUni / Bidi / MaxData / MaxStreamData each
    individually drain to a SentPacket carrying the matching token
  - multiplePending: all four pending types drain into one packet
    (writer drains them sequentially, no fan-out)
  - noPending: drain produces no extension tokens
  - pendingDrainBeforeThresholdCheck_supersedeOrderObservable:
    writer drains the pending value as-is — supersede check is the
    setter's responsibility (step 6), not the drain's
  - pendingClearedAcrossDrains: a cleared pending stays cleared on
    subsequent drains

Mirrors neqo's `fc.rs` retransmit tests in the plan
(`need_max_allowed_frame_after_loss`, `lost_after_increase`,
`multiple_retries_after_frame_pending_is_set`,
`new_retired_before_loss`). The supersede-check tests
(`no_max_allowed_frame_after_old_loss`,
`set_max_active_equal_does_not_set_frame_pending`) belong to step 6
and land there.

https://claude.ai/code/session_01PYYez8a6sjiakyjAxsfCEQ
2026-05-04 22:48:40 +00:00
Claude 0ced269b27 feat(quic): step 3 of RFC 9002 retransmit — drain SentPacket on ACK
QuicConnectionParser's AckFrame handler now drains
state.sentPackets of every entry whose packet number is covered by
the ACK's ranges. The drained SentPackets are returned but
discarded for now; step 5 will route them to loss detection / RTT.

New helper file `connection/recovery/AckedPackets.kt`:

  - forEachAckedPacketNumber(ack, block): inline iterator over an
    AckFrame's ranges in RFC 9000 §19.3.1 order. Walks first range
    [largestAcked - firstAckRange, largestAcked] then each
    additional range with `nextLargest = previousSmallest - gap - 2`,
    `nextSmallest = nextLargest - ackRangeLength`. Defensive clamp
    at PN 0 against malformed peer ACKs.
  - drainAckedSentPackets(sentPackets, ack): walks via
    forEachAckedPacketNumber and removes each matching entry from
    the map. Returns the drained list.

Wired into QuicConnectionParser.kt:165 alongside the existing
ackTracker.purgeBelow call.

Tests added (9, all pass):
  - simpleRange / multipleRanges / singlePacketAck: range walking
    for typical ACK shapes
  - ackForUnsentPn_isNoOp / emptyMap_returnsEmptyDrain: defensive
    paths
  - ackBoundary_pn0Inclusive: PN 0 is correctly included, no
    underflow
  - forEachAckedPacketNumber_iteratesDescending /
    forEachAckedPacketNumber_acrossMultipleRanges_descending:
    iterator semantics
  - returnedDrain_preservesTokens: drained SentPacket retains its
    full tokens list — step 5+ will dispatch these to RTT / loss

Mirror of neqo's `recovery/mod.rs::remove_acked` (one of the 20
recovery tests we owe per
`quic/plans/2026-05-04-control-frame-retransmit.md`). The remaining
loss-detection tests land in step 5.

Full :quic test suite + nestsClient moq-lite tests pass.

https://claude.ai/code/session_01PYYez8a6sjiakyjAxsfCEQ
2026-05-04 22:45:10 +00:00
Claude ea15a9afa1 feat(quic): step 2 of RFC 9002 retransmit — record SentPacket per outbound
Plumbs sent-packet retention into the writer. From now on every
Application packet emission stores a SentPacket in
`LevelState.sentPackets`, keyed by packet number, carrying:

  - the packet number (from `pnSpace.allocateOutbound()`)
  - the writer's `nowMillis` send time
  - whether the packet is ack-eliciting (RFC 9000 §13.2.1)
  - the encrypted on-wire size (or 0 if encrypt threw)
  - a list of RecoveryTokens — one per retransmittable frame in the
    packet (Ack token for ACK frames; MaxStreamsUni / MaxStreamsBidi
    / MaxData / MaxStreamData for the corresponding flow-control
    extensions)

`appendFlowControlUpdates` now takes a parallel `tokens: MutableList`
and writes lock-step with `frames`. The writer's existing semantics
are unchanged — same frames go on the wire, same advertised-cap
bookkeeping. Step 2 only adds the retention; nothing reads
`sentPackets` yet (steps 3–6 do that).

Order of operations on packet emission:
  1. Allocate packet number
  2. runCatching the encrypt step
  3. Record SentPacket regardless of encrypt outcome (sizeBytes=0 if
     it threw — the bookkeeping survives so loss detection can later
     declare the gap lost on the time threshold)
  4. Re-throw the encrypt exception so the driver loop sees the same
     error it did before this change

Tests added (3, all pass):
  - writer_records_sent_packet_with_max_streams_uni_token: cross
    half-window, drain, observe a SentPacket whose tokens contain
    MaxStreamsUni with maxStreams matching advertisedMaxStreamsUni
  - ack_only_outbound_records_sent_packet_with_ack_token_and_not_ack_eliciting:
    pending ACK only, observe a SentPacket with single Ack token and
    ackEliciting=false
  - successive_drains_record_distinct_packet_numbers: two drains
    record disjoint PN sets

Full :quic test suite passes (no regressions). nestsClient moq-lite
tests pass.

https://claude.ai/code/session_01PYYez8a6sjiakyjAxsfCEQ
2026-05-04 22:42:15 +00:00
Claude 9e6fa3d3f0 feat(quic): step 1 of RFC 9002 retransmit — RecoveryToken + SentPacket types
First step of `quic/plans/2026-05-04-control-frame-retransmit.md`.
Pure type definitions, no behavior change yet — sets up the data
shape the next steps will populate from the writer (step 2) and
drain from ACK / loss-detection paths (steps 3–6).

RecoveryToken: sealed class mirroring neqo's
neqo-transport/src/recovery/token.rs:21 StreamRecoveryToken.

  - Ack (singleton object): tracked but never retransmitted, so the
    sent-packet map invariant ("every retained entry has at least
    one token") holds for ACK-only packets too
  - MaxStreamsUni / MaxStreamsBidi: receive-side stream-id cap
    extension (RFC 9000 §19.11) — the frame whose loss tripped
    the moq-rs cliff
  - MaxData: connection-level data cap extension (§19.9)
  - MaxStreamData: per-stream data cap extension (§19.10)

SentPacket: data class mirroring neqo's
neqo-transport/src/recovery/sent.rs::Packet. Held in a per-pn-space
map on QuicConnection (step 2 wires this up). Carries packet number,
send time, ack-eliciting flag, on-wire size, and the token list to
dispatch on loss.

Tests: 6 token tests (data-class equality, sealed-hierarchy
exhaustiveness, Ack singleton-ness) + 6 SentPacket tests (equality
across fields, copy semantics, ACK-only-with-Ack-token convention,
multi-token packet shape). All pass; full :quic test suite still
passes — types are purely additive.

Out of scope as documented in the plan: STREAM data retransmit,
CRYPTO retransmit, RESET_STREAM/STOP_SENDING/etc, congestion control,
0-RTT. Those are separate follow-ups.

https://claude.ai/code/session_01PYYez8a6sjiakyjAxsfCEQ
2026-05-04 22:31:07 +00:00
Claude d391ae1db9 quic: emit MAX_STREAMS_* to extend peer's stream-id cap (fixes prod cliff at frame ~99)
flowControlSnapshot dump from sweep_frames_200 against nostrnests.com proved
the speaker side was perfectly healthy (~5 EB conn-credit unused, 10 000-stream
peer cap unused, 0 bytes stuck in send buffers, all 200 streams opened) yet
the listener cliffed at frame 99. The constraint was on the listener's
receive side: our :quic was advertising initial_max_streams_uni = 100 at
handshake and never extending it, so the relay could only open 100 uni
streams to us *for the lifetime of the connection*. Each Opus frame the
relay forwards is a fresh peer-initiated uni stream, so any broadcast
longer than ~100 frames silently truncated at the audience.

QuicConnection:
  - peerInitiatedUniCount / peerInitiatedBidiCount counters (incremented
    in getOrCreatePeerStreamLocked).
  - advertisedMaxStreamsUni / advertisedMaxStreamsBidi tracking, starting
    at config.initialMaxStreams* and raised by the writer.

QuicConnectionWriter.appendFlowControlUpdates:
  - When peerInitiatedUniCount + cfg.initialMaxStreamsUni / 2 >=
    advertisedMaxStreamsUni, emit MaxStreamsFrame(bidi=false, newCap)
    where newCap = peerInitiatedUniCount + cfg.initialMaxStreamsUni.
    Same pattern the existing MaxDataFrame / MaxStreamDataFrame
    extension uses; same half-window threshold so we don't spam the
    peer.
  - Symmetric branch for bidi.

PeerStreamCreditExtensionTest:
  - Writer DOES emit MAX_STREAMS_UNI when peer's lifetime uni-stream count
    crosses the half-window threshold; advertisedMaxStreamsUni updates.
  - Writer DOES NOT emit MAX_STREAMS_UNI below the threshold;
    advertisedMaxStreamsUni stays at the initial cap.

Updated nestsClient/plans/2026-05-01-quic-stream-cliff-investigation.md
with the fc-snapshot dump that proved the cause and the fix that
addresses it. The framesPerGroup=5 mitigation in NestMoqLiteBroadcaster
can be reverted to 1 once the prod sweep confirms the cliff is gone.

All :quic and :nestsClient JVM tests pass.
2026-05-01 14:21:52 +00:00
Claude a0e5e04964 quic: expose flow-control snapshot for prod cliff investigation
Adds a read-only diagnostic surface that lets a test (or any caller)
read the peer's transport parameters, the live connection-level send
credit / consumed counters, the current peer-granted MAX_STREAMS_*
values, and the total bytes sitting in stream send buffers but not
yet handed to STREAM frames.

Goal: pin which budget runs out at the production "stream cliff"
described in nestsClient/plans/2026-05-01-quic-stream-cliff-investigation.md.
The plan flagged three candidates — connection-level MAX_DATA, per-
stream MAX_STREAM_DATA, or the relay's MAX_STREAMS_UNI extension
policy. The snapshot makes it possible to attribute the stall by
reading the diff between the pre-pump, post-pump, and post-grace
snapshots from the test's stdout.

QuicConnection:
  - flowControlSnapshot(): suspend, lock-protected, returns
    QuicFlowControlSnapshot (new data class) — peer TPs + live
    accounting + sum of enqueued-not-sent bytes across streams.

QuicWebTransportSession (jvmAndroid adapter):
  - quicFlowControlSnapshot() passthrough so the test can downcast
    its WebTransportSession and read the underlying connection's
    state without poking through the common transport interface.

SendTraceScenario:
  - Optional flowControlSnapshot lambda parameter; when supplied,
    logs three checkpoints — fc-pre, fc-post-pump, fc-post-grace —
    each on a single line with the full snapshot.

NostrnestsProdAudioTransmissionTest + NostrNestsSustainedSendOutcomesInteropTest:
  - withProdSpeakerAndListeners / withHarnessSpeakerAndListeners now
    yield a snapshot lambda to the scenario block. Every sweep test
    automatically dumps fc-* lines to the JUnit XML system-out.

FlowControlSnapshotTest:
  - Pre-handshake: peer TP fields are null, counters zero.
  - Post-handshake: every TP field reflects what the in-process TLS
    server advertised; sendConnectionFlowCredit equals
    initial_max_data; consumed = 0.
  - Post-allocate-and-enqueue: nextLocalUni/BidiIndex advance,
    totalEnqueuedNotSentBytes sums the buffered chunks, and
    streamsWithPendingBytes counts only streams with > 0 pending.

Reading the production sweep output after this commit:
  - fc-pre dumps what the relay grants on handshake (initial_max_data,
    initial_max_stream_data_uni, initial_max_streams_uni).
  - fc-post-pump shows whether sendConnectionFlowConsumed has
    plateaued at the cap (peer didn't extend MAX_DATA) or whether
    bytes are stuck in stream send buffers.
  - The diff between fc-post-pump and fc-post-grace tells us
    whether the relay's eventual MAX_DATA / MAX_STREAMS update did
    or didn't arrive during the 30-60 s grace window.
2026-05-01 14:14:06 +00:00
Claude 96a585a67e Revert "quic: suspend openUni/BidiStream on peer-cap exhaustion + emit STREAMS_BLOCKED"
This reverts commit f0705e3ab1.
2026-05-01 13:59:38 +00:00
Claude f0705e3ab1 quic: suspend openUni/BidiStream on peer-cap exhaustion + emit STREAMS_BLOCKED
Production sweep showed every audio scenario opening >100 client-initiated
uni streams cliffed at received=99/N (one-line summary
[sweep-30s] sub[0] received=99/1500 missing=[99-1499]).
Same shape across every cadence, payload, and frame-count sweep variant —
the relay's initial_max_streams_uni=100 was being silently exhausted, after
which openUniStream threw QuicStreamLimitException, which the production
NestMoqLiteBroadcaster swallowed via its outer runCatching, dropping every
subsequent frame on the floor.

Fix:

  - QuicConnection.openBidiStream / openUniStream now SUSPEND when the
    peer-granted cap is reached, instead of throwing. They re-acquire
    the connection lock on each retry, so the parser's MAX_STREAMS update
    is observed atomically. Closing the connection wakes blocked openers
    with QuicConnectionClosedException so they don't hang.

  - QuicConnection.streamCapNotifier — single CompletableDeferred swapped
    after each fire so all blocked openers wake at once rather than
    serialising through Channel.receive.

  - QuicConnectionParser fires the notifier whenever an inbound
    MAX_STREAMS frame raises peerMaxStreams{Bidi,Uni}.

  - QuicConnectionWriter emits a STREAMS_BLOCKED frame (RFC 9000 §19.14)
    when an opener registers itself blocked, draining the slot once
    written so we send at most one STREAMS_BLOCKED per cap value.
    Frame.kt gains a real StreamsBlockedFrame class — previously the
    inbound bytes were just consumed and discarded.

  - QuicConnectionDriver.start wires connection.sendWakeupHook so an
    internal opener-blocked event nudges the send loop without callers
    needing a driver reference.

PeerStreamLimitTest rewritten:
  - "throws QuicStreamLimitException" → "suspends with withTimeoutOrNull"
  - Added: MAX_STREAMS_UNI frame wakes a suspended opener
  - Added: openUniStream queues the STREAMS_BLOCKED slot
  - Added: closing the connection unblocks waiters with the closed
    exception
  - Added: StreamsBlockedFrame round-trips through encode/decode

All :quic and :nestsClient JVM tests pass.
2026-05-01 12:52:22 +00:00
Vitor Pamplona f63e3b1c67 Merge branch 'main' of https://github.com/vitorpamplona/amethyst 2026-04-29 17:28:38 -04:00
Vitor Pamplona df98235d31 Minor adjustments to remove warnings 2026-04-29 17:16:14 -04:00
Claude a84fbd2e57 test(quic): add concurrent producer/consumer regression for SendBuffer
The previous SendBuffer suite (FlowControlEnforcementTest) is entirely
single-threaded — every test calls enqueue and takeChunk sequentially
on the same coroutine, so the race that crashed the audio path in
production (NoSuchElementException from chunks.first() under
concurrent enqueue + takeChunk) stayed invisible. The whole :quic
commonTest tree had no concurrent test at all.

Three new tests run real-thread races on Dispatchers.Default:

  - concurrent_enqueue_and_takeChunk_does_not_throw drives multiple
    producer coroutines + a consumer coroutine and asserts the buffer
    drains cleanly with no exception.
  - concurrent_takeChunk_callers_never_double_drain_a_chunk fans out
    multiple consumers against a pre-populated buffer; asserts the
    sum of bytes handed out equals the bytes enqueued (i.e. no chunk
    is double-counted by overlapping head-peel paths).
  - concurrent_finish_with_inflight_enqueue_emits_correct_fin races
    finish() against in-flight writes and asserts the FIN comes
    AFTER every enqueued byte.

Tests pass against the synchronised SendBuffer; running them against
the pre-fix unsynchronised version corrupts state badly enough that
the consumer wedges (an explicit "this is what the bug looked like"
demonstration). With internal synchronisation in place the suite
finishes in <0.2 s.

Documents the concurrent-access contract so a future "let's drop the
sync, it's hot" refactor immediately fails CI.
2026-04-29 21:08:03 +00:00
Claude 7f05fd6e2a fix(quic): round-5 audit fixes — concurrency, ack-eliciting flags, scope leaks
Two parallel audit agents inspected the round-4 commits for regressions and
concurrency hazards. Major findings:

ackEliciting regression (HIGH from core-regression report):
  Round-4's ACK gating optimization (only emit ACKs when something
  ack-eliciting was received) didn't update the parser's per-frame
  handling. MaxDataFrame, MaxStreamDataFrame, MaxStreamsFrame,
  NewConnectionIdFrame, HandshakeDoneFrame, ResetStreamFrame, StopSendingFrame,
  NewTokenFrame all need ackEliciting=true per RFC 9000 §13.2.1. Pre-fix a
  packet carrying only one of these would record the PN but never trigger
  an ACK, causing the peer to PTO-retransmit forever.

HandshakeDoneFrame conditional (HIGH):
  Pre-fix the dispatcher unconditionally set status=CONNECTED; if
  applyPeerTransportParameters had just called markClosedExternally
  (e.g. CID-validation failure), a later HANDSHAKE_DONE in the same
  payload would resurrect the connection. Now only sets CONNECTED when
  status is HANDSHAKING.

WT scope leak (CRITICAL from concurrency report):
  QuicWebTransportSessionState.close() never cancelled the scope holding
  the demux pump and capsule reader coroutines; both kept running past
  close, retaining QuicStream / chunk channels indefinitely. Memory
  growth on long sessions that opened/closed many WT sessions.

WtPeerStreamDemux.route() collector leak (CRITICAL):
  The route function launches a coroutine to drain stream.incoming into
  chunkChannel (UNLIMITED). Four early-return paths (truncated stream
  type, mismatched WT signal, foreign session id) returned without
  closing chunkChannel — collector kept running, channel grew unbounded.
  Now wrapped in coroutineScope{} so the collector is joined on every
  exit. Also explicitly cancels collector on the catch path.

RESET_STREAM stream-id ownership (HIGH):
  Pre-fix the dispatcher closed the local read side on whatever stream
  the peer named. RFC 9000 §3.5: the peer can only RESET_STREAM streams
  where it owns a send side. A peer RESETting a CLIENT_UNI is
  STREAM_STATE_ERROR (we own the only side). Now closes the connection
  in that case.

looksLikeIpLiteral tightening (HIGH):
  Pre-fix accepted "1.2.3.4.5", "1.2", "1." as IP literals — Java's
  InetAddress.getByName resolves all of those via DNS, defeating
  audit-4 #4's SNI-leak fix. Now strict: 4 dot-separated octets each
  in 0..255, or contains a colon (IPv6).

Signal channels closed on teardown (MEDIUM):
  closeAllSignals() helper closes peerStreamSignal +
  incomingDatagramSignal alongside closedSignal; pre-fix only closedSignal
  was closed and racing parser frames could still trySend into
  never-consumed channels. Centralised the call so close() and
  markClosedExternally both invoke it.

Driver.close() idempotency (HIGH):
  A second concurrent close() (common: session close + read-loop death
  racing) used to launch a parallel teardown that called scope.cancel()
  while the first's joinAll was mid-flight. Now memoizes the launched
  Job behind a synchronized block.

Driver.close() flush detection (MEDIUM):
  Pre-fix spun on `pendingDatagrams.isEmpty()` to detect
  CONNECTION_CLOSE flush, but the writer's CLOSING branch bypasses
  pendingDatagrams entirely. Now spins on `connection.status ==
  CLOSING`, which transitions to CLOSED only after drainOutbound builds
  the close packet.

@Volatile on peerMaxStreams* (MEDIUM):
  peerMaxStreamsBidi/Uni snapshots are documented lock-free; without
  @Volatile, JLS allows long-tearing on 32-bit JVMs and the JIT may
  cache stale values.

CertificateFactory parse inside try (MEDIUM):
  Malformed cert chain bytes used to throw raw CertificateException
  through the read loop. Now wrapped, so parse failure becomes a clean
  CONNECTION_CLOSE.

GOAWAY id-regression observability (MEDIUM):
  Pre-fix the QuicCodecException thrown on increasing GOAWAY id was
  silently swallowed by route()'s catch. Now also surfaces via
  peerGoawayProtocolError so the application/QUIC layer can act.

appendFlowControlUpdates uses streamsListLocked (perf):
  The round-4 perf #10 fix introduced streamsListLocked (no
  entries.toList per drain) but appendFlowControlUpdates still iterated
  the Map. Now also uses the index-friendly view.

peerCloseDeferred completion on session close (MEDIUM):
  awaitPeerClose() used to hang forever if the local side called close()
  before any peer-initiated WT_CLOSE_SESSION arrived. Now cancelled with
  CancellationException on local close.

Tests:
  AckElicitingFramesTest pins the ackEliciting contract on every round-5
  fix plus the ResetStream-on-CLIENT_UNI rejection.
  JdkCertificateValidatorIpLiteralTest pins the tightened pattern via
  reflection.

https://claude.ai/code/session_01EC1tfXfap8k8GyKvrxkxZx
2026-04-26 01:06:52 +00:00
Claude 920b36cdd6 perf(quic): round-4 perf-audit fixes — ACK gating, flow-control dirty-set, streams list view
Four perf wins from the round-4 audit, all in the steady-state hot path
(audio rooms run at ~50 datagrams/sec/participant).

Perf #1 — ACK frame gating (saves a frame + bandwidth on every drain):
  AckTracker.buildAckFrame now returns null when no new ack-eliciting
  packet has arrived since the last build. Pre-fix the writer emitted a
  redundant ACK frame on every outbound packet (~50/sec each direction)
  even when the only inbound traffic since last drain was ACK-only.
  RFC 9000 §13.2 only requires ACKs in response to ack-eliciting
  packets within max_ack_delay; gating on ackElicitingPending satisfies
  that without delay-timer machinery.

Perf #11 — AckTracker.purgeBelow short-circuit:
  Common case: peer ACKs a high PN, we already pruned below it. Pre-fix
  triggered a full ListIterator walk anyway. Now bails out when the
  tail's start is already above the threshold.

Perf #9 — Flow-control dirty-set:
  QuicStream gains a receiveDirtyForFlowControl flag set by the parser
  when readContiguous advances the frontier. The writer's
  appendFlowControlUpdates now skips per-direction-window lookup +
  threshold comparison for streams whose flag is unset. Big win for
  multi-stream sessions (audio rooms with N×M streams per participant).

Perf #10 — Streams list view:
  QuicConnection maintains a parallel insertion-ordered List<QuicStream>
  alongside the streams Map. Writer's round-robin scan reads the list
  directly instead of allocating `entries.toList()` per drain.
  No removal path exists today; the insert points (openBidi/UniStream,
  getOrCreatePeerStreamLocked) update both.

AckTrackerGatingTest pins the new gating contract: first build returns
a frame; second build without new reception returns null; subsequent
ack-eliciting reception re-arms; non-ack-eliciting receptions alone
don't.

https://claude.ai/code/session_01EC1tfXfap8k8GyKvrxkxZx
2026-04-26 00:49:20 +00:00
Claude 21da61ad64 test(quic): comprehensive regression tests for round-4 fixes + coverage holes
Pins every behavioural change made in the round-4 audit-fix commit so a
future regression can't quietly resurrect any of the bugs.

FrameRoutingTest (new):
  * RESET_STREAM / STOP_SENDING / NEW_TOKEN round-trip and don't kill
    the connection on arrival
  * Peer attempting CLIENT_BIDI / CLIENT_UNI stream IDs closes the
    connection (STREAM_STATE_ERROR)
  * MaxDataFrame raises sendConnectionFlowCredit; lower values ignored
  * CONNECTION_CLOSE returns immediately — frames after it are not
    dispatched (would otherwise create phantom streams on a closed
    connection)
  * HANDSHAKE_DONE at APPLICATION level is legal (ensures the level-
    validation guard didn't over-fire)
  * incomingDatagrams queue caps at MAX_INCOMING_DATAGRAM_QUEUE; oldest
    entries dropped on overflow

ReceiveBufferFinTest (new): the audit-4 #4 silent-truncation fix
  * isFullyRead() stays false when FIN arrives before a gap fills
  * isFullyRead() flips true only after contiguous-end reaches finOffset
  * Zero-length FIN frame at exact end marks stream complete
  * finOffset is pinned at first observation (RFC 9000 §4.5)

QuicConnectionWriterTest (new): drainOutbound paths previously untested
  * CLOSING-status drain produces a CONNECTION_CLOSE packet
  * appendFlowControlUpdates raises stream.receiveLimit after consumer
    drains > half window
  * Writer enforces sendConnectionFlowCredit cap (audit-4 #9 — never
    exceeds the peer's initial_max_data even with more bytes queued)

JcaAesGcmAeadTest (new, jvmTest): JVM-platform AEAD round-trip
  * seal → open round-trip
  * Different nonces produce different ciphertexts
  * Rebuild path (same nonce twice) uses fallback cipher and still opens
  * Corrupted ciphertext / wrong AAD → null
  * key/nonce/tag length constants

ChaCha20Poly1305AeadTest (new): the seal side that prior tests skipped
  * seal → open round-trip
  * Bad tag / wrong AAD → null
  * Wrong-size key/nonce throws IllegalArgumentException

WtPeerStreamDemuxTest: GOAWAY id-regression branch (audit-4 #5)
  * Increasing GOAWAY id is rejected; previously recorded id stays put

CapsuleReaderTest: new strictness assertions
  * WT_CLOSE_SESSION body < 4 bytes throws QuicCodecException
  * Reason > 8192 bytes throws QuicCodecException

Notes:
  * QuicConnectionDriver direct unit tests would require turning UdpSocket
    from `expect class` into an interface; deferred — driver paths are
    exercised end-to-end by InteropRunner and indirectly via every pipe-
    based test.
  * Decrypting client-emitted packets in tests requires server-side keys
    (server's RX = client's TX with different cached cipher state in JCA);
    writer tests assert side-effects on connection state instead.

https://claude.ai/code/session_01EC1tfXfap8k8GyKvrxkxZx
2026-04-26 00:38:55 +00:00