aioquic multiplexing qlog showed:
t=31424: still receiving STREAM frames
t=31493: PTO timer expired
t=31507: connection_closed (owner: local)
We were still actively transferring when our 30s timeout fired. The
multiplexing testcase generates many small files (3431 sim packets
captured for the run) and download throughput on Mac+Rosetta is
dominated by per-write filesystem overhead in the Docker volume mount.
60s gives enough headroom without making fast-completing tests slower.
Doesn't fix any real protocol issue — just lets the test budget match
the workload.
If multiplexing still fails after this, the next investigation is the
sequential `.map { it.await() }` pattern in runTransferTest: a single
hung stream blocks the await chain even though others completed.
https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
aioquic's retry test result, surfaced via qlog:
Check of downloaded files succeeded.
Client reset the packet number. Check failed for PN 0
Our applyRetry called LevelState.resetForVersionNegotiation, which
creates a fresh PacketNumberSpaceState() — resetting PN to 0. The
qlog confirmed: PN=0 sent at t=388 (pre-Retry ClientHello), then PN=0
again at t=1468 (post-Retry retried ClientHello). Same PN reused
across the boundary.
RFC 9001 §5.7 + RFC 9000 §17.2.5: the Initial PN namespace CONTINUES
across the Retry boundary. The new Initial keys are derived from the
new DCID, but PN doesn't reset. Reusing a PN under different keys
makes the runner's pcap-decryption check fail (it's also a security
concern in the general case, hence the strict spec rule).
Fix: new LevelState.resetForRetry that's identical to
resetForVersionNegotiation EXCEPT it preserves pnSpace. applyRetry
calls resetForRetry. Two regression tests updated to assert the
post-Retry Initial uses PN=1 (continues from PN=0 of the pre-Retry
attempt) rather than PN=0 (the buggy reset behavior).
For Version Negotiation the original semantics still apply (RFC 9000
§6.2: client treats VN as if the original Initial was never sent;
PN reset to 0 is correct).
This should bring the retry testcase from ✕(S) to ✓(S) against
servers that exercise the Retry path. The handshake / transfer
already succeeded over the Retry per the qlog (the server's check
"Check of downloaded files succeeded." passed); only the PN-reuse
flag was failing the test.
https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
Direct evidence from aioquic interop qlog:
packet_sent PN=0 frames=[crypto] (ClientHello)
packet_sent PN=1 frames=[ping] (PTO probe — bare PING)
packet_received frames=[connection_close]
reason="Packet contains no CRYPTO frame"
aioquic strictly rejects pre-handshake Initials that contain no
CRYPTO frame. Our PTO probe was a bare PING, not a CRYPTO retransmit.
Agent 2's c9e036f72 implemented the right behavior: on PTO, the
driver requeues all inflight CRYPTO at the highest pre-application
level, so the next drain emits a fresh CRYPTO frame at the original
offset. The pendingPing flag stays as a fallback only used when no
CRYPTO is available to retransmit. That logic was wiped during the
qlog observer merge — driver's PTO branch only set pendingPing,
nothing requeued anything, probe was always bare PING.
Restored. Should fix the post-flush-fix regression where aioquic
went 1/7 (only transfer green) — once a packet was dropped or PTO
fired, the next probe was bare PING, aioquic closed with
"no CRYPTO frame", the test failed. Sequential tests on the same
matrix invocation likely intermittently passed/failed depending on
whether PTO fired before the response arrived.
The :quic helpers requeueAllInflightCrypto + highestPreApplicationLevel
already exist on the branch — just the driver call wasn't wired.
https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
aioquic still 0/7 after the close-race fix because there's a deeper
issue. QlogWriter.writeLineLocked() called writer.flush() on every
event. On macOS Docker Desktop the filesystem is virtualized and
per-event flush is multi-millisecond. The qlog hooks fire from
QuicConnectionWriter.drainOutbound() which runs INSIDE
conn.lock.withLock { ... } on the send loop. Every flush stalls
that lock; the receive loop blocks indefinitely trying to acquire
it for feedDatagram().
Symptom matched: client.sqlog showed our ClientHello packet_sent
at t=400ms, then NOTHING for the rest of the test — no
packet_received (server's response never decoded), no PTO probe
(send loop stuck behind disk IO).
Fix: remove per-event flush. close() flushes once at the end of
the run, which gives us the trace for everything except a hard
JVM kill mid-test (acceptable trade). BufferedWriter still buffers
in-memory; we don't lose events, just don't sync them to disk
synchronously.
Should restore aioquic to 5/7. picoquic was already 5/7 because
its server-response timing apparently let our send loop drain
between flushes; aioquic's faster turnaround tripped the lock more
reliably.
https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
aioquic full-matrix run came back 0/7 with stack traces:
java.io.IOException: Stream closed
at QlogWriter.writeLineLocked(QlogWriter.kt:289)
at QlogWriter.onPacketSent(QlogWriter.kt:126)
at QuicConnectionWriter.emitQlogSent(...)
The send loop runs concurrently with InteropClient's teardown sequence:
1. driver.close() — launches CLOSING → CLOSED async
2. qlogWriter?.close() — closes the file
3. delay(50)
4. scope.cancel() — kills coroutines
Between (2) and (4), the send loop is still alive and emits
packet_sent for the CONNECTION_CLOSE Initial. QlogWriter.writeLineLocked
threw IOException into the send loop, propagated up, took down the
connection — including connections that had completed transfer
successfully.
Fix: QlogWriter swallows post-close IOExceptions silently. An observer
must NOT break the connection it's observing — that's the whole NoOp
contract. Adds @Volatile closed flag, latches it both on explicit
close() and on the first IOException; subsequent emits short-circuit.
picoquic was 5/7 (predictable: M + S still open) so the regression was
specifically aioquic-timing-sensitive. With this fix aioquic should
return to 5/7 too.
https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
Documents:
- Three overnight agents merged (VN defense in :quic, qlog observer
+ JSON-NDJSON writer wired into the prod endpoint, peer-uni-stream
drainer fixing the multiplexing tear-down).
- The agent-A versionnegotiation testcase mismatch — runner doesn't
have that name; v2 is the closest, tests QUIC v2 which we don't
speak. VN code stays as defensive support.
- Open issues for tomorrow: retry test failure (qlog now available
for diagnosis), v2 (deferred), re-runs against all three peers
with the new fixes.
- Documented run-matrix.sh's concurrency limit.
https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
Agent B's QlogWriter previously lived only in :quic's jvmTest scope,
hooked into the standalone InteropRunner. Brings it into :quic-interop
proper:
- QlogWriter.kt + QlogWriterTest.kt copied into :quic-interop with
package com.vitorpamplona.quic.interop.runner.
- Jackson dep added to :quic-interop's build.gradle.kts.
- InteropClient reads $QLOGDIR (the runner sets it inside the
container per docker-compose.yml). When set, opens a QlogWriter
at <QLOGDIR>/client.sqlog, passes as QuicConnection.qlogObserver,
closes on every exit path (handshake_failed / udp_failed /
transfer_timeout / ok).
- QUIC_INTEROP_DEBUG=1 now prints qlogdir alongside the other env
fields.
Net effect: any future runner-driven test failure dumps a structured
qlog file the user can drag straight into qvis (qvis.quictools.info)
to see frame-by-frame what the client decided to do. The retry test
failure on aioquic + picoquic that we still need to debug now produces
a usable artifact.
NOTE: the QlogWriter copy in :quic's jvmTest stays in place as the
helper for the standalone InteropRunner main(). Slight duplication;
acceptable while :quic-interop is its own module.
https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
Two corrections from the user's just-completed run:
1. The runner does not have a 'versionnegotiation' testcase. Available
list output: handshake, transfer, longrtt, chacha20, multiplexing,
retry, resumption, zerortt, http3, blackhole, keyupdate, ecn,
amplificationlimit, handshakeloss, transferloss, handshakecorruption,
transfercorruption, ipv6, v2, rebind-port, rebind-addr,
connectionmigration. (`v2` exists but tests QUIC v2 protocol
support — we're v1-only, so it would correctly fail.)
The :quic-side VN code from agent A stays as defensive support for
any future server that throws a VN at us; just no testcase wires
it up.
2. run-matrix.sh is NOT safe to run concurrently. The runner's
docker-compose.yml hardcodes container_name: sim/server/client —
Docker enforces those globally regardless of COMPOSE_PROJECT_NAME.
Documented the limitation + the recommended sequential loop in
the script's comments.
https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
Same merge-from-main shape as the prior agent A integration: the qlog
agent's worktree didn't carry the version-negotiation work (QuicVersion
import, currentVersion / vnConsumed fields, applyVersionNegotiation),
the Retry work (extraSecretsListener / cipherSuites / applyRetry), or
the version-negotiation testcase wiring. Merge with -X theirs took the
qlog version of QuicConnection.kt + QuicConnectionWriter.kt + the
existing InteropRunner.kt wholesale, dropping those.
Restored:
- QuicVersion import in QuicConnection.kt + QuicConnectionWriter.kt +
the test-side InteropRunner.kt (also touched by agent B's qlog
hooks).
- extraSecretsListener / cipherSuites / initialVersion ctor params
on QuicConnection (qlogObserver kept; new qlog work landed).
Net result: all three overnight agents (A versionnegotiation, B qlog
observer, C peer-uni-stream drainer) now coexist on the branch with no
references missing. Full :quic:jvmTest green; :quic-interop:test +
installDist green.
https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
Two integrations from overnight agents:
1. agent C — peer-uni-stream drainer for the H3 multiplexing fix.
Http3GetClient.init() now takes a CoroutineScope and calls
conn.drainPeerInitiatedUniStreamsIntoBlackHole(scope) to consume +
discard the server's three uni streams (control, qpack-encoder,
qpack-decoder). RFC 9114 §6.2 mandates the client read these;
pre-fix their bytes accumulated in 64-chunk per-stream channels
until parser overflow tore down the connection with INTERNAL_ERROR
~4.5s into a multiplexing run.
2. agent A — versionnegotiation testcase wired into dispatch.
Adds the testcase to the runTransferTest route and threads
QuicVersion.FORCE_VERSION_NEGOTIATION as the initial version when
testcase=='versionnegotiation'. The QuicConnection's RFC 9000 §6
VN flow (applyVersionNegotiation) takes over from there, retrying
with v1 once the server replies with its supported list.
https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
The agent A worktree was based on main, so its QuicConnection.kt
didn't carry the Retry handling from d03e17981. Merging with
-X theirs replaced the file wholesale, dropping applyRetry +
extraSecretsListener / cipherSuites constructor params + the
RetryPacket import in the parser.
Restored:
- extraSecretsListener + cipherSuites ctor params (used in tlsListener
+ the TlsClient construction site).
- applyRetry method, now using the version-negotiation-introduced
LevelState.resetForVersionNegotiation helper (functionally
equivalent to the prior restoreFromRetry it replaced).
- RetryPacket import in QuicConnectionParser.
Also dedupes a duplicate `originalClientHello` field that the merge
left both copies of (one from agent 3's Retry work, one from agent A's
VN work). Single field now serves both reset paths.
https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
Variant (B) from the three-way fix menu in the multiplexing-interop
investigation: keep `:quic` strict about per-stream backpressure (the
audit-4 #3 "INTERNAL_ERROR: stream … consumer overflowed" tear-down
stays the contract for app-data overflow on bidi streams) but expose
an explicit, opt-in helper for peer-initiated UNI streams that the
application has decided it does not need to interpret.
Root cause confirmed in QuicConnectionParser.kt:290: when the server
opens its three RFC 9114 §6.2.1 peer-uni streams (CONTROL +
QPACK_ENCODER + QPACK_DECODER) and the H3 client does not consume
them, the parser routes their bytes into each stream's bounded
incomingChannel (capacity 64). Once the QPACK encoder issues
dynamic-table inserts beyond 64 chunks the next chunk overflows
trySend, sets QuicStream.overflowed, and the parser maps that to
markClosedExternally — the entire connection dies.
Notes on scope:
- The `Http3GetClient` and `:quic-interop` runner mentioned in the
investigation prompt do NOT exist on the `main` worktree this
branch starts from. The fix here is therefore `:quic`-only: the
public `awaitIncomingPeerStream` API was already sufficient for
an integrator to write the accept loop themselves; this commit
wraps the common case in `drainPeerInitiatedUniStreamsIntoBlackHole`
and updates the doc on `awaitIncomingPeerStream` so the next
integrator landing the H3 GET client doesn't hit the same trap.
- Variant (C) — silent default drain in `:quic` itself — was
deliberately rejected: defaults that swallow application bytes
are the misconfiguration we want type-system-or-API-explicit.
The new helper requires the caller to pass a CoroutineScope, so
opt-in is unmistakable in any callsite.
Regression test coverage in PeerUniStreamDrainTest:
- pre_fix_no_consumer_overflows_and_tears_down_connection — pushes
65 chunks (capacity + 1) on a SERVER_UNI stream with no consumer;
asserts the connection transitions to CLOSED. Pins the existing
backpressure contract.
- drainPeerInitiatedUniStreamsIntoBlackHole_keeps_connection_alive
— same setup but with the new helper running on a side scope;
pushes 256 chunks (4× capacity) and asserts the connection stays
CONNECTED. With the helper sabotaged, this test fails at
line 119 with status=CLOSED, confirming it actually exercises
the fix.
Full quic test suite: 295 tests, 0 failures, 0 errors.
https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
Adds the client-side VN flow needed for the interop runner's
`versionnegotiation` testcase:
- `QuicConnection` accepts an `initialVersion` constructor parameter
(default `QuicVersion.V1`) and exposes a mutable `currentVersion`
the writer stamps into outbound long-headers. `start()` now caches
the ClientHello bytes for VN-driven re-emission.
- `applyVersionNegotiation(supportedVersions)` validates per §6.2
(anti-downgrade: reject if list contains the offered version),
picks v1 from the offered set, regenerates DCID, re-derives
Initial keys against the new DCID, resets the Initial level via
`LevelState.resetForVersionNegotiation`, re-enqueues the cached
ClientHello, and latches `vnConsumed` so a second VN is dropped.
Failure to find a mutually supported version closes the connection
with `QuicVersionNegotiationException`.
- `QuicConnectionParser.feedDatagram` detects `version == 0` long
headers BEFORE peekHeader (whose layout assumes v1) and dispatches
to a new `feedVersionNegotiationPacket` that parses the §17.2.1
shape and validates the echoed DCID.
- `QuicConnectionWriter` reads `conn.currentVersion` instead of the
hardcoded `QuicVersion.V1`.
- `QuicVersion.FORCE_VERSION_NEGOTIATION = 0x1a2a3a4a` for the
interop runner.
- `InteropRunner` honors `TESTCASE=versionnegotiation` (or
`-DinteropTestcase=`) and offers the force-VN version.
Regression coverage in `VersionNegotiationTest`:
- happy path: VN switches `currentVersion` to v1, regenerates DCID,
resets PN, and the next drain emits a v1 Initial on the wire.
- downgrade defense: VN listing the offered version is dropped.
- unsupported list: VN whose versions we can't speak fails the
handshake and closes the connection.
- second VN: post-consumption VN is ignored.
- DCID mismatch: spoofed VN with wrong echoed DCID is dropped.
- backward compatibility: default `initialVersion` keeps v1 behavior
for existing callers.
https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
Three agents in flight in worktrees:
A — versionnegotiation testcase + configurable initial-version writer
B — qlog observer infrastructure (QlogObserver interface in :quic,
JSON-NDJSON writer in :quic-interop, hooks at packet/key/recovery
sites, reads $QLOGDIR the runner already sets)
C — multiplexing channel-saturation fix (consume server's peer uni
streams in Http3GetClient — RFC 9114 §6.2 says clients MUST
process incoming control + QPACK streams)
Plan doc now also captures the latest test matrix (pre-acfe815e1
multi-ALPN-offer fix) and the predictions for the next run.
https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
The previous commit (e5bbf8509) switched ALPN per testcase based on
quic-interop-runner convention (h3 for http3/multiplexing, hq-interop
otherwise). That broke picoquic, which had been 4/4 green: picoquic-qns
strictly registers h3 ALPN and rejects hq-interop. Different servers
disagree on which ALPN they configure for the same testcase:
- quic-go-qns — strictly hq-interop for non-http3
- aioquic-qns — accepts either
- picoquic-qns — strictly h3 for all testcases
There's no per-testcase convention all peers honor. Right move is the
TLS-spec-supported one: offer BOTH h3 and hq-interop in the ClientHello,
let the server pick whichever matches its config, then dispatch the GET
client by `tls.negotiatedAlpn` after the handshake completes. http3 and
multiplexing still restrict to h3 only since they exercise H3 framing.
Brings picoquic back to fully green and should unblock quic-go for the
non-http3 testcases.
https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
Two more testcases now dispatched through runTransferTest:
- retry — exercises the RFC 9000 §17.2.5 + RFC 9001 §5.8 Retry
handling that landed in d03e17981 / 671f9c705 (DCID swap,
integrity-tag verify, key re-derivation, token threading).
- ipv6 — same flow over an IPv6 socket. JDK's DatagramChannel.connect
handles the v6 address resolution natively; if anything
breaks it'll be an actual bug worth surfacing.
Plan doc updated to reflect Phase 3 landings (ALPN per testcase,
HqInteropGetClient, multi-stream FIN delivery fix) and the current
validation matrix across aioquic / picoquic / quic-go.
https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
quic-go-qns interop revealed two coupled gaps in our endpoint:
1. ALPN per testcase. The runner convention (followed strictly by
quic-go-qns, lazily by aioquic / picoquic) is:
- testcase 'http3' / 'multiplexing' → ALPN 'h3' (full HTTP/3)
- everything else → ALPN 'hq-interop'
(HTTP/0.9 over QUIC)
We had hardcoded 'h3'. quic-go closed every non-http3 connection
with CRYPTO_ERROR 0x178 (TLS no_application_protocol).
2. We had no HQ-interop client. HQ is dead simple: open bidi, send
`GET /path\r\n` raw, FIN, read body verbatim until server FINs.
No framing, no QPACK, no control stream.
This commit adds:
- GetClient interface + GetResponse data class extracted from
Http3GetClient so the runner code dispatches uniformly.
- HqInteropGetClient — the 30-line HQ-interop GET implementation.
- Alpn enum (H3, HQ_INTEROP) wired through main() → runTransferTest
→ QuicConnection.alpnList. Picked from the testcase name per the
convention above.
Validated locally: :quic-interop:test green; :quic-interop:installDist
clean. Will need a real run against quic-go to confirm the fix.
https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
When QuicConnection tears down (CONNECTION_CLOSE, read-loop death,
INTERNAL_ERROR from a saturated stream channel, etc.) the per-stream
incomingChannel objects were left open, so any application coroutine
suspended on `stream.incoming.collect { … }` hung forever waiting for
a FIN that would never come. The connection-wide signal channels
(closedSignal, peerStreamSignal, incomingDatagramSignal) all closed
cleanly, but the per-stream Flows did not — surfacing in the
quic-interop-runner `multiplexing` case as 677 collectors stuck after
the connection died mid-response, so zero of the 1999 expected files
landed.
Fix: closeAllSignals() now also calls closeIncoming() on every stream
in streamsList. Channel.close() is idempotent, and consumeAsFlow drains
already-buffered chunks before honouring the close, so any bytes the
parser had already pushed are still surfaced to the collector before
the Flow terminates.
Adds MultiStreamFinDeliveryTest covering: (a) FIN delivery to N parallel
client-bidi streams, (b) connection-teardown unblocks every per-stream
Flow, (c) buffered bytes survive a teardown without an explicit FIN.
make smoke was the bisector for "is the bug in :quic or in the runner
environment?" while we were debugging the initial close-path / PTO /
padding issues. Now that the runner reliably runs the matrix end-to-end
(handshake + chacha20 green vs aioquic), smoke's only job — running
picoquic outside the runner — is unneeded, and the picoquic image
keeps changing its entrypoint / required args in ways that make smoke
finicky to maintain.
Removes:
- make smoke / smoke-down targets
- SMOKE_NET / SMOKE_PICOQUIC / SMOKE_CLIENT vars
- SMOKE_MODE handling in run_endpoint.sh (just always tolerates
/setup.sh failure now — same effect, less ceremony)
- Plan doc note about make smoke updated to reflect removal.
If we ever need a non-runner bisector again, it's two `docker run`
commands; not worth permanent maintenance.
https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
The runner aggregates container stderr verbatim, which gave us a
~50-line block of routing setup, NIC checksum offload toggles,
container lifecycle messages, and the long Command: WAITFORSERVER=...
line for every single testcase. Useful once for triage; pure noise
across a multi-test matrix run.
Two changes:
- run-matrix.sh pipes the runner output through a grep -Ev filter
that drops the boilerplate while keeping outcomes (Test: ... took /
status, the summary table, server's Starting server, sim scenario
+ capture lines, and any Python tracebacks). VERBOSE=1 bypasses
the filter for debugging.
- InteropClient drops its own pre-test header dump unless
QUIC_INTEROP_DEBUG=1 is set; the per-GET success line goes silent
too — failures still print as before.
Net result: a passing test reduces from ~50 noise lines to ~5
meaningful lines (test name, time, status, summary table).
https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
Two coupled bugs in our endpoint that the runner just surfaced:
1. The runner mounts \$CLIENT_DOWNLOADS to /downloads as a Docker volume
(per quic-interop-runner's docker-compose.yml `client.volumes`).
There is no DOWNLOADS env var. Our endpoint was reading \$DOWNLOADS,
getting null, and (for handshake/chacha20) skipping any download path.
Hard-code /downloads.
2. The handshake / chacha20 / handshakeloss testcases don't just verify
the handshake completes — they also require the requested file at
/downloads/<basename>. The runner's _check_files validator
reported "Missing files: ['intense-tremendous-firefighter']" while
our client side reported `handshake ok`. Route handshake-flavor
testcases through runTransferTest so the full H3 GET pipeline runs.
`chacha20` keeps the cipher-suite override (ChaCha20-only ClientHello);
`handshakeloss` reuses the same H3 flow against a lossy sim.
Also drops the now-dead runHandshakeTest + parseFirstTarget helpers.
https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
pyshark's pcap reader calls asyncio.get_event_loop_policy().get_event_loop(),
which raises RuntimeError on Python 3.14 (deprecated in 3.12, removed in
3.14). Symptom: the matrix run completes the actual interop test
successfully but the runner's pcap-validation step crashes before it can
emit the result, dropping a 100+ line traceback.
run-matrix.sh now scans for python3.13 → 3.12 → 3.11 → python3, picking
the first that's both installed and < 3.14, and creates the venv with it.
Existing venvs created with 3.14 keep working (only matters for fresh
clones); for an existing broken venv, rm -rf .venv and re-run.
https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
Confirmed via `docker run --entrypoint find privateoctopus/picoquic:latest
/ -name picoquicdemo` — the binary isn't on PATH inside the image.
Use the absolute path so smoke runs without depending on PATH defaults.
https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
The privateoctopus/picoquic:latest image now wraps its CMD in the
runner's /run_endpoint.sh, which expects ROLE/TESTCASE/CERTS env vars
and exits 127 with "Unsupported test case:" if they're absent. So the
old `... privateoctopus/picoquic:latest picoquicdemo -p 4433` form
silently doesn't start picoquicdemo at all — picoquic exits, our
client times out trying to talk to a dead container.
Override the entrypoint so picoquicdemo runs directly with our args.
https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
The Linux build path of the vlc-setup plugin pulls vlc-plugins-linux from
Maven Central (ir.mahozad:vlc-plugins-linux), which only ships 3.0.20 and
3.0.20-2 — there is no 3.0.21 artifact yet. With vlcVersion set to 3.0.21,
both the CI pre-fetch step and the in-Gradle vlcDownload task hit a 404.
Match VLC_VERSION in build.yml and document the lag so future bumps wait
for the Maven artifact to catch up.
https://claude.ai/code/session_01GrZLMi3sdp6frwREmQ9cUi
The Retry parser + integrity-tag verifier already existed in
RetryPacket.kt, but feedDatagram dropped Retry packets on the floor.
Hook them up:
- QuicConnectionParser.feedLongHeaderPacket detects RETRY type before
the standard parse-and-decrypt path, parses via RetryPacket, and
dispatches to QuicConnection.applyRetry.
- QuicConnection.start() now caches the ClientHello bytes (TLS only
emits ClientHello once; we need to re-queue the same bytes on the
fresh Initial keys after Retry). New applyRetry method:
verifies the integrity tag, swaps DCID to Retry's SCID, re-derives
Initial keys, resets the Initial PN space + sentPackets +
cryptoSend, re-enqueues the cached ClientHello, stores the Retry
token, and latches retryConsumed so a second Retry is dropped.
- LevelState.restoreFromRetry / PacketNumberSpaceState.resetForRetry
give applyRetry an in-place reset (the level reference is a `val`,
so we mirror discardKeys' field-reset pattern).
- QuicConnectionWriter.buildLongHeaderFromFrames threads
conn.retryToken through the Initial header's Token field on every
Initial we emit after Retry.
Per RFC 9001 §5.8, a Retry with a bad integrity tag is silently
dropped; per RFC 9000 §17.2.5.2, only one Retry is honored per
connection. Both invariants are tested.
New test: RetryHandlingTest covers the happy path (DCID swap, PN
reset, token threading, ClientHello replay, ≥1200-byte padding),
the bad-tag path, and the second-retry path.
Even with /setup.sh skipped, the base image
(martenseemann/quic-network-simulator-endpoint) ships a /etc/resolv.conf
primed for the runner's sim that breaks Docker bridge name resolution
inside the JVM on macOS. Bypass DNS entirely: docker inspect the
picoquic container's IP and pass it directly via REQUESTS.
https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
Pre-handshake PTO previously only set `pendingPing`, which collapsed to
either nothing (the bug 86b6c609a fixed in a sibling branch) or a bare
PING. Against aioquic in the quic-interop-runner ns-3 sim, the first
ClientHello can be dropped (server not ready at t≈0.5s); a follow-up
PING with the same DCID is silently ignored because the server has no
state for that DCID. We need to retransmit the actual ClientHello bytes
so the server sees a full connection attempt.
Implementation:
- SendBuffer.requeueAllInflight(): walks the inflight list and moves
every sent-but-not-ACK'd range to the retransmit queue, preserving
offset and FIN. Mirrors markLost's per-range path but applies to all
inflight ranges in one shot. Idempotent + best-effort safe.
- QuicConnection.requeueAllInflightCrypto(level): thin wrapper that
drives the per-level cryptoSend buffer's new method.
- QuicConnectionDriver.sendLoop: PTO branch now calls the new helper
at the highest active pre-application level (Handshake > Initial)
when 1-RTT keys aren't installed. The next drain naturally emits a
CRYPTO frame at the original offset (takeChunk drains the
retransmit queue first per existing semantics).
- QuicConnectionWriter.collectHandshakeLevelFrames: honors pendingPing
pre-handshake at the highest active level — but skips the PING when
a CRYPTO frame is already in the same level's frame list (the
CRYPTO retransmit covers the ack-eliciting requirement). Bare PING
still goes out when there's nothing to retransmit.
Old RecoveryToken.Crypto entries in sentPackets for the original PNs
remain harmless: when loss detection eventually declares them lost,
markLost re-runs against ranges that have already moved on, which is
itself idempotent (clamps to flushedFloor / no-op on already-queued).
Test: PtoCryptoRetransmitTest reproduces the wire scenario — first drain
emits ClientHello in a ≥1200-byte Initial datagram; simulated PTO
calls requeueAllInflightCrypto + sets pendingPing; second drain must
contain a CRYPTO frame at the same offset with the same payload bytes,
not a bare PING. Datagram size still ≥1200 (RFC 9000 §14.1).
https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
The padding-rebuild branch in QuicConnectionWriter.drainOutbound computed
`padBytes = 1200 - natural`, but the QUIC long-header Length field is a
varint (RFC 9000 §16). When the natural-size payload was small enough for
Length to fit in 1 byte (body ≤ 63 bytes), the rebuild's larger body
crossed the 64-byte threshold and Length grew to 2 bytes — adding 1 wire
byte that wasn't in `natural`. PING-only PTO probe Initials therefore went
out at exactly 1199 bytes, one short of the §14.1 floor.
Fix: rebuild iteratively. After the first rebuild, measure the actual
datagram size; if still < 1200, bump padBytes by the residual and rebuild
once more. PADDING bytes inside the AEAD envelope add 1:1 to the wire
size and the Length varint grows monotonically, so the loop terminates
in ≤ 2 iterations for any reachable payload.
Same fix is applied to buildClosingDatagram so close-only Initial probes
on the boundary aren't tripped by future varint-growth changes.
Tightens the existing PTO-probe regression test to assert ≥ 1200 (was
relaxed to ≥ 1199 in 86b6c609a) and adds a new boundary test that builds
a single-byte-payload Initial and checks 1200 ≤ size ≤ 1203 — strict
floor with a tight ceiling so over-correction would also fail.
https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
The base image's /setup.sh configures routing for the runner's ns-3 sim
*and* points the container's DNS resolver at the sim's nameserver. When
we source it inside `make smoke` (which uses a plain Docker bridge
network without the sim), DNS resolution for the bridge container names
breaks: the JVM client can't resolve `amethyst-quic-interop-smoke-picoquic`
and aborts before any QUIC traffic.
Adds a SMOKE_MODE=1 env var; run_endpoint.sh skips /setup.sh entirely
when it's set. The Makefile smoke target sets it. Inside the runner
(unset, default), /setup.sh runs as before.
https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
Split the previously global `AudioFormat.CHANNELS = 1` into a
`DEFAULT_CHANNELS` constant + per-call-site `channelCount` parameters
so a single broadcast can advertise stereo Opus without forcing every
mono call site to grow a new argument. Generalises the catalog factory
to `MoqLiteHangCatalog.opus48k(name, channels)` with memoised JSON
bytes per shape, threads a new `AudioBroadcastConfig(channelCount)`
through `connectNestsSpeaker` / `connectReconnectingNestsSpeaker` /
`MoqLiteNestsSpeaker`, and adds a `channelCount` parameter to
`MediaCodecOpusEncoder`. Production behaviour is unchanged for
mono callers (the new config defaults to mono); the listener side
already discovers the channel count from the catalog via
`NestViewModel.awaitAudioPipelineConfig`. No test or wire changes.
Phase 1 of `nestsClient/plans/2026-05-06-i4-stereo-cross-stack-scenario.md`.
The hang-interop test scaffolding (`HangInteropTest`, `runSpeakerToHangListen`,
Rust `hang-listen` / `hang-publish`, `JvmOpusEncoder`) doesn't exist on
this branch yet, so the I4 forward + reverse scenarios are deferred
until the parent T16 plan lands.
https://claude.ai/code/session_01EqJEADzH9yjSuoP5L9js8i
The aioquic interop run revealed bug #3 (after the close-padding and
close-frame-type fixes): when PTO fires before the handshake completes,
the driver sets `pendingPing = true` but the writer only consumed that
flag in the 1-RTT path. Pre-handshake the flag was silently discarded,
so the second drain produced no Initial datagram — the connection sat
mute through every subsequent PTO. Symptom on the wire: exactly one
Initial packet (the close at PN=1, after our internal handshake
timeout), zero retransmits across the full 10-second budget, no chance
for the peer to recover from a dropped first ClientHello.
Fix routes pendingPing through to whatever encryption level is the
highest currently active — preferring 1-RTT, falling through Handshake,
finally Initial. Adds a regression test that drains a fresh connection
with `pendingPing = true` and verifies an Initial-level padded probe
datagram comes out (vs. null pre-fix).
Test relaxes the size assertion to ≥ 1199 due to a separate pre-existing
off-by-one in the writer's padding deficit calculation when the natural
payload uses a 1-byte Length varint that grows to 2 after padding —
that's a follow-up; the regression we care about here is "no probe at
all," not the byte-precise padding edge.
Outstanding from this run:
- Strict ≥ 1200 padding for tiny payloads (PING-only Initial = 1199)
- PTO should retransmit unacked CRYPTO bytes, not just emit a PING
(current PING gets ACK + relies on packet-number-threshold loss
detection to trigger CRYPTO retransmit; works but suboptimal)
https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
The quic-interop-runner against aioquic surfaced two real bugs in the
writer's CLOSING-status branch, both visible on the wire as a single
~45-byte UDP datagram instead of a properly framed close.
§10.2.3 — at Initial / Handshake levels only CONNECTION_CLOSE (Transport,
0x1c) is allowed. The application-level form (0x1d) leaks app state
pre-handshake. The writer was unconditionally building a 0x1d frame and
shipping it inside an Initial packet; aioquic dropped it silently.
§14.1 — any client datagram containing an Initial MUST be ≥ 1200 bytes
in UDP-payload terms. The CLOSING branch bypassed the existing padding
logic, so close-only Initial datagrams went out at ~45 bytes and
servers correctly rejected them (also as malformed).
Fix replaces the CLOSING branch with a dedicated `buildClosingDatagram`
helper:
- Application keys present → 0x1d, original error code + reason.
- Pre-1-RTT (Handshake or Initial) → 0x1c with errorCode =
APPLICATION_ERROR (0x0c), frameType=0, empty reason.
- Initial level: build at natural size, rewind PN if < 1200, rebuild
with PADDING-frame deficit inside the AEAD envelope.
Plus a regression test covering both: pre-handshake close datagram size
≥ 1200, and ConnectionCloseFrame round-trips 0x1c vs 0x1d for the right
constructor inputs.
NOT addressed yet: why our ClientHello at PN=0 doesn't appear in the
runner pcap (only PN=1 close does). With this fix the close packet is
now well-formed; the next runner run will tell us whether the missing
ClientHello is a sim/capture artifact or a separate writer bug.
https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT
NestViewModel.connect() launches an infinite cliff-detector loop in
viewModelScope (`while(true) { delay(...) }`). The existing tests in
NestViewModelTest call connect() but never disconnect/onCleared, so
that loop stays alive. Under runTest's virtual scheduler each delay
returns instantly, the loop spins millions of iterations per second of
real time, and runTest never reaches the idle state — every test
wedges until the per-test deadline (60 s × 17 tests => :commons:jvmTest
hangs for ~17 minutes).
Wrap each test body with runVmTest, which runs the body inside a
try/finally that calls disconnect() on every VM created by
newViewModel before runTest tries to drain. teardown() cancels
cliffDetectorJob, the scheduler becomes idle, and the test returns.
A 10 s runTest timeout is the safety net — healthy tests now finish
in milliseconds, and tripping the timeout signals a new viewModelScope
coroutine that needs its own teardown call.
Verified: :commons:jvmTest --rerun-tasks now completes in 52 s
(409 tests, 0 failures); NestViewModelTest's 17 tests run in 0.083 s
total versus the previous indefinite hang.
https://claude.ai/code/session_01R9wUxnRJrr299W8TzRyFs7
Two coupled fixes that surface together when running the smoke target
outside the runner's privileged sim environment:
1. run_endpoint.sh now tolerates /setup.sh failures.
The base image's /setup.sh manipulates routes for the runner's ns-3
sim. Outside the runner (e.g., make smoke without --privileged), it
fails with "netlink error: Operation not permitted" — and our
set -euo pipefail bailed before launching the JVM. The setup is
genuinely not needed for smoke; tolerate the failure and continue.
2. make smoke uses a private Docker bridge instead of --network host.
--network host on Docker Desktop for Mac is *not* the Mac host's
network — it's the LinuxKit VM's, and 127.0.0.1 isn't routed back
to other Docker containers. A dedicated bridge with DNS-by-name
works identically on Mac and Linux. Bonus: smoke-down target
reliably tears down the env.
Inside the runner these changes are no-ops: the runner provides the
privileges /setup.sh needs, and uses its own compose network not
--network host.
https://claude.ai/code/session_01HcvfQq1ttPV9PkRoJb4nyT