Repository navigation
stream: keep a pending FIN schedulable after data is acknowledged - #2804
akasakariko wants to merge 1 commit into
Conversation
The FIN was never sent in two cases: - a lost STREAM frame carrying FIN whose data was acknowledged in another packet, since only lost zero-length FINs were requeued - data retransmitted before an already acknowledged tail, since the stream stopped being flushable once that data was sent Track whether the final size still needs to be sent in SendBuf: - set it when the final size becomes known, or when a FIN-bearing frame is lost before the FIN is acknowledged - clear it when a FIN is emitted or acknowledged, or when the stream is reset - treat a stream as flushable when only its FIN is left to send This replaces the zero-length FIN special cases, avoids resending a FIN that is still in flight, and no longer sends a STREAM frame after RESET_STREAM when a FIN-bearing frame is lost Fixes cloudflare#2800 Co-authored-by: Cursor <cursoragent@cursor.com>
|
@lixing-lx Thanks for offering to help validate! The fix is up in this PR — would you mind running it against your TrustTunnel test setup to confirm the stranded FIN case is resolved? |
|
Thanks @akasakariko. I validated Deterministic comparison (same dependency lock, BoringSSL 5.2.0, Rust 1.99.0): I appended the two original Issue #2800 reproducer functions unchanged in behavior to separate base/head checkouts. On base TrustTunnel downstream validation: I compiled both the client and the server against this exact PR's quiche production sources, checking the changed production files against the commit. Client/server Cargo target directories are separate. The server uses the source from TrustTunnel's dependency-upgrade PR #155, H2/H3 smoke tests pass, covering TCP/UDP, authentication and certificate rejection, half-close and concurrent byte-checked transfers. The sustained H3 run uses client BBRv2 with a 32-packet initial window and the server's unchanged default congestion settings, in a Linux arm64 Docker namespace with
Total: 8,817 content-and-EOF-verified transfers / 577,830,912 bytes, no missing EOF or transfer timeout. The same run passes 20 session recreations, cancellation of an old read during a blackhole, bounded blackhole dial failure and subsequent recovery. Kernel UDP receive-buffer/send-buffer error counters remain zero; the acceptance validator checks all phase durations/counts and recreation cases, not just the test harness exit code. The latency/loss relay uses approximately 80 ms ± 10 ms delay per direction and a 10 Mbps cap per direction. The fault conditions, source fingerprints, binary SHA256 values and per-run logs are retained locally. This confirms the standalone scheduling failures and the controlled TrustTunnel path on this head; it is not a claim of public-network or production-server acceptance. |
|
Thanks a lot @lixing-lx for the thorough validation, especially the downstream TrustTunnel run under loss and latency. Really appreciate it Maintainers, this is ready for review whenever you have a chance. Happy to address any feedback |
Summary
Testing
cargo +nightly-2026-09-15 fmt -- --checkcargo clippy --features=async,ffi,qlog,rpk --workspace -- -D warningscargo test --all-targets --features=async,ffi,qlog,rpk --workspacecargo test --doc --features=async,ffi,qlog,rpk --workspaceFixes #2800
Made with Cursor