[PATCH] tcp: Don't fast re-transmit if only our FIN is outstanding
From: Aris Konstantoulas
On Mon, 28 Sep 2026 13:31:20 +0300
Aris Konstantoulas
[...]
Don't consider a duplicate ACK as a fast re-transmit trigger if the only outstanding sequence number is the FIN, and leave it to the timer. Fast re-transmit of data is unaffected, with or without a FIN queued after it.
Thanks for the investigation and for the patch, I'll review it in a bit! Just a couple of quick (hopefully helpful) remarks for now:
[...]
- Open question: I haven't established how a real flow first reaches the state where the peer ACKs exactly up to, but not past, our FIN.
If I recall correctly, a while ago, I managed to make this happen with a Linux kernel as peer, under memory pressure, with the receiver updating the remaining buffer estimation right after processing a large segment (the last data segment) that's assembled from smaller segments. You would need a slow userspace receiver as well (maybe reading exactly one byte after accepting the connection and then nothing else... something like that). In any case, I don't think it really matters to find a "natural" reproducer of this case, I wouldn't go crazy with it, your patch looks obviously correct to me anyway.
[...] This patch breaks the loop in either case, but not that possible entry path. I'm happy to look into it further, possibly as part of bug 125.
Wow, yes, if you could also tackle bug #125, or even a part of it, that would be very warmly appreciated! -- Stefano
On Mon, 28 Sep 2026 13:31:20 +0300
Aris Konstantoulas
From: Aris Konstantoulas
In the TAP_FIN_RCVD path of tcp_tap_handler(), a bare segment from the guest acknowledging exactly seq_ack_from_tap, with an unchanged window, is taken as a duplicate ACK and triggers a fast re-transmit.
If the only unacknowledged sequence number is our own FIN, that's harmful: tcp_rewind_seq() rewinds seq_to_tap and clears TAP_FIN_SENT, so tcp_data_from_sock() immediately sends the FIN again. If the guest answers that FIN with the same bare ACK, as a socket in TIME-WAIT will, we loop at packet rate:
- conn->retries is never incremented on this path, so we never reach TCP_MAX_RETRIES and tcp_rst()
- ACK_FROM_TAP_DUE is re-armed on every iteration, so the backed-off re-transmission in tcp_timer_handler() never fires
- TAP_FIN_ACKED can't be set, as it requires TAP_FIN_SENT, which the rewind just cleared
On an idle Podman host (rootless, pasta), this showed up as a single flow exchanging ~45,000 54-byte segments per second between pasta and a container whose socket was in TIME-WAIT, with pasta using ~75% of one core, until the socket was killed by hand. It recurred on the idle teardown of an HTTP/2 connection to an ACME server.
With a raw-socket peer driving the same sequence against pasta at f8df3f1, pasta re-sent the FIN 727,509 times in 10 seconds. With this change it's re-transmitted by the timer at 1, 3, 7, 15, 31, 63 and 127 seconds, and the connection is reset once TCP_MAX_RETRIES is reached.
Don't consider a duplicate ACK as a fast re-transmit trigger if the only outstanding sequence number is the FIN, and leave it to the timer. Fast re-transmit of data is unaffected, with or without a FIN queued after it.
Fixes: bde1847960cf ("tcp: Fast re-transmit if half-closed, make TAP_FIN_RCVD path consistent") Link: https://bugs.passt.top/show_bug.cgi?id=125 Assisted-by: Claude:claude-opus-5-5 Signed-off-by: Aris Konstantoulas
---
Applied, thanks for fixing this, and welcome to the git log! -- Stefano
participants (2)
-
Aris Konstantoulas
-
Stefano Brivio