[PATCH] tcp_splice: avoid delay on certain transfers
without this patch, sending buffers of certain sizes
caused a delay of 200ms because tcp_splice_forward()
passes SPLICE_F_MORE to the writer
when a read filled at least 90% of the pipe.
This is easy to hit when pipes are small: once a user exceeds
fs.pipe-user-pages-soft (64 MiB by default, which a few long-running
pasta instances with large pipes reach on their own),
tcp_set_pipe_size() settles on the 8 KiB minimum, and every message
whose length is 7372 to 8192 bytes (90% of the pipe and more) past a
multiple of 8 KiB stalls.
This change leaves bulk throughput unchanged.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Bernhard M. Wiedemann
Bernard, thanks for the investigation and for the patch. Just one doubt:
On Fri, 2 Oct 2026 18:39:25 +0200
"Bernhard M. Wiedemann"
without this patch, sending buffers of certain sizes caused a delay of 200ms because tcp_splice_forward() passes SPLICE_F_MORE to the writer when a read filled at least 90% of the pipe.
This is easy to hit when pipes are small: once a user exceeds fs.pipe-user-pages-soft (64 MiB by default, which a few long-running pasta instances with large pipes reach on their own), tcp_set_pipe_size() settles on the 8 KiB minimum, and every message whose length is 7372 to 8192 bytes (90% of the pipe and more) past a multiple of 8 KiB stalls.
This change leaves bulk throughput unchanged.
Assisted-by: Claude:claude-opus-5-5 Signed-off-by: Bernhard M. Wiedemann
--- Notes: The change was slightly tested and benchmarked. Results look decent. Not sure if we actually need the flow_trace for error handling there. This issue was accidentally found by running PostgreSQL in a rootless podman container for benchmarking. tcp_splice.c | 15 ++++++++++++++- 1 file changed, 14 insertions(+), 1 deletion(-)
diff --git a/tcp_splice.c b/tcp_splice.c index 4b01f1a..f7a3913 100644 --- a/tcp_splice.c +++ b/tcp_splice.c @@ -488,6 +488,7 @@ static int tcp_splice_forward(struct ctx *c, { uint8_t lowat_set_flag = RCVLOWAT_SET(fromsidei); uint8_t lowat_act_flag = RCVLOWAT_ACT(fromsidei); + bool corked = false;
while (1) { ssize_t readlen, written; @@ -517,8 +518,17 @@ static int tcp_splice_forward(struct ctx *c, * there's nothing in the pipe so there's nothing to do * write side either. */ - if (!conn->pending[fromsidei]) + if (!conn->pending[fromsidei]) { + /* Setting TCP_NODELAY again flushes data held + * back by SPLICE_F_MORE
...nice, I didn't know about that trick. But wouldn't it be more natural to not use SPLICE_F_MORE if the pipe is small enough? Do you have a stand-alone reproducer that could help figuring this out? I'm a bit worried we might cause unnecessary setsockopt() calls in some corner cases if we go this way, even though it's a rather minor concern (we just called splice(), and that setsockopt() is not _that_ expensive in comparison).
+ */ + if (corked && + setsockopt(conn->s[!fromsidei], SOL_TCP, + TCP_NODELAY, &((int){ 1 }), + sizeof(int))) + flow_trace(conn, "failed to push data"); break; + } } else { conn->pending[fromsidei] += readlen;
@@ -549,6 +559,9 @@ static int tcp_splice_forward(struct ctx *c, if (written < 0) break;
+ if (written > 0) + corked = more; + conn->pending[fromsidei] -= written;
if (!conn->pending[fromsidei] && readlen <= 0) {
-- Stefano
On 05/10/2026 21.37, Stefano Brivio wrote:
Bernard, thanks for the investigation and for the patch. Just one doubt: [...] I'm a bit worried we might cause unnecessary setsockopt() calls in some corner cases if we go this way, even though it's a rather minor concern (we just called splice(), and that setsockopt() is not _that_ expensive in comparison). The 'corked' variable should ensure that we only add the setsockopt when it would otherwise add the unneccesary 200ms delay.
I attached the reproducers and benchmarking tools. The finding from running it was that SPLICE_F_MORE gives plenty extra performance even for 8K pipes and setsockopt hardly costs anything. To test: # maybe with adjustment to /proc/sys/fs/pipe-user-pages-soft pasta --config-net -t 127.0.0.1/25432:15432 -- python3 rr.py serve 15432 python3 rr.py 25432 ┌─────────┬─────────┬───────────────┬─────────────────┬────────────────┐ │ Build │ Pipes │ Worst latency │ Upload (Gbit/s) │ Download/Gbit │ ├─────────┼─────────┼───────────────┼─────────────────┼────────────────┤ │ base │ 8 KiB │ 208 ms │ 89–91 │ 87–89 │ ├─────────┼─────────┼───────────────┼─────────────────┼────────────────┤ │ patched │ 8 KiB │ 0.1 ms │ 88–89 │ 86–89 │ ├─────────┼─────────┼───────────────┼─────────────────┼────────────────┤ │ nomore │ 8 KiB │ 0.1 ms │ 42–43 │ 40 │ ├─────────┼─────────┼───────────────┼─────────────────┼────────────────┤ │ base │ default │ 0.3 ms │ 203–204 │ 189–198 │ ├─────────┼─────────┼───────────────┼─────────────────┼────────────────┤ │ patched │ default │ 0.1 ms │ 203–205 │ 185–190 │ ├─────────┼─────────┼───────────────┼─────────────────┼────────────────┤ │ nomore │ default │ 0.1 ms │ 193–203 │ 177–190 │ └─────────┴─────────┴───────────────┴─────────────────┴────────────────┘ Ciao Bernhard M.
participants (2)
-
Bernhard M. Wiedemann
-
Stefano Brivio