Bernard, thanks for the investigation and for the patch. Just one doubt:
On Fri, 2 Oct 2026 18:39:25 +0200
"Bernhard M. Wiedemann"
without this patch, sending buffers of certain sizes caused a delay of 200ms because tcp_splice_forward() passes SPLICE_F_MORE to the writer when a read filled at least 90% of the pipe.
This is easy to hit when pipes are small: once a user exceeds fs.pipe-user-pages-soft (64 MiB by default, which a few long-running pasta instances with large pipes reach on their own), tcp_set_pipe_size() settles on the 8 KiB minimum, and every message whose length is 7372 to 8192 bytes (90% of the pipe and more) past a multiple of 8 KiB stalls.
This change leaves bulk throughput unchanged.
Assisted-by: Claude:claude-opus-5-5 Signed-off-by: Bernhard M. Wiedemann
--- Notes: The change was slightly tested and benchmarked. Results look decent. Not sure if we actually need the flow_trace for error handling there. This issue was accidentally found by running PostgreSQL in a rootless podman container for benchmarking. tcp_splice.c | 15 ++++++++++++++- 1 file changed, 14 insertions(+), 1 deletion(-)
diff --git a/tcp_splice.c b/tcp_splice.c index 4b01f1a..f7a3913 100644 --- a/tcp_splice.c +++ b/tcp_splice.c @@ -488,6 +488,7 @@ static int tcp_splice_forward(struct ctx *c, { uint8_t lowat_set_flag = RCVLOWAT_SET(fromsidei); uint8_t lowat_act_flag = RCVLOWAT_ACT(fromsidei); + bool corked = false;
while (1) { ssize_t readlen, written; @@ -517,8 +518,17 @@ static int tcp_splice_forward(struct ctx *c, * there's nothing in the pipe so there's nothing to do * write side either. */ - if (!conn->pending[fromsidei]) + if (!conn->pending[fromsidei]) { + /* Setting TCP_NODELAY again flushes data held + * back by SPLICE_F_MORE
...nice, I didn't know about that trick. But wouldn't it be more natural to not use SPLICE_F_MORE if the pipe is small enough? Do you have a stand-alone reproducer that could help figuring this out? I'm a bit worried we might cause unnecessary setsockopt() calls in some corner cases if we go this way, even though it's a rather minor concern (we just called splice(), and that setsockopt() is not _that_ expensive in comparison).
+ */ + if (corked && + setsockopt(conn->s[!fromsidei], SOL_TCP, + TCP_NODELAY, &((int){ 1 }), + sizeof(int))) + flow_trace(conn, "failed to push data"); break; + } } else { conn->pending[fromsidei] += readlen;
@@ -549,6 +559,9 @@ static int tcp_splice_forward(struct ctx *c, if (written < 0) break;
+ if (written > 0) + corked = more; + conn->pending[fromsidei] -= written;
if (!conn->pending[fromsidei] && readlen <= 0) {
-- Stefano