[PATCH net-next v10 7/7] tls: Preserve sk_err across recvmsg() when data has been copied

Chuck Lever <[email protected]> Mon, 11 May 2026 19:25:58 -0400
Newsgroups dev.linux.lists.kernel-tls-handshake,org.kernel.vger.netdev
Message-ID <[email protected]>
From: Chuck Lever <[email protected]>

Both sk_err checks in tls_rx_rec_wait() consume the error via
sock_error(), which clears sk_err atomically. When the caller
(tls_sw_recvmsg, tls_sw_splice_read, or tls_sw_read_sock) already
has bytes copied to userspace, it returns those bytes and discards
the error from this call. sk_err is now zero on the socket, so the
next read syscall observes only RCV_SHUTDOWN and reports a clean
EOF instead of the actual error (typically -ECONNRESET).

The race was reachable before this series via tls_read_flush_backlog()
when its periodic sk_flush_backlog() triggered tcp_reset() in the
middle of a multi-record read. The earlier patch in this series that
flushes the backlog inside tls_rx_rec_wait() widens the window: the
flush now runs on every iteration of every wait, not only when the
periodic threshold fires.

Have tls_rx_rec_wait() report sk_err without clearing it, using
READ_ONCE() to keep the read explicit. Each caller's return path
consumes sk_err only when no data is being returned and the err
about to surface matches the pending sk_err. This mirrors the
tcp_recvmsg() preserve-and-surface pattern, and also handles
tls_rx_one_record()'s decrypt-abort path: it raises sk_err to
EBADMSG via tls_err_abort() before returning a different errno
(-EFAULT from tls_setup_from_iter() on zero-copy receive, -ENOMEM
from decrypt setup). The gate keeps the actual error on this read
and lets the EBADMSG surface on the next, matching pre-series
behavior.

Fixes: c46b01839f7a ("tls: rx: periodically flush socket backlog")
Signed-off-by: Chuck Lever <[email protected]>
---
 net/tls/tls_sw.c | 34 +++++++++++++++++++++++++++++-----
 1 file changed, 29 insertions(+), 5 deletions(-)

diff --git a/net/tls/tls_sw.c b/net/tls/tls_sw.c
index 2b7093d27eb6..6a0ac2ccde56 100644
--- a/net/tls/tls_sw.c
+++ b/net/tls/tls_sw.c
@@ -1376,8 +1376,14 @@ tls_rx_rec_wait(struct sock *sk, struct sk_psock *psock, bool nonblock,
 		if (!sk_psock_queue_empty(psock))
 			return 0;
 
+		/* Report sk_err without clearing it. The caller may
+		 * discard the error return from this function in favor
+		 * of bytes already copied; leaving sk_err set ensures
+		 * the next read syscall surfaces the error instead of
+		 * a spurious EOF.
+		 */
 		if (sk->sk_err)
-			return sock_error(sk);
+			return -READ_ONCE(sk->sk_err);
 
 		if (ret < 0)
 			return ret;
@@ -1399,7 +1405,7 @@ tls_rx_rec_wait(struct sock *sk, struct sk_psock *psock, bool nonblock,
 		 * actual error rather than a clean EOF.
 		 */
 		if (sk->sk_err)
-			return sock_error(sk);
+			return -READ_ONCE(sk->sk_err);
 		if (sk->sk_shutdown & RCV_SHUTDOWN)
 			return 0;
 
@@ -1430,6 +1436,18 @@ tls_rx_rec_wait(struct sock *sk, struct sk_psock *psock, bool nonblock,
 	return 1;
 }
 
+/* Clear sk_err only when it matches the err about to be returned.
+ * tls_rx_one_record() can raise sk_err to EBADMSG via tls_err_abort()
+ * while returning a different errno; preserving sk_err in that case
+ * lets the EBADMSG surface on the next read.
+ */
+static int tls_sw_consume_matching_sk_err(struct sock *sk, int err)
+{
+	if (err < 0 && -err == READ_ONCE(sk->sk_err))
+		return sock_error(sk);
+	return err;
+}
+
 static int tls_setup_from_iter(struct iov_iter *from,
 			       int length, int *pages_used,
 			       struct scatterlist *to,
@@ -2285,7 +2303,9 @@ int tls_sw_recvmsg(struct sock *sk,
 	tls_rx_reader_unlock(sk, ctx);
 	if (psock)
 		sk_psock_put(sk, psock);
-	return copied ? : err;
+	if (copied)
+		return copied;
+	return tls_sw_consume_matching_sk_err(sk, err);
 }
 
 ssize_t tls_sw_splice_read(struct socket *sock,  loff_t *ppos,
@@ -2350,7 +2370,9 @@ ssize_t tls_sw_splice_read(struct socket *sock,  loff_t *ppos,
 
 splice_read_end:
 	tls_rx_reader_unlock(sk, ctx);
-	return copied ? : err;
+	if (copied)
+		return copied;
+	return tls_sw_consume_matching_sk_err(sk, err);
 
 splice_requeue:
 	__skb_queue_head(&ctx->rx_list, skb);
@@ -2444,7 +2466,9 @@ int tls_sw_read_sock(struct sock *sk, read_descriptor_t *desc,
 
 read_sock_end:
 	tls_rx_reader_release(sk, ctx);
-	return copied ? : err;
+	if (copied)
+		return copied;
+	return tls_sw_consume_matching_sk_err(sk, err);
 
 read_sock_requeue:
 	__skb_queue_head(&ctx->rx_list, skb);

-- 
2.54.0