git: 6563dcb6b1f5 - main - unix: allow connectat(2) to name the peer socket by descriptor

Mark Johnston <[email protected]>
Newsgroups gmane.os.freebsd.devel.cvs.src
Message-ID <6a7a0d14.40711.1eee9a74__21100.9468462848$1786383806$gmane$org@gitrepo.freebsd.org>
The branch main has been updated by markj:

URL: https://cgit.FreeBSD.org/src/commit/?id=6563dcb6b1f57e51db63854f3774b52e672232ed

commit 6563dcb6b1f57e51db63854f3774b52e672232ed
Author:     John Ericson <[email protected]>
AuthorDate: 2026-08-10 15:04:29 +0000
Commit:     Mark Johnston <[email protected]>
CommitDate: 2026-08-10 17:31:22 +0000

    unix: allow connectat(2) to name the peer socket by descriptor
    
    Accept an empty `sun_path` when `fd` is not `AT_FDCWD`: the descriptor
    then names the peer unix socket directly, instead of being the starting
    directory for a pathname lookup.  The held file reference keeps the peer
    PCB stable, playing the role `unp_vp_mtxpool` plays in the pathname
    path.
    
    The descriptor must carry `CAP_CONNECTAT` and refer to an `AF_UNIX`
    socket (`EPROTOTYPE` otherwise, `ENOTSOCK` for non-sockets).  As with a
    pathname, a stream/seqpacket peer must be listening.  No filesystem
    permission or MAC vnode check applies on this path: possession of the
    descriptor is the authorization, as with descriptor passing.
    
    Note this makes it possible to connect a datagram socket to an unbound
    peer, which no pathname could previously name.
    
    `connect(2)` and the implicit-connect send path pass `AT_FDCWD` and
    still reject an empty path with `EINVAL`.
    
    The `unp_sun_path()` call is hoisted out of `unp_connectat()` because the
    early exit conditions for the two system calls (`connect(2)` and
    `connectat(2)`) are slightly different.
    
    Additionally, support `/dev/fd/<N>`. In a world with `connectat(2)`,
    this is largely overkill, but this also allows me to add support for
    direct peer connections with plain `connect(2)`. I think that is a wise
    choice because this will allow me to propose this functionality for
    Linux too without a new system call (saving that conversation for
    later). Ultimately, I want to see multiple operating systems support
    this to foster broader userland adoption, which should benefit everyone
    including FreeBSD --- it's nicer if more 3rd party in addition to 1st
    party software uses the new kernel functionality. Therefore, I hope this
    additional feature is also acceptable.
    
    Signed-off-by: John Ericson <[email protected]>
    Assisted-by: Claude Code (Claude Opus 4.8 and Fable 5)
    
    Reviewed by:    markj
    MFC after:      2 months
    Differential Revision:  https://reviews.freebsd.org/D58405
---
 lib/libsys/connectat.2 |  52 ++++++++++++++-
 share/man/man4/unix.4  |  94 ++++++++++++++++++++++++++-
 sys/kern/uipc_usrreq.c | 172 +++++++++++++++++++++++++++++++++++++++----------
 3 files changed, 281 insertions(+), 37 deletions(-)

diff --git a/lib/libsys/connectat.2 b/lib/libsys/connectat.2
index 64fd805549b7..f40d2d24c370 100644
--- a/lib/libsys/connectat.2
+++ b/lib/libsys/connectat.2
@@ -24,7 +24,7 @@
 .\" OUT OF THE USE OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF
 .\" SUCH DAMAGE.
 .\"
-.Dd February 13, 2013
+.Dd July 25, 2026
 .Dt CONNECTAT 2
 .Os
 .Sh NAME
@@ -66,6 +66,48 @@ If the file path stored in the
 field of the sockaddr_un structure is a relative path, it is located relative
 to the directory associated with the file descriptor
 .Fa fd .
+.Pp
+.It
+If the
+.Fa sun_path
+field is empty, that is,
+.Fa namelen
+equals
+.Li offsetof(struct sockaddr_un, sun_path) ,
+then
+.Fa fd
+does not resolve a path but instead names the peer socket directly.
+.Pp
+The descriptor may be the peer socket itself, or a descriptor for a file, such
+as an
+.Dv O_PATH
+handle opened with
+.Xr open 2 .
+In the latter case, the
+.Dv O_PATH
+file description may point either to a traditional socket file bound to the
+peer, or to a
+.Pa /dev/fd/N
+node.
+The
+.Pa /dev/fd/N
+node must name the socket descriptor itself; the named descriptor is resolved a
+single level and not chased further.
+With a standard
+.Xr fdescfs 5
+mount this means a node naming a traditional socket file, or another
+.Pa /dev/fd/N
+node, does not resolve to a peer.
+.Pp
+In all cases,
+.Fa fd
+must carry the
+.Dv CAP_CONNECTAT
+capability right, and the target of a
+.Pa /dev/fd/N
+node must also carry that right; see
+.Xr unix 4
+for the details of this form.
 .El
 .Sh RETURN VALUES
 .Rv -std connectat
@@ -92,6 +134,14 @@ field is not an absolute path and
 is neither
 .Dv AT_FDCWD
 nor a file descriptor associated with a directory.
+.It Bq Er EINVAL
+The
+.Fa sun_path
+field is empty but
+.Fa fd
+is
+.Dv AT_FDCWD ,
+so no descriptor names the peer.
 .El
 .Sh SEE ALSO
 .Xr bindat 2 ,
diff --git a/share/man/man4/unix.4 b/share/man/man4/unix.4
index f83f9ddffea3..497b26ff49b5 100644
--- a/share/man/man4/unix.4
+++ b/share/man/man4/unix.4
@@ -25,7 +25,7 @@
 .\" OUT OF THE USE OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF
 .\" SUCH DAMAGE.
 .\"
-.Dd June 3, 2026
+.Dd July 25, 2026
 .Dt UNIX 4
 .Os
 .Sh NAME
@@ -125,6 +125,95 @@ of a
 or
 .Xr sendto 2
 must be writable.
+.Ss Naming a peer
+A peer can be named either by a pathname or, using
+.Xr connectat 2 ,
+directly by a descriptor.
+Across these two forms a peer can be named five ways, all of which reach the
+same socket:
+.Bl -column "A bound socket's file" "an O_PATH descriptor" "an ordinary pathname" -offset indent
+.It Em Target Ta Em "By descriptor" Ta Em "By pathname"
+.It "The peer socket" Ta "the socket's fd" Ta "(none)"
+.It "A bound socket's file" Ta "an O_PATH descriptor" Ta "an ordinary pathname"
+.It "A /dev/fd/N node" Ta "an O_PATH descriptor" Ta Pa /dev/fd/N
+.El
+.Pp
+A pathname, accepted by both
+.Xr connect 2
+and
+.Xr connectat 2 ,
+names either a bound socket's file created by
+.Xr bind 2 ,
+or the
+.Pa /dev/fd/N
+node that
+.Xr fdescfs 5
+provides for descriptor
+.Ar N .
+In the latter case the
+.Pa /dev/fd/N
+node must name the socket descriptor itself; the named descriptor is resolved a
+single level and not chased further.
+With a standard
+.Xr fdescfs 5
+mount this means a node naming a traditional socket file, or another
+.Pa /dev/fd/N
+node, does not resolve to a peer; a
+.Cm nodup
+mount, however, dereferences such a descriptor before the lookup completes.
+.Pp
+.Xr connectat 2
+also supports both of the above cases.
+The path is allowed to be empty
+\(em that is, a
+.Fa namelen
+equal to
+.Li offsetof(struct sockaddr_un, sun_path)
+\(em
+in which case the given file descriptor will be used directly, like with most
+.Sy *at
+system calls.
+(Note:
+.Pa /dev/fd/N
+nodes are not recommended to be used for the
+.Xr connectat 2 ,
+since the descriptor can just be named directly, and this is both more
+efficient and less userland code, but since
+.Xr connect 2
+is implemented in terms of
+.Xr connectat 2
+this is provided for free.)
+.Pp
+Besides supporting cases that
+.Xr connect 2
+also supports,
+.Xr connectat 2
+additionally supports directly naming the peer by the
+.Fa fd
+descriptor (and an empty path).
+The descriptor may be the peer socket itself.
+In all cases, the descriptor passed must carry the
+.Dv CAP_CONNECTAT
+capability right, and in the
+.Pa /dev/fd/N
+form the target descriptor must
+.Em also
+carry that capability.
+A descriptor limited to
+.Dv CAP_CONNECTAT
+alone is thus a pure
+.Dq connect-to-me
+token, usable as a connection target but not listened on, accepted from, or
+read.
+.Pp
+For the descriptor forms, possession of the descriptor is the authorization,
+as with descriptor passing over
+.Dv SCM_RIGHTS ;
+the file system access-control and
+.Xr mac 4
+checks that apply to the pathname forms are not repeated.
+Because a descriptor alone suffices, a datagram socket may connect to an
+unbound peer, which no pathname could name.
 .Sh CONTROL MESSAGES
 The
 .Ux Ns -domain
@@ -460,6 +549,7 @@ chronological order they were sent.
 The order is preserved for writes coming through a particular connection.
 .Sh SEE ALSO
 .Xr connect 2 ,
+.Xr connectat 2 ,
 .Xr dup 2 ,
 .Xr fchmod 2 ,
 .Xr fcntl 2 ,
@@ -471,6 +561,8 @@ The order is preserved for writes coming through a particular connection.
 .Xr socket 2 ,
 .Xr CMSG_DATA 3 ,
 .Xr intro 4 ,
+.Xr mac 4 ,
+.Xr fdescfs 5 ,
 .Xr sysctl 8
 .Rs
 .%T "An Introductory 4.3 BSD Interprocess Communication Tutorial"
diff --git a/sys/kern/uipc_usrreq.c b/sys/kern/uipc_usrreq.c
index 60b0f3b8706e..93add7494644 100644
--- a/sys/kern/uipc_usrreq.c
+++ b/sys/kern/uipc_usrreq.c
@@ -290,9 +290,7 @@ static struct mtx	unp_defers_lock;
 
 static int	uipc_connect2(struct socket *, struct socket *);
 static int	uipc_ctloutput(struct socket *, struct sockopt *);
-static int	unp_connect(struct socket *, struct sockaddr *,
-		    struct thread *);
-static int	unp_connectat(int, struct socket *, struct sockaddr *,
+static int	unp_connectat(int, struct socket *, const char *, int,
 		    struct thread *, struct socket **);
 static int	unp_connect_peer(struct socket *, struct unpcb *,
 		    struct sockaddr **, struct thread *, bool);
@@ -728,22 +726,38 @@ uipc_bind(struct socket *so, struct sockaddr *nam, struct thread *td)
 static int
 uipc_connect(struct socket *so, struct sockaddr *nam, struct thread *td)
 {
-	int error;
+	const char *path;
+	int error, len;
 
 	KASSERT(td == curthread, ("uipc_connect: td != curthread"));
-	error = unp_connect(so, nam, td);
-	return (error);
+
+	error = unp_sun_path(nam, &path, &len);
+	if (error != 0)
+		return (error);
+	/*
+	 * unp_connectat() does not early exit on empty paths, because that is
+	 * explicitly supported when naming the peer by file descriptor, but
+	 * connect(2) only ever passes AT_FDCWD, so reject it here. This
+	 * preserves historical behavior.
+	 */
+	if (len == 0)
+		return (EINVAL);
+	return (unp_connectat(AT_FDCWD, so, path, len, td, NULL));
 }
 
 static int
 uipc_connectat(int fd, struct socket *so, struct sockaddr *nam,
     struct thread *td)
 {
-	int error;
+	const char *path;
+	int error, len;
 
 	KASSERT(td == curthread, ("uipc_connectat: td != curthread"));
-	error = unp_connectat(fd, so, nam, td, NULL);
-	return (error);
+
+	error = unp_sun_path(nam, &path, &len);
+	if (error != 0)
+		return (error);
+	return (unp_connectat(fd, so, path, len, td, NULL));
 }
 
 static void
@@ -2086,7 +2100,12 @@ uipc_sosend_dgram(struct socket *so, struct sockaddr *addr, struct uio *uio,
 	SOCK_SENDBUF_UNLOCK(so);
 
 	if (addr != NULL) {
-		if ((error = unp_connectat(AT_FDCWD, so, addr, td, &peer)))
+		const char *path;
+		int len;
+
+		if ((error = unp_sun_path(addr, &path, &len)))
+			goto out3;
+		if ((error = unp_connectat(AT_FDCWD, so, path, len, td, &peer)))
 			goto out3;
 		UNP_PCB_LOCK_ASSERT(unp);
 		unp2 = unp->unp_conn;
@@ -2905,16 +2924,10 @@ uipc_ctloutput(struct socket *so, struct sockopt *sopt)
 	return (error);
 }
 
-static int
-unp_connect(struct socket *so, struct sockaddr *nam, struct thread *td)
-{
-
-	return (unp_connectat(AT_FDCWD, so, nam, td, NULL));
-}
-
 /*
- * Connect socket 'so' to the unix-domain peer named by 'nam', resolved
- * relative to descriptor 'fd' (AT_FDCWD for connect(2)).
+ * Connect socket 'so' to the unix-domain peer named by the 'len'-byte 'path'
+ * (an empty path names the peer directly by descriptor), resolved relative to
+ * descriptor 'fd' (AT_FDCWD for connect(2)).
  *
  * 'referenced_peerp' selects how the peer is returned.  If NULL, on exit the
  * peer's PCB is unlocked and the peer is unreferenced, symmetrically releasing
@@ -2931,24 +2944,18 @@ unp_connect(struct socket *so, struct sockaddr *nam, struct thread *td)
  * path, which enqueues under the peer's PCB lock.
  */
 static int
-unp_connectat(int fd, struct socket *so, struct sockaddr *nam,
+unp_connectat(int fd, struct socket *so, const char *path, int len,
     struct thread *td, struct socket **referenced_peerp)
 {
 	struct socket *so2;
 	struct unpcb *unp;
 	char buf[SOCK_MAXADDRLEN];
 	struct sockaddr *sa;
-	const char *path;
-	int error, len;
+	int error;
 	bool connreq;
 
 	CURVNET_ASSERT_SET();
 
-	error = unp_sun_path(nam, &path, &len);
-	if (error != 0)
-		return (error);
-	if (len == 0)
-		return (EINVAL);
 	bcopy(path, buf, len);
 	buf[len] = 0;
 
@@ -2994,6 +3001,7 @@ unp_connectat(int fd, struct socket *so, struct sockaddr *nam,
 		sa = malloc(sizeof(struct sockaddr_un), M_SONAME, M_WAITOK);
 	else
 		sa = NULL;
+
 	error = unp_connectat_peer(td, fd, buf, &so2);
 	if (error != 0)
 		goto out;
@@ -3017,9 +3025,77 @@ out:
 }
 
 /*
- * Resolve a connectat(2) target -- descriptor 'fd' and the pathname in 'buf' --
- * to a referenced peer unix socket in '*so2p'.  The caller must release it with
- * sorele().
+ * Resolve descriptor 'fd' to the referenced unix-domain socket it *is* (as
+ * opposed to one it names through the file system) in '*so2p'.  Returns
+ * ENOTSOCK if 'fd' is not a socket -- letting an empty-path caller fall back to
+ * a vnode lookup -- or EPROTOTYPE if it is a socket of another domain.  The
+ * caller must release the returned socket with sorele().
+ */
+static int
+unp_socket_fd_peer(struct thread *td, int fd, struct socket **so2p)
+{
+	struct socket *so2;
+	struct file *fp;
+	cap_rights_t rights;
+	int error;
+
+	error = getsock(td, fd, cap_rights_init_one(&rights, CAP_CONNECTAT),
+	    &fp);
+	if (error != 0)
+		return (error);
+	so2 = fp->f_data;
+	if (so2->so_proto->pr_domain->dom_family != AF_UNIX)
+		error = EPROTOTYPE;
+	else {
+		soref(so2);
+		*so2p = so2;
+	}
+	fdrop(fp, td);
+	return (error);
+}
+
+/*
+ * Resolve a synthetic descriptor vnode -- as fdescfs fabricates for a /dev/fd/N
+ * path -- to the peer socket named by the descriptor it stands for.
+ *
+ * Such a node has no object of its own; VOP_OPEN reports the underlying
+ * descriptor in td_dupfd and fails with ENODEV, the same convention open(2)
+ * follows via dupfdopen() for /dev/fd.  We honour it here and resolve that
+ * descriptor as the peer, so a plain connect(2) to /dev/fd/N reaches the
+ * socket.  Does not consume 'vp'.
+ */
+static int
+unp_dupfd_peer(struct vnode *vp, struct thread *td, struct socket **so2p)
+{
+	int dupfd, error;
+
+	ASSERT_VOP_LOCKED(vp, __func__);
+
+	td->td_dupfd = -1;
+	error = VOP_OPEN(vp, FREAD, td->td_ucred, td, NULL);
+	dupfd = td->td_dupfd;
+	td->td_dupfd = 0;
+	if (error == ENODEV && dupfd >= 0)
+		return (unp_socket_fd_peer(td, dupfd, so2p));
+	if (error == 0) {
+		/* Not the dupfd convention: an openable node is not a peer. */
+		(void)VOP_CLOSE(vp, FREAD, td->td_ucred, td);
+		error = ECONNREFUSED;
+	}
+	return (error);
+}
+
+/*
+ * Resolve a connectat(2) target -- descriptor 'fd' together with the pathname
+ * in 'buf' (null when len == 0) -- to a referenced peer unix socket in
+ * '*so2p', covering all four ways a peer can be named:
+ *
+ *	empty path + socket fd		the descriptor is the peer socket
+ *	empty path + O_PATH vnode	EMPTYPATH resolves the socket's vnode
+ *	/dev/fd/N pathname		fdescfs names a descriptor
+ *	ordinary pathname		a bound socket looked up by path
+ *
+ * The caller must release the returned socket with sorele().
  */
 static int
 unp_connectat_peer(struct thread *td, int fd, const char *buf,
@@ -3029,6 +3105,19 @@ unp_connectat_peer(struct thread *td, int fd, const char *buf,
 	cap_rights_t rights;
 	int error;
 
+	/*
+	 * An empty sun_path means 'fd' names the peer directly.  If it is a
+	 * socket, it is the peer, so return success (or its error) with no
+	 * fallback; if not, it may be an O_PATH handle for a bound socket's
+	 * vnode, so fall through to an EMPTYPATH lookup.
+	 */
+	if (*buf == '\0') {
+		error = unp_socket_fd_peer(td, fd, so2p);
+		if (error != ENOTSOCK)
+			return (error);
+	}
+
+	/* Resolve the path to a vnode. */
 	NDINIT_ATRIGHTS(&nd, LOOKUP, FOLLOW | LOCKSHARED | LOCKLEAF |
 	    (fd == AT_FDCWD ? 0 : EMPTYPATH), UIO_SYSSPACE, buf, fd,
 	    cap_rights_init_one(&rights, CAP_CONNECTAT));
@@ -3038,12 +3127,25 @@ unp_connectat_peer(struct thread *td, int fd, const char *buf,
 	NDFREE_PNBUF(&nd);
 
 	/*
-	 * Resolve the vnode to a referenced peer socket and drop the vnode
-	 * before connecting: the reference keeps the peer stable, so no vnode
-	 * lock is held across unp_connect_peer() (which matters for the
-	 * return_locked datagram fast path).
+	 * Dispatch on the resolved vnode, then drop it: for the socket cases the
+	 * returned reference keeps the peer stable, so the caller holds no vnode
+	 * lock across unp_connect_peer() (which matters for the return_locked
+	 * datagram fast path).
+	 *
+	 * A synthetic descriptor node -- as fdescfs fabricates for a /dev/fd/N
+	 * path -- carries no type of its own (VNON); opening it yields the
+	 * descriptor it stands for, which we resolve as the peer socket.
+	 * Otherwise the path must name a bound socket's vnode (VSOCK), which
+	 * unp_vnode_peer() connects to, rejecting any other type with ENOTSOCK.
+	 *
+	 * unp_dupfd_peer() resolves that descriptor exactly once: if it is not
+	 * a socket the connect fails, so this does *not* recur through a chain
+	 * of O_PATH handles of /dev/fd nodes, which could be arbitrarily long.
 	 */
-	error = unp_vnode_peer(nd.ni_vp, td, so2p);
+	if (nd.ni_vp->v_type == VNON)
+		error = unp_dupfd_peer(nd.ni_vp, td, so2p);
+	else
+		error = unp_vnode_peer(nd.ni_vp, td, so2p);
 	vput(nd.ni_vp);
 	return (error);
 }
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.