git: 6563dcb6b1f5 - main - unix: allow connectat(2) to name the peer socket by descriptor

From: Mark Johnston <markj_at_FreeBSD.org>
Date: Mon, 10 Aug 2026 17:40:36 UTC
The branch main has been updated by markj:

URL: https://cgit.FreeBSD.org/src/commit/?id=6563dcb6b1f57e51db63854f3774b52e672232ed

commit 6563dcb6b1f57e51db63854f3774b52e672232ed
Author:     John Ericson <John.Ericson@Obsidian.Systems>
AuthorDate: 2026-08-10 15:04:29 +0000
Commit:     Mark Johnston <markj@FreeBSD.org>
CommitDate: 2026-08-10 17:31:22 +0000

    unix: allow connectat(2) to name the peer socket by descriptor
    
    Accept an empty `sun_path` when `fd` is not `AT_FDCWD`: the descriptor
    then names the peer unix socket directly, instead of being the starting
    directory for a pathname lookup.  The held file reference keeps the peer
    PCB stable, playing the role `unp_vp_mtxpool` plays in the pathname
    path.
    
    The descriptor must carry `CAP_CONNECTAT` and refer to an `AF_UNIX`
    socket (`EPROTOTYPE` otherwise, `ENOTSOCK` for non-sockets).  As with a
    pathname, a stream/seqpacket peer must be listening.  No filesystem
    permission or MAC vnode check applies on this path: possession of the
    descriptor is the authorization, as with descriptor passing.
    
    Note this makes it possible to connect a datagram socket to an unbound
    peer, which no pathname could previously name.
    
    `connect(2)` and the implicit-connect send path pass `AT_FDCWD` and
    still reject an empty path with `EINVAL`.
    
    The `unp_sun_path()` call is hoisted out of `unp_connectat()` because the
    early exit conditions for the two system calls (`connect(2)` and
    `connectat(2)`) are slightly different.
    
    Additionally, support `/dev/fd/<N>`. In a world with `connectat(2)`,
    this is largely overkill, but this also allows me to add support for
    direct peer connections with plain `connect(2)`. I think that is a wise
    choice because this will allow me to propose this functionality for
    Linux too without a new system call (saving that conversation for
    later). Ultimately, I want to see multiple operating systems support
    this to foster broader userland adoption, which should benefit everyone
    including FreeBSD --- it's nicer if more 3rd party in addition to 1st
    party software uses the new kernel functionality. Therefore, I hope this
    additional feature is also acceptable.
    
    Signed-off-by: John Ericson <John.Ericson@Obsidian.Systems>
    Assisted-by: Claude Code (Claude Opus 4.8 and Fable 5)
    
    Reviewed by:    markj
    MFC after:      2 months
    Differential Revision:  https://reviews.freebsd.org/D58405
---
 lib/libsys/connectat.2 |  52 ++++++++++++++-
 share/man/man4/unix.4  |  94 ++++++++++++++++++++++++++-
 sys/kern/uipc_usrreq.c | 172 +++++++++++++++++++++++++++++++++++++++----------
 3 files changed, 281 insertions(+), 37 deletions(-)

diff --git a/lib/libsys/connectat.2 b/lib/libsys/connectat.2
index 64fd805549b7..f40d2d24c370 100644
--- a/lib/libsys/connectat.2
+++ b/lib/libsys/connectat.2
@@ -24,7 +24,7 @@
 .\" OUT OF THE USE OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF
 .\" SUCH DAMAGE.
 .\"
-.Dd February 13, 2013
+.Dd July 25, 2026
 .Dt CONNECTAT 2
 .Os
 .Sh NAME
@@ -66,6 +66,48 @@ If the file path stored in the
 field of the sockaddr_un structure is a relative path, it is located relative
 to the directory associated with the file descriptor
 .Fa fd .
+.Pp
+.It
+If the
+.Fa sun_path
+field is empty, that is,
+.Fa namelen
+equals
+.Li offsetof(struct sockaddr_un, sun_path) ,
+then
+.Fa fd
+does not resolve a path but instead names the peer socket directly.
+.Pp
+The descriptor may be the peer socket itself, or a descriptor for a file, such
+as an
+.Dv O_PATH
+handle opened with
+.Xr open 2 .
+In the latter case, the
+.Dv O_PATH
+file description may point either to a traditional socket file bound to the
+peer, or to a
+.Pa /dev/fd/N
+node.
+The
+.Pa /dev/fd/N
+node must name the socket descriptor itself; the named descriptor is resolved a
+single level and not chased further.
+With a standard
+.Xr fdescfs 5
+mount this means a node naming a traditional socket file, or another
+.Pa /dev/fd/N
+node, does not resolve to a peer.
+.Pp
+In all cases,
+.Fa fd
+must carry the
+.Dv CAP_CONNECTAT
+capability right, and the target of a
+.Pa /dev/fd/N
+node must also carry that right; see
+.Xr unix 4
+for the details of this form.
 .El
 .Sh RETURN VALUES
 .Rv -std connectat
@@ -92,6 +134,14 @@ field is not an absolute path and
 is neither
 .Dv AT_FDCWD
 nor a file descriptor associated with a directory.
+.It Bq Er EINVAL
+The
+.Fa sun_path
+field is empty but
+.Fa fd
+is
+.Dv AT_FDCWD ,
+so no descriptor names the peer.
 .El
 .Sh SEE ALSO
 .Xr bindat 2 ,
diff --git a/share/man/man4/unix.4 b/share/man/man4/unix.4
index f83f9ddffea3..497b26ff49b5 100644
--- a/share/man/man4/unix.4
+++ b/share/man/man4/unix.4
@@ -25,7 +25,7 @@
 .\" OUT OF THE USE OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF
 .\" SUCH DAMAGE.
 .\"
-.Dd June 3, 2026
+.Dd July 25, 2026
 .Dt UNIX 4
 .Os
 .Sh NAME
@@ -125,6 +125,95 @@ of a
 or
 .Xr sendto 2
 must be writable.
+.Ss Naming a peer
+A peer can be named either by a pathname or, using
+.Xr connectat 2 ,
+directly by a descriptor.
+Across these two forms a peer can be named five ways, all of which reach the
+same socket:
+.Bl -column "A bound socket's file" "an O_PATH descriptor" "an ordinary pathname" -offset indent
+.It Em Target Ta Em "By descriptor" Ta Em "By pathname"
+.It "The peer socket" Ta "the socket's fd" Ta "(none)"
+.It "A bound socket's file" Ta "an O_PATH descriptor" Ta "an ordinary pathname"
+.It "A /dev/fd/N node" Ta "an O_PATH descriptor" Ta Pa /dev/fd/N
+.El
+.Pp
+A pathname, accepted by both
+.Xr connect 2
+and
+.Xr connectat 2 ,
+names either a bound socket's file created by
+.Xr bind 2 ,
+or the
+.Pa /dev/fd/N
+node that
+.Xr fdescfs 5
+provides for descriptor
+.Ar N .
+In the latter case the
+.Pa /dev/fd/N
+node must name the socket descriptor itself; the named descriptor is resolved a
+single level and not chased further.
+With a standard
+.Xr fdescfs 5
+mount this means a node naming a traditional socket file, or another
+.Pa /dev/fd/N
+node, does not resolve to a peer; a
+.Cm nodup
+mount, however, dereferences such a descriptor before the lookup completes.
+.Pp
+.Xr connectat 2
+also supports both of the above cases.
+The path is allowed to be empty
+\(em that is, a
+.Fa namelen
+equal to
+.Li offsetof(struct sockaddr_un, sun_path)
+\(em
+in which case the given file descriptor will be used directly, like with most
+.Sy *at
+system calls.
+(Note:
+.Pa /dev/fd/N
+nodes are not recommended to be used for the
+.Xr connectat 2 ,
+since the descriptor can just be named directly, and this is both more
+efficient and less userland code, but since
+.Xr connect 2
+is implemented in terms of
+.Xr connectat 2
+this is provided for free.)
+.Pp
+Besides supporting cases that
+.Xr connect 2
+also supports,
+.Xr connectat 2
+additionally supports directly naming the peer by the
+.Fa fd
+descriptor (and an empty path).
+The descriptor may be the peer socket itself.
+In all cases, the descriptor passed must carry the
+.Dv CAP_CONNECTAT
+capability right, and in the
+.Pa /dev/fd/N
+form the target descriptor must
+.Em also
+carry that capability.
+A descriptor limited to
+.Dv CAP_CONNECTAT
+alone is thus a pure
+.Dq connect-to-me
+token, usable as a connection target but not listened on, accepted from, or
+read.
+.Pp
+For the descriptor forms, possession of the descriptor is the authorization,
+as with descriptor passing over
+.Dv SCM_RIGHTS ;
+the file system access-control and
+.Xr mac 4
+checks that apply to the pathname forms are not repeated.
+Because a descriptor alone suffices, a datagram socket may connect to an
+unbound peer, which no pathname could name.
 .Sh CONTROL MESSAGES
 The
 .Ux Ns -domain
@@ -460,6 +549,7 @@ chronological order they were sent.
 The order is preserved for writes coming through a particular connection.
 .Sh SEE ALSO
 .Xr connect 2 ,
+.Xr connectat 2 ,
 .Xr dup 2 ,
 .Xr fchmod 2 ,
 .Xr fcntl 2 ,
@@ -471,6 +561,8 @@ The order is preserved for writes coming through a particular connection.
 .Xr socket 2 ,
 .Xr CMSG_DATA 3 ,
 .Xr intro 4 ,
+.Xr mac 4 ,
+.Xr fdescfs 5 ,
 .Xr sysctl 8
 .Rs
 .%T "An Introductory 4.3 BSD Interprocess Communication Tutorial"
diff --git a/sys/kern/uipc_usrreq.c b/sys/kern/uipc_usrreq.c
index 60b0f3b8706e..93add7494644 100644
--- a/sys/kern/uipc_usrreq.c
+++ b/sys/kern/uipc_usrreq.c
@@ -290,9 +290,7 @@ static struct mtx	unp_defers_lock;
 
 static int	uipc_connect2(struct socket *, struct socket *);
 static int	uipc_ctloutput(struct socket *, struct sockopt *);
-static int	unp_connect(struct socket *, struct sockaddr *,
-		    struct thread *);
-static int	unp_connectat(int, struct socket *, struct sockaddr *,
+static int	unp_connectat(int, struct socket *, const char *, int,
 		    struct thread *, struct socket **);
 static int	unp_connect_peer(struct socket *, struct unpcb *,
 		    struct sockaddr **, struct thread *, bool);
@@ -728,22 +726,38 @@ uipc_bind(struct socket *so, struct sockaddr *nam, struct thread *td)
 static int
 uipc_connect(struct socket *so, struct sockaddr *nam, struct thread *td)
 {
-	int error;
+	const char *path;
+	int error, len;
 
 	KASSERT(td == curthread, ("uipc_connect: td != curthread"));
-	error = unp_connect(so, nam, td);
-	return (error);
+
+	error = unp_sun_path(nam, &path, &len);
+	if (error != 0)
+		return (error);
+	/*
+	 * unp_connectat() does not early exit on empty paths, because that is
+	 * explicitly supported when naming the peer by file descriptor, but
+	 * connect(2) only ever passes AT_FDCWD, so reject it here. This
+	 * preserves historical behavior.
+	 */
+	if (len == 0)
+		return (EINVAL);
+	return (unp_connectat(AT_FDCWD, so, path, len, td, NULL));
 }
 
 static int
 uipc_connectat(int fd, struct socket *so, struct sockaddr *nam,
     struct thread *td)
 {
-	int error;
+	const char *path;
+	int error, len;
 
 	KASSERT(td == curthread, ("uipc_connectat: td != curthread"));
-	error = unp_connectat(fd, so, nam, td, NULL);
-	return (error);
+
+	error = unp_sun_path(nam, &path, &len);
+	if (error != 0)
+		return (error);
+	return (unp_connectat(fd, so, path, len, td, NULL));
 }
 
 static void
@@ -2086,7 +2100,12 @@ uipc_sosend_dgram(struct socket *so, struct sockaddr *addr, struct uio *uio,
 	SOCK_SENDBUF_UNLOCK(so);
 
 	if (addr != NULL) {
-		if ((error = unp_connectat(AT_FDCWD, so, addr, td, &peer)))
+		const char *path;
+		int len;
+
+		if ((error = unp_sun_path(addr, &path, &len)))
+			goto out3;
+		if ((error = unp_connectat(AT_FDCWD, so, path, len, td, &peer)))
 			goto out3;
 		UNP_PCB_LOCK_ASSERT(unp);
 		unp2 = unp->unp_conn;
@@ -2905,16 +2924,10 @@ uipc_ctloutput(struct socket *so, struct sockopt *sopt)
 	return (error);
 }
 
-static int
-unp_connect(struct socket *so, struct sockaddr *nam, struct thread *td)
-{
-
-	return (unp_connectat(AT_FDCWD, so, nam, td, NULL));
-}
-
 /*
- * Connect socket 'so' to the unix-domain peer named by 'nam', resolved
- * relative to descriptor 'fd' (AT_FDCWD for connect(2)).
+ * Connect socket 'so' to the unix-domain peer named by the 'len'-byte 'path'
+ * (an empty path names the peer directly by descriptor), resolved relative to
+ * descriptor 'fd' (AT_FDCWD for connect(2)).
  *
  * 'referenced_peerp' selects how the peer is returned.  If NULL, on exit the
  * peer's PCB is unlocked and the peer is unreferenced, symmetrically releasing
@@ -2931,24 +2944,18 @@ unp_connect(struct socket *so, struct sockaddr *nam, struct thread *td)
  * path, which enqueues under the peer's PCB lock.
  */
 static int
-unp_connectat(int fd, struct socket *so, struct sockaddr *nam,
+unp_connectat(int fd, struct socket *so, const char *path, int len,
     struct thread *td, struct socket **referenced_peerp)
 {
 	struct socket *so2;
 	struct unpcb *unp;
 	char buf[SOCK_MAXADDRLEN];
 	struct sockaddr *sa;
-	const char *path;
-	int error, len;
+	int error;
 	bool connreq;
 
 	CURVNET_ASSERT_SET();
 
-	error = unp_sun_path(nam, &path, &len);
-	if (error != 0)
-		return (error);
-	if (len == 0)
-		return (EINVAL);
 	bcopy(path, buf, len);
 	buf[len] = 0;
 
@@ -2994,6 +3001,7 @@ unp_connectat(int fd, struct socket *so, struct sockaddr *nam,
 		sa = malloc(sizeof(struct sockaddr_un), M_SONAME, M_WAITOK);
 	else
 		sa = NULL;
+
 	error = unp_connectat_peer(td, fd, buf, &so2);
 	if (error != 0)
 		goto out;
@@ -3017,9 +3025,77 @@ out:
 }
 
 /*
- * Resolve a connectat(2) target -- descriptor 'fd' and the pathname in 'buf' --
- * to a referenced peer unix socket in '*so2p'.  The caller must release it with
- * sorele().
+ * Resolve descriptor 'fd' to the referenced unix-domain socket it *is* (as
+ * opposed to one it names through the file system) in '*so2p'.  Returns
+ * ENOTSOCK if 'fd' is not a socket -- letting an empty-path caller fall back to
+ * a vnode lookup -- or EPROTOTYPE if it is a socket of another domain.  The
+ * caller must release the returned socket with sorele().
+ */
+static int
+unp_socket_fd_peer(struct thread *td, int fd, struct socket **so2p)
+{
+	struct socket *so2;
+	struct file *fp;
+	cap_rights_t rights;
+	int error;
+
+	error = getsock(td, fd, cap_rights_init_one(&rights, CAP_CONNECTAT),
+	    &fp);
+	if (error != 0)
+		return (error);
+	so2 = fp->f_data;
+	if (so2->so_proto->pr_domain->dom_family != AF_UNIX)
+		error = EPROTOTYPE;
+	else {
+		soref(so2);
+		*so2p = so2;
+	}
+	fdrop(fp, td);
+	return (error);
+}
+
+/*
+ * Resolve a synthetic descriptor vnode -- as fdescfs fabricates for a /dev/fd/N
+ * path -- to the peer socket named by the descriptor it stands for.
+ *
+ * Such a node has no object of its own; VOP_OPEN reports the underlying
+ * descriptor in td_dupfd and fails with ENODEV, the same convention open(2)
+ * follows via dupfdopen() for /dev/fd.  We honour it here and resolve that
+ * descriptor as the peer, so a plain connect(2) to /dev/fd/N reaches the
+ * socket.  Does not consume 'vp'.
+ */
+static int
+unp_dupfd_peer(struct vnode *vp, struct thread *td, struct socket **so2p)
+{
+	int dupfd, error;
+
+	ASSERT_VOP_LOCKED(vp, __func__);
+
+	td->td_dupfd = -1;
+	error = VOP_OPEN(vp, FREAD, td->td_ucred, td, NULL);
+	dupfd = td->td_dupfd;
+	td->td_dupfd = 0;
+	if (error == ENODEV && dupfd >= 0)
+		return (unp_socket_fd_peer(td, dupfd, so2p));
+	if (error == 0) {
+		/* Not the dupfd convention: an openable node is not a peer. */
+		(void)VOP_CLOSE(vp, FREAD, td->td_ucred, td);
+		error = ECONNREFUSED;
+	}
+	return (error);
+}
+
+/*
+ * Resolve a connectat(2) target -- descriptor 'fd' together with the pathname
+ * in 'buf' (null when len == 0) -- to a referenced peer unix socket in
+ * '*so2p', covering all four ways a peer can be named:
+ *
+ *	empty path + socket fd		the descriptor is the peer socket
+ *	empty path + O_PATH vnode	EMPTYPATH resolves the socket's vnode
+ *	/dev/fd/N pathname		fdescfs names a descriptor
+ *	ordinary pathname		a bound socket looked up by path
+ *
+ * The caller must release the returned socket with sorele().
  */
 static int
 unp_connectat_peer(struct thread *td, int fd, const char *buf,
@@ -3029,6 +3105,19 @@ unp_connectat_peer(struct thread *td, int fd, const char *buf,
 	cap_rights_t rights;
 	int error;
 
+	/*
+	 * An empty sun_path means 'fd' names the peer directly.  If it is a
+	 * socket, it is the peer, so return success (or its error) with no
+	 * fallback; if not, it may be an O_PATH handle for a bound socket's
+	 * vnode, so fall through to an EMPTYPATH lookup.
+	 */
+	if (*buf == '\0') {
+		error = unp_socket_fd_peer(td, fd, so2p);
+		if (error != ENOTSOCK)
+			return (error);
+	}
+
+	/* Resolve the path to a vnode. */
 	NDINIT_ATRIGHTS(&nd, LOOKUP, FOLLOW | LOCKSHARED | LOCKLEAF |
 	    (fd == AT_FDCWD ? 0 : EMPTYPATH), UIO_SYSSPACE, buf, fd,
 	    cap_rights_init_one(&rights, CAP_CONNECTAT));
@@ -3038,12 +3127,25 @@ unp_connectat_peer(struct thread *td, int fd, const char *buf,
 	NDFREE_PNBUF(&nd);
 
 	/*
-	 * Resolve the vnode to a referenced peer socket and drop the vnode
-	 * before connecting: the reference keeps the peer stable, so no vnode
-	 * lock is held across unp_connect_peer() (which matters for the
-	 * return_locked datagram fast path).
+	 * Dispatch on the resolved vnode, then drop it: for the socket cases the
+	 * returned reference keeps the peer stable, so the caller holds no vnode
+	 * lock across unp_connect_peer() (which matters for the return_locked
+	 * datagram fast path).
+	 *
+	 * A synthetic descriptor node -- as fdescfs fabricates for a /dev/fd/N
+	 * path -- carries no type of its own (VNON); opening it yields the
+	 * descriptor it stands for, which we resolve as the peer socket.
+	 * Otherwise the path must name a bound socket's vnode (VSOCK), which
+	 * unp_vnode_peer() connects to, rejecting any other type with ENOTSOCK.
+	 *
+	 * unp_dupfd_peer() resolves that descriptor exactly once: if it is not
+	 * a socket the connect fails, so this does *not* recur through a chain
+	 * of O_PATH handles of /dev/fd nodes, which could be arbitrarily long.
 	 */
-	error = unp_vnode_peer(nd.ni_vp, td, so2p);
+	if (nd.ni_vp->v_type == VNON)
+		error = unp_dupfd_peer(nd.ni_vp, td, so2p);
+	else
+		error = unp_vnode_peer(nd.ni_vp, td, so2p);
 	vput(nd.ni_vp);
 	return (error);
 }