Re: Fun with RDMA and NFS
- Reply: Konstantin Belousov : "Re: Fun with RDMA and NFS"
- In reply to: Konstantin Belousov : "Re: Fun with RDMA and NFS"
- Go to: [ bottom of page ] [ top of archives ] [ this month ]
Date: Sun, 09 Aug 2026 19:53:39 UTC
On Sun, Aug 9, 2026 at 12:20 PM Konstantin Belousov <kib@freebsd.org> wrote:
>
> Moving this to arch@.
>
> On Sat, Aug 08, 2026 at 05:03:56PM -0700, Rick Macklem wrote:
> > On Sat, Aug 8, 2026 at 2:53 PM Konstantin Belousov <kib@freebsd.org> wrote:
> > >
> > > On Sat, Aug 08, 2026 at 01:15:00PM -0700, Rick Macklem wrote:
> > > > Hi,
> > > >
> > > > As you all no doubt remember, in around 1986, Kirk proposed
> > > > something like this:
> > > > #ifdef notdef
> > > > /* Not Yet.. */
> > > > if (uiop->uio_iov->iov_op != NULL)
> > > > (*(uiop->uio_iov->iov_op))(mbufcp, uiocp, xfer);
> > > > else
> > > > #endif
> > > > The idea was that iov_op would "copy" data without
> > > > actually doing a memory->memory copying. It never
> > > > happened, afaik.
> > > >
> > > > Well, after only 40years, it looks like NFS might be
> > > > able to do that.
> > > >
> > > > For example, with RDMA a Read will look something like:
> > > > - NFS VOP_READ() allocates a buffer cache block and
> > > > passes the pages (in b_pages[]) to the RPC code.
> > > > - The RPC code will map the pages into bus dma space
> > > > and make the RPC call (referring to the memory region
> > > > for these pages). (What ofed calls FRWR.)
> > > > On the NFS server...
> > > > - The server will allocate pages for the read reply data and
> > > > map those into bus address space.
> > > > It will make a VOP_READ() call to copy the data into these
> > > > pages.
> > > > The RPC code will get the NIC to copy the data from the pages
> > > > to the pages on the NFS client (the buffer cache block in the client
> > > > using that FRWR stuff) and then send a small RPC reply noting
> > > > the Read has completed.
> > > > A write would be similar, except the data is copied into the NFS
> > > > server's pages before the VOP_WRITE().
> > > >
> > > > So, at this point, there is still a memory->memory copy done
> > > > in the exported file system's VOP_READ()/VOP_WRITE().
> > > >
> > > > I think that it should be possible to implement a couple of
> > > > new VOP_xxx() calls to avoid this memory->memory copy.
> > > > VOP_PAGEIO() - Would return an array of pages with the
> > > > read data in them or where the data can be written into
> > > > them (it could just return a "struct iovec iov[]"
> > > > with the vm_page_t for each page in the iov_base entries).
> > > > Presumably the pages would be buffer cache or ARC or ??
> > > > for the server file system.
> > > > VOP_DONE_PAGEIO() - Would be called once the transfer
> > > > is done to release the pages.
> > > >
> > > > Does this sound feasible?
> > > Yes, I think it is reasonable.
> > >
> > > Perhaps the following signature would be ok:
> > > int VOP_PAGEIO(struct vnode *vp, foff_t offset, size_t nbytes,
> > > enum uio_rw dir, vm_page_t ma[], int max_ma_len, int *ma_len);
> > > The VOP would return the ma array filled with the busied pages, which
> > > provide the file content from trunc_page(offset) to round_page(offset+nbytes).
> > >
> > > The twist there is that ma_len returns the number of pages that are
> > > actually filled into the ma array. For instance, for UFS, it is not
> > > feasible to fill more than a single buffer in time.
> > >
> > > From the first look, this is quite easy for UFS and tmpfs.
> > > If you want I can write the prototype.
> > Sounds fine to me. No rush, it'll be a while before I get through
> > initial testing and, since that is the server side, I won't be looking
> > at it until after that.
> >
> > >
> > > >
> > > > On the NFS client end, it should also be possible for O_DIRECT
> > > > to have the process's buffer mapped in, so that there is no
> > > > need for a buffer cache block. (This is more overhead than
> > > > the FRWR that can be used for kernel pages, but still might
> > > > be worth the effort.)
> > > But then we loose the cache coherency.
> > > Otherwise yes, vm_fault_quick_hold_pages() would provide the ma array for
> > > the user io buffer. The easiest example is probably kern/kern_physio.c.
> > I'm not worried about O_DIRECT, at least for now. It currently works
> > through the buffer cache by default, and that should work the same as
> > without O_DIRECT (except when the read and write RPCs are done).
> > (To be honest, my current prototype code just refuses to allow
> > RDMA mounts when newnfs_directio_enable (the sysctl variable) when
> > it is set. It is 0 by default.
>
> I started looking at implementing this stuff, and realized that really
> it is not needed in the proposed form. I am not saying that some helpers
> should be written, but I currently think that the proposed form is not
> useful.
>
> So lets state the desired operation:
> given the vnode, and the array of pointers to vm_page, the KPI should
> fill the array with actual pointers to the pages owned by the vnode,
> with the content. That is, the pages have valid content for read, or
> could be written to, in this case what is written to pages become the
> file content. Pages should be available until some explicit operation
> is done freeing them.
>
> First, I do not see why this would need to be a VOP. In fact, there is
> nothing that requires that the operation is done on vnode, and not on
> a vm_object. VM subsystem already handles similar operations, creating
> and filling the pages for fault, or preparing the valid pages for paging
> out.
>
> Second, I believe that we should specify the KPI somewhat modern by
> making it async. After all, the supposed usage of it is RDMA, where
> the main operation is async by its nature, so it might be a useful
> option to not block the caller.
>
> Right now I wrote the following KPI definition:
>
> diff --git a/sys/vm/vm_object.h b/sys/vm/vm_object.h
> index 5a3a78b5b5a1..6a13989117dd 100644
> --- a/sys/vm/vm_object.h
> +++ b/sys/vm/vm_object.h
> @@ -391,6 +391,20 @@ void vm_object_unwire(vm_object_t object, vm_ooffset_t offset,
> vm_size_t length, uint8_t queue);
> struct vnode *vm_object_vnode(vm_object_t object);
> bool vm_object_is_active(vm_object_t obj);
> +
> +#define VM_OBJECT_PGREQ_MAXPG 32
> +struct vm_object_pgreq {
> + vm_object_t obj;
> + vm_prot_t prot;
> + vm_page_t ma[VM_OBJECT_PGREQ_MAXPG];
> + vm_ooffset_t start;
> + vm_ooffset_t end;
> + int ma_len;
> + void *priv;
> + void (*comp)(int error, struct vm_object_pgreq *req);
> +};
> +int vm_object_get_pages_async(struct vm_object_pgreq *req);
> +
> #endif /* _KERNEL */
>
> #endif /* _VM_OBJECT_ */
>
> The function is called with the struct vm_object_pgreq which encodes the
> request, and in future will carry an auxillary data needed to execute
> the request. On completion, successful or not, the req->comp method is
> called. The priv pointer is available for caller to stuff its own data.
>
> Does it sound ok?
What about ZFS?
I'm just glancing at zfs_read() and it starts out doing a zfs_rangelock_enter()
and then assorted things I don't currently understand.
If you look at dmu_read_uio_dnode(), it appears that
dmu_buf_hold_array_by_dnode()
gets a list of buffers and then dmu_buf_rele_array() releases them.
--> It might be possible to get the first VOP_xxx() call to go as far as
dmu_buf_hold_array_by_dnode() and the second to do
dmu_buf_rele_array()?
(Take this with a large grain of salt, since I know diddly about ZFS.)
I was envisioning the new VOP_xxx() calls would basically be
VOP_READ()/VOP_WRITE() split up into two calls.
- The first would do the same as VOP_READ()/VOP_WRITE() up until
the actual data copying to/from the buffer/arc/?? and then return page
references instead of doing the actual copying.
- The second would do whatever VOP_READ()/VOP_WRITE() does
after the copying.
Basically, there's more to read/write than just finding the pages.
But if others know that the vm_object pages are what ZFS uses,
then I'll believe folk who know how ZFS behaves.
I'd like to come up with a patch for ZFS that is minimal, since any
commit to ZFS gets scary for me.
I'll admit I don't see any need for it being asynchronous, but I
don't have a strong opinion w.r.t. that. (I find most of the asynchronous
nature of ofed more bother than it's worth. All the steps to set up a qp/id
seem to be "start step, wait for completion, start next step..)
rick
>