Re: Fun with RDMA and NFS
- Reply: Rick Macklem : "Re: Fun with RDMA and NFS"
- In reply to: Rick Macklem : "Re: Fun with RDMA and NFS"
- Go to: [ bottom of page ] [ top of archives ] [ this month ]
Date: Sun, 09 Aug 2026 20:01:28 UTC
On Sun, Aug 09, 2026 at 12:53:39PM -0700, Rick Macklem wrote:
> On Sun, Aug 9, 2026 at 12:20 PM Konstantin Belousov <kib@freebsd.org> wrote:
> >
> > Moving this to arch@.
> >
> > On Sat, Aug 08, 2026 at 05:03:56PM -0700, Rick Macklem wrote:
> > > On Sat, Aug 8, 2026 at 2:53 PM Konstantin Belousov <kib@freebsd.org> wrote:
> > > >
> > > > On Sat, Aug 08, 2026 at 01:15:00PM -0700, Rick Macklem wrote:
> > > > > Hi,
> > > > >
> > > > > As you all no doubt remember, in around 1986, Kirk proposed
> > > > > something like this:
> > > > > #ifdef notdef
> > > > > /* Not Yet.. */
> > > > > if (uiop->uio_iov->iov_op != NULL)
> > > > > (*(uiop->uio_iov->iov_op))(mbufcp, uiocp, xfer);
> > > > > else
> > > > > #endif
> > > > > The idea was that iov_op would "copy" data without
> > > > > actually doing a memory->memory copying. It never
> > > > > happened, afaik.
> > > > >
> > > > > Well, after only 40years, it looks like NFS might be
> > > > > able to do that.
> > > > >
> > > > > For example, with RDMA a Read will look something like:
> > > > > - NFS VOP_READ() allocates a buffer cache block and
> > > > > passes the pages (in b_pages[]) to the RPC code.
> > > > > - The RPC code will map the pages into bus dma space
> > > > > and make the RPC call (referring to the memory region
> > > > > for these pages). (What ofed calls FRWR.)
> > > > > On the NFS server...
> > > > > - The server will allocate pages for the read reply data and
> > > > > map those into bus address space.
> > > > > It will make a VOP_READ() call to copy the data into these
> > > > > pages.
> > > > > The RPC code will get the NIC to copy the data from the pages
> > > > > to the pages on the NFS client (the buffer cache block in the client
> > > > > using that FRWR stuff) and then send a small RPC reply noting
> > > > > the Read has completed.
> > > > > A write would be similar, except the data is copied into the NFS
> > > > > server's pages before the VOP_WRITE().
> > > > >
> > > > > So, at this point, there is still a memory->memory copy done
> > > > > in the exported file system's VOP_READ()/VOP_WRITE().
> > > > >
> > > > > I think that it should be possible to implement a couple of
> > > > > new VOP_xxx() calls to avoid this memory->memory copy.
> > > > > VOP_PAGEIO() - Would return an array of pages with the
> > > > > read data in them or where the data can be written into
> > > > > them (it could just return a "struct iovec iov[]"
> > > > > with the vm_page_t for each page in the iov_base entries).
> > > > > Presumably the pages would be buffer cache or ARC or ??
> > > > > for the server file system.
> > > > > VOP_DONE_PAGEIO() - Would be called once the transfer
> > > > > is done to release the pages.
> > > > >
> > > > > Does this sound feasible?
> > > > Yes, I think it is reasonable.
> > > >
> > > > Perhaps the following signature would be ok:
> > > > int VOP_PAGEIO(struct vnode *vp, foff_t offset, size_t nbytes,
> > > > enum uio_rw dir, vm_page_t ma[], int max_ma_len, int *ma_len);
> > > > The VOP would return the ma array filled with the busied pages, which
> > > > provide the file content from trunc_page(offset) to round_page(offset+nbytes).
> > > >
> > > > The twist there is that ma_len returns the number of pages that are
> > > > actually filled into the ma array. For instance, for UFS, it is not
> > > > feasible to fill more than a single buffer in time.
> > > >
> > > > From the first look, this is quite easy for UFS and tmpfs.
> > > > If you want I can write the prototype.
> > > Sounds fine to me. No rush, it'll be a while before I get through
> > > initial testing and, since that is the server side, I won't be looking
> > > at it until after that.
> > >
> > > >
> > > > >
> > > > > On the NFS client end, it should also be possible for O_DIRECT
> > > > > to have the process's buffer mapped in, so that there is no
> > > > > need for a buffer cache block. (This is more overhead than
> > > > > the FRWR that can be used for kernel pages, but still might
> > > > > be worth the effort.)
> > > > But then we loose the cache coherency.
> > > > Otherwise yes, vm_fault_quick_hold_pages() would provide the ma array for
> > > > the user io buffer. The easiest example is probably kern/kern_physio.c.
> > > I'm not worried about O_DIRECT, at least for now. It currently works
> > > through the buffer cache by default, and that should work the same as
> > > without O_DIRECT (except when the read and write RPCs are done).
> > > (To be honest, my current prototype code just refuses to allow
> > > RDMA mounts when newnfs_directio_enable (the sysctl variable) when
> > > it is set. It is 0 by default.
> >
> > I started looking at implementing this stuff, and realized that really
> > it is not needed in the proposed form. I am not saying that some helpers
> > should be written, but I currently think that the proposed form is not
> > useful.
> >
> > So lets state the desired operation:
> > given the vnode, and the array of pointers to vm_page, the KPI should
> > fill the array with actual pointers to the pages owned by the vnode,
> > with the content. That is, the pages have valid content for read, or
> > could be written to, in this case what is written to pages become the
> > file content. Pages should be available until some explicit operation
> > is done freeing them.
> >
> > First, I do not see why this would need to be a VOP. In fact, there is
> > nothing that requires that the operation is done on vnode, and not on
> > a vm_object. VM subsystem already handles similar operations, creating
> > and filling the pages for fault, or preparing the valid pages for paging
> > out.
> >
> > Second, I believe that we should specify the KPI somewhat modern by
> > making it async. After all, the supposed usage of it is RDMA, where
> > the main operation is async by its nature, so it might be a useful
> > option to not block the caller.
> >
> > Right now I wrote the following KPI definition:
> >
> > diff --git a/sys/vm/vm_object.h b/sys/vm/vm_object.h
> > index 5a3a78b5b5a1..6a13989117dd 100644
> > --- a/sys/vm/vm_object.h
> > +++ b/sys/vm/vm_object.h
> > @@ -391,6 +391,20 @@ void vm_object_unwire(vm_object_t object, vm_ooffset_t offset,
> > vm_size_t length, uint8_t queue);
> > struct vnode *vm_object_vnode(vm_object_t object);
> > bool vm_object_is_active(vm_object_t obj);
> > +
> > +#define VM_OBJECT_PGREQ_MAXPG 32
> > +struct vm_object_pgreq {
> > + vm_object_t obj;
> > + vm_prot_t prot;
> > + vm_page_t ma[VM_OBJECT_PGREQ_MAXPG];
> > + vm_ooffset_t start;
> > + vm_ooffset_t end;
> > + int ma_len;
> > + void *priv;
> > + void (*comp)(int error, struct vm_object_pgreq *req);
> > +};
> > +int vm_object_get_pages_async(struct vm_object_pgreq *req);
> > +
> > #endif /* _KERNEL */
> >
> > #endif /* _VM_OBJECT_ */
> >
> > The function is called with the struct vm_object_pgreq which encodes the
> > request, and in future will carry an auxillary data needed to execute
> > the request. On completion, successful or not, the req->comp method is
> > called. The priv pointer is available for caller to stuff its own data.
> >
> > Does it sound ok?
> What about ZFS?
>
> I'm just glancing at zfs_read() and it starts out doing a zfs_rangelock_enter()
> and then assorted things I don't currently understand.
> If you look at dmu_read_uio_dnode(), it appears that
> dmu_buf_hold_array_by_dnode()
> gets a list of buffers and then dmu_buf_rele_array() releases them.
> --> It might be possible to get the first VOP_xxx() call to go as far as
> dmu_buf_hold_array_by_dnode() and the second to do
> dmu_buf_rele_array()?
> (Take this with a large grain of salt, since I know diddly about ZFS.)
>
> I was envisioning the new VOP_xxx() calls would basically be
> VOP_READ()/VOP_WRITE() split up into two calls.
> - The first would do the same as VOP_READ()/VOP_WRITE() up until
> the actual data copying to/from the buffer/arc/?? and then return page
> references instead of doing the actual copying.
> - The second would do whatever VOP_READ()/VOP_WRITE() does
> after the copying.
> Basically, there's more to read/write than just finding the pages.
>
> But if others know that the vm_object pages are what ZFS uses,
> then I'll believe folk who know how ZFS behaves.
>
> I'd like to come up with a patch for ZFS that is minimal, since any
> commit to ZFS gets scary for me.
>
> I'll admit I don't see any need for it being asynchronous, but I
> don't have a strong opinion w.r.t. that. (I find most of the asynchronous
> nature of ofed more bother than it's worth. All the steps to set up a qp/id
> seem to be "start step, wait for completion, start next step..)
The proposed design does not require any changes or additions to any
of the existing fs. It is just a VM wrapper.
I can drop the async option.