Re: Fun with OFED and RDMA for NFS
- Go to: [ bottom of page ] [ top of archives ] [ this month ]
Date: Tue, 22 Sep 2026 01:03:16 UTC
On Mon, Sep 21, 2026 at 1:47 PM Konstantin Belousov <kib@kib.kiev.ua> wrote: > > On Sun, Sep 20, 2026 at 04:40:48PM -0700, Rick Macklem wrote: > > Hi, > > > > I have been debugging client code for NFS over RDMA and it is > > going pretty well. > > > > However, with repeated cycles of.. > > # make -j8 buildkernel > > where both the sources and obj are on the mount, I can get.. > > > > Sep 12 03:28:24 mercat1 kernel: mlx5_1: WARN: dump_cqe:273:(pid > > 100228): dump error cqe > > Sep 12 03:28:24 mercat1 kernel: 00000000 00000000 00000000 00000000 > > Sep 12 03:28:24 mercat1 syslogd: last message repeated 2 times > > Sep 12 03:28:24 mercat1 kernel: 00000000 08007806 25000907 1b0405d3 > > Sep 12 03:28:24 mercat1 kernel: rpcrdma_send_done: failed opcode=0 > > status=6 (Status 6 is IB_WC_MW_BIND_ERR.) > > Sep 12 03:28:24 mercat1 kernel: rpcrdma_send_done: pg=0x400003163 > > sgeaddr=0x0 len=0 lkey=0x0 num_sge=0 send_flags=0x0 > > > > Obviously, it has been trashed, but it is not obvious why? > > > > I don't get it on every "make -j8 buildkernel" and never seem to get > > it for "make -j8 buildworld". > > > > I have a KASSERT() just before the ib_post_send() that sanity > > checks the WR and this KASSERT() is never triggered, so the WR > > seems ok when ib_post_send() is called. > > > > So, I thought I'd try an "options KASAN" kernel, to see if it might > > find something. What happened? > > - Nothing. I've now run 10 cycles of "make -j8 buildkernel" and none > > of the above. (Without KASAN, it happens within the first 5 cycles.) > > So, I'm guessing that KASAN slows things down enough that it > > never occurs. > > > > I can's see anything in my code that would walk over the WR and > > the CQE that precedes it in the same structure is ok, since the done > > event handler gets called. > > > > Any ideas w.r.t. tracking this down further? > > > > Thanks in advance for any suggestions, rick > > ps: The NICs are ConnectX4 and have firmware that is a few years old. > > I've been paranoid w.r.t. asking Netperf admin for permission to update > > their firmware, since I can test effectively now and if (not likely) newer > > firmware broke them, I'd be stuck. > > I cannot guarantee anything of course, but usually updating firmware is not > that stressful. I agree. The upgrade I did on mercat1 (to the same fw as mercat2) went smoothly. (I think it also provides a "fallback to old firmware" path if there is trouble.) Actually, this quirk was very useful for testing. It provided a simple way to debug the reconnect code, so it now reconnects and continues on, at least for my testing. (It might have a memory leak during the reconnect. I'm not sure of I clear up everything on the old defunct connection (QP) quite right, so I can still use if for testing.) At some point, I think I will ask netperf-admin for permission to do an upgrade to the newest firmware on both of them. Have a good week, rick