[Bug 299135] vmm(4): IOMMU host domain omits ARI functions > 7 (e.g. SR-IOV VFs); their DMA is blocked once PCI passthrough enables VT-d
- Reply: bugzilla-noreply_a_freebsd.org: "[Bug 299135] vmm(4): IOMMU host domain omits ARI functions > 7 (e.g. SR-IOV VFs); their DMA is blocked once PCI passthrough enables VT-d"
- Reply: bugzilla-noreply_a_freebsd.org: "[Bug 299135] vmm(4): IOMMU host domain omits ARI functions > 7 (e.g. SR-IOV VFs); their DMA is blocked once PCI passthrough enables VT-d"
- Go to: [ bottom of page ] [ top of archives ] [ this month ]
Date: Mon, 05 Oct 2026 01:38:44 UTC
https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=299135
Bug ID: 299135
Summary: vmm(4): IOMMU host domain omits ARI functions > 7
(e.g. SR-IOV VFs); their DMA is blocked once PCI
passthrough enables VT-d
Product: Base System
Version: 15.1-RELEASE
Hardware: amd64
OS: Any
Status: New
Severity: Affects Some People
Priority: ---
Component: kern
Assignee: bugs@FreeBSD.org
Reporter: eppower@umich.edu
Created attachment 275386
--> https://bugs.freebsd.org/bugzilla/attachment.cgi?id=275386&action=edit
This patch forces vmm to scan function numbers >> PCI_FUNCMAX, thereby catching
the function numbers used by SR-IOV VFs and preventing them from getting wiped
out during IOMMU initialization
Disclaimer: AI (Claude Opus 5.5) was used to help me identify and solve this
corner case of bad behavior in vmm. Below is the bug report that Claude drafted
for this issue. I've reviewed it and it matches with the behavior I saw and the
steps taken to fix the problem. I believe it to be accurate, however if you
find there's information missing please let me know and I'll do what I can to
fill in the blanks.
Description:
I believe this to be similar to bug 234073.
When the first bhyve VM with a PCI passthrough device starts, vmm.ko
initializes the IOMMU (hw.vmm.iommu.initialized goes 0 -> 1). It builds a
1:1 "host domain" for all devices still owned by the host and then turns
on DMA remapping. iommu_init() in sys/amd64/vmm/io/iommu.c finds those
devices with:
for (bus = 0; bus <= PCI_BUSMAX; bus++)
for (slot = 0; slot <= PCI_SLOTMAX; slot++)
for (func = 0; func <= PCI_FUNCMAX; func++)
dev = pci_find_dbsf(0, bus, slot, func);
PCI_FUNCMAX is 7. With ARI, every function of a device is numbered in
slot 0 with function numbers up to 255 (PCIE_ARI_FUNCMAX). SR-IOV VFs
usually live there. Any host-owned device at a function number above 7 is
therefore never added to the host domain. Once remapping is enabled, its
DMA is blocked.
In this case, all host-side X710 VFs (iavf) failed the moment the first
passthrough VM started, and every vnet jail using a VF lost networking.
The passthrough device itself was unrelated to the NIC: assigning only the
HDA audio function of a GPU was enough. Unloading vmm.ko afterwards did
not recover the VFs; only a host reboot did (here I (human) diverge from Claude
- I could have torn down and re-created the VFs, I was just too lazy. A reboot
was simpler).
The same loop is present in releng/14.4, releng/15.1 and main (checked
early October 2026). The loop is generic code, so AMD-Vi hosts are
probably affected too, but I (human) have only tested Intel VT-d.
Devices that appear after IOMMU initialization are added to the host
domain by the pci_add_device event handler (iommu_pci_add()), so VFs
created after the first passthrough VM starts are probably unaffected. I
(Claude) inferred that from the code and have not tested it.
Environment:
- FreeBSD 15.1-RELEASE amd64 (also seen on 14.4-RELEASE)
- 2x Intel Xeon Platinum 8260, Supermicro X11DPH-i, latest BIOS, VT-d on
- Intel X710-T2L (8086:15ff), ixl(4) PF with SR-IOV VFs created via
iovctl(8). ARI is in use, so the VFs enumerate as e.g.
iavf0@pci0:94:0:16 ... iavf8@pci0:94:0:26 (PF ixl2, pci0:94:0:0)
iavf9@pci0:94:0:80 ... iavf13@pci0:94:0:85 (PF ixl3, pci0:94:0:1)
Some VFs are host-owned (iavf, assigned to vnet jails); a few are
configured passthrough=true and owned by ppt(4).
- Passthrough devices reserved via pptdevs in loader.conf (a discrete GPU
and its HDA function).
How to reproduce:
1. Create SR-IOV VFs on an ARI-capable NIC so that host-owned VFs have
function numbers > 7, and use them on the host (e.g. in vnet jails).
2. Confirm hw.vmm.iommu.initialized is 0 and VF networking works.
3. Start any bhyve VM with a passthrough device, for example:
bhyve -c 1 -m 256M -AHPw -S \
-l bootrom,/usr/local/share/uefi-firmware/BHYVE_UEFI.fd \
-s 0,hostbridge -s 31,lpc -s 2:0,passthru,137/0/0 testvm
4. hw.vmm.iommu.initialized becomes 1 and, at that moment, dmesg shows:
iavf3: ARQ Critical Error detected
iavf0: ARQ Critical Error detected
iavf1: ARQ Critical Error detected
iavf3: ASQ Critical Error detected
iavf0: ASQ Critical Error detected
iavf1: ASQ Critical Error detected
iavf2: ARQ Critical Error detected
iavf2: ASQ Critical Error detected
ixl2: Malicious Driver Detection event 1 on RX queue 35, pf number 0
(PF-0), (VF-2)
followed by a continuous stream of Malicious Driver Detection events
for the VFs. Networking over all host-owned VFs is dead. Snapshots taken
while the VM was running and after it was destroyed show the failure
starts at VM start (IOMMU enable), not at teardown.
Fix:
The attached patch (vmm-iommu-ari-vfs.patch, against releng/15.1) scans
function numbers up to PCIE_ARI_FUNCMAX when slot == 0, so ARI functions
are added to the host domain like everything else. ppt(4)-owned devices
are still skipped as before.
(In Claude's opinion) A cleaner alternative would be to walk pci_devq directly
instead of
probing every bus/slot/function combination; I (Claude) kept the change
minimal.
Testing with the patch:
- vmm.ko rebuilt from releng/15.1 + patch; GENERIC kernel otherwise stock.
- Started and destroyed passthrough VMs repeatedly (HDA function alone,
then GPU + HDA, then a full Windows guest and a Linux guest). In every
case hw.vmm.iommu.initialized became 1, no ARQ/ASQ errors and no
Malicious Driver Detection events were logged, and a vnet jail on a
host-owned VF kept working throughout, e.g.
# jexec -l jail1 ping -c 3 peer.example.org
3 packets transmitted, 3 packets received, 0.0% packet loss
measured while the VM was running and again after it was destroyed.
- Without the patch, the same sequence reproducibly killed all host-owned
VFs on this machine (on both 14.4 and 15.1).
Attachment: vmm-iommu-ari-vfs.patch
Back story, if you've read this far: I am not a software developer and could
never have found or fixed this on my own - I'm a hobbyist and tinkerer with a
day job that doesn't involve software development. Here's the human
perspective: I have a Supermicro X11DPH-I mainboard (dual Cascade Lake Xeon
board) with a discrete Intel X710-T2L 10GbE NIC installed. The X710 supports
SR-IOV, and I use multiple VFs on this NIC to support my jail networking. I
also have a Sparkle A310 ECO GPU installed, which is an Intel Alchemist
generation dGPU. The GPU is intended to provide hardware-accelerated
transcoding for a VM running jellyfin. The host machine was originally brought
online with 14.3-RELEASE, then upgraded to 14.4-RELEASE, and most recently
upgraded to 15.1-RELEASE. Across all versions I noted the same behavior - when
attempting to pass through the GPU to a VM (either with a native bhyve call or
using vm-bhyve), all VFs on my NIC immediately died a horrible death with
messages like "Malicious driver detection" and complete loss of networking
functionality. The root cause Claude identified and the fix it proposed are
simple: vmm is not counting high enough when mapping host-owned devices during
IOMMU initialization. The fix is a couple of lines, and I can only say that it
works _for my specific hardware + OS combination_. With the fix in place I can
now pass a ppt device through to a bhyve VM and the VFs on my X710 NIC continue
to chug along with no interruption.
--
You are receiving this mail because:
You are the assignee for the bug.