git: 8404c535d9bc - main - ice: Rebuild VF VSIs after PF resets

From: Kevin Bowling <kbowling_at_FreeBSD.org>
Date: Thu, 17 Sep 2026 06:42:45 UTC
The branch main has been updated by kbowling:

URL: https://cgit.FreeBSD.org/src/commit/?id=8404c535d9bc75287f56e353cc7bff9d9ccd85b8

commit 8404c535d9bc75287f56e353cc7bff9d9ccd85b8
Author:     Kevin Bowling <kbowling@FreeBSD.org>
AuthorDate: 2026-08-18 03:25:11 +0000
Commit:     Kevin Bowling <kbowling@FreeBSD.org>
CommitDate: 2026-09-17 06:42:33 +0000

    ice: Rebuild VF VSIs after PF resets
    
    PF and device resets discard all hardware VSI state.  ice_rebuild()
    only recreated the main PF VSI before replaying configuration for every
    VSI.  As a result, replay used stale VF VSI handles.  Firmware rejected
    it with AQ_RC_EACCES and the failure aborted the entire PF rebuild.
    
    Re-add each VF VSI before replaying its configuration, and do not
    publish VFACTIVE until both operations succeed.  Clear the initialized
    state so that the VF must negotiate again.  If a VF rebuild fails, leave
    it inactive and continue so that one guest cannot prevent the PF or
    sibling VFs from recovering.
    
    Mark it initialized only after the resource response is submitted
    successfully, so the state follows completion of the mailbox handshake.
    
    Make that failed state authoritative in the mailbox path.  While the
    firmware VSI is invalid, permit only VERSION and RESET_VF and reject
    operations which require VSI state.  A VFR may complete the hardware
    reset, but cannot make the PF-owned VSI valid or publish the VF active.
    
    Preserve accumulated VF statistics while establishing a new raw
    hardware sample after reconstruction.  This keeps the cumulative totals
    returned to iavf monotonic across PF resets.  Return after rejecting a
    GET_STATS request for the wrong VSI so that it cannot receive a second
    success reply.
    
    Before a locally initiated reset, notify initialized VFs with
    RESET_IMPENDING while the mailbox control queue is still alive.  Ignore
    individual send failures so one VF cannot prevent notification of its
    siblings or the reset itself.  Send the event from the common reset
    preparation path and before directly triggering CORE and GLOBAL resets.
    
    Preserve the inactive state across later VF resets.  The zero-queue
    Disable LAN Tx AQ remains mandatory to complete every VFR, but do not
    clear VFSWR or publish VFACTIVE while the PF-owned VSI remains invalid.
    
    Linux ice uses the same separation: generic rebuild excludes VF VSIs,
    the VF reset path rebuilds them separately, and reset preparation
    notifies initialized VFs before tearing down the control queues.
    
    Validated on an E810-XXV with one and eight host-attached iavf VFs.
    Repeated PF and CORE resets recovered every VF under traffic without
    watchdog, MDD, or persistent data-path errors.  An additional one-VF
    test kept traffic active across PF and CORE resets; carrier and traffic
    recovered automatically and both PF and VF watchdog counters stayed at
    zero.
    
    Two four-queue VFs passed through to a Linux 7.0 iavf guest also
    recovered carrier and traffic automatically after PF and CORE resets.
    Simultaneous traffic on both VFs resumed without intervention and the PF
    watchdog counter remained zero.
    
    DPDK 25.11 testpmd, using vfio no-IOMMU and two queues per VF, sustained
    about 4.55 Mpps per VF before reset.  It received reset events for both
    VFs after PF and CORE resets.  The documented ethdev stop, reset,
    reconfigure, and start sequence restored traffic after each reset.
    Physical bus mastering remained enabled and the PF watchdog counter
    remained zero.
    
    MFC after:      2 weeks
    Sponsored by:   BBOX.io
    Differential Revision:  https://reviews.freebsd.org/D58905
---
 share/man/man4/ice.4       |  10 +++-
 sys/dev/ice/ice_iov.c      | 122 ++++++++++++++++++++++++++++++++++++++++++---
 sys/dev/ice/ice_iov.h      |   5 +-
 sys/dev/ice/ice_lib.c      |  14 ++++++
 sys/dev/ice/if_ice_iflib.c |   5 ++
 5 files changed, 146 insertions(+), 10 deletions(-)

diff --git a/share/man/man4/ice.4 b/share/man/man4/ice.4
index 1ce49c4fe632..51c172e08fcb 100644
--- a/share/man/man4/ice.4
+++ b/share/man/man4/ice.4
@@ -32,7 +32,7 @@
 .\"
 .\" * Other names and brands may be claimed as the property of others.
 .\"
-.Dd August 11, 2026
+.Dd August 19, 2026
 .Dt ICE 4
 .Os
 .Sh NAME
@@ -1145,6 +1145,14 @@ The VF's default mac address does not count towards this limit.
 By default, this is set to 64.
 .El
 .Pp
+PF and device resets discard the hardware state of every VF.
+The driver keeps each configured VF inactive while reconstructing its VSI and
+allows it to renegotiate resources only after reconstruction succeeds.
+If one VF cannot be reconstructed, it remains configured but uninitialized
+while the PF and its sibling VFs continue recovering.
+A VF function-level reset cannot recreate this PF-owned VSI state; another PF
+or device reset, or SR-IOV configuration recreation, is required.
+.Pp
 An up to date list of parameters and their defaults can be found by using
 .Xr iovctl 8
 with the
diff --git a/sys/dev/ice/ice_iov.c b/sys/dev/ice/ice_iov.c
index c5a3e1060e44..11d6bbed9b2d 100644
--- a/sys/dev/ice/ice_iov.c
+++ b/sys/dev/ice/ice_iov.c
@@ -494,6 +494,36 @@ ice_iov_handle_vflr(struct ice_softc *sc)
 	}
 }
 
+/**
+ * ice_iov_notify_vfs_reset - Notify initialized VFs of an impending reset
+ * @sc: device softc structure
+ *
+ * Give VF drivers advance notice while the mailbox control queue is still
+ * alive. Ignore individual send failures so one VF cannot prevent the PF from
+ * notifying its siblings or proceeding with the reset.
+ */
+void
+ice_iov_notify_vfs_reset(struct ice_softc *sc)
+{
+	struct virtchnl_pf_event event = {};
+	struct ice_hw *hw = &sc->hw;
+	struct ice_vf *vf;
+
+	if (!ice_check_sq_alive(hw, &hw->mailboxq))
+		return;
+
+	event.event = VIRTCHNL_EVENT_RESET_IMPENDING;
+	event.severity = PF_EVENT_SEVERITY_CERTAIN_DOOM;
+	for (int i = 0; i < sc->num_vfs; i++) {
+		vf = &sc->vfs[i];
+		if ((atomic_load_acq_32(&vf->vf_flags) &
+		    VF_FLAG_INITIALIZED) == 0)
+			continue;
+		(void)ice_aq_send_msg_to_vf(hw, vf->vf_num, VIRTCHNL_OP_EVENT,
+		    VIRTCHNL_STATUS_SUCCESS, (u8 *)&event, sizeof(event), NULL);
+	}
+}
+
 /**
  * ice_iov_ready_vf - Setup VF interrupts and mark it as ready
  * @sc: device softc structure
@@ -523,6 +553,54 @@ ice_iov_ready_vf(struct ice_softc *sc, struct ice_vf *vf)
 	ice_flush(hw);
 }
 
+/**
+ * ice_iov_rebuild_vf - Rebuild a VF VSI after a PF or device reset
+ * @sc: device softc structure
+ * @vsi: VF VSI to rebuild
+ *
+ * PF and device resets discard the hardware VSI and interrupt state for every
+ * VF. Re-add the VSI and replay its configuration before reporting the VF as
+ * active. A failed rebuild leaves the VF inactive while allowing the PF and
+ * other VFs to recover.
+ */
+int
+ice_iov_rebuild_vf(struct ice_softc *sc, struct ice_vsi *vsi)
+{
+	struct ice_eth_stats accumulated_stats;
+	struct ice_hw *hw = &sc->hw;
+	struct ice_vf *vf;
+	int error, status;
+
+	MPASS(vsi->type == ICE_VSI_VF);
+	vf = ice_iov_get_vf(sc, vsi->vf_num);
+	atomic_clear_32(&vf->vf_flags, VF_FLAG_INITIALIZED);
+	atomic_set_32(&vf->vf_flags, VF_FLAG_REBUILD_FAILED);
+
+	/* A new hardware VSI starts a new raw statistics epoch. */
+	accumulated_stats = vsi->hw_stats.cur;
+	error = ice_initialize_vsi(vsi);
+	if (error != 0) {
+		device_printf(sc->dev,
+		    "Unable to re-initialize VF %d VSI, err %s\n",
+		    vf->vf_num, ice_err_str(error));
+		return (error);
+	}
+	vsi->hw_stats.cur = accumulated_stats;
+
+	status = ice_replay_vsi(hw, vsi->idx);
+	if (status != 0) {
+		device_printf(sc->dev,
+		    "Failed to replay VF %d VSI, err %s aq_err %s\n",
+		    vf->vf_num, ice_status_str(status),
+		    ice_aq_str(hw->adminq.sq_last_status));
+		return (EIO);
+	}
+
+	atomic_clear_32(&vf->vf_flags, VF_FLAG_REBUILD_FAILED);
+	ice_iov_ready_vf(sc, vf);
+	return (0);
+}
+
 /**
  * ice_reset_vf - Perform a hardware reset (VFR) on a VF
  * @sc: device softc structure
@@ -575,14 +653,14 @@ ice_reset_vf(struct ice_softc *sc, struct ice_vf *vf, bool trigger_vflr)
 		device_printf(sc->dev,
 			"VF-%d PCI transactions stuck\n", vf->vf_num);
 
-	/* Disable TX queues, which is required during VF reset */
-	status = ice_dis_vsi_txq(hw->port_info, vf->vsi->idx, 0, 0, NULL, NULL,
-			NULL, ICE_VF_RESET, vf->vf_num, NULL);
+	/* This zero-queue command is required to complete every VF reset. */
+	status = ice_dis_vsi_txq(hw->port_info, vf->vsi->idx, 0, 0,
+	    NULL, NULL, NULL, ICE_VF_RESET, vf->vf_num, NULL);
 	if (status)
 		device_printf(sc->dev,
-			      "%s: Failed to disable LAN Tx queues: err %s aq_err %s\n",
-			      __func__, ice_status_str(status),
-			      ice_aq_str(hw->adminq.sq_last_status));
+		    "%s: Failed to disable LAN Tx queues: err %s aq_err %s\n",
+		    __func__, ice_status_str(status),
+		    ice_aq_str(hw->adminq.sq_last_status));
 
 	/* Then check for the VF reset to finish in HW */
 	for (i = 0; i < ICE_VPGEN_VFRSTAT_WAIT_COUNT; i++) {
@@ -596,6 +674,11 @@ ice_reset_vf(struct ice_softc *sc, struct ice_vf *vf, bool trigger_vflr)
 		device_printf(sc->dev,
 			"VF-%d Reset is stuck\n", vf->vf_num);
 
+	/* A VFR cannot recover PF-owned VSI state lost during PF rebuild. */
+	if ((atomic_load_acq_32(&vf->vf_flags) &
+	    VF_FLAG_REBUILD_FAILED) != 0)
+		return;
+
 	ice_iov_ready_vf(sc, vf);
 }
 
@@ -620,6 +703,7 @@ ice_vc_get_vf_res_msg(struct ice_softc *sc, struct ice_vf *vf, u8 *msg_buf)
 	struct virtchnl_vsi_resource *vsi_res;
 	u16 vf_res_len;
 	u32 vf_caps;
+	int status;
 
 	/* XXX: Only support one VSI per VF, so this size doesn't need adjusting */
 	vf_res_len = sizeof(struct virtchnl_vf_resource);
@@ -653,8 +737,15 @@ ice_vc_get_vf_res_msg(struct ice_softc *sc, struct ice_vf *vf, u8 *msg_buf)
 	if (!ETHER_IS_ZERO(vf->mac))
 		memcpy(vsi_res->default_mac_addr, vf->mac, ETHER_ADDR_LEN);
 
-	ice_aq_send_msg_to_vf(hw, vf->vf_num, VIRTCHNL_OP_GET_VF_RESOURCES,
-	    VIRTCHNL_STATUS_SUCCESS, (u8 *)vf_res, vf_res_len, NULL);
+	status = ice_aq_send_msg_to_vf(hw, vf->vf_num,
+	    VIRTCHNL_OP_GET_VF_RESOURCES, VIRTCHNL_STATUS_SUCCESS,
+	    (u8 *)vf_res, vf_res_len, NULL);
+	if (status == 0)
+		atomic_set_32(&vf->vf_flags, VF_FLAG_INITIALIZED);
+	else
+		device_printf(sc->dev,
+		    "Unable to send VF-%u resource response, err %s\n",
+		    vf->vf_num, ice_status_str(status));
 
 	free(vf_res, M_ICE);
 }
@@ -1489,6 +1580,7 @@ ice_vc_get_stats_msg(struct ice_softc *sc, struct ice_vf *vf, u8 *msg_buf)
 		    __func__, vf->vf_num, vqs->vsi_id, vsi->idx);
 		ice_aq_send_msg_to_vf(hw, vf->vf_num, VIRTCHNL_OP_GET_STATS,
 		    VIRTCHNL_STATUS_ERR_PARAM, NULL, 0, NULL);
+		return;
 	}
 
 	ice_update_vsi_hw_stats(vf->vsi);
@@ -1667,6 +1759,7 @@ ice_vc_handle_vf_msg(struct ice_softc *sc, struct ice_rq_event_info *event)
 	device_t dev = sc->dev;
 	struct ice_vf *vf;
 	int err = 0;
+	u32 vf_flags;
 
 	u32 v_opcode = event->desc.cookie_high;
 	u16 v_id = event->desc.retval;
@@ -1690,6 +1783,19 @@ ice_vc_handle_vf_msg(struct ice_softc *sc, struct ice_rq_event_info *event)
 		return;
 	}
 
+	vf_flags = atomic_load_acq_32(&vf->vf_flags);
+	if ((vf_flags & VF_FLAG_ENABLED) == 0 || vf->vsi == NULL)
+		return;
+
+	/* Only a later PF rebuild can restore an invalid firmware VSI. */
+	if ((vf_flags & VF_FLAG_REBUILD_FAILED) != 0 &&
+	    v_opcode != VIRTCHNL_OP_VERSION &&
+	    v_opcode != VIRTCHNL_OP_RESET_VF) {
+		ice_aq_send_msg_to_vf(hw, v_id, v_opcode,
+		    VIRTCHNL_STATUS_ERR_ADMIN_QUEUE_ERROR, NULL, 0, NULL);
+		return;
+	}
+
 	switch (v_opcode) {
 	case VIRTCHNL_OP_VERSION:
 		ice_vc_version_msg(sc, vf, msg);
diff --git a/sys/dev/ice/ice_iov.h b/sys/dev/ice/ice_iov.h
index c4fb3e932e3f..68c897461a13 100644
--- a/sys/dev/ice/ice_iov.h
+++ b/sys/dev/ice/ice_iov.h
@@ -69,6 +69,8 @@ enum ice_vf_flags {
 	VF_FLAG_VLAN_CAP		= BIT(2),
 	VF_FLAG_PROMISC_CAP		= BIT(3),
 	VF_FLAG_MAC_ANTI_SPOOF		= BIT(4),
+	VF_FLAG_INITIALIZED		= BIT(5),
+	VF_FLAG_REBUILD_FAILED		= BIT(6),
 };
 
 /**
@@ -114,12 +116,13 @@ int ice_iov_detach(struct ice_softc *sc);
 
 int ice_iov_init(struct ice_softc *sc, uint16_t num_vfs, const nvlist_t *params);
 int ice_iov_add_vf(struct ice_softc *sc, uint16_t vfnum, const nvlist_t *params);
+int ice_iov_rebuild_vf(struct ice_softc *sc, struct ice_vsi *vsi);
 void ice_iov_uninit(struct ice_softc *sc);
 
 void ice_iov_handle_vflr(struct ice_softc *sc);
+void ice_iov_notify_vfs_reset(struct ice_softc *sc);
 
 void ice_vc_handle_vf_msg(struct ice_softc *sc, struct ice_rq_event_info *event);
 void ice_vc_notify_all_vfs_link_state(struct ice_softc *sc);
 
 #endif /* _ICE_IOV_H_ */
-
diff --git a/sys/dev/ice/ice_lib.c b/sys/dev/ice/ice_lib.c
index 4c1b48db18e1..6af0d2e550d2 100644
--- a/sys/dev/ice/ice_lib.c
+++ b/sys/dev/ice/ice_lib.c
@@ -6462,6 +6462,9 @@ ice_sysctl_request_reset(SYSCTL_HANDLER_ARGS)
 	 * interrupt on all PFs. Initiate the reset now. Preparation and
 	 * rebuild logic will be handled by the admin status task.
 	 */
+#ifdef PCI_IOV
+	ice_iov_notify_vfs_reset(sc);
+#endif
 	status = ice_reset(hw, reset_type);
 
 	/*
@@ -7860,6 +7863,17 @@ ice_replay_all_vsi_cfg(struct ice_softc *sc)
 		if (!vsi)
 			continue;
 
+#ifdef PCI_IOV
+		if (vsi->type == ICE_VSI_VF) {
+			status = ice_iov_rebuild_vf(sc, vsi);
+			if (status != 0)
+				device_printf(sc->dev,
+				    "Failed to rebuild VF %d VSI; leaving VF disabled\n",
+				    vsi->vf_num);
+			continue;
+		}
+#endif
+
 		status = ice_replay_vsi(hw, vsi->idx);
 		if (status) {
 			device_printf(sc->dev, "Failed to replay VSI %d, err %s aq_err %s\n",
diff --git a/sys/dev/ice/if_ice_iflib.c b/sys/dev/ice/if_ice_iflib.c
index 4a8f35ac9aad..7f6bd1e1e0a3 100644
--- a/sys/dev/ice/if_ice_iflib.c
+++ b/sys/dev/ice/if_ice_iflib.c
@@ -2556,6 +2556,11 @@ ice_prepare_for_reset(struct ice_softc *sc)
 	if (ice_test_state(&sc->state, ICE_STATE_RECOVERY_MODE))
 		return;
 
+#ifdef PCI_IOV
+	/* Notify initialized VFs while the mailbox queue is still available. */
+	ice_iov_notify_vfs_reset(sc);
+#endif
+
 	/* Restore identification while the control queues are still usable. */
 	ice_led_restore(sc);