From nobody Sat Aug 29 23:13:37 2026 X-Original-To: dev-commits-src-all@mlmmj.nyi.freebsd.org Received: from mx1.freebsd.org (mx1.freebsd.org [IPv6:2610:1c1:1:606c::19:1]) by mlmmj.nyi.freebsd.org (Postfix) with ESMTP id 4hXWJy0rbPz6pZYV for ; Sat, 29 Aug 2026 23:13:38 +0000 (UTC) (envelope-from git@FreeBSD.org) Received: from mxrelay.nyi.freebsd.org (mxrelay.nyi.freebsd.org [IPv6:2610:1c1:1:606c::19:3]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (4096 bits) server-digest SHA256 client-signature RSA-PSS (4096 bits) client-digest SHA256) (Client CN "mxrelay.nyi.freebsd.org", Issuer "YR2" (not verified)) by mx1.freebsd.org (Postfix) with ESMTPS id 4hXWJy0Fbtz3NFT for ; Sat, 29 Aug 2026 23:13:38 +0000 (UTC) (envelope-from git@FreeBSD.org) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=freebsd.org; s=dkim; t=1788045218; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding; bh=hwzT5qe4Gw/9OHYNAG5LQajTr0QQIcoEbtjb6mxm9tE=; b=iqX38wMMLrNu7ipixJ+IIJg2TWKnWgUzrKFM77VSxQCBKYT8u8HAfjMZVxLu/IUpU9AhLq yzSpvwBSrgUQWlah3raQVukHDx0RizwXeues4cGSabAfxGgAz96E/1QjCopVV84Zg5+rs0 +EOMRtkbGaiZ1EIPTILE31Sn18s465oyyAjBT7msi/eCrpzsLUsCjr4DcNSaWa6ZGNaHNZ 06g4l78bgISopuB/nthWYHP1DFuKTv0WhYQ7xU5Ll8m9TaY5ENjovZq5d+xvUj0WCIrUcf xTjzGeg72g9w6a5nOB2Hrim1T+V8PQksILB+gBmQLiO78rJisD+tXXWvX4ihdA== ARC-Seal: i=1; s=dkim; d=freebsd.org; t=1788045218; a=rsa-sha256; cv=none; b=hfHvaO7TxRJgn/QuariXQq09k3eiNVfRTd9S2OpHGUXLxhk/d+x3yt3+50u7OoPEkxpA67 EfdD3aYGYKTEuajbNkFTEd8A9NdVG6IyBkZb+kcaVsCun9s6X72Jc8baEx8ciMkASnZR+d Qcy4IreMDwdorO0iarr5whC1EMv9LO1ZlhyXz4WOtJL/k2/TZ/vV+84dfwwFliGns3vOEN /hCAW9G56J+5QcP5ZTXlS6TAP9UL8hnKtbg6/Ex2ZO5ERWuNLEtY6ihwR0lItvZKJUuetO qXjTQ4Y9iBqpXK7ihcthQ8Z5ts/OFTJoxb6fZbAA/j7i93WRLG6+tg1nvhGXvg== ARC-Authentication-Results: i=1; mx1.freebsd.org; none ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=freebsd.org; s=dkim; t=1788045218; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding; bh=hwzT5qe4Gw/9OHYNAG5LQajTr0QQIcoEbtjb6mxm9tE=; b=tmVLNj//BkT1V572rBwYh4nzC0JWVF/sEjJimL3rH3k9dl/dwJStK7CZ/fiRCBgZMp9bzO aud+zTxZgd/suNvNjjyVoR7ajuxVKX1X4Q2RnFekiVZOXvweHjgWJn4imUfd1q8ca1buqm Q6GxiEAR2fzEJaX0mKPvKlrEwszUVe2LwGk1SpodDLN+rIVnPpblDhMo4rbYV7K8tyRWP1 MBUYjb/Azp3isK01eQe97bCQW7eUbzC8tAebEGOXJnVq4aeTaqbMDfLbKfYFBU5Lfn1JTr kP9tdlGCwau7lqZ7K7AC5zyn3NO1Z29+hH4N1XlGUK6dbg3myoom8yhmocXpsg== Received: from gitrepo.freebsd.org (gitrepo.freebsd.org [IPv6:2610:1c1:1:6068::e6a:5]) by mxrelay.nyi.freebsd.org (Postfix) with ESMTP id 4hXWJx69ymz10p7 for ; Sat, 29 Aug 2026 23:13:37 +0000 (UTC) (envelope-from git@FreeBSD.org) Received: from git (uid 1279) (envelope-from git@FreeBSD.org) id 277f1 by gitrepo.freebsd.org (DragonFly Mail Agent v0.13+ on gitrepo.freebsd.org); Sat, 29 Aug 2026 23:13:37 +0000 To: src-committers@FreeBSD.org, dev-commits-src-all@FreeBSD.org, dev-commits-src-main@FreeBSD.org From: Kevin Bowling Subject: git: 4674cc87a8d0 - main - e1000: Recover from igb(4) controller DEV_RST List-Id: Commit messages for all branches of the src repository List-Archive: https://lists.freebsd.org/archives/dev-commits-src-all List-Help: List-Post: List-Subscribe: List-Unsubscribe: X-BeenThere: dev-commits-src-all@freebsd.org Sender: owner-dev-commits-src-all@FreeBSD.org List-Id: List-Post: List-Help: List-Subscribe: List-Unsubscribe: List-Owner: Precedence: list MIME-Version: 1.0 Content-Type: text/plain; charset=utf-8 Content-Transfer-Encoding: 8bit X-Git-Committer: kbowling X-Git-Repository: src X-Git-Refname: refs/heads/main X-Git-Reftype: branch X-Git-Commit: 4674cc87a8d013b0f2337329c4314cb9fc29e02b Auto-Submitted: auto-generated Date: Sat, 29 Aug 2026 23:13:37 +0000 Message-Id: <6a9367a1.277f1.63c80e82@gitrepo.freebsd.org> The branch main has been updated by kbowling: URL: https://cgit.FreeBSD.org/src/commit/?id=4674cc87a8d013b0f2337329c4314cb9fc29e02b commit 4674cc87a8d013b0f2337329c4314cb9fc29e02b Author: Kevin Bowling AuthorDate: 2026-08-29 02:44:59 +0000 Commit: Kevin Bowling CommitDate: 2026-08-29 23:12:15 +0000 e1000: Recover from igb(4) controller DEV_RST CTRL.DEV_RST resets every port on an 82580 and newer igb device. Hardware reports the event to each affected function through ICR.DRSTA and requires software to reinitialize the port registers and descriptor rings. The driver neither enabled nor handled this cause, so a reset initiated by another function could leave a running interface using stale state. In FreeBSD, we do not currently send this, but other OSes including Linux do, so a PF passed through to such a guest or FreeBSD as a guest running with a passthrough PF on the same controller can be wedged. Enable DRSTA for 82580 and newer PFs in both MSI-X and shared MSI/legacy modes. Bit 30 is reserved on 82575 and is the TCP timer on 82576, so leave it masked on those parts. Latch the event without programming port registers from the interrupt filter. Defer the iflib reset request to admin-task context because the request takes STATE_LOCK. Use IAM to auto-mask shared interrupts on the first ICR read. Keep admin, queue-vector, and interrupt-rearm work quiesced until initialization succeeds. Follow the device reset handshake before the first CTX-owned register programming: wait for GCR to report that the device reset and pending PCIe transactions have completed, then acknowledge STATUS.DEV_RST_SET. Also verify EEPROM autoload and PF reset completion on I350 and newer parts. This follows the Intel 82580 Datasheet, section 4.3.2, and the Intel I350 Datasheet, section 4.3.4. Detect a completed reset from STATUS even when the interface was down. A bounded timeout falls back to the requested port reset. Retain staged reset state through the complete initialization. Check ICR.DRSTA, GCR, and STATUS.DEV_RST_SET after all registers and rings have been rebuilt. If another reset arrived while interrupts were masked, reject the incomplete initialization and repeat the handshake. If both MMIO and PCI configuration space have disappeared, leave the interface stopped rather than queueing an endless recovery loop. Validated on an I210 with INVARIANTS and WITNESS. Injecting CTRL.DEV_RST while igb0 was running recovered through a full iflib initialization in both MSI-X and MSI modes without a panic. STATUS.DEV_RST_SET and GCR.DEV_RST_IN_PROGRESS cleared, enabled DMA coalescing was restored. Injecting the reset while igb0 was down left DEV_RST_SET latched; the first ifconfig up consumed it. On a dual-port 82580, synthetic DRSTA injection produced one complete reinitialization in four-queue MSI-X and shared-MSI modes. Five repeated events produced five clean reinitializations without a panic or watchdog. A raw CTRL.DEV_RST test is not counted because Intel's shared code deliberately avoids that unreliable operation on 82580. On an I350, a real device-wide reset initiated by a sibling function produced one reinitialization while igb0 was running and restored its carrier and enabled state. With igb0 down, STATUS.DEV_RST_SET remained latched until the first up, which consumed it and restored the correct state. The host remained healthy in both cases. MFC after: 2 weeks Sponsored by: BBOX.io --- sys/dev/e1000/e1000_82575.h | 1 + sys/dev/e1000/if_em.c | 299 +++++++++++++++++++++++++++++++++++++++++--- sys/dev/e1000/if_em.h | 1 + 3 files changed, 283 insertions(+), 18 deletions(-) diff --git a/sys/dev/e1000/e1000_82575.h b/sys/dev/e1000/e1000_82575.h index c919a8064476..430eed4202e3 100644 --- a/sys/dev/e1000/e1000_82575.h +++ b/sys/dev/e1000/e1000_82575.h @@ -55,6 +55,7 @@ #define E1000_RAR_ENTRIES_I350 32 #define E1000_SW_SYNCH_MB 0x00000100 #define E1000_STAT_DEV_RST_SET 0x00100000 +#define E1000_GCR_DEV_RST_IN_PROGRESS 0x80000000 struct e1000_adv_data_desc { __le64 buffer_addr; /* Address of the descriptor's data buffer */ diff --git a/sys/dev/e1000/if_em.c b/sys/dev/e1000/if_em.c index 7b1ada08e89e..3acc4b03f34c 100644 --- a/sys/dev/e1000/if_em.c +++ b/sys/dev/e1000/if_em.c @@ -457,6 +457,11 @@ static int igb_if_rx_queue_intr_enable(if_ctx_t, uint16_t); static int igb_if_tx_queue_intr_enable(if_ctx_t, uint16_t); static void em_handle_fatal_error_intr(struct e1000_softc *, u32); static bool em_handle_fatal_error_admin(struct e1000_softc *); +static u32 igb_device_reset_intr_mask(struct e1000_softc *); +static bool igb_device_reset_pending(struct e1000_softc *); +static bool igb_handle_device_reset(struct e1000_softc *, u32); +static void igb_prepare_device_reset(struct e1000_softc *); +static bool igb_finish_device_reset(struct e1000_softc *, u32); static void em_prepare_fatal_error_reset(struct e1000_softc *); static void em_finish_fatal_error_reset(struct e1000_softc *); static void em_configure_peind_memory_errors(struct e1000_softc *); @@ -516,6 +521,15 @@ enum em_fatal_error_state { EM_FATAL_ERROR_RESET_PREPARED, }; +enum igb_device_reset_state { + IGB_DEVICE_RESET_NONE, + IGB_DEVICE_RESET_DETECTED, + IGB_DEVICE_RESET_REQUESTED, + IGB_DEVICE_RESET_PREPARED, +}; + +#define IGB_DEVICE_RESET_TIMEOUT_MS 100 + /* MSI-X handlers */ static int em_if_msix_intr_assign(if_ctx_t, int); static int em_msix_link(void *); @@ -1990,12 +2004,6 @@ em_if_init(if_ctx_t ctx) if (sc->hw.mac.type >= igb_mac_min) igb_initialize_interrupt_rate(sc); - if (!sc->vf_ifp) { - /* Clear pending PF interrupts and request a link check. */ - E1000_READ_REG(&sc->hw, E1000_ICR); - E1000_WRITE_REG(&sc->hw, E1000_ICS, E1000_ICS_LSC); - } - /* AMT based hardware can now take control from firmware */ if (sc->has_manage && sc->has_amt) em_get_hw_control(sc); @@ -2011,8 +2019,24 @@ em_if_init(if_ctx_t ctx) em_configure_peind_memory_errors(sc); em_configure_82575_memory_errors(sc); em_configure_82580_memory_errors(sc); - if (sc->vf_ifp) + if (sc->vf_ifp) { sc->vf_reset_pending = false; + } else { + u32 icr; + + /* + * Drain stale causes only after register reconstruction is + * complete. DRSTA and DEV_RST_SET together close the window in + * which another device reset can arrive while interrupts are + * masked. + */ + icr = E1000_READ_REG(&sc->hw, E1000_ICR); + if (igb_finish_device_reset(sc, icr)) { + iflib_init_failed(ctx); + return; + } + E1000_WRITE_REG(&sc->hw, E1000_ICS, E1000_ICS_LSC); + } } /* @@ -2883,6 +2907,202 @@ em_handle_fatal_error_admin(struct e1000_softc *sc) return (true); } +/* + * ICR bit 30 is reserved on 82575 and is the TCP timer on 82576. It becomes + * the Device Reset Asserted interrupt starting with 82580. + */ +static u32 +igb_device_reset_intr_mask(struct e1000_softc *sc) +{ + + return (sc->hw.mac.type >= e1000_82580 ? E1000_IMS_DRSTA : 0); +} + +/* Keep interrupt-side work quiesced until device-reset recovery completes. */ +static bool +igb_device_reset_pending(struct e1000_softc *sc) +{ + + return (!sc->vf_ifp && igb_device_reset_intr_mask(sc) != 0 && + atomic_load_acq_32(&sc->device_reset_state) != + IGB_DEVICE_RESET_NONE); +} + +/* + * CTRL.DEV_RST resets every port in the device. ICR.DRSTA tells the other + * ports that their registers and descriptor rings must be reinitialized. + */ +static bool +igb_handle_device_reset(struct e1000_softc *sc, u32 icr) +{ + u32 state; + + if (sc->vf_ifp || igb_device_reset_intr_mask(sc) == 0 || + (icr & E1000_ICR_DRSTA) == 0) + return (false); + state = atomic_swap_32(&sc->device_reset_state, + IGB_DEVICE_RESET_DETECTED); + if (state == IGB_DEVICE_RESET_DETECTED) + return (true); + + iflib_admin_intr_deferred(sc->ctx); + return (true); +} + +/* + * A device reset can leave a sibling port accessible before its internal + * reset and PCIe transactions have completed. For 82580 and newer parts, + * wait for that device-wide reset to finish and acknowledge it before any + * ordinary port register programming. I350 and newer parts also publish + * explicit EEPROM autoload and PF-reset completion indications. + * + * The wait is bounded because the only useful fallback for a controller + * that never completes the device reset is the port reset already requested + * by the interrupt handler. + */ +static void +igb_prepare_device_reset(struct e1000_softc *sc) +{ + struct e1000_hw *hw; + u32 state; + u32 eecd, gcr, status; + int i; + + hw = &sc->hw; + state = atomic_load_acq_32(&sc->device_reset_state); + if (state != IGB_DEVICE_RESET_DETECTED && + state != IGB_DEVICE_RESET_REQUESTED && + hw->mac.type >= e1000_82580) { + /* + * A reset can start while this interface has interrupts disabled. + * GCR is the documented gate before ordinary port accesses. STATUS + * also detects a reset that completed while this interface was down + * or after an earlier preparation pass. + */ + gcr = E1000_READ_REG(hw, E1000_GCR); + if (gcr != 0xffffffff && + (gcr & E1000_GCR_DEV_RST_IN_PROGRESS) != 0) { + atomic_store_rel_32(&sc->device_reset_state, + IGB_DEVICE_RESET_DETECTED); + state = IGB_DEVICE_RESET_DETECTED; + } else if (gcr != 0xffffffff) { + status = E1000_READ_REG(hw, E1000_STATUS); + if (status != 0xffffffff && + (status & E1000_STAT_DEV_RST_SET) != 0) { + atomic_store_rel_32(&sc->device_reset_state, + IGB_DEVICE_RESET_DETECTED); + state = IGB_DEVICE_RESET_DETECTED; + } + } + } + if (state != IGB_DEVICE_RESET_DETECTED && + state != IGB_DEVICE_RESET_REQUESTED) + return; + + if (hw->mac.type >= e1000_82580) { + for (i = 0; i < IGB_DEVICE_RESET_TIMEOUT_MS; i++) { + gcr = E1000_READ_REG(hw, E1000_GCR); + if (gcr != 0xffffffff && + (gcr & E1000_GCR_DEV_RST_IN_PROGRESS) == 0) + break; + msec_delay(1); + } + if (i == IGB_DEVICE_RESET_TIMEOUT_MS) { + device_printf(sc->dev, + "device-wide reset did not complete; " + "attempting port reset\n"); + goto prepared; + } + + /* STATUS.DEV_RST_SET is write-one-to-clear. */ + E1000_WRITE_REG(hw, E1000_STATUS, E1000_STAT_DEV_RST_SET); + + if (hw->mac.type >= e1000_i350) { + for (i = 0; i < IGB_DEVICE_RESET_TIMEOUT_MS; i++) { + eecd = E1000_READ_REG(hw, E1000_EECD); + status = E1000_READ_REG(hw, E1000_STATUS); + if (eecd != 0xffffffff && status != 0xffffffff && + (eecd & E1000_EECD_AUTO_RD) != 0 && + (status & E1000_STATUS_RST_DONE) != 0) + break; + msec_delay(1); + } + if (i == IGB_DEVICE_RESET_TIMEOUT_MS) + device_printf(sc->dev, + "device-wide reset did not finish EEPROM " + "autoload or port reset; attempting port " + "reset\n"); + } + } + +prepared: + atomic_store_rel_32(&sc->device_reset_state, + IGB_DEVICE_RESET_PREPARED); +} + +/* + * A second device reset can arrive while the port is being initialized. + * Leave its status latched for the next preparation pass and do not let + * iflib publish this incomplete initialization as a running datapath. + */ +static bool +igb_finish_device_reset(struct e1000_softc *sc, u32 icr) +{ + bool reset_again; + u32 gcr, state, status; + + if (igb_device_reset_intr_mask(sc) == 0) + return (false); + + state = atomic_load_acq_32(&sc->device_reset_state); + reset_again = icr != 0xffffffff && + (icr & E1000_ICR_DRSTA) != 0; + if (sc->hw.mac.type >= e1000_82580) { + gcr = E1000_READ_REG(&sc->hw, E1000_GCR); + if (gcr != 0xffffffff && + (gcr & E1000_GCR_DEV_RST_IN_PROGRESS) != 0) + reset_again = true; + status = E1000_READ_REG(&sc->hw, E1000_STATUS); + if (status == 0xffffffff && + state != IGB_DEVICE_RESET_NONE) { + /* + * MMIO can disappear briefly while SR-IOV is changing, but + * config space remains readable. If both are gone, retain the + * stopped state without queueing an endless reset loop. + */ + if (pci_read_config(sc->dev, PCIR_VENDOR, 2) == 0xffff) { + atomic_store_rel_32(&sc->device_reset_state, + IGB_DEVICE_RESET_DETECTED); + device_printf(sc->dev, + "device unavailable after device-wide reset; " + "leaving interface stopped\n"); + return (true); + } + reset_again = true; + } else if (status != 0xffffffff && + (status & E1000_STAT_DEV_RST_SET) != 0) + reset_again = true; + } + if (state == IGB_DEVICE_RESET_DETECTED || + state == IGB_DEVICE_RESET_REQUESTED) + reset_again = true; + if (!reset_again) { + if (state == IGB_DEVICE_RESET_PREPARED && + !atomic_cmpset_rel_32(&sc->device_reset_state, + IGB_DEVICE_RESET_PREPARED, IGB_DEVICE_RESET_NONE)) + return (true); + return (false); + } + + state = atomic_swap_32(&sc->device_reset_state, + IGB_DEVICE_RESET_DETECTED); + if (state != IGB_DEVICE_RESET_DETECTED) { + iflib_request_reset_if_up(sc->ctx); + iflib_admin_intr_deferred(sc->ctx); + } + return (true); +} + /* * A PCIe-region parity failure stops PCIe and DMA traffic. I350, I354, * I210, and I211 require a port reset before master disable in this case. @@ -3086,14 +3306,18 @@ em_intr(void *arg) if (hw->mac.type >= e1000_82571 && (reg_icr & E1000_ICR_INT_ASSERTED) == 0) return FILTER_STRAY; + if (igb_handle_device_reset(sc, reg_icr)) + return (FILTER_HANDLED); + if (igb_device_reset_pending(sc)) + return (FILTER_HANDLED); /* - * Only MSI-X interrupts have one-shot behavior by taking advantage - * of the EIAC register. Thus, explicitly disable interrupts. This - * also works around the MSI message reordering errata on certain - * systems. + * IAM auto-masks igb shared interrupts when ICR is read. Older em + * hardware still needs an explicit disable, which also works around + * MSI message reordering errata on certain systems. */ - IFDI_INTR_DISABLE(ctx); + if (sc->vf_ifp || hw->mac.type < igb_mac_min) + IFDI_INTR_DISABLE(ctx); /* Link status change */ if (reg_icr & (E1000_ICR_RXSEQ | E1000_ICR_LSC)) @@ -3136,6 +3360,8 @@ igb_if_rx_queue_intr_enable(if_ctx_t ctx, uint16_t rxqid) struct e1000_softc *sc = iflib_get_softc(ctx); struct em_rx_queue *rxq = &sc->rx_queues[rxqid]; + if (igb_device_reset_pending(sc)) + return (0); E1000_WRITE_REG(&sc->hw, E1000_EIMS, rxq->eims); return (0); } @@ -3146,6 +3372,8 @@ igb_if_tx_queue_intr_enable(if_ctx_t ctx, uint16_t txqid) struct e1000_softc *sc = iflib_get_softc(ctx); struct em_tx_queue *txq = &sc->tx_queues[txqid]; + if (igb_device_reset_pending(sc)) + return (0); E1000_WRITE_REG(&sc->hw, E1000_EIMS, txq->eims); return (0); } @@ -3164,6 +3392,8 @@ em_msix_que(void *arg) ++que->irqs; + if (igb_device_reset_pending(sc)) + return (FILTER_HANDLED); em_newitr(sc, que, rxr); return (FILTER_SCHEDULE_THREAD); @@ -3195,6 +3425,8 @@ em_msix_link(void *arg) } reg_icr = E1000_READ_REG(&sc->hw, E1000_ICR); + if (igb_device_reset_pending(sc)) + return (FILTER_HANDLED); /* * Enabling or disabling SR-IOV can briefly make PF MMIO reads return @@ -3203,6 +3435,8 @@ em_msix_link(void *arg) */ if (__predict_false(reg_icr == 0xffffffff)) goto rearm; + if (igb_handle_device_reset(sc, reg_icr)) + return (FILTER_HANDLED); if (reg_icr & E1000_ICR_RXO) sc->rx_overruns++; @@ -3219,7 +3453,8 @@ rearm: /* Re-arm unconditionally */ if (sc->hw.mac.type >= igb_mac_min) { E1000_WRITE_REG(&sc->hw, E1000_IMS, - E1000_IMS_LSC | igb_iov_intr_mask(sc) | + E1000_IMS_LSC | igb_device_reset_intr_mask(sc) | + igb_iov_intr_mask(sc) | em_fatal_error_intr_mask(sc)); E1000_WRITE_REG(&sc->hw, E1000_EIMS, sc->link_mask); } else if (sc->hw.mac.type == e1000_82574) { @@ -3565,6 +3800,23 @@ em_if_update_admin_status(if_ctx_t ctx) KASSERT(!sc->vf_ifp, ("%s called for a VF", __func__)); if (em_handle_fatal_error_admin(sc)) return; + /* A sibling-port reset invalidated the registers and VF mailboxes. */ + if (atomic_cmpset_acq_32(&sc->device_reset_state, + IGB_DEVICE_RESET_DETECTED, IGB_DEVICE_RESET_REQUESTED)) { + if (sc->link_state == EM_LINK_STATE_UP) + iflib_link_state_change(ctx, LINK_STATE_DOWN, 0); + sc->link_speed = 0; + sc->link_duplex = 0; + sc->link_state = EM_LINK_STATE_DOWN_RESET_PENDING; + /* Request the reset here; interrupt filters cannot take STATE_LOCK. */ + iflib_request_reset_if_up(ctx); + /* Re-enter the admin task so it observes the reset request. */ + iflib_admin_intr_deferred(ctx); + return; + } + if (atomic_load_acq_32(&sc->device_reset_state) != + IGB_DEVICE_RESET_NONE) + return; if (atomic_readandclear_32(&sc->promisc_pending) != 0) (void)em_if_set_promisc_impl(ctx, @@ -6007,6 +6259,8 @@ igb_if_intr_enable(if_ctx_t ctx) struct e1000_hw *hw = &sc->hw; u32 mask, reg; + if (igb_device_reset_pending(sc)) + return; if (__predict_true(sc->intr_type == IFLIB_INTR_MSIX)) { mask = (sc->que_mask | sc->link_mask); /* @@ -6020,11 +6274,16 @@ igb_if_intr_enable(if_ctx_t ctx) igb_iov_intr_drain_stale(sc); E1000_WRITE_REG(hw, E1000_EIMS, mask); E1000_WRITE_REG(hw, E1000_IMS, - E1000_IMS_LSC | igb_iov_intr_mask(sc) | + E1000_IMS_LSC | igb_device_reset_intr_mask(sc) | + igb_iov_intr_mask(sc) | em_fatal_error_intr_mask(sc)); - } else - E1000_WRITE_REG(hw, E1000_IMS, - IMS_ENABLE_MASK | em_fatal_error_intr_mask(sc)); + } else { + mask = IMS_ENABLE_MASK | igb_device_reset_intr_mask(sc) | + em_fatal_error_intr_mask(sc); + /* Reading ICR masks every shared interrupt before the filter runs. */ + E1000_WRITE_REG(hw, E1000_IAM, mask); + E1000_WRITE_REG(hw, E1000_IMS, mask); + } E1000_WRITE_FLUSH(hw); } @@ -6035,6 +6294,9 @@ igb_if_intr_disable(if_ctx_t ctx) struct e1000_hw *hw = &sc->hw; u32 mask, reg; + /* This is the first CTX-owned register access after ICR.DRSTA. */ + igb_prepare_device_reset(sc); + if (__predict_true(sc->intr_type == IFLIB_INTR_MSIX)) { /* * Do not use a blanket EIMC write here. VF interrupt controls @@ -6049,7 +6311,8 @@ igb_if_intr_disable(if_ctx_t ctx) E1000_WRITE_REG(hw, E1000_EIMC, mask); reg = E1000_READ_REG(hw, E1000_EIAC); E1000_WRITE_REG(hw, E1000_EIAC, reg & ~mask); - } + } else + E1000_WRITE_REG(hw, E1000_IAM, 0); E1000_WRITE_REG(hw, E1000_IMC, 0xffffffff); E1000_WRITE_FLUSH(hw); } diff --git a/sys/dev/e1000/if_em.h b/sys/dev/e1000/if_em.h index 5e72be998078..e6ece2a9a689 100644 --- a/sys/dev/e1000/if_em.h +++ b/sys/dev/e1000/if_em.h @@ -625,6 +625,7 @@ struct e1000_softc { u32 phy_hang_count; u32 promisc_pending; u32 stats_pending; + u32 device_reset_state; u32 fatal_error_state; u32 fatal_error_icr; u32 fatal_error_pbeccsts;