From nobody Mon Sep 14 20:22:52 2026 X-Original-To: dev-commits-src-main@mlmmj.nyi.freebsd.org Received: from mx1.freebsd.org (mx1.freebsd.org [IPv6:2610:1c1:1:606c::19:1]) by mlmmj.nyi.freebsd.org (Postfix) with ESMTP id 4hkGmX3mF7z6rgZw for ; Mon, 14 Sep 2026 20:22:52 +0000 (UTC) (envelope-from git@FreeBSD.org) Received: from mxrelay.nyi.freebsd.org (mxrelay.nyi.freebsd.org [IPv6:2610:1c1:1:606c::19:3]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange x25519 server-signature RSA-PSS (4096 bits) server-digest SHA256 client-signature RSA-PSS (4096 bits) client-digest SHA256) (Client CN "mxrelay.nyi.freebsd.org", Issuer "YR2" (not verified)) by mx1.freebsd.org (Postfix) with ESMTPS id 4hkGmX3WKCz4hBW for ; Mon, 14 Sep 2026 20:22:52 +0000 (UTC) (envelope-from git@FreeBSD.org) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=freebsd.org; s=dkim; t=1789417372; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding; bh=R+r0QWNKu28MDLuffpRu6neTz4GndSVizucQ5/MlhpU=; b=CB+RGAdP5P+zRDw6naCYEsSFm1RNtzXWNaOoJxhC9wo09d+JSDLm5FkaYnBofIGFgHFRnL 2riML3QfkaruzOGS/vnqCICBGDiIx3zVdFvKMkbXh1dcbLoT3K1FE1O4Ify3uipmrK+wse oEXK/MNBPbYPwZ35hMW3P4emuazxgq/s7lWUP7KsvnegqcJkZmrhLymCusY00USnuhXh7Y T5jO171TiwW1Eh8h7RnyVF7AMGHWpgXA6R/hkG61QhmtXV1g9eygKbZNgsreLtGSlo1Jna LgbTxXzF3SvQBBzv12jwKknd/kIxSbdQYPqeQ7273Cp5nZIc6ucpOs//3j/rXg== ARC-Seal: i=1; a=rsa-sha256; d=freebsd.org; s=dkim; cv=none; t=1789417372; b=t6EPhJ4185vOD+ZL99FtiYofnKB61VYvZHLpqVGhSnD3moitKPyqgyCKfwjHsL0VdSD+QG wxS3ykccsFlsMrUITZDWErtIype56NxeIE8Wt169GbLEO6R3B5Q1ifzAmoTUiDMNHeXjyd t7aC3S0P/NSV1Q8bJsysYliy6qb3Qn9bJqN2dr3IclmQDWZlP3B9hWWjPyblan/YZbGiUu xNHw6AJg8P5mVm6XHE2PHhfXc/MFDibE1kavqFN34YuKDLCaV+8mmW3ORp3M7pEcR275FS +KDFSbS7YHL+Zmiy0CbnEjOwe9XAeLwRMNSJocufPX+UPGdJdTI16rIujGAkJA== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=freebsd.org; s=dkim; t=1789417372; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding; bh=R+r0QWNKu28MDLuffpRu6neTz4GndSVizucQ5/MlhpU=; b=wjIIiiy1RLxN5bmaNKEsmeWSQiWrGpdQ0krxnjsaMEhkdnwENvS/x/hHh1Y+1Z9djer2d2 BLgUKund1uGAqPO5xrFQ2ZAVgGcBkJhcEIzHLvdSoDRMCwxpRpbCGDqqFXxI/IzqWbUXcn 3pG1dEgf9/R934ScTHRlwElVfSJmVqLu9ntXdx3AJKOjPq57g4ZFK0yTMg3IMWoyJi/gTo 4RvXWNPA1UlZB0Z83G2orxHrEm6Ft35cgxptHnfZ2oV7Wgw3cwNTpjxY8Fee3AFyxkK8My keAtzKOvJ6ceY+ajcXLx38irONb7xeFGFVbFxe1Vst6GoBHnLzraX4ViTZpBEg== ARC-Authentication-Results: i=1; mx1.freebsd.org; none Received: from gitrepo.freebsd.org (gitrepo.freebsd.org [IPv6:2610:1c1:1:6068::e6a:5]) by mxrelay.nyi.freebsd.org (Postfix) with ESMTP id 4hkGmX2YDqzSql for ; Mon, 14 Sep 2026 20:22:52 +0000 (UTC) (envelope-from git@FreeBSD.org) Received: from git (uid 1279) (envelope-from git@FreeBSD.org) id 317dc by gitrepo.freebsd.org (DragonFly Mail Agent v0.13+ on gitrepo.freebsd.org); Mon, 14 Sep 2026 20:22:52 +0000 To: src-committers@FreeBSD.org, dev-commits-src-all@FreeBSD.org, dev-commits-src-main@FreeBSD.org From: Andrew Gallatin Subject: git: 7e2a425f7ea3 - main - iflib: Use a bounded buf_ring for simple_tx List-Id: Commit messages for the main branch of the src repository List-Archive: https://lists.freebsd.org/archives/dev-commits-src-main List-Help: List-Post: List-Subscribe: List-Unsubscribe: X-BeenThere: dev-commits-src-main@freebsd.org Sender: owner-dev-commits-src-main@FreeBSD.org List-Id: List-Post: List-Help: List-Subscribe: List-Unsubscribe: List-Owner: Precedence: list MIME-Version: 1.0 Content-Type: text/plain; charset=utf-8 Content-Transfer-Encoding: 8bit X-Git-Committer: gallatin X-Git-Repository: src X-Git-Refname: refs/heads/main X-Git-Reftype: branch X-Git-Commit: 7e2a425f7ea31c422569887b6189651b0ad4f89f Auto-Submitted: auto-generated Date: Mon, 14 Sep 2026 20:22:52 +0000 Message-Id: <6aa8579c.317dc.1b2dbef4@gitrepo.freebsd.org> The branch main has been updated by gallatin: URL: https://cgit.FreeBSD.org/src/commit/?id=7e2a425f7ea31c422569887b6189651b0ad4f89f commit 7e2a425f7ea31c422569887b6189651b0ad4f89f Author: Andrew Gallatin AuthorDate: 2026-09-14 19:14:52 +0000 Commit: Andrew Gallatin CommitDate: 2026-09-14 20:18:40 +0000 iflib: Use a bounded buf_ring for simple_tx Implement buf_ring/drbr deferred transmit in iflib. This is intended to allow the new simpler code path to replace mp_ring. This patch makes the simple_tx outperform mp_ring by a wide margin when CPU is the bottleneck (eg, cannot fill the NIC). See graphs at: https://people.freebsd.org/~gallatin/mpring_vs_simple_tx Note that the buf ring is used for contention, not capacity. Eg, it is used as a place for contending threads to put packets without waiting for a mutex. It is not designed to act as a software ring on top of the hardware descriptors provided by the underlying NIC driver. "stranded packets" are exceedingly rare due to the fact that if there is enough load to use the buf_ring, there will probably be more load coming that can be a drainer. Not scheduling a gtask to drain is intentional, and we really on the timer as a fallback. One thing I noticed while developing this patch is that a simple mutex with no deferral generally outperformed both mp_ring and drbr at high levels of contention for the same queue. This is inherent in a bounded MPSC queue where multiple producers are contending on claiming ring entries. So I came up with the idea of bounding the number of producers such that the deferral ring would devolve to a mutex when contention was high. Identifying the crossover point in a general way was hard. On different machines, the point between a mutex and a deferral ring was very different and also depended on the placement of the producers. I eventually realized that on the large AMD EPYC servers that I was testing on, the crossover point generally coincided with the work spilling into another CCX. So I developed an approach where we limit the number of producers by limiting simultanious producers to the same L3 AND by bounding the number of simultanious producers. This limit can be adjusted via the sysctl net.iflib.max_producers. The number of packets drained from the deferral ring is limited by net.iflib.simple_drain_quota. The intent is to process just enough packets in the gtaskq context so as to make space to allow threads to make progress. The task also uses a trylock so as to avoid blocking, waiting for a thread to drain. This change also enables ALTQ support for simple tx. Reviewed by: kbowling Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D58901 --- share/man/man4/iflib.4 | 51 ++++- sys/net/iflib.c | 611 +++++++++++++++++++++++++++++++++++++++++++++---- 2 files changed, 619 insertions(+), 43 deletions(-) diff --git a/share/man/man4/iflib.4 b/share/man/man4/iflib.4 index fb8bb37413c7..ca1b7918412e 100644 --- a/share/man/man4/iflib.4 +++ b/share/man/man4/iflib.4 @@ -1,4 +1,4 @@ -.Dd August 26, 2026 +.Dd September 14, 2026 .Dt IFLIB 4 .Os .Sh NAME @@ -65,7 +65,11 @@ core. When set to a non-zero value, TX queues are assigned to cores following the last RX queue. .It Va simple_tx -When set to one, iflib uses a simple transmit routine with no queuing at all. +When set to one, iflib uses a simple transmit routine that sends packets +directly while the transmit queue mutex is available. +During contention, packets are deferred through a multi-producer +.Xr buf_ring 9 +and drained by the mutex holder. By default, iflib uses a highly optimized, lockless, transmit queue called mp_ring. This performs well when there are more CPU cores than NIC @@ -116,8 +120,20 @@ Zero (the default) indicates the default (currently 16) should be used. .El .Pp There are also some global sysctls which can change behaviour for all drivers, -and may be changed at any time. +and, unless noted otherwise, may be changed at any time. .Bl -tag -width indent +.It Va net.iflib.max_producers +Limits the number of callers that may concurrently enqueue packets to a +simple transmit deferral ring without taking the transmit queue mutex. +The limit is applied separately to each transmit queue. +If lockless producers from one last-level cache are active, a producer from +another last-level cache waits for the transmit queue mutex regardless of the +limit. +Setting this to zero forces all contending producers to wait for the mutex. +The default is 8. +This setting applies only when +.Va simple_tx +is enabled. .It Va net.iflib.min_tx_latency If this is set to a non-zero value, iflib will avoid any attempt to combine multiple transmits, and notify the hardware as quickly as possible of @@ -128,6 +144,35 @@ Some NICs allow processing completed transmit descriptors in batches. Doing so usually increases the transmit throughput by reducing the number of transmit interrupts. Setting this to a non-zero value will disable the use of this feature. +.It Va net.iflib.simple_drain_quota +Sets the maximum number of packets that the transmit task may remove from a +simple transmit deferral ring in one drain operation. +The default is 8. +This setting applies only when +.Va simple_tx +is enabled. +.It Va net.iflib.simple_drain_quota_thread +Sets the maximum number of packets that a transmitting thread holding the +transmit queue mutex may remove from a simple transmit deferral ring in one +drain operation. +A smaller value bounds the time for which a transmitting thread can hold the +mutex while other producers replenish the ring. +The default is 65536. +This setting applies only when +.Va simple_tx +is enabled. +.It Va net.iflib.simple_txbr_size +Sets the number of entries in each +.Xr buf_ring 9 +used by the simple transmit path to defer packets. +The value must be a power of two and at least 64. +Invalid values are replaced with the default of 1024. +This is a read-only tunable that must be set before +.Nm +is loaded. +This setting applies only when +.Va simple_tx +is enabled. .It Va net.iflib.tx_watchdog_periods Number of consecutive .Va net.iflib.timer_default diff --git a/sys/net/iflib.c b/sys/net/iflib.c index d3d25bc13f85..9979d17777ef 100644 --- a/sys/net/iflib.c +++ b/sys/net/iflib.c @@ -33,6 +33,7 @@ #include #include #include +#include #include #include #include @@ -40,6 +41,7 @@ #include #include #include +#include #include #include #include @@ -132,7 +134,20 @@ static MALLOC_DEFINE(M_IFLIB, "iflib", "ifnet library"); #define IFLIB_RXEOF_MORE (1U << 0) #define IFLIB_RXEOF_EMPTY (2U << 0) - +#define IFLIB_TXQ_QUIESCING (1U << 31) +#define IFLIB_TXQ_PRODUCER_LLC_SHIFT 20 +#define IFLIB_TXQ_PRODUCER_MAX \ + ((1U << IFLIB_TXQ_PRODUCER_LLC_SHIFT) - 1) +#define IFLIB_TXQ_PRODUCER(_state) \ + ((_state) & IFLIB_TXQ_PRODUCER_MAX) +#define IFLIB_TXQ_PRODUCER_LLC_MASK \ + (~(IFLIB_TXQ_QUIESCING | IFLIB_TXQ_PRODUCER_MAX)) +#define IFLIB_TXQ_PRODUCER_LLC(_state) \ + (((_state) & IFLIB_TXQ_PRODUCER_LLC_MASK) >> \ + IFLIB_TXQ_PRODUCER_LLC_SHIFT) + +CTASSERT(MAXCPU - 1 <= (IFLIB_TXQ_PRODUCER_LLC_MASK >> + IFLIB_TXQ_PRODUCER_LLC_SHIFT)); struct iflib_txq; typedef struct iflib_txq *iflib_txq_t; struct iflib_rxq; @@ -177,9 +192,9 @@ enum iflib_pm_state { static void iru_init(if_rxd_update_t iru, iflib_rxq_t rxq, uint8_t flid); static void iflib_timer(void *arg); static void iflib_tqg_detach(if_ctx_t ctx); -#ifndef ALTQ static int iflib_simple_transmit(if_t ifp, struct mbuf *m); -#endif +static void iflib_simple_if_start(if_t ifp); +static void iflib_simple_txq_drain(iflib_txq_t txq); typedef struct iflib_filter_info { driver_filter_t *ifi_filter; @@ -423,6 +438,15 @@ struct iflib_txq { uint64_t ift_txd_encap_efbig; uint64_t ift_pullups; uint64_t ift_last_timer_tick; + uint64_t ift_drbr_direct; + uint64_t ift_drbr_stall; + counter_u64_t ift_drbr_deferred; + counter_u64_t ift_drbr_drops; + counter_u64_t ift_drbr_blocked; + counter_u64_t ift_drbr_remote; + + /* Lockless producer count and current llc. */ + volatile u_int ift_producers __aligned(CACHE_LINE_SIZE); struct mtx ift_mtx; struct mtx ift_db_mtx; @@ -430,6 +454,7 @@ struct iflib_txq { /* constant values */ if_ctx_t ift_ctx; struct ifmp_ring *ift_br; + struct buf_ring *ift_drbr; struct grouptask ift_task; qidx_t ift_size; qidx_t ift_pad; @@ -648,6 +673,48 @@ SYSCTL_INT(_net_iflib, OID_AUTO, tx_watchdog_periods, CTLFLAG_RWTUN, "consecutive frozen timer periods under demand before a TX queue is " "checked for a hang (0 disables the check)"); +#define IFLIB_SIMPLE_TXBR_MIN 64 +#define IFLIB_SIMPLE_TXBR_SIZE 1024 +static int iflib_simple_txbr_size = IFLIB_SIMPLE_TXBR_SIZE; +SYSCTL_INT(_net_iflib, OID_AUTO, simple_txbr_size, CTLFLAG_RDTUN, + &iflib_simple_txbr_size, 0, + "number of entries in the simple tx deferral ring"); +static u_int iflib_simple_drain_quota = 8; +SYSCTL_UINT(_net_iflib, OID_AUTO, simple_drain_quota, CTLFLAG_RWTUN, + &iflib_simple_drain_quota, 0, + "maximum packets sent per simple tx deferral ring drain from the tx task"); +static u_int iflib_simple_drain_quota_thread = 65536; +SYSCTL_UINT(_net_iflib, OID_AUTO, simple_drain_quota_thread, CTLFLAG_RWTUN, + &iflib_simple_drain_quota_thread, 0, + "maximum packets sent per simple tx deferral ring drain from another thread"); +static u_int iflib_max_producers = 8; +static bool iflib_single_llc __read_mostly; +static bool iflib_producer_gate __read_mostly; + +static int +iflib_sysctl_max_producers(SYSCTL_HANDLER_ARGS) +{ + u_int max_producers; + int error; + + max_producers = iflib_max_producers; + error = sysctl_handle_int(oidp, &max_producers, 0, req); + if (error != 0 || req->newptr == NULL) + return (error); + + max_producers = MIN(max_producers, IFLIB_TXQ_PRODUCER_MAX); + iflib_max_producers = max_producers; + if (iflib_single_llc) + iflib_producer_gate = mp_ncpus > max_producers; + return (0); +} +SYSCTL_PROC(_net_iflib, OID_AUTO, max_producers, + CTLTYPE_UINT | CTLFLAG_RWTUN | CTLFLAG_MPSAFE, &iflib_max_producers, 0, + iflib_sysctl_max_producers, "IU", + "maximum concurrent lockless producers per transmit queue"); + +/* Encoded llc for each CPU. */ +static uint16_t iflib_cpu_llc[MAXCPU] __read_mostly; #if IFLIB_DEBUG_COUNTERS @@ -1919,6 +1986,31 @@ iflib_txq_destroy(iflib_txq_t txq) txq->ift_br = NULL; } + /* Free any mbufs stranded in the deferral ring */ + if (txq->ift_drbr != NULL) { + mtx_lock(&txq->ift_mtx); + drbr_flush(NULL, txq->ift_drbr); + mtx_unlock(&txq->ift_mtx); + buf_ring_free(txq->ift_drbr, M_IFLIB); + txq->ift_drbr = NULL; + } + if (txq->ift_drbr_deferred != NULL) { + counter_u64_free(txq->ift_drbr_deferred); + txq->ift_drbr_deferred = NULL; + } + if (txq->ift_drbr_blocked != NULL) { + counter_u64_free(txq->ift_drbr_blocked); + txq->ift_drbr_blocked = NULL; + } + if (txq->ift_drbr_remote != NULL) { + counter_u64_free(txq->ift_drbr_remote); + txq->ift_drbr_remote = NULL; + } + if (txq->ift_drbr_drops != NULL) { + counter_u64_free(txq->ift_drbr_drops); + txq->ift_drbr_drops = NULL; + } + mtx_destroy(&txq->ift_mtx); if (txq->ift_sds.ifsd_map != NULL) { @@ -2570,8 +2662,11 @@ iflib_timer(void *arg) txq->ift_outstanding_prev = outstanding; txq->ift_processed_prev = txq->ift_processed; } - /* handle any laggards */ - if (txq->ift_db_pending) + /* Handle any laggards */ + if (txq->ift_db_pending || + (txq->ift_drbr != NULL && + (!if_altq_is_enabled(ctx->ifc_ifp) || txq->ift_id == 0) && + !drbr_empty(ctx->ifc_ifp, txq->ift_drbr))) GROUPTASK_ENQUEUE(&txq->ift_task); sctx->isc_pause_frames = 0; @@ -2660,7 +2755,6 @@ iflib_init_locked(if_ctx_t ctx) CALLOUT_UNLOCK(txq); (void)iflib_netmap_txq_init(ctx, txq); } - /* * Calculate a suitable Rx mbuf size prior to calling IFDI_INIT, so * that drivers can use the value when setting up the hardware receive @@ -2711,9 +2805,14 @@ iflib_init_locked(if_ctx_t ctx) if_setdrvflagbits(ctx->ifc_ifp, IFF_DRV_RUNNING, IFF_DRV_OACTIVE); IFDI_INTR_ENABLE(ctx); txq = ctx->ifc_txqs; - for (i = 0; i < scctx->isc_ntxqsets; i++, txq++) + for (i = 0; i < scctx->isc_ntxqsets; i++, txq++) { callout_reset_on(&txq->ift_timer, iflib_timer_default, iflib_timer, txq, txq->ift_timer.c_cpu); + if (ctx->ifc_sysctl_simple_tx) { + atomic_clear_rel_int(&txq->ift_producers, + IFLIB_TXQ_QUIESCING); + } + } /* Re-enable txsync/rxsync. */ netmap_enable_all_rings(ifp); @@ -2788,6 +2887,14 @@ iflib_stop(if_ctx_t ctx) ("recursive iflib stop")); stop_hardware = ctx->ifc_datapath_state != IFLIB_DP_STOPPED; + if (ctx->ifc_sysctl_simple_tx && stop_hardware) { + /* close deferral rings to new traffic */ + for (i = 0; i < scctx->isc_ntxqsets; i++) { + atomic_set_int(&txq[i].ift_producers, + IFLIB_TXQ_QUIESCING); + } + } + /* Tell the stack that the interface is no longer active */ if_setdrvflagbits(ctx->ifc_ifp, IFF_DRV_OACTIVE, IFF_DRV_RUNNING); @@ -2819,9 +2926,13 @@ iflib_stop(if_ctx_t ctx) #endif /* DEV_NETMAP */ CALLOUT_UNLOCK(txq); + /* clean any enqueued buffers */ if (!ctx->ifc_sysctl_simple_tx) { - /* clean any enqueued buffers */ iflib_ifmp_purge(txq); + } else { + mtx_lock(&txq->ift_mtx); + drbr_flush(ctx->ifc_ifp, txq->ift_drbr); + mtx_unlock(&txq->ift_mtx); } /* Free any existing tx buffers. */ for (j = 0; j < txq->ift_size; j++) { @@ -2842,6 +2953,13 @@ iflib_stop(if_ctx_t ctx) txq->ift_closed = txq->ift_mbuf_defrag = txq->ift_mbuf_defrag_failed = 0; txq->ift_no_tx_dma_setup = txq->ift_txd_encap_efbig = txq->ift_map_failed = 0; txq->ift_pullups = 0; + txq->ift_drbr_direct = txq->ift_drbr_stall = 0; + if (ctx->ifc_sysctl_simple_tx) { + counter_u64_zero(txq->ift_drbr_deferred); + counter_u64_zero(txq->ift_drbr_drops); + counter_u64_zero(txq->ift_drbr_blocked); + counter_u64_zero(txq->ift_drbr_remote); + } ifmp_ring_reset_stats(txq->ift_br); for (j = 0, di = txq->ift_ifdi; j < sctx->isc_ntxqs; j++, di++) bzero((void *)di->idi_vaddr, di->idi_size); @@ -4017,7 +4135,7 @@ iflib_completed_tx_reclaim(iflib_txq_t txq, struct mbuf **m_defer) int reclaim; reclaim = iflib_txq_can_reclaim(txq); - if (reclaim == 0) + if (reclaim <= 0) return (0); _iflib_completed_tx_reclaim(txq, m_defer, reclaim); return (reclaim); @@ -4243,12 +4361,10 @@ _task_fn_tx(void *context) netmap_tx_irq(ifp, txq->ift_id)) goto skip_ifmp; #endif - if (ctx->ifc_sysctl_simple_tx) { - mtx_lock(&txq->ift_mtx); - (void)iflib_completed_tx_reclaim(txq, NULL); - mtx_unlock(&txq->ift_mtx); - goto skip_ifmp; - } + if (ctx->ifc_sysctl_simple_tx) { + iflib_simple_txq_drain(txq); + goto skip_ifmp; + } #ifdef ALTQ if (if_altq_is_enabled(ifp)) iflib_altq_if_start(ifp); @@ -4592,13 +4708,18 @@ iflib_altq_if_start(if_t ifp) static int iflib_altq_if_transmit(if_t ifp, struct mbuf *m) { + if_ctx_t ctx = if_getsoftc(ifp); int err; if (if_altq_is_enabled(ifp)) { IFQ_ENQUEUE(&ifp->if_snd, m, err); /* XXX - DRVAPI */ if (err == 0) - iflib_altq_if_start(ifp); - } else + if_start(ifp); + return (err); + } + if (ctx->ifc_sysctl_simple_tx) + err = iflib_simple_transmit(ifp, m); + else err = iflib_if_transmit(ifp, m); return (err); @@ -4615,9 +4736,16 @@ iflib_if_qflush(if_t ifp) STATE_LOCK(ctx); ctx->ifc_flags |= IFC_QFLUSH; STATE_UNLOCK(ctx); - for (i = 0; i < NTXQSETS(ctx); i++, txq++) + for (i = 0; i < NTXQSETS(ctx); i++, txq++) { + if (txq->ift_drbr != NULL) { + mtx_lock(&txq->ift_mtx); + drbr_flush(ifp, txq->ift_drbr); + mtx_unlock(&txq->ift_mtx); + continue; + } while (!(ifmp_ring_is_idle(txq->ift_br) || ifmp_ring_is_stalled(txq->ift_br))) iflib_txq_check_drain(txq, 0); + } STATE_LOCK(ctx); ctx->ifc_flags &= ~IFC_QFLUSH; STATE_UNLOCK(ctx); @@ -5416,13 +5544,12 @@ iflib_device_register(device_t dev, void *sc, if_shared_ctx_t sctx, if_ctx_t *ct scctx = &ctx->ifc_softc_ctx; ifp = ctx->ifc_ifp; if (ctx->ifc_sysctl_simple_tx) { + /* if_start drives the same drbr drain when ALTQ is active. */ #ifndef ALTQ if_settransmitfn(ifp, iflib_simple_transmit); - device_printf(dev, "using simple if_transmit\n"); -#else - device_printf(dev, "ALTQ prevents using simple if_transmit\n"); - ctx->ifc_sysctl_simple_tx = 0; #endif + if_setstartfn(ifp, iflib_simple_if_start); + device_printf(dev, "using simple transmit\n"); } iflib_reset_qvalues(ctx); CTX_LOCK(ctx); @@ -6126,6 +6253,49 @@ iflib_device_iov_add_vf(device_t dev, uint16_t vfnum, const nvlist_t *params) * **********************************************************************/ +static void +iflib_cpu_llc_init(void) +{ +#ifdef SMP + struct cpu_group *cg, *llc; + uint16_t llc_id, first_llc_id; + int cpu; +#endif + + iflib_single_llc = true; +#ifdef SMP + first_llc_id = USHRT_MAX; + for (cpu = 0; cpu <= mp_maxid; cpu++) { + if (CPU_ABSENT(cpu)) + continue; + + /* + * Select the outermost shared-cache group containing this + * CPU. On AMD systems this is the L3/CCX group. This also + * gives sensible behavior when the llc is not L3. + */ + llc = NULL; + for (cg = smp_topo_find(cpu_top, cpu); cg != NULL; + cg = cg->cg_parent) { + if (cg->cg_level != CG_SHARE_NONE) + llc = cg; + } + + /* cg_first is a stable llc identifier. */ + llc_id = llc != NULL ? llc->cg_first : cpu; + iflib_cpu_llc[cpu] = llc_id; + if (first_llc_id == USHRT_MAX) + first_llc_id = llc_id; + else if (llc_id != first_llc_id) + iflib_single_llc = false; + } +#else + iflib_cpu_llc[0] = 0; +#endif + iflib_producer_gate = !iflib_single_llc || + mp_ncpus > iflib_max_producers; +} + /* * - Start a fast taskqueue thread for each core * - Start a taskqueue for control operations @@ -6134,6 +6304,15 @@ static int iflib_module_init(void) { iflib_timer_default = hz / 2; + iflib_cpu_llc_init(); + + if (iflib_simple_txbr_size < IFLIB_SIMPLE_TXBR_MIN || + !powerof2(iflib_simple_txbr_size)) { + printf("iflib: simple_txbr_size %d is not a power of 2 >= %d " + "- using default value of %d\n", iflib_simple_txbr_size, + IFLIB_SIMPLE_TXBR_MIN, IFLIB_SIMPLE_TXBR_SIZE); + iflib_simple_txbr_size = IFLIB_SIMPLE_TXBR_SIZE; + } return (0); } @@ -6403,6 +6582,14 @@ iflib_queues_alloc(if_ctx_t ctx) device_printf(dev, "Unable to allocate buf_ring\n"); goto err_tx_desc; } + if (ctx->ifc_sysctl_simple_tx) { + txq->ift_drbr = buf_ring_alloc(iflib_simple_txbr_size, + M_IFLIB, M_WAITOK, &txq->ift_mtx); + txq->ift_drbr_deferred = counter_u64_alloc(M_WAITOK); + txq->ift_drbr_drops = counter_u64_alloc(M_WAITOK); + txq->ift_drbr_blocked = counter_u64_alloc(M_WAITOK); + txq->ift_drbr_remote = counter_u64_alloc(M_WAITOK); + } txq->ift_reclaim_thresh = ctx->ifc_sysctl_tx_reclaim_thresh; } @@ -7564,6 +7751,27 @@ iflib_add_device_sysctl_post(if_ctx_t ctx) SYSCTL_ADD_COUNTER_U64(ctx_list, queue_list, OID_AUTO, "r_abdications", CTLFLAG_RD, &txq->ift_br->abdications, "# of consumer abdications in the mp_ring for this queue"); + if (txq->ift_drbr == NULL) + continue; + SYSCTL_ADD_UQUAD(ctx_list, queue_list, OID_AUTO, + "drbr_direct", CTLFLAG_RD, &txq->ift_drbr_direct, + "# of packets sent without touching the deferral ring"); + SYSCTL_ADD_UQUAD(ctx_list, queue_list, OID_AUTO, + "drbr_stall", CTLFLAG_RD, &txq->ift_drbr_stall, + "# of times the drain stopped with no descriptors free"); + SYSCTL_ADD_COUNTER_U64(ctx_list, queue_list, OID_AUTO, + "drbr_deferred", CTLFLAG_RD, &txq->ift_drbr_deferred, + "# of packets deferred after losing the tx trylock"); + SYSCTL_ADD_COUNTER_U64(ctx_list, queue_list, OID_AUTO, + "drbr_drops", CTLFLAG_RD, &txq->ift_drbr_drops, + "# of packets dropped because the deferral ring was full"); + SYSCTL_ADD_COUNTER_U64(ctx_list, queue_list, OID_AUTO, + "drbr_blocked", CTLFLAG_RD, &txq->ift_drbr_blocked, + "# of times a full deferral ring forced a tx lock wait"); + SYSCTL_ADD_COUNTER_U64(ctx_list, queue_list, OID_AUTO, + "drbr_remote", CTLFLAG_RD, &txq->ift_drbr_remote, + "# of times a different llc or producer limit forced a tx " + "lock wait"); } if (scctx->isc_nrxqsets > 100) @@ -7768,12 +7976,16 @@ iflib_debugnet_poll(if_t ifp, int count) } #endif /* DEBUGNET */ -#ifndef ALTQ static inline iflib_txq_t iflib_simple_select_queue(if_ctx_t ctx, struct mbuf *m) { int qidx; +#ifdef ALTQ + /* ALTQ-enabled interfaces always use queue 0. */ + if (if_altq_is_enabled(ctx->ifc_ifp)) + return (&ctx->ifc_txqs[0]); +#endif if ((NTXQSETS(ctx) > 1) && M_HASHTYPE_GET(m)) qidx = QIDX(ctx, m); else @@ -7781,36 +7993,317 @@ iflib_simple_select_queue(if_ctx_t ctx, struct mbuf *m) return (&ctx->ifc_txqs[qidx]); } +enum iflib_txq_producer_status { + IFLIB_TXQ_PRODUCER_ENTERED, + IFLIB_TXQ_PRODUCER_QUIESCING, + IFLIB_TXQ_PRODUCER_REMOTE, +}; + +/* + * Keep the most recent producer llc when the count reaches zero. Producers + * in that llc can use fetchadd without contending on a compare-and-swap. + * Another llc can take ownership only while the producer count is zero. + */ +static __inline enum iflib_txq_producer_status +iflib_txq_producer_enter(iflib_txq_t txq, bool *pinned) +{ + u_int count, llc_id, max_producers, newstate, old, state; + + if (!iflib_producer_gate) { + state = atomic_fetchadd_int(&txq->ift_producers, 1); + if (__predict_false((state & IFLIB_TXQ_QUIESCING) != 0)) { + atomic_subtract_int(&txq->ift_producers, 1); + return (IFLIB_TXQ_PRODUCER_QUIESCING); + } + *pinned = false; + return (IFLIB_TXQ_PRODUCER_ENTERED); + } + + max_producers = iflib_max_producers; + if (iflib_single_llc) { + state = atomic_fetchadd_int(&txq->ift_producers, 1); + if (__predict_false((state & IFLIB_TXQ_QUIESCING) != 0)) { + atomic_subtract_int(&txq->ift_producers, 1); + return (IFLIB_TXQ_PRODUCER_QUIESCING); + } + count = IFLIB_TXQ_PRODUCER(state); + if (count >= max_producers || + count == IFLIB_TXQ_PRODUCER_MAX) { + atomic_subtract_int(&txq->ift_producers, 1); + return (IFLIB_TXQ_PRODUCER_REMOTE); + } + *pinned = false; + return (IFLIB_TXQ_PRODUCER_ENTERED); + } + + sched_pin(); + llc_id = iflib_cpu_llc[curcpu]; + state = atomic_load_acq_int(&txq->ift_producers); + for (;;) { + if ((state & IFLIB_TXQ_QUIESCING) != 0) { + sched_unpin(); + return (IFLIB_TXQ_PRODUCER_QUIESCING); + } + + count = IFLIB_TXQ_PRODUCER(state); + if (count >= max_producers || + count == IFLIB_TXQ_PRODUCER_MAX) { + sched_unpin(); + return (IFLIB_TXQ_PRODUCER_REMOTE); + } + if (IFLIB_TXQ_PRODUCER_LLC(state) == llc_id) { + /* + * The llc can change between the load and fetchadd only + * if the count was zero. Validate the returned + * state and undo the increment if ownership changed. + */ + old = atomic_fetchadd_int(&txq->ift_producers, 1); + if ((old & IFLIB_TXQ_QUIESCING) != 0) { + atomic_subtract_int(&txq->ift_producers, 1); + sched_unpin(); + return (IFLIB_TXQ_PRODUCER_QUIESCING); + } + count = IFLIB_TXQ_PRODUCER(old); + if (IFLIB_TXQ_PRODUCER_LLC(old) == llc_id && + count < max_producers && + count != IFLIB_TXQ_PRODUCER_MAX) { + *pinned = true; + return (IFLIB_TXQ_PRODUCER_ENTERED); + } + + atomic_subtract_int(&txq->ift_producers, 1); + sched_unpin(); + return (IFLIB_TXQ_PRODUCER_REMOTE); + } + + if (count != 0) { + sched_unpin(); + return (IFLIB_TXQ_PRODUCER_REMOTE); + } + + newstate = llc_id << IFLIB_TXQ_PRODUCER_LLC_SHIFT; + newstate |= 1; + if (atomic_fcmpset_acq_int(&txq->ift_producers, &state, + newstate)) { + *pinned = true; + return (IFLIB_TXQ_PRODUCER_ENTERED); + } + } +} + +static __inline void +iflib_txq_producer_exit(iflib_txq_t txq, bool pinned) +{ + u_int state __diagused; + + if (!pinned) { + atomic_subtract_rel_int(&txq->ift_producers, 1); + return; + } + + atomic_thread_fence_rel(); + state = atomic_fetchadd_int(&txq->ift_producers, -1); + KASSERT(IFLIB_TXQ_PRODUCER(state) != 0, + ("%s: producer count underflow", __func__)); + KASSERT(IFLIB_TXQ_PRODUCER_LLC(state) == iflib_cpu_llc[curcpu], + ("%s: producer llc changed", __func__)); + sched_unpin(); +} + +/* Consumes the mbuf in all cases */ +static int +iflib_simple_encap(iflib_txq_t txq, struct mbuf *m, int *bytes, int *pkts, + int *mcasts) +{ + if_t ifp; + int error; + + mtx_assert(&txq->ift_mtx, MA_OWNED); + ifp = txq->ift_ctx->ifc_ifp; + + error = iflib_encap(txq, &m, bytes, pkts); + if (__predict_false(error != 0)) { + /* iflib_encap() always frees the mbuf on failures */ + if (error == ENOBUFS) + if_inc_counter(ifp, IFCOUNTER_OQDROPS, 1); + else + if_inc_counter(ifp, IFCOUNTER_OERRORS, 1); + return (error); + } + *mcasts += !!(m->m_flags & M_MCAST); + DBG_COUNTER_INC(tx_sent); + ETHER_BPF_MTAP(ifp, m); + (void)iflib_txd_db_check(txq, false); + return (0); +} + +/* Drain the deferral ring into the hardware. */ +static void +iflib_simple_drbr_drain(iflib_txq_t txq, u_int quota, int *bytes, int *pkts, + int *mcasts) +{ + if_ctx_t ctx; + struct mbuf *m; + if_t ifp; + u_int i; + + mtx_assert(&txq->ift_mtx, MA_OWNED); + ctx = txq->ift_ctx; + ifp = ctx->ifc_ifp; + if (__predict_false(!(if_getdrvflags(ifp) & IFF_DRV_RUNNING) || + !LINK_ACTIVE(ctx))) + return; + +#ifdef ALTQ + /* We only drain from txq0 when altq is enabled. */ + if (__predict_false(if_altq_is_enabled(ifp) && txq->ift_id != 0)) + return; +#endif + for (i = 0; TXQ_AVAIL(txq) >= MAX_TX_DESC(ctx); i++) { + if (i == quota) + return; + m = drbr_dequeue(ifp, txq->ift_drbr); + if (m == NULL) + return; + (void)iflib_simple_encap(txq, m, bytes, pkts, mcasts); + } + if (!drbr_empty(ifp, txq->ift_drbr)) { + txq->ift_drbr_stall++; + if (quota != iflib_simple_drain_quota && + (txq->ift_task.gt_task.ta_flags & TASK_ENQUEUED) == 0) + GROUPTASK_ENQUEUE(&txq->ift_task); + } +} + +/* + * Reclaim completed descriptors and push out anything that was deferred while + * the tx lock was held. Called from tx completion and from the timer. + * A thread already holding the lock is draining, and will re-arm us if it + * cannot finish, so never wait for it here. + */ +static void +iflib_simple_txq_drain(iflib_txq_t txq) +{ + if_t ifp; + int bytes_sent = 0, pkt_sent = 0, mcast_sent = 0; + + ifp = txq->ift_ctx->ifc_ifp; + + if (!mtx_trylock(&txq->ift_mtx)) + return; + + if ((atomic_load_acq_int(&txq->ift_producers) & IFLIB_TXQ_QUIESCING) + != 0) { + mtx_unlock(&txq->ift_mtx); + return; + } + + (void)iflib_completed_tx_reclaim(txq, NULL); + if (!drbr_empty(ifp, txq->ift_drbr)) + iflib_simple_drbr_drain(txq, iflib_simple_drain_quota, + &bytes_sent, &pkt_sent, &mcast_sent); + if (txq->ift_db_pending != 0) + (void)iflib_txd_db_check(txq, true); + mtx_unlock(&txq->ift_mtx); + + if_inc_counter(ifp, IFCOUNTER_OBYTES, bytes_sent); + if_inc_counter(ifp, IFCOUNTER_OPACKETS, pkt_sent); + if (mcast_sent) + if_inc_counter(ifp, IFCOUNTER_OMCASTS, mcast_sent); +} + +/* + * When nothing is queued ahead of us there is no ordering constraint, + * so the mbuf goes straight to the hardware and the deferral ring is + * never touched. + */ +static int +iflib_simple_transmit_locked(iflib_txq_t txq, struct mbuf *m, int *bytes, + int *pkts, int *mcasts) +{ + if_ctx_t ctx; + if_t ifp; + int error; + + mtx_assert(&txq->ift_mtx, MA_OWNED); + ctx = txq->ift_ctx; + ifp = ctx->ifc_ifp; + + if (__predict_true(!drbr_needs_enqueue(ifp, txq->ift_drbr) && + TXQ_AVAIL(txq) >= MAX_TX_DESC(ctx))) { + txq->ift_drbr_direct++; + error = iflib_simple_encap(txq, m, bytes, pkts, mcasts); + } else { + error = buf_ring_enqueue(txq->ift_drbr, m); + if (__predict_false(error != 0)) { + m_freem(m); + DBG_COUNTER_INC(tx_frees); + counter_u64_add(txq->ift_drbr_drops, 1); + if_inc_counter(ifp, IFCOUNTER_OQDROPS, 1); + } + } + + /* + * Other transmitters may have deferred to the ring while we were in + * iflib_encap(), so always check again before dropping the lock. We + * are the only thread that can drain it. + */ + if (!drbr_empty(ifp, txq->ift_drbr)) + iflib_simple_drbr_drain(txq, iflib_simple_drain_quota_thread, + bytes, pkts, mcasts); + return (error); +} + static int iflib_simple_transmit(if_t ifp, struct mbuf *m) { if_ctx_t ctx; iflib_txq_t txq; struct mbuf **m_defer; + enum iflib_txq_producer_status producer_status; + bool pinned; int error, i, reclaimable; int bytes_sent = 0, pkt_sent = 0, mcast_sent = 0; - ctx = if_getsoftc(ifp); if (__predict_false((if_getdrvflags(ifp) & IFF_DRV_RUNNING) == 0 - || !LINK_ACTIVE(ctx))) { - DBG_COUNTER_INC(tx_frees); - m_freem(m); - return (ENETDOWN); - } + || !LINK_ACTIVE(ctx))) + goto net_down; txq = iflib_simple_select_queue(ctx, m); - mtx_lock(&txq->ift_mtx); - error = iflib_encap(txq, &m, &bytes_sent, &pkt_sent); - if (error == 0) { - mcast_sent += !!(m->m_flags & M_MCAST); - (void)iflib_txd_db_check(txq, true); - } else { - if (error == ENOBUFS) - if_inc_counter(ifp, IFCOUNTER_OQDROPS, 1); - else - if_inc_counter(ifp, IFCOUNTER_OERRORS, 1); + /* + * Avoid blocking behind another transmitter; the ring is drained by + * whoever holds ift_mtx, by tx completion, or by the watchdog timer. + */ + if (__predict_false(!mtx_trylock(&txq->ift_mtx))) { + producer_status = iflib_txq_producer_enter(txq, &pinned); + if (producer_status == IFLIB_TXQ_PRODUCER_QUIESCING) + goto net_down; + + if (producer_status == IFLIB_TXQ_PRODUCER_ENTERED) { + error = buf_ring_enqueue(txq->ift_drbr, m); + iflib_txq_producer_exit(txq, pinned); + if (__predict_true(error == 0)) { + counter_u64_add(txq->ift_drbr_deferred, 1); + return (0); + } + counter_u64_add(txq->ift_drbr_blocked, 1); + } else { + counter_u64_add(txq->ift_drbr_remote, 1); + } + mtx_lock(&txq->ift_mtx); + } + + if (__predict_false(atomic_load_acq_int(&txq->ift_producers) & + IFLIB_TXQ_QUIESCING)) { + mtx_unlock(&txq->ift_mtx); + goto net_down; } + + error = iflib_simple_transmit_locked(txq, m, &bytes_sent, &pkt_sent, + &mcast_sent); + if (txq->ift_db_pending != 0) + (void)iflib_txd_db_check(txq, true); m_defer = NULL; reclaimable = iflib_txq_can_reclaim(txq); if (reclaimable != 0) { @@ -7845,5 +8338,43 @@ iflib_simple_transmit(if_t ifp, struct mbuf *m) if_inc_counter(ifp, IFCOUNTER_OMCASTS, mcast_sent); return (error); + +net_down: + m_freem(m); + DBG_COUNTER_INC(tx_frees); + return (ENETDOWN); +} + +/* + * ALTQ entry point. drbr_dequeue() pulls from ifp->if_snd when a discipline + * is attached, so the drain loop is shared with the if_transmit path. Only + * queue zero is used, matching the queue ALTQ itself selects. + */ +static void +iflib_simple_if_start(if_t ifp) +{ + if_ctx_t ctx; + iflib_txq_t txq; + bool retry; + int bytes_sent = 0, pkt_sent = 0, mcast_sent = 0; + + ctx = if_getsoftc(ifp); + txq = &ctx->ifc_txqs[0]; + + mtx_lock(&txq->ift_mtx); + (void)iflib_completed_tx_reclaim(txq, NULL); + iflib_simple_drbr_drain(txq, UINT_MAX, &bytes_sent, &pkt_sent, + &mcast_sent); + if (txq->ift_db_pending != 0) + (void)iflib_txd_db_check(txq, true); + retry = !drbr_empty(ifp, txq->ift_drbr); + mtx_unlock(&txq->ift_mtx); + + if (retry) + GROUPTASK_ENQUEUE(&txq->ift_task); + + if_inc_counter(ifp, IFCOUNTER_OBYTES, bytes_sent); + if_inc_counter(ifp, IFCOUNTER_OPACKETS, pkt_sent); + if (mcast_sent) + if_inc_counter(ifp, IFCOUNTER_OMCASTS, mcast_sent); } -#endif