[Bug 298862] LinuxKPI/iwlwifi: reboot panic from out-of-bounds lsta->kc[65535] in lkpi_hw_crypto_prepare
Date: Fri, 25 Sep 2026 22:45:16 UTC
https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=298862
Bug ID: 298862
Summary: LinuxKPI/iwlwifi: reboot panic from out-of-bounds
lsta->kc[65535] in lkpi_hw_crypto_prepare
Product: Base System
Version: 16.0-CURRENT
Hardware: amd64
OS: Any
Status: New
Severity: Affects Only Me
Priority: ---
Component: wireless
Assignee: wireless@FreeBSD.org
Reporter: oleglelchuk@gmail.com
I encountered a general protection fault immediately after issuing reboot on
FreeBSD 16.0-CURRENT amd64. The saved dump shows a Wi-Fi TX worker using
IEEE80211_KEYIX_NONE (65535) as an index into the four-entry lsta->kc array,
then dereferencing the invalid pointer it obtained in lkpi_hw_crypto_prepare().
I have observed this specific panic once. I have not reproduced it with an
unmodified kernel or established a reliable reproducer. My kernel contains
local patches, as detailed below. An LLM analyzed the saved dump with kgdb,
matching debug symbols, disassembly, and comparisons against unmodified FreeBSD
sources. The confirmed invalid lookup is distinguished below from the inferred
race.
Environment and trigger
-----------------------
* FreeBSD 16.0-CURRENT, amd64, GENERIC configuration, based on
b9811d13572b93adff2da16508eef18872b24c85.
* if_iwlwifi with LinuxKPI 802.11; hardware crypto was enabled
(lkpi_hwcrypto=true in the dump).
* I issued reboot while Wi-Fi was in use. A Wi-Fi state-transition thread was
taking the interface from RUN to INIT when the TX worker faulted.
* Earlier in the same boot, bounded CPU/build and memory-pressure tests had
run, and desktop GPU hangs had occurred. The test allocations had been released
before reboot. At the panic, approximately 5.51 GiB of RAM was free. This is
context, not an established trigger or evidence of RAM exhaustion.
Confirmed fault and stacks
--------------------------
Panic: general protection fault (amd64 trap 9).
The relevant faulting stack, innermost first, is:
lkpi_hw_crypto_prepare() [inlined into lkpi_80211_txq_tx_one]
lkpi_80211_txq_tx_one()
lkpi_80211_txq_task()
taskqueue_run_locked()
taskqueue_thread_loop()
The fault is at switch (kc->cipher), following:
kc = lsta->kc[k->wk_keyix];
info = IEEE80211_SKB_CB(skb);
info->control.hw_key = kc;
if (kc == NULL) { ... return (ENXIO); }
switch (kc->cipher) { ... }
The declaration is:
struct ieee80211_key_conf *kc[IEEE80211_WEP_NKID];
IEEE80211_WEP_NKID is 4. The saved values and pointer relationships are:
k == &lsta->ni->ni_ucastkey
k->wk_keyix == 0xffff == IEEE80211_KEYIX_NONE
k->wk_cipher == &ieee80211_cipher_none
k->wk_cipher->ic_name == "NONE"
k->wk_flags == 3 [IEEE80211_KEY_XMIT | IEEE80211_KEY_RECV]
lsta->ni->ni_refcnt == 3
lsta->ni->ni_drv_data == lsta
lsta->state == 4 [IEEE80211_STA_AUTHORIZED]
lsta->txq_ready == false
lsta->added_to_drv == true
lsta->in_mgd == true
lsta->kc[0] and lsta->kc[1] are non-NULL; slots 2 and 3 are NULL
kc == 0x0f0f0f0f0f0f0f0f
The compiler loads the index with movzwl, so this is a zero-extended index of
65535, not a signed -1 lookup. It reads far outside the four-entry array. The
resulting noncanonical pointer passes the NULL test, then causes a general
protection fault when its cipher member is read. The 0x0f pattern is the value
obtained from the out-of-bounds location; this is not evidence that a
legitimate key allocation was freed or poisoned.
Disassembly of the actual faulting binary shows an earlier index check and a
later independent reload. Offsets below are relative to lkpi_80211_txq_tx_one
in this binary; comments summarize the inspected control flow:
+132 call ieee80211_crypto_get_txkey
+142 movzwl 0x8(%rax),%ecx # earlier wk_keyix load
+146 cmp $0xffff,%ecx
+152 je ... # software-crypto branch
+160 cmpq $0,0x78(%r14,%rcx,8) # key-slot check
...
+1198 movzwl 0x8(%rdi),%eax # reload wk_keyix
+1202 mov 0x78(%r14,%rax,8),%rbx # kc[index], with RAX=65535
+1207 mov %rbx,0x68(%r12) # info->control.hw_key
+1212 test %rbx,%rbx
+1215 je ... # NULL-key error path
+1221 mov 0x10(%rbx),%r10d # fault: kc->cipher
A different thread was waiting for this same lsta's TX task:
taskqueue_drain(taskqueue_thread, &lsta->txq_task)
lkpi_80211_flush_tx(lsta, ...)
lkpi_sta_run_to_assoc(..., arg=3)
lkpi_sta_run_to_init(...)
lkpi_iv_newstate(..., IEEE80211_S_INIT, arg=3)
ieee80211_newstate_cb()
taskqueue_run_locked()
taskqueue_thread_loop()
Race interpretation and limits
------------------------------
The key's saved state matches ieee80211_crypto_resetkey(...,
IEEE80211_KEYIX_NONE), which is called after the key is cleared in
_ieee80211_crypto_delkey(). Together with the earlier successful hardware-key
selection and the later index reload, this strongly suggests key reset/deletion
racing an in-progress TX operation. The dump does not identify the exact
instruction/thread that reset the key or retain the historical interleaving.
The out-of-bounds lookup itself is directly established by the registers,
disassembly, and array declaration.
A fix likely needs to consider both index validation and synchronization/key
lifetime across reset/deletion and TX, including any queued skb references to
hw_key. A bounds check at this lookup alone would not establish that all key
lifetime races are fixed. I have not implemented or tested a fix.
Local changes and upstream comparison
-------------------------------------
My kernel includes local PFN/shmem, Wi-Fi recovery, USB/Bluetooth, and
USB-Ethernet changes. Diagnostic drm-kmod modules and a shutdown recorder were
also loaded. The recorder's observer was asleep and had attempted no EFI writes
at the panic. These qualifications do not establish that the local changes had
no effect on timing.
The LLM compared the complete relevant function bodies with unmodified
b9811d13572b. These are unchanged:
lkpi_hw_crypto_prepare
lkpi_hw_crypto_tailroom
lkpi_80211_txq_tx_one
lkpi_iv_key_update_begin / lkpi_iv_key_update_end
lkpi_iv_key_delete
lkpi_lsta_free
lkpi_80211_flush_tx
lkpi_sta_run_to_assoc
The outer lkpi_80211_txq_task has local firmware-recovery gating/requeue
changes, but the hardware's restart_flags were zero in the dump, so the added
restart gate was inactive. I am reporting an unsafe access present in upstream
code, not claiming that I have reproduced the entire incident on a vanilla
kernel.
Upstream main was also fetched at 4bfe52650c140694505dd136f24bb3649153124a.
Across the 34 commits after b9811d13572b, linux_80211.c/.h and the net80211
crypto, node, protocol, output, and ioctl files have identical Git blob IDs.
The only net80211 change in that interval is a sysctl macro conversion. The
unchecked lookup is still at linux_80211.c:5950 and the cipher dereference at
:5962 in that upstream revision:
https://cgit.freebsd.org/src/tree/sys/compat/linuxkpi/common/src/linux_80211.c?id=4bfe52650c140694505dd136f24bb3649153124a#n5950
Related work
------------
* D49256 addresses an earlier key-removal/update locking race. Its commit
b8dfc3ecf703 is already in my kernel's base.
https://reviews.freebsd.org/D49256
* D49791 / bug 285729 addresses sleeping while holding a non-sleepable lock
during shutdown. Its commit a6165709e3c8 is also already in the base; its panic
mechanism differs from this invalid-index access.
https://reviews.freebsd.org/D49791
https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=285729
* Bug 298369 concerns a TX task synchronously draining itself. In this incident
the TX worker faults, while a separate thread waits for it. I have not
established a common cause with that report or with silent shutdown freezes.
https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=298369
I retained the matching kernel, debug symbols, and dump locally. The report
includes only selected technical findings; no raw dump, complete crash report,
network identifiers, encryption-key contents, or identifying host/path
information is attached.
--
You are receiving this mail because:
You are the assignee for the bug.