Re: nvme device errors & zfs

From: Dave Cottlehuber <dch_at_FreeBSD.org>
Date: Thu, 06 Mar 2025 16:10:36 UTC
On Tue, 5 Nov 2024, at 17:21, Warner Losh wrote:
> On Tue, Nov 5, 2024 at 4:20 AM Dave Cottlehuber <dch@freebsd.org> wrote:
>> On Tue, 5 Nov 2024, at 11:10, Tomek CEDRO wrote:
>> > Magician software can upgrade firmware and perform other checks, works
>> > on Windoze macOS and Android:
>> 
>> that will be difficult, I don't have an nvme capable thing for any
>> of those.
>
> nvmecontrol updates firmware just fine, though.
> 
>> > Another idea is maybe disk overheats and resets itself to cool down?
>> 
>> that is a great point, the mainboard comes with inbuilt heatsinks, but
>> when I assembled it, the 2nd nvme slot heatsink looked a lot less
>> bulky than the other one, I remember distinctly wondering if it would
>> cope. If it happens again I'll see if I can get a temp measurement
>> at the time of failure.
>> 
>> I would hope temperature throttling would not be quite so brutal, to
>> remove itself from the bus entirely, but its a reasonable explanation.
>
> What's supposed to happen is that the temperature climbs slowly enough
> that there's a chance for it to kick in. It might be thermals, but I'd expect
> at least some indications that it's thermal. Heat sinks are cheap enough,
> if it's really thermal. log page 2 has the temperature.
>
> How often does the reset happen? A firmware upgrade might solve that
> problem if they are older. There might just be a bug that causes the firmware
> to 'trap' and it takes several seconds for the SoC / controller to reboot.
>
> Warner

Thanks Warner

TLDR the motherboard[1] firmware[2] needed a specific fix to resolve this,
the new firmware was published just a fortnight ago! Details below.

To my surprise, its the easily-reached NVMe slot that has issues, and this
one I should be able to put a 3rd party heatsink on it, if necesssary. The
other one is squeezed under the CPU cooler, and there is not really
sufficient space for a larger heatsink.

Prior to firmware update, I can trigger the failure with `zpool scrub 
root` or poudriere reproducibly any day.

Over the last couple of months of fiddling. I:

- swapped the 2 nvme drives: failure stays in the same slot
- replaced the drives entirely: failure is present in same slot
- checked that the drive firmware is latest
- physically reset the vendor-provided heatsinks
- updated mainboard firmware, it has a Samsung-specific fix from
  a couple of weeks ago:

Version 2806 2025/02/21 Fixed compatibility issue with Samsung EVO 990

Since this firmware update, I've had only 1 failure, and 12 successful
scrub while concurrently rebuilding ports.

I will follow up with Samsung support about the temperature warnings below,
maybe they can explain how to get timestamps out of this as well. The other
device has zero for the highlighted temperature threshold transitions.

$ doas nvmecontrol logpage -p 2 nvme1
Log
============================
Critical Warning State:         0x00
 Available spare:               0
 Temperature:                   0
 Device reliability:            0
 Read only:                     0
 Volatile memory backup:        0
Temperature:                    312 K, 38.85 C, 101.93 F
Available spare:                100
Available spare threshold:      10
Percentage used:                0
Data units (512,000 byte) read: 63747843
Data units written:             29591310
Host read commands:             324928116
Host write commands:            668650979
Controller busy time (minutes): 951
Power cycles:                   23
Power on hours:                 1814
Unsafe shutdowns:               15
Media errors:                   0
No. error info log entries:     0
Warning Temp Composite Time:    9 <----------
Error Temp Composite Time:      0
Temperature Sensor 1:           312 K, 38.85 C, 101.93 F
Temperature Sensor 2:           315 K, 41.85 C, 107.33 F
Temperature 1 Transition Count: 2 <----------
Temperature 2 Transition Count: 1 <----------
Total Time For Temperature 1:   20 <---------
Total Time For Temperature 2:   553 <--------

So perhaps a 3rd party heatsink would be advised, if these metrics
go up at all in future.

[1]: https://www.asus.com/motherboards-components/motherboards/proart/proart-x670e-creator-wifi/
[2]: https://dlcdnets.asus.com/pub/ASUS/mb/BIOS/ProArt-X670E-CREATOR-WIFI-ASUS-2806.zip?model=ProArt%20X670E-CREATOR%20WIFI

A+
Dave
———
O for a muse of fire, that would ascend the brightest heaven of invention!