I have a Supermicro SSLH12-CT moatherboard. using the on board nics. after the new update i get bnxt_en errors and it cant connect with the firm ware.
Ai summary
Title: bnxt_en: firmware hang loop (hwrm_ring_free failed / HWRM timeouts) on 7.2.7-zabbly+, survives reboot
Summary
On kernel 7.2.7-zabbly+, the onboard Broadcom NIC (bnxt_en, PCI 0000:46:00.0, eno1np0) enters an endless recovery loop. The driver repeatedly fails to free TX/RX rings, and the NIC firmware then stops responding to HWRM commands entirely. The problem survived a cold reboot. Unbinding/rebinding the driver fails with firmware timeouts.
Environment
- Kernel:
7.2.7-zabbly+
- Distro: Ubuntu 24.04 (sudo 1.9.15p5-3ubuntu5.24.04.2)
- NIC:
<lspci -nn -s 46:00.0>
- NIC firmware:
<ethtool -i eno1np0 → firmware-version>
- Motherboard / BIOS:
<dmidecode -s baseboard-product-name / bios-version>
- CPU:
<lscpu | grep 'Model name'>
- Last known good kernel:
<version, or "unknown">
- Kernel cmdline:
<cat /proc/cmdline>
- Usage: host runs Incus containers bridged on eno1np0 + VLAN 40 subinterface
Symptoms
Dmesg floods with:
bnxt_en 0000:46:00.0 eno1np0: Resp cmpl intr err msg: 0x51
bnxt_en 0000:46:00.0 eno1np0: hwrm_ring_free type 1 failed. rc:fffffff0 err:0
bnxt_en 0000:46:00.0 eno1np0: Resp cmpl intr err msg: 0x51
bnxt_en 0000:46:00.0 eno1np0: hwrm_ring_free type 2 failed. rc:fffffff0 err:0
(repeating every ~1-2 s, starting around t=1516 s after boot)
Attempting recovery via sysfs:
# echo 0000:46:00.0 > /sys/bus/pci/drivers/bnxt_en/unbind
bnxt_en 0000:46:00.0 eno1np0 (unregistered): Error (timeout: 500015) msg {0x1a 0x1351} len:0
# echo 0000:46:00.0 > /sys/bus/pci/drivers/bnxt_en/bind
bnxt_en 0000:46:00.0 (unnamed net_device) (uninitialized): Error (timeout: 500015) msg {0x0 0x0} len:0
tee: /sys/bus/pci/drivers/bnxt_en/bind: No such device
modprobe -r bnxt_en fails with "Module bnxt_en is in use". (bnxt_re is not built in this kernel.)
What was tried
- Cold reboot: problem returned
- Driver unbind/bind: firmware does not respond (HWRM timeout, even on VER_GET)
- PCI reset / remove+rescan: Fail
- Full power drain (PSU unplugged 60 s): Fail
- Booting older / stock Ubuntu kernel: `Pending
Notes
According to the Broadcom maintainer's comments on similar reports (https://lists.openwall.net/netdev/2020/11/30/177), this pattern indicates IRQ or firmware issues. Here the firmware ends up completely unresponsive, which may indicate a driver regression in 7.2.x triggering a firmware fault, or an incompatibility with the installed NIC firmware.
The host recently went from using two ports of this dual-port NIC to only eno1np0 (eno2np1 removed from the config), shortly before the issue appeared: <confirm timing or delete this line>.
Full dmesg from boot until the loop starts: <attach dmesg.txt>
I have a Supermicro SSLH12-CT moatherboard. using the on board nics. after the new update i get bnxt_en errors and it cant connect with the firm ware.
Ai summary
Title: bnxt_en: firmware hang loop (hwrm_ring_free failed / HWRM timeouts) on 7.2.7-zabbly+, survives reboot
Summary
On kernel
7.2.7-zabbly+, the onboard Broadcom NIC (bnxt_en, PCI0000:46:00.0,eno1np0) enters an endless recovery loop. The driver repeatedly fails to free TX/RX rings, and the NIC firmware then stops responding to HWRM commands entirely. The problem survived a cold reboot. Unbinding/rebinding the driver fails with firmware timeouts.Environment
7.2.7-zabbly+<lspci -nn -s 46:00.0><ethtool -i eno1np0 → firmware-version><dmidecode -s baseboard-product-name / bios-version><lscpu | grep 'Model name'><version, or "unknown"><cat /proc/cmdline>Symptoms
Dmesg floods with:
Attempting recovery via sysfs:
modprobe -r bnxt_enfails with "Module bnxt_en is in use". (bnxt_reis not built in this kernel.)What was tried
Notes
According to the Broadcom maintainer's comments on similar reports (https://lists.openwall.net/netdev/2020/11/30/177), this pattern indicates IRQ or firmware issues. Here the firmware ends up completely unresponsive, which may indicate a driver regression in 7.2.x triggering a firmware fault, or an incompatibility with the installed NIC firmware.
The host recently went from using two ports of this dual-port NIC to only
eno1np0(eno2np1removed from the config), shortly before the issue appeared:<confirm timing or delete this line>.Full dmesg from boot until the loop starts:
<attach dmesg.txt>