Thread (8 messages) flat view 8 messages, 3 authors, 2026-07-03

Re: [PATCH iwl v3] ice: retry reading NVM if admin queue returns EBUSY

From: Robert Malz <hidden>
Date: 2026-07-03 10:46:41
Also in: intel-wired-lan

Hey Przemek,
I have submitted a new patch to solve the current issue:
https://lists.osuosl.org/pipermail/intel-wired-lan/Week-of-Mon-20260629/055891.html
I ran numerous tests on it and also verified it against the setup
which initially triggered the issue.
This system had a problem where ice_discover_flash_size read time was
greater that 3 seconds and as a result failed with EBUSY status (FW at
asome point releases the lock).
With new patch issue was not reproduced even when the read time was
much higher than normally expected:
[   75.999353] ice 0000:05:00.0: CQ CMD: opcode 0x0008, flags 0x2000,
datalen 0x0000, retval 0x0000 <- First lock acquisition
[   82.152231] ice 0000:05:00.0: Predicted flash size is 10485760
bytes <- ice_discover_flash_size finished reading (7 seconds).
In this case, the driver properly read the NVM contents without
needing any retries.

I also analyzed the overhead of adding an acquire/release lock between
each 0x0701 call and based on my observations it did not negatively
impact the driver in any way.
Execution time for ice_discover_flash_size in normal scenario
increased from around 0,02s to 0,05s which should still be acceptable.

Regards,
Robert


On Fri, Jun 26, 2026 at 1:46 PM Przemek Kitszel
[off-list ref] wrote:
On 6/26/26 10:15, Robert Malz wrote:
quoted
Hey Przemek,
I ran some tests and unfortunately, the following sentence from the
datasheet is true:
"For specific resources, such as Change Lock (0x0003) and Global Config Lock
(0x0004), this field is used by software to override the default timeout for the
operation, and also to specify the timeout used for this operation."

This means we can only change a default timeout for 0x0003 and 0x0004
but not for 0x0001 (NVM resource).
Whatever timeout I provide FW defaults to 0xB88
Input:
[ 2209.656758] ice 0000:31:00.0: CQ CMD: opcode 0x0008, flags 0x2000,
datalen 0x0000, retval 0x0000
[ 2209.656760] ice 0000:31:00.0:        cookie (h,l) 0x00000000 0x00000000
[ 2209.656761] ice 0000:31:00.0:        param (0,1)  0x00010001 0x00000BB9
Output:
[ 2209.656927] ice 0000:31:00.0: CQ CMD: opcode 0x0008, flags 0x2003,
datalen 0x0000, retval 0x0000
[ 2209.656929] ice 0000:31:00.0:        cookie (h,l) 0x00000000 0x00000000
[ 2209.656931] ice 0000:31:00.0:        param (0,1)  0x00010001 0x00000BB8

Correct me If I'm wrong, but the only way to properly handle it is to
ensure the resource is locked and released between every
ice_acquire_nvm call.
I'll start working on this.
thank you for checking out!

I agree that simple retries with improved (refactored) locking will be
good solution.
Failure to lock should count as an unsuccessful attempt, with possible
retry after a sleep.
quoted
Regards,
Robert
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help