Thread (23 messages) flat view 23 messages, 3 authors, 1d ago

Re: [PATCH net-next v6 00/13] ax88179_178a: Add support for AX88179A-based chips

From: Jianhui Xu <hidden>
Date: 2026-08-10 01:36:03
Also in: linux-usb, lkml

Sorry for the late reply. I think the previous experimental workaround may
have some race conditions, so I was seeking for a better solution.
I tested the 100MBit connections mainly with a AX88772E 100MBit adapter
(UGREEN CR110), which has the same firmware (1.3.0.0) as your and my
AX88179B adapter.
My adapter is an AX88179B with firmware v1.3.0.3, not v1.3.0.0.
You did not mention which device is used on the other side of the
Ethernet link (or maybe I missed that), could you specify this?
The link partner was the Ethernet port of a ZTE ZXHN F7005MV3 gateway, not
another USB Ethernet adapter. ethtool reported autonegotiation support and
no advertised pause frames.
The only way this could be coming from the driver that I see is via a call
to ax88179a_stop(), which would clear exactly that bit.
Have you traced this and can exclude that this function is called somehow?
Yes. I added a dynamic kprobe on ax88179_write_cmd(), filtered to
AX_MEDIUM_STATUS_MODE writes, and recorded the value and caller stack.

During 30 100baseT/Full-to-1000baseT/Full cycles, four transitions to
100baseT/Full lost RX. In all four cases:

- ax88179a_mac_link_up() first wrote 0x0102;
- there was no intervening Linux write to AX_MEDIUM_STATUS_MODE;
- about one second later the delayed worker read the register with
  AX_MEDIUM_RECEIVE_EN clear and restored 0x0102.

All 71 traced writes to AX_MEDIUM_STATUS_MODE had
AX_MEDIUM_RECEIVE_EN set. The complete 1,209-entry trace contained no
ax88179a_stop() or ax88179_change_mtu() caller.

The earlier failed-state dump also retained AX_RX_CTL at 0x0198 with carrier
up. I therefore think ax88179a_stop() can be excluded as the direct source
of these clears.

I also repeated the test with the same AX88179B and ZTE link partner using
the ASIX vendor driver on the Arch host. All 30 further
100baseT/Full-to-1000baseT/Full cycles passed. All 4,415 carrier-up
100baseT/Full samples retained receive-enable at 0x0132. A separate trace
recorded 30 writes of 0x0132 and 30 writes of 0x0133, with no write clearing
AX_MEDIUM_RECEIVE_EN.

Dense sampling did show the adapter changing 0x0133 to 0x0033, or 0x0132 to
0x0032, while the link was down during renegotiation, without a corresponding
vendor-driver write. The vendor link-setting path then restored 0x0132 or
0x0133 before the link became stably up.

So the device can clear AX_MEDIUM_RECEIVE_EN without a corresponding host
write. In the failing v6 case, the clear likewise was not caused by a Linux
write and appears to be an autonomous device-side change.

I do not think this proves an unconditional device-side bug, though. The v6
driver failed four times in 30 cycles after writing 0x0102, while the vendor
driver had no failures after writing 0x0132.

I therefore tested a focused v6 variant that writes 0x0132 instead of 0x0102
at 100baseT/Full. Failures still occurred, so retaining the RX/TX
flow-control bits alone is not sufficient to prevent the problem.

This points to some other difference in the vendor driver's link-setting
sequence, possibly register ordering or timing. The physical xHCI host in
my tests versus QEMU's emulated xHCI is another uncontrolled difference.

At this point this looks like a device-side quirk exposed by the driver's
link-setting sequence, but I cannot distinguish firmware behavior from
autonomous MAC hardware behavior.
If this can indeed be attributed to a bug in the firmware of the
adapters, I would add your patch to the series with an "Authored-by" you,
as this sounds like a good solution for this issue.
I found several race conditions in the patch. For example,
`cancel_delayed_work(...)` only cancels pending work. If the callback has
already started running, it may still be running when
`cancel_delayed_work(...)` returns. Therefore, if another execution context
performs a read-modify-write operation on `MEDIUM_STATUS`, there can be
a race: the link is brought down, but the worker subsequently writes the
RX-enable bit back. This particular issue can be fixed by using
`cancel_delayed_work_sync(...)`, but there are still other races.

One possible solution would be to add a mutex to serialize accesses to the
medium register. However, I am reluctant to add too much synchronization
machinery for what is essentially a defensive workaround, especially given
how infrequently these network configuration operations occur.

I see two options:
1. Leave the code mostly as it is, changing only
   `cancel_delayed_work(...)` to `cancel_delayed_work_sync(...)`. This keeps
   the main code path simple and clear, at the cost of leaving a few rare
   corner cases unresolved.
2. Add stronger synchronization to eliminate these races completely, at the
   cost of making the code considerably more complex for cases that are
   unlikely to occur in practice.

What do you think?

Thanks,
Jianhui
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help