configs/cubie_a7a/linux-6.6/board.dts line 2144 sets:
phy-mode = "rgmii";
tx-delay = <12>;
rx-delay = <10>;
On my Cubie A7A, tx-delay = <12> corrupts transmitted frames. A sweep of all
32 possible values found the board only carries real traffic at 9, 10 or 11,
and only 9 has zero loss at every frame size.
Why it is easy to miss
It does not present as a network fault. The link negotiates 1000/full, ethtool -S reports no TX errors, and short pings reply normally. What actually happens
is that SSH hangs immediately after key exchange and apt stalls part-way
through a download — so it reads as a bad cable, and it survives changing the
cable.
Measurement
All 32 delay values, 50 ICMP echoes per frame size, counting a reply only when
the payload returned byte-identical. Loss %, generic PHY driver:
| tx_delay |
64 B |
300 B |
700 B |
1200 B |
| 8 |
2 |
10 |
20 |
32 |
| 9 |
0 |
0 |
0 |
0 |
| 10 |
0 |
0 |
2 |
0 |
| 11 |
0 |
0 |
4 |
6 |
| 12 (current) |
4 |
8 |
24 |
42 |
| 13 |
22 |
86 |
86 |
100 |
Outside 7–13 it is 100% loss at every size. Loss climbs with frame size rather
than switching on at a threshold, which is why ping succeeds and apt does
not.
The part that explains why this shipped
The corruption is data-dependent. At tx-delay = 12, a payload of one
repeated byte passes with zero loss at every frame size, while random data
loses up to 42%. The constant-byte payload works across delays 8–13; real
traffic only works across 9–11.
Any link test using a simple generated payload will therefore report this board
as healthy. Testing needs high-entropy payloads.
Ruled out: the missing PHY driver
The PHY reports 0x7b744412, which drivers/net/phy/maxio.c names
"MAE0621A/B-Q3C(I)". CONFIG_MAXIO_PHY is not enabled in the BSP defconfig, so
it binds to the generic PHY driver and maxio_mae0621aq3ci_config_init() —
27 vendor registers including TXCST OFF and IO DS=1 — never runs.
I built CONFIG_MAXIO_PHY as a module, loaded it, and repeated the identical
sweep. It makes delay 12 worse (0 / 90 / 100 / 100 %) and narrows the usable
range to 8–9. So the delay value is the fault, not the absent PHY driver —
though enabling that driver may be worth doing separately.
Suggested fix
- tx-delay = <12>;
+ tx-delay = <9>;
More usefully: bsp/drivers/stmmac/dwmac-sunxi.c already supports a
delay-maps property — <soc_version, rx_delay, tx_delay> triplets applied per
silicon revision. No board device tree in the tree uses it. If the correct value
varies by stepping, that mechanism already exists and would be a better fix than
a single constant.
Scope
Measured on one board, one cable, one switch, single run. Individual cells vary
by a few percent between runs; the position of the window and the
constant-versus-random gap reproduce.
The same tx-delay = <12> appears in cubie_a7s (5.15 and 6.6), QA and
evb1. I have not tested those and am not proposing a change to them — but
since the value is identical across boards with different layouts, it may be
worth checking.
Full sweep, both payload types, both PHY drivers, and the raw data and script:
https://github.com/Rabs9/radxa-cubie-a7a-kernel/blob/main/docs/ETHERNET-TX-DELAY.md
Happy to send a PR for the one-line change if that is useful, or to re-run the
sweep with any parameters you would find more convincing.
/cc @vamrs-feng @RadxaStephen
configs/cubie_a7a/linux-6.6/board.dtsline 2144 sets:On my Cubie A7A,
tx-delay = <12>corrupts transmitted frames. A sweep of all32 possible values found the board only carries real traffic at 9, 10 or 11,
and only 9 has zero loss at every frame size.
Why it is easy to miss
It does not present as a network fault. The link negotiates 1000/full,
ethtool -Sreports no TX errors, and short pings reply normally. What actually happensis that SSH hangs immediately after key exchange and
aptstalls part-waythrough a download — so it reads as a bad cable, and it survives changing the
cable.
Measurement
All 32 delay values, 50 ICMP echoes per frame size, counting a reply only when
the payload returned byte-identical. Loss %, generic PHY driver:
Outside 7–13 it is 100% loss at every size. Loss climbs with frame size rather
than switching on at a threshold, which is why
pingsucceeds andaptdoesnot.
The part that explains why this shipped
The corruption is data-dependent. At
tx-delay = 12, a payload of onerepeated byte passes with zero loss at every frame size, while random data
loses up to 42%. The constant-byte payload works across delays 8–13; real
traffic only works across 9–11.
Any link test using a simple generated payload will therefore report this board
as healthy. Testing needs high-entropy payloads.
Ruled out: the missing PHY driver
The PHY reports
0x7b744412, whichdrivers/net/phy/maxio.cnames"MAE0621A/B-Q3C(I)".
CONFIG_MAXIO_PHYis not enabled in the BSP defconfig, soit binds to the generic PHY driver and
maxio_mae0621aq3ci_config_init()—27 vendor registers including
TXCST OFFandIO DS=1— never runs.I built
CONFIG_MAXIO_PHYas a module, loaded it, and repeated the identicalsweep. It makes delay 12 worse (0 / 90 / 100 / 100 %) and narrows the usable
range to 8–9. So the delay value is the fault, not the absent PHY driver —
though enabling that driver may be worth doing separately.
Suggested fix
More usefully:
bsp/drivers/stmmac/dwmac-sunxi.calready supports adelay-mapsproperty —<soc_version, rx_delay, tx_delay>triplets applied persilicon revision. No board device tree in the tree uses it. If the correct value
varies by stepping, that mechanism already exists and would be a better fix than
a single constant.
Scope
Measured on one board, one cable, one switch, single run. Individual cells vary
by a few percent between runs; the position of the window and the
constant-versus-random gap reproduce.
The same
tx-delay = <12>appears incubie_a7s(5.15 and 6.6),QAandevb1. I have not tested those and am not proposing a change to them — butsince the value is identical across boards with different layouts, it may be
worth checking.
Full sweep, both payload types, both PHY drivers, and the raw data and script:
https://github.com/Rabs9/radxa-cubie-a7a-kernel/blob/main/docs/ETHERNET-TX-DELAY.md
Happy to send a PR for the one-line change if that is useful, or to re-run the
sweep with any parameters you would find more convincing.
/cc @vamrs-feng @RadxaStephen