Post-mortem of a dying PHY

A pole controller that was healthy on July 29 was unreachable by August 4. Three sessions across two nights chased it: the first convicted the wrong components with real measurements, the second acquitted them with a better experiment and found a fault that worsens with 100 Mbit operation and rewinds with a register-level reset, and a third proved that the scariest symptom, a 274 ms ping, was never this board's fault at all. This is the post-mortem: what the fault looks like, where it probably lives (the root cause is not yet proven; the bench session that would prove it is queued), and an itemized list of the investigative flaws, because every one of them is the reusable part.
The board
"Top Support" is one deployed pole controller: an ESP32 driving a LAN8720A Ethernet PHY into discrete magnetics, on the previous-generation two-layer board (the four-layer redesign this project has been revising all month is its successor and has not been fabricated yet). That matters twice over: the failure is field evidence feeding the redesign rather than a defect in it, and any bench work has to re-map its probe points from the two-layer board's own schematic, because the component designators in the planning notes were read off the four-layer netlist. It sat healthy in the bench records on July 29. Six days later its link LED was blinking and nothing answered ping. The processor itself was fine: uptime monotonic, no crashes, no brownouts. Whatever had died, died in the network path.
What the console said, and why it was wrong
The boot console reported a link that came up for 2.000 seconds, dropped for 8.000, and repeated every 10.000 seconds with zero drift over 25 cycles. Zero drift is normally a strong signal: deterministic process, not a loose connector. The inference was right and the number was still wrong.
The console's view of the link is a 2000 ms poll. Sampling the PHY's own registers at 250 ms (the measurement boundary here is MDIO register reads, not the driver's events) showed the real cycle: up 1.0 s at 100BASE-TX full duplex, down 1.5 s, period 2.5 s. A 2.5 s square wave sampled every 2.0 s beats at the least common multiple: exactly 10.0 s.
The console's 2-second poll reported 2 s up, 8 s down, 10 s period. The PHY's registers sampled at 250 ms show the true cycle: 1.0 s up, 1.5 s down, 2.5 s period. The 10-second cycle was the beat frequency between the fault and the poll.
The tell was visible in the very first capture: every transition timestamp landed on a multiple of the poll interval. A periodic-looking result at a suspicious multiple of your sample rate is a claim about your sampler first and the system second.
Two register facts from the same dump carried the diagnosis forward. ENERGYON never dropped, not once, including mid-outage, so the partner drives the pair continuously and the receive path passes energy. And symbol errors ran about 2 per link cycle, single digits, which retired the board's long-standing RMII-clock-margin suspicion on the spot: a marginal 50 MHz reference produces hundreds to thousands, not nineteen in 25 seconds. That retirement turned out to be too cheap. The hundreds-to-thousands threshold was assumed rather than measured, the same counter later climbed to 450 per second as the fault progressed, and the next day's layout review put the clock's actual route on the two-layer board (61 mm, two vias, no termination) back on the suspect list. The clock is back in scope for the bench.
The first conviction, and the experiment that overturned it
Autonegotiation completed cleanly every cycle; the link then died within a second. Autoneg runs on Fast Link Pulses: slow, widely spaced, easy to transmit. 100BASE-TX idle is 125 MBaud MLT-3. A transmit path too degraded for 125 MBaud can still pass FLPs perfectly, and dropping the link to 10BASE-T (about a sixth the symbol rate, larger amplitude, and a different line coding entirely, so a change of mode rather than of one variable) held rock solid: zero transitions, zero packet loss. Conclusion of night one: a degraded analog transmit path between the PHY and the jack. Prime suspects: the two TX bias resistors, their solder joints, the magnetics.
That conviction did not survive the second session's experiment. The LAN8720A's register 27 allows forcing the MDI/MDIX channel assignment manually, meaning which physical wire pair the transmitter actually drives. Sweeping {100M, 10M} × {MDI, MDIX} in 60-second legs, counted as link-down polls per 240 samples:
At 100 Mbit the link flaps hard on both forced channel assignments: 152 to 165 link-down polls out of 240 on MDI across three cycles, 152 to 155 on MDIX across two. At 10 Mbit both assignments hold: 7 to 8 down-polls, all within the leg-start renegotiation window, with zero ping loss.
Both physical pairs fail at 100 Mbit, at comparable rates: 152 to 165 down-polls of 240 on one assignment, 152 to 155 on the other. Both hold at 10 Mbit. A single drifted resistor, a cracked joint, or one damaged winding cannot do that: the two pairs share nothing downstream of the PHY except the analog supply, the magnetics package, and its cable-side common network. The pair-specific suspects were acquitted by symmetry; everything the pairs share stayed on the list.
What the fault actually looks like
The same sweep produced the finding that reframed everything. The 100 Mbit symbol-error rate is not constant: it ramps with time since the last PHY reset.
During the first session's flap the counter recorded 19 symbol errors in 25 seconds, under one per second. In the sweep session, seconds after a PHY soft reset it ran 12 to 90 per second; after roughly eight minutes of accumulated 100 Mbit operation across the sweep's legs it had climbed to 450 to 500 per second, on both pairs equally. A BMCR soft reset alone, with power never cycled and the board still warm, dropped it straight back to about 12 per second, and it began climbing again.
The sharpest control is the last bar: a register-level soft reset, no power cycle, board still warm, rewinds the error rate to its fresh value, and the climb starts over. That brackets the drift to state a reset clears, with a caveat the bench still owes an answer: a BMCR reset also drops the link and forces the partner to retrain, so "inside the PHY" is the leading reading, not the proven one. Bulk board temperature is not the whole story (the board stayed warm through the rewind), though junction-level and activity-driven heating remain live, which is what the freeze-spray step exists to test. The strongest correlate on the record is accumulated post-reset 100 Mbit operation, with the honest caveat that time, activity, and reset age were not varied independently. And a board that was healthy in the bench records on July 29 was unreachable on August 4. Whatever this is, it is getting worse.
This shape has an ugly corollary: every reboot makes the board look freshly plausible. "It works after a power cycle" means nothing when severity is a function of time-since-reset.
Root cause hypotheses
As of the night of August 4, ranked, each with the measurement that would convict it:
- The PHY itself (most likely). Internal analog drift, in bias, regulator, or adaptive state, matches a rate-dependent, pair-symmetric, soft-reset-rewindable fault. Discriminator: freeze-spray the chip alone mid-flap; if cold stabilizes the link or slows the ramp, the drift is thermal and local, though freeze also flexes the QFN's joints and exposed pad, so a positive convicts a region before it convicts the silicon. Repair is a hot-air QFN swap either way, which is also why a successful swap will not by itself say which it was.
- The bias-setting resistor (RBIAS). One 12.1k resistor sets every analog bias current in the chip; drift there is common-mode by construction. Discriminator: 30 seconds with a meter, board off; it should read ≈12.1k in circuit.
- The internal-regulator capacitors. The PHY's 1.2 V core comes from an internal regulator with two external caps. Discriminator: meter on the rail, expecting ~1.2 V flat; drift over the first minutes after power-up that tracks the error ramp is a conviction.
- The analog-supply network: never actually eliminated. Night one "eliminated power" by swapping the upstream supply. Both supplies fed the PHY's analog rail through the same ferrite bead and the same bypass network, so a fault in any of them was invisible to that test by construction. Discriminator: millivolts across the ferrite during a flap window, watched against the error ramp; a bead reading ohms end-to-end is dead.
- The RMII clock route, readmitted the next day. The two-layer board's 50 MHz reference travels 61 mm through two vias with no termination, and night one retired it on an assumed error threshold that the ramp's own 450 per second later embarrassed. Discriminator: the drive-strength sweep already built into the diagnostic firmware, run reset-bracketed, ideally with a scope on the pin.
The bench sequence, expected values written down before each probe, is queued as the next hands-on session.
The red herring with the biggest number
Across the investigation the board's ping ran anywhere from about 170 ms to nearly 400, depending on the hour and, it turned out, on the traffic. By the final controlled measurement it sat at 274 ms from a host on the same switch, where a healthy pole answers in under half a millisecond. It was tempting to hang that on the dying PHY too. It does not belong there.
With the full 16.79 Mbit/s multicast flood running, the board answers ping in 274 ms. Stop both multicast senders: 0.58 ms. Start them again: 274 ms. Same board, same firmware, no reflash between arms. Stopping only the larger sender leaves 4.3 Mbit/s offered and 2.5 ms.
The board's shipping firmware had (correctly) withdrawn 100BASE-TX and fallen back to a stable 10 Mbit link, a feature written and verified live during this same investigation. But the bench network carries 16.79 Mbit/s of unsubscribed multicast, flooded to every port, which is how a switch behaves when no IGMP snooping is in effect. That is 168% of a 10 Mbit link: the port's queue can never drain. Stop the flood and the "dying" board pings in half a millisecond. The full teardown of that interaction (including the retraction of a conclusion that had already been written into the docs) is its own post, Six sessions were queued for the same bench, and the bench was wrong. The board's fault never created that latency; it supplied the slow port the flood then saturated. The fallback itself does carry a real, smaller cost, about 6.5 ms of added wire time per full frame at 10 Mbit, which is the price of staying on stage and has nothing to do with the 274.
Flaws and corrections
The itemized list, because the instrument errors are the part that generalizes. Each one was committed or narrowly dodged in this investigation; each correction is now standing practice in the lab notes.
- Trusting a period the sampler cannot resolve. The console's 10.000 s cycle with zero drift was the beat between a 2.5 s fault and a 2.0 s poll. Correction: when every transition lands on the poll grid, measure faster; report the sampler's period as the sampler's, not the system's.
- A diagnostic that would have erased its own evidence. The obvious register to watch during a link flap has a latching link-status bit that the driver consumes to detect the flap. A second reader eats the latch and hands the driver a link that never dropped, while appearing to work. Correction, dodged by design: sample only non-latching registers, or latching ones the driver never touches.
- Claiming "changes exactly one variable" while changing two. The 10BASE-T discriminator restricted the advertisement and restarted autonegotiation, which can also re-resolve which physical pair the link rides, and the resolved channel is not readable. The experiment could not distinguish "rate-dependent" from "one bad pair," which was exactly the distinction the conclusion rested on. Correction: force the channel explicitly; invoking the one-change rule by name while breaking it moves the apparent evidence level up when it should move down.
- A counter that fossilizes instead of zeroing. The PHY's symbol-error counter neither counts nor clears at 10BASE-T: it re-serves its last 100 Mbit value forever. The dataset's "cleanest" reading, one lone error at 10 Mbit, was a stale byte read 240 times. All 240 reads returning exactly the same nonzero value is not what a physical noise process does; it is what a frozen register does. Correction: only 100 Mbit-mode values of that counter are real, and identical repeated reads of a "noisy" quantity are a red flag in any instrument.
- An elimination only as wide as what was varied. "Power eliminated" tested two upstream supplies that converge on the same downstream network. Correction: write down what a swap actually varies before crediting it with an elimination; the analog-supply network stayed on the suspect list for exactly this reason.
- A latch that outlives the experiment. The sweep's manual channel selection survives the driver's soft reset; the next session's measurements would have run with the channels silently crossed, and only a true power cycle reliably clears it. Correction: the diagnostic build now restores auto-negotiation of the channel at startup and logs when it found a stale latch.
- Reading a shared broadcast bus as a private line. The wire-side verification listens on a UDP port that every pole on the subnet broadcasts to; the first version-line captured belonged to a different board. Correction: filter by source IP, always.
- Confounding link speed with offered load. The 10 Mbit latency floor was measured carefully, with a concurrent quiet-LAN check that was honestly run and blind anyway: the capture ran on the Mac's Wi-Fi interface, a different physical path whose access point does not forward the bench segment's multicast groups. The wire carried a flood the check could not see. Correction: the 2×2 (speed × flood) separated the interaction, the conclusion in the docs was retracted with the measurement that replaced it, and a quiet-LAN check now names its interface and runs on the path under test.
- Letting one session's safety feature rewrite another session's experiment. The shipping link-fallback and the register sweep both write the same negotiation registers; run together they fight, and the sweep would have measured the fight. Correction: the fallback is gated off in sweep builds, and the build system now stamps every diagnostic knob into the artifact so an arm can be proven from the bytes on the board rather than from the source it was allegedly built from.
Where it stands
The board runs a shipping-config image with the 10BASE-T fallback and answers at its usual address. Read off the wire (the report port, filtered to this board's source IP):
PoleFX_version,Top Support,2026.07,7a8d1d44,Aug 4 2026 22:37:55,ota_0,valid
PoleFX_link,Top Support,mode=10F,fallback=1,downs=4
downs=4 is the entire story of the fallback working: the trip threshold is four link-down polls in a rolling minute, it reached exactly four, re-advertised, and the counter stopped; a fallback that had not worked would show it still climbing.
The repair is queued for the bench: meter the rails against the error ramp, freeze-spray the chip, meter the passives, and in all likelihood replace the PHY. The pass signal is already defined: link at 100 Mbit full duplex that stays, symbol errors flat at zero for ten minutes, and ping back to sub-millisecond from inside the bench segment with the flood accounted for, because the one thing this investigation kept re-learning is that a pass signal chosen after the repair tends to be whatever the repaired board happens to do.