If your ASIC runs fine for a while, then drops to 0 GH/s, and a reboot brings it straight back, you are almost certainly chasing a heat-related or power-related fault that gets worse the longer the machine runs. The reboot doesn’t fix anything — it just resets the clock on however long the miner takes to fail again. This guide walks the real causes in order of likelihood and gives you a bench procedure to isolate the fault instead of swapping parts blind.
Why “zero then recovers on reboot” is a distinct fault pattern
A hashrate that reads exactly zero (not just low or unstable) after a warm-up period is the firmware’s own protection logic doing its job. When an ASIC’s on-die temperature diode reports a chip crossing the thermal limit, the mining firmware holds every chip in reset and drives the domain voltage to 0 V — the board is powered but not hashing. On the next reboot the chips start cold, hash normally, and heat back up until the same trip fires again. The give-away is the time-to-fail: consistent runtime before the drop points at thermal buildup; instant or random drops point at power or a mechanical connection.
Tools you’ll need
- The miner’s web dashboard / kernel log (chip temps, per-chain status, error codes)
- A known-good PSU rated for your model, plus a multimeter for AC input and DC output
- Compressed air or a soft brush, and a clean workspace
- Optional but decisive: a thermal camera (FLIR C5 or similar, 160×120 minimum) and fresh thermal gel (Fujipoly SPG-30B)
- An ambient thermometer for the room the miner runs in
Cause 1 — High-temperature protection (most common)
This is the default explanation for “works cold, dies warm.” The firmware’s thermal-recovery routine trips when a sensed chip temperature crosses its limit (roughly the 75 °C+ chip-sensor range on Antminer-class hardware), holds the chains in reset, and zeroes the domain voltage. Track down why the board runs hot:
- Dust and airflow. Power down and blow out both heatsink faces and the fans. A clogged heatsink is the single most common cause. Confirm the fans actually spin up on boot and hit full RPM under load — a dashboard “fan speed” reading of 0 or a stuck fan will cook a board in minutes.
- Ambient temperature. Intake air should be well below the chip limit — aim for a hashroom ambient under ~35 °C. A miner that only fails in the afternoon heat is telling you the room, not the machine, is the problem.
- Dried or missing thermal interface. After heavy runtime the gel between chip and heatsink dries out and stops conducting. If cleaning and airflow don’t hold the temps, the fix is to pull the heatsink, clean off the old compound, and reapply fresh thermal gel. Never run a hashboard with a detached or dry heatsink even for a test — the die will overheat in seconds.
- Over-aggressive tuning. If the machine is overclocked or on an autotuned profile pushing high frequency, it generates more heat than the cooling can shed. Voltage and frequency on modern firmware are calculated at runtime by the autotuner, not fixed presets — dial the target hashrate/power down a notch and see if the trip stops.
The false-positive trap: a failed temperature sensor
A stuck or dead temperature sensor (I²C bus fault or a failed sensor IC) makes the firmware believe the board is overheating when it is stone cold, tripping thermal protection for no real reason. If your chip temps read impossibly high, wildly inconsistent between chains, or the board is cool to a thermal camera while the log screams overheat, suspect the sensor — not the cooling. This is a board-level repair, not a cleaning job.
Cause 2 — PSU degradation and voltage sag
If the miner status shows all chips as “x” (whole chains failing to enumerate) rather than a clean thermal trip, look hard at power. An aging or marginal PSU delivers full voltage cold, then sags as its own components heat under load — the hashboards brown out, chips drop off the chain, and a reboot recovers it until the PSU warms again. Same symptom, completely different root cause from a thermal trip.
- Verify AC input. Most S17/S19-class Antminer PSUs (APW9, APW9+, APW12) are rated for a 200–240 V AC input and their undervoltage protection shuts the unit down when AC drops below roughly 180 V. S9-era units (APW3++/APW7) accept a wider 100–264 V. A sagging wall circuit, a long/undersized extension run, or too many miners on one breaker can drag input voltage down under load and cause exactly this drop-and-recover behaviour. (Field note: APW9/APW9+ will physically run on 110/120 V despite the 200–240 V label, but at reduced output — under-fed input voltage makes sag far more likely.)
- Swap in a known-good PSU. The fastest isolation step: substitute a PSU you trust that is rated for the model. If the fault vanishes, the original PSU is the culprit.
- Check the DC side. With a meter on the PSU’s hashboard output under load, a healthy 12–15 V rail (APW12-class) that collapses as the unit heats confirms a failing supply.
Cause 3 — Loose power/ribbon connections and cold solder joints
Thermal expansion is mechanical. A power cable that is slightly loose, an oxidized bus-bar screw, or a hairline BGA/CLK-resistor solder crack can conduct fine cold and open as the board heats, then re-seat as it cools on reboot. Reseat both the heavy PSU cables and the 18-pin signal ribbons to each chain (power off, firm click, correct orientation). If a specific chain breaks at the same chip position every time it warms up, that points at a cracked joint or dead chip on that board — a hashboard repair, not a config fix. A thermal camera scan under load will show a cold chip among hot neighbours (dead chip) or a hot spot (short).
Bench procedure — isolate before you replace
- Log the time-to-fail and open the kernel log the instant it drops — note whether it’s a clean thermal trip or chains showing “x”.
- Power down, clean heatsinks and fans, confirm all fans reach full RPM.
- Watch per-chip temps for one full warm-up cycle. Impossible/inconsistent readings ⇒ sensor fault. Even, climbing temps that hit the limit ⇒ real cooling problem.
- If chains drop as “x” with normal temps, swap the PSU for a known-good unit and re-test.
- Reseat power and ribbon cables; if one chain always fails at the same chip when warm, flag that board for hashboard-level repair.
- Change one variable at a time and re-run the full warm-up — never swap two parts at once or you’ll never know which one mattered.
How to confirm it’s actually fixed
A reboot proving the miner hashes is not a fix — it always hashes cold. The only valid confirmation is a sustained run: let the machine hold its full target hashrate for several hours past its previous time-to-fail, with all chips enumerated, stable chip temps below the limit, and no thermal or chain errors in the log. If it clears its old failure window and stays flat, the fault is closed.
When to escalate
If temps are sane, the PSU is known-good, cables are reseated, and a chain still zeroes out when warm, you’re into board-level territory — dead chips, cracked BGA joints, failed temperature sensors, or a shorted power stage. That’s a workbench repair with the right test fixtures, not a field fix.
Related: Look up your exact fault symptom in the ASIC fault finder, follow the model-specific teardown in our repair manuals, read the deep-dive on thermal imaging for hashboard faults, or if the board needs bench work, start a repair and source components from our ASIC repair parts catalog.
