Unplanned downtime is the single most expensive line item most hosted mining operators never see coming. A single ASIC that overheats silently for three weeks before it trips offline doesn’t just lose that machine’s hashrate — it signals that dozens of neighboring units on the same power and cooling loop are drifting toward the same failure. Reactive maintenance, where a technician responds after a miner goes dark, is the most expensive way to run a fleet. Predictive maintenance — using thermal imaging, vibration analysis, and transformer oil testing to catch degradation before it causes an outage — is how large-scale hosting operations protect uptime without over-staffing the floor.
Why Reactive Maintenance Fails at Scale
In a garage with ten miners, a technician can eyeball every unit weekly and catch problems by ear and by hand. At hundreds or thousands of units per facility, that approach breaks down. Fan bearing wear, thermal paste degradation, and loose power connections all produce measurable warning signs weeks before a unit actually fails — but only if someone is instrumented to see them. Without a monitoring program, the first signal most operators get is a drop in pool reported hashrate, by which point the hardware has often already sustained damage.
ASIC chips are generally rated to run in the 60–85°C range under normal load. Sustained operation above roughly 90°C accelerates electromigration in the chip’s internal interconnects — a slow, cumulative degradation that permanently reduces achievable hashrate even after the unit is brought back into a healthy thermal range. That means a hashboard limping along at elevated temperature for a month isn’t just underperforming today; it is losing terahashes it will never get back.
Thermal Imaging: Finding Heat Before It Finds You
Infrared thermography is the fastest way to screen a row of miners for developing problems. A handheld thermal camera or a fixed infrared sensor array reveals hotspots invisible to the naked eye — a hashboard running several degrees hotter than its neighbors, a heatsink with failed thermal paste, or a power connector developing resistance from a loosening crimp.
- Hashboard hotspots: uneven chip temperatures across a board usually indicate a failing fan, blocked airflow, or a chip beginning to draw abnormal current.
- Power connector heating: a PDU or PSU connector that runs noticeably hotter than identical connectors nearby is a loose-connection fire and failure risk, not just an efficiency loss.
- Immersion and hydro-cooled loops: thermal imaging of manifold inlets and outlets catches flow restriction and localized fouling before it shows up as a fleet-wide temperature creep.
A monthly or quarterly thermal pass across every row, logged and compared against baseline images from when the row was commissioned, turns an invisible problem into a flagged work order days or weeks before the unit would otherwise trip offline. For facilities running at scale, this is a far cheaper investment than the hashrate lost to even a handful of silent thermal failures. Our own facilities run tiered cooling — see how that pairs with immersion economics in our breakdown of immersion cooling ROI timelines.
Vibration Analysis: Listening for Bearing Failure
Every ASIC miner depends on high-RPM fans to move enormous volumes of air through a dense hashboard stack. Fan bearings are consumable parts, and bearing wear produces a measurable vibration signature long before the fan actually fails. Industrial predictive-maintenance programs commonly use piezoelectric accelerometers, following vibration severity guidance such as ISO 10816/20816, to flag bearings trending toward failure.
At mining-facility scale, this doesn’t require instrumenting every single fan with a dedicated sensor. A rotating handheld vibration check during routine floor walks, combined with acoustic monitoring for the telltale grinding or whine of a failing bearing, catches the large majority of fan failures before they cascade into a thermal event. Replacing a $15 fan on a scheduled basis is trivial next to replacing a hashboard that cooked itself because a dead fan went unnoticed for two weeks.
Transformer Oil Testing: Protecting the Electrical Backbone
Thermal and vibration monitoring protect individual miners. Dissolved gas analysis (DGA) of oil-filled transformer insulation protects the electrical infrastructure an entire facility depends on. As a transformer ages or experiences internal stress — overheating, arcing, or insulation breakdown — it generates characteristic dissolved gases (hydrogen, methane, acetylene, and others) that dissolve into the insulating oil in predictable ratios.
Periodic DGA sampling, typically done annually for critical mining-facility transformers and more frequently for units showing early warning signs, catches developing faults — overheated windings, partial discharge, arcing — months or years before they progress to catastrophic failure. A transformer failure at a mining facility isn’t a single-unit problem; it can take an entire row or building offline until a replacement unit is sourced and installed, which can take weeks. For the electrical fundamentals DGA programs are built around, see our guide to transformer vector groups and power quality.
Building a Predictive Maintenance Cadence
A practical program for a hosting-scale facility combines three cadences:
- Continuous: automated hashrate and temperature telemetry per unit, flagging statistical outliers against the row average in real time.
- Monthly/quarterly: thermal imaging passes and spot vibration/acoustic checks across every row, logged against baseline.
- Annual (or trend-triggered): dissolved gas analysis on critical transformers, oil testing, and switchgear thermal scans during scheduled maintenance windows.
None of this requires exotic tooling — a quality handheld thermal camera, a basic vibration meter, and a relationship with an electrical testing lab for DGA sampling cover the large majority of failure modes that take ASICs and infrastructure offline. The return on that modest investment is measured in avoided downtime, extended hardware lifespan, and fewer emergency electrical repairs — the difference between a facility that loses a few percent of fleet uptime to attrition and one that loses it to a 2 a.m. breaker trip nobody saw coming.
What the Program Actually Costs
Operators sometimes assume predictive maintenance means an expensive sensor-network retrofit. In practice, the tooling floor is low. A commercial-grade handheld thermal camera suitable for row-by-row screening typically runs a few hundred to a couple thousand dollars depending on resolution and sensitivity. A basic vibration meter for spot-checking fan bearings costs even less. The largest real cost is labor time — building the habit of walking every row on a fixed schedule, logging readings, and comparing them against a baseline rather than relying on memory.
Dissolved gas analysis is the one component that typically requires an outside electrical testing lab, since it involves drawing an oil sample and running laboratory chromatography. Even so, annual DGA testing on a facility’s critical transformers usually costs a small fraction of what a single unplanned transformer replacement and the associated downtime would cost. Framed as insurance against a multi-day outage affecting an entire building, the economics are straightforward.
Common Mistakes That Undermine a Monitoring Program
- No baseline. A thermal or vibration reading is only useful compared against what that same unit looked like when it was healthy. Facilities that skip baseline imaging at commissioning lose the ability to spot gradual drift later.
- Inconsistent intervals. A quarterly thermal pass that slips to “whenever someone has time” stops catching problems before they become outages. Treat the schedule like a compliance requirement, not a nice-to-have.
- Ignoring outliers that “still work.” A hashboard running 8°C hotter than its row average but still hashing normally is not a non-issue — it’s the earliest and cheapest point to intervene, before performance actually degrades.
- Treating DGA as one-and-done. A single clean oil test tells you the transformer was healthy on that date. The value comes from trending results year over year to catch slow degradation.
Hosting Where Maintenance Discipline Is Built In
Facility-level maintenance discipline is part of what separates reliable hosting from a warehouse full of unmonitored machines. Rax Mining’s hosted colocation includes 24/7 facility monitoring, scheduled maintenance, and infrastructure built around uptime — not just cheap power. If hardware failure and unplanned downtime are eating into your mining returns, explore our NatGas MDU colocation options or talk to our team about what a monitored, maintained hosting environment looks like for your fleet.
Explore Rax Mining
- Bitcoin Miner Hosting — Competitive rates from $0.075/kWh
- NatGas MDU Units — 1MW modular datacenter containers
- Mining Profitability Calculator — Estimate your mining returns
- Our Facility — Tour our mining infrastructure

