A heat pump is a rotating machine plus heat exchangers plus controls, so nothing in it is new to a reliability engineer. The compressor is the heart and the biggest risk; the refrigerant charge, the heat exchangers and the expansion device are the slow-degrading rest.
Almost every fault announces itself early through a small set of vital signs: suction and discharge pressure, superheat, subcooling, discharge temperature, compressor current, vibration, condenser approach, and the COP/EER trend. Read together, they tell you which fault is developing, not just that something is wrong.
Each of those signals is a P-F curve waiting to be trended. Watch it continuously, alarm at the potential-failure point P, and you intervene on a planned work order long before functional failure F. That is exactly what the Bluestream PdM platform does โ and its live demo already monitors an air-handling unit.
1 · HVAC-R is just rotating & heat-transfer assets
It is tempting to treat refrigeration as its own specialism with its own vocabulary โ superheat, subcooling, TXV, lift โ and stop there. But reframe it and the mystery evaporates. A heat pump, a chiller, a rooftop unit and a cold-store pack are all the same three things bolted together:
- A compressor โ a rotating (or reciprocating) machine driven by an electric motor. Everything the Academy teaches about pumps, bearings and motors applies directly.
- Two heat exchangers โ the evaporator and the condenser, subject to the same fouling and approach-temperature physics as any other exchanger.
- A controls loop โ an expansion device plus sensors and a controller that hold superheat and capacity where they belong, with all the instrument-drift and loop-tuning failure modes that come with control.
Once you see it that way, the whole condition-based-maintenance framework snaps into place. The refrigerant cycle from the opening guide tells you what should be happening; condition-based maintenance tells you how to notice, early, when it stops. This article is the bridge between the two: the failure modes of an HVAC-R asset, and the signals that catch each one on the way down.
The organising idea is the P-F curve. A fault does not appear at failure โ it appears at a potential-failure point P, degrades along a curve, and only reaches functional failure F later. The gap between them is the P-F interval: your warning time. Condition monitoring exists to find P, and the whole game is choosing signals with a long, readable P-F interval. If that idea is new, read the predictive-maintenance primer first โ everything below is an application of it.
2 · The compressor — the heart and the biggest risk
The compressor is the one component that turns electricity into work, runs hot, spins fast and cannot be repaired in place. It is both the most expensive part to replace and the one whose failure strands the whole machine. It earns the highest criticality rating in the system and the closest watch. Treated as a machine in its own right, its failure modes are familiar:
- Valve wear / broken reeds (reciprocating and scroll) — leaking or fractured discharge/suction valves let high-pressure gas re-expand, so the machine pumps less mass per stroke. Capacity falls, discharge temperature climbs, and the pressure ratio drifts. On a scroll, tip-seal or flank wear does the same.
- Bearing wear — the same story as any rotating asset. Rising vibration at bearing defect frequencies is the classic leading indicator, with a P-F interval of weeks to months if you are trending it.
- Motor-winding insulation breakdown — a hermetic compressor's motor is cooled by the returning suction gas, so anything that raises discharge temperature or starves the return also bakes the windings. This is the same insulation-ageing story told in motor overheating: every 10 °C over the insulation-class limit roughly halves winding life. Watch current imbalance, insulation resistance and running temperature.
- Liquid slugging — the refrigeration-specific killer. If superheat collapses โ from a flooded evaporator, an overcharge, or a failed expansion valve โ liquid refrigerant reaches the compressor. Liquid does not compress, so a slug spikes cylinder or scroll forces far beyond design, snapping reeds, bending rods or cracking scrolls. It also washes oil off the bearings. Low superheat is the upstream warning; a sudden vibration or current transient is the event itself.
- Oil loss / poor lubrication — refrigerant carries oil around the loop, and it must return. Oil logging in a flooded evaporator, migration during long off-cycles, or refrigerant dilution of the oil all starve the bearings. Wear metals and viscosity shift show up in oil analysis where sampling is feasible; on sealed machines, vibration and temperature are the proxies.
- Short-cycling — a compressor that starts and stops too often never settles into steady lubrication, and every start is a thermal and mechanical shock (the same reasoning as motor starting). Causes range from an oversized machine to a fouled coil tripping on pressure. Starts-per-hour is itself a condition signal worth counting.
How the compressor tells on itself
Three signal families cover almost all of it. Vibration catches the mechanical faults โ bearings, looseness, imbalance, and the shock of slugging โ exactly as covered in the vibration-analysis guide. Motor current (and its signature analysis, MCSA) catches electrical faults, load changes and valve problems without a single wire into the refrigerant circuit. And discharge temperature is the cheap, powerful catch-all: it rises with valve leakage, undercharge, high lift and poor motor cooling, so a discharge-temperature trend flags a remarkable share of developing faults early.
3 · Refrigerant charge faults
The mass of refrigerant in the loop is a design quantity, and the cycle only performs at its design charge. Both directions off it are faults, and the diagnostic pair that separates them is the one introduced in article 1: superheat (how far the suction vapour is above its boiling point โ proof it fully boiled) and subcooling (how far the liquid leaving the condenser is below its condensing point โ the primary read on charge).
Undercharge — almost always a leak
Refrigerant does not get consumed, so a falling charge means a leak. As charge drops you see low subcooling (not enough liquid to stack up in the condenser), high superheat (the evaporator runs dry at the outlet), reduced capacity, and a rising discharge temperature as the starved, superheated suction gas gives the motor less cooling. A slow leak is the textbook slow P-F degradation: subcooling drifts down over weeks, capacity sags, discharge temperature creeps up โ a long, gentle slope you can trend and act on with a planned recovery-and-recharge before the compressor over-temps. Left alone, it ends in a starved, overheating compressor and an emergency call-out.
Overcharge — too much liquid
Too much refrigerant backs liquid into the condenser, raising subcooling and head pressure. High head pressure increases the lift the compressor fights (so power rises and COP falls), pushes discharge temperature up, and โ if liquid backs far enough โ threatens the low-superheat slugging described above. Overcharge and a fouled condenser look similar on head pressure alone; subcooling is what tells them apart (high subcooling points to charge, near-normal subcooling with high head points to fouling).
Superheat and subcooling are a coordinate pair, not two numbers. Superheat reads the evaporator/expansion-valve side; subcooling reads the condenser/charge side. Plot where a fault moves each one and the diagnosis is usually unique: high superheat + low subcooling → undercharge; low superheat + high subcooling → overcharge; high superheat + normal subcooling → a starving expansion valve, not a charge problem. Streamed continuously, that pair is a live diagnosis rather than a one-off gauge reading.
4 · Heat-exchanger fouling
The evaporator and condenser are just heat exchangers, and they degrade the way all heat exchangers do โ by losing the ability to move heat across the wall. In a heat pump that loss is expensive twice over: it wrecks efficiency and it drives the compressor into a punishing lift.
Condenser side — lift goes up, COP falls
An air-cooled condenser fouls with dust and debris; a water-cooled one scales or biofouls; either way airflow or water flow can also simply drop (a dirty filter, a failing fan, a throttled pump). Heat rejection worsens, so the refrigerant has to condense at a higher temperature to shed the same duty. That raises head pressure and the temperature lift, and โ straight off the Carnot relationship from article 1 โ COP falls and power rises. Where the condenser rejects to a cooling tower, the tower's own fouling and approach degradation feed straight into this, which is why the cooling-tower guide matters to compressor health.
Evaporator side — the cycle gets starved
Evaporator fouling, low airflow, or a blocked filter starves the low-pressure side: suction pressure and temperature drop, capacity falls, and on a cold coil you get frosting that insulates the surface and chokes airflow further โ a self-reinforcing spiral that can end in the low-superheat, liquid-return regime that threatens the compressor.
The on-condition signals
You do not need to open a coil to know it is fouling. The approach temperature โ the gap between the refrigerant's saturation temperature and the air or water leaving the exchanger โ widens as heat transfer degrades; it is the cleanest single measure of exchanger cleanliness. And the COP/EER trend, computed from the same data, falls steadily. A gently rising condenser approach with a gently falling COP over a season is a fouling P-F curve you can read weeks ahead of a high-head-pressure trip, and schedule a coil clean on your terms.
5 · Expansion device & controls
The expansion valve is the cycle's throttle, and the controller that drives it is trying to hold superheat at a setpoint low enough for efficiency but high enough to keep liquid out of the compressor. When that control misbehaves, the whole cycle wanders.
- TXV / EEV hunting — the valve over- and under-feeds the evaporator in an oscillation the controller cannot damp, so superheat swings instead of holding. Beyond the efficiency loss, the low excursions of a hunting cycle are exactly when liquid can reach the compressor. Hunting is a control-stability problem โ the same family as the oscillation and integral-windup issues in erratic controller output โ and it shows up as periodic swings on superheat, suction pressure and current.
- Sensor drift feeding the controller bad data — the controller is only as good as the superheat it thinks it sees. A drifting suction-temperature or pressure sensor makes it hold the wrong superheat: drift low and it starves the evaporator (high real superheat, lost capacity); drift the other way and it floods toward the compressor. This is the classic instrument-drift hazard โ the reading looks fine while the real quantity walks away โ and it is why sensor calibration is itself a condition-monitoring task. Cross-checking two independent signals (does measured superheat agree with the pressure-temperature pair?) is how you catch a drifting sensor before it drives a real fault.
6 · The vital-signs table
Here is the whole article in one place: the signals a monitored heat pump exposes, what each one tells you, and which faults it catches. This is the fault-signature map โ read a row to understand a sensor, read a column of the last table to work backwards from a symptom to a cause.
| Signal | What it tells you | Faults it catches |
|---|---|---|
| Suction pressure / temp | The state of the low-pressure side and how well the evaporator is fed. | Evaporator fouling, low airflow, frosting, undercharge, a starving or hunting expansion valve. |
| Discharge pressure (head) | How hard the high side is working to reject heat; sets the lift. | Condenser fouling, low condenser flow, overcharge, non-condensables in the loop. |
| Superheat | Whether the suction is fully boiled and safely clear of liquid; the evaporator/valve-side vital sign. | Undercharge (high), overcharge or flooding (low → slugging risk), expansion-valve hunting or failure. |
| Subcooling | Whether the liquid fully condensed; the primary read on refrigerant charge. | Undercharge (low), overcharge (high), condenser not rejecting (high with high head). |
| Discharge temperature | The cheap catch-all: rises with almost anything that stresses the compressor. | Valve leakage, undercharge, high lift, poor motor cooling, high pressure ratio. |
| Compressor current | Load, electrical health and โ via MCSA โ mechanical faults reflected into the motor. | Winding faults, current imbalance, valve/load problems, short-cycling, locked rotor. |
| Vibration | The mechanical health of the rotating element; the earliest warning on bearings. | Bearing wear, imbalance, looseness, misalignment, and the shock transient of liquid slugging. |
| Condenser approach | Cleanliness of the high-side heat exchanger, independent of load. | Condenser / cooling-tower fouling, scaling, low airflow or water flow. |
| COP / EER trend | The whole-machine efficiency roll-up; the slow integrator of every degradation. | Fouling, charge drift, worn valves, rising lift — anything that quietly costs efficiency. |
Every one of these is a P-F curve. A signal sits in a normal band (healthy), begins to drift when a fault initiates (the potential-failure point P), and degrades along a slope until the machine can no longer do its job (functional failure F). The art is picking signals whose P-F interval โ the warning time between P and F โ is long enough to plan around. A slow leak read on subcooling gives you weeks; a bearing read on vibration gives you weeks to months; a slugging event read only on vibration gives you seconds, which is why you also watch its upstream cause, superheat. If the P-F idea needs refreshing, it is the backbone of the predictive-maintenance guide.
How hard you watch each signal is a criticality decision, not a uniform rule. A single rooftop unit over a storeroom might warrant a monthly trend review; the chiller cooling a data hall, or a process refrigeration pack whose loss stops production, earns continuous streaming with tight alarms. The disciplined way to set that per asset is a FMECA โ list the failure modes above, score their effect and likelihood, and let the risk ranking decide which signals are streamed, which are checked on rounds, and which are left to run-to-failure because the consequence is trivial. That is condition monitoring applied with judgement rather than blanket sensors.
7 · From signal to work order
A trend on a screen is not maintenance. The value only lands when a drifting signal becomes a planned intervention that happens before the failure. The chain is short and it is the same for every fault above:
- Trend continuously. Stream the vital signs, establish each asset's healthy baseline, and let the history define what normal looks like for this machine at these conditions โ because a "normal" head pressure depends on the day's lift.
- Alarm at P. When a signal crosses out of its band โ subcooling sliding, approach widening, vibration rising, superheat swinging โ raise a condition alarm at the potential-failure point, not at the trip.
- Diagnose with the pair. Use the vital-signs table to turn the symptom into a cause: high head with high subcooling is charge; high head with normal subcooling is fouling. The diagnosis writes the work scope.
- Plan the intervention before F. Because you found it at P, you have the P-F interval to work in โ order the part, book the window, do the coil clean or the leak repair or the bearing change on your schedule, not at 2 a.m. on a failure.
This is precisely what the Bluestream PdM platform is built to do for these assets: ingest the streaming signals, hold each asset's baseline and criticality, trend toward the alarm at P, and raise a work order with the evidence attached. The platform's live demo already monitors an air-handling unit โ a compressor, coils and controls, exactly the machine this article describes โ so the loop from "a signal drifted" to "a scheduled work order" is not a diagram, it is running. The heat-pump series began by explaining how the cycle should behave; it ends here, showing how to notice โ early, and on your terms โ when it stops.
Key takeaways
- A heat pump is a rotating machine plus heat exchangers plus controls โ so the entire condition-based-maintenance toolkit already applies, with no special HVAC-R magic required.
- The compressor carries the highest risk; watch it with vibration, motor current and discharge temperature, and guard against liquid slugging by watching superheat upstream.
- Superheat and subcooling are a diagnostic pair โ together they separate undercharge, overcharge, fouling and valve faults; a slow leak is a textbook slow P-F degradation.
- Fouling shows up as a widening approach and a falling COP long before a high-head trip โ a P-F curve you can schedule a coil clean against.
- Every vital sign is a P-F curve; criticality and a FMECA decide how hard you watch each one, and the platform closes the loop from alarm at P to planned work order before F.