Why stable temperatures could be hiding a cooling problem

Rupesh Mainali
Rupesh Mainali
Senior Member of Technical Staff at Reliability Engine

A liquid-cooled rack can appear perfectly healthy while its operating margin quietly disappears. Rupesh Mainali, Senior Member of Technical Staff at Reliability Engine, explains what operators should be looking for.

A liquid-cooled rack can be within temperature limits and still be moving away from its commissioned state. During a morning review, supply and processor temperatures are green. Yet the coolant distribution unit (CDU) is driving its pump harder than at handover, filter differential pressure has risen at the same flow, and a top-up appears only in the maintenance log. Nothing is overheating. Something has changed.

The controls are doing their job. Pumps and valves compensate for variation and preserve the thermal result, potentially hiding the early stages of a restriction, flow redistribution, sensor drift or coolant change. By the time temperature moves, the simplest maintenance window may have narrowed.

Stable temperature is not the same as stable condition

Temperature remains the first line of protection. It shows whether the system is meeting its immediate thermal duty, but not how difficult that duty has become. A variable-speed pump can offset rising resistance, a CDU valve can compensate at the facility-water boundary, and a rack average can hide one branch losing flow.

The better question is whether the loop is producing the same result with the same effort. Once a pump or valve reaches the end of its range, temperature will move, but the opportunity to intervene without disturbing compute may have passed.

Compare the result with the effort behind it

This does not require another wall of charts. Pair flow with differential pressure, pump command with flow and valve state, component-to-coolant temperature difference with power, and fluid results with fills, top-ups, filter changes and component service. Each pair answers a question that temperature alone cannot.

Compare like with like. Pressure drop normally increases with flow. If flow, valve position and coolant temperature are unchanged, however, a higher pressure drop signals a different condition. A warmer processor during a heavier run is expected; a wider processor-to-coolant temperature difference at comparable power and flow needs investigation.

Location matters. CDU total flow can remain acceptable while one rack or tray branch falls behind, and a single return temperature can hide the hottest path. Branch measurements and peer comparison expose local changes that a fleet average smooths away.

Know how coolant risks enter the loop

Coolant risk often enters during ordinary work. A top-up can introduce oxygen or change the inhibitor or glycol concentration. Replacing a hose or cold plate can introduce debris, leave trapped gas or add a new wetted material to the loop. Construction residue and corrosion products can load filters or collect in narrow passages. None necessarily changes temperature immediately.

Fluid checks should match the coolant, treatment plan and materials. Depending on the loop, they may include pH, conductivity, inhibitor or glycol strength, particles or turbidity, microbial activity, dissolved oxygen and dissolved metals. Define the sample point, fluid temperature, collection method and confirmation step. A sample from a stagnant leg may not represent the circulating fluid.

These are diagnostic signals, not automatic shutdown inputs. Use them to decide whether to repeat a sample, inspect a filter, review a top-up or check a serviced branch.

Turn commissioning into a usable baseline

A handover pack can contain hundreds of readings and still leave the shift team without a useful reference. Operators need a repeatable definition of normal. After flushing, filling, venting, filtration, balancing and sensor checks, record stable points at low, typical and high load. Include rack power, temperatures, flow, differential pressure, pump command, valve state and filter differential pressure.

Keep coolant identity, fluid-condition results, sample details and sensor calibration status with those points. The goal is not a perfect day-one report. It is a comparison the team can reproduce after a rack expansion, filter replacement, controls change or hardware refresh.

Issue a new, dated baseline after rebalancing, a pump-sequence change or a rack addition. Otherwise, today’s plant is being judged against a configuration that no longer exists.

Put maintenance events on the same timeline

Many cooling anomalies begin with an event recorded elsewhere: a hose change, filter replacement, top-up, new rack or controls update. If those events remain in work orders while telemetry remains in the CDU or building platform, the graph loses its most useful context.

Add the time, affected loop, action, component, fluid identity and quantity to the trend history. Run an event-based check after a fill, top-up, leak repair, component replacement, prolonged shutdown, abnormal fluid loss or coolant-batch change. Calendar-only sampling misses some high-risk moments.

Make every alert actionable

Before enabling an alert, define the failure mechanism, misleading conditions, independent confirmation and next action. Filter differential pressure must be read with flow. A fluid result needs a confirmation method. A branch-flow alert needs a team that can inspect the path.

Reserve automatic protection for conditions that threaten hardware faster than a person can investigate, such as critical flow loss, a confirmed leak or a thermal limit. Slow fluid-condition changes usually call for confirmation and planned intervention. Giving both the same severity creates outages or teaches operators to ignore alarms.

Ownership must cross the facilities and IT boundary. Facilities may own the facility-water connection and part of the CDU; hardware teams may own the technology loop, rack service and component telemetry. Agree on shared measurements, the escalation path and the evidence record during handover, not during the first incident.

A ten-minute review while the dashboard is green

  • Match the operating point. Compare similar rack power, coolant supply temperature, flow and control mode.
  • Check control effort. Look for higher pump command, filter pressure drop or valve travel at the same flow.
  • Compare peers. Find the rack, tray or branch that no longer follows otherwise similar equipment.
  • Review recent work. Check fills, top-ups, filter changes, component service, controls updates and new rack connections.
  • Choose a confirmation. Verify the sensor, inspect the path or collect a controlled sample before escalating the response.

No single trend proves a failure. Confidence comes from independent evidence that points to the same mechanism: more pump effort without more flow, a branch separating from its peers, a thermal path worsening at the same power, or a fluid result changing after service. Confirm the change, choose the smallest safe intervention and preserve the record.

Temperature should stay at the centre of protection. By also watching the effort behind that temperature, operators gain time to act while the rack is still doing its job.

Categories

Related Articles

More Opinions

It takes just one minute to register for the leading twice weekly B2B newsletter for the data centre industry, and it's free.