The Systems That Pass Commissioning but Fail Under Real Operating Conditions

PASSED. NOT PROVEN. The weakness shows up after turnover.

Part 2 of 4 —Accepted is not the same as proven.

The commissioning report looked clean.

Pumps started. Valves stroked. Alarms came in. The chiller hit leaving-water temperature. The standby unit rotated on command. The checklist was signed, filed, and handed to the owner.

Three months later, during the first real heat event, the same system could not hold stable operation for six hours.

That is one of the most common cooling system commissioning gaps in mission-critical facilities. The system passed the test. It was accepted. It was not proven.

There is a difference.

1. Commissioning proves a condition, not the whole operating life

Commissioning is necessary. Nobody serious argues against it.

But commissioning usually tests the system under defined conditions: controlled load, known ambient, clean strainers, fresh filters, new sensors, tuned controls, available contractors, and a project team still watching every point on the screen.

That is not how a data center, lab, imaging suite, semiconductor support system, or ATE floor operates after turnover.

Real operation brings drifting loads, dirty filters, operators responding to other issues, valves that do not land where the graphic says they landed, glycol chemistry changes, towers carrying biological load, and standby equipment that has not run under pressure since acceptance testing.

A commissioning checklist tells you what happened during the test window. It does not tell you how the system behaves after six weeks of partial-load operation, after the first maintenance transition, or after the first hot afternoon when every connected load asks for cooling at the same time.

That is where accepted systems start showing their real character.

2. Design intent, commissioning conditions, and operating behavior are three different things

Design intent is the engineer’s target. It lives in drawings, sequences, calculations, pump curves, and equipment schedules.

Commissioning conditions are the conditions available during startup. They live in temporary load banks, seasonal weather, staged test scripts, and whatever facility state exists before full occupancy.

Operating behavior is what the system does when nobody is watching.

Those three are not always aligned.

I have seen systems designed for stable full-load operation spend most of their lives at 35–55% load. At that point, a pump that looked fine on paper runs left of its best efficiency point, differential pressure control starts hunting, and the plant burns operator attention trying to correct a problem that began as a design assumption.

I have seen standby chillers start exactly as specified, then fail to stabilize because the control sequence never gave the system enough time to settle before opening load. The test showed “start successful.” The building saw supply temperature swing.

I have seen medical imaging cooling loops pass initial flow verification, then lose margin after strainers collected enough construction debris to drop pressure below what the MRI cooling skid needed during peak scan schedules.

The issue was not that commissioning missed everything. The issue was that the test condition was too clean, too short, and too narrow to represent real life.

3. The failures usually appear after turnover

A lot of cooling failures do not show up on day one. They show up after the construction team leaves, the warranty clock starts, and the operating team inherits a system that still has project DNA inside it.

That is when the weak spots surface.

A filter loads up. A strainer plugs. A control valve begins to lose stroke range. A sensor starts reading two degrees low. A standby pump starts, but the check valve does not seat cleanly. A bypass valve that was quiet at startup begins to fight the primary control loop. A tower basin picks up debris after the first storm. The glycol loop chemistry shifts. The operator rotation changes, and the new team follows the written sequence without knowing what normal sounds like.

None of those items feel dramatic by themselves.

Together, they turn a passed system into an unstable one.

In an AI data center, that instability shows up as thermal throttling, where processors slow themselves down because temperatures are climbing. In a lab, it shows up as a process tool that cannot hold temperature band. In a hospital imaging suite, it shows up as cancelled scans and a service call that starts with the phrase nobody wants to hear: “The chiller is running, but we still do not have cooling.”

That is the part owners need to understand. The chiller running is not the same as the system delivering stable cooling to the load.

4. Short-duration tests miss long-duration problems

A 20-minute test will not prove a 20-year system.

Short-duration tests are good at finding hard failures. They catch the pump that does not start, the actuator wired backward, the alarm that does not report, the sensor that is obviously wrong, and the control point that is mapped to the wrong graphic.

They are not as good at finding slow instability.

A variable-frequency drive on a secondary pump cycling between 42 and 61 Hz every 90 seconds tells you something about the control loop, the pump curve, and the operating envelope. But if the test script only verifies that the pump starts and responds, that behavior gets missed.

A return water temperature climbing from 50°F to 58°F over the course of an afternoon tells you something about heat rejection, load profile, flow balance, and margin. But if the system hits setpoint during a staged test, the commissioning report still looks clean.

A plugged strainer does not always announce itself on day one. It builds pressure drop. It steals flow. It raises delta-T, the temperature difference between supply and return water. The system keeps running, so nobody treats it like a failure until the load notices.

That is why trend duration matters. Cooling reliability is not proven by seeing the right value once. It is proven by seeing the system hold the right behavior across time, load, rotation, and disturbance.

5. Operators know things the checklist does not capture

Good operators see patterns before the report shows a failure.

They know when a tower fan sounds wrong. They know when a pump takes too long to stabilize. They know when a valve is “always around 80% open now” even though nobody changed the setpoint. They know when a standby chiller technically starts but makes the room nervous.

That feedback belongs in a reliability review.

Trend data matters. Alarm history matters. Operator notes matter. Maintenance records matter. So do the quiet comments people make during a walkthrough: “That one has been noisy since startup,” or “We try not to run that pump unless we have to,” or “The BAS says it is fine, but the room gets warm every Friday afternoon.”

Those comments are not complaints. They are field data.

A reliability-focused review takes that field data and compares it against actual system behavior. It looks at trends, not snapshots. It compares design flow to measured flow. It checks whether standby equipment reaches stable operating temperature, not just whether it starts. It looks for nuisance alarms that operators have learned to ignore. It reviews whether the control sequence works at night, at partial load, during maintenance, during lead-lag rotation, and during heat events.

That is where the hidden risk usually sits.

6. Edge cases are not rare in real facilities

A lot of teams treat edge cases like unusual events. In mission-critical cooling, edge cases become normal sooner than people expect.

A hot day combined with a maintenance window. A standby start during partial load. A utility interruption followed by a rapid restart. A process load coming online while filters are already loaded. A lab expansion tied into an older loop. A semiconductor support system operating outside the original load profile. A medical imaging suite adding scan volume without revisiting cooling margin.

Each event looks reasonable by itself.

The system fails when two or three of them land together.

That is why “it passed commissioning” is not enough of an answer when the owner asks whether the system is ready. Ready for what? Ready for the scripted test? Ready for full load? Ready for low load? Ready for August? Ready for maintenance? Ready for a failed sensor? Ready for a plugged strainer and a standby start at the same time?

Those are different questions.

A system that passes one of them has not passed all of them.

7. What a reliability-focused review looks for

A reliability-focused review does not repeat the commissioning script and call it insight.

It asks how the system behaves now.

It looks at actual trend data: temperatures, valve positions, pump speeds, differential pressure, alarm frequency, equipment rotation, starts, stops, and recovery time. It compares the control sequence against what operators actually do. It checks whether maintenance activities change system stability. It looks for components that are “working” but no longer working with margin.

That includes strainers, filters, check valves, balancing valves, pressure sensors, temperature sensors, chemical treatment, tower performance, pump control, bypass behavior, and standby sequencing.

It also asks a simple question owners should ask more often:

What has to be true for this system to keep the load safe when conditions are not ideal?

That question changes the conversation.

It moves the team away from paperwork and toward resilience. It separates equipment status from system performance. It exposes the gap between “accepted” and “operationally proven.”

Accepted is paperwork. Proven is performance under pressure.

Commissioning matters. Acceptance matters. Documentation matters.

But none of them should be confused with operational proof.

Operational proof comes from sustained behavior, trend evidence, alarm review, operator feedback, maintenance history, and testing the conditions that actually threaten uptime. It comes from seeing whether the cooling system can hold stable operation when the easy assumptions are removed.

That is where high-risk facilities protect themselves: before the plugged strainer becomes a thermal event, before the standby chiller fails to stabilize, before the MRI schedule gets cancelled, before the ATE floor loses temperature control, before the data center finds out that a clean commissioning report did not describe August.

The commissioning report describes one day. Operations describes every day after that.

Part 3 should go one level deeper: the warning signs owners miss before cooling failures become downtime events — the small drifts, repeated nuisance alarms, unstable valves, chemical changes, and operator workarounds that tell you the system is already asking for attention.

Passing a commissioning test is not the same thing as proving operational resilience.