Mission-Critical Commissioning: Why “Passed” Does Not Mean Proven

Industrial chiller room with overhead pipe arrays and foreground pressure gauge near red zone

Part 1 of 4 — The Failures That Begin After Sign-Off

Why a system can pass startup and functional testing, then fail under the operating conditions that actually matter.

A mission-critical compressed dry air compressor failed every four to six weeks for nearly two years.

Not because the operations team ignored it.

Not because maintenance had been deferred.

Not because the machine was defective.

It failed because it had been specified for roughly a 70% duty cycle—then installed in a facility that needed it to operate near continuously.

One module at a time, the team replaced internal elements at roughly $10,000 each. Every failure also brought six to eight hours of production disruption, emergency coordination, vendor calls, repair decisions, and a growing concern that no one could fully explain.

The failures kept coming.

By the time the pattern was understood, nearly every major internal element in the machine had been replaced.

The root cause had been present from day one.

The compressor was being asked to operate outside the envelope it had been designed to support. It had passed commissioning anyway.

And by the time the owner understood the full consequence of that original design assumption, the system was only about two years old.

Later in this article, I will share what that one missed assumption ultimately cost the Fortune 500 owner.

Mission-critical commissioning is not proven when equipment starts, alarms clear, and a closeout report is signed. It is proven when the system performs reliably under the actual conditions the owner will depend on for years.

The Difference Between Startup and Operational Proof

Most teams understand that equipment must be started, tested, balanced, and documented before turnover.

That is necessary.

It is not always sufficient.

A startup test may prove that a compressor starts. A functional performance test may prove that a control sequence responds correctly under a controlled scenario. A commissioning report may confirm that a chiller meets the specified leaving-water temperature at the load and ambient conditions present on the day it was tested.

But mission-critical equipment does not live on a commissioning day.

It lives through August heat.

It lives through winter free-cooling transitions.

It lives through production ramps, changing process loads, equipment failures, operator turnover, maintenance windows, and the slow performance drift that rarely appears in a project closeout meeting.

Mission-critical commissioning must test more than startup performance. It must establish whether critical equipment, controls, and redundancy can survive the operating conditions the owner will actually face.

A system can pass every startup and functional test in the book while still carrying a hidden operational weakness.

That weakness may be a duty-cycle mismatch.

It may be a pump selected too close to the edge of its curve.

It may be a chiller plant with enough nameplate capacity but insufficient capacity at peak ambient.

It may be a control sequence that works when one piece of equipment is operating, but becomes unstable when the lead unit fails and the lag unit takes over.

It may be redundancy that looks complete on a one-line drawing but cannot actually support the load under real operating conditions.

The issue is not that commissioning has no value. The issue is that conventional commissioning is often concentrated near the end of construction, when the facility has experienced the least demanding conditions of its life.

How Mission-Critical Facilities Actually Get Built

An owner hires an architect.

The architect assembles the design team: mechanical engineers, electrical engineers, controls integrators, specialty consultants, and sometimes a commissioning authority.

The project is engineered, bid, issued for construction, built by a general contractor and subcontractors, tested, documented, and turned over.

Everyone involved may be capable.

Everyone may be acting in good faith.

But the project structure often has one dominant finish line:

Get the project completed.

Get it accepted.

Get it signed off.

Get paid.

That structure creates a predictable gap.

The people who selected the equipment, wrote the sequences, coordinated the drawings, and built the system are often not the people who will live with its operating behavior for the next five or ten years.

The operations team inherits the facility after the major project milestones are complete.

They inherit the alarms.

They inherit the maintenance burden.

They inherit the seasonal failures.

They inherit the production risk.

And when something begins failing six months after turnover, the original project team may be difficult to reassemble, contractually disconnected, or focused on the next project.

This is not an accusation against design firms, contractors, or commissioning providers.

It is a recognition that the standard project model is optimized for delivery—not necessarily for proving long-term operational reliability.

Those are not the same thing.

Why a System Can Look Fine at Turnover

Commissioning usually happens under conditions that are not representative of the facility’s hardest operating day.

The building may be partially occupied.

Process equipment may not be fully installed.

A data hall may not be at design load.

The outdoor temperature may be mild.

The water loop may be clean and new.

The controls system may still be under close attention from the integrator.

The operations team may not yet have assumed responsibility for daily adjustments, alarm response, and emergency recovery.

In other words, the facility is often tested during its most forgiving period.

That is where mission-critical commissioning must go beyond startup testing and ask whether the system can remain stable under real load, real weather, and real failure conditions.

That is why the question should never be limited to:

Did the system pass the test?

The better question is:

What real operating condition has this system not yet experienced—and what would happen if it occurred tomorrow?

For a cooling system, that may mean asking:

  • What happens when peak outdoor ambient coincides with maximum process load?
  • What happens when the lead chiller fails while the lag chiller is already operating near capacity?
  • What happens when a pump must move from low demand to high demand quickly?
  • What happens when a VFD, sensor, actuator, or control valve provides bad feedback?
  • What happens when a system transitions from mechanical cooling to economizer or free-cooling mode?
  • What happens when filters load, heat exchangers foul, or water chemistry begins to drift?
  • What happens when the facility operates at the load profile it will actually see—not the load profile assumed during design?

These are operational questions.

They are also reliability questions.

And they should be part of the commissioning strategy before construction begins.

The Duty-Cycle Problem Nobody Challenged

The CDA compressor in this story was not a mystery.

The equipment had a defined duty-cycle capability.

The facility had a real operating requirement.

Those two facts were not adequately aligned.

A compressor designed for intermittent or moderate-duty use can operate successfully for a period of time in a near-continuous application. That does not mean it is suitable for the application.

It means the failure mechanism has not fully revealed itself yet.

When equipment operates beyond its intended duty cycle, the damage is often gradual.

Heat builds.

Wear accelerates.

Lubrication performance changes.

Internal elements see more cycles or run hours than anticipated.

Maintenance intervals become meaningless because the machine is no longer operating within the conditions used to establish them.

At first, the team may see an isolated failure.

Then another.

Then a third.

Each repair can appear unrelated because the failed component is different.

That is where many organizations get trapped.

They treat each event as a service issue instead of stepping back and asking whether the entire application is wrong for the equipment selected.

The expensive question is not:

Which component failed this time?

It is:

What is this machine being asked to do that it was never designed to do?

Why Mission-Critical Commissioning Fails After Sign-Off

When a mission-critical system fails after turnover, the response often follows a familiar pattern.

The owner calls the service contractor.

The contractor identifies the failed component.

The OEM provides a repair recommendation.

The team gets the system back online.

The incident closes.

Then it happens again.

What is often missing is a formal requirement to examine the system-level cause.

Was the equipment correctly selected for actual demand?

Did the design assumptions match the final operating profile?

Was the control sequence validated under all meaningful modes?

Did the redundancy strategy function under real load?

Did the commissioning plan include seasonal validation?

Was there enough trend data to understand what happened before the failure?

Was anyone contractually required to stay engaged long enough to answer those questions?

Without that accountability, an organization can spend months repairing symptoms while the root cause continues to operate every hour of every day.

That is not a maintenance failure.

It is a project-delivery failure that has been transferred to operations.

What Smart Owners Require Before Design Begins

Mission-critical commissioning should be treated as an operational reliability process—not simply a construction closeout requirement.

The strongest owners do not wait until closeout to think about commissioning. They treat long-term operational validation as part of the original project scope.

That starts before the first design package is issued.

An independent pre-construction design review can identify equipment, controls, and capacity mismatches before they are purchased, installed, and placed into service.

KMC² provides reliability reviews, commissioning frameworks, and design validation support for mission-critical cooling and utility systems.

Mission-critical commissioning validation from startup through seasonal operating conditions

1. Design Against Actual Operating Profiles

Do not validate equipment solely against a design-day calculation.

Validate it against how the facility will actually operate.

For a CDA system, that means understanding minimum demand, normal demand, maximum demand, run hours, load variation, standby requirements, and the consequence of losing the lead machine.

For a chilled-water or process-cooling system, it means understanding actual load diversity, expected expansion, tool heat rejection, water temperatures, ambient conditions, and whether critical equipment will operate near its limits for sustained periods.

The design basis should reflect operating reality—not a convenient average.

2. Select Equipment for Duty Cycle, Margin, and Degradation

Nameplate capacity alone is not enough.

A compressor can have enough capacity and still be inappropriate for the required runtime.

A pump can meet flow and head requirements and still be poorly positioned on its curve.

A chiller can meet the load at moderate ambient and fail to carry it at peak condenser conditions.

A heat exchanger can be adequate when clean and become marginal after normal fouling.

Critical equipment should be selected with margin for operational reality, not simply to comply with the minimum design calculation.

3. Build BAS Trending Into the Original Scope

Owners should not have to retrofit the data they need after a problem develops.

Trending requirements should be included in the mechanical and controls specifications before construction.

That means identifying the critical points needed to understand system behavior:

  • Supply and return temperatures
  • Differential pressure
  • Flow, where reliable measurement is available
  • Pump speed and VFD status
  • Compressor run hours and load state
  • Chiller loading, condenser conditions, and alarm history
  • Valve positions
  • Filter differential pressure
  • Equipment lead-lag status
  • Mode changes and sequence transitions

The goal is not to trend everything.

The goal is to trend the points that allow an operator or reliability reviewer to reconstruct what happened before, during, and after a failure.

Without data, every post-event review becomes a debate based on memory.

4. Require Seasonal and Failure-Mode Validation

Some mission-critical tests cannot be completed during initial turnover.

Seasonal testing should be written into the project requirements.

ASHRAE’s commissioning guidance recognizes that system verification can extend beyond initial turnover to include operational testing, delayed testing, and seasonal conditions necessary to confirm that the system meets the owner’s requirements.

Depending on the facility, that may include:

  • Peak summer cooling validation
  • Winter or economizer-mode testing
  • Lead-lag rotation and failure recovery
  • Loss-of-sensor and bad-signal response
  • Pump and valve control stability at low, medium, and high demand
  • Failure of a lead chiller, compressor, pump, or control component
  • Recovery after utility interruption or emergency shutdown
  • Review of alarm thresholds and escalation paths

Not every facility needs every scenario.

But every mission-critical facility should identify the scenarios most likely to expose hidden weaknesses.

5. Keep Accountability Alive Beyond Startup

The objective is not to create unreasonable liability for the project team.

It is to create shared accountability for proving that the system works beyond a one-day acceptance test.

Reasonable holdbacks tied to defined seasonal or operational milestones change the conversation.

They keep the relevant parties engaged.

They give the owner leverage to resolve open items.

They make it harder for unresolved controls issues, missing trend points, incomplete sequences, and questionable equipment selections to disappear into project closeout.

The exact structure should be developed with legal and procurement support.

But the principle is simple:

A mission-critical system should not be considered fully proven merely because it starts.

6. Involve Operations Before Turnover

The people responsible for operating the system should not receive it as a surprise.

They should be involved while decisions can still be changed at reasonable cost.

That includes reviewing sequences of operation, alarm philosophy, maintenance access, sensor locations, trend points, service requirements, vendor capabilities, and failure-recovery procedures.

Operations teams see risks that are invisible on drawings.

They know which alarms will be ignored because there are too many.

They know which valves cannot be accessed.

They know which OEM response times are unrealistic.

They know where a control sequence will create confusion at 2:00 AM.

Their input does not replace engineering.

It improves the chance that the design will survive actual operation.

The Cost of One Missed Assumption

In this case, the facility spent two years replacing compressor elements, responding to repeated production disruptions, and trying to repair a machine that was being pushed beyond its intended operating envelope.

Eventually, the owner had no responsible option left except to replace the system.

The original system was only about two years old.

The final cost to remove, replace, and correct the installation was just under $1 million.

That number was not caused by one bad service call.

It was not caused by one missed maintenance interval.

It was the accumulated cost of a single assumption that was never fully challenged:

That equipment designed for a 70% duty cycle would be acceptable in a near-continuous-duty application.

It was a design assumption.

A specification assumption.

A commissioning assumption.

And ultimately, an owner-risk assumption.

The most expensive failures in mission-critical facilities often begin long before the first alarm.

They begin when a project is treated as complete before the system has proven itself under the conditions that actually matter.

The Owner’s Real Question

Before approving a design, owners should ask:

What operating condition is most likely to expose the weakness this project has not yet tested?

For some facilities, the answer is peak summer ambient.

For others, it is winter free-cooling transition, full production loading, loss of lead equipment, variable process demand, utility interruption, control-loop instability, or gradual thermal degradation.

The answer will vary by facility.

The question should not.

That is the purpose of mission-critical commissioning: to identify the conditions that could expose a hidden weakness before the owner inherits the operating risk.

Because once the project team has been paid and moved on, the owner is usually the one left carrying the cost of a system that looked good on paper.

This Is Part 1 of 4

This mission-critical commissioning series gives owners, operators, and design teams a practical framework for proving reliability beyond startup—from the original design contract through the first full year of operation.

The purpose of this mission-critical commissioning series is to help owners identify reliability risks before those risks become expensive operating problems.

Part 2 will cover the contractual language and accountability structure that keeps key parties engaged past startup.

Part 3 will address BAS trending data: who controls it, what should be trended, and what happens when the data needed to understand a failure is unavailable.

Part 4 will make the financial case for treating seasonal and operational validation as a core project requirement—not an optional add-on after turnover.

KMC² will publish a free Owner/Designer Playbook alongside the final article in this series: a practical framework owners can use before their next mission-critical project goes out to design.

The goal is not to make commissioning more complicated.

The goal is to make sure the system you accept is the system you actually need.