By Abdullah Showaib
Facilities and maintenance strategy professional based in Dubai
A maintenance team can repair the same fault three times and still never solve it.
The first visit restores service. The second replaces a component. The third closes another work order. On paper, each job may look successful. Yet the building continues to experience the same cooling complaint, pump trip, drainage overflow or electrical interruption.
That is the point at which ordinary troubleshooting should become root cause analysis.
Root cause analysis in maintenance is a structured way to identify why a failure occurred, not merely what failed. The objective is to distinguish the visible symptom from the direct cause, contributing conditions and the underlying issue that must change if recurrence is to stop.
This matters in buildings because failures rarely occur in a vacuum. Equipment condition, operating environment, controls, maintenance quality, access, loading, installation, previous repairs and human decisions can all contribute to the same event.
In Dubai’s commercial, hospitality and residential buildings, the need for evidence-led RCA is amplified by intensive cooling demand and tightly interconnected HVAC, electrical, plumbing and control systems. A repeat fault can therefore travel quickly across system boundaries, increasing disruption, call-outs and maintenance cost if the underlying condition is not removed.
The strongest maintenance organisations therefore develop a low tolerance for repeat failure. A repeated fault is not simply another ticket. It is evidence that either the diagnosis, corrective action or maintenance strategy may be incomplete.
Repair Is Not the Same as Root Cause Removal
A repair answers the question: what must we do to restore the asset or service now?
Root cause analysis answers a different question: what condition allowed this failure to happen, and what must change to prevent it from happening again?
Those two questions can produce very different actions.
If an FCU stops because a condensate drain is blocked, clearing the drain restores service. But if the same blockage returns every month, the investigation may need to examine drain routing, access, cleaning frequency, insulation condition, negative pressure, contamination or whether previous maintenance ever reached the actual restriction.
The repair may be correct and still be incomplete as a reliability solution.
When Should a Maintenance Team Trigger an RCA?
Not every minor defect needs a formal investigation. Root cause analysis is most valuable when the consequence or recurrence justifies deeper work.
- the same asset or system fails repeatedly
- a fault causes significant downtime or occupant disruption
- the same symptom returns after apparently successful repair
- a failure crosses more than one building system
- the cost of repeated corrective work is increasing
- a safety, compliance or business-continuity risk is involved
- the failure mechanism remains uncertain after normal troubleshooting
- multiple similar assets begin showing the same pattern
The Repeat Failure Evidence Chain
A useful RCA for building maintenance should move through six evidence stages. The framework below is designed to stop teams from jumping from symptom to assumption.
1. Symptom: What was observed or reported? Record the failure without embedding a diagnosis into the description.
2. Event evidence: What readings, alarms, photographs, timestamps, trips, leaks, temperatures or physical conditions were present at the time?
3. Direct cause: What immediately caused the service to stop or degrade?
4. Contributing conditions: What made the direct cause more likely, such as loading, contamination, access, control settings, environment, maintenance gaps or installation conditions?
5. Root cause: What underlying physical, procedural, design or management condition must change to prevent recurrence?
6. Effectiveness check: What evidence will prove that the corrective action worked over time?
The sixth stage is frequently missed. An RCA is not complete when a corrective action is assigned. It is complete when the team can demonstrate that the action changed the failure pattern.
Start With a Neutral Problem Statement
Weak root cause analysis often begins with a conclusion disguised as a problem statement.
“Pump failed because maintenance was poor” already assumes both the cause and responsibility.
A stronger statement is specific and neutral:
“Booster pump No. 2 tripped on overload three times in 21 days during peak demand, causing temporary low water pressure to occupied floors.”
That statement gives the team something testable. It identifies the asset, event, frequency, operating condition and consequence without deciding the cause in advance.
Build the Failure Timeline Before Asking Why
The popular “5 Whys” method can be useful, but asking why too early can produce a chain of guesses.
Before repeatedly asking why, reconstruct what changed between normal operation and failure.
Review work-order history, alarms, readings, previous replacements, maintenance dates, operator observations, BMS trends, environmental conditions and any recent changes to the system.
A timeline is particularly valuable for intermittent faults. If a breaker trips only during a specific load period, or a drainage problem appears only after prolonged cooling operation, timing may reveal dependencies that a static inspection misses.
Separate Direct Cause, Contributing Cause and Root Cause
| Level | Example | What It Means |
| Direct cause | Motor protection tripped | The immediate event that stopped operation |
| Contributing condition | Bearing temperature was high and lubrication condition was poor | A condition that increased failure likelihood |
| Root cause | Lubrication task was absent, unsuitable or not being verified in the maintenance strategy | The underlying condition that must change to prevent recurrence |
This distinction prevents a common RCA failure: stopping at the first technically correct explanation. “The motor tripped” may be true, but it is not enough if the goal is prevention.
Cross-System Failures Need Wider Evidence
Dubai buildings are especially vulnerable to incomplete RCA where HVAC, mechanical, electrical, plumbing and controls interact. Effective MEP maintenance services should therefore investigate the dependency chain around the affected asset, not only the trade named on the work order.
Recurring cooling complaint: Possible evidence may include airflow, coil condition, chilled-water flow, valve position, sensor accuracy, control logic, electrical supply and condensate condition.
Repeated pump trip: The pump may need review alongside pressure conditions, valve position, level controls, sensors, contactors, protection settings and duty/standby sequencing.
Water staining above a ceiling: The visible mark may require tracing plumbing, condensate, waterproofing, drainage routes and any electrical or finish damage created by moisture.
The aim is not to make every fault unnecessarily complex. It is to expand the investigation only where the evidence shows that the failure crosses system boundaries.
Use Fault History as Evidence, Not Background Noise
A single work order describes an event. A sequence of work orders describes a pattern.
Repeat-failure analysis should therefore compare asset ID, symptom, fault code, corrective action, parts replaced, technician notes and the interval between failures.
If three different technicians use different wording for the same recurring fault, the pattern can disappear inside the maintenance system. Consistent asset identification and fault coding make RCA substantially stronger.
This is one reason work-order closure rate alone can mislead management. High closure volume can coexist with poor reliability if the same assets keep returning.
Corrective Action Must Change the Failure Mechanism
Corrective actions generally fall into four categories:
Physical: repair, redesign, replacement, improved access, different component or changed installation
Maintenance: new inspection task, changed frequency, condition monitoring, measurement or verification
Operational: revised setpoint, loading, sequencing, operating procedure or user practice
Management: approval route, vendor accountability, training, spare strategy or escalation rule
The action should match the identified mechanism. Increasing inspection frequency will not solve a design restriction. Replacing a component will not solve an incorrect control sequence. Training will not solve an inaccessible valve.
The RCA Closure Test
Before closing a repeat-failure RCA, ask five questions:
- Was the root cause supported by evidence rather than assumption?
- Did the corrective action address that cause directly?
- Was responsibility and completion evidence recorded?
- Was the preventive maintenance plan changed where necessary?
- Has the asset operated long enough without recurrence to demonstrate effectiveness?
If the fifth answer is unknown, the investigation may be administratively closed but technically unverified.
Example: Repeated FCU Water Leakage
Consider an FCU associated with repeated ceiling water leakage. A weak maintenance record might show three visits, three drain cleanings and three closed jobs.
An RCA would instead separate the evidence:
Symptom: Water staining and dripping from the ceiling during extended cooling operation.
Direct cause: Condensate tray overflow.
Contributing conditions: Restricted drain flow and difficult access for full cleaning.
Further evidence: Drain routing, insulation, unit pressure condition, previous cleaning records and whether the same section repeatedly blocks.
Corrective action: Remove the confirmed restriction, correct any physical defect, improve access or cleaning method where needed, and revise the PM task if the previous scope was inadequate.
Effectiveness check: No recurrence across the next defined operating period, supported by inspection evidence.
The important change is that the work order stops being a record of attendance and becomes a record of learning.
Root Cause Analysis Should Improve Preventive Maintenance
RCA has limited value if the finding stays inside an incident report.
The result should feed back into preventive maintenance, asset criticality, spare planning, technician instructions, inspection standards or capital replacement decisions.
IBM describes root cause analysis as a way to identify the underlying cause of asset failure rather than focusing on symptoms, and notes that resolving the root cause can help prevent future disruption. IBM’s maintenance guidance supports the same principle: the finding matters most when it changes what the maintenance team does next.
From Fast Repair to Reliable Repair
Maintenance performance is often measured by speed: response time, attendance time and work-order closure.
Those metrics matter. But they do not answer whether the fault stayed fixed.
A mature maintenance operation should therefore treat repeat failure as a separate signal. When the same asset, symptom or failure mode returns, the organisation should escalate from repair to investigation.
The goal of root cause analysis in maintenance is not to make every breakdown into a lengthy engineering study. It is to know when ordinary repair is no longer enough.
The most useful maintenance question after a repeat fault is not “How quickly did we close it?”
It is:
“What has to change so we do not have to close this same fault again?”
Frequently Asked Questions
What is root cause analysis in maintenance?
Root cause analysis in maintenance is a structured process for identifying the underlying cause of an equipment or building-system failure so corrective action can prevent recurrence.
When should RCA be used in facility maintenance?
Use it for repeat failures, high-impact breakdowns, uncertain failure mechanisms, cross-system faults, safety or compliance events, and recurring problems that normal troubleshooting has not eliminated.
What is the difference between a direct cause and a root cause?
A direct cause is the immediate condition that produced the failure. A root cause is the underlying physical, procedural, design or management condition that must change to prevent the failure from returning.
Is the 5 Whys method enough for maintenance RCA?
It can help with simple problems, but complex building failures often require a timeline, work-order history, readings, system dependencies and evidence testing before a cause is accepted.
How do you know whether corrective action worked?
Define an effectiveness check, such as no recurrence over a suitable operating period, improved readings, reduced fault frequency or verified removal of the failure mechanism.
Why is root cause analysis important for building maintenance in Dubai?
Dubai buildings often operate under intensive cooling demand, high occupancy and interconnected MEP systems. RCA helps maintenance teams identify why recurring faults return rather than repeatedly restoring the same symptom.
Author
Abdullah Showaib is a facilities and maintenance strategy professional based in Dubai.
About SnapFixNow FMC
SnapFixNow FMC is a Dubai-based engineering-led facility and building maintenance company providing AMC, facility management, HVAC, MEP and outsourced engineering support across residential, commercial, hospitality and industrial assets.



























