Replacing a failed part can put a machine back into service without ever answering why the part failed — and if the underlying condition survives the repair, so does the failure.
When a bearing fails, replacing the bearing may be exactly what is required to get the machine running again.
But it answers only one question: how do we restore the machine?
It does not necessarily answer why the bearing failed.
Those are different engineering problems.
The distinction becomes obvious when the same machine returns with the same problem weeks or months later.
- Another bearing is fitted.
- Another seal is replaced.
- Another coupling is changed.
The repair was successful in the short term because the machine returned to service.
But if the condition that damaged the original component remains, the maintenance team may simply have reset the failure clock.
I have encountered this kind of situation in practical engineering work: a machine or component can be repaired, yet a similar problem later returns. That experience makes the difference between restoring function and eliminating the cause of failure particularly important.
Root Cause Analysis, or RCA, is intended to address the second problem.
“What sequence of physical and system conditions allowed this failure to happen, and what must change to stop it happening again?”
Replacing a component is not the same as solving a failure
Suppose a pump bearing is damaged.
The bearing is removed. A new one is installed. The pump runs normally again.
From a maintenance-response perspective, the repair worked.
But consider several possible underlying conditions:
- Shaft misalignment
- Contaminated lubricant
- Incorrect lubrication quantity
- Poor installation
- Excessive loading
- An unsuitable fit
- Damaged sealing
Any of these could remain after the replacement.
If so, the new bearing begins operating inside the same damaging environment as the old one.
This is why repeated component replacement is such an important warning sign.
The component may be the victim of the problem rather than its origin.
A useful RCA principle is therefore:
“Do not confuse the location where failure became visible with the reason the failure occurred.”
Failure mode, mechanism and cause are different
Good machinery RCA requires precise language.
ISO 14224:2016 provides standardized reliability and maintenance terminology and failure information for equipment used in petroleum, petrochemical and natural-gas operations. Although its formal industry scope is specific, its failure-data distinctions are useful more broadly when thinking about equipment reliability.
Consider a damaged rolling-element bearing.
Failure event
The machine can no longer continue normal operation because of the bearing condition.
Failure mode
This describes the way the required function is lost or degraded.
Failure mechanism
This is the physical process that produced the damage. Examples could include:
- Fatigue
- Wear
- Overheating
- Lubrication breakdown
- Corrosion
Failure cause
“Why did that mechanism develop?”
Possible causes might include:
- Alignment
- Contamination
- Lubrication practice
- Fit
- Installation
- Excessive loading
- Operating environment
These levels should not be collapsed into one statement.
Saying "the bearing failed" is not a root-cause conclusion. It describes where investigation should begin.
Failure mode → failure mechanism → candidate causes → verified cause → corrective action
A failure mechanism gives the investigation direction
Imagine a bearing surface shows evidence consistent with severe overheating.
That observation immediately changes the questions worth asking. The investigation might examine:
- Lubrication condition
- Lubricant quantity
- Fit
- Load
- Speed
- Cooling
- Installation
- Operating temperature history
If instead the dominant damage pattern indicates another mechanism, the investigation follows a different path.
This is why understanding the physical mechanism is so valuable.
The mechanism helps connect visible damage to plausible conditions that could have created it.
Without this step, RCA can become little more than brainstorming.
Restore production — but preserve the evidence
Maintenance teams often operate under pressure.
The machine is down. Production is waiting.
The obvious priority is to remove the failed component and get the asset running.
But repair can alter the evidence.
A damaged part may be cleaned, dismantled, discarded or repositioned.
Lubricant may be drained. Alignment may change during disassembly. Controller alarms may be reset. The exact operating state may be forgotten.
By the time someone asks for an RCA two days later, some of the strongest evidence may have disappeared.
Where safety, production requirements and practicality allow, evidence should therefore be captured before it is destroyed.
Useful information may include:
- Photographs
- Component orientation
- Wear pattern
- Fracture or damaged surfaces
- Lubricant appearance
- Deposits
- Looseness
- Alignment measurements
- Operating state
- Alarm history
- Vibration trends
- Temperature trends
- Recent maintenance activity
The principle is not: delay a critical repair indefinitely.
It is: restore production while preserving enough evidence to understand what happened.
Start with a precise problem statement
Compare these two ways of describing the same event:
| Vague | Precise |
|---|---|
| Pump keeps failing. | Drive-end bearing on the pump has experienced a second high-vibration event followed by overheating under normal production operation. |
The precise version is more useful because it begins defining:
- Which asset
- Which component
- What happened
- Whether it has recurred
- What symptoms were observed
- Under what broad operating condition
A problem statement should describe the event. It should not prematurely insert the cause.
For example: "Bearing failed because the operator did not lubricate it."
That is already a conclusion.
If the investigation has not demonstrated inadequate lubrication or why it occurred, the statement introduces bias before the evidence is analysed.
Reconstruct what happened before the failure
A machine does not usually begin failing at the moment it stops.
The observable failure may be the final event in a much longer sequence.
A timeline can therefore be extremely useful. For example:
- Bearing replaced
- Machine returned to operation
- Production loading changed
- Vibration gradually increased
- Temperature began rising
- Abnormal noise reported
- Machine stopped
Now investigators can ask:
- Which event happened first?
- Which variables changed before the symptoms appeared?
- Was anything modified?
- Did the same sequence occur during the previous failure?
Useful timeline information can come from maintenance records, operator reports, PLC/controller events, condition-monitoring data, work orders, lubrication records and production records.
The timeline converts a failure from a single moment into a process.
Physical evidence should come before confident explanations
A meeting room can generate dozens of possible causes.
That is useful for expanding the investigation. It is not evidence.
A stronger RCA begins with what is known. For machine failures, useful evidence might include:
| Evidence type | Examples |
|---|---|
| Physical | Failed component, wear, fracture, discoloration, deposits, deformation |
| Condition-monitoring | Vibration, temperature, oil condition, motor current/load |
| Operating | Speed, load, duty cycle, process changes |
| Maintenance | Previous repairs, component replacement, lubrication, alignment, inspection |
| Control-system | Alarms, machine state, trips, operating sequence |
The investigator then asks:
“Which causal explanation is consistent with all of this evidence?”
Condition monitoring and RCA answer different questions
Condition monitoring may tell us that vibration increased before the machine stopped.
That is valuable. But it does not necessarily tell us why.
Possible causes could include imbalance, alignment, looseness, bearing damage or process loading.
Condition monitoring helps identify what changed and when. RCA attempts to explain what created the change.
This is why good condition-monitoring history can be so valuable after a failure.
A single damaged component shows the final state. A trend may show how the machine arrived there.
The Five Whys can be useful
The Five Whys is one of the best-known RCA techniques.
ASQ describes it as a method of repeatedly asking why in order to move beyond symptoms and explore deeper causes. Importantly, ASQ also notes that it may take more or fewer than five questions; the number five is not a rigid stopping rule.
Consider an illustrative example.
| Question | Answer |
|---|---|
| Problem | Bearing repeatedly overheats. |
| Why? | Friction increased. |
| Why? | The lubrication condition was inadequate. |
| Why? | The bearing was not receiving the required lubrication at the required interval. |
| Why? | The existing lubrication schedule did not reflect current operating duty. |
| Why? | The maintenance task had not been reviewed after the machine's operating conditions changed. |
Now the possible issue has moved from replace the bearing to review how lubrication requirements are controlled when duty changes.
That can produce a more durable improvement.
But there is a major caution.
Five Whys does not prove the answer
Every statement in the previous chain could be wrong.
Perhaps lubrication was adequate.
Perhaps the bearing overheated because fit was incorrect, shaft loading changed, alignment was poor, or another fault increased load.
The Five Whys is therefore best treated as a way of constructing candidate causal reasoning. It is not an evidence generator.
ASQ's 8D problem-solving guidance makes this point especially clearly: causes should be verified or proved, rather than simply determined by brainstorming.
So after developing the chain, investigators should ask:
“What evidence demonstrates each important step?”
Complex failures may have more than one cause
Machine failure is not always a clean chain where A caused B and B caused C.
Imagine a bearing operating under slight misalignment, marginal lubrication and increased process load.
None of those conditions alone may have produced rapid failure. Together, they may.
Modern RCA therefore needs to tolerate multiple causal paths.
Even ASQ's broader problem-solving guidance notes that complex problems may have more than one root cause and emphasizes asking what other potential causes were studied and eliminated.
This is where a single Five Whys chain can become too narrow.
Fishbone diagrams help broaden the search
A cause-and-effect or fishbone diagram can be useful when several families of cause are possible.
ASQ describes it as a method for identifying and organizing many possible causes into useful categories.
For machinery investigations, possible categories might include machine, method, material, measurement, environment and human factors.
For a repeated bearing failure, candidate causes might include:
| Category | Candidate causes |
|---|---|
| Machine | Alignment, shaft condition, housing, fit |
| Method | Installation, lubrication procedure, alignment procedure |
| Material/component | Bearing specification, lubricant specification |
| Measurement | Inadequate monitoring, incorrect measurement location |
| Environment | Contamination, temperature |
| Human/system factors | Unclear responsibility, incomplete instructions, missed maintenance task |
The diagram helps ensure that one attractive explanation does not dominate too early.
But it still does not prove anything.
A fishbone is a hypothesis map
This distinction is important.
A fishbone diagram containing misalignment does not mean the machine was misaligned. It means misalignment is one candidate cause worth evaluating.
“A plausible explanation is a hypothesis; a candidate root cause becomes credible only once it has been checked against evidence.”
The next step is evidence.
- Was alignment measured?
- What was found?
- Was it sufficiently abnormal to explain the damage?
- Does the failure pattern support it?
- Were competing explanations eliminated?
The RCA process should move from possible, to probable, to supported.
"Human error" is often an incomplete root cause
Suppose an investigation concludes: the operator failed to lubricate the bearing.
Perhaps that action happened. But the investigation should usually continue. Ask:
- Was lubrication responsibility clear?
- Was the task scheduled?
- Was the interval appropriate?
- Was lubricant available?
- Was the point accessible?
- Was the requirement documented?
- Had operating conditions changed?
- Was completion verified?
This is not about removing individual responsibility. It is about understanding the entire causal system.
ASQ's problem-solving guidance specifically encourages moving away from simply determining who was at fault and toward understanding how the process can be improved.
“Human action can be part of the failure chain without necessarily being the deepest useful cause.”
Mechanical causes must not disappear either
There is an opposite RCA mistake.
A team may quickly jump from "bearing failed" to "maintenance-management system problem."
But the physical failure still has to make engineering sense.
A robust analysis can connect several levels. For example:
| Level | Explanation |
|---|---|
| Physical mechanism | Bearing damage associated with inadequate lubrication |
| Immediate condition | Lubricant supply to the bearing was insufficient |
| Maintenance-system issue | Lubrication interval or execution did not match actual operating requirements |
The physical and organizational explanations are not competitors. They are different levels of the same causal chain.
Correlation is not automatically cause
Suppose vibration increased shortly before failure. That shows correlation in time.
Did vibration cause the failure? Not necessarily.
The vibration may itself have been another symptom of a developing fault.
Similarly, a temperature increase does not mean high temperature initiated the failure. Perhaps increased friction caused both the temperature rise and the surface damage.
RCA should therefore distinguish a signal associated with the failure from a condition that materially caused the failure.
This is especially important as maintenance becomes more data-driven.
More data produce more correlations. They do not automatically produce causal explanations.
Root cause should be verified against evidence
Suppose the proposed cause is: repeated coupling failure was caused by shaft misalignment.
What evidence should exist if that explanation is correct? Potential evidence could include:
- Measured alignment outside acceptable limits
- Wear consistent with the condition
- Vibration behaviour consistent with misalignment
- Recurrence associated with installation or alignment changes
Then ask:
“What evidence would contradict this explanation?”
That second question is particularly useful.
Good investigation should try to disprove a favourite hypothesis, not only collect evidence that supports it.
Corrective action should logically follow from the cause
Imagine the verified problem is alignment.
Then replacing the coupling is repair. Correcting alignment is closer to corrective action.
But even that may be incomplete.
Why was alignment lost? Perhaps foundation movement, poor installation procedure, thermal growth, or inadequate verification after maintenance.
A stronger corrective package may therefore include:
- Restore correct alignment
- Address the condition creating alignment drift
- Define measurement/verification after future work
The best corrective action breaks the causal chain at a practical and controllable point.
Verify the corrective action
RCA should not finish when the report is closed.
The final question is:
“Did the change actually prevent the problem we were trying to prevent?”
Useful verification might include:
- Condition returns toward baseline
- Operating temperature stabilizes
- Alignment remains acceptable
- Contamination reduces
- Failure does not recur during an appropriate observation period
ASQ's structured 8D process places verification of root causes and corrective actions inside the formal problem-solving sequence rather than treating corrective action as an assumption.
NASA failure-analysis guidance expresses a similar engineering principle: root-cause investigation should proceed to a level sufficient to support effective corrective action and recurrence control, while recognizing that additional analysis is not always justified when further insight would not be cost-effective.
Not every failure needs the same investigation depth
A complete RCA costs time, engineering effort, production attention and specialist resources.
A low-consequence consumable reaching its expected wear limit may not justify the same investigation as:
- Repeated critical bearing failures
- Major production loss
- Safety-related failures
- Expensive component damage
- Unexplained catastrophic failure
Investigation depth should therefore reflect consequence, recurrence, uncertainty and learning value.
NASA's failure-analysis guidance explicitly recognizes that deeper root-cause investigation may be stopped when further analysis is not cost-effective relative to the risk reduction or insight gained, provided the rationale is documented.
“Use enough RCA to make the right engineering decision — not the maximum possible investigation for every event.”
An illustrative repeated bearing failure
Consider a fictional production machine.
A bearing fails. It is replaced. Two months later, the replacement bearing begins showing abnormal vibration.
The immediate response could be: replace it again.
Instead, the maintenance team preserves the bearing and operating information. The timeline shows:
- Bearing replacement
- Machine returned to service
- Gradual vibration increase
- Temperature rise
- Eventual damage
Inspection suggests a lubrication-related failure mechanism.
Now possible causes are investigated:
- Wrong lubricant?
- Contamination?
- Insufficient quantity?
- Incorrect interval?
- Blocked delivery path?
- Excessive operating load?
Oil and maintenance records show no lubricant-specification change.
Inspection finds no strong evidence of external contamination.
The maintenance history, however, shows that the machine's production duty increased significantly while the lubrication schedule remained unchanged.
That still does not automatically prove causation.
The team evaluates whether the increased duty could materially change lubrication requirements and whether the observed damage is consistent with inadequate lubrication under those conditions.
Suppose the evidence supports the hypothesis.
The corrective action is no longer replace the bearing. It becomes:
- Replace damaged component
- Restore correct lubrication condition
- Revise lubrication task based on actual duty
- Monitor temperature/vibration following return to service
The example demonstrates the difference between repair and recurrence prevention.
The example is illustrative, not a claim about a specific real failure.
My perspective on recurring failure
In practical engineering work, encountering a machine or component that is repaired and then develops a similar problem again changes the way maintenance is viewed.
The first repair naturally focuses attention on the broken component.
Recurrence forces a different question:
“What did we leave unchanged?”
That question is central to Root Cause Analysis.
A repeated failure does not necessarily mean the replacement was done badly.
It may mean the repair restored the component while leaving the damaging operating condition intact.
This is why RCA is most valuable when it moves attention from "what part should we replace" to "what allowed this part to fail."
The Machine Failure RCA Chain
A practical machinery RCA can be organized into ten stages.
1. Stabilize
First:
- Make the equipment safe
- Control immediate consequences
- Isolate energy as required
- Protect people and surrounding equipment
Investigation does not take priority over safety.
2. Preserve
Before evidence disappears, capture photographs, component condition, operating state, trends, alarms and relevant samples.
3. Define
State precisely what failed, where, when, how it was detected and its operational impact. Avoid premature causal language.
4. Reconstruct
Develop the timeline. Include operation, maintenance, modifications, and previous symptoms or failures.
5. Explain the failure mechanism
Ask how the component physically reached this condition. This connects the visible damage to engineering physics.
6. Generate candidate causes
Use physical reasoning, Five Whys, fishbone diagrams, expert knowledge and maintenance history. Do not select a root cause yet.
7. Verify
For each strong candidate: identify supporting evidence, identify contradictory evidence, test where practical, and compare with competing explanations.
8. Correct
Choose actions that interrupt the verified causal path. The action may involve the machine, its design, installation, maintenance, operation or monitoring.
9. Confirm
After implementation: monitor expected indicators, verify that the condition improved, and check for recurrence.
10. Learn
Use the investigation beyond one machine where appropriate. Update maintenance tasks, inspection criteria, design standards, training, condition monitoring and failure databases.
That is how one failure can produce wider reliability improvement.
Common RCA mistakes
- Replacing the part and calling the issue closed. Repair and RCA have different purposes.
- Starting the investigation after evidence has disappeared. Capture critical information early.
- Calling the failed component the root cause. The failed component may be the consequence.
- Stopping after exactly five Whys. ASQ explicitly notes that fewer or more questions may be required.
- Treating brainstorming as proof. Candidate causes must be evaluated.
- Selecting "operator error" and stopping there. Look at the conditions that shaped the action as well.
- Ignoring physical failure mechanisms. Organizational explanations still need to connect to the machine physics.
- Searching only for one cause. Complex failures may involve multiple interacting conditions.
- Implementing an action without verifying the cause. The machine may improve temporarily for an unrelated reason.
- Closing RCA before checking recurrence. Corrective action needs follow-up.
How RCA improves maintenance strategy
The value of a good RCA can extend beyond one repair.
Suppose an investigation finds that a critical failure was preceded by a detectable vibration pattern. The organization might improve its condition monitoring.
If the investigation identifies lubrication intervals that do not match operating duty, it may improve preventive maintenance.
If degradation can be reliably detected and trended, the asset may become a candidate for predictive maintenance.
If recurring failure substantially reduces production, the improvement may appear later through better Availability and OEE.
This shows how the maintenance topics connect.
RCA is not separate from maintenance strategy. It supplies knowledge that can make the strategy better.
Key takeaway
When a machine fails, replacing the broken component may be necessary. Sometimes it is all that is economically justified.
But when failure is serious, recurrence is expensive, the same component keeps failing, or the cause is unclear, repair alone is not enough.
A strong investigation asks:
- What failed?
- How did it fail?
- What conditions produced that mechanism?
- What evidence supports that explanation?
- What change will prevent recurrence?
Finally: did the change work?
Tools such as Five Whys and fishbone diagrams can help organize the investigation. They should never replace evidence.
“A root cause is not the most convincing story about a failure. It is a causal explanation strong enough to support effective corrective action and survive comparison with the evidence.”
That is the difference between repeatedly repairing a machine and learning from why it failed.
References and further reading
- ISO 14224:2016 — Petroleum, petrochemical and natural gas industries — Collection and exchange of reliability and maintenance data for equipment. Its formal scope is petroleum, petrochemical and natural-gas industries, not general machinery RCA — but its standardized reliability and failure-mode terminology and data-quality principles are useful more broadly when structuring a machinery investigation.
- ASQ — Five Whys and Five Hows. Primary guidance on using repeated "why" questions to move beyond symptoms; ASQ notes that the method may require more or fewer than five iterations.
- ASQ — Eight Disciplines (8D) Problem Solving. Requires identifying and verifying root causes rather than selecting them through unsupported brainstorming, and places corrective-action verification inside the formal problem-solving sequence.
- ASQ — Fishbone / Cause-and-Effect Diagram. Guidance on using a fishbone diagram to organize multiple potential causes into categories for further investigation.
- NASA-HDBK-8739.18 — NASA Root Cause Analysis guidance within problem/failure reporting. Engineering guidance on taking root-cause investigation far enough to support effective corrective action and recurrence control, while recognizing practical limits to investigation depth.




