Skip to main content
Business software? Visit Kipeo Digital ↗
HarunLucas.com
Article
Maintenance and ReliabilityPublished

Root Cause Analysis for Machine Failures: Moving Beyond Replacing the Broken Part

18 min read

Replacing a failed part can restore a machine to service without ever explaining why the part failed. This article works through how to separate failure mode, mechanism and cause, preserve evidence before it disappears, verify candidate causes rather than assume them, and design corrective action that actually prevents recurrence.

Maintenance engineer using a flashlight and magnifying glass to inspect a disassembled bearing on a workbench, with additional bearing components, a caliper, a handwritten inspection log and a tablet showing a vibration waveform, beside a partially disassembled industrial gearbox and motor.

Replacing a failed part can put a machine back into service without ever answering why the part failed — and if the underlying condition survives the repair, so does the failure.

When a bearing fails, replacing the bearing may be exactly what is required to get the machine running again.

But it answers only one question: how do we restore the machine?

It does not necessarily answer why the bearing failed.

Those are different engineering problems.

The distinction becomes obvious when the same machine returns with the same problem weeks or months later.

  • Another bearing is fitted.
  • Another seal is replaced.
  • Another coupling is changed.

The repair was successful in the short term because the machine returned to service.

But if the condition that damaged the original component remains, the maintenance team may simply have reset the failure clock.

I have encountered this kind of situation in practical engineering work: a machine or component can be repaired, yet a similar problem later returns. That experience makes the difference between restoring function and eliminating the cause of failure particularly important.

Root Cause Analysis, or RCA, is intended to address the second problem.

What sequence of physical and system conditions allowed this failure to happen, and what must change to stop it happening again?

Replacing a component is not the same as solving a failure

Suppose a pump bearing is damaged.

The bearing is removed. A new one is installed. The pump runs normally again.

From a maintenance-response perspective, the repair worked.

But consider several possible underlying conditions:

  • Shaft misalignment
  • Contaminated lubricant
  • Incorrect lubrication quantity
  • Poor installation
  • Excessive loading
  • An unsuitable fit
  • Damaged sealing

Any of these could remain after the replacement.

If so, the new bearing begins operating inside the same damaging environment as the old one.

This is why repeated component replacement is such an important warning sign.

The component may be the victim of the problem rather than its origin.

A useful RCA principle is therefore:

Do not confuse the location where failure became visible with the reason the failure occurred.

Failure mode, mechanism and cause are different

Good machinery RCA requires precise language.

ISO 14224:2016 provides standardized reliability and maintenance terminology and failure information for equipment used in petroleum, petrochemical and natural-gas operations. Although its formal industry scope is specific, its failure-data distinctions are useful more broadly when thinking about equipment reliability.

Consider a damaged rolling-element bearing.

Failure event

The machine can no longer continue normal operation because of the bearing condition.

Failure mode

This describes the way the required function is lost or degraded.

Failure mechanism

This is the physical process that produced the damage. Examples could include:

  • Fatigue
  • Wear
  • Overheating
  • Lubrication breakdown
  • Corrosion

Failure cause

Why did that mechanism develop?

Possible causes might include:

  • Alignment
  • Contamination
  • Lubrication practice
  • Fit
  • Installation
  • Excessive loading
  • Operating environment

These levels should not be collapsed into one statement.

Saying "the bearing failed" is not a root-cause conclusion. It describes where investigation should begin.

Failure modefailure mechanismcandidate causesverified causecorrective action

The damaged component is where a failure becomes visible — the mode, mechanism and underlying causes explain why it happened, and only the last step, corrective action, changes the outcome.

A failure mechanism gives the investigation direction

Imagine a bearing surface shows evidence consistent with severe overheating.

That observation immediately changes the questions worth asking. The investigation might examine:

  • Lubrication condition
  • Lubricant quantity
  • Fit
  • Load
  • Speed
  • Cooling
  • Installation
  • Operating temperature history

If instead the dominant damage pattern indicates another mechanism, the investigation follows a different path.

This is why understanding the physical mechanism is so valuable.

The mechanism helps connect visible damage to plausible conditions that could have created it.

Without this step, RCA can become little more than brainstorming.

Restore production — but preserve the evidence

Maintenance teams often operate under pressure.

The machine is down. Production is waiting.

The obvious priority is to remove the failed component and get the asset running.

But repair can alter the evidence.

A damaged part may be cleaned, dismantled, discarded or repositioned.

Lubricant may be drained. Alignment may change during disassembly. Controller alarms may be reset. The exact operating state may be forgotten.

By the time someone asks for an RCA two days later, some of the strongest evidence may have disappeared.

Where safety, production requirements and practicality allow, evidence should therefore be captured before it is destroyed.

Useful information may include:

  • Photographs
  • Component orientation
  • Wear pattern
  • Fracture or damaged surfaces
  • Lubricant appearance
  • Deposits
  • Looseness
  • Alignment measurements
  • Operating state
  • Alarm history
  • Vibration trends
  • Temperature trends
  • Recent maintenance activity
Repair can alter or destroy the evidence a failure left behind — capturing physical, condition, operating, maintenance and control-system evidence early keeps the investigation possible.

The principle is not: delay a critical repair indefinitely.

It is: restore production while preserving enough evidence to understand what happened.

Start with a precise problem statement

Compare these two ways of describing the same event:

VaguePrecise
Pump keeps failing.Drive-end bearing on the pump has experienced a second high-vibration event followed by overheating under normal production operation.

The precise version is more useful because it begins defining:

  • Which asset
  • Which component
  • What happened
  • Whether it has recurred
  • What symptoms were observed
  • Under what broad operating condition

A problem statement should describe the event. It should not prematurely insert the cause.

For example: "Bearing failed because the operator did not lubricate it."

That is already a conclusion.

If the investigation has not demonstrated inadequate lubrication or why it occurred, the statement introduces bias before the evidence is analysed.

Reconstruct what happened before the failure

A machine does not usually begin failing at the moment it stops.

The observable failure may be the final event in a much longer sequence.

A timeline can therefore be extremely useful. For example:

  • Bearing replaced
  • Machine returned to operation
  • Production loading changed
  • Vibration gradually increased
  • Temperature began rising
  • Abnormal noise reported
  • Machine stopped

Now investigators can ask:

  • Which event happened first?
  • Which variables changed before the symptoms appeared?
  • Was anything modified?
  • Did the same sequence occur during the previous failure?

Useful timeline information can come from maintenance records, operator reports, PLC/controller events, condition-monitoring data, work orders, lubrication records and production records.

The timeline converts a failure from a single moment into a process.

Physical evidence should come before confident explanations

A meeting room can generate dozens of possible causes.

That is useful for expanding the investigation. It is not evidence.

A stronger RCA begins with what is known. For machine failures, useful evidence might include:

Evidence typeExamples
PhysicalFailed component, wear, fracture, discoloration, deposits, deformation
Condition-monitoringVibration, temperature, oil condition, motor current/load
OperatingSpeed, load, duty cycle, process changes
MaintenancePrevious repairs, component replacement, lubrication, alignment, inspection
Control-systemAlarms, machine state, trips, operating sequence

The investigator then asks:

Which causal explanation is consistent with all of this evidence?

Condition monitoring and RCA answer different questions

Condition monitoring may tell us that vibration increased before the machine stopped.

That is valuable. But it does not necessarily tell us why.

Possible causes could include imbalance, alignment, looseness, bearing damage or process loading.

Condition monitoring helps identify what changed and when. RCA attempts to explain what created the change.

This is why good condition-monitoring history can be so valuable after a failure.

A single damaged component shows the final state. A trend may show how the machine arrived there.

The Five Whys can be useful

The Five Whys is one of the best-known RCA techniques.

ASQ describes it as a method of repeatedly asking why in order to move beyond symptoms and explore deeper causes. Importantly, ASQ also notes that it may take more or fewer than five questions; the number five is not a rigid stopping rule.

Consider an illustrative example.

QuestionAnswer
ProblemBearing repeatedly overheats.
Why?Friction increased.
Why?The lubrication condition was inadequate.
Why?The bearing was not receiving the required lubrication at the required interval.
Why?The existing lubrication schedule did not reflect current operating duty.
Why?The maintenance task had not been reviewed after the machine's operating conditions changed.

Now the possible issue has moved from replace the bearing to review how lubrication requirements are controlled when duty changes.

That can produce a more durable improvement.

But there is a major caution.

Five Whys does not prove the answer

Every statement in the previous chain could be wrong.

Perhaps lubrication was adequate.

Perhaps the bearing overheated because fit was incorrect, shaft loading changed, alignment was poor, or another fault increased load.

The Five Whys is therefore best treated as a way of constructing candidate causal reasoning. It is not an evidence generator.

ASQ's 8D problem-solving guidance makes this point especially clearly: causes should be verified or proved, rather than simply determined by brainstorming.

So after developing the chain, investigators should ask:

What evidence demonstrates each important step?

Complex failures may have more than one cause

Machine failure is not always a clean chain where A caused B and B caused C.

Imagine a bearing operating under slight misalignment, marginal lubrication and increased process load.

None of those conditions alone may have produced rapid failure. Together, they may.

Modern RCA therefore needs to tolerate multiple causal paths.

Even ASQ's broader problem-solving guidance notes that complex problems may have more than one root cause and emphasizes asking what other potential causes were studied and eliminated.

This is where a single Five Whys chain can become too narrow.

A cause-and-effect or fishbone diagram can be useful when several families of cause are possible.

ASQ describes it as a method for identifying and organizing many possible causes into useful categories.

For machinery investigations, possible categories might include machine, method, material, measurement, environment and human factors.

For a repeated bearing failure, candidate causes might include:

CategoryCandidate causes
MachineAlignment, shaft condition, housing, fit
MethodInstallation, lubrication procedure, alignment procedure
Material/componentBearing specification, lubricant specification
MeasurementInadequate monitoring, incorrect measurement location
EnvironmentContamination, temperature
Human/system factorsUnclear responsibility, incomplete instructions, missed maintenance task

The diagram helps ensure that one attractive explanation does not dominate too early.

But it still does not prove anything.

A fishbone is a hypothesis map

This distinction is important.

A fishbone diagram containing misalignment does not mean the machine was misaligned. It means misalignment is one candidate cause worth evaluating.

A plausible explanation is a hypothesis; a candidate root cause becomes credible only once it has been checked against evidence.

The next step is evidence.

  • Was alignment measured?
  • What was found?
  • Was it sufficiently abnormal to explain the damage?
  • Does the failure pattern support it?
  • Were competing explanations eliminated?

The RCA process should move from possible, to probable, to supported.

"Human error" is often an incomplete root cause

Suppose an investigation concludes: the operator failed to lubricate the bearing.

Perhaps that action happened. But the investigation should usually continue. Ask:

  • Was lubrication responsibility clear?
  • Was the task scheduled?
  • Was the interval appropriate?
  • Was lubricant available?
  • Was the point accessible?
  • Was the requirement documented?
  • Had operating conditions changed?
  • Was completion verified?

This is not about removing individual responsibility. It is about understanding the entire causal system.

ASQ's problem-solving guidance specifically encourages moving away from simply determining who was at fault and toward understanding how the process can be improved.

Human action can be part of the failure chain without necessarily being the deepest useful cause.

Mechanical causes must not disappear either

There is an opposite RCA mistake.

A team may quickly jump from "bearing failed" to "maintenance-management system problem."

But the physical failure still has to make engineering sense.

A robust analysis can connect several levels. For example:

LevelExplanation
Physical mechanismBearing damage associated with inadequate lubrication
Immediate conditionLubricant supply to the bearing was insufficient
Maintenance-system issueLubrication interval or execution did not match actual operating requirements

The physical and organizational explanations are not competitors. They are different levels of the same causal chain.

Correlation is not automatically cause

Suppose vibration increased shortly before failure. That shows correlation in time.

Did vibration cause the failure? Not necessarily.

The vibration may itself have been another symptom of a developing fault.

Similarly, a temperature increase does not mean high temperature initiated the failure. Perhaps increased friction caused both the temperature rise and the surface damage.

RCA should therefore distinguish a signal associated with the failure from a condition that materially caused the failure.

This is especially important as maintenance becomes more data-driven.

More data produce more correlations. They do not automatically produce causal explanations.

Root cause should be verified against evidence

Suppose the proposed cause is: repeated coupling failure was caused by shaft misalignment.

What evidence should exist if that explanation is correct? Potential evidence could include:

  • Measured alignment outside acceptable limits
  • Wear consistent with the condition
  • Vibration behaviour consistent with misalignment
  • Recurrence associated with installation or alignment changes

Then ask:

What evidence would contradict this explanation?

That second question is particularly useful.

Good investigation should try to disprove a favourite hypothesis, not only collect evidence that supports it.

Corrective action should logically follow from the cause

Imagine the verified problem is alignment.

Then replacing the coupling is repair. Correcting alignment is closer to corrective action.

But even that may be incomplete.

Why was alignment lost? Perhaps foundation movement, poor installation procedure, thermal growth, or inadequate verification after maintenance.

A stronger corrective package may therefore include:

  • Restore correct alignment
  • Address the condition creating alignment drift
  • Define measurement/verification after future work

The best corrective action breaks the causal chain at a practical and controllable point.

Verify the corrective action

RCA should not finish when the report is closed.

The final question is:

Did the change actually prevent the problem we were trying to prevent?

Useful verification might include:

  • Condition returns toward baseline
  • Operating temperature stabilizes
  • Alignment remains acceptable
  • Contamination reduces
  • Failure does not recur during an appropriate observation period

ASQ's structured 8D process places verification of root causes and corrective actions inside the formal problem-solving sequence rather than treating corrective action as an assumption.

NASA failure-analysis guidance expresses a similar engineering principle: root-cause investigation should proceed to a level sufficient to support effective corrective action and recurrence control, while recognizing that additional analysis is not always justified when further insight would not be cost-effective.

Not every failure needs the same investigation depth

A complete RCA costs time, engineering effort, production attention and specialist resources.

A low-consequence consumable reaching its expected wear limit may not justify the same investigation as:

  • Repeated critical bearing failures
  • Major production loss
  • Safety-related failures
  • Expensive component damage
  • Unexplained catastrophic failure

Investigation depth should therefore reflect consequence, recurrence, uncertainty and learning value.

NASA's failure-analysis guidance explicitly recognizes that deeper root-cause investigation may be stopped when further analysis is not cost-effective relative to the risk reduction or insight gained, provided the rationale is documented.

Use enough RCA to make the right engineering decision — not the maximum possible investigation for every event.

An illustrative repeated bearing failure

Consider a fictional production machine.

A bearing fails. It is replaced. Two months later, the replacement bearing begins showing abnormal vibration.

The immediate response could be: replace it again.

Instead, the maintenance team preserves the bearing and operating information. The timeline shows:

  • Bearing replacement
  • Machine returned to service
  • Gradual vibration increase
  • Temperature rise
  • Eventual damage
Replacing the bearing alone keeps the failure on a repeating loop — breaking the loop means preserving evidence, verifying a cause and correcting the condition that damaged the original part. This example is illustrative, not a specific real case.

Inspection suggests a lubrication-related failure mechanism.

Now possible causes are investigated:

  • Wrong lubricant?
  • Contamination?
  • Insufficient quantity?
  • Incorrect interval?
  • Blocked delivery path?
  • Excessive operating load?

Oil and maintenance records show no lubricant-specification change.

Inspection finds no strong evidence of external contamination.

The maintenance history, however, shows that the machine's production duty increased significantly while the lubrication schedule remained unchanged.

That still does not automatically prove causation.

The team evaluates whether the increased duty could materially change lubrication requirements and whether the observed damage is consistent with inadequate lubrication under those conditions.

Suppose the evidence supports the hypothesis.

The corrective action is no longer replace the bearing. It becomes:

  • Replace damaged component
  • Restore correct lubrication condition
  • Revise lubrication task based on actual duty
  • Monitor temperature/vibration following return to service

The example demonstrates the difference between repair and recurrence prevention.

The example is illustrative, not a claim about a specific real failure.

My perspective on recurring failure

In practical engineering work, encountering a machine or component that is repaired and then develops a similar problem again changes the way maintenance is viewed.

The first repair naturally focuses attention on the broken component.

Recurrence forces a different question:

What did we leave unchanged?

That question is central to Root Cause Analysis.

A repeated failure does not necessarily mean the replacement was done badly.

It may mean the repair restored the component while leaving the damaging operating condition intact.

This is why RCA is most valuable when it moves attention from "what part should we replace" to "what allowed this part to fail."

The Machine Failure RCA Chain

A practical machinery RCA can be organized into ten stages.

1. Stabilize

First:

  • Make the equipment safe
  • Control immediate consequences
  • Isolate energy as required
  • Protect people and surrounding equipment

Investigation does not take priority over safety.

2. Preserve

Before evidence disappears, capture photographs, component condition, operating state, trends, alarms and relevant samples.

3. Define

State precisely what failed, where, when, how it was detected and its operational impact. Avoid premature causal language.

4. Reconstruct

Develop the timeline. Include operation, maintenance, modifications, and previous symptoms or failures.

5. Explain the failure mechanism

Ask how the component physically reached this condition. This connects the visible damage to engineering physics.

6. Generate candidate causes

Use physical reasoning, Five Whys, fishbone diagrams, expert knowledge and maintenance history. Do not select a root cause yet.

7. Verify

For each strong candidate: identify supporting evidence, identify contradictory evidence, test where practical, and compare with competing explanations.

8. Correct

Choose actions that interrupt the verified causal path. The action may involve the machine, its design, installation, maintenance, operation or monitoring.

9. Confirm

After implementation: monitor expected indicators, verify that the condition improved, and check for recurrence.

10. Learn

Use the investigation beyond one machine where appropriate. Update maintenance tasks, inspection criteria, design standards, training, condition monitoring and failure databases.

That is how one failure can produce wider reliability improvement.

Common RCA mistakes

  • Replacing the part and calling the issue closed. Repair and RCA have different purposes.
  • Starting the investigation after evidence has disappeared. Capture critical information early.
  • Calling the failed component the root cause. The failed component may be the consequence.
  • Stopping after exactly five Whys. ASQ explicitly notes that fewer or more questions may be required.
  • Treating brainstorming as proof. Candidate causes must be evaluated.
  • Selecting "operator error" and stopping there. Look at the conditions that shaped the action as well.
  • Ignoring physical failure mechanisms. Organizational explanations still need to connect to the machine physics.
  • Searching only for one cause. Complex failures may involve multiple interacting conditions.
  • Implementing an action without verifying the cause. The machine may improve temporarily for an unrelated reason.
  • Closing RCA before checking recurrence. Corrective action needs follow-up.

How RCA improves maintenance strategy

The value of a good RCA can extend beyond one repair.

Suppose an investigation finds that a critical failure was preceded by a detectable vibration pattern. The organization might improve its condition monitoring.

If the investigation identifies lubrication intervals that do not match operating duty, it may improve preventive maintenance.

If degradation can be reliably detected and trended, the asset may become a candidate for predictive maintenance.

If recurring failure substantially reduces production, the improvement may appear later through better Availability and OEE.

This shows how the maintenance topics connect.

RCA is not separate from maintenance strategy. It supplies knowledge that can make the strategy better.

Key takeaway

When a machine fails, replacing the broken component may be necessary. Sometimes it is all that is economically justified.

But when failure is serious, recurrence is expensive, the same component keeps failing, or the cause is unclear, repair alone is not enough.

A strong investigation asks:

  • What failed?
  • How did it fail?
  • What conditions produced that mechanism?
  • What evidence supports that explanation?
  • What change will prevent recurrence?

Finally: did the change work?

Tools such as Five Whys and fishbone diagrams can help organize the investigation. They should never replace evidence.

A root cause is not the most convincing story about a failure. It is a causal explanation strong enough to support effective corrective action and survive comparison with the evidence.

That is the difference between repeatedly repairing a machine and learning from why it failed.

References and further reading

  • ISO 14224:2016 — Petroleum, petrochemical and natural gas industries — Collection and exchange of reliability and maintenance data for equipment. Its formal scope is petroleum, petrochemical and natural-gas industries, not general machinery RCA — but its standardized reliability and failure-mode terminology and data-quality principles are useful more broadly when structuring a machinery investigation.
  • ASQ — Five Whys and Five Hows. Primary guidance on using repeated "why" questions to move beyond symptoms; ASQ notes that the method may require more or fewer than five iterations.
  • ASQ — Eight Disciplines (8D) Problem Solving. Requires identifying and verifying root causes rather than selecting them through unsupported brainstorming, and places corrective-action verification inside the formal problem-solving sequence.
  • ASQ — Fishbone / Cause-and-Effect Diagram. Guidance on using a fishbone diagram to organize multiple potential causes into categories for further investigation.
  • NASA-HDBK-8739.18 — NASA Root Cause Analysis guidance within problem/failure reporting. Engineering guidance on taking root-cause investigation far enough to support effective corrective action and recurrence control, while recognizing practical limits to investigation depth.
02Frequently Asked Questions

A few common questions

Root Cause Analysis is a structured investigation intended to identify the causes that produced a machinery failure so that effective corrective actions can reduce or prevent recurrence.

Usually not by itself. The damaged bearing identifies the failed component or failure location. RCA asks what physical mechanism damaged it and what conditions produced that mechanism.

Failure mode describes how equipment loses or degrades its required function. Root cause addresses why the failure developed. ISO 14224 provides standardized reliability and maintenance failure terminology useful for structuring this information.

It can be useful for relatively straightforward problems, but it does not automatically prove causation. ASQ notes that the method may involve fewer or more than five questions, and structured RCA should verify candidate causes using evidence.

A fishbone diagram is useful for organizing multiple potential causes into categories, particularly when the problem may have several contributing factors. The candidate causes should then be investigated rather than accepted automatically.

A replacement can restore function while leaving conditions such as misalignment, contamination, lubrication problems, overload or maintenance-system weaknesses unchanged. Recurrence is therefore a strong reason to investigate beyond the component itself.

No. Investigation depth should reflect factors such as safety consequence, production impact, cost, recurrence, criticality and uncertainty. Failure-analysis guidance recognizes that deeper investigation may not always be justified when the additional insight or risk reduction does not justify the effort.

Develop evidence that should exist if the proposed cause is true, compare it with actual physical, operational and historical evidence, examine alternative causes, and where possible test the hypothesis.

A strong RCA is not complete simply when a cause is written in a report. The corrective action should be implemented and its effectiveness monitored to determine whether the damaging condition and recurrence have actually been controlled.

05About the Author
Harun Lucas working at his desk, reviewing code and systems dashboards across multiple monitors

Harun Lucas

Mechanical Engineer · Technology Education Researcher · Engineering Systems Developer

Harun writes from the same practice covered on this site — mechanical engineering, technology education research, and engineering systems development — connecting hands-on work with the ideas behind it.

More About Harun