Root Cause Analysis
Root Cause Analysis: After a significant problem occurs, systematically trace backwards from the symptom to the underlying causal factors. Fix the root cause β the deepest addressable systemic failure β not just the symptom. RCA distinguishes between "what happened" (symptom) and "why the system allowed it to happen" (root cause), preventing recurrence rather than just recovery.
What Is Root Cause Analysis?β
Root Cause Analysis is a family of methods β not a single technique β united by one principle: when a significant problem occurs, understanding and fixing the immediate trigger is insufficient. The goal is to identify the deepest causal factor that, if corrected, would prevent the problem from recurring.
RCA emerged from industrial safety engineering in the mid-20th century. Major industrial disasters β Texas City refinery explosion (2005), Three Mile Island (1979), Challenger Space Shuttle (1986) β were studied using RCA methodologies and revealed that nearly every major accident has multiple contributing causes, typically including organisational and systemic failures, not just technical ones.
The key distinction in RCA is between symptom, proximate cause, and root cause. A hospital patient receives the wrong medication (symptom). The nurse picked up the wrong vial (proximate cause). The medications were stored in identical-looking vials adjacent to each other, with no double-check protocol, in an understaffed unit during a 12-hour shift (root causes β multiple). Firing the nurse addresses the proximate cause; redesigning storage, implementing double-checks, and reviewing shift length addresses root causes.
RCA is often confused with blame assignment. The best RCA practice (derived from the aviation and nuclear industries) assumes that systems fail before people fail: when a person makes an error, ask what the system should have done to prevent a competent person from making that error, rather than attributing causality to individual incompetence.
How It Worksβ
Step 1: Contain the immediate problem
β Stop the bleeding before diagnosing
β RCA happens after stabilisation, not during it
Step 2: Define the problem precisely
β What happened? When? Where? How severe?
β Quantify impact: how many affected, what cost, what duration?
Step 3: Collect data and preserve evidence
β Gather logs, records, witness accounts, measurements
β Time-sequence the events leading to the failure
Step 4: Identify causal factors (use 5 Whys, Fishbone, or fault trees)
β Map all contributing causes, not just the final trigger
β Distinguish physical causes, human causes, and systemic/latent causes
Step 5: Identify root cause(s)
β The deepest cause(s) that, if corrected, prevent recurrence
β Often: a missing or defective system, policy, training, or design element
Step 6: Design and implement corrective actions
β Address root causes, not just symptoms
β Assign owners, timelines, and verification criteria
Step 7: Verify and monitor
β Did the fix work? Has the problem recurred?
Three Real-World Examplesβ
Aviation: Air France 447 (2009)β
Air France Flight 447 disappeared over the Atlantic. The proximate cause: pilots failed to maintain safe airspeed and pulled back on the stick during a stall, causing the aircraft to descend into the ocean. But RCA revealed multiple root causes: pitot tube icing gave false airspeed readings; the aircraft switched to alternate law (limiting automated protections) without sufficient crew awareness; pilot training for manual flight in high-altitude stalls was inadequate; and cockpit crew coordination protocols failed during the emergency. Fixes addressed all root causes: pitot tube design changes across the fleet, revised training requirements, and updated crew communication protocols β preventing similar incidents, not just punishing crew error.
Healthcare: Central Line Infections at Johns Hopkinsβ
Researcher Peter Pronovost identified that central-line bloodstream infections (CLABSIs) killed approximately 31,000 patients annually in the US. RCA revealed root causes: no standardised insertion checklist, inconsistent hand hygiene compliance, and no empowerment for nurses to halt non-compliant procedures. Pronovost's intervention β a five-item checklist plus nurse authority to stop violations β reduced CLABSI rates at Michigan ICUs by 66% within 18 months. The root cause was not physician negligence; it was absent systems for consistent practice.
Software: GitLab Database Incident (2017)β
GitLab accidentally deleted a production database. Proximate cause: an engineer ran the wrong command under pressure. RCA: no backup restoration had ever been tested; automated backups existed but weren't monitored for completion; a script ran on the wrong database; and there was no "dry run" mode for destructive operations. The incident post-mortem (conducted publicly) produced 20+ systemic corrective actions β tested backups, deployment safeguards, runbook improvements β rather than punishing one engineer. GitLab's transparent RCA has become an industry reference for post-mortem culture.
When to Use Itβ
β RCA is essential for:
- Any significant failure with potential for recurrence (outages, defects, safety incidents)
- Post-mortem reviews after project failures or near-misses
- Quality improvement initiatives targeting persistent, recurring problems
- Regulatory requirements in healthcare, aviation, nuclear, and finance
β Less appropriate for:
- Novel problems with no prior pattern (not enough data)
- Very minor incidents where RCA cost exceeds benefit
- Problems requiring creative solutions rather than causal analysis
| Pairs well with | Why |
|---|---|
| 5 Whys | 5 Whys is the simplest RCA technique |
| Fishbone Diagram | Fishbone maps all causal candidates before drilling |
| Black Box Thinking | Black Box Thinking is the cultural mindset that makes RCA valuable |
| Pre-mortem | Pre-mortems apply RCA logic before failure occurs |
Common Misuses and Limitationsβ
Stopping at the human error. Human error is almost always a proximate cause, not a root cause. The system that put a person in a position to make a consequential error is the root cause.
Single root cause thinking. Major failures almost always have multiple contributing root causes at different levels. Swiss cheese model: each layer of defence has holes; the accident occurs when holes align. Identifying one root cause and stopping misses the systemic complexity.
RCA as blame assignment. Cultures where RCA is used to punish produce RCAs designed to protect people from punishment β not RCAs that reveal systemic truth. Blameless (or "just culture") post-mortems, as practised in aviation and high-quality software organisations, produce far more honest and useful analyses.
Inadequate corrective action follow-through. Studies of hospital RCA programmes found that the majority of recommended corrective actions are never implemented. RCA without accountability for implementation is wasted effort.
Related Modelsβ
| Model | Relationship |
|---|---|
| 5 Whys | 5 Whys is the most common RCA technique |
| Fishbone Diagram | Fishbone is a complementary RCA brainstorming tool |
| Black Box Thinking | Black Box Thinking is the cultural context in which RCA thrives |
| Proximate vs Root Cause | RCA operationalises the distinction between proximate and root causes |
Frequently Asked Questionsβ
What is "blameless post-mortem" culture and why does it matter?
Blameless post-mortems, popularised by Google's Site Reliability Engineering practice, proceed from the assumption that engineers act in good faith with the information available to them. When they make mistakes, the investigation asks: what information, tools, or processes should the system have provided to prevent this? Not: why did this person fail? This approach produces honest incident reports and systemic fixes. Blame-based cultures produce cover-ups and surface-level recommendations. The aviation industry's adoption of blameless reporting (especially for near-miss incidents) has been a major driver of its extraordinary safety record.
How many root causes should an RCA identify?
As many as genuinely contributed to the problem β typically 3β7 for significant incidents. The goal is not parsimony; it's completeness. Each identified root cause should have a corresponding corrective action. However, not all root causes are equally important: prioritise those with the highest recurrence risk and the most actionable fixes. A single well-implemented fix for the highest-leverage root cause beats seven incomplete fixes across all contributing factors.
How long should an RCA take?
It scales with incident severity. A minor software bug might warrant a 30-minute team retrospective. A production outage affecting millions of users warrants 1β2 days of thorough investigation before publishing a post-mortem. A major aviation accident warrants months or years (the NTSB takes 12β18 months for major incidents). The investment should be proportional to the recurrence risk and the severity of consequences β spending 3 days on a $500 incident is not value-adding; spending 3 hours on a $500,000 outage is essential.
Further Readingβ
- Reason, J. (1990). Human Error β the Swiss Cheese Model and systems approach to accidents
- Leveson, N. (2011). Engineering a Safer World β systems-theoretic accident analysis
- Google SRE Book, Chapter 15: "Postmortem Culture: Learning from Failure" (free online)
Apply with AIβ
π Run a root cause analysis with MindMax β
This page is part of the MindMax Mental Models Knowledge Base.