FMEA
Failure Mode and Effects Analysis (FMEA) is a structured, team-based risk assessment methodology used to identify potential failure modes in product designs or manufacturing processes. Teams evaluate failure effects and assign 1 to 10 ratings across three dimensions: Severity, Occurrence, and Detection. Multiplying these factors produces a Risk Priority Number ranging from 1 to 1,000. By quantifying risks, organizations rank operational vulnerabilities, allocate engineering resources, and deploy preventative countermeasures before defects reach production or customers.
- Three judgments, one row each
- Every failure mode gets three 1–10 scores: Severity — how bad the effect is; Occurrence — how often it happens; Detection — how likely it escapes before being caught (10 = escapes). The scores come from the team that knows the process, anchored by shared scales.
- RPN: the multiplication
- Risk Priority Number = severity × occurrence × detection. The bent pin scores 7 × 5 × 6 = 210. Multiplying means one terrible dimension is enough to surface a risk even when the other two look mild.
- The sheet exists to rank
- Rows sort worst-first, and the top row becomes the assignment: a named action with an owner and a date. An FMEA that produces no action has not done its job.
- Attack occurrence, not consequence
- The action is a keyed insertion fixture — a poka-yoke that makes off-axis insertion physically impossible. Rescored after it lands, occurrence falls 5 → 2 and the RPN 210 → 84, but severity stays 7: a bent pin is exactly as bad whenever it happens. Good actions move O and D; S almost never moves.
- Detection 7
- The off-center gasket only ranks second by RPN, but its detection score of 7 means it is the failure most likely to reach a customer. Once the pin fixture is installed and that row is rescored, this one tops the sheet — the ranking always shows the next priority.
Key facts
- Origin
- Late 1940s (US Military standard MIL-P-1629)
- Core dimensions
- Severity, Occurrence, Detection
- Scoring scale
- 1 to 10 for each dimension
- Formula
- Severity × Occurrence × Detection
- Score range
- 1 to 1,000
- Primary variations
- Design, Process, System, Service FMEA
By Matthew Savas — Founder of Kaizumi. Reviewed 30 August 2026.
Failure Mode and Effects Analysis (FMEA) is a structured risk assessment method that identifies all potential failure modes in a design or process, evaluates their effects, and assigns a 1 to 10 score across three dimensions: Severity, Occurrence, and Detection. Multiplying these factors produces a Risk Priority Number, which is calculated as Severity multiplied by Occurrence multiplied by Detection. This quantitative output allows cross-functional teams to rank risks, allocate engineering resources, and prioritize preventative countermeasures before defects reach production or the end customer.
Historical background
The United States military originally developed Failure Mode and Effects Analysis in the late 1940s under the military standard MIL-P-1629, titled Procedures for Performing a Failure Mode, Effects and Criticality Analysis. The primary objective was to evaluate the impact of system and equipment failures on mission success and personnel safety in aerospace and defense applications.
The National Aeronautics and Space Administration later adopted the methodology during the Apollo space program in the 1960s to ensure the reliability of complex flight systems. In the 1970s, the automotive industry adopted FMEA following high-profile reliability and safety challenges. Ford Motor Company integrated the tool into its product design and manufacturing workflows, leading to broader industry adoption. By the 1990s, the Automotive Industry Action Group and the German Association of the Automotive Industry standardized FMEA manuals across the global supply chain, embedding the technique into quality management systems and Six Sigma continuous improvement methodologies.
Core dimensions and scoring criteria
An FMEA relies on a team-based evaluation using standardized rating scales from 1 to 10 for each of the three dimensions. The cross-functional team must establish clear scoring guidelines before conducting the analysis to maintain consistency across different products and processes.
Severity
Severity measures the seriousness of the effect of a potential failure mode on the customer, downstream operations, or regulatory compliance. A failure mode that causes a minor cosmetic blemish scores low, whereas a failure that risks operator safety, violates regulatory standards, or causes complete system shutdown receives a high score, typically from 8 to 10. Severity ratings cannot be reduced unless the product design or the underlying process architecture is fundamentally altered.
A typical 1 to 10 Severity scale follows this distribution:
- 1: No noticeable effect on product performance, process, or customer.
- 2 to 3: Slight customer annoyance; slight defect that can be remedied with minor rework.
- 4 to 6: Moderate customer dissatisfaction; reduced system performance or significant process rework required.
- 7 to 8: High customer dissatisfaction; system becomes inoperable or major component fails, but without safety risks.
- 9 to 10: Critical failure involving non-compliance with safety regulations, hazardous operating conditions, or danger to the operator or end user.
Occurrence
Occurrence measures the likelihood or estimated frequency that a specific failure mode and its associated cause will take place during the planned life of the product or process. The score is assigned from 1 to 10 based on observed historical defect data, statistical process control capability metrics, or engineering estimates.
A standard Occurrence scale includes:
- 1: Failure is extremely unlikely; process is under strict statistical control with historical defect rates approaching zero.
- 2 to 3: Low failure rate; isolated defects occur infrequently under similar operational parameters.
- 4 to 6: Moderate failure rate; process experiences occasional instability or documented historical occurrences.
- 7 to 8: High failure rate; process produces frequent non-conformances during standard operational cycles.
- 9 to 10: Very high failure rate; failure is almost inevitable under current operating conditions.
Detection
Detection measures the capability of current control methods, inspection systems, or verification tests to identify a failure mode or its cause before the item leaves the workstation, manufacturing facility, or design gate. The scale is inverse: a score of 1 indicates virtually certain detection, such as an automated sensor that stops the line, while a score of 10 indicates that existing controls are unlikely or unable to catch the defect.
A standard Detection scale is structured as follows:
- 1 to 2: Automated error-proofing, interlocking sensors, or 100 percent automated inspection that prevents defective parts from proceeding.
- 3 to 4: High probability of detection using automated gauging or rigorous multi-stage manual verification at the workstation.
- 5 to 6: Moderate probability of detection through statistical sampling, end-of-line functional tests, or standard visual inspection.
- 7 to 8: Low probability of detection; controls rely on random spot checks, visual inspections under poor conditions, or subjective assessment criteria.
- 9 to 10: No existing controls; inspection cannot detect the defect, or testing is not performed prior to shipping.
Risk Priority Number
These three ratings are multiplied together to calculate the Risk Priority Number:
Risk Priority Number equals Severity multiplied by Occurrence multiplied by Detection.
The resulting score ranges from 1 to 1,000. The calculated score provides an initial numerical rank to focus engineering attention on the highest risk failure modes. A core operational rule in standard practice dictates that the top Risk Priority Number becomes a named action assigned to an owner with a strict implementation deadline.
Variations of FMEA
Organizations apply FMEA across different phases of the product life cycle. The primary variations address distinct operational domains:
- Design FMEA: Focuses on product design, material selection, engineering tolerances, and component geometry before releasing designs to production. Design FMEA evaluates potential product malfunctions, shortened service life, and maintenance vulnerabilities caused by design choices.
- Process FMEA: Evaluates manufacturing and assembly processes. It assumes the product design is established and examines how human error, machine variation, tooling wear, environmental conditions, and material handling can cause process non-conformances.
- System FMEA: Analyzes high-level system interactions, sub-system interfaces, and integration failures across complex software, electrical, and mechanical assemblies.
- Service FMEA: Identifies failure modes in transactional, administrative, or customer service processes before rolling out new operational procedures.
Step-by-step implementation process
Executing an FMEA requires a disciplined sequence of steps managed by a team that includes design engineers, quality specialists, maintenance technicians, operators, and manufacturing engineers.
- Assemble the cross-functional team and define the scope: Establish the boundary of the analysis using process maps, engineering schematics, and functional block diagrams.
- List process steps or design functions: Break down the target process into sequential operational steps or map every functional requirement of the design.
- Identify potential failure modes: Brainstorm every conceivable manner in which the component or step could fail to meet its functional specification.
- Determine effects and assign Severity: Analyze the consequences of each failure mode on the customer, downstream operations, and regulatory compliance, then assign a 1 to 10 Severity score.
- Identify causes and assign Occurrence: Perform root cause analysis to determine the underlying drivers behind each failure mode and score the Occurrence from 1 to 10 based on observed or expected frequency. Teams often apply a fishbone diagram to isolate the contributing variables.
- Evaluate existing controls and assign Detection: Document existing inspection, testing, or process controls and assign a Detection rating from 1 to 10 based on their reliability.
- Calculate the Risk Priority Number: Multiply Severity, Occurrence, and Detection for each identified failure mode.
- Prioritize and execute preventative countermeasures: Target high scores by implementing physical improvements, such as poka-yoke mechanisms, tool redesigns, or tighter process controls.
- Recalculate the Risk Priority Number: After countermeasures are deployed and verified, re-evaluate Occurrence and Detection to quantify risk reduction and update the documentation.
Teams can build your own FMEA with the free template to standardize this workflow across internal projects.
Worked example: electronic connector assembly
Consider a manufacturing workstation responsible for inserting a multi-pin wiring harness into a primary circuit board housing. During the initial Process FMEA, the team evaluates the assembly operation and identifies the top risk.
The initial baseline assessment is recorded as follows:
- Process step: Manual harness connector insertion into board housing.
- Potential failure mode: Bent connector pin during manual alignment.
- Potential effect of failure: Intermittent electrical open circuit during final vehicle operation, leading to dashboard warning lights and loss of sensor feedback.
- Potential cause: Operator inserts the connector at an off-center angle under tight cycle time constraints.
- Existing controls: Visual inspection by the operator prior to passing the board to the next workstation.
The team scores the initial conditions:
- Severity is scored at 7 because the defect leads to component loss of function and customer dissatisfaction.
- Occurrence is scored at 5 because manual alignment errors occur at a moderate historical defect frequency.
- Detection is scored at 6 because visual inspection of enclosed connector pins under standard cycle times misses minor pin deflections.
The initial Risk Priority Number calculation is: 7 multiplied by 5 multiplied by 6 equals 210.
Because 210 is the top risk shown across the assembly process, it becomes a named action. The engineering team designs and installs a precision alignment fixture and guide block. The fixture physically prevents the harness from engaging the pins if the insertion angle deviates by more than 0.5 degrees.
After the fixture is installed and tested on the line:
- Severity remains 7 because the consequence of an open circuit has not changed.
- Occurrence drops from 5 to 2 because the mechanical guide prevents off-angle insertion.
- Detection remains 6.
The revised Risk Priority Number calculation is: 7 multiplied by 2 multiplied by 6 equals 84.
The countermeasure reduces the total risk score by 126 points, demonstrating a measurable improvement in process capability.
Analysis and limitations of the Risk Priority Number
While the Risk Priority Number provides a straightforward ranking metric, relying solely on total score can produce misleading risk assessments. Because the calculation uses simple multiplication, mathematical anomalies arise where distinct risk profiles yield identical scores.
Two failure modes can share an identical Risk Priority Number yet present fundamentally different risk profiles:
- A failure with Severity 8, Occurrence 3, and Detection 7 produces a Risk Priority Number of 168 and represents a critical safety or field-escape hazard.
- A failure with Severity 3, Occurrence 8, and Detection 7 also produces a Risk Priority Number of 168 but represents a minor, frequent inconvenience.
Treating both failure modes with equal urgency based solely on the total score can lead teams to prioritize high-frequency minor defects over low-frequency catastrophic failures. To prevent this, standard operating guidelines require that any failure mode with a Severity rating of 8, 9, or 10 receives immediate corrective action regardless of its total Risk Priority Number.
Additionally, the distribution of numbers between 1 and 1,000 is non-linear. Many numbers cannot be formed by the product of three integers between 1 and 10, creating statistical clusters and gaps across the spectrum. Modern quality standards, such as the harmonized AIAG-VDA FMEA standard, frequently replace the simple Risk Priority Number with an Action Priority logic table that gives higher structural weight to Severity first, Occurrence second, and Detection third.
Integration with lean and quality systems
An FMEA is not a static document; it functions as a live repository of organizational process knowledge. To maintain operational stability, organizations integrate FMEA outcomes into daily manufacturing and continuous improvement routines:
- Risk prioritization: Teams use a Pareto chart to rank Risk Priority Numbers across an entire facility or line, directing capital expenditure toward the 20 percent of failure modes that generate 80 percent of overall operational risk.
- Process control plans: The Detection controls established during the FMEA directly populate the plant Control Plan, specifying gauge types, sample frequencies, and containment actions.
- Shop floor verification: Quality audits, including kamishibai card systems, incorporate verification checks on designated FMEA high-risk controls to ensure error-proofing devices remain active and operational during production shifts.
- Continuous engineering cycles: When a customer complaint, internal scrap spike, or audit finding occurs, engineers open the FMEA to verify whether the failure mode was anticipated, update the Occurrence and Detection scores based on new empirical data, and implement permanent structural countermeasures.