FMECA: how to run a failure modes, effects and criticality analysis
FMECA works through how equipment can fail, what each failure does and which failures matter most, so maintenance effort goes where the risk is. The method step by step, how to score criticality, where the risk priority number falls short, and a free worksheet template to start from.

Key takeaways
- FMECA works through each way an asset can fail, what the failure does and how much it matters, so maintenance effort is spent where the risk is.
- FMEA stops at failure modes and their effects. FMECA adds a criticality ranking, most often severity against occurrence on a risk matrix.
- Write failure modes as the specific event that causes the failure, at the level a maintenance task can act on. "Pump fails" is a symptom, not a failure mode.
- The risk priority number multiplies severity, occurrence and detection, and it can rank a nuisance level with a hazard. Use it to sort within a band, not to set the band.
- An FMECA pays off only when every ranked failure mode ends in a task, a design change or a recorded decision to run to failure, loaded into the CMMS.
Failure modes, effects and criticality analysis, usually shortened to FMECA, is a structured way of working out how equipment can fail, what each failure does, and which failures matter most. It is the analytical core of most maintenance strategy work. Done well, it gives every maintenance task a reason, in the form of the failure mode it manages, and gives every significant failure mode a deliberate decision.
The method is old and well documented. It came out of United States military reliability work in the late 1940s, and IEC 60812 now sets out how to plan, run and document it. What goes wrong in practice is rarely the method itself. It is scope that is too wide, failure modes that are really symptoms, scores nobody calibrated, and results that never reach the CMMS.
This guide covers what FMECA is and where it fits, the standards, how to run one step by step, the ways to rank criticality and where the risk priority number falls short, how the results become maintenance tasks, and the mistakes that turn an analysis into shelfware. A free FMECA worksheet template comes with it, set up with the worked example used throughout.
Free download
FMECA worksheet template
The worksheet, scales and task selection guide in one workbook, with formulas for criticality, band and RPN, and the worked example from this guide filled in.
The difference between FMEA and FMECA
A failure modes and effects analysis (FMEA) lists the ways an item can fail and the effects of each failure. A failure modes, effects and criticality analysis (FMECA) goes one step further and ranks each failure mode by how critical it is, so the analysis ends in a priority order rather than a long list.
IEC 60812 treats FMECA as a variant of FMEA, in which the ranking takes in at least the severity of the consequences and often other measures, such as how likely the failure is. In practice the names are used loosely, and many documents called an FMEA include a criticality ranking. What matters is whether the ranking drives decisions.
The same method is used in three quite different settings, and it helps to know which one you are in.
| Type | What it analyses | Typical use |
|---|---|---|
| Design FMEA | How a product or system design could fail to meet its requirements | Engineering design, before the item is built |
| Process FMEA | How a manufacturing or assembly process could produce a defect | Manufacturing quality, common in automotive supply chains |
| Equipment or maintenance FMECA | How installed equipment fails in its operating context, and what each failure does to the operation | Maintenance strategy and RCM for plant, mobile fleets and infrastructure |
This guide is about the third kind, where the question is not how to improve a design but how failures are prevented, detected or tolerated once the equipment is running.
Where FMECA fits in maintenance strategy
FMECA sits between the criticality assessment and the maintenance plan. An asset criticality assessment ranks whole assets or systems by the consequence of their failure. FMECA then goes inside the systems that rank highest and works out how each one fails. Reliability centred maintenance decision logic takes those failure modes and chooses a task for each.
The order matters because FMECA is expensive. A thorough analysis takes a facilitated team through every function and failure of a system, and it is rarely worth doing on equipment whose failure costs little. Ranking criticality first, with a tool such as the Criticality Matrix Calculator, keeps the deep analysis on the systems that carry the risk.
FMECA is also used on its own, outside a full RCM program. Typical uses are reviewing the strategy for a problem asset, preparing the maintenance plan for new equipment during operational readiness, deciding which spares to hold, and understanding what a proposed modification would do to reliability.
The standards and references
No single standard is mandatory for a maintenance FMECA. The references below are the ones most often used, and many sites pair one of them with their own risk framework.
| Reference | What it covers | Useful for |
|---|---|---|
| IEC 60812:2018 | The general method for FMEA and FMECA, from planning to reporting, including criticality methods such as the risk priority number and the criticality matrix. Adopted in Australia as AS/NZS IEC 60812:2020. | The default reference for any FMECA |
| MIL-STD-1629A | The original military procedure, with a quantitative criticality number built from failure rates. Cancelled in 1998 but still widely cited. | Quantitative criticality where failure rate data exists |
| AIAG and VDA FMEA Handbook (2019) | A seven step method for design and process FMEA, which replaced the risk priority number with action priority tables. | Manufacturing quality and automotive supply |
| SAE JA1011 and JA1012 | The criteria a process must meet to be called RCM, and a guide to applying them. | FMECA as part of a full RCM analysis |
| ISO 14224:2016 | Standard taxonomies for equipment, failure modes, failure mechanisms and causes, with failure mode codes by equipment class. | Consistent failure mode wording and failure data |
Before the first session
Most of the quality of an FMECA is set before the team meets. Decide these things first and write them down.
- The system and its boundary. Where it starts and stops, and whether interfaces such as the power supply and the control system are in or out of scope.
- The operating context. Duty, throughput, operating hours, redundancy, environment and constraints. The same pump can deserve a different strategy in a different service.
- The level of analysis. Record failure modes at the maintainable item, such as a pump, a gearbox or a filter, rather than at every part.
- The scales. Severity, occurrence and, if used, detection, with anchors the site already recognises. Borrowing the site risk matrix avoids a second, competing definition of risk.
- The information. Drawings and P&IDs, manufacturer manuals, the current maintenance plan, and failure history from the CMMS, ideally coded consistently enough to count failures by mode.
- The team. A facilitator who knows the method, the operators and maintainers who know how the equipment behaves, and an engineer who can make decisions stick.
An hour spent agreeing the boundary and the scales saves days of arguing about them in the sessions.
How to run an FMECA, step by step
The steps below follow the structure of IEC 60812 and the questions RCM asks of each asset. The worksheet template has a column for each, and the examples come from its first row.
| Step | What to record | Example |
|---|---|---|
| 1. Functions | What the item must do, with a performance standard where one applies | Deliver filtered oil to the crusher bearings at the required pressure |
| 2. Functional failures | Each way it can fail to do that, fully or partly | Fails to deliver any oil flow |
| 3. Failure modes | The specific event that causes each functional failure | Pump bearing seizes |
| 4. Causes | Why the failure mode happens, the mechanism behind it | Bearing fatigue at end of life |
| 5. Effects | What happens locally and at the level of the operation, and whether the failure is evident or hidden | Low oil pressure trip stops the crusher. With no spare pump the plant is down until the pump is replaced |
| 6. Current controls | What already prevents or detects the failure mode | Low oil pressure trip. Operator round each shift |
| 7. Scores | Severity, occurrence and detection on the agreed scales, with the reasoning | Severity 4, occurrence 3, detection 3 |
| 8. Criticality | The ranking, from severity and occurrence on the risk matrix | 12 of 25, high |
| 9. Actions | A task, design change or recorded decision, with an owner | Monthly vibration route on the pump and motor bearings. Hold a spare pump on site |
Writing failure modes that are useful
The failure mode is the most important entry in the worksheet and the one most often written badly. A failure mode is the event that causes a functional failure. It should name the part that fails and how it fails, specifically enough to point at a cause and a task.
"Pump fails" is not a failure mode. It is a functional failure at best, and it gives the team nothing to act on. "Pump bearing seizes from fatigue", "mechanical seal leaks from face wear" and "impeller erodes in abrasive slurry" each suggest a different detection method and a different task.
- Record failure modes at the level where a task can be applied, usually the maintainable item. Going deeper multiplies the rows without changing the decisions.
- Include failure modes that have happened, those that happen on similar equipment elsewhere, and those that have not happened yet but reasonably could, especially where the consequence would be severe.
- Include operating and human causes, such as an incorrect start-up or overloading, where they are credible. Not every failure starts in the equipment.
- Use consistent wording. The failure mode codes in ISO 14224, such as FTS for fail to start on demand and ELP for external leakage of process medium, keep the analysis comparable with the failure data the CMMS collects.
If a failure mode does not suggest what you would do about it, it is not specific enough yet.
Describing effects, and spotting hidden failures
Effects describe what happens when the failure mode occurs, assuming nothing is done to stop it. Record the local effect on the item and the end effect on the system or the operation, in enough detail to judge severity. Say what the operators would see, how long the loss lasts and what else is damaged.
Note whether each failure is evident or hidden. An evident failure makes itself known to the operating crew in normal operation. A hidden failure does not, usually because the item is a protective device that only acts on demand, such as a trip switch, a relief valve or a standby pump. Its consequence is felt only when something else fails as well, which is why it needs its own treatment.
The low oil pressure switch in the worked example shows how this plays out. While oil pressure is normal, a stuck switch changes nothing. If the pump then fails, the crusher runs on without lubrication and the main bearings are destroyed. No monitoring of the switch in service would reveal the fault. It has to be tested.
Scoring severity, occurrence and detection
Each failure mode is scored on three scales. Severity rates the worst credible end effect. Occurrence rates how often the failure mode is expected to happen. Detection rates how likely the current controls are to catch it before the effect occurs.
Anchor every level with a description the team can test a score against, and keep the scales short. Five levels are easier to apply consistently than ten, because people can actually tell neighbouring levels apart. The worksheet uses the same five level severity and occurrence scales as the Criticality Matrix Calculator, so failure modes rank on the same basis as the assets they belong to, and adds the detection scale below.
Occurrence is where failure history earns its keep. Where the CMMS records failures against the right equipment with consistent codes, occurrence can be counted rather than guessed, and the MTBF / MTTR Calculator turns counts into rates. Where history is thin, use experience on similar equipment, manufacturer data or published reliability data, and note that the score is an estimate.
| Detection score | Meaning |
|---|---|
| 1. Almost certain | Continuous monitoring with an alarm, or obvious to operators at once |
| 2. High | Found by current inspections or condition monitoring with time to act |
| 3. Moderate | Detectable, but only if a check falls inside the warning period |
| 4. Low | Little warning, and current checks are unlikely to catch it |
| 5. None | Hidden until it fails on demand or causes the effect |
Score against the anchors, not against the row above. A team that compares rows drifts toward rating everything a three.
Ways to rank criticality
The criticality in FMECA can be worked out in several ways. IEC 60812 describes more than one, and the right choice depends on the data available and what the ranking is for.
| Method | How it works | Strengths | Weaknesses |
|---|---|---|---|
| Criticality matrix | Severity against occurrence on a risk matrix, often the site's own, giving a score and a band | Simple, visual and consistent with how the site already describes risk | Ignores detection unless it is recorded separately |
| Risk priority number (RPN) | Severity × occurrence × detection, each on the same scale, giving one number | Quick to calculate and sort, and takes detection into account | Equal numbers can hide very different risks, and a low score can mask a severe consequence |
| Action priority | A lookup table in the AIAG and VDA handbook assigns high, medium or low priority to every combination of scores, weighted toward severity | A severe consequence is never ranked low | Written for design and process FMEA, and the tables sit in the handbook |
| Quantitative criticality number | Criticality from the failure rate, the share of failures in each mode, the probability of the effect and the operating time, as in MIL-STD-1629A | Objective where good failure rate data exists | Needs failure mode data that most maintenance teams do not have |
- Most maintenance FMECAs use a criticality matrix, because it speaks the same risk language as the rest of the site and the band it produces maps directly to a strategy.
- Recording detection alongside is still worthwhile. It shows where the current controls are weak, which is often where the quickest improvement lies.
Why the risk priority number misleads
The risk priority number is the most familiar criticality measure and the most criticised. The scores it multiplies are rankings, not measurements, so the product behaves in ways that surprise people.
- Different risks score the same. On five point scales, a failure mode with severity 5, occurrence 1 and detection 2 scores 10, exactly the same as a nuisance with severity 1, occurrence 5 and detection 2.
- Detection can bury severity. A severe failure mode with good detection can rank below a trivial one with poor detection, even though a detection control can fail too.
- The numbers are lumpy. Only some values between 1 and 125 can occur and many combinations crowd the middle, so a one point change in a single score can move a failure mode a long way up or down the list.
- Thresholds invite gaming. A rule such as acting on anything over 50 encourages teams to shade scores to land just under it.
If you use RPN at all, use it to sort failure modes within a band, and let severity and criticality decide the band.
A worked example on a crusher lubrication system
The worksheet template comes with this example filled in. It is the lubrication system for a primary crusher with a single, unspared lube oil pump, the same asset scored in the Criticality Matrix Calculator example. Six failure modes are shown, one for each item, to illustrate the range of outcomes rather than to be complete.
Ranked by criticality, the pump bearing comes first and earns monitoring and a spare. Ranked by RPN, the low oil pressure switch comes first, because its failure is hidden. Both views are telling you something. The switch does not need a higher band. It needs a test, because no other kind of task can find a hidden failure.
| Failure mode | S | O | D | Criticality | Band | RPN | Task type |
|---|---|---|---|---|---|---|---|
| Lube oil pump, pump bearing seizes | 4 | 3 | 3 | 12 | High | 36 | Condition-based task |
| Oil supply filter, filter element blocks | 3 | 3 | 3 | 9 | Medium | 27 | Condition-based task |
| Low oil pressure switch, switch contacts stick closed | 5 | 2 | 5 | 10 | Medium | 50 | Failure-finding task |
| Oil cooler, cooler tube leaks | 3 | 2 | 2 | 6 | Medium | 12 | Condition-based task |
| Return line hose, hose fitting leaks | 2 | 3 | 1 | 6 | Medium | 6 | Condition-based task |
| Local pressure gauge, gauge mechanism fails | 1 | 2 | 2 | 2 | Low | 4 | Run to failure |
- The pump bearing ranks highest on criticality. A monthly vibration route gives warning of a developing fault, and a spare pump on site shortens the outage if the warning is missed.
- The pressure switch ranks highest on RPN because its failure is hidden. A six-monthly function test against a calibrated source is the only task that can find it.
- The filter and the cooler are managed on condition, through the differential pressure indicator and oil analysis, because both failures develop slowly and give warning.
- The hose leak is caught on the operator round. The local gauge is run to failure, because the pressure transmitter and trip already provide the protection.
From failure modes to maintenance tasks
The analysis earns its cost at this step. Every failure mode above the lowest band needs a decision, and the decision should follow from how the failure mode behaves rather than from habit. RCM decision logic, as described in SAE JA1012, is the usual framework, and the table summarises it in the order the logic considers each option.
| Task type | When it fits |
|---|---|
| Condition-based task | The failure mode gives detectable warning, with a P-F interval long enough to plan and act on. |
| Scheduled restoration | The failure mode is age related, and reworking the item restores its resistance to failure. |
| Scheduled replacement | The failure mode is age related, and fitting a new item restores the original reliability. |
| Failure-finding task | The failure is hidden, typically in a protective device that only acts on demand, so it has to be tested. |
| Run to failure | The consequence is acceptable and no task is worth its cost. Record it as a decision. |
| Redesign or change | No task brings the risk to an acceptable level. Where the consequence is safety or environmental, a change is compulsory. |
- Write each task so it can be loaded into the CMMS as it stands, with what to check, the acceptance limit, the interval and the trade.
- Set condition-based intervals from the P-F interval, and choose techniques as the condition monitoring guide describes.
- Set failure-finding intervals from how often the protective device fails and how much risk of an undetected failure the site will accept, not from convenience.
- Record run to failure as a decision with its reason, so it is not mistaken for an oversight at the next review.
- Keep the link. Each task in the CMMS should reference the failure mode it manages, so the strategy can be tested against the failures that actually occur.
A task with no failure mode behind it is a habit. A failure mode with no decision is a gap.
Common FMECA mistakes
- Analysing every asset to the same depth, so the program never finishes. Criticality should decide where FMECA goes.
- Going too deep, into every bolt and gasket, which multiplies rows without changing a single decision.
- Writing symptoms or functional failures as failure modes, so there is nothing specific to act on.
- Copying a generic failure mode library without testing it against the operating context and the site's own history.
- Scoring without anchors, or letting the loudest voice in the room set the scores.
- Leaving operators and maintainers out of the sessions, and with them most of the knowledge of how the equipment really fails.
- Stopping at the ranking. A prioritised list that never becomes tasks, spares and design changes in the CMMS has cost the effort and delivered nothing.
- Never revisiting it. An FMECA should be reviewed after significant failures, modifications and changes in duty.
Using the FMECA worksheet template
The FMECA worksheet template is a free Excel workbook laid out in the order of the steps above. It opens in Excel, Google Sheets and LibreOffice.
- The Worksheet tab has a column for each step, drop-down lists for the scores, severity driver and task type, and formulas that calculate criticality, band and RPN before and after action.
- The Scales tab holds the example severity, occurrence and detection scales, the band rules and the task selection guide. Replace the anchors with your own risk framework before relying on the ranking.
- Bands follow the same rule as the Criticality Matrix Calculator, high from 12 and medium from 5, and always high where a severity of 5 is driven by safety or environment.
- The six example rows are the crusher lubrication system from this guide. Delete them, or keep them as a guide to how entries are worded.
Running an FMECA
Scope
System, boundary, operating context and level of analysis.
Functions and failures
What each item must do, and each way it can fail to.
Failure modes and effects
Specific causes, local and end effects, evident or hidden.
Score and rank
Severity, occurrence and detection on anchored scales.
Decide
A task, a design change or run to failure for each mode.
Load and review
Tasks in the CMMS, reviewed after failures and changes.
Scope
System, boundary, operating context and level of analysis.
Functions and failures
What each item must do, and each way it can fail to.
Failure modes and effects
Specific causes, local and end effects, evident or hidden.
Score and rank
Severity, occurrence and detection on anchored scales.
Decide
A task, a design change or run to failure for each mode.
Load and review
Tasks in the CMMS, reviewed after failures and changes.
Common questions
What is FMECA?
FMECA, or failure modes, effects and criticality analysis, is a structured method for working out how an item can fail, what each failure causes, and how critical each failure is. It is an FMEA with a criticality ranking added, and in maintenance it is the usual way to decide which failure modes need a task.
What is the difference between FMEA and FMECA?
An FMEA identifies failure modes and their effects. An FMECA adds criticality analysis, ranking each failure mode by severity and likelihood, and sometimes detection, so the most important ones are dealt with first. IEC 60812 covers both.
How is criticality calculated in an FMECA?
Most maintenance FMECAs multiply a severity score by an occurrence score and read the result off a risk matrix, often the site's own. Other methods include the risk priority number, which also multiplies in a detection score, and the quantitative criticality number in MIL-STD-1629A, which is built from failure rates.
What is a risk priority number?
The risk priority number, or RPN, is severity multiplied by occurrence multiplied by detection. It is simple and widely used, but different combinations give the same number and a low score can hide a severe consequence, which is why the AIAG and VDA FMEA Handbook replaced it with action priority tables in 2019.
What is a good RPN threshold?
There is no universal one. RPN values are rankings multiplied together, so a fixed cut-off treats very different risks the same way and invites teams to shade scores to land under it. Rank by severity and criticality first, and use RPN to sort within a band.
How does FMECA relate to RCM?
FMECA supplies the failure modes, effects and consequences that reliability centred maintenance decision logic works through. RCM then chooses a task for each failure mode, such as condition monitoring, scheduled replacement, failure finding or run to failure. SAE JA1011 sets the criteria a process must meet to be called RCM.
How detailed should failure modes be?
Detailed enough to point at a specific cause and a task that could manage it, and no deeper. For most maintenance work that means the maintainable item, such as a pump bearing or a filter element, rather than every bolt and gasket.
Who should take part in an FMECA?
A facilitator who knows the method, the operators and maintainers who work on the equipment, a reliability or maintenance engineer, and specialists such as electrical or instrumentation where the system needs them. Operators and maintainers know how the equipment actually fails, which no drawing shows.
Which standards cover FMECA?
IEC 60812:2018, adopted in Australia as AS/NZS IEC 60812:2020, is the general standard for FMEA and FMECA. MIL-STD-1629A, cancelled in 1998, still underpins many criticality methods. The AIAG and VDA FMEA Handbook covers design and process FMEA in manufacturing, and ISO 14224 provides standard failure mode codes for equipment.
Is there a free FMECA worksheet template?
Yes. The FMECA worksheet template on this page is a free Excel workbook with the columns, scales, drop-down lists and formulas set up, and the crusher lubrication example from this guide filled in. Adapt the scales to your own risk framework before relying on the ranking.
Key terms
Plain-language definitions from our glossary for the concepts this article leans on.
Standards and further reading
- IEC 60812:2018 Failure modes and effects analysis (FMEA and FMECA) (IEC)
- AS/NZS IEC 60812:2020 Failure modes and effects analysis (FMEA and FMECA) (Standards Australia)
- MIL-STD-1629A Procedures for performing a failure mode, effects and criticality analysis (US Department of Defense)
- AIAG and VDA FMEA Handbook, first edition (2019) (AIAG)
- SAE JA1011 Evaluation criteria for RCM processes (2024 revision) (SAE International)
- SAE JA1012 A guide to the reliability-centered maintenance (RCM) standard (SAE International)
- ISO 14224:2016 Collection and exchange of reliability and maintenance data for equipment (ISO)
Related case studies and tools
- Criticality Matrix Calculator (tool)
- Reliability engineering and asset strategy program (case study)
Related reading
Working through something like this?
See how we approach these initiatives, or tell us what you are dealing with.