September 2026 · 17 min read

FMECA: how to run a failure modes, effects and criticality analysis

FMECA works through how equipment can fail, what each failure does and which failures matter most, so maintenance effort goes where the risk is. The method step by step, how to score criticality, where the risk priority number falls short, and a free worksheet template to start from.

Reliability EngineeringMaintenance StrategyAsset Management
Hydraulic hoses, fittings and cylinders on a drilling rig against a clear blue sky

Key takeaways

  • FMECA works through each way an asset can fail, what the failure does and how much it matters, so maintenance effort is spent where the risk is.
  • FMEA stops at failure modes and their effects. FMECA adds a criticality ranking, most often severity against occurrence on a risk matrix.
  • Write failure modes as the specific event that causes the failure, at the level a maintenance task can act on. "Pump fails" is a symptom, not a failure mode.
  • The risk priority number multiplies severity, occurrence and detection, and it can rank a nuisance level with a hazard. Use it to sort within a band, not to set the band.
  • An FMECA pays off only when every ranked failure mode ends in a task, a design change or a recorded decision to run to failure, loaded into the CMMS.

Failure modes, effects and criticality analysis, usually shortened to FMECA, is a structured way of working out how equipment can fail, what each failure does, and which failures matter most. It is the analytical core of most maintenance strategy work. Done well, it gives every maintenance task a reason, in the form of the failure mode it manages, and gives every significant failure mode a deliberate decision.

The method is old and well documented. It came out of United States military reliability work in the late 1940s, and IEC 60812 now sets out how to plan, run and document it. What goes wrong in practice is rarely the method itself. It is scope that is too wide, failure modes that are really symptoms, scores nobody calibrated, and results that never reach the CMMS.

This guide covers what FMECA is and where it fits, the standards, how to run one step by step, the ways to rank criticality and where the risk priority number falls short, how the results become maintenance tasks, and the mistakes that turn an analysis into shelfware. A free FMECA worksheet template comes with it, set up with the worked example used throughout.

Free download

FMECA worksheet template

The worksheet, scales and task selection guide in one workbook, with formulas for criticality, band and RPN, and the worked example from this guide filled in.

Download (Excel, 25 KB)

The difference between FMEA and FMECA

A failure modes and effects analysis (FMEA) lists the ways an item can fail and the effects of each failure. A failure modes, effects and criticality analysis (FMECA) goes one step further and ranks each failure mode by how critical it is, so the analysis ends in a priority order rather than a long list.

IEC 60812 treats FMECA as a variant of FMEA, in which the ranking takes in at least the severity of the consequences and often other measures, such as how likely the failure is. In practice the names are used loosely, and many documents called an FMEA include a criticality ranking. What matters is whether the ranking drives decisions.

The same method is used in three quite different settings, and it helps to know which one you are in.

TypeWhat it analysesTypical use
Design FMEAHow a product or system design could fail to meet its requirementsEngineering design, before the item is built
Process FMEAHow a manufacturing or assembly process could produce a defectManufacturing quality, common in automotive supply chains
Equipment or maintenance FMECAHow installed equipment fails in its operating context, and what each failure does to the operationMaintenance strategy and RCM for plant, mobile fleets and infrastructure

This guide is about the third kind, where the question is not how to improve a design but how failures are prevented, detected or tolerated once the equipment is running.

Where FMECA fits in maintenance strategy

FMECA sits between the criticality assessment and the maintenance plan. An asset criticality assessment ranks whole assets or systems by the consequence of their failure. FMECA then goes inside the systems that rank highest and works out how each one fails. Reliability centred maintenance decision logic takes those failure modes and chooses a task for each.

The order matters because FMECA is expensive. A thorough analysis takes a facilitated team through every function and failure of a system, and it is rarely worth doing on equipment whose failure costs little. Ranking criticality first, with a tool such as the Criticality Matrix Calculator, keeps the deep analysis on the systems that carry the risk.

FMECA is also used on its own, outside a full RCM program. Typical uses are reviewing the strategy for a problem asset, preparing the maintenance plan for new equipment during operational readiness, deciding which spares to hold, and understanding what a proposed modification would do to reliability.

The standards and references

No single standard is mandatory for a maintenance FMECA. The references below are the ones most often used, and many sites pair one of them with their own risk framework.

ReferenceWhat it coversUseful for
IEC 60812:2018The general method for FMEA and FMECA, from planning to reporting, including criticality methods such as the risk priority number and the criticality matrix. Adopted in Australia as AS/NZS IEC 60812:2020.The default reference for any FMECA
MIL-STD-1629AThe original military procedure, with a quantitative criticality number built from failure rates. Cancelled in 1998 but still widely cited.Quantitative criticality where failure rate data exists
AIAG and VDA FMEA Handbook (2019)A seven step method for design and process FMEA, which replaced the risk priority number with action priority tables.Manufacturing quality and automotive supply
SAE JA1011 and JA1012The criteria a process must meet to be called RCM, and a guide to applying them.FMECA as part of a full RCM analysis
ISO 14224:2016Standard taxonomies for equipment, failure modes, failure mechanisms and causes, with failure mode codes by equipment class.Consistent failure mode wording and failure data

Before the first session

Most of the quality of an FMECA is set before the team meets. Decide these things first and write them down.

  • The system and its boundary. Where it starts and stops, and whether interfaces such as the power supply and the control system are in or out of scope.
  • The operating context. Duty, throughput, operating hours, redundancy, environment and constraints. The same pump can deserve a different strategy in a different service.
  • The level of analysis. Record failure modes at the maintainable item, such as a pump, a gearbox or a filter, rather than at every part.
  • The scales. Severity, occurrence and, if used, detection, with anchors the site already recognises. Borrowing the site risk matrix avoids a second, competing definition of risk.
  • The information. Drawings and P&IDs, manufacturer manuals, the current maintenance plan, and failure history from the CMMS, ideally coded consistently enough to count failures by mode.
  • The team. A facilitator who knows the method, the operators and maintainers who know how the equipment behaves, and an engineer who can make decisions stick.

An hour spent agreeing the boundary and the scales saves days of arguing about them in the sessions.

How to run an FMECA, step by step

The steps below follow the structure of IEC 60812 and the questions RCM asks of each asset. The worksheet template has a column for each, and the examples come from its first row.

StepWhat to recordExample
1. FunctionsWhat the item must do, with a performance standard where one appliesDeliver filtered oil to the crusher bearings at the required pressure
2. Functional failuresEach way it can fail to do that, fully or partlyFails to deliver any oil flow
3. Failure modesThe specific event that causes each functional failurePump bearing seizes
4. CausesWhy the failure mode happens, the mechanism behind itBearing fatigue at end of life
5. EffectsWhat happens locally and at the level of the operation, and whether the failure is evident or hiddenLow oil pressure trip stops the crusher. With no spare pump the plant is down until the pump is replaced
6. Current controlsWhat already prevents or detects the failure modeLow oil pressure trip. Operator round each shift
7. ScoresSeverity, occurrence and detection on the agreed scales, with the reasoningSeverity 4, occurrence 3, detection 3
8. CriticalityThe ranking, from severity and occurrence on the risk matrix12 of 25, high
9. ActionsA task, design change or recorded decision, with an ownerMonthly vibration route on the pump and motor bearings. Hold a spare pump on site

Writing failure modes that are useful

The failure mode is the most important entry in the worksheet and the one most often written badly. A failure mode is the event that causes a functional failure. It should name the part that fails and how it fails, specifically enough to point at a cause and a task.

"Pump fails" is not a failure mode. It is a functional failure at best, and it gives the team nothing to act on. "Pump bearing seizes from fatigue", "mechanical seal leaks from face wear" and "impeller erodes in abrasive slurry" each suggest a different detection method and a different task.

  • Record failure modes at the level where a task can be applied, usually the maintainable item. Going deeper multiplies the rows without changing the decisions.
  • Include failure modes that have happened, those that happen on similar equipment elsewhere, and those that have not happened yet but reasonably could, especially where the consequence would be severe.
  • Include operating and human causes, such as an incorrect start-up or overloading, where they are credible. Not every failure starts in the equipment.
  • Use consistent wording. The failure mode codes in ISO 14224, such as FTS for fail to start on demand and ELP for external leakage of process medium, keep the analysis comparable with the failure data the CMMS collects.

If a failure mode does not suggest what you would do about it, it is not specific enough yet.

Describing effects, and spotting hidden failures

Effects describe what happens when the failure mode occurs, assuming nothing is done to stop it. Record the local effect on the item and the end effect on the system or the operation, in enough detail to judge severity. Say what the operators would see, how long the loss lasts and what else is damaged.

Note whether each failure is evident or hidden. An evident failure makes itself known to the operating crew in normal operation. A hidden failure does not, usually because the item is a protective device that only acts on demand, such as a trip switch, a relief valve or a standby pump. Its consequence is felt only when something else fails as well, which is why it needs its own treatment.

The low oil pressure switch in the worked example shows how this plays out. While oil pressure is normal, a stuck switch changes nothing. If the pump then fails, the crusher runs on without lubrication and the main bearings are destroyed. No monitoring of the switch in service would reveal the fault. It has to be tested.

Scoring severity, occurrence and detection

Each failure mode is scored on three scales. Severity rates the worst credible end effect. Occurrence rates how often the failure mode is expected to happen. Detection rates how likely the current controls are to catch it before the effect occurs.

Anchor every level with a description the team can test a score against, and keep the scales short. Five levels are easier to apply consistently than ten, because people can actually tell neighbouring levels apart. The worksheet uses the same five level severity and occurrence scales as the Criticality Matrix Calculator, so failure modes rank on the same basis as the assets they belong to, and adds the detection scale below.

Occurrence is where failure history earns its keep. Where the CMMS records failures against the right equipment with consistent codes, occurrence can be counted rather than guessed, and the MTBF / MTTR Calculator turns counts into rates. Where history is thin, use experience on similar equipment, manufacturer data or published reliability data, and note that the score is an estimate.

Detection scoreMeaning
1. Almost certainContinuous monitoring with an alarm, or obvious to operators at once
2. HighFound by current inspections or condition monitoring with time to act
3. ModerateDetectable, but only if a check falls inside the warning period
4. LowLittle warning, and current checks are unlikely to catch it
5. NoneHidden until it fails on demand or causes the effect

Score against the anchors, not against the row above. A team that compares rows drifts toward rating everything a three.

Ways to rank criticality

The criticality in FMECA can be worked out in several ways. IEC 60812 describes more than one, and the right choice depends on the data available and what the ranking is for.

MethodHow it worksStrengthsWeaknesses
Criticality matrixSeverity against occurrence on a risk matrix, often the site's own, giving a score and a bandSimple, visual and consistent with how the site already describes riskIgnores detection unless it is recorded separately
Risk priority number (RPN)Severity × occurrence × detection, each on the same scale, giving one numberQuick to calculate and sort, and takes detection into accountEqual numbers can hide very different risks, and a low score can mask a severe consequence
Action priorityA lookup table in the AIAG and VDA handbook assigns high, medium or low priority to every combination of scores, weighted toward severityA severe consequence is never ranked lowWritten for design and process FMEA, and the tables sit in the handbook
Quantitative criticality numberCriticality from the failure rate, the share of failures in each mode, the probability of the effect and the operating time, as in MIL-STD-1629AObjective where good failure rate data existsNeeds failure mode data that most maintenance teams do not have
  • Most maintenance FMECAs use a criticality matrix, because it speaks the same risk language as the rest of the site and the band it produces maps directly to a strategy.
  • Recording detection alongside is still worthwhile. It shows where the current controls are weak, which is often where the quickest improvement lies.

Why the risk priority number misleads

The risk priority number is the most familiar criticality measure and the most criticised. The scores it multiplies are rankings, not measurements, so the product behaves in ways that surprise people.

  • Different risks score the same. On five point scales, a failure mode with severity 5, occurrence 1 and detection 2 scores 10, exactly the same as a nuisance with severity 1, occurrence 5 and detection 2.
  • Detection can bury severity. A severe failure mode with good detection can rank below a trivial one with poor detection, even though a detection control can fail too.
  • The numbers are lumpy. Only some values between 1 and 125 can occur and many combinations crowd the middle, so a one point change in a single score can move a failure mode a long way up or down the list.
  • Thresholds invite gaming. A rule such as acting on anything over 50 encourages teams to shade scores to land just under it.

If you use RPN at all, use it to sort failure modes within a band, and let severity and criticality decide the band.

A worked example on a crusher lubrication system

The worksheet template comes with this example filled in. It is the lubrication system for a primary crusher with a single, unspared lube oil pump, the same asset scored in the Criticality Matrix Calculator example. Six failure modes are shown, one for each item, to illustrate the range of outcomes rather than to be complete.

Ranked by criticality, the pump bearing comes first and earns monitoring and a spare. Ranked by RPN, the low oil pressure switch comes first, because its failure is hidden. Both views are telling you something. The switch does not need a higher band. It needs a test, because no other kind of task can find a hidden failure.

Failure modeSODCriticalityBandRPNTask type
Lube oil pump, pump bearing seizes43312High36Condition-based task
Oil supply filter, filter element blocks3339Medium27Condition-based task
Low oil pressure switch, switch contacts stick closed52510Medium50Failure-finding task
Oil cooler, cooler tube leaks3226Medium12Condition-based task
Return line hose, hose fitting leaks2316Medium6Condition-based task
Local pressure gauge, gauge mechanism fails1222Low4Run to failure
  • The pump bearing ranks highest on criticality. A monthly vibration route gives warning of a developing fault, and a spare pump on site shortens the outage if the warning is missed.
  • The pressure switch ranks highest on RPN because its failure is hidden. A six-monthly function test against a calibrated source is the only task that can find it.
  • The filter and the cooler are managed on condition, through the differential pressure indicator and oil analysis, because both failures develop slowly and give warning.
  • The hose leak is caught on the operator round. The local gauge is run to failure, because the pressure transmitter and trip already provide the protection.

From failure modes to maintenance tasks

The analysis earns its cost at this step. Every failure mode above the lowest band needs a decision, and the decision should follow from how the failure mode behaves rather than from habit. RCM decision logic, as described in SAE JA1012, is the usual framework, and the table summarises it in the order the logic considers each option.

Task typeWhen it fits
Condition-based taskThe failure mode gives detectable warning, with a P-F interval long enough to plan and act on.
Scheduled restorationThe failure mode is age related, and reworking the item restores its resistance to failure.
Scheduled replacementThe failure mode is age related, and fitting a new item restores the original reliability.
Failure-finding taskThe failure is hidden, typically in a protective device that only acts on demand, so it has to be tested.
Run to failureThe consequence is acceptable and no task is worth its cost. Record it as a decision.
Redesign or changeNo task brings the risk to an acceptable level. Where the consequence is safety or environmental, a change is compulsory.
  • Write each task so it can be loaded into the CMMS as it stands, with what to check, the acceptance limit, the interval and the trade.
  • Set condition-based intervals from the P-F interval, and choose techniques as the condition monitoring guide describes.
  • Set failure-finding intervals from how often the protective device fails and how much risk of an undetected failure the site will accept, not from convenience.
  • Record run to failure as a decision with its reason, so it is not mistaken for an oversight at the next review.
  • Keep the link. Each task in the CMMS should reference the failure mode it manages, so the strategy can be tested against the failures that actually occur.

A task with no failure mode behind it is a habit. A failure mode with no decision is a gap.

Common FMECA mistakes

  • Analysing every asset to the same depth, so the program never finishes. Criticality should decide where FMECA goes.
  • Going too deep, into every bolt and gasket, which multiplies rows without changing a single decision.
  • Writing symptoms or functional failures as failure modes, so there is nothing specific to act on.
  • Copying a generic failure mode library without testing it against the operating context and the site's own history.
  • Scoring without anchors, or letting the loudest voice in the room set the scores.
  • Leaving operators and maintainers out of the sessions, and with them most of the knowledge of how the equipment really fails.
  • Stopping at the ranking. A prioritised list that never becomes tasks, spares and design changes in the CMMS has cost the effort and delivered nothing.
  • Never revisiting it. An FMECA should be reviewed after significant failures, modifications and changes in duty.

Using the FMECA worksheet template

The FMECA worksheet template is a free Excel workbook laid out in the order of the steps above. It opens in Excel, Google Sheets and LibreOffice.

  • The Worksheet tab has a column for each step, drop-down lists for the scores, severity driver and task type, and formulas that calculate criticality, band and RPN before and after action.
  • The Scales tab holds the example severity, occurrence and detection scales, the band rules and the task selection guide. Replace the anchors with your own risk framework before relying on the ranking.
  • Bands follow the same rule as the Criticality Matrix Calculator, high from 12 and medium from 5, and always high where a severity of 5 is driven by safety or environment.
  • The six example rows are the crusher lubrication system from this guide. Delete them, or keep them as a guide to how entries are worded.

Running an FMECA

  1. Scope

    System, boundary, operating context and level of analysis.

  2. Functions and failures

    What each item must do, and each way it can fail to.

  3. Failure modes and effects

    Specific causes, local and end effects, evident or hidden.

  4. Score and rank

    Severity, occurrence and detection on anchored scales.

  5. Decide

    A task, a design change or run to failure for each mode.

  6. Load and review

    Tasks in the CMMS, reviewed after failures and changes.

Common questions

What is FMECA?

FMECA, or failure modes, effects and criticality analysis, is a structured method for working out how an item can fail, what each failure causes, and how critical each failure is. It is an FMEA with a criticality ranking added, and in maintenance it is the usual way to decide which failure modes need a task.

What is the difference between FMEA and FMECA?

An FMEA identifies failure modes and their effects. An FMECA adds criticality analysis, ranking each failure mode by severity and likelihood, and sometimes detection, so the most important ones are dealt with first. IEC 60812 covers both.

How is criticality calculated in an FMECA?

Most maintenance FMECAs multiply a severity score by an occurrence score and read the result off a risk matrix, often the site's own. Other methods include the risk priority number, which also multiplies in a detection score, and the quantitative criticality number in MIL-STD-1629A, which is built from failure rates.

What is a risk priority number?

The risk priority number, or RPN, is severity multiplied by occurrence multiplied by detection. It is simple and widely used, but different combinations give the same number and a low score can hide a severe consequence, which is why the AIAG and VDA FMEA Handbook replaced it with action priority tables in 2019.

What is a good RPN threshold?

There is no universal one. RPN values are rankings multiplied together, so a fixed cut-off treats very different risks the same way and invites teams to shade scores to land under it. Rank by severity and criticality first, and use RPN to sort within a band.

How does FMECA relate to RCM?

FMECA supplies the failure modes, effects and consequences that reliability centred maintenance decision logic works through. RCM then chooses a task for each failure mode, such as condition monitoring, scheduled replacement, failure finding or run to failure. SAE JA1011 sets the criteria a process must meet to be called RCM.

How detailed should failure modes be?

Detailed enough to point at a specific cause and a task that could manage it, and no deeper. For most maintenance work that means the maintainable item, such as a pump bearing or a filter element, rather than every bolt and gasket.

Who should take part in an FMECA?

A facilitator who knows the method, the operators and maintainers who work on the equipment, a reliability or maintenance engineer, and specialists such as electrical or instrumentation where the system needs them. Operators and maintainers know how the equipment actually fails, which no drawing shows.

Which standards cover FMECA?

IEC 60812:2018, adopted in Australia as AS/NZS IEC 60812:2020, is the general standard for FMEA and FMECA. MIL-STD-1629A, cancelled in 1998, still underpins many criticality methods. The AIAG and VDA FMEA Handbook covers design and process FMEA in manufacturing, and ISO 14224 provides standard failure mode codes for equipment.

Is there a free FMECA worksheet template?

Yes. The FMECA worksheet template on this page is a free Excel workbook with the columns, scales, drop-down lists and formulas set up, and the crusher lubrication example from this guide filled in. Adapt the scales to your own risk framework before relying on the ranking.

Key terms

Plain-language definitions from our glossary for the concepts this article leans on.

Standards and further reading

Related case studies and tools

Related reading

Working through something like this?

See how we approach these initiatives, or tell us what you are dealing with.