Reliability centred maintenance, and when criticality-led beats it
Reliability centred maintenance is rigorous, defensible and expensive. What RCM analysis involves, how PMO actually compares, and when a criticality-led approach gets you most of the value for a fraction of the work.

Key takeaways
- RCM derives maintenance from functions and failure modes. Its rigour is real, and so is its cost, measured in person-days per system.
- PMO is faster because it starts from the existing task list, and blind for the same reason: it cannot find failure modes no task has ever covered.
- A criticality-led approach gives every asset a defensible strategy within months, reserving full analysis for the systems that carry the risk.
- Whichever method you choose, without consistent failure coding and a review loop the analysis decays back into an inherited task list.
Reliability centred maintenance is a structured method for deciding what maintenance an asset actually needs, based on how it can fail and what happens when it does. It came out of civil aviation in the 1970s, where the question was why adding more overhauls was not making aircraft safer.
The answer was that most failure modes are not age related, so more frequent intrusive maintenance introduced about as much risk as it removed. That finding still holds on a plant or a mine site, and it is why the method survives. What gets less discussion is what it costs to do properly, and whether every asset deserves it.
What RCM analysis involves
RCM works through seven questions for each asset in its operating context. The SAE JA1011 criteria are a useful benchmark for whether a process genuinely qualifies as RCM rather than borrowing the label. Questions two to five are usually answered with an FMECA, which records the failure modes, their effects and their consequences.
- What are the functions and performance standards of the asset as it is actually used here?
- In what ways can it fail to deliver those functions?
- What causes each functional failure?
- What happens when each failure occurs?
- In what way does each failure matter, in safety, environmental, operational or cost terms?
- What proactive task will predict or prevent it, and at what interval?
- What happens if no suitable proactive task exists?
Consequence decides the strategy
The fifth question does most of the work. RCM separates failures by consequence: hidden failures that give no evidence until something else fails, safety and environmental consequences, operational consequences that cost production, and non-operational consequences that cost only the repair.
That separation changes what a task is worth. A hidden failure in a protective system needs a failure-finding task at an interval set by the risk you are prepared to carry, not by how quickly the item wears out. An operational consequence justifies spending up to what the downtime would cost. A non-operational one often justifies nothing at all.
Choosing the task, not just the interval
Where a condition-based task is technically feasible, the P-F interval governs inspection frequency: the time between a detectable potential failure and functional failure. Inspect well inside that window and you get useful warning. Inspect outside it and you are collecting data after the point it could have changed anything.
Where no condition-based task fits, the logic steps through scheduled restoration, scheduled discard, failure-finding and, honestly, run to failure. Run to failure is a legitimate answer for a non-critical asset, and a strategy that never selects it is not being applied properly.
What it costs
A rigorous analysis takes a facilitated group through every functional failure of a system, and that group has to include the people who operate and maintain the equipment. Budget person-days per system, not hours. On a large plant, analysing everything becomes a program that runs for years and outlives the interest of whoever sponsored it.
That is the real constraint, and it is why so many RCM programs stall halfway. The method is sound. The scope was never affordable.
Most stalled RCM programs did not fail on method. They failed on scope: everything was critical, so nothing got finished.
PMO is not cheap RCM
Planned maintenance optimisation is the alternative clients ask about most, usually framed as RCM results at a fraction of the cost. The comparison is worth making carefully, because the two methods do not start from the same place and do not find the same things.
RCM starts from the asset's functions and derives what maintenance is deserved, whether or not any of it exists today. PMO starts from the maintenance program you already have: gather the existing tasks, ask what failure mode each one manages, then challenge every task on that basis. Duplicates surface, tasks with no credible failure mode behind them get removed, invasive work gets converted to condition-based where the failure gives warning, and intervals get reset against evidence rather than habit.
That inversion is why PMO is faster, commonly by severalfold per system, and on a mature plant it removes or improves a large share of the task list. It is also the method's blind spot. PMO can only interrogate what the current program and the failure history put in front of it. A failure mode nobody has experienced and no task addresses is invisible to it, and those missing tasks are precisely what a functions-first analysis exists to find. On equipment where an unmanaged failure mode carries safety or production consequence, that gap is the whole argument.
- Reach for PMO on a mature plant with an established program, where the pain is bloat: too many tasks, too invasive, intervals nobody can defend.
- Reach for RCM where coverage is the question: new or heavily modified equipment, no trustworthy history, or consequence severe enough that a missed failure mode is not survivable.
- The honest answer is often staged: PMO to strip the waste out of the long tail quickly, full analysis reserved for the systems that carry the risk.
- Either way the output only holds if failure coding and the review loop hold, which is the same discipline problem RCM has.
Where an asset criticality assessment fits
A criticality-led approach inverts the effort differently. Rank equipment by consequence of failure and likelihood, band the results, apply a standard strategy template to each band, and reserve full RCM analysis for the top one.
It is less rigorous, and it will miss failure modes a facilitated session would have caught. What it buys is coverage: every asset gets a defensible strategy within months, rather than a handful of assets getting an excellent one over years.
- Consistent equipment classes make the banding comparable, whether the taxonomy is ISO 14224, an internal standard or a blend of both.
- Criticality has to reflect operating context, so the same pump model can sit in different bands in different services.
- Every band needs a documented strategy template, otherwise the assessment produces a ranked list and nothing changes.
- Reassess after material changes to duty, throughput or configuration.
Choosing between them
| RCM | PMO | Criticality-led | |
|---|---|---|---|
| Starts from | Functions and failure modes | The existing task list | Consequence ranking |
| Best at | Coverage of critical assets | Stripping bloat quickly | Coverage of the whole base |
| Blind spot | Cost and duration | Failure modes never yet covered | Depth on complex assets |
| Effort | Person-days per system | Severalfold faster than RCM | Weeks to rank, then templates |
| Use when | Consequence is severe or the asset is new | A mature program has grown bloated | The long tail needs a defensible strategy |
- Use full RCM where failure carries safety or environmental consequence, where the asset is genuinely critical to production, or where an existing strategy is demonstrably not working and you need to know why.
- Use PMO where a mature program exists and the problem is bloat rather than gaps, accepting that it will not surface failure modes the current program never covered.
- Use a criticality-led approach across the long tail, where the value is having any defensible strategy rather than the optimal one.
- Use RCM where a regulator, insurer or internal standard wants the traceability, since it produces a documented line from failure mode to task.
- Do none of them on an asset you are about to replace.
Making either one hold
Both methods fail the same way. The analysis gets done, the tasks get loaded into the CMMS, and nothing reviews them again. Failure coding drifts, the strategy stops matching the equipment, and within a few years the task list is back to being inherited rather than designed.
The review loop is the part worth protecting: failure history coded consistently enough to test whether a strategy is working, a trigger to revisit when duty changes, and a named owner for the decision. Without that, the analysis is a document rather than a strategy.
How an asset strategy review runs
Understand
Work through the current strategy and where it hurts with the people who operate and maintain the equipment.
Assess
Benchmark current maturity and failure data quality against your own standards and recognised industry practice.
Criticality
Rank systems by consequence and likelihood in their real operating context, then band the results.
Analyse
Apply RCM decision logic to the critical band and strategy templates across the rest.
Embed
Load tasks with traceability back to the failure mode, then set the review trigger and the measures that prove it is working.
Understand
Work through the current strategy and where it hurts with the people who operate and maintain the equipment.
Assess
Benchmark current maturity and failure data quality against your own standards and recognised industry practice.
Criticality
Rank systems by consequence and likelihood in their real operating context, then band the results.
Analyse
Apply RCM decision logic to the critical band and strategy templates across the rest.
Embed
Load tasks with traceability back to the failure mode, then set the review trigger and the measures that prove it is working.
Common questions
Is PMO cheaper than RCM?
Per system, considerably, because it interrogates an existing task list rather than deriving strategy from functions. The saving is real; the trade is that PMO cannot surface failure modes the current program has never covered.
Do we need RCM to satisfy auditors or insurers?
Where traceability from failure mode to task is the requirement, RCM produces it natively. The SAE JA1011 criteria are the usual benchmark for whether a process qualifies as RCM rather than borrowing the label.
Can the approaches be combined?
That is usually the honest answer: a criticality ranking first, full RCM on the top band, PMO or strategy templates across the rest, and one review loop over all of it.
Key terms
Plain-language definitions from our glossary for the concepts this article leans on.
Standards and further reading
- JA1011: Evaluation Criteria for RCM Processes (SAE International)
- ISO 14224:2016, Collection and exchange of reliability and maintenance data for equipment (ISO)
- ISO 55001:2024, Asset management systems (ISO)
- Best Practice Metrics Guidelines (SMRP)
- Asset Management Landscape v3 (GFMAM)
- Operations and Maintenance Best Practices Guide (US DOE FEMP)
Related case studies and tools
- Reliability engineering and asset strategy program (case study)
- MTBF / MTTR Calculator (tool)
- Cost of Downtime Calculator (tool)
- Criticality Matrix Calculator (tool)
Related reading
Working through something like this?
See how we approach these initiatives, or tell us what you are dealing with.