Reliability engineering and asset strategy program
Maintenance strategies accumulate rather than get designed. Why asset strategy programs matter, and how FMECA and reliability centred maintenance logic decide what each asset actually deserves.

work moves to the failure modes that matter
stop introducing the failures you prevent
every task traces to a failure mode
The challenge
Maintenance strategies commonly accumulate rather than get designed. Tasks are added after incidents, carried across from vendor manuals and inherited from sites that ran different equipment, and they are rarely removed. Many are invasive enough to introduce as much failure risk as they prevent, and the link between a task and the failure mode it manages is often long lost.
Failure history usually exists but is coded inconsistently, so the data cannot settle which strategies are working and the argument falls back to opinion. Reliability and maintenance personnel require a defensible, repeatable basis for deciding what each asset deserves, and a review loop that keeps those decisions current as the operating context changes.
When it's time to act
The signs we commonly see when this initiative is due.
- The PM program has only ever grown, and nobody can say what half the tasks prevent.
- Invasive overhauls running on equipment whose failures are not age related.
- Failure history exists but is coded too inconsistently to answer anything.
- The same failures recur on critical systems despite full PM compliance.
- Strategy debates settled by the loudest opinion in the room.
How we deliver it
Criticality and data
Rank systems by consequence and fix the failure coding needed for analysis.
FMECA
Identify failure modes, effects and criticality on the systems that carry the risk.
Strategy build
Select tasks with structured reliability centred maintenance decision logic: condition-based where justified, run-to-failure where honest.
Deploy
Load strategies to the CMMS as executable plans with intervals, trades and materials.
Living program
Review triggers and KPIs keep strategies aligned to operating context, not history.
Criticality and data
Rank systems by consequence and fix the failure coding needed for analysis.
FMECA
Identify failure modes, effects and criticality on the systems that carry the risk.
Strategy build
Select tasks with structured reliability centred maintenance decision logic: condition-based where justified, run-to-failure where honest.
Deploy
Load strategies to the CMMS as executable plans with intervals, trades and materials.
Living program
Review triggers and KPIs keep strategies aligned to operating context, not history.
Our approach
- Understand the failures driving cost, downtime and risk, working through them with the people who repair them rather than from the register alone.
- Assess current strategy maturity against your own asset management framework and, where relevant, published practice, so the program starts from an agreed picture of where the gaps are.
- Rank systems by safety and production consequence so analysis effort lands where failure actually hurts.
- Sort out failure coding first, because strategy decisions are only as good as the event data behind them. That means a consistent taxonomy, whether the client's own or a published one such as ISO 14224.
- Run FMECA on critical systems and apply structured RCM decision logic: every retained task answers a specific failure mode, and every failure mode gets a deliberate decision, including run-to-failure where that is the honest answer. The SAE JA1011 criteria are a useful benchmark for whether a process genuinely qualifies as RCM.
- Use age-to-failure analysis where the data supports it, so intervals are set on evidence rather than habit.
- Deploy strategies to the CMMS as complete, executable plans and stand up a governance loop with defined review triggers: incidents, context changes and time.
Tools and methods
The value it creates
- The PM program gets rationalised: duplicates removed, invasive low-value tasks eliminated, and real gaps on genuine failure modes closed.
- Maintenance effort moves to the failure modes that drive downtime and risk, which is where availability improvements come from.
- Strategy decisions become documented and auditable, and the review loop keeps them current instead of frozen at whenever the last incident happened.
What changes
Tasks inherited from manuals and incidents
Every task traced to a failure mode
Invasive PM introducing risk
Condition-based where justified, run-to-failure where honest
Intervals set by habit
Intervals set on evidence where the data supports it
A strategy frozen at the last incident
A living program with defined review triggers
Where these initiatives fail
The failure modes we design against.
- Analysing everything: full RCM across a whole plant becomes a program that outlives its sponsor.
- Skipping failure coding, so the analysis stands on data that cannot support it.
- Producing strategies that never reach the CMMS as executable, kitted plans.
- No review loop, which returns the program to inherited tasks within a few years.
Common questions
Is this full RCM on every asset?
No. Analysis effort follows criticality: full FMECA and RCM decision logic on the systems that carry the risk, and templated strategies across the long tail. Analysing everything is how these programs stall.
What if our failure data is poor?
That is common, and it is why coding gets fixed early. Analysis can start from the knowledge in the room while the data improves, and intervals are then refined on evidence as it accumulates.
How do we know the program is working?
Recurring failures on critical systems, the balance of planned to unplanned work, and the share of tasks traceable to a failure mode all move. The review loop keeps measuring them after we leave.
Key terms
Plain-language definitions from our glossary for the concepts this page leans on.
Standards and further reading
Reference points we draw on where they suit the work. We also work to client internal standards and established site practice.
- SAE JA1011 Evaluation criteria for RCM processes (SAE International)
- ISO 14224:2016 Collection and exchange of reliability and maintenance data for equipment (ISO)
- ISO 55001:2024 Asset management system requirements (ISO)
- The Asset Management Landscape, third edition (GFMAM)
- Best practices, metrics and guidelines for maintenance and reliability (SMRP)
- Asset Management Council, Australian asset management community (Asset Management Council)
Further reading
Articles and calculators on the methods behind this work.
- MTBF / MTTR Calculator (tool)
- Cost of Downtime Calculator (tool)
- The asset management plan, and the strategy it comes from (article)
- FMECA: how to run a failure modes, effects and criticality analysis (article)
- Condition monitoring: techniques, intervals and building a program that works (article)
- Running an asset criticality assessment that holds up (article)
- Reliability centred maintenance, and when criticality-led beats it (article)
- Operational readiness for new assets: methods, standards and the CMMS build (article)
- Criticality Matrix Calculator (tool)
- Mining solutions (solution)
- Oil and gas solutions (solution)
Related projects
Facing a similar challenge?
Tell us what you are working through and we will bring the right mix of engineering, data and hands-on experience.
Contact us