Predictive maintenance and condition monitoring
Breakdown work costs more, takes longer and carries more risk than planned work. Why condition monitoring programs pay for themselves, and how we build them.

faults found early become planned work
planned repairs cost less than breakdowns
fewer secondary failures and safety exposures
The challenge
Critical equipment is commonly maintained to fixed intervals set years earlier, with condition data collected on separate routes, held in vendor portals and reviewed in monthly reports. Bearing and gearbox defects develop between those inspections, and on a connected circuit a single unplanned stoppage cascades through the value chain at real cost to production.
The gap is rarely sensing, it is the distance between detection and a planned work order. That distance has to fit inside the P-F interval, the window between a failure becoming detectable and becoming functional. Maintenance and reliability personnel require monitoring they can trust, and a path from alert to scheduled work that holds up under pressure.
When it's time to act
The signs we commonly see when this initiative is due.
- Breakdowns on equipment that was inspected recently, because the failure developed between routes.
- Condition data spread across vendor portals that only the contractor ever opens.
- Alerts arriving as emails that quietly expire without becoming work orders.
- Planned work repeatedly displaced by urgent repairs on the same handful of assets.
- Monthly condition reports describing last month's plant rather than today's.
How we deliver it
Criticality assessment
Rank assets by production and safety consequence to target monitoring where failure hurts most.
Monitoring strategy
Select techniques per dominant failure mode, drawing on RCM logic and program guidance such as ISO 17359.
Platform build
Consolidate vibration, thermography, oil and process data into one time-series platform.
Baselines and alerts
Establish per-asset baselines and staged alert limits tuned to the P-F interval of each failure mode.
Close the loop
Route confirmed alerts into the CMMS as planned work, with feedback tightening the model.
Criticality assessment
Rank assets by production and safety consequence to target monitoring where failure hurts most.
Monitoring strategy
Select techniques per dominant failure mode, drawing on RCM logic and program guidance such as ISO 17359.
Platform build
Consolidate vibration, thermography, oil and process data into one time-series platform.
Baselines and alerts
Establish per-asset baselines and staged alert limits tuned to the P-F interval of each failure mode.
Close the loop
Route confirmed alerts into the CMMS as planned work, with feedback tightening the model.
Our approach
- Understand the failures that actually hurt, with the maintainers and reliability engineers who see them, and assess how the current monitoring program performs against your own standards and relevant published guidance. That benchmark turns a list of opinions into a measured starting point.
- Rank the asset base by consequence of failure, so monitoring investment follows criticality rather than convenience.
- Match each critical failure mode to a detection technique and interval, guided by your own program standards or published guidance such as ISO 17359, whichever suits the operation.
- Consolidate vendor systems into one platform using an open, layered data architecture, from acquisition through state detection to health assessment. The layered model in ISO 13374 is a useful reference point here.
- Set staged alarm levels so an alert always lands early enough in the P-F interval to plan, kit and schedule the corrective work.
- Integrate alerts with the CMMS so a confirmed diagnosis becomes a prioritised, planned work order rather than an email.
Tools and methods
The value it creates
- Failures are found while they are still cheap: repairs are planned into existing windows with parts on hand, instead of breaking the schedule and risking secondary damage.
- The scale of the prize is well documented: the US DOE FEMP O&M best practices guide reports downtime reductions of 35 to 45 per cent from properly functioning predictive maintenance programs.
- Maintenance shifts from reacting to managing, with one health view per critical asset and a measured, auditable path from alert to work order.
What changes
Fixed intervals set years ago
Techniques and intervals matched to each failure mode
Condition data siloed in vendor portals
One health view per critical asset
Alerts landing in inboxes
Confirmed diagnoses routed into the CMMS as planned work
Failures discovered at failure
Faults found early enough to plan, kit and schedule
Where these initiatives fail
The failure modes we design against.
- Instrumenting what is easy to sense rather than what is critical, so the program monitors the wrong assets well.
- Alarm limits set generically, so false positives train people to ignore the channel that eventually matters.
- No route from alert to work order, which leaves good diagnoses stranded as emails.
- Declaring victory at installation: without a review loop, limits drift and credibility erodes.
Common questions
Do we need new sensors to start a condition monitoring program?
Usually not at the start. Most operations already collect vibration, oil and process data; the early value is consolidating what exists and closing the gap between detection and planned work. New sensing is added where criticality justifies it.
Which assets should be monitored first?
The ones where failure hurts most. A criticality ranking across production and safety consequence decides where monitoring investment lands, and it is usually a small fraction of the asset base.
How is the success of a program measured?
By conversion, lead time and outcomes: the share of alerts that become planned work orders, the warning time achieved against the P-F interval, and the trend in unplanned failures on monitored assets.
Key terms
Plain-language definitions from our glossary for the concepts this page leans on.
Standards and further reading
Reference points we draw on where they suit the work. We also work to client internal standards and established site practice.
- ISO 17359:2018 Condition monitoring and diagnostics of machines, general guidelines (ISO)
- ISO 13374-1 Condition monitoring data processing, communication and presentation (ISO)
- SAE JA1011 Evaluation criteria for RCM processes (SAE International)
- Best practices, metrics and guidelines for maintenance and reliability (SMRP)
- Operations and maintenance best practices guide (US DOE Federal Energy Management Program)
- Guidelines library for mining technology and interoperability (Global Mining Guidelines Group)
Further reading
Articles and calculators on the methods behind this work.
- Condition monitoring: techniques, intervals and building a program that works (article)
- Running an asset criticality assessment that holds up (article)
- Condition monitoring people actually act on (article)
- OEE Calculator (tool)
- MTBF / MTTR Calculator (tool)
- Cost of Downtime Calculator (tool)
- Mining solutions (solution)
Related projects
Facing a similar challenge?
Tell us what you are working through and we will bring the right mix of engineering, data and hands-on experience.
Contact us