September 2026 · 13 min read

Condition monitoring: techniques, intervals and building a program that works

Condition monitoring finds failures while they are still developing, so the repair can be planned instead of rushed. The main techniques and what each one detects, how to choose between them, how often to check, and how to build a program that pays for itself.

Reliability EngineeringMaintenance StrategyAsset ManagementIoT & Telemetry
Engineer in high visibility gear inspecting a large machine rotor in a workshop

Key takeaways

  • Condition monitoring measures the health of equipment while it runs, so a developing failure is found early enough to plan the repair rather than react to a breakdown.
  • Each technique sees different failure modes. Vibration finds mechanical faults in rotating equipment, oil analysis finds wear and contamination, thermography finds heat, and ultrasound finds friction, leaks and electrical discharge.
  • Choose techniques failure mode by failure mode, starting from criticality. Monitoring only pays where a failure gives warning and the consequence justifies the cost.
  • The check interval comes from the P-F interval, the warning time between detectable and failed. A common rule of thumb is to check at no more than half of it.
  • Judge a program by the planned work it creates and the failures it prevents, not by the number of readings it collects.

Condition monitoring is the practice of measuring the health of equipment while it is running, and using the trend to decide when maintenance is needed. It works because most failures do not happen instantly. A bearing starts to spall, a gearbox picks up wear metal in its oil, an electrical joint starts to run hot, and each of those changes can be measured weeks or months before the equipment stops.

Finding the failure early changes what kind of work it becomes. A breakdown is unplanned, usually arrives at the worst time and often does collateral damage. The same fault found early is a planned job, with the parts on the shelf, the crew booked and the outage window chosen by you rather than by the machine.

This guide covers the main techniques and what each one can see, how to choose techniques and intervals asset by asset, online versus route-based monitoring, the standards worth knowing, and how to build a program that keeps paying. Once a program is running, the harder problem is acting on what it finds, which condition monitoring people actually act on covers in depth.

Condition monitoring, predictive maintenance and preventive maintenance

The three terms overlap and are often used interchangeably. Preventive maintenance services or replaces equipment on a fixed interval of time or use, whatever its condition. Condition monitoring measures condition. Predictive maintenance, sometimes called condition-based maintenance, is the strategy of scheduling work from what the monitoring shows.

Fixed intervals suit failure modes that wear out predictably with age or use. A large share of failures in complex equipment are not like that. They occur at random with respect to age, so replacing parts on a calendar spends money on healthy components and still misses the failure. That is the case for monitoring, and it is the reasoning at the centre of reliability centred maintenance.

ApproachWhat triggers the workBest suited to
Run to failureThe failure itselfLow-consequence items that cost less to replace than to monitor
Preventive (time or usage based)A calendar interval, operating hours or throughputFailure modes that wear out predictably with age or use
Condition based (predictive)A measured change in conditionFailure modes that give warning, where the consequence justifies monitoring

Monitoring does not stop failures. It moves them from the middle of a shift to a time you choose.

The main condition monitoring techniques

Each technique detects a particular physical symptom, so each one sees some failure modes and is blind to others. Most programs combine several, layered on top of operator rounds, which remain the cheapest and broadest check of all.

TechniqueWhat it detectsTypical equipment
Vibration analysisImbalance, misalignment, looseness, bearing and gear defects, resonanceMotors, pumps, fans, gearboxes, conveyor drives and pulleys
Oil analysisWear metals, contamination by dirt or water, and breakdown of the lubricant itselfGearboxes, hydraulic systems, engines, compressors
Infrared thermographyAbnormal heat from loose or overloaded electrical joints, failing bearings, blocked coolers and damaged refractorySwitchboards and motor control centres, conveyor idlers, kilns, heat exchangers
UltrasoundFriction from poor lubrication, early bearing faults, compressed air and gas leaks, electrical discharge, valves passingSlow-speed bearings, compressed air systems, steam traps, switchgear
Motor current and electrical testingRotor bar and winding faults, insulation degradation, supply problemsInduction motors and their supply
Thickness and wear measurementWall loss from corrosion and erosion, liner and wear plate lossChutes, bins, pipework, tanks, pressure vessels
Process and performance dataEfficiency loss from wear or fouling, seen in flow, pressure, power and temperaturePumps, compressors, fans, heat exchangers
Operator and inspector roundsLeaks, noise, smell, heat, visible damageEverything, as the first and cheapest layer

Vibration analysis

Vibration analysis is the backbone of most programs for rotating equipment, because so many mechanical faults change how a machine vibrates. Accelerometers on the bearing housings measure overall vibration levels and the frequency spectrum. Overall levels show that something has changed. The spectrum shows what, because each fault leaves a signature at a characteristic frequency, such as once per revolution for imbalance and the defect frequencies set by a bearing's geometry.

Severity is judged in two ways. Absolute limits, such as the evaluation zones in ISO 20816-3 for industrial machines above 15 kW, show whether a vibration level is acceptable for that class of machine. Change from the machine's own baseline often matters more, because a machine that has doubled its vibration is telling you something even while it sits inside the acceptable zone.

Slow-speed equipment is harder. At low running speeds a developing bearing defect releases little energy at conventional frequencies, so techniques such as high-frequency demodulation, shock pulse and ultrasound usually find it sooner. Grinding mill girth gears, slow conveyor pulleys and large crushers often need that specialist set-up.

The analyst matters as much as the instrument. Reading spectra is a skill built over years, and the analyst categories in ISO 18436-2 are a useful benchmark for hiring, training and judging a contractor.

Oil analysis

Oil analysis reads the condition of the machine and the condition of the lubricant from the same sample. Wear metals such as iron, copper and chromium show which components are wearing, and a rising trend points to a developing problem long before it can be heard or felt. Particle counts, water content and viscosity show whether contamination or degradation is doing the damage.

Contamination is often the cause rather than the symptom, which makes oil analysis a proactive tool as well as a diagnostic one. Fluid cleanliness is commonly specified with the ISO 4406 particle code, and setting and holding target cleanliness levels on hydraulic and gear systems is among the cheapest reliability improvements available.

Sampling decides whether the results mean anything. Take samples from the same live point, at the same operating condition, into clean bottles, and use a laboratory that trends results by asset rather than reporting each sample in isolation. ISO 14830-1 sets out general requirements for this kind of tribology-based monitoring.

Thermography and ultrasound

Infrared thermography turns heat into an image. It is the standard way to find loose or overloaded electrical connections, which run hot long before they fail, and it is fast enough to cover a whole switchroom or a long run of conveyor idlers in one route. Readings must be corrected for the surface's emissivity and for reflected heat, and electrical surveys are most useful when circuits are carrying a meaningful load. ISO 18434-1 covers the general procedures.

Ultrasound listens above the range of human hearing. Airborne instruments find compressed air and gas leaks, steam traps and valves passing, and the corona and tracking that come before electrical failure. Contact instruments pick up the friction of a bearing running short of grease, which makes ultrasound the natural tool for lubrication routes. The technician greases until the reading falls, rather than giving every bearing the same number of shots. ISO 29821 covers the method.

Choosing techniques failure mode by failure mode

The right first question is never which technology to buy. It is which failure modes matter, and which of those can be detected in time. Work through it in order.

  • Start with criticality. Monitoring effort belongs on the equipment whose failure hurts most, not spread evenly across the site.
  • List the dominant failure modes for each critical asset, drawing on failure history, similar equipment elsewhere and a structured analysis such as FMECA.
  • For each failure mode, ask whether it gives warning. A failure with no detectable warning, or a warning too short to act on, cannot be managed by monitoring and needs a design change, redundancy or another strategy.
  • Match a technique that can see that warning, at a cost the consequence justifies. A pump in a duty and standby pair may justify nothing more than an operator round. A single-line mill gearbox justifies online vibration and routine oil analysis.
  • Write the result down as a task with a technique, a measurement point, a frequency, alarm limits and the action expected when a limit is crossed.

Buying sensors before listing failure modes is how sites end up with a lot of data about the wrong things.

How often to check: the P-F interval

The P-F interval is the time between the point a failure first becomes detectable (P) and the point of functional failure (F). It sets the monitoring frequency. Check less often than the P-F interval and a failure can start and finish between readings.

A common rule of thumb is to check at no more than half the P-F interval, so that at least one reading falls inside the warning period with time to spare. The time left matters as much as the detection. If planning, parts and access take three weeks, the program has to find the fault at least three weeks before it fails, so the useful warning is the P-F interval less the time needed to respond.

P-F intervals differ by technique for the same failure. A bearing defect may show in ultrasound and high-frequency vibration months before it shows in overall vibration levels, well before it produces measurable heat, and long before anyone can hear it. Earlier techniques buy more planning time, which is why critical equipment often carries more than one.

If the P-F interval isCheck at leastTypical approach
MonthsMonthly, or every few weeksRoute-based portable collection or periodic oil sampling
WeeksWeekly or more oftenFrequent routes, or online monitoring on critical assets
Days or hoursContinuouslyOnline monitoring with alarms, or a design or redundancy solution

Online monitoring or route-based collection

Route-based monitoring sends a technician with portable instruments around a fixed route, often monthly for vibration. It is flexible and cheap per point, and it puts a trained person beside the machine, who notices things no sensor measures. Its limits are frequency, access, and equipment that cannot be safely reached while it runs.

Online monitoring fixes sensors to the machine and streams data continuously, increasingly over wireless networks. It suits critical equipment with short P-F intervals, machines in hazardous or hard-to-reach places, and faults that only appear under particular loads. The data usually lands in a data historian and needs its own alarm logic and review routine, which is where many online systems quietly stop being used. Sizing the storage for high-frequency data is its own exercise, and the time-series data volume calculator gives a first estimate.

Remote and distributed assets, such as pump stations and rail-side equipment, are often only practical to monitor online, and the remote telemetry case study covers how that is set up. Most mature programs use both approaches, with routes covering the broad population and online systems covering the few assets where a missed failure is unaffordable.

Condition monitoring in mining, rail and processing

  • On conveyors, idler and pulley bearings are numerous and individually cheap, so thermography routes, acoustic monitoring and belt inspection systems cover them far more economically than a sensor on every idler. Drive trains and critical pulleys justify vibration and oil analysis.
  • Crushers and grinding mills carry gearboxes, girth gears and large slow-speed bearings that need specialist vibration set-ups and oil analysis, while liner wear is tracked by thickness measurement or scanning.
  • Mobile fleets such as haul trucks and loaders lean heavily on oil analysis and on the equipment's own onboard health data, with component replacement decided on trend rather than hours alone.
  • Rail networks use wayside systems, including hot bearing detectors, wheel impact load detectors and acoustic bearing monitors, to check every wagon as it passes, alongside track and wheel measurement.
  • Process plants follow the same pattern for rotating equipment, while static equipment such as vessels and piping is managed through inspection and thickness monitoring, often planned with risk based inspection.

Building a condition monitoring program

Understand the current state first. Many sites already have routes, a contractor or an online system, and the useful question is what those have caught and missed over the last year. Assess coverage against criticality, and the program against your own standards, with ISO 17359 as a helpful reference for the overall structure.

Then define the tasks for critical assets, take baselines while equipment is healthy, and agree alarm limits, starting from standards and manufacturer guidance and tightening them against your own trend data. For a new asset this work belongs in operational readiness, with routes, baselines and limits in place before handover rather than after the first failure.

Connect findings to work management. A condition finding should raise a notification or work request in the CMMS, against the right equipment record and with the evidence attached, so it is planned like any other job. Programs that report findings only in a separate portal or a monthly PDF lose most of their value in the handover.

Keep competency in-house even when collection is contracted. Someone on site needs to understand the reports, challenge them and own the response. And review the program against failures. Every significant failure should prompt the question of whether monitoring should have caught it, and if not, why not.

Measuring whether the program works

Count outcomes, not readings. A program that takes thousands of readings and converts few of them into planned work is expensive data collection. The measures below show whether monitoring is earning its keep.

Put a value on the failures it prevents with the Cost of Downtime Calculator, and track the reliability of monitored equipment over time with the MTBF / MTTR Calculator.

MeasureWhat it tells you
Coverage of critical assetsWhether monitoring is where the risk is
Route or collection complianceWhether the planned checks get done
Findings converted to planned workWhether detection turns into action
Warning time achievedWhether faults are found early enough to plan, kit and schedule
Failures missed by monitoringWhere techniques, intervals or limits need to change
False alarm rateWhether people will keep trusting the alarms
Unplanned failures on monitored assetsThe trend that justifies the program's cost

Why condition monitoring programs fail

  • Technology bought before failure modes were understood, so the sensors watch the wrong things.
  • Every asset monitored the same way, with no link to criticality, which spreads effort thinnest where it matters most.
  • Alarm limits left at a vendor default and never tuned, producing either constant false alarms or silence.
  • Findings reported outside the CMMS, so they never become planned work.
  • Nobody on site owns the program, so a contractor's report arrives, gets filed and is forgotten.
  • Intervals set by habit rather than by the P-F interval, so failures develop between readings.
  • No review after failures, so the same miss happens twice.

Building a condition monitoring program

  1. Rank by criticality

    Focus monitoring where failure consequences are highest.

  2. List failure modes

    Dominant failure modes for each critical asset, from history and FMECA.

  3. Match techniques

    A technique that detects each failure mode in time, at a justified cost.

  4. Set intervals and limits

    Frequency from the P-F interval, alarms from standards and baselines.

  5. Connect to work

    Findings raise work in the CMMS and are planned like any other job.

  6. Review against failures

    Check every miss, then tune techniques, intervals and limits.

Common questions

What is condition monitoring?

Condition monitoring is measuring the health of equipment while it runs, using techniques such as vibration analysis, oil analysis, thermography and ultrasound, and trending the results to find developing failures early. The aim is to plan the repair before the equipment fails, rather than react to a breakdown.

What is the difference between condition monitoring and predictive maintenance?

Condition monitoring is the measurement. Predictive maintenance, also called condition-based maintenance, is the strategy of scheduling work from what the measurements show. Monitoring without a process to act on the findings is not yet predictive maintenance.

What are the main condition monitoring techniques?

Vibration analysis for mechanical faults in rotating equipment, oil analysis for wear and contamination, infrared thermography for abnormal heat, and ultrasound for friction, leaks and electrical discharge. Motor current analysis, thickness measurement, process performance data and operator rounds complete most programs.

How often should condition monitoring be done?

Often enough to detect a failure with time left to act. The interval comes from the P-F interval, the warning time between a failure becoming detectable and the equipment failing. A common rule of thumb is to check at no more than half the P-F interval, and more often still when planning and parts take a long time.

What is the P-F interval?

The P-F interval is the time between the point a developing failure can first be detected (P) and the point the equipment stops performing its function (F). It varies by failure mode and by technique, and it sets how often monitoring has to happen to catch the failure in time.

Should condition monitoring be online or route-based?

Usually both, for different equipment. Route-based collection with portable instruments is flexible and economical for most assets. Online monitoring suits critical equipment with short P-F intervals, hazardous or inaccessible locations, and faults that only appear under particular operating conditions.

Which equipment should be monitored?

Start from criticality. Monitor the equipment whose failure has the highest consequence, for the failure modes that give detectable warning. Low-consequence items that are cheap to replace are often better run to failure or covered by operator rounds.

Which standards apply to condition monitoring?

ISO 17359 gives general guidelines for setting up a program. Technique standards include ISO 20816 for evaluating vibration, ISO 14830-1 for oil and lubricant analysis, ISO 18434-1 for thermography and ISO 29821 for ultrasound, with the ISO 18436 series covering personnel competence. Most sites combine these with their own internal standards.

Key terms

Plain-language definitions from our glossary for the concepts this article leans on.

Standards and further reading

Related case studies and tools

Related reading

Working through something like this?

See how we approach these initiatives, or tell us what you are dealing with.