Summary
In short
- Reliability-centred maintenance analysis consistently finds that most failure modes are not age-related, which invalidates the assumption behind calendar-based replacement for the majority of components.
- Intrusive maintenance introduces failure. Every disturbance carries a probability of incorrect reassembly, contamination and damage, and the period immediately after a major intervention shows elevated failure rates.
- Condition-based tasks are preferable where a detectable warning exists and the interval fits inside it. The relevant question is how long between the first detectable indication and functional failure.
- Task selection should follow from consequence: safety, environmental, operational and economic consequences of failure each warrant different approaches.
- Run-to-failure is a legitimate strategy for non-critical components with no safety or environmental consequence, and it is a decision rather than a default.
- Corrective work on assets under a preventive regime is the feedback signal. Where it is not analysed, the plan cannot be known to be working.
What it is
What it is
What is a preventive maintenance plan?
The schedule of maintenance tasks performed on an asset before failure: inspection, condition monitoring, servicing, adjustment and scheduled replacement, with intervals and the reasoning that set them.
Why does the failure pattern matter?
Because it determines whether a time-based task can work at all. Scheduled replacement or overhaul only prevents failure where the failure mode is age-related. Where failure is random with respect to age, replacing a component on a calendar does not reduce the failure rate, and the intrusive work introduces a fresh period of elevated infant mortality.
When to use it
When to use it, and when not to
This plan sets out what maintenance is performed before failure, and why. Work performed sits separately.
Use it for
- Designing the maintenance regime for an asset or asset class
- Reviewing intervals and task content against actual failure history
- Introducing new equipment, where the initial regime comes from the manufacturer and needs adjustment
- After a change in duty, operating environment or criticality
- Where corrective work on assets under the regime indicates the plan is not working
Not for
- Corrective work orders, which record unplanned repairs and provide the feedback this plan needs
- The asset register, which holds identity, criticality and statutory obligations
- Condition monitoring data, which is an input to condition-based decisions
- Statutory inspection regimes, which are mandated and not subject to optimisation
- Spares strategy, which follows from criticality and lead time
Standards
What it is built against
Maintenance planning sits under asset management standards, with methodology from the reliability literature.
| Clause | Requirement | Where it lands |
|---|---|---|
| ISO 55001 cl.8.1 | Operational planning and control of activities to achieve asset management objectives | Header |
| ISO 55001 cl.6.2.2 | Planning to achieve objectives, informed by asset criticality and risk | Tasks |
| ISO 55001 cl.9.1 | Evaluation of asset performance and effectiveness of the asset management system | Plan quality |
| SAE JA1011 | Evaluation criteria for reliability-centred maintenance processes | Tasks |
| ISO 14224 | Collection and exchange of reliability and maintenance data, supporting interval decisions | Plan quality |
| ISO 9001 cl.7.1.3 | Infrastructure maintained to ensure conformity of products and services | Header |
| 21 CFR 117.40 | Plant and equipment maintained in a condition preventing contamination of food | Tasks |
| NFPA 70B 2023 | Electrical maintenance programme with intervals from equipment condition assessment | Tasks |
What it does not cover
- Corrective work orders, recording unplanned repairs and providing the feedback the plan depends on.
- The asset register, holding identity, criticality and statutory obligations.
- Condition monitoring data, an input to condition-based task decisions.
- Statutory inspection regimes, which are mandated and not subject to optimisation.
- Spares strategy, which follows from criticality and lead time.
Filling it in
Filling it in well
Choose tasks from failure mode and consequence, not from habit, and let the failure history correct the plan.
What actually fails, how it fails, and whether that failure mode has an identifiable wear-out age. A pump does not fail; a seal fails through wear, a bearing fails through contamination or misalignment, an impeller fails through erosion. Each has a different pattern and warrants a different task, and a task specified against the pump addresses none of them precisely.
Where there is a detectable indication before functional failure, and the interval between detection and failure is long enough to act, condition monitoring beats scheduled replacement. The question that decides it is how much warning the indication actually gives, which determines the inspection interval.
Safety and environmental consequences justify tasks regardless of cost. Operational consequence justifies tasks up to the cost of the downtime. Economic consequence justifies tasks up to the cost of the failure. And where there is no meaningful consequence, run-to-failure is a legitimate choice that should be recorded as a decision rather than arrived at by neglect.
Failures on assets under a preventive regime tell you the plan is wrong: the task did not cover that component, the interval was too long, or the task was not actually performed. That feedback loop is the only mechanism by which a plan improves, and in most operations the two record sets are never compared.
Audit findings
Common audit findings
Maintenance plan findings concentrate on the basis for tasks and intervals.
| Finding | Clause | What fixes it |
|---|---|---|
| Calendar-based replacement applied where the failure mode is not age-related. | SAE JA1011 | Time-based tasks work only where wear-out is identifiable; otherwise use condition-based. |
| Intervals inherited from the manufacturer and never reviewed against failure history. | ISO 55001 cl.9.1 | Manufacturer intervals are a starting point set without knowledge of your duty. |
| Tasks specified against assets rather than failure modes. | SAE JA1011 | Specify against what actually fails and how. |
| Corrective failures on assets under PM not analysed. | ISO 55001 cl.9.1 | These are the feedback signal; without them the plan cannot be known to work. |
| Intrusive overhaul performed on components with random failure patterns. | SAE JA1011 | The disturbance introduces infant mortality without reducing the failure rate. |
| Run-to-failure applied by default rather than by decision. | ISO 55001 cl.6.2.2 | It is a legitimate strategy where consequence is low, and should be recorded as chosen. |
| Condition monitoring interval longer than the warning period. | SAE JA1011 | The interval must fit inside the time between detectable indication and functional failure. |
| Hidden failures with no failure-finding task. | SAE JA1011 | Protective devices fail unnoticed; scheduled function tests are the only way to know. |
| Plan not adjusted after a change in duty or operating environment. | ISO 55001 cl.8.1 | Duty determines wear; a change in running hours or product changes the regime. |
| PM compliance measured without any measure of effectiveness. | ISO 55001 cl.9.1 | Completion rate says the tasks were done, not that they were the right tasks. |
Worked case
Case in point: the overhaul that caused the failures
A plant overhauled a set of gearboxes annually, stripping and rebuilding each one during the summer shutdown. The regime had been in place for years and was regarded as good practice.
Analysis of failure dates showed a clear pattern: failures clustered in the weeks immediately following the shutdown, then fell away for the remainder of the year. The gearboxes were not wearing out; they were being disturbed. Seal seating, bearing preload and contamination introduced during reassembly were producing failures the overhaul was intended to prevent.
Moving to condition monitoring, with vibration and oil analysis and intervention on indication, reduced both the failure rate and the maintenance cost. The gearboxes that had run for years without being opened were the ones that had never failed.
Definitions
Definitions and key terms
- Failure mode
- The specific way a component fails, which determines whether a time-based task can prevent it.
- Age-related failure
- A failure mode showing an identifiable wear-out point, where scheduled replacement can work.
- Infant mortality
- Elevated failure probability immediately after installation or intervention, introduced by intrusive maintenance.
- Condition-based maintenance
- Intervention triggered by a detectable indication of impending failure rather than by elapsed time.
- P-F interval
- The time between the first detectable indication of failure and functional failure, which sets the inspection interval.
- Failure-finding task
- A scheduled function test of a protective device whose failure would otherwise be hidden until needed.
- Run to failure
- A deliberate strategy of not intervening before failure, legitimate where consequences are low.
- Hidden failure
- A failure not evident to operators under normal conditions, typically in standby and protective systems.
FAQ
Frequently asked questions
Why is calendar-based maintenance often wrong?+
Because it assumes failure is age-related, and reliability analysis consistently finds that most failure modes are not. Where failure is random with respect to operating age, replacing a component on a schedule swaps one component for another with the same failure probability, and adds the disturbance risk of the intervention. Time-based tasks belong where a genuine wear-out age exists.
Can maintenance cause failures?+
Yes, and it is a well-documented effect. Every intrusive intervention carries a probability of incorrect reassembly, contamination, damage to adjacent components and disturbed settings, which produces elevated failure rates in the period immediately afterwards. Plotting failures against time since last intervention will show it where it is happening.
How do we decide which tasks to perform?+
From the failure mode and the consequence. Identify what actually fails and how; establish whether there is a detectable warning and how much time it gives; and set the task by consequence, with safety and environmental consequences justifying tasks regardless of cost and economic consequences justifying tasks up to the cost of the failure.
Is run-to-failure acceptable?+
For components where failure has no safety or environmental consequence and the operational and economic consequences are low, yes, and it is a legitimate strategy rather than an admission. What matters is that it is recorded as a decision with reasoning, rather than being what happens to components nobody got round to including in the plan.
How do we know the plan is working?+
By analysing corrective work on assets under the regime. A failure on an asset with a preventive task covering that component means the task did not cover it, the interval was too long, or it was not performed. Measuring PM completion rate tells you the tasks were done; it says nothing about whether they were the right tasks.
The agents
What the agents do with it
The plan sets what is done before failure. What fails is intervals nobody reviewed and corrective work nobody read back.
Holds tasks against failure modes with the reasoning for each interval, and compares corrective failures to the regime covering that component.
Plots failures against time since last intervention, which directly tests whether the regime is preventing or producing them.
Identifies protective and standby systems requiring failure-finding tasks, where failure is hidden until the device is needed.
Connects maintenance to product safety where equipment condition affects contamination risk, and to validated configurations.
This template lives in KnowMaintain — asset maintenance. Work orders, planned maintenance, calibration, reliability and shutdowns.
Meet KnowMaintain→Sources
Sources
- ISO 55001:2014, asset management, clauses 6.2.2, 8.1 and 9.1
- SAE JA1011, evaluation criteria for reliability-centred maintenance processes
- ISO 14224:2016, collection and exchange of reliability and maintenance data
- 21 CFR 117.40, plant and equipment, FDA
- NFPA 70B 2023, standard for electrical equipment maintenance