Purpose
Apply RCM decision logic in strict order: on-condition first, scheduled restoration next, then scheduled discard, then failure-finding, then redesign or run-to-failure. The consequence category from Phase 5 governs the cost/feasibility test.
Scope
- On-condition tasks (predictive / condition-based)
- Scheduled restoration (overhaul at defined age)
- Scheduled discard (replace at defined age)
- Failure-finding tasks (for hidden failures only)
- Default actions: no scheduled maintenance, or redesign
Process Steps
- 01
Test for an on-condition task
IntentIs there a detectable warning of failure (potential-failure point P), and is the P–F interval long enough to plan and act before functional failure?
Inputs- Failure mode + effect + consequence from Phase 5
- Available condition-monitoring technology (vibration, oil, IR, ultrasonic, DCS trend, visual)
- Manufacturer or reference P–F intervals; internal history
Actions- Confirm a detectable P point exists for this mode
- Confirm technology and access are available and cost-effective
- Set inspection interval at less than half the P–F interval
- Document who does the inspection, with what, and the action limit
PeopleReliability engineer owns the logic. Condition-monitoring specialist confirms feasibility. Planner confirms it can actually be executed.
ProcessOn-condition is tested first for every mode. Only when it fails all four feasibility criteria does the team move to Step 6.2.
TechnologyRCM decision worksheet with task-type column and interval column. Condition-monitoring database.
What Good Looks Like- Every reliability, maintenance and operations lead can point to where this step lives — decisions, evidence and outputs are in one place, not in someone's inbox.
- The opposite of: choosing time-based maintenance without testing on-condition first
- The opposite of: setting inspection interval equal to or greater than the P–F interval
- The opposite of: no documented action limit
What Bad Looks Like- Choosing time-based maintenance without testing on-condition first
- Setting inspection interval equal to or greater than the P–F interval
- No documented action limit — inspectors don't know when to react
Outputs- Accepted on-condition task with interval, technology, action limit and responsible role
- Rejected candidates flagged with the reason (no P, no technology, uneconomic)
- 02
Test for scheduled restoration or discard
IntentOnly valid when there is clear evidence that conditional probability of failure increases sharply at a definable age — do not assume wear-out.
Inputs- Weibull analysis or reliable historical failure-age data
- Cost of restoration / replacement vs cost of failure
- Manufacturer end-of-life guidance
Actions- Require Weibull or reliable historical evidence — do not assume wear-out
- Set the intervention age just before the wear-out zone
- Confirm the task actually restores initial resistance (restoration) or removes the component (discard)
PeopleReliability engineer runs the analysis. Trades confirm restoration is technically feasible.
ProcessAge-based tasks are the exception, not the default — RCM data shows only ~11% of failures are age-related.
TechnologyWeibull / reliability-analysis tool. CMMS failure history filtered by mode.
What Good Looks Like- Every reliability, maintenance and operations lead can point to where this step lives — decisions, evidence and outputs are in one place, not in someone's inbox.
- The opposite of: applying calendar-based PMs to random failures
- The opposite of: overhaul intervals set by 'we've always done it that way'
- The opposite of: discarding parts that show no age-related failure pattern
What Bad Looks Like- Applying calendar-based PMs to random failures — burns labour with no benefit
- Overhaul intervals set by 'we've always done it that way'
- Discarding parts that show no age-related failure pattern
Outputs- Accepted restoration or discard task with age-based interval
- Rejected candidates with a note that no age-related failure exists
- 03
Test for failure-finding (hidden failures)
IntentProve the hidden protective function still works, at an interval that meets the tolerable unavailability of the protected system.
Inputs- Hidden-flag failures from Phase 5
- MTIVE (mean time between failure of the protected function)
- Tolerable unavailability target for the protected system
Actions- Use FFI = 2 × MTIVE × unavailability tolerance (or the site formula)
- Confirm the test actually exercises the protective function end-to-end
- Never confuse failure-finding with condition monitoring
PeopleReliability engineer + HSE. Instrument technician confirms the test method actually exercises the function.
ProcessFailure-finding intervals are calculated, not chosen. A hidden failure without a failure-finding task is a compliance gap.
TechnologyRCM decision worksheet; SIS proof-test schedule; CMMS PM plan.
What Good Looks Like- Every reliability, maintenance and operations lead can point to where this step lives — decisions, evidence and outputs are in one place, not in someone's inbox.
- The opposite of: missing failure-finding tasks for hidden failures
- The opposite of: tests that check power to the device but not the trip action itself
- The opposite of: intervals guessed rather than calculated
What Bad Looks Like- Missing failure-finding tasks for hidden failures
- Tests that check power to the device but not the trip action itself
- Intervals guessed rather than calculated
Outputs- Failure-finding task with interval and test method
- Traceability from the task back to the protective function it verifies
- 04
Default action
IntentIf no proactive task is technically feasible and worth doing: run-to-failure (non-safety) or redesign (safety / hidden that cannot be tolerably reduced).
Inputs- Rejected proactive candidates from Steps 6.1–6.3
- Redesign register from Phase 5
- Cost of run-to-failure over asset life
Actions- Document the rationale for run-to-failure explicitly (not silence)
- Log redesign proposals into the engineering-change register with a business case
- Confirm run-to-failure only ever applies to non-safety, non-hidden modes
PeopleReliability engineer records. Asset owner endorses run-to-failure. Engineering-change board owns redesign delivery.
ProcessDefault action is a deliberate decision, documented per mode — never a silent 'nothing scheduled'.
TechnologyRCM decision worksheet; engineering-change register; run-to-failure register.
What Good Looks Like- Every reliability, maintenance and operations lead can point to where this step lives — decisions, evidence and outputs are in one place, not in someone's inbox.
- The opposite of: silent run-to-failure
- The opposite of: redesigns logged but never sponsored, so they never happen
- The opposite of: applying run-to-failure to safety or hidden modes
What Bad Looks Like- Silent run-to-failure — the mode simply drops off the worksheet
- Redesigns logged but never sponsored, so they never happen
- Applying run-to-failure to safety or hidden modes
Outputs- Run-to-failure register with justification per mode
- Redesign register with proposal, sponsor and target date
People
Reliability engineer owns the decision logic. Trades confirm technical feasibility. Planner confirms task can actually be executed.
Process
Follow the logic in order — do not shortcut to time-based tasks. Every accepted task has a stated interval, resource, duration and skill.
Technology
RCM decision worksheet with columns for task type, interval, duration, trade and instructions.
Phase Outputs
- Proposed maintenance task list with type, interval and resource
- Redesign register
- Run-to-failure register (with justification)
What Good Looks Like
- Every reliability, maintenance and operations lead can point to where this phase lives — decisions, evidence and outputs are all in one place, not in someone's inbox.
- The opposite of: defaulting to time-based PMs
- The opposite of: missing failure-finding tasks for hidden failures
- The opposite of: setting intervals without reference to the P–F interval
Common Pitfalls
- Defaulting to time-based PMs — most failures are random
- Missing failure-finding tasks for hidden failures
- Setting intervals without reference to the P–F interval
