RCM Process · Phase 4

Failure Modes & Effects (FMEA)

Answer RCM Questions 3 & 4: what causes each functional failure, and what happens when it occurs?

Purpose

The FMEA lists every reasonably likely cause of each functional failure and describes what actually happens — evidence of failure, safety and environmental impact, effect on production and secondary damage.

Scope

  • Component-level failure causes
  • Human-error and design-error causes where credible
  • Wear-out, random and infant-mortality mechanisms
  • Failure effects: evidence, safety, environment, operations, secondary damage

Process Steps

  1. 01

    List failure modes per functional failure

    Intent

    Enumerate every reasonably likely cause of each functional failure — mechanical, corrosion, contamination, human error, design error, loss of utilities.

    Inputs
    • Validated functional-failure list from Phase 3
    • Prior FMEAs on similar assets
    • Bad-actor and RCFA history from the CMMS
    • OEM failure-mode libraries and industry data (OREDA / ISO 14224)
    Actions
    • Use trade knowledge and prior FMEAs to seed the list
    • Import bad-actor history and RCFA outputs
    • Include human-error and design-error causes where credible
    • Stop at 'reasonably likely' — do not chase every theoretical cause
    People

    Trades and operators contribute causes from lived experience. Reliability engineer facilitates and records. Asset engineer challenges credibility.

    Process

    Modes are captured per functional failure — never a free-for-all list. Cadence is timeboxed per system so momentum is kept.

    Technology

    FMEA worksheet linked upward to the failure register. Reference library of prior FMEAs.

    What Good Looks Like
    • Every reliability, maintenance and operations lead can point to where this step lives — decisions, evidence and outputs are in one place, not in someone's inbox.
    • The opposite of: listing components instead of failure modes
    • The opposite of: analysis paralysis on unlikely modes
    • The opposite of: ignoring human error and design shortfall as valid causes
    What Bad Looks Like
    • Listing components instead of failure modes
    • Analysis paralysis on unlikely modes
    • Ignoring human error and design shortfall as valid causes
    Outputs
    • Populated failure-mode column on the FMEA worksheet, one row per mode
    • Traceability from each mode back up to its functional failure and function
  2. 02

    Describe failure effects

    Intent

    For each mode, describe evidence to the operator, safety, environmental, operational and secondary-damage effects — Phase 5 depends on this.

    Inputs
    • Failure-mode list from Step 4.1
    • Safety, environmental and operational history
    • Downtime and repair-cost data from the CMMS
    Actions
    • For each mode, describe what the operator sees / hears / measures
    • State the safety impact — injury, fatality, exposure
    • State the environmental impact — spill, emission, breach of permit
    • State the operational impact — throughput hit, quality, cycle time
    • State secondary damage and what is required to restore
    • Quantify downtime and repair time where data supports it
    People

    Operators describe evidence. Trades describe restoration. HSE describes safety / environmental impact. Reliability engineer captures.

    Process

    Effects are written in plain language — an operator should recognise the description. Empty effects rows are not permitted to pass into Phase 5.

    Technology

    FMEA worksheet with linked columns for evidence, S/E/O impact, secondary damage and restoration.

    What Good Looks Like
    • Every reliability, maintenance and operations lead can point to where this step lives — decisions, evidence and outputs are in one place, not in someone's inbox.
    • The opposite of: skipping the effects column
    • The opposite of: effects written in engineering language operators can't recognise
    • The opposite of: no restoration time / cost estimate
    What Bad Looks Like
    • Skipping the effects column — Phase 5 becomes guesswork
    • Effects written in engineering language operators can't recognise
    • No restoration time / cost estimate — Phase 6 can't weigh options
    Outputs
    • Effects column populated for every failure mode
    • Restoration time / cost estimates carried onto the worksheet

People

Trades and operators contribute causes and effects from lived experience. Reliability engineer facilitates and captures.

Process

This is the deepest phase. Keep pace by timeboxing per system and finalising one function-failure at a time.

Technology

FMEA worksheet — link each row up to its functional failure and function.

Phase Outputs

  • Populated FMEA worksheet (function → functional failure → failure mode → effect)
  • Priority bad-actor shortlist ready for consequence analysis

What Good Looks Like

  • Every reliability, maintenance and operations lead can point to where this phase lives — decisions, evidence and outputs are all in one place, not in someone's inbox.
  • The opposite of: listing components instead of failure modes
  • The opposite of: analysis paralysis on unlikely modes
  • The opposite of: skipping the effects column

Common Pitfalls

  • Listing components instead of failure modes
  • Analysis paralysis on unlikely modes
  • Skipping the effects column — Phase 5 depends on it