Understand
Who this role is and what they own.
Engineering Manager / Reliability Manager
- Supervisor — sponsor of investigations and strategy changes
- Planner — implements Standard Job Plan and PM changes
- Operations Leader — provides consequence and production context
- Own the failure code catalogue and its integrity
- Run the monthly bad-actor review
- Lead FMEA / RCM reviews and document per-mode decisions
Deep dive · phase-by-phase detail
P1Identify WorkOpen ↓Close ↑
The Reliability Engineer turns condition-monitoring signals and repeat-failure patterns into work requests before Operations feel the consequence. If they wait until a bad actor breaks again to act, the strategy has already failed — the whole point of condition-based work is to act inside the P-F interval, not after it.
- Scan overnight condition-monitoring alarms every morning and confirm P-F interval actions are raised as work requests
- Screen new work requests for repeat descriptions and cluster them for bad-actor review
- Own the criticality register — every new asset gets a criticality rating before it goes live, not retrospectively
- Set condition-monitoring routes and thresholds based on failure mode, not vendor default
- Flag any asset that has failed twice in a rolling year as a candidate for the bad-actor list, before Operations complain
- Condition-based work is raised inside the P-F interval and scheduled in time to prevent functional failure
- Repeat failures on the same asset are surfaced within one week, not one quarter
- The criticality register is current — no asset in production without a rating
- Condition-monitoring alarms sit unread; repeat failures are re-raised as fresh reactive work each cycle
- Reliability Engineer becomes a reporting function — publishes trends but never raises the work
- Criticality register hasn't been touched in two years — new assets get a default rating and nobody notices
P2Plan WorkOpen ↓Close ↑
The Reliability Engineer sets the strategy content of Standard Job Plans on critical assets — the precision maintenance standards, hold points and acceptance criteria that turn a job from 'we did the work' into 'we can prove the asset was returned to specification'. Generic checklists cannot deliver precision maintenance.
- Co-author Standard Job Plans for critical assets with the Planner — do not delegate the reliability content
- Specify torque, alignment, cleanliness and vibration acceptance criteria in measurable terms on every critical Standard Job Plan
- Define hold points where the job must stop and be verified before proceeding — signed by name
- Review the PM task list on any asset added to the bad-actor list — often the wrong task is being done at the wrong frequency
- Retire PM tasks that no longer address a real failure mode; add ones for failure modes emerging in the field
- Standard Job Plans on critical assets carry measurable acceptance criteria — not just a list of steps
- PM task content on critical assets is traceable to a documented failure mode
- Precision-maintenance evidence exists on every critical asset repair — torque figures, alignment sheets, vibration baselines
- Standard Job Plans are generic checklists — no way to prove precision maintenance actually happened
- PM tasks copied from the OEM manual and never questioned, even after multiple in-service failures
- Reliability sets strategy on a whiteboard but never lands it in the CMMS — nothing changes on the floor
P3Schedule WorkOpen ↓Close ↑
The Reliability Engineer defends condition-based work against being casually deferred. Every CBM work order has a P-F interval; deferring beyond it converts predictive work into reactive work and breaks the whole strategy. The Reliability Engineer must make this cost visible to the Scheduler and the Supervisor.
- Flag any condition-based work order about to breach its P-F interval to the Scheduler with the consequence in plain language
- Attend the weekly scheduling meeting when critical-asset or bad-actor work is on the candidate list
- Publish CBM breach rate weekly alongside schedule compliance so both KPIs are visible together
- Escalate any condition-based order deferred more than once — the strategy is broken, not just the schedule
- Confirm strategy-change work orders (PM adds, retires, frequency changes) actually get scheduled within a month of approval
- CBM breach rate is reported weekly and stays near zero
- Strategy-change work lands on the schedule inside a month of the decision, not the next quarter
- The Scheduler treats condition-based work as non-negotiable, not 'nice to have if we have time'
- CBM work rolls week to week until it becomes reactive — strategy invisible, effort wasted
- Approved strategy changes sit in the CI register for months waiting to be scheduled
- P-F breaches never quantified, so no one knows the strategy is failing until an asset trips
P4Execute & RecordOpen ↓Close ↑
The Reliability Engineer sets and audits the precision maintenance standard at the asset itself. Standards on paper mean nothing if the actual work on the floor doesn't meet them. Their job in this phase is to coach through presence — attend the jobs, audit the readings, baseline the repairs — not to police from a spreadsheet.
- Attend selected precision jobs and audit against the Standard Job Plan acceptance criteria in real time
- Baseline vibration, thermography or oil analysis after every major repair before the asset is released to production
- Coach the tradesperson at the job face when acceptance criteria are missed — not weeks later in a report
- Review handback statements on critical assets — is the condition claim actually supported by evidence?
- Sample-audit permits and isolations on critical-asset work — the safety envelope is a reliability issue too
- Post-repair baselines exist on every critical repair; infant-mortality failures are trending down
- Precision maintenance evidence is captured against the work order, not in a private folder
- Coaching happens at the asset — tradespeople know the Reliability Engineer as a person, not a report
- 'Return to service' with no baseline — no way to detect early re-failure until it re-fails
- Reliability Engineer never seen on the floor — precision maintenance stays a paper concept
- Audits done from the office reading history text — real workmanship never inspected
P5Close-out & AnalysisOpen ↓Close ↑
The failure-code catalogue and the history entered on close-out are the raw material of every reliability insight the site will ever produce. If the Reliability Engineer does not police this data quality, every Pareto chart, MTBF calculation and RCM decision downstream is built on garbage.
- Sample-audit close-outs weekly for failure-code accuracy and history depth — publish the audit result
- Retire dead failure codes; add missing ones; keep the catalogue small and used, not encyclopaedic and ignored
- Review all 'Other' or 'Unknown' failure-code entries monthly and re-classify — do not let the bucket grow
- Feed close-out quality trends back to the Supervisor for coaching — this is not just a reliability problem
- Trigger a formal root-cause investigation on any failure that damages a critical asset, without waiting for it to reach the bad-actor list
- Failure-code catalogue has under 100 active codes and over 90% coverage in real use
- History quality on critical assets is audited monthly and trends up quarter over quarter
- Every critical-asset failure has an investigation record — even the ones that looked minor at the time
- 1,000-code catalogue that everyone bypasses with 'Other' — no analysis possible
- History quality accepted as 'as good as it gets' — the site regresses to reactive quietly
- Investigations only happen after the second or third failure — the first one is lost forever
P6Metrics & Continuous ImprovementOpen ↓Close ↑
The Reliability Engineer owns the full continuous improvement loop — from Pareto through investigation through strategy change through verification. The measure of success is not how many reports were produced but how many failure modes have been permanently eliminated. If the same bad actors appear quarter after quarter, the loop is not closing.
- Run the monthly bad-actor review; assign a named investigator per top item, with a due date
- Publish weekly and monthly KPI packs with trend commentary — say what the numbers actually mean, not just what they are
- Push approved strategy changes to the Planner and verify implementation in the CMMS within 30 days
- Escalate any CI action overdue more than 60 days to the Maintenance Manager — do not let the register rot
- Verify strategy-change effectiveness at 6 and 12 months — has the failure mode actually reduced, or did we just move it?
- Run at least one full RCM or FMEA review per quarter on a critical asset class
- Every bad actor has an investigation, a decision, an implementation date and a verification date
- Strategy changes show up as measurable failure-mode reductions 6-12 months later — the loop is closed
- The CI register is short and active, not long and stale
- Reports produced, no actions closed — 'analysis theatre' becomes the whole function
- KPI definitions redefined quarterly so no long-term trend is possible
- Investigations end with 'need more data' every time — the same bad actors return year after year
- Condition-based work is raised inside the P-F interval and scheduled in time to prevent functional failure
- Repeat failures on the same asset are surfaced within one week, not one quarter
- The criticality register is current — no asset in production without a rating
- Condition-monitoring alarms sit unread; repeat failures are re-raised as fresh reactive work each cycle
- Reliability Engineer becomes a reporting function — publishes trends but never raises the work
- Criticality register hasn't been touched in two years — new assets get a default rating and nobody notices
Learn
Courses, guides and reading to build the capability.
Competencies
Implement
Tools you run daily, weekly, monthly.
Cadenced checklists
Tick items as you go. Progress is saved in your browser — reset any time.
Improve
Measure how you're doing and where to sharpen next.
Run an audit
KPIs to own
If you see this, fix it
- Reports that produce actions no one closes — analysis theatre
- Redefining KPIs quarterly — no trend possible
- Investigations that end with 'need more data' every time
