⚙️ Work Management by Role

Reliability Engineer

Runs the intelligence loop — bad-actor analysis, FMEA / RCM, strategy change and continuous improvement.

01

Understand

Who this role is and what they own.

Reports to

Engineering Manager / Reliability Manager

Key relationships
  • Supervisor — sponsor of investigations and strategy changes
  • Planner — implements Standard Job Plan and PM changes
  • Operations Leader — provides consequence and production context
Top accountabilities
  • Own the failure code catalogue and its integrity
  • Run the monthly bad-actor review
  • Lead FMEA / RCM reviews and document per-mode decisions

Deep dive · phase-by-phase detail

P1Identify Work
Open ↓

The Reliability Engineer turns condition-monitoring signals and repeat-failure patterns into work requests before Operations feel the consequence. If they wait until a bad actor breaks again to act, the strategy has already failed — the whole point of condition-based work is to act inside the P-F interval, not after it.

What this role does
  • Scan overnight condition-monitoring alarms every morning and confirm P-F interval actions are raised as work requests
  • Screen new work requests for repeat descriptions and cluster them for bad-actor review
  • Own the criticality register — every new asset gets a criticality rating before it goes live, not retrospectively
  • Set condition-monitoring routes and thresholds based on failure mode, not vendor default
  • Flag any asset that has failed twice in a rolling year as a candidate for the bad-actor list, before Operations complain
Good looks like
  • Condition-based work is raised inside the P-F interval and scheduled in time to prevent functional failure
  • Repeat failures on the same asset are surfaced within one week, not one quarter
  • The criticality register is current — no asset in production without a rating
Bad looks like
  • Condition-monitoring alarms sit unread; repeat failures are re-raised as fresh reactive work each cycle
  • Reliability Engineer becomes a reporting function — publishes trends but never raises the work
  • Criticality register hasn't been touched in two years — new assets get a default rating and nobody notices
P2Plan Work
Open ↓

The Reliability Engineer sets the strategy content of Standard Job Plans on critical assets — the precision maintenance standards, hold points and acceptance criteria that turn a job from 'we did the work' into 'we can prove the asset was returned to specification'. Generic checklists cannot deliver precision maintenance.

What this role does
  • Co-author Standard Job Plans for critical assets with the Planner — do not delegate the reliability content
  • Specify torque, alignment, cleanliness and vibration acceptance criteria in measurable terms on every critical Standard Job Plan
  • Define hold points where the job must stop and be verified before proceeding — signed by name
  • Review the PM task list on any asset added to the bad-actor list — often the wrong task is being done at the wrong frequency
  • Retire PM tasks that no longer address a real failure mode; add ones for failure modes emerging in the field
Good looks like
  • Standard Job Plans on critical assets carry measurable acceptance criteria — not just a list of steps
  • PM task content on critical assets is traceable to a documented failure mode
  • Precision-maintenance evidence exists on every critical asset repair — torque figures, alignment sheets, vibration baselines
Bad looks like
  • Standard Job Plans are generic checklists — no way to prove precision maintenance actually happened
  • PM tasks copied from the OEM manual and never questioned, even after multiple in-service failures
  • Reliability sets strategy on a whiteboard but never lands it in the CMMS — nothing changes on the floor
P3Schedule Work
Open ↓

The Reliability Engineer defends condition-based work against being casually deferred. Every CBM work order has a P-F interval; deferring beyond it converts predictive work into reactive work and breaks the whole strategy. The Reliability Engineer must make this cost visible to the Scheduler and the Supervisor.

What this role does
  • Flag any condition-based work order about to breach its P-F interval to the Scheduler with the consequence in plain language
  • Attend the weekly scheduling meeting when critical-asset or bad-actor work is on the candidate list
  • Publish CBM breach rate weekly alongside schedule compliance so both KPIs are visible together
  • Escalate any condition-based order deferred more than once — the strategy is broken, not just the schedule
  • Confirm strategy-change work orders (PM adds, retires, frequency changes) actually get scheduled within a month of approval
Good looks like
  • CBM breach rate is reported weekly and stays near zero
  • Strategy-change work lands on the schedule inside a month of the decision, not the next quarter
  • The Scheduler treats condition-based work as non-negotiable, not 'nice to have if we have time'
Bad looks like
  • CBM work rolls week to week until it becomes reactive — strategy invisible, effort wasted
  • Approved strategy changes sit in the CI register for months waiting to be scheduled
  • P-F breaches never quantified, so no one knows the strategy is failing until an asset trips
P4Execute & Record
Open ↓

The Reliability Engineer sets and audits the precision maintenance standard at the asset itself. Standards on paper mean nothing if the actual work on the floor doesn't meet them. Their job in this phase is to coach through presence — attend the jobs, audit the readings, baseline the repairs — not to police from a spreadsheet.

What this role does
  • Attend selected precision jobs and audit against the Standard Job Plan acceptance criteria in real time
  • Baseline vibration, thermography or oil analysis after every major repair before the asset is released to production
  • Coach the tradesperson at the job face when acceptance criteria are missed — not weeks later in a report
  • Review handback statements on critical assets — is the condition claim actually supported by evidence?
  • Sample-audit permits and isolations on critical-asset work — the safety envelope is a reliability issue too
Good looks like
  • Post-repair baselines exist on every critical repair; infant-mortality failures are trending down
  • Precision maintenance evidence is captured against the work order, not in a private folder
  • Coaching happens at the asset — tradespeople know the Reliability Engineer as a person, not a report
Bad looks like
  • 'Return to service' with no baseline — no way to detect early re-failure until it re-fails
  • Reliability Engineer never seen on the floor — precision maintenance stays a paper concept
  • Audits done from the office reading history text — real workmanship never inspected
P5Close-out & Analysis
Open ↓

The failure-code catalogue and the history entered on close-out are the raw material of every reliability insight the site will ever produce. If the Reliability Engineer does not police this data quality, every Pareto chart, MTBF calculation and RCM decision downstream is built on garbage.

What this role does
  • Sample-audit close-outs weekly for failure-code accuracy and history depth — publish the audit result
  • Retire dead failure codes; add missing ones; keep the catalogue small and used, not encyclopaedic and ignored
  • Review all 'Other' or 'Unknown' failure-code entries monthly and re-classify — do not let the bucket grow
  • Feed close-out quality trends back to the Supervisor for coaching — this is not just a reliability problem
  • Trigger a formal root-cause investigation on any failure that damages a critical asset, without waiting for it to reach the bad-actor list
Good looks like
  • Failure-code catalogue has under 100 active codes and over 90% coverage in real use
  • History quality on critical assets is audited monthly and trends up quarter over quarter
  • Every critical-asset failure has an investigation record — even the ones that looked minor at the time
Bad looks like
  • 1,000-code catalogue that everyone bypasses with 'Other' — no analysis possible
  • History quality accepted as 'as good as it gets' — the site regresses to reactive quietly
  • Investigations only happen after the second or third failure — the first one is lost forever
P6Metrics & Continuous Improvement
Open ↓

The Reliability Engineer owns the full continuous improvement loop — from Pareto through investigation through strategy change through verification. The measure of success is not how many reports were produced but how many failure modes have been permanently eliminated. If the same bad actors appear quarter after quarter, the loop is not closing.

What this role does
  • Run the monthly bad-actor review; assign a named investigator per top item, with a due date
  • Publish weekly and monthly KPI packs with trend commentary — say what the numbers actually mean, not just what they are
  • Push approved strategy changes to the Planner and verify implementation in the CMMS within 30 days
  • Escalate any CI action overdue more than 60 days to the Maintenance Manager — do not let the register rot
  • Verify strategy-change effectiveness at 6 and 12 months — has the failure mode actually reduced, or did we just move it?
  • Run at least one full RCM or FMEA review per quarter on a critical asset class
Good looks like
  • Every bad actor has an investigation, a decision, an implementation date and a verification date
  • Strategy changes show up as measurable failure-mode reductions 6-12 months later — the loop is closed
  • The CI register is short and active, not long and stale
Bad looks like
  • Reports produced, no actions closed — 'analysis theatre' becomes the whole function
  • KPI definitions redefined quarterly so no long-term trend is possible
  • Investigations end with 'need more data' every time — the same bad actors return year after year
Good looks like
  • Condition-based work is raised inside the P-F interval and scheduled in time to prevent functional failure
  • Repeat failures on the same asset are surfaced within one week, not one quarter
  • The criticality register is current — no asset in production without a rating
Bad looks like
  • Condition-monitoring alarms sit unread; repeat failures are re-raised as fresh reactive work each cycle
  • Reliability Engineer becomes a reporting function — publishes trends but never raises the work
  • Criticality register hasn't been touched in two years — new assets get a default rating and nobody notices
02

Learn

Courses, guides and reading to build the capability.

Competencies

Engineering degree plus reliability certification (CMRP / CRE or equivalent)RCM, FMEA and root-cause analysis methodologiesData analysis — Pareto, Weibull, trendingFacilitation — able to run investigations with cross-functional teams
03

Implement

Tools you run daily, weekly, monthly.

Cadenced checklists

Tick items as you go. Progress is saved in your browser — reset any time.

Daily
0 / 3
Weekly
0 / 4
Monthly
0 / 5
Quarterly
0 / 3
Yearly
0 / 3
04

Improve

Measure how you're doing and where to sharpen next.

KPIs to own

MTBF (top assets)PM effectiveness% failure modes with strategyRCFA close-out rateBad-actor reduction

If you see this, fix it

  • Reports that produce actions no one closes — analysis theatre
  • Redefining KPIs quarterly — no trend possible
  • Investigations that end with 'need more data' every time