Case study · Reliability analytics · KPI design · Data fluency
Fleet Analytics
The reliability dashboard I built on a 111,000-row maintenance dataset — recreated here page for page, with synthetic data. It started as a year-end slide for my manager and ended up changing what the team worked on.
- Work
- Pump Fleet Reliability Dashboard
- Role
- Designer, builder, and chief user
- Dataset
- 111k maintenance records, fleet-wide
- Pages
- Year comparison · fleet health · failures · vendor KPIs · maintenance mix
- Audience
- Reliability team, vendor QBRs, management reviews
I actually created this dashboard at the end of the year. I wanted to make a PowerPoint for my manager on what we needed to focus on in the upcoming January, and while building it I noticed how we’d been picking our work: the team focused on whichever tool groups showed up most often when we reviewed the previous day’s tickets. Very knee-jerk-reaction culture.
So I stopped counting tickets. I normalized by fleet size for each process group and looked at failure rates and median run life instead — and that view unveiled the real pain points: wafer impact, rebuild costs from frequent failures, downtime. They were not the toolsets we’d been chasing. I highlighted the most painful pump models and process groups for each of those KPIs and put them in front of the team, so we could tackle the real issues. The dashboard changed what the team worked on.
The recreation below mirrors the production tool page for page — the dark canvas, the page tabs along the bottom, the panel layout, the goal lines, even the table stubs. Vendors became letters, units became numbers, and every value was re-rolled synthetically; the structure is the real thing.
| root-cause note | wl qty |
|---|---|
| fault with alarm, wafers at risk | 1 |
| pump faulted mid-process | 1 |
| exchange from tool PM window | 2 |
| stranded wafers on transfer | 2 |
| no alarm present at fault | 1 |
| verbal pass-down, day shift | 1 |
| high temp during process | 2 |
| production interruption, batch | 3 |
| toolset | unit | target mtbf (wk) | live health |
|---|---|---|---|
| TOOL-A01 | UNIT-01 | 47 | Healthy |
| TOOL-B02 | UNIT-02 | 21 | Healthy |
| TOOL-C03 | UNIT-03 | 54 | Healthy |
| TOOL-D04 | UNIT-04 | 97 | Healthy |
| TOOL-E05 | UNIT-05 | 6 | Healthy |
| TOOL-A06 | UNIT-06 | 14 | Healthy |
| TOOL-B07 | UNIT-07 | 137 | Healthy |
| TOOL-C08 | UNIT-08 | 77 | Healthy |
| TOOL-D09 | UNIT-09 | 14 | Healthy |
| TOOL-E10 | UNIT-10 | 47 | Healthy |
| TOOL-A11 | UNIT-11 | 85 | Healthy |
| TOOL-B12 | UNIT-12 | 6 | Healthy |
| TOOL-C13 | UNIT-13 | 143 | Healthy |
| TOOL-D14 | UNIT-14 | 77 | Healthy |
| exchange date | unit | root-cause detail | runtime (h) |
|---|---|---|---|
| 2025-09-23 | UNIT-03 | bearing wear at inlet stage | 20,683 |
| 2025-04-16 | UNIT-22 | sensor drift vs manual read | 10,693 |
| 2025-08-19 | UNIT-30 | preventive exchange | 12,248 |
| 2025-05-08 | UNIT-26 | overtemp trip during ramp | 8,398 |
| 2025-02-19 | UNIT-10 | preventive exchange | 11,655 |
| 2025-12-15 | UNIT-10 | seal leak, process side | 4,268 |
| 2025-09-14 | UNIT-06 | motor fault, phase loss | 5,380 |
| 2025-08-14 | UNIT-02 | seal leak, process side | 18,687 |
| 2025-10-26 | UNIT-29 | motor fault, phase loss | 11,545 |
| date | unit | model | runtime (h) | exchange | in warranty |
|---|---|---|---|---|---|
| 2025-12-20 | UNIT-19 | MODEL γ | 21,448 | vendor exchange | false |
| 2025-12-29 | UNIT-28 | MODEL δ | 57,175 | scheduled swap | false |
| 2025-12-04 | UNIT-16 | MODEL δ | 13,158 | vendor exchange | true |
| 2025-12-07 | UNIT-15 | MODEL β | 19,408 | scheduled swap | true |
| 2025-12-04 | UNIT-01 | MODEL β | 75,335 | vendor exchange | false |
| 2025-12-20 | UNIT-01 | MODEL α | 32,256 | scheduled swap | true |
| 2025-12-21 | UNIT-09 | MODEL γ | 83,941 | scheduled swap | false |
| 2025-12-04 | UNIT-04 | MODEL δ | 66,078 | scheduled swap | false |
| unit | exchange type | reason |
|---|---|---|
| UNIT-05 | bm | noise/vibration |
| UNIT-01 | bm | overtemp |
| UNIT-14 | bm | overtemp |
| UNIT-01 | pm | overtemp |
| UNIT-10 | bm | drift |
| UNIT-09 | pm | overtemp |
| UNIT-02 | pm | seal leak |
02.1What each page is for
Year comparison answers “are we better than last year?” with wafer-loss events against a weekly goal and stop-loss hours against a contract ceiling. Fleet health bins every unit against a bathtub-curve guideline so “watch” units get their PM pulled in before the wear-out wall. Pump fail is where the treemaps live — one look separates a vendor problem from a model problem. Vendor KPI pages put stop-loss hours, wafer scrap, and warranty failures on one sheet per supplier; they became the agenda of every quarterly business review. BM weekly is the fleet’s pulse: the planned-to-breakdown ratio, and a Pareto that turns a bad week into a root-cause assignment.
02.2How I read a fleet
Same method every time a new pump model or configuration enters the fleet:
- Wafer quality data first. If the change could touch product, nothing else matters yet — that check comes before any reliability claim.
- One unit, then phases. Prove it on a single pump, scale in phases. Never the whole fleet on faith.
- Routine trend monitoring. Proactive, scheduled looks at the trends — not waiting for a ticket to tell me something moved.
- Verify with run life. Track run life against the previous model, and only call it a real improvement once there are enough data points to be sure it isn’t noise.
02.3When the vendor says it’s our process
The vendor KPI pages earn their keep when a supplier pushes back. I provide the basis for my opinion — the failure trend, normalized, with run life attached — and then we collaborate with the rebuild data they hold to find the correlation. Many times they say it’s our process, that we run the pumps too harsh. Sometimes that’s right. But when the rebuild teardown shows there isn’t build-up — no process signature at all — the issue relates to something specific on their side, and that’s what we ask them to work on. The teardown settles the argument either way.
02.4Design notes
Three rules made this tool trusted. Every view answers a question someone actually asked; anything “interesting” but unactionable was cut. Every chart keeps the same vocabulary — a unit, a failure, an exchange mean the same thing on every page, which is harder than it sounds across 111k rows of free-text maintenance logs. And the dashboard never editorializes: the treemap doesn’t say a vendor is bad, it says early-life failures concentrate in one supplier’s install base — with a warranty clause attached.