Sohaib Aldayem

Case study · Reliability analytics · KPI design · Data fluency

Fleet Analytics

The reliability dashboard I built on a 111,000-row maintenance dataset — recreated here page for page, with synthetic data. It started as a year-end slide for my manager and ended up changing what the team worked on.

Work
Pump Fleet Reliability Dashboard
Role
Designer, builder, and chief user
Dataset
111k maintenance records, fleet-wide
Pages
Year comparison · fleet health · failures · vendor KPIs · maintenance mix
Audience
Reliability team, vendor QBRs, management reviews

I actually created this dashboard at the end of the year. I wanted to make a PowerPoint for my manager on what we needed to focus on in the upcoming January, and while building it I noticed how we’d been picking our work: the team focused on whichever tool groups showed up most often when we reviewed the previous day’s tickets. Very knee-jerk-reaction culture.

So I stopped counting tickets. I normalized by fleet size for each process group and looked at failure rates and median run life instead — and that view unveiled the real pain points: wafer impact, rebuild costs from frequent failures, downtime. They were not the toolsets we’d been chasing. I highlighted the most painful pump models and process groups for each of those KPIs and put them in front of the team, so we could tackle the real issues. The dashboard changed what the team worked on.

The recreation below mirrors the production tool page for page — the dark canvas, the page tabs along the bottom, the panel layout, the goal lines, even the table stubs. Vendors became letters, units became numbers, and every value was re-rolled synthetically; the structure is the real thing.


pump dashboard · recreation · synthetic data
Wafer Loss
020406080weekly wafer goal < 135254263272284293308031232333334353653738139340JUNJULAUGSEPWAFER-LOSS EVENTS PER WORK WEEK — colored by month · synthetic
Pump Change YoY
0204060801005353Jan4654Feb6937Mar5839Apr7466May5498Jun6879Jul6871Aug5564Sep6467Oct4656Nov5462DecYEAR 1YEAR 2PUMP CHANGES YEAR-OVER-YEAR BY MONTH — synthetic counts
MRO Table for WL
root-cause notewl qty
fault with alarm, wafers at risk1
pump faulted mid-process1
exchange from tool PM window2
stranded wafers on transfer2
no alarm present at fault1
verbal pass-down, day shift1
high temp during process2
production interruption, batch3
Pump & Acc Yearly Stop-Loss Comparison
0115230345460goal ≤ 67 hours1591317193212052524645918829220198333718841454953PUMP + ACCESSORY STOP-LOSS HOURS PER WORK WEEK — colored by month · synthetic
Acc Yearly Stop-Loss Comparison
04080120160goal ≤ 56 hoursJanFebMarAprMayJunJulAugSepOctNovDecACCESSORY STOP-LOSS HOURS BY MONTH — synthetic
1,768 of 111k rows270+ columns0 marked
Recreation · synthetic dataPage-for-page recreation of the production dashboard I built and ran at work — same pages, same panel layout, same chart grammar. Every number is invented; vendor names are letters; units and tool groups are renumbered.

02.1What each page is for

Year comparison answers “are we better than last year?” with wafer-loss events against a weekly goal and stop-loss hours against a contract ceiling. Fleet health bins every unit against a bathtub-curve guideline so “watch” units get their PM pulled in before the wear-out wall. Pump fail is where the treemaps live — one look separates a vendor problem from a model problem. Vendor KPI pages put stop-loss hours, wafer scrap, and warranty failures on one sheet per supplier; they became the agenda of every quarterly business review. BM weekly is the fleet’s pulse: the planned-to-breakdown ratio, and a Pareto that turns a bad week into a root-cause assignment.

02.2How I read a fleet

Same method every time a new pump model or configuration enters the fleet:

  1. Wafer quality data first. If the change could touch product, nothing else matters yet — that check comes before any reliability claim.
  2. One unit, then phases. Prove it on a single pump, scale in phases. Never the whole fleet on faith.
  3. Routine trend monitoring. Proactive, scheduled looks at the trends — not waiting for a ticket to tell me something moved.
  4. Verify with run life. Track run life against the previous model, and only call it a real improvement once there are enough data points to be sure it isn’t noise.

02.3When the vendor says it’s our process

The vendor KPI pages earn their keep when a supplier pushes back. I provide the basis for my opinion — the failure trend, normalized, with run life attached — and then we collaborate with the rebuild data they hold to find the correlation. Many times they say it’s our process, that we run the pumps too harsh. Sometimes that’s right. But when the rebuild teardown shows there isn’t build-up — no process signature at all — the issue relates to something specific on their side, and that’s what we ask them to work on. The teardown settles the argument either way.

02.4Design notes

Three rules made this tool trusted. Every view answers a question someone actually asked; anything “interesting” but unactionable was cut. Every chart keeps the same vocabulary — a unit, a failure, an exchange mean the same thing on every page, which is harder than it sounds across 111k rows of free-text maintenance logs. And the dashboard never editorializes: the treemap doesn’t say a vendor is bad, it says early-life failures concentrate in one supplier’s install base — with a warranty clause attached.