Case study · Program management · Vendor management · Reliability
Pump Fleet Conversion
A fleet of process-vacuum pumps kept dying in infancy. I ran the program that converted 98 of them to a modular design — they now last about twice as long — and built the business case that paid for it.
- Work
- Modular Pump Conversion Program
- Role
- Program lead — equipment reliability engineer
- Timeline
- 12-month cross-functional rollout
- Vendors
- 6 suppliers evaluated and managed
- Outcome
- ≈$330K/yr savings, ≈$1.2M cumulative; ≈1.9× MTBF
- Breakeven
- ≈2.8 years on conversion capital
The fleet’s dry pumps were failing early — corrosion-driven, well before their design life. Each unplanned failure risks scrapping everything the process tool is working on and takes the tool down for hours. The failure data pointed to a modular pump architecture that isolates the corrosion-prone stage into a replaceable block.
My job was everything between that insight and 98 converted pumps: the failure analysis, the capital justification, selecting among six vendors, the FMEA that gated each rollout phase, PO and spares tracking, and the warranty program that kept suppliers accountable afterward.
01.1The reliability problem
New pumps should fail like old machines: rarely, and only at the end. These failed at the beginning — the left wall of the bathtub curve. Early-life failure is a program problem, not a maintenance problem: no PM schedule can fix a unit that dies before its first PM.
01.2The money chart
The boxplot below is the entire program in one image — the same paneled-boxplot view the real analysis used. Before: median life around 8,000 hours. After conversion: around 15,600. Every hour in between is a failure that didn’t happen and product that didn’t get scrapped.
01.3The business case
Reliability arguments don’t fund programs; cashflow arguments do. The proposal priced each conversion at roughly $17K per pump and weighed it against avoided failures, avoided scrap risk, and reduced rebuild spend. It cleared at about a 2.8-year breakeven and six-figure annual savings — then beat the estimate in production.
| Line item | Value | Basis |
|---|---|---|
| Conversion cost per pump | ~$17K | hardware + labor + qualification |
| Initial program capital | ~$200K | first conversion phase |
| Annual savings at full fleet | ~$330K/yr | avoided failures, rebuilds, scrap-risk hours |
| Cumulative to date | ≈$1.2M | tracked against baseline failure rate |
| Breakeven | ≈2.8 yr | conservative case as approved |
01.4An FMEA-gated rollout
Ninety-eight pumps is too many to convert on faith. A failure-mode and effects analysis scored each risk at three stages — current state, at install, and post-implementation — and each rollout phase only proceeded when its high-RPN items had owners and closure evidence.
| Failure mode | Effect | S | O | D | RPN | Mitigation | Post-RPN |
|---|---|---|---|---|---|---|---|
| Corrosive attack on exposed stage | Early-life pump failure | 9 | 7 | 4 | 252 | Modular corrosion-resistant block | 72 |
| Wrong block variant installed | Repeat failure, lost conversion | 8 | 4 | 3 | 96 | Variant matrix by tool type; install checklist | 32 |
| Spares gap during phase-in | Extended downtime on failure | 7 | 5 | 2 | 70 | Dual-sourced spares; rebuild loop sized to fleet | 28 |
| Vendor rebuild quality drift | Reduced converted-unit life | 7 | 4 | 4 | 112 | Warranty tracking + quarterly KPI reviews | 42 |
| Install-window overrun | Production schedule impact | 6 | 3 | 2 | 36 | Phased by tool availability; standard work | 18 |
01.5What I’d do differently
Start the spares and rebuild-loop conversation with vendors before the pilot, not after — the phase-in spares gap was the closest the program came to a schedule slip. And publish the vendor KPI scorecard from month one; suppliers behave differently when they know the chart exists. That scorecard became its own tool, and it lives in Fleet Analytics.
The discipline I hold hardest came from a call I got wrong. A long-down event: we swapped components, exchanged the pump, exchanged it again — and the pump was never the problem. A recipe upstream was flowing more gas than the pump could handle, and my team wasn’t aware of it, because we’re a subsystem and nobody tells the pump crew about flow changes. The fix wasn’t hardware; it was a meeting with the adjacent process team to lessen the load. Next time — and every time since — I ask better questions first and work to understand the issue externally from my direct system.