Predictive Maintenance Without the Buzzwords: What It Takes to Ship
Most predictive maintenance projects stall long before machine learning matters, on sensors, data plumbing, and trust. What shipping one actually involves, with the unglamorous parts left in.
Predictive maintenance might be the most oversold phrase in industrial software. The pitch is always the same: AI watches your machines and tells you a bearing will fail next Tuesday. The reality on a plant floor is messier, and the projects that fail usually fail for reasons that have nothing to do with algorithms. We shipped a predictive maintenance system for a manufacturer running production lines with aging equipment, and the honest breakdown of that project was roughly 20 percent modeling and 80 percent everything else. If you're evaluating predictive maintenance software development for your own plant, that ratio is the single most useful thing to know going in.
First Problem: Your Machines Weren't Built to Talk
The brochure version assumes clean sensor streams. The plant version is a 1998 CNC machine with no network port, a newer line that speaks Modbus, a packaging machine on a proprietary protocol, and one critical compressor with no instrumentation at all. Before any prediction happens, someone has to get vibration, temperature, current draw, and cycle counts off that mixed fleet reliably. On our project that meant retrofit vibration sensors on the oldest machines, an edge gateway per line translating three different protocols into one stream, and a month of fixing dropouts, clock drift between devices, and a sensor that reported perfect readings while hanging loose from its mount. Ten to twelve weeks for trustworthy data collection is a normal number for a mid-sized plant. Any proposal that skips this phase is a proposal to predict from garbage.
Second Problem: You Probably Don't Have Failure History
Machine learning for failure prediction wants examples of failures. Most plants, reasonably, try hard not to have failures, and the ones that happened live in a paper logbook as 'bearing changed, line 3, March.' So the first year of a serious deployment is mostly not deep learning. It's condition monitoring with statistical baselines: learn each machine's normal vibration signature and temperature envelope per shift and per product, then alert on drift. That sounds modest. It caught a misaligned coupling on our project in month two, weeks before it would have taken a line down, and unplanned downtime on that line ran to thousands of dollars per hour. Anomaly detection on a good baseline delivers most of the early value while your labeled failure history slowly accumulates for the fancier models later.
Third Problem: Nobody Trusts a Black Box on the Shop Floor
A maintenance supervisor with 20 years on these machines will not schedule a teardown because a dashboard turned amber. Nor should he. The system earned trust on our project because every alert shipped with its evidence: the vibration trend against the machine's own 90-day baseline, the recent alerts on that asset, and what was found the last time someone opened it up. We also tuned deliberately for fewer, better alerts. An early version alerted eagerly, operators started ignoring it within two weeks, and we cut alert volume by more than half before anyone acted on the system consistently. False alarms are not a cosmetic issue. They're how these systems die, quietly, while the licence keeps getting paid.
Start With Two Machines, Not the Whole Plant
The other consistent failure mode is scope. A plant has 60 machines, so the project plans to instrument 60 machines, and 18 months later there's a lot of hardware and no results anyone acts on. We did the opposite: picked the two assets with the worst downtime history, instrumented them properly, and ran the full loop, sensing, baselining, alerting, and maintenance response, on just those two for a quarter. That pilot surfaced every real-world problem cheaply. It told us which sensors earned their cost, and vibration did while two others didn't. It produced the first caught failure, which did more for internal buy-in than any projection deck. And it gave us a per-machine cost and effort number that made the rollout budget honest instead of hopeful. Scaling from a working two-machine loop to 20 machines is an operations exercise. Scaling from a slide deck to 60 machines is how these projects end up as expensive dashboards nobody opens.
What Shipping Actually Requires
- A sensor and connectivity audit per machine before anyone writes a proposal with a number in it
- An edge layer that buffers locally, because plant networks fail more often than plant managers admit
- Statistical baselines and drift alerts in phase one; save the ML roadmap for when failure labels exist
- Every alert accompanied by the evidence behind it, in terms a technician recognizes
- Alert volume tuned against operator behavior, measured by actions taken, not alerts sent
- Integration into the existing maintenance workflow and ERP, so predictions become work orders, not emails
The Payoff, Measured Honestly
Done in this order, the economics work. Our manufacturing deployment reached positive return inside the first year on avoided downtime alone, and the plant's planning got calmer in ways that never show up in an ROI slide: parts ordered before they were urgent, maintenance scheduled into planned stops instead of interrupting runs. None of it required believing a single buzzword. It required sensors that report the truth, plumbing that doesn't lose data, baselines that know each machine, and alerts a supervisor can argue with. If a vendor's pitch spends more time on the model architecture than on how they'll get clean data off your 1998 machines, ask harder questions, because the model is the part that was never going to be your problem. That's the whole trick, and it's harder and more valuable than the brochure version.
Have a project that needs this kind of thinking applied to it?