Manufacturing Intelligence
On the gap between production data and operational decisions — traceability, bottleneck analysis, and the intelligence layer factories are missing.
Writing
Essays on manufacturing intelligence, industrial AI, and connected vehicles — grounded in how these systems actually work in the field.
On the gap between production data and operational decisions — traceability, bottleneck analysis, and the intelligence layer factories are missing.
Condition monitoring, predictive maintenance, and the hard problem of turning signals into actions people will actually take.
Telematics hardware, battery analytics, and the intelligence layer for EV fleets and OEMs across the full vehicle lifecycle.
Industrial AI
Why most predictive maintenance implementations fail at the last mile — and what the decision architecture looks like when they don't.
The model fires. The alert appears. Somebody sees it — maybe. And then nothing happens.
This is not a failure mode. It is the default outcome for most predictive maintenance implementations. The model is usually fine. The problem sits at the handoff from prediction to decision.
Three questions stall action every time the alert appears: How confident is this prediction? Can the line afford to stop right now? What happens if we wait until the weekend shutdown? These are not bad questions. They are the right questions. The problem is the system does not answer them — it fires the alert and leaves the rest to the operator standing in front of a production line that is running.
When the system cannot answer those questions, inaction is the rational response. Not laziness. Rationality. The operator has no basis to stop production on an alert from a system they cannot interrogate. So they wait. And the failure happens anyway, at a time and cost the system was specifically supposed to prevent.
What a real decision architecture looks like is not a binary alert. It is a complete answer. Likelihood of failure given current operating conditions. Estimated time window before the failure becomes critical — not just 'soon,' but 'within the next 72 hours at current load.' Business impact of waiting: what does an unplanned stoppage on this asset cost versus a planned two-hour maintenance window next shift? And a specific recommended action, not a score.
The system should make it easier to act than to ignore.
The wrong response to implementation failure is to generate more alerts. Alert volume is not intelligence. It is noise with a confidence score attached. An operator who has learned to dismiss alerts because most of them do not require action has not failed to adopt the technology. They have correctly adapted to a system that gives them no way to distinguish signal from noise.
The question that matters is not whether the model predicts correctly. It is whether the prediction creates a decision. Build the system so the decision is the output, not the alert.
Manufacturing Intelligence
On the gap between data collection and operational intelligence — and what it takes to actually close it.
Every plant I have worked with collects production data. Event logs, machine states, cycle times, yield results, defect codes. The data exists. Often there is too much of it.
What does not exist is the answer to the question the shift manager asked this morning: which station is actually constraining throughput right now?
This is the core problem with manufacturing data, and it has nothing to do with collection. The data is there. The intelligence layer — the structured analysis that turns raw events into answers to specific operational questions — is not.
Three questions come up consistently in every manufacturing operation, and they consistently go unanswered despite the data existing. The first: which station is actually constraining throughput? Not which station has the longest queue. Not which station has the highest alarm count. Which station, if it ran one cycle faster, would move more product through the line? Answering that requires bottleneck analysis — understanding the dependency structure of the line, not just the utilisation of individual stations.
The second: why does one shift consistently underperform another, and is it a machine issue, a process issue, or a people pattern? Shift comparison data exists. But comparing shift-level OEE numbers without decomposing them by station, by product type, by the actual production events that drove the delta tells you nothing actionable. The difference could be a machine that runs hotter in the evening. It could be parameter drift that correlates with crew changeover. You need the traceability, not the summary.
The third: when a defect surfaces at end-of-line, can it be traced back to where it originated? This is the hardest problem because it requires consistent ID propagation from raw material through every production event to the final product. Most plants have partial traceability. They can tell you which station inspected the unit. They cannot tell you which parameters were running on that station when the unit passed through it.
These questions go unanswered not because the data does not exist but because there is no intelligence layer between the raw production events and the people who need to act. The data sits in MES databases, in historian tables, in alarm logs. It is not connected. It is not structured to answer questions. It is not indexed to the physical objects moving through the plant.
The architecture that closes this gap is not another dashboard. It is a connected analysis layer — from production event to cause to business impact to recommended action, traceable from the individual cell through to the vehicle it will go into.
Visibility is a prerequisite, not the outcome. The outcome is the decision that comes after someone can actually answer the question.
Industrial AI
Why scheduled maintenance is structurally incapable of preventing the failures that matter most.
Scheduled maintenance is built on a false premise: that machines degrade by time. They do not. They degrade by load, by operating condition, by the accumulated effect of duty cycles that vary every shift.
A pump that has run at 80% load for six months has experienced a fundamentally different wear profile than a pump that has run at 40% load for the same period. Servicing both at the same interval — because the calendar says it is time — is not maintenance. It is a ritual.
The failure modes of scheduled maintenance are well understood but rarely named clearly. It over-services healthy equipment, which introduces new failure modes through unnecessary inspection and reassembly. It misses random failures that fall between service intervals — precisely the failures that cause unplanned downtime. And it creates a false confidence that a maintained asset is a safe asset, which it is, until the failure that does not care about the service record.
Reactive maintenance is worse. Wait for the failure, pay emergency costs, take unplanned downtime, and run a root cause analysis that usually finds a degradation pattern that was detectable weeks earlier.
Condition-based monitoring changes the logic. Continuous signals — vibration, current draw, temperature — give you a degradation profile, not a calendar date. The physics is different for every asset type: the vibration signature of a failing bearing in a gearbox is different from the thermal pattern of a motor running out of spec, which is different from the current signature of a conveyor under increased load. But the principle is the same. The machine tells you what it knows about its own condition, continuously, and the analytics layer translates that into a time window and a specific recommended action.
The business case is simple: one prevented failure on a critical production line pays for years of monitoring investment. The barrier is not cost. It is inertia — the accumulated weight of how maintenance has always been done, and the trust gap between a health score on a screen and the judgment of the engineer who has maintained that machine for fifteen years.
The question is not whether condition monitoring pays for itself. It is why the industry waited this long to ask.