back to overview

Why root cause analysis takes weeks

Root cause analysis

CONTENT

  • The hidden friction before analysis begins
  • The technical mechanism that causes delay
  • What happens to the quality of the analysis
  • What changes when context is already structured

A multi-site manufacturer asks a simple question: why does the same recurring fault appear at four of its eight plants, but not the other four? The question is strategically important. The answer requires comparing process data, machine states, production order patterns, and maintenance histories across sites that each use slightly different naming conventions, historian configurations, and downtime categorization schemes. The investigation starts with the question and spends its first two weeks on data logistics. That pattern repeats every time.

The hidden friction before analysis begins

Root cause analysis has a cost that rarely appears in project plans or post-incident reviews: the cost of assembling the analytical context before the actual investigation can begin. In most industrial organizations, the relevant data for a complex investigation is distributed across a historian, a MES system, a maintenance platform, alarm logs from SCADA, and sometimes operator shift notes. Each of those systems was implemented independently, optimized for its own purpose, and designed without the assumption that its data would ever need to be joined with the others.

The result is that every RCA begins with an implicit sub-project: retrieve data from each source, align timestamps, reconcile naming differences, agree on which records from which system reflect the same underlying event, and then build a provisional timeline that the team can share. That phase is not analysis. It is reconstruction. And it is so normalized in most organizations that it is no longer recognized as the bottleneck it actually is. Hypothetically, between thirty and fifty percent of the total elapsed time in a typical multi-system RCA disappears before the first hypothesis about the actual cause is formed.

The technical mechanism that causes delay

The delay is a structural consequence of how industrial data is typically organized. Each system stores data in its own schema, with its own identifiers and its own event model. The historian records tag values against timestamps. The MES records production orders against machines and time windows. The maintenance system records work orders against assets and completion timestamps. None of those systems shares a common event model. None of them knows about the others.

When those datasets need to be combined, the join logic has to be constructed manually. An engineer decides that machine tag "L4_CNV_Motor_A" in the historian corresponds to asset "Line 4 Conveyor" in the maintenance system, and that both should be aligned to production order "PO-20240318-04" from the MES. That alignment may be obvious to someone who knows the plant. It is not obvious to the data. And across sites, it becomes exponentially harder, because each plant has made its own naming decisions over years of independent operation.

What happens to the quality of the analysis

The reconstruction phase does more than consume time. It degrades the quality of the conclusions. When analysts spend their first week aligning datasets, they inevitably make pragmatic choices: use data from the last four days because that is what is easy to retrieve, focus on the tags that are available in all systems rather than the ones that would be most informative, accept approximate timestamp alignment because exact synchronization takes too long. Each of those choices introduces interpretive risk that compounds across the investigation.

Teams then fill the gaps with expertise and intuition. Experienced engineers compensate for data limitations with process knowledge. That expertise is genuinely valuable. But it makes the analysis less reproducible. A different team on a different site, facing the same underlying failure mechanism, may reach a different conclusion because they had different data available and different experience to fill the gaps. The organization does not learn the same lesson twice. It learns a locally adapted version of it, shaped as much by data availability as by the actual physics of the failure.

What changes when context is already structured

Root cause analysis becomes structurally faster when the context that analysts currently reconstruct is part of the data architecture from the beginning. That means each machine has a consistent identity across systems. Each production event knows which asset it belongs to. Each alarm and maintenance intervention is linked to the same operational context. When a deviation occurs, an engineer starts from an event that already carries its surrounding context: the asset, the production order, the process state, the recent interventions.

The investigation can then focus on interpretation rather than assembly. Cross-site comparisons become possible because the underlying data describes the same reality in the same terms. The four plants with the recurring fault and the four without can be compared directly, because their data shares a common model. The question shifts from "where can I find this information?" to "what does this pattern mean?"

Capture structures industrial data around exactly that principle. Machines, process parameters, production orders, and operational events are connected within a shared context from the moment data is collected. When an investigation starts, the relevant event timeline is already navigable rather than assembled from scratch. Root cause analysis does not become simpler as a problem. It becomes faster and more consistent as a process, because the architecture stops forcing teams to rebuild reality before they can begin to understand it.

Want to get your reporting straight?