Selection guide / Updated August 2026 / 8 min read

Build or buy industrial analytics, written for an asset owner

By Bhavik Modi / CEO & Co-Founder LinkedIn

Instrumentation and process engineering, electrolyser technology and machine learning, with experience at Siemens, L&T, Mitsubishi and Newtrace.

An energy asset owner is answering a different question from a software buyer. Acquisition and storage are buildable by a competent team. The domain layer needs process engineering rather than software engineering, and the cost that decides the case is ownership in year three, once the plant has changed and the person who built it has moved on.

Build vs buyProcurementAnalyticsAsset management

Why the software framing does not fit

The familiar framing for this decision is feature parity, time to market and engineering headcount. An asset owner is answering a different question, because what is being built is an internal capability rather than a product, and it has to survive a control system upgrade and a change of staff.

The relevant question is narrower: which parts of an analytics capability can this organisation own competently for ten years, and which parts will quietly rot when the person who built them moves on.

That reframing matters because it changes the answer. Most build decisions in this sector are made on capability at the start and regretted on ownership in year three.

The four layers, and which are genuinely buildable

Acquisition and storage. Getting data off the plant and into somewhere queryable. Buildable by a competent team, well-served by open tooling, and routinely underestimated by a factor most teams find uncomfortable once tag mapping starts.

Contextualisation. Knowing what each tag means, its units, its sign convention, which asset it belongs to and how those assets relate. Unglamorous, plant-specific, and a common place for a schedule to disappear.

Domain modelling. The physics, the degradation mechanisms, the attribution logic. This needs process and control engineering rather than software engineering, and it is where build decisions most often fail.

Presentation and workflow. Getting a finding to the person who acts on it, in a form they will act on. Deceptively hard, because the requirement is behavioural rather than technical, and dashboards are the graveyard of this layer.

Data engineering: buildable, and underestimated

The connection itself is a week. Making sense of what comes down it is not. A plant of any age has thousands of tags with inconsistent naming, missing or wrong units, undocumented sign conventions, and duplicates left by successive control system upgrades. Establishing what each signal actually represents is where the schedule goes.

Historian compression is the second surprise. Many historians discard points judged uninformative, and a signal stored as a five-minute average cannot be un-averaged. Discovering this after a year of collection is common and expensive.

Neither of those is a reason not to build this layer. They are reasons to scope it honestly, and to note that the work is a one-off asset the plant keeps regardless of what happens to any analytics vendor. That is the strongest argument for building here: the contextualised data model is the thing with lasting value, and it should not be locked inside somebody else's product.

The interface question comes first and is separate: what can be read, through which path, and what enforces it. That is covered in reading plant data without modifying the control system.

Domain modelling: the layer that decides it

This is where a build either works or produces a dashboard nobody opens. Detecting that a value has changed is straightforward. Explaining why, in terms an engineer will accept, requires modelling the mechanism, and modelling the mechanism requires someone who knows it.

Plants that succeed here have a specific characteristic: a process or control engineer who understands both the equipment and the analysis, and who stays. Where that person exists, the build case is real. Where the plan is to hire data scientists and have them learn electrochemistry, it usually is not.

There is a middle failure mode worth naming. A team builds statistical anomaly detection, it works in testing, and it produces alerts nobody can interpret because there is no mechanism attached. The system is then blamed for false positives when the actual gap is attribution. The distinction is set out in physics-driven AI against generic industrial analytics.

Honest test for this layer: can the team say, today, what physical relationships they would encode and what would falsify each one. If the answer is that they would start with the data and see what emerges, the domain layer is not being built, it is being hoped for.

The maintenance burden nobody costs

Plants change. Setpoints move, equipment is replaced, operating strategy shifts, control systems are upgraded and tag names change with them. Every one of those breaks something in an analytics stack, and someone has to fix it.

In year one the builder is present and fixes things in an afternoon. In year three the builder has moved to another role or another company, the documentation describes an earlier version, and the system produces output nobody quite trusts. This is a common ending for a built industrial analytics capability, and it is organisational rather than technical.

The cost to carry is not the build cost. It is a standing allocation of engineering time, indefinitely, plus a succession plan for the person who holds the knowledge. Costing that honestly at the start changes many build decisions, which is presumably why it is so often omitted.

A bought system has the mirror-image risk, which is worth stating fairly: the vendor changes direction, is acquired, or ends the product, and the plant is left with a dependency it did not choose. The mitigation is the same in both directions, which is owning the data layer and the contextualisation regardless.

Splits that work

The arrangement that holds up is a boundary drawn along the skills. The plant owns acquisition, storage and contextualisation, because that is data engineering and it is the asset with lasting value. The vendor owns the domain models and the attribution, because that is where specialised knowledge is amortised across more than one site.

This has a practical test attached: if the vendor relationship ended tomorrow, does the plant still have its data, its tag model and its history in a form another party could use. If yes, the dependency is manageable. If no, the split was not really made.

The arrangement that does not work is the plant building the domain layer while buying the data layer. It inverts the skills, puts the specialised work where the specialism is thinnest, and leaves the commodity work with a supplier.

Fleet size moves the line. One asset rarely justifies building the domain layer, because the modelling effort is amortised over nothing. Twenty similar assets might, because the same model is reused and the standing engineering cost is spread. Somewhere between those the calculation flips, and it is worth doing rather than assuming.

Questions to answer before deciding

Who, by name, owns this in year three, and what happens when they leave.

What is the standing annual engineering allocation, and has it been agreed by whoever controls that budget.

Which physical relationships will the domain layer encode, and who knows them.

If the vendor relationship ended, or the internal builder left, what remains usable.

How many assets does this apply to now, and how many in five years.

And what decision is meant to change as a result. A capability that cannot name the decision it improves is a project looking for a justification, whichever way it is sourced.

Questions teams ask

Frequently asked questions

Should an asset owner build or buy industrial analytics?

Usually a split. The data acquisition, storage and contextualisation layer is worth owning because it has lasting value and is portable between vendors. The domain modelling layer is where builds most often fail, because it needs process and control engineering rather than software engineering.

What is most underestimated in a build?

Tag mapping and contextualisation. Establishing what each of several thousand signals actually represents, with inconsistent naming, missing units and undocumented sign conventions, consumes more schedule than the modelling. Historian compression is the second surprise, and it is discovered late.

Why do built systems fail in year three?

Because plants change and the person who built the system has moved on. Setpoints move, equipment is replaced, tag names change with a control system upgrade, and a system nobody owns degrades quietly until it is wrong often enough to be ignored. It is an organisational failure rather than a technical one.

Does fleet size change the answer?

Materially. Domain modelling effort is amortised across assets, so one plant rarely justifies building that layer while twenty similar assets might. The data layer is worth owning at any size because it is the part with lasting value.

What is the test for whether a split was really made?

If the vendor relationship ended tomorrow, does the plant still hold its data, its tag model and its history in a form another party could use. If yes the dependency is manageable. If no, the plant bought a system rather than making a split.

Which split does not work?

Building the domain layer while buying the data layer. It puts the specialised work where the in-house specialism is thinnest and outsources the commodity work, which is the opposite of how the skills are usually distributed.