Technology guide / Updated August 2026 / 9 min read

What a digital twin of an electrolyser actually models

Engineering leader with experience at GE, Mitsubishi and Alstom, specialising in advanced controls, industrial process and multi-physics modelling, with R&D and patent-pending work behind the Yunify engine.

An account for someone who has to build or buy one: which states are estimated, which are measured, where the model is calibrated, and what it is expected to answer on a shift. A stack twin is mostly a state estimator, because several of the variables that decide the operating decision are not measured on any plant.

Digital twinElectrolyserModellingOperations

Why the generic twin taxonomy does not transfer

The three-way split that works for a power plant, between a geometric twin, a data-driven twin and a hybrid one, only half applies here. A geometric model of an electrolyser hall answers layout and outage-planning questions and has nothing to say about the stack. A purely data-driven twin can be built, and on a plant that runs a steady duty it is adequate. Neither is what people mean when they ask for a twin of an electrolyser.

What they mean is a model that says what the stack is doing right now, including the parts no instrument reports, and what it will be doing in a few weeks if the operating pattern continues. That is a state estimation problem sitting on top of an electrochemical model, and it is closer to a soft sensor than to a simulation.

The distinction matters commercially because the three are sold under one word. The wider taxonomy is set out in digital twins for power plants. The electrochemical case adds something specific to it.

What is measured, what is estimated, and what is neither

A typical electrolyser plant measures stack current, stack voltage, inlet and outlet temperature, system pressure, differential pressure across the separator, hydrogen content in the oxygen stream, feedwater conductivity, and the electrical and hydraulic state of the balance of plant. On a well-instrumented plant, individual cell voltages as well.

What it does not measure includes almost everything that decides a maintenance decision. Membrane or diaphragm condition. Catalyst activity. Current distribution across the electrode area. Electrolyte concentration in an alkaline system, on most plants, between manual samples. Flow distribution between cells. Local temperature inside the assembly rather than in the manifold.

The gap between those two lists is the twin's job. Everything in the second list has to be inferred from the first through a model that knows how they are related, and the quality of that inference is what separates a twin from a trend viewer.

A third category is worth naming: quantities that are neither measured nor estimable from what is available. Where the model cannot distinguish two mechanisms from the instrumentation present, the honest output is a ranked pair and the measurement that would separate them, not a single answer.

The electrochemical core

The backbone is the relationship between cell voltage and current density at a given temperature, pressure and electrolyte condition. Its shape separates the loss mechanisms: activation losses dominate at low current density, ohmic losses grow roughly with current, and transport limitations appear at the top of the range.

That shape is what makes attribution possible. A stack whose ohmic term has grown is a different problem from one whose activation term has grown, and the two are distinguishable from a curve in a way they are not distinguishable from a single operating point. A twin without a polarisation model is fitting a trend, not modelling a mechanism.

Thermal coupling comes next. Cell voltage above the thermoneutral point generates heat, that heat has to be removed, and the removal capacity depends on flow and ambient conditions. A stack running hotter than reference has a lower cell voltage and a higher degradation rate at the same time, which is why an efficiency improvement can be a life reduction wearing a disguise.

Gas transport closes the loop. Permeation through the separator does not scale with current, so the hydrogen fraction in the oxygen stream rises as load falls, and its trend at fixed load describes separator condition. That single relationship carries both a safety constraint and a condition indicator, and a twin has to keep them apart.

Why the model has to cross into balance of plant

The number an operator sees is plant-level specific energy. The stack is the largest term in it, but it is not the only one, and the terms around it are what move the figure without the stack changing. Rectifier efficiency falls away from its design point at deep part load. Circulation pumps, thermal management, gas drying and any compression consume energy that is close to constant regardless of production. A model that stops at the stack boundary cannot explain the number the plant is actually being judged on.

This is the same attribution problem as a plant reading more kWh per kg than its datasheet, and the twin is the mechanism for solving it continuously rather than during an investigation.

Crossing the boundary also catches the interaction cases. A circulation pump losing performance changes flow distribution, which changes temperature spread, which changes cell behaviour. Monitored separately, that chain gets diagnosed as a stack problem, and the stack is the expensive thing to replace.

Calibration, and the reference that cannot be recreated

A model with the right structure and the wrong parameters produces confident nonsense. The parameters come from a reference measurement taken while the plant is known-good: a polarisation curve at defined conditions, a documented performance test, and a thermal profile across the assembly.

None of those can be captured retrospectively. A curve taken after something started drifting describes the drifted state, and every later comparison is against it. This is the single cheapest thing to do at commissioning and the most expensive omission to discover afterwards.

Re-calibration is a separate discipline. A model tuned to a drifting instrument will confidently reproduce the drift, so a periodic controlled measurement is needed to distinguish model error from process change. The interval is a judgement, but the principle is not: a twin nobody re-references degrades quietly until it is wrong often enough to be ignored.

The questions it should answer on a shift

What is the condition of the stack now, including the variables no sensor reports. What will it be in a few weeks if the current operating pattern continues. What is the best operating point given today's power price, resource and constraints. And what should be done in the next few hours.

Each of those has a decision attached, which is the test worth applying to any proposed output. A twin that produces a fidelity metric, or a health score between zero and one, has produced a display rather than a decision, and the control room stops opening it within a fortnight.

The operating point question is the one most often skipped and frequently the most valuable. System specific energy against load is typically a shallow curve with an optimum in the middle of the range rather than at either end, and that optimum moves as the stack ages. Knowing where it sits today is worth more than knowing the rated figure.

A workable architecture

Acquisition sits closest to the plant, reading from interfaces that already exist rather than adding new ones, which is a separate problem covered in reading plant data without modifying the control system. Nothing in a twin should write anywhere near control logic.

The estimation layer runs the physical model against incoming data and produces the states that are not measured, together with a residual: the difference between what the model expected and what the plant reported. Learning components work on the residual rather than on the raw signal, so ordinary variation caused by load and ambient conditions is explained rather than treated as anomalous.

The attribution layer takes a residual that has grown and ranks the mechanisms that could produce it, along with the measurement or inspection that would separate them. This is where the output stops being a number and becomes something an engineer can act on or argue with.

All of it runs on site. The data volumes are modest, the latency requirement is real, and for assets with public or sovereign participation an on-premise deployment is frequently a procurement condition rather than a preference.

Questions teams ask

Frequently asked questions

What does a digital twin of an electrolyser actually model?

The electrochemical behaviour of the stack, principally the relationship between cell voltage and current density at a given temperature and pressure, coupled to thermal and gas transport behaviour, and extended across the balance of plant so the plant-level energy figure can be attributed. Most of what it produces are estimates of quantities no instrument reports.

How is it different from a digital twin of a power plant?

The taxonomy is the same but the useful kind is narrower. A geometric twin has nothing to say about a stack, and a purely statistical twin fails where the plant spends most of its life, which on a renewable-coupled asset is at operating points that are unusual by construction. What is left is a physics-based state estimator.

What data does it need?

Stack current and voltage, inlet and outlet temperature, pressure and differential pressure, hydrogen in oxygen, feedwater quality, and the electrical and hydraulic state of the balance of plant. Cell-level voltage where it exists, because the distribution carries information the aggregate hides. Plus a commissioning reference to calibrate against.

Can a twin be built after the plant has been running for years?

Yes, but with a weaker foundation. The calibration reference has to come from a controlled measurement taken now, which describes the current condition rather than the as-new one, so absolute degradation cannot be recovered. Trends from that point forward are still useful, and the sooner the reference is taken the more it is worth.

Does a twin replace the control system?

No, and it should not touch it. A twin reads from interfaces the plant already exposes and produces advisory information. Existing alarms, trips, interlocks and OEM operating limits remain the authoritative layers, and anything modelled on top of them sits alongside rather than in the path.

How do you tell a real twin from a trend viewer with a good name?

Ask what it estimates that is not measured, and how it was calibrated. A system that only reports transformations of tags it received is a trend viewer. A model that reports membrane condition, current distribution or an internal temperature, and can say what reference it was calibrated against, is doing something else.