Operational guide / September 2026 / 11 min read

Pump condition monitoring from the signals a plant already has

Engineering leader with experience at GE, Mitsubishi and Alstom, specialising in advanced controls, industrial process and multi-physics modelling, with R&D and patent-pending work behind the Yunify engine.

Protection systems watch a pump against fixed limits and act when one is crossed. A reliability decision needs something else: the same signals assembled against the operating state that produced them, including the start-up window where vibration alarms are commonly delayed or suppressed and where a developing fault shows itself first.

Condition MonitoringRotating EquipmentPredictive MaintenanceReliability

A threshold crossing is not an explanation

A protection system exists to stop a machine before it damages itself, and it does that with fixed limits and a short reaction time. That job is well served by a single number compared against a high and a high-high setting. It is the correct design for protection, and nothing here displaces it: alarms, trips, interlocks and OEM operating limits remain authoritative, and anything modelled on top of them is advisory.

The reliability question is a different one and the same number answers it poorly. When a vibration reading rises, the maintenance planner needs to know whether the machine has changed or whether the duty has, because those two lead to opposite decisions. A pump running against a partly closed discharge valve, or with a suction condition that has drifted, can produce a vibration signature that looks like a mechanical fault and is not.

Three questions decide the response, and none of them can be answered by a limit. Is the X and Y vibration changing at the same operating condition, or is the condition doing the changing? What happened before, during and after the last trip? Was the machine abnormal, or was it responding to hydraulic load exactly as it should?

The pump is a system, not a bearing

Condition monitoring on rotating equipment is often reduced to vibration because vibration is where the dedicated instrumentation sits. On a pump train, the signals that make a vibration reading interpretable are spread across four domains, and a plant that has a control system already has most of them.

The hydraulic domain carries suction and discharge pressure, flow, fluid temperature and valve position. It is what says whether the pump is operating near its best efficiency point or far from it, and whether the net positive suction head available has fallen towards the level where cavitation begins.

The electrical domain carries motor current and phase values and winding temperature. Current is an indirect but responsive measure of load, and its relationship to flow is a useful check on both: a current that has risen while flow has not is a different problem from one where both moved together.

The mechanical domain carries X and Y vibration, bearing temperature and, on instrumented machines, shaft position and phase data. The operational domain carries run state, speed, trips and direction of rotation, and it is the domain that tells the other three which regime they were recorded in.

Read together, these separate the cases that a single vibration number merges. Rising vibration with falling suction pressure and unchanged current points at the hydraulics. Rising vibration with unchanged process conditions across several days, at the same load, points at the machine.

Where the alarm layer is switched off

Vibration on a centrifugal pump rises legitimately during a start. The rotor accelerates through its own dynamics, hydraulic conditions have not settled, and readings pass through values that would be alarming during steady operation and are ordinary during a run-up. A fixed high-high limit applied through that period would trip the machine on every start.

The standard engineering answer is to delay or suppress the vibration alarms for a defined window after the start command. That answer is correct and it has a consequence: the part of the operating cycle that stresses the machine most, and that reveals a change earliest, is the part the alarm layer is configured not to react to.

A model conditioned on the operating state does not have that problem, because it is not comparing against a fixed limit. It compares against the behaviour expected for a start: the shape of the run-up, the values at each point along it, and the plant's own history of starts on that machine. A deviation from that envelope is meaningful while the reading is still well below the high alarm, which is the entire point of watching it.

The distinction is worth stating plainly, because it is the difference between the two layers. Fixed thresholds protect. A state-aware envelope warns, and it can warn earlier because it is asking a harder question than whether a number is large.

A baseline that does not know the operating state is not a baseline

Comparing today's vibration to last month's average assumes the machine was doing the same thing on both days. On a plant with a variable load, and particularly on one following a renewable profile, that assumption fails routinely. The pump ran at a different speed, against a different head, with a different fluid temperature, for a different duration.

What survives that variability is a baseline held per operating state. Starting, ramping, stable operation at each load band and shutdown each get their own expected behaviour, and a reading is compared against the state it belongs to. A deviation then means the machine departed from what it does in that state, rather than from what it does on average across all of them.

Persistence is the second half of it. A single excursion in a noisy measurement is not evidence, and treating it as evidence is how a monitoring system loses an operator's attention. A deviation that repeats across successive starts, or that holds for hours at the same operating point, is a different object entirely, and separating the two is a large part of what keeps the advisory list short enough to be read.

Reverse rotation, and events worth reconstructing whole

When a running pump trips against a static head, the column of liquid downstream can drive flow backwards through the impeller until the check valve seats. The pump turns backwards while that happens. How far backwards, for how long, and how quickly the hydraulics decay are all determined by the check valve's behaviour and the system it sits in.

This is a high-value event because several things are correlated across it. Direction and duration describe the reverse itself. The decay of discharge pressure describes the check valve. And the vibration on the next start describes whether anything moved as a result. A check valve that has begun to seat slowly is visible in that sequence long before it is visible anywhere else.

Reading it requires the event to be reconstructed as an event, with the signals from all four domains aligned on one timeline and at their original resolution. An averaged historian trend cannot do this, because the averaging happens on exactly the timescale the event occupies. The same applies to a trip, a swap between duty and standby machines, and a start into an unusual system alignment.

Two depths of signal, and starting with the one you have

There are two useful depths of pump data and they are not equally available. The first is what the control system and the historian already carry: pressures, flows, temperatures, currents, run states and the overall vibration values passed up from the monitoring rack. This depth supports operating-state baselines, persistent-change detection and event context, and it needs no new instrumentation.

The second is detailed vibration data: waveforms, spectra, run-up and coast-down plots, order tracking and envelope features. This is what identifies a mechanism by name, separating an unbalance from a misalignment from a bearing defect frequency. Where a machine has a dedicated vibration monitoring platform, such as a Bently Nevada 3500 series rack or a comparable system, that depth exists on the plant already and frequently stays inside the rack.

The practical order is to start with the first depth, because it is available on every critical pump on the site rather than on the few that carry proximity probes, and to bring in the second where the interface allows it. Both are read through interfaces the plant already exposes, which keeps this on the same footing as any other acquisition question: read-only, no modification to control logic, and the connection agreed against the site's management-of-change position before anything is connected. That ground is covered in reading plant data without modifying the control system.

What it changes about the maintenance decision

The output that matters is not an alarm count. It is a shift in what the planner is holding when the conversation about an outage happens: which component has changed, under which operating conditions the change appears, how far it has moved from that machine's own history, and whether it is still moving.

That is enough to convert a finding into a date. A deviation that appears only on starts and has grown across the last twelve of them supports a different plan from one that appears at high flow and has been stable for a month. Neither is available from a threshold, because a threshold has one bit of information in it.

Scoping follows the same logic. The critical pump, the one whose failure stops production or trips a unit, is where the work pays for itself, and the pattern extends across the rest of the plant afterwards without new instrumentation, since the signals were already being recorded.

Questions teams ask

Frequently asked questions

Does pump condition monitoring require new sensors?

Not to begin. Suction and discharge pressure, flow, fluid and bearing temperature, motor current, run state and overall vibration are typically already recorded by the control system and the historian, and that set supports operating-state baselines, persistent-change detection and event reconstruction. Detailed vibration data sharpens the diagnosis where a dedicated monitoring rack is installed, and the interface to it is a scoping question rather than a new installation.

Why can the DCS not do this already?

Because it is doing a different job. The control and protection layers compare readings against fixed limits so they can react quickly and predictably, which is the correct design for protection. Assembling several signals against the operating state that produced them, and comparing that to the machine's own history, is an analysis task rather than a control task, and putting it in the control layer would compromise the property that makes the control layer trustworthy.

Why does start-up matter so much?

Vibration rises legitimately during a run-up, so fixed alarms have to be delayed or suppressed through that window or the machine would trip on every start. That leaves the most mechanically demanding and most diagnostic part of the cycle without an alarm watching it. A model comparing against the expected shape of a start rather than against a fixed limit can evaluate that window, and can flag a change while the readings are still far below the high alarm.

What does reverse rotation indicate?

It occurs when flow reverses through the impeller after a trip, before the check valve seats. Direction, duration, the decay of discharge pressure and the vibration on the following start are correlated, and together they describe how the check valve is behaving. A valve seating more slowly than it did shows up in that sequence well before it shows up as a process problem.

Does this work alongside an existing vibration monitoring system?

Yes, and that is the intended arrangement. A Bently Nevada 3500 series rack or a comparable platform continues to perform its protection function unchanged. What is added is the reading of its outputs alongside the hydraulic, electrical and operational signals, through the DCS, industrial or vendor-supported interfaces available on the site, so the vibration data is interpreted in the operating context the rack does not see.

Is this a protection function?

No. It is read-only and advisory. Nothing writes to control logic, no trip depends on it, and the existing protection settings are unchanged by its presence. The safety instrumented system stays independent of it, as IEC 61511 requires of anything sharing a boundary with a safety function.

Which pumps are worth starting with?

The ones whose failure stops production or trips a unit, and the ones with a repair history nobody can currently explain. Those two sets usually overlap, and they are where the available data has the most to say, because the trips and excursions already sit in the historian waiting to be replayed.