AI in Power Plants: What It Actually Returns, What Your Data Must Contain, and Where Projects Fail

Power Generation

February 23, 2024

24 minutes read

power plant optimization

Three AI applications earn their keep in a power plant. The rest are marketed.

Combustion optimisation recovers 0.5 to 2.0 percent of heat rate and consistently delivers the highest return of any AI application in a thermal plant. Condition monitoring finds developing faults earlier than threshold alarms. Degradation tracking separates losses you can wash out from losses you cannot.

The returns are real and they are not revolutionary. On a 500 MW plant, a 1 percent heat rate improvement is worth roughly $1.2 to $2.4 million a year in fuel, typically achieved within 6 to 12 months.

The prerequisite is less demanding than most operators assume, and different from what they expect. A baseline performance model needs 6 to 12 months of DCS and historian data at 1 to 5 minute intervals, and most operating plants already hold 2 to 5 years of it. What they frequently do not hold is a record of which events were failures, and that absence stops more projects than data volume ever has.

This guide covers the three applications, the returns, the data requirement, the advisory-to-closed-loop decision, and where these projects go wrong.

What AI Actually Does in a Power Plant

Three applications operate at plant level and change an operating decision. Everything else on a typical vendor list is either a different reader's problem or a dashboard.

Combustion optimisation adjusts fuel-to-air ratio and burner biasing continuously against measured conditions, improving heat rate and reducing emissions.

Condition monitoring and anomaly detection identify developing equipment faults from patterns across vibration, temperature and process data, earlier than fixed thresholds respond.

Performance degradation tracking measures actual against design performance continuously, separating recoverable losses such as fouling from unrecoverable ones such as clearance growth.

What a digital twin is, and is not

A digital twin is a model of the plant that runs alongside the plant, fed by live data, producing an expected output against which the real output is compared. The gap is the finding.

It is not the same as a physics-based simulation, which models design behaviour without live data, and it is not the same as a dashboard, which displays data without modelling anything. Hybrid approaches combining physics-based models with data-driven ones are increasingly recognised as more reliable than either alone, for reasons the false positives section covers.

Where to start

Start with combustion optimisation if you burn fuel continuously, because it has the shortest path to a measurable number. Start with condition monitoring if your problem is unplanned outages rather than fuel cost.

Starting with both at once is how projects stall, because the data preparation, the operator engagement and the validation effort all double before anything has proved itself.

The Returns, With Numbers

AI in power plants delivers single-digit percentage improvements, and on a plant burning fuel continuously those are large absolute numbers.

Heat rate

Published deployments report 0.5 to 2.0 percent heat rate recovery from AI combustion optimisation, and 0.1 to 0.4 percent from steam temperature control. Plants typically operate 200 to 600 kcal/kWh above their design heat rate, which is the gap these systems address.

The strongest non-vendor case on record: a large utility in the southern United States improved heat rate by 1 to 3 percent and deployed more than 400 AI models to reduce forced outages across 67 generation units, both coal and gas.

What that is worth

A 1 percent heat rate improvement on a 500 MW coal plant is worth approximately $1.2 to $2.4 million annually in fuel, typically achieved within 6 to 12 months of deployment. On a 500 MW combined cycle plant, published figures put the same 1 percent at $1 to $3 million a year.

Heat rate is the fuel energy required per unit of electrical output, in Btu/kWh or kcal/kWh. Lower is better, and because fuel is the largest lifetime cost of a thermal plant, a fractional improvement compounds across every operating hour.

Emissions

Deployments report NOx reductions of 10 to 15 percent alongside the heat rate gain, because the same combustion adjustments that improve efficiency change the formation conditions for NOx.

Payback

Industry survey data indicates that 27 percent of organisations implementing predictive maintenance achieve full payback within the first year. On a power plant, payback frequently occurs after preventing a single major forced outage, because the avoided cost of one turbine or generator failure exceeds the programme cost.

Treat vendor-published forced outage and maintenance cost improvements with the same scepticism you would apply to any supplier's own performance data, and require a validated baseline before and after.

For what availability, forced outage rate and maintenance cost per MWh actually mean and how they are defined, see our guide to power plant O&M metrics and contracts.

Combustion Optimisation: The Highest-Return Application

Combustion optimisation consistently delivers the highest return on investment of any AI application in a thermal power plant, and it is also the oldest.

It works by adjusting fuel-to-air ratio and burner biasing continuously against measured conditions, rather than by the periodic manual tuning a conventional control system relies on.

Fuel-to-air ratio is the proportion of combustion air to fuel. Too little air and combustion is incomplete, producing CO and unburned fuel. Too much and the excess air carries heat up the stack. The optimum moves with load, fuel composition, ambient conditions and equipment condition, which is why a fixed setting is wrong most of the time.

Burner biasing distributes fuel and air unevenly across multiple burners to compensate for differences between them, flattening temperature distribution.

Why it outperforms manual tuning

A conventional distributed control system holds a static PID setting, tuned periodically and manually. A model-based optimiser adjusts continuously, which captures the gap between the one setting that was correct at the time of tuning and the many settings that are correct across a year of operation.

The emissions interaction

NOx formation is temperature-dependent, so the adjustments that improve heat rate also change NOx output, generally downward. On units with selective catalytic reduction, optimisation can also reduce ammonia reagent consumption while holding emissions within limit.

Where a unit is permitted under 40 CFR Part 60 Subpart KKKKa or an equivalent standard, any change to combustion operation has a compliance dimension, and emissions monitoring must demonstrate the unit remains within its envelope.

For conventional combustion tuning on dry low NOx systems, see our overview of DLN tuning services.

Condition Monitoring and Predictive Maintenance

AI-based condition monitoring identifies developing equipment faults from patterns across vibration, temperature and process data, earlier than fixed threshold alarms respond.

Traditional condition monitoring triggers when a measured value crosses a set limit. Pattern-based detection triggers when the relationship between several values departs from its learned normal, which happens earlier because the relationship degrades before any single value reaches its limit.

Remaining useful life (RUL) estimation projects how long a component will continue to function at its current degradation rate. It is the output that converts a detection into a scheduling decision.

What it changes operationally

It changes when you intervene, not whether you do. A detection with eight weeks of warning lets the work go into a planned outage; the same detection with eight hours does not.

That is the entire value proposition, and it is more modest and more defensible than the framing most vendors use.

Where it does not help

On components that fail without a detectable degradation path, pattern detection has nothing to detect. Random failures, manufacturing defects and foreign object damage all occur without a warning signature.

Matching the monitoring method to the failure mode is the prerequisite, and applying detection uniformly across a plant wastes effort on assets that will not benefit. For how to make that match asset by asset, see our comparison of predictive versus preventive maintenance. For the task-level programme underneath it, see our rotating equipment maintenance field guide.

Availability reporting

Outages avoided do not appear in any metric. Outages that occur are reported as forced outages and feed the forced outage rate used in capacity markets and O&M contracts, which means the benefit of this work shows up as an absence and has to be measured against a baseline rather than observed directly.

Degradation Tracking: Separating Recoverable From Permanent

Performance models calculate actual against design heat rate continuously, which lets an operator distinguish losses that washing recovers from losses that it does not.

This is the application that changes a specific recurring decision, and it is the clearest example of AI producing an operating instruction rather than a chart.

The compressor wash example

Heat rate degradation and output loss from compressor fouling are measurable in DCS data before they become visible operationally.

A performance model calculates the efficiency loss continuously and triggers an offline wash recommendation when the loss crosses the economic threshold where washing cost is less than continued fuel waste.

That replaces a fixed quarterly wash schedule with a condition-based one. In practice, condition-based washing adds 2 to 3 additional washes per year in high-fouling conditions while deferring washes in clean conditions, improving both fuel cost and compressor wear against a fixed interval programme.

Fouling rate varies significantly by season, inlet air quality and load profile, which is precisely why a calendar interval is wrong for most of the year.

What it finds that an operator would not

A single coked fuel nozzle on an F-class gas turbine can reduce that unit's heat rate by 0.8 to 1.4 percent per outage interval. That is a measurable loss with no alarm attached, and it is invisible without a continuous performance comparison.

Recoverable and unrecoverable

Recoverable: fouling, which washing restores. Unrecoverable: erosion, blade profile change, tip clearance growth, which washing does not touch.

If a wash does not restore performance, the loss is mechanical, and that distinction is the single most useful thing a degradation model produces. For the mechanisms and the washing economics, see our guide to gas turbine casings, clearance and fouling. For scheduling the intervention, see our outage planning guide.

The Data Question

A baseline AI performance model needs 6 to 12 months of DCS and historian data at 1 to 5 minute intervals, and most operating plants already have 2 to 5 years of it.

That is less demanding than most operators expect, and it means the resolution objection is usually wrong.

What is actually missing

Labelled failure events. A model that predicts failures needs to know what failure looked like. Most plants record that a work order was raised and closed, not that a specific component failed in a specific mode on a specific date, which means the failure signatures sit in the data unlabelled.

Tag naming consistency across sites. The same sensor named three ways across a fleet prevents a model trained on one unit from transferring to another. On a single-site deployment this does not bite; on a fleet programme it is the whole project.

Maintenance records linked to the asset. The CMMS record and the historian tag have to resolve to the same component, or the model cannot connect an intervention to a change in behaviour.

The honest limitation

Predictive models based solely on SCADA data face significant challenges including noisy sensor readings and missing data, which can compromise their reliability.

Integrating data-driven and knowledge-driven techniques is increasingly recognised as critical to improving reliability and reducing false positive alarms. In practice that means a model informed by engineering understanding of the machine performs better than one trained on data alone.

If your data has gaps

Gaps are normal and are handled by preprocessing rather than by abandoning the project. Systematic gaps are different from random ones: a historian that stopped logging during every outage has removed exactly the periods the model most needs.

If failures were never recorded

Start the record now and use unsupervised anomaly detection in the meantime. Unsupervised methods identify departures from normal without needing labelled failures, and they become supervised models as the failure record accumulates.

For the historian, protocols and control system architecture this all depends on, see our guide to SCADA systems, architecture and protocols.

Advisory or Closed Loop: The Decision That Matters Most

AI systems operate in advisory mode, recommending setpoint changes that operators approve, or in closed loop, changing setpoints directly. The two are completely different propositions.

The standard sequence is advisory first. Models run in shadow mode, operators see predictions alongside their own judgement, and the system's recommendations are evaluated against what the operators would have done anyway. Transition to closed loop follows for specific optimisation loops once confidence is established.

Plants typically see measurable improvement within 4 to 8 weeks of closed-loop activation, which is the point at which the system starts capturing the gap between the recommended setting and the setting a busy operator actually makes.

Why the boundary is not just a maturity milestone

Closed loop means software changes plant operating conditions without a human in the path. That has three consequences a maturity framing obscures.

Functional safety. Any system that can influence a safety-related function is inside the safety case. IEC 61511 governs safety instrumented systems in the process industries, and a control optimiser must be demonstrably outside that boundary or inside it with the attendant rigour.

Cybersecurity. An AI system pulling data from OT and writing setpoints back crosses the IT and OT boundary in both directions. ISA/IEC 62443 governs that crossing, and a bidirectional path is a materially different risk from a read-only one.

Insurance and liability. When a closed-loop system makes a setting that contributes to damage, the question of who is responsible is answered by the contract, and most AI contracts are written without that question in view.

If operators ignore the recommendations

That is the most common advisory-stage outcome and it is information. Either the recommendations are wrong, or they are right and the operators do not trust them. The first is a model problem; the second is an engagement problem, and treating one as the other wastes the deployment.

False Positives and Model Decay

False positive rates below 10 percent are achievable in mature deployments, and a system that cries wolf is worse than no system at all.

A false positive is an alert for a condition that is not developing. Each one consumes investigation time, and enough of them teach operators to ignore the system entirely.

Mature deployments report false positive rates falling below 10 percent, with vendors claiming figures under 4 percent through multi-parameter cross-validation and per-asset tuning during a pilot phase.

Require a measured false positive rate from a pilot on your own plant, not a published figure from someone else's.

Why rates start high and fall

Early models flag everything unusual, including unusual-but-normal conditions such as seasonal operation, fuel changes and planned derates. Rates fall as the model learns which departures from normal are benign, which requires feedback on every alert.

A deployment with no feedback loop does not improve. The confirmed failures and the confirmed false alarms are both training data, and only one of them is usually captured.

Model drift

Model drift is the degradation of a model's accuracy as the plant changes beneath it. Component replacements, overhauls, fuel changes and ageing all shift the normal the model learned.

Retraining cadence belongs in the contract. A model trained once and never updated becomes less useful every month, and the deterioration is gradual enough that nobody notices until the alerts stop being believed.

Integration and the Standards That Apply

AI systems integrate with plant control systems through established industrial data interfaces, and the integration is usually the longer part of the work.

Integration is typically through a plant information historian, OPC UA, or direct DCS application interfaces. OPC UA, standardised as IEC 62541, is the platform-independent industrial data access standard and the usual route for a new system.

The standards that govern the boundary

ISA-95, also published as IEC 62264, defines the enterprise and control system integration hierarchy. It is the reference for where an analytics layer sits relative to the control system.

ISA/IEC 62443 governs industrial automation and control system security, using zones and conduits to define segmentation. An AI system is a conduit between zones, and it must be treated as one.

ISO 55001 provides the asset management system framework. Predictive maintenance outputs including failure probability records, intervention justification and remaining useful life estimates can be structured to support it.

IEC 61511 governs functional safety for safety instrumented systems, and defines the boundary a control optimiser must respect.

The cybersecurity consequence

An analytics platform that reads from the historian is a low-risk addition. One that writes setpoints into the control system is not. The direction of data flow, not the sophistication of the model, determines the security assessment.

Where These Projects Fail

Five failure modes account for most unsuccessful AI deployments in power plants, and only one of them is about the model.

No labelled failure history. The data exists and nobody recorded which events were failures, so there is nothing for a supervised model to learn from.

No feedback loop. Alerts are raised, investigated and closed without the outcome being fed back, so the model never learns which alarms were real.

Operator disengagement. Recommendations arrive without explanation, operators cannot evaluate them, and the system is ignored into irrelevance.

No validated baseline. The savings cannot be proven because nobody established what performance looked like before. This is the one that kills the second phase of funding.

Scope sprawl. Three applications attempted at once, none of them deployed well enough to prove itself.

Staffing

These systems require someone who understands both the plant and the data to act as the interface. Without that person the alerts go to a queue and the recommendations go unread, regardless of model quality.

Insurance

Insurers assess maintenance regime and documented condition monitoring when pricing machinery breakdown risk. A documented predictive programme with recorded findings and actions is underwriting evidence, and an undocumented one is not.

For structured root cause work when something does fail, see our guide to gas turbine troubleshooting and unplanned downtime.

How to Evaluate a Proposal

Six questions separate a deployable proposal from a demonstration, and none of them is about the algorithm.

What data does it need, and do we have it? Specify the tags, the resolution and the history period, and audit against the historian before signature rather than during mobilisation.

Is the model trained on our plant or on a generic baseline? A model trained on your machines, your fuel and your operating profile behaves differently from one trained on an industry average.

What false positive rate did it achieve in the pilot, on our data? Published figures are marketing; a pilot number is evidence.

How is the baseline established and the saving validated? Agree the measurement method before the project starts, because agreeing it afterwards is a negotiation.

Who owns the models, the configuration and the data at contract end? A vendor holding your trained models controls your next tender, in exactly the way a contractor holding your CMMS data or your relay settings does.

What is the retraining cadence and whose cost is it? Model drift is certain, and an unretrained model decays.

Build, buy or partner

Buy where the application is standard and the vendor has a comparable reference. Build where the requirement is specific and the organisation has data engineering capability it can keep busy. Partner where the capability is needed for a defined programme rather than permanently.

Deployment timeline

Published deployment timelines vary widely, from a five-week fixed pilot to six to eighteen months for full historian integration and model configuration. The variable is data condition, not software, which is why the data audit belongs before the proposal rather than after.

How Applications Differ by Plant Type

The three applications are constant. Their relative value changes with the machine.

Gas turbine simple cycle. Degradation tracking leads, because fouling is the dominant recoverable loss and the wash decision is the highest-frequency operating call. Condition monitoring on the rotating fleet follows.

Combined cycle. Combustion optimisation and heat rate tracking carry the most value, because the plant runs more hours and the HRSG and steam cycle add parameters worth optimising.

Coal and steam plant. Combustion optimisation is the single largest recoverable lever, and the heat rate gap against design is typically largest.

Reciprocating engine plant. Many units, each with its own condition data, which suits fleet-level anomaly detection and comparative benchmarking across machines.

Smaller plants. The economics are harder because the absolute saving scales with fuel consumption. A 1 percent heat rate improvement on a 20 MW plant is a fraction of the same percentage on a 500 MW one, and the deployment cost does not scale down proportionally.

Outside the United States. The applications and the ISA, IEC and ISO standards are international. Emissions permitting frameworks and grid codes are not, and both affect what a combustion optimiser is permitted to do.

For how data centre demand is reshaping generation requirements, see our analysis of AI and data centre power needs. For grid-side integration standards, see our guide to smart grid and distributed energy resources.

What Prismecs Does

Prismecs operates and maintains power generation plant, provides instrumentation and control services including SCADA and DCS support, and delivers technology and consulting for performance assessment.

Delivered project scope includes four TM2500 units totalling 110 MW at Duqm, Oman, kept grid-ready with resident O&M crews, CMMS and parts support; eight TM2500 dual-fuel units totalling 260 MW at Birr, Switzerland, online in six months with a new 220 kV interconnection; an LM2500XPRESS plant at Miaoli, Taiwan delivered in ten months; and an LM6000 fleet decommissioned in Norway, transported and recommissioned at a new site.

The Duqm arrangement is the relevant reference for this subject. Resident crews with a maintained CMMS and parts support is the operating model that produces the labelled failure history and the maintenance records any plant analytics programme depends on. The data problem described in this article is solved by operating discipline, not by software.

Capability spans technology and consulting for performance assessment and options analysis, O&M services for the operating phase, I&C services for instrumentation, controls, SCADA and DCS support, power generation asset services for the equipment, and owner's engineering for independent review.

Prismecs is OEM-agnostic, which on a performance assessment matters because the party measuring the degradation is not the party selling the parts.

Apply this article's criteria to any proposal, including ours. Ask what data it needs and audit the historian against that list before signing. Ask whether the model trains on your plant or on a generic baseline. Ask for a pilot false positive rate measured on your own data. Ask who owns the models at contract end.

To discuss plant performance, data readiness or an O&M scope, send your plant configuration, current heat rate against design, historian platform and retention period, and the outcome you need to sales@prismecs.com or call +1 (888) 774-7632.

Frequently Asked Questions

What does AI actually do in a power plant?

Three applications operate at plant level and change an operating decision. Combustion optimisation adjusts fuel-to-air ratio and burner biasing continuously against measured conditions. Condition monitoring identifies developing faults from patterns across vibration, temperature and process data earlier than fixed thresholds. Degradation tracking measures actual against design performance continuously, separating recoverable losses such as fouling from unrecoverable ones such as clearance growth.

How much heat rate improvement is realistic?

Published deployments report 0.5 to 2.0 percent heat rate recovery from AI combustion optimisation and 0.1 to 0.4 percent from steam temperature control. A large US utility reported 1 to 3 percent across its fleet, alongside more than 400 AI models deployed to reduce forced outages across 67 generation units. Plants typically operate 200 to 600 kcal/kWh above their design heat rate.

What is a 1 percent heat rate improvement worth?

Approximately $1.2 to $2.4 million annually on a 500 MW coal plant, typically achieved within 6 to 12 months of deployment. Published figures put the same 1 percent improvement at $1 to $3 million a year on a 500 MW combined cycle plant. Because fuel is the largest lifetime cost of a thermal plant, a fractional efficiency gain compounds across every operating hour.

Which AI application has the best return?

Combustion optimisation, consistently, in thermal plants. It adjusts fuel-to-air ratio and burner biasing continuously where a conventional control system holds a static setting tuned periodically and manually. The gain comes from capturing the difference between one setting that was correct at tuning and the many settings correct across a year of varying load, fuel and ambient conditions. Deployments also report 10 to 15 percent NOx reduction.

How much plant data do I need before starting?

A minimum of 6 to 12 months of DCS and historian data at 1 to 5 minute intervals is sufficient to train a baseline performance model, and most operating plants already hold 2 to 5 years of it. Data volume and resolution are rarely the blocker. The common gap is labelled failure history: plants record that a work order was raised and closed, not that a specific component failed in a specific mode.

What if our failures were never recorded?

Start the record now and use unsupervised anomaly detection in the meantime. Unsupervised methods identify departures from learned normal behaviour without requiring labelled failures, and they become supervised models as the failure record accumulates. The absence of labelled history delays the more accurate class of model; it does not prevent a project from starting.

What is the difference between advisory and closed loop?

Advisory means the system recommends setpoint changes that operators approve. Closed loop means it changes setpoints directly. The standard sequence is advisory first, with models running in shadow mode so operators can evaluate predictions, then transition to closed loop for specific optimisation loops. Plants typically see measurable improvement within 4 to 8 weeks of closed-loop activation.

Why does the advisory to closed loop decision matter beyond maturity?

Because closed loop means software changes plant operating conditions without a human in the path. That has functional safety implications under IEC 61511 if the system can influence a safety-related function, cybersecurity implications under ISA/IEC 62443 because the data path becomes bidirectional across the IT and OT boundary, and liability implications that most AI contracts are written without considering.

What false positive rate should I expect?

Mature deployments report false positive rates below 10 percent, with some vendors claiming under 4 percent through multi-parameter cross-validation and per-asset tuning during a pilot. Require a measured rate from a pilot on your own plant rather than a published figure from someone else's. Rates start higher and fall as the model learns which departures from normal are benign.

What is model drift and how is it managed?

Model drift is the degradation of a model's accuracy as the plant changes beneath it, through component replacement, overhauls, fuel changes and ageing. All of these shift the normal behaviour the model learned. Retraining cadence belongs in the contract, including whose cost it is, because a model trained once and never updated decays gradually enough that nobody notices until the alerts stop being believed.

How does AI change compressor washing?

Heat rate degradation and output loss from fouling are measurable in DCS data before becoming visible operationally. A performance model calculates the loss continuously and triggers a wash recommendation when it crosses the threshold where washing cost is less than continued fuel waste. In practice this adds 2 to 3 washes per year in high-fouling conditions and defers them in clean conditions, improving both fuel cost and compressor wear.

What is a digital twin?

A model of the plant that runs alongside it, fed by live data, producing an expected output against which real output is compared. The gap is the finding. It differs from a physics-based simulation, which models design behaviour without live data, and from a dashboard, which displays data without modelling. Hybrid approaches combining physics-based and data-driven models are increasingly recognised as more reliable than either alone.

Which standards apply to plant AI systems?

ISA-95, also published as IEC 62264, defines the enterprise and control system integration hierarchy. ISA/IEC 62443 governs industrial control system security through zones and conduits, and an AI system crossing between them is a conduit. IEC 61511 governs functional safety for safety instrumented systems. ISO 55001 provides the asset management framework that predictive maintenance outputs can be structured to support.

Why do power plant AI projects fail?

Five reasons, and only one concerns the model. No labelled failure history, so there is nothing for a supervised model to learn. No feedback loop, so alerts are closed without the outcome returning to the model. Operator disengagement, where recommendations arrive without explanation and get ignored. No validated baseline, so savings cannot be proven. And scope sprawl, attempting three applications at once.

Tags: AI in Power Plants Heat Rate Optimization Predictive Maintenance Combustion Optimization Plant Data