O&M Services
July 24, 2025
29 minutes read
O&M is two functions carrying one name, and most owners can only describe one of them.
Maintenance is well documented. Operations is what you are actually handing over when you sign an O&M agreement: the control room, the procedures, the shift roster, the alarm response and the daily decisions that determine whether your plant runs. Most owners discover where the line falls when something goes wrong.
This guide covers what an operations scope contains, how operators are qualified and how competency differs from training, what ISA-18.2 and EEMUA 191 say about how many alarms one operator can actually handle, which regulatory obligations legally cannot transfer to a contractor, and how to measure the whole arrangement.
Operations means running the plant. Maintenance means keeping its equipment fit to run. An O&M agreement covers both, and the boundary between them is where most contractual disputes originate.
Operations is continuous and people-dependent. It covers the control room, startup and shutdown, load following and dispatch compliance, operator rounds, fuel handling, water chemistry, alarm response, permit issuance, and environmental compliance reporting. It runs every hour the plant exists, whether or not the plant is generating.
Maintenance is episodic and asset-dependent. It covers preventive tasks on intervals, condition monitoring, corrective repair, outage execution and spares management. It is driven by asset condition and calendar rather than by the shift clock.
The distinction matters commercially because the two carry different risks. Maintenance risk is largely technical and shows up as equipment failure. Operations risk is largely human and shows up as an error, a missed alarm, a procedure not followed, or a compliance report not filed.
An O&M agreement is the contract under which a third party provides one or both functions. It differs from a long term service agreement, which covers maintenance of specified equipment rather than operation of the facility. Many owners hold both, and the boundary between the two instruments needs to be defined by system and tag number. For how to choose and structure the maintenance side, see our guide to choosing a power plant O&M provider. For how to select the right maintenance strategy per asset, see our comparison of predictive versus preventive maintenance.
An operations scope contains nine functions, and a contract that does not name all nine has left at least one of them unassigned.
Operator rounds are the scheduled physical inspection routes an operator walks each shift, taking local readings and observing equipment condition that instrumentation does not capture. They are the oldest diagnostic tool in a plant and the first thing cut when staffing is thin.
Note what is not in this list. Preventive maintenance execution, outage planning, spares procurement and condition monitoring are maintenance functions. They may sit with the same contractor, but they are a separate scope with separate deliverables.
Three contract models exist, and they differ in who employs the operators and who carries the operational decision.
Full O&M transfers the most and is the only model where a single party can be held to a plant-level availability guarantee. Staff augmentation transfers the least: the owner keeps every obligation and buys headcount.
The middle model is the most common and the most frequently underspecified, because both parties assume the other holds the functions that sit between operating and maintaining. Fuel handling, permit issuance, environmental reporting and CMMS data entry are the four that most often fall in the gap.
Keep operations in-house when your plant is one of several you already operate, when your process knowledge is genuinely proprietary, or when operator continuity is itself the asset. Outsource when you do not have an operating organisation, when recruiting into your location is the constraint, or when you need round-the-clock coverage you cannot staff.
The test is not cost. It is whether you can recruit, qualify, retain and supervise a shift roster indefinitely. If the answer is no, a staffing model that depends on you doing so will fail regardless of the rate.
Mobilising a new operations contractor takes 8 to 16 weeks from award to full shift coverage, and rushing it is the most common cause of an early-term incident.
The sequence that works: procedure review and gap closure, operator recruitment and screening, site-specific training and plant familiarisation, supervised shadow shifts alongside the incumbent, competency assessment and sign-off, permit authority handover with named authorised persons, CMMS and records transfer, then phased shift assumption with the incumbent on site during ramp-up.
Where transition coincides with an outage, plan the outage separately and do not compress both into one window. See our guide to gas turbine outage planning.
Training records prove attendance. Competency assessment proves the operator can perform the task correctly under the conditions the task will actually occur in, and only the second one protects your plant.
Require both, and require them separately. A provider that supplies only training certificates has not demonstrated that anyone can start your unit at three in the morning during a grid disturbance.
TRIR is the OSHA total recordable incident rate. EMR is the experience modification rate used by workers compensation insurers, where 1.0 is the industry average and below 1.0 indicates a better-than-average loss history. Both are verifiable and both are what site prequalification actually turns on.
Operator turnover is the largest hidden risk in an operations contract, because competency is site-specific and takes months to rebuild. Write in a minimum overlap period for replacements, a cap on simultaneous departures from a single shift, and a requirement that any replacement completes the full site qualification before unsupervised duty. Then track actual turnover as a contract KPI, not just as a conversation.
Operating procedures are the instrument through which the owner retains control of a plant they no longer staff, and they should be owner-approved documents that the contractor executes rather than contractor documents the owner receives.
A complete set contains standard operating procedures for normal operation, startup and shutdown procedures for each operating mode, abnormal operating procedures covering the deviations an operator must recognise and correct, emergency operating procedures, and isolation and return-to-service procedures.
Abnormal operating procedures are the ones that matter most and the ones most often missing. They cover the state between normal running and an emergency, where an operator has time to act and the outcome depends on whether they recognise the condition. Every alarm in a rationalised system should map to a documented operator response, which is exactly the link between procedures and alarm management.
29 CFR 1910.119, OSHA's Process Safety Management standard, defines fourteen elements that together form the most widely recognised operations governance framework: employee participation, process safety information, process hazard analysis, operating procedures, training, contractors, pre-startup safety review, mechanical integrity, hot work permit, management of change, incident investigation, emergency planning and response, compliance audits, and trade secrets.
Whether PSM legally applies to your site is a separate question, addressed in the next section. Whether the framework is useful is not. Management of change, pre-startup safety review and incident investigation are the three that most improve an operations contract regardless of applicability, because each one is a decision gate where the owner and contractor must both sign.
Management of change is the process by which any modification to equipment, procedures, staffing or operating limits is reviewed and approved before implementation. Without it, an operations contractor can progressively change how your plant runs and nobody records that it happened.
Pre-startup safety review is the check performed before a modified or newly commissioned system is returned to service, confirming that procedures exist, training is complete and the modification was executed as designed.
Permit to work is the system controlling who may perform what work on which equipment and under what isolations. Name the authorised persons in a register, define authority levels, and require that the register is maintained and auditable. Permit authority is the single most consequential thing an operations contract delegates.
Procedures decay. Require an annual review with the review date recorded on each document, require that every management of change generates a procedure revision where relevant, and spot-check three procedures a year against what operators actually do. The gap between the written procedure and the practised method is the most reliable early indicator of an operations problem.
A well-managed alarm system presents no more than one alarm per operator per ten minutes during normal operations, roughly 144 to 150 alarms per day, and no more than ten alarms in the first ten minutes following a major upset.
Those benchmarks come from EEMUA Publication 191, Alarm Systems: A Guide to Design, Management and Procurement, now in its 4th Edition (November 2024), and are reflected in ANSI/ISA-18.2, the North American standard for alarm management in the process industries, and IEC 62682, its international counterpart.
This matters to an O&M contract for one reason: alarm rate determines how many operators you need. An operator cannot respond to alarms faster than they can process them, so a poorly rationalised alarm system silently converts into either a staffing cost or an unacknowledged risk.
EEMUA 191 characterises operator load bands directly. Fewer than ten alarms in a ten-minute period should be manageable, though difficult if several require a complex response. Twenty to a hundred alarms per operator is hard to cope with. More than a hundred is excessive and very likely to lead to the operator abandoning use of the system.
Sustained rates above roughly twelve alarms per hour are treated across the benchmark family as more than an operator can reliably manage.
Real-world performance sits well outside those targets. Benchmarking work by the Abnormal Situation Management Consortium across 37 operator consoles found a monthly overall average of 2.3 alarms per ten-minute period and a median of 1.77, meaning just over half of consoles reached the manageable band. About one third achieved the recommended rate of under one alarm per ten minutes. The earlier UK Health and Safety Executive study that led to the EEMUA guidance found five alarms per ten-minute window.
An alarm flood is conventionally defined as more than ten alarms in any ten-minute period for one operator, a rate at which an operator can no longer meaningfully process each one. ANSI/ISA-18.2 uses this definition.
Averages hide floods. A plant can pass a daily alarm target while its operators are overwhelmed during every upset, which is precisely when the alarm system matters. Track the percentage of time above ten alarms per ten minutes as a weekly operations KPI with a named owner.
Alarm rationalisation is the process of assessing every configured alarm against a documented alarm philosophy, confirming that it has a genuine purpose, a defined operator response and a verified setpoint. The output is a master alarm database recording each alarm's setpoint, priority, classification, cause, consequence and expected operator action.
Rationalisation typically eliminates 30 to 60 percent of an existing alarm database's configured alarms. The work is also far less daunting than it sounds, because alarm activity is heavily skewed: the top ten to twenty alarm tags in a typical unrationalised system produce most of the daily load. Start there.
ANSI/ISA-18.2 treats rationalisation as an ongoing work process within a ten-stage alarm management lifecycle, not a one-time cleanup, and requires periodic review as equipment and process change.
Require a baseline alarm performance assessment at mobilisation, monthly reporting of average rate, flood percentage and standing alarm count, a documented alarm philosophy, and a rationalisation programme with a defined completion horizon. Where alarms are credited as a layer of protection in a process hazard analysis or under IEC 61511 and ANSI/ISA 84 for safety instrumented systems, that credit is only valid if the alarm performance supports it.
This is the clearest example of a discipline that sits squarely in operations, is fully standardised, directly drives staffing and safety, and appears in almost no O&M contract.
Signing an O&M agreement transfers work. It does not automatically transfer regulatory liability, and four obligations in particular tend to stay with the owner whether or not anyone says so.
NERC registers entities under functional categories including Generator Owner and Generator Operator, and the registered entity carries compliance obligations under the applicable Reliability Standards. An O&M contractor performing the operating work does not by itself become the registered Generator Operator.
Establish explicitly which party holds registration, who performs the compliance activities, who maintains the evidence, and who responds to an audit. Confirm your current registration status with your Regional Entity rather than assuming, because registration criteria have been revised and the position is entity-specific.
NERC's Generating Availability Data System, GADS, requires performance reporting from registered entities, and the data behind it comes from the operating crew. Assign the reporting duty and the data quality duty separately, because they are different jobs.
Title V operating permits are issued to the source and the permit holder retains responsibility for compliance. Continuous emission monitoring under 40 CFR Part 75 generates the data, but the certification, quality assurance and reporting obligations sit with the permit holder.
This is the function most often left unassigned in the delegation matrix. The contractor operates the CEMS. The owner holds the permit. Nobody has been told who files the deviation report.
Both the owner and the contractor have duties on a shared worksite, and an O&M agreement does not extinguish the owner's. 29 CFR 1910.119 contains a specific contractors element addressing exactly this, requiring the host employer to evaluate contractor safety performance and inform contractors of hazards.
This is where most power plant operators get the boundary wrong in both directions.
29 CFR 1910.119 applies to processes involving a chemical listed in Appendix A at or above its threshold quantity, or a Category 1 flammable gas or a flammable liquid with a flashpoint below 100°F present in one location at 10,000 pounds or more. Appendix A lists 137 chemicals with threshold quantities ranging from 100 pounds to 15,000 pounds.
But 29 CFR 1910.119(a)(1)(ii)(A) excludes hydrocarbon fuels used solely for workplace consumption as a fuel, provided they are not part of a process containing another covered highly hazardous chemical. OSHA's interpretation confirms that natural gas, fuel oils, heating oils, bunker oil, blast furnace gas and coke oven gas are examples of materials that qualify for this exclusion.
So a gas-fired plant burning natural gas as fuel generally falls outside PSM for that gas. A refinery or petrochemical site does not.
The trap is ammonia. Appendix A lists anhydrous ammonia at a 10,000-pound threshold and ammonia solutions above 44 percent by weight at 15,000 pounds. Adding selective catalytic reduction with anhydrous ammonia storage can therefore pull a plant into full PSM coverage that was previously outside it. This is why SCR installations commonly specify aqueous ammonia below that concentration, and it is an operations consequence of an emissions decision that must be evaluated at design rather than discovered at an inspection.
Establish which framework applies before the contract is written, because PSM applicability changes the entire procedure, training and audit burden on the operating scope.
Every availability figure in an O&M contract must name its IEEE Std 762 metric, or the guarantee cannot be enforced in any dispute that matters.
IEEE Std 762, Definitions for Reporting Electric Generating Unit Reliability, Availability, and Productivity, is the terminology standard that makes availability comparable between parties. NERC's GADS calculates its statistics using IEEE 762 procedures, which gives you an auditable basis rather than a house definition.
The last three are leading indicators. The first six are lagging. An operations contract measured only on lagging indicators tells you about failures after they have happened, which is late.
Operations fees are typically structured as a fixed monthly management fee plus reimbursable staff costs, or as a fully fixed fee. Fixed fee transfers cost risk to the contractor and buys budget certainty at a premium. Cost-plus keeps margin visible and suits owners with the capacity to govern.
Where a bonus and penalty regime is attached, tie it to a named IEEE Std 762 metric, state the measurement period, list every exclusion in full, and express liquidated damages as a stated amount per percentage point against your actual revenue exposure rather than a notional figure.
Most availability disputes are definitional rather than factual. Preventing them requires four things agreed before mobilisation: the metric and its IEEE Std 762 definition, the data source and who records it, the exclusion list, and the dispute procedure including who holds the tie-breaking data. Agree all four in writing, because the party holding the historian holds the argument.
Your maintenance and operating history is an asset, and the contract must state that you own it in an exportable format regardless of how the relationship ends.
A computerised maintenance management system, or CMMS, is the platform that holds the equipment register, work orders, PM schedules, asset history and often the operating log. It is where the evidence of how your plant has been run lives.
Who owns the licence and who holds administrative credentials. Whether the equipment register, work orders, inspection findings, condition data, operating logs and procedure library are exportable and in what format. Who is responsible for data quality and completeness, which is a different obligation from data entry. And that everything is delivered at termination regardless of cause.
A contractor holding your operating history controls your position in the next tender and depresses your asset's value in any technical due diligence. This is the least negotiated and most consequential clause in a typical O&M agreement.
ISO 14224, Collection and exchange of reliability and maintenance data for equipment, defines the failure and maintenance data taxonomy that makes records comparable across sites and across contractors. Specifying it at mobilisation means the data you get back is usable rather than merely voluminous.
ISO 55001:2024, Asset management, Asset management systems, Requirements, is the management system standard the whole arrangement sits inside where an owner runs a formal asset management framework.
Operating procedures follow the same rule. Specify that the procedure library is owner-owned, owner-approved and delivered at termination in an editable format. Procedures written by a departing contractor and taken with them is a recurring and entirely avoidable failure.
At transition, count and verify rather than accept. Check the CMMS export against the physical equipment register, check open work orders against the deferred maintenance backlog, check the authorised person register against who is actually on site, and check the procedure library against the equipment list for gaps.
Shift handover is the highest-risk routine event in plant operations, because it is the only moment when responsibility for a running plant transfers between people.
Continuous coverage requires a minimum of four shift crews to sustain a 24-hour, 7-day roster with allowance for leave, training and absence. Five-crew rosters are common where training load or leave entitlement is higher. A plant that staffs three crews is running on overtime and will show it in fatigue-related error.
Crew size per shift is driven by three things rather than by plant capacity: the number of operator rounds the site requires, the alarm load per console, and whether permit issuance and field operation can be performed by the same person during an event. That last question is the one most often answered optimistically.
Remote operation and unmanned running are viable for some simple-cycle and reserve plants and unacceptable for others, and the deciding factors are the time to reach site, the consequence of a failed automatic response, and whether any manual local action appears in your emergency procedures.
Where a plant runs unmanned, define the maximum response time to site, confirm that no procedure requires a local action inside that window, and test the assumption rather than asserting it. Prismecs runs seasonal O&M teams that keep a TM2500 safe and ready for rapid restarts during peak demand, which is a staffing model built around readiness rather than continuous presence.
Plant status and operating mode, work in progress and its stage, open permits and active isolations, abnormal conditions and standing alarms, any deviation from procedure in force, communications received from the grid operator or the owner, and anything the incoming shift must watch. Require it in a structured written format, face to face, with both parties signing.
The nine functions stay the same. The weighting and the regulatory framework change.
Dispatch compliance, grid code obligations and NERC reporting dominate, and the operating crew generates the data that proves all three. Start reliability matters more than availability on peaking and reserve duty, and it is the KPI most often omitted.
PSM applies in full at most of these sites, so the fourteen-element framework is a legal requirement rather than good practice. Alarm management carries higher consequence because alarms are frequently credited as protection layers in the hazard analysis. Permit to work and confined space entry under 29 CFR 1910.146 carry the heaviest operational load.
Operations is dominated by change control and by the discipline of not touching anything during a critical window. Concurrent maintainability means work proceeds while load is live, which makes permit control and isolation verification the highest-risk activities on site. For the underlying redundancy design, see our guide to data center power redundancy.
Remote siting means long response times and thin local labour markets, which pushes toward larger resident crews and greater self-sufficiency. Cross-trained operators who can perform first-line maintenance are worth more here than anywhere else.
Smaller sites often cannot justify four shift crews, so the realistic models are day-shift operation with remote monitoring out of hours, or shared coverage across multiple sites. Define the out-of-hours response obligation precisely, because that is the scope that actually differs.
Below roughly 25 MW, a full operations contract is frequently uneconomic. A practical alternative is owner-employed operators with contracted supervision, procedures and competency assessment, which buys the governance without the headcount.
Prismecs provides operations and maintenance with resident crews, not only maintenance visits, and the operations evidence is on units Prismecs did not manufacture.
Verified scope includes full O&M staffing and specialist crews operating three TM2500 units delivering primary power; multi-year O&M crews with CMMS sustaining eight TM2500 dual-fuel units totalling 260 MW at Birr, Switzerland as reserve capacity; four TM2500 units at Duqm, Oman totalling 110 MW kept grid-ready with O&M teams, CMMS and parts support; seasonal O&M teams maintaining a TM2500 in a safe and ready state for rapid restarts during peak demand; and installation and commissioning of three LM6000PC units adding 150 MW of fast-start reserve.
Prismecs is OEM-agnostic, which means the operating and maintenance recommendation is not tied to one manufacturer's aftermarket, and the same organisation that commissions a unit can operate it afterwards.
Apply this article's criteria to any provider, including us. Ask for the competency assessment records rather than the training certificates, the authorised person register and authority levels, the alarm performance baseline method, the named IEEE Std 762 metric behind any availability figure, and the shift roster with crew count and rotation.
To request an operations scope and delegation review, send your unit type and capacity, current staffing arrangement, contract expiry date, applicable permits and registration status to sales@prismecs.com or call +1 (888) 774-7632. We return a function-by-function delegation matrix, a gap list and a KPI framework. Full service detail is on our O&M services page.
Operations means running the plant: control room duty, startup and shutdown, operator rounds, fuel handling, water chemistry, permit issuance, alarm response and compliance reporting. Maintenance means keeping equipment fit to run through preventive tasks, condition monitoring, repair and outage execution. Operations is continuous and people-dependent; maintenance is episodic and asset-dependent. The boundary between them is where most O&M disputes originate.
Nine functions: control room operation, startup and shutdown execution, operator rounds, fuel management, water and chemistry control, permit to work, environmental compliance reporting, performance monitoring, and emergency response. A contract that does not name all nine has left at least one unassigned. Environmental reporting, fuel handling, permit issuance and CMMS data entry are the four most commonly left in the gap between owner and contractor.
EEMUA 191 and ANSI/ISA-18.2 converge on no more than one alarm per operator per ten minutes during normal operation, roughly 144 to 150 per day, and no more than ten alarms in the first ten minutes of a major upset. Sustained rates above roughly twelve alarms per hour exceed what an operator can reliably manage. More than a hundred alarms per operator is likely to cause the operator to abandon the system.
ANSI/ISA-18.2 defines an alarm flood as more than ten alarms activating within any ten-minute period for one operator, a rate at which the operator can no longer meaningfully process each alarm individually. Floods matter more than averages, because a plant can pass a daily alarm target while operators are overwhelmed during every upset. Track the percentage of time above ten alarms per ten minutes as a weekly KPI.
Rationalisation typically eliminates 30 to 60 percent of an existing alarm database's configured alarms. The work is less daunting than the database size suggests, because alarm activity is heavily skewed and the top ten to twenty alarm tags in an unrationalised system usually produce most of the daily load. Start with those, then work through the remainder as part of the ANSI/ISA-18.2 lifecycle.
Usually not for the fuel itself. 29 CFR 1910.119(a)(1)(ii)(A) excludes hydrocarbon fuels used solely for workplace consumption as a fuel, and OSHA interpretation confirms natural gas and fuel oils qualify, provided they are not part of a process containing another covered chemical. However, adding selective catalytic reduction with anhydrous ammonia can trigger coverage, since Appendix A lists anhydrous ammonia at a 10,000-pound threshold.
Not automatically the contractor. NERC registers entities under functional categories including Generator Owner and Generator Operator, and the registered entity carries the compliance obligations under applicable Reliability Standards. An O&M contractor performing the operating work does not by itself become the registrant. Establish in the contract which party holds registration, who performs compliance activities, who maintains evidence and who responds to an audit.
The permit holder, unless the contract says otherwise and the regulator accepts it. Title V operating permits are issued to the source and the holder retains compliance responsibility. Continuous emission monitoring under 40 CFR Part 75 generates the data, but certification, quality assurance and reporting obligations sit with the permit holder. This function is the single most commonly unassigned item in an O&M scope.
A minimum of four shift crews to sustain 24-hour, 7-day coverage with allowance for leave, training and absence, with five-crew rosters common where training load or leave entitlement is higher. A three-crew roster runs on overtime and produces fatigue-related error. Crew size per shift is driven by round frequency, alarm load per console, and whether permit issuance and field operation can be performed by the same person during an event.
Training records prove attendance at a course. Competency assessment proves the operator can perform a specific task correctly, on your plant, under the conditions the task will occur in. Only the second protects your asset. Require both separately, with task-based assessment against your plant rather than a generic syllabus, and require sign-off per operator per unit before unsupervised duty.
Through lagging metrics defined under IEEE Std 762, including equivalent availability factor, equivalent forced outage rate and its demand-weighted variant, mean time between failures and mean time to repair, plus heat rate deviation and start reliability. Add three leading indicators that most contracts omit: safety metrics including TRIR and EMR, alarm performance, and procedure and training compliance rates.
Only what your contract says. Establish that the equipment register, work orders, inspection findings, condition data, operating logs and procedure library belong to you, are held in an exportable format, and are delivered at termination regardless of cause. Specify ISO 14224 as the failure data taxonomy at mobilisation so the records are comparable and usable. A contractor holding your history controls your next tender.
Typically 8 to 16 weeks from award to full shift coverage. The sequence runs procedure review and gap closure, operator recruitment and screening, site-specific training and plant familiarisation, supervised shadow shifts alongside the incumbent, competency assessment and sign-off, permit authority handover, CMMS and records transfer, then phased shift assumption. Compressing this is the most common cause of an early-term incident.
Keep them in-house when you already operate several plants, when your process knowledge is genuinely proprietary, or when operator continuity is itself the asset. Outsource when you have no operating organisation, when recruiting into your location is the binding constraint, or when you need round-the-clock coverage you cannot staff. The test is not cost. It is whether you can recruit, qualify, retain and supervise a shift roster indefinitely.
Tags: Plant Operations and Maintenance O&M Scope of Work Control Room Operations Alarm Management ISA-18.2 Operator Qualification
Power Utilities
20 minutes read
Fast-Track Power Plants Deployment: How to Reach Speed to Power in Under 12 Months
Fast-track power plant deployment can hit first power in weeks. See real TM2500 timelines from permitting to COD, proven in Oman and Taiwan. Plan your...
O&M Services
11 minutes read
Gas Turbine Outage Planning: Schedule, Scope Rules, and Checklist
Gas turbine outage planning starts 18 months out. Get the T-minus schedule, scope freeze rules, parts readiness gates and checklist. Talk to a Prismec...
O&M Services
15 minutes read
Predictive vs Preventive Maintenance: How to Choose Per Asset
Most plants apply predictive vs preventive maintenance facility wide and overspend. Match each asset by failure mode, criticality and P-F interval. Ta...
Data Centers
59 minutes read
Data Center Power Redundancy: N, N+1, 2N and 2N+1- What Each Level Takes to Build and Prove
Most data center power redundancy claims fail under test. See what N+1, 2N and 2N+1 truly require, how Level 5 testing proves them, and what to demand...