O&M Services
August 24, 2026
36 minutes read
A boiler feed pump trips at 2 a.m. and the operator on shift has fifteen minutes to decide whether to restart it or wake up the plant manager. The maintenance history in the CMMS, the last vibration reading, and the seal leakage log will decide which call gets made. That decision, and hundreds like it every month, is what rotating equipment maintenance actually governs.
Rotating equipment maintenance is the systematic inspection, lubrication, condition monitoring, and planned intervention on machinery with rotating components, including pumps, compressors, turbines, motors, and gearboxes, performed to prevent unplanned failure and maximize asset availability. It spans daily operator rounds, scheduled preventive tasks, predictive diagnostics, and major overhauls across the full asset lifecycle.
Most published guidance on this subject comes from repair shops or software vendors. This guide takes a different position: it is written from the perspective of teams that hold multi-year operations and maintenance contracts on gas turbine fleets across three continents, where the uptime number is a contractual obligation, not a marketing line. Everything that follows is what works when equipment availability is the deliverable and downtime carries a penalty clause.
The most effective maintenance strategy for rotating equipment is not a single strategy at all. It is a criticality-matched mix in which each asset receives the strategy its consequence of failure justifies: predictive maintenance for critical machines, time-based preventive maintenance for mid-tier assets, and deliberate run-to-failure for cheap, spared, low-consequence equipment. Plants that apply one strategy uniformly either overspend on monitoring or bleed availability through preventable trips.
The four strategies differ in when work is triggered. Corrective or reactive maintenance responds after failure. Time-based maintenance (TBM) intervenes at fixed calendar or run-hour intervals regardless of condition. Condition-based maintenance (CBM) intervenes when monitored parameters cross a threshold. Reliability-centered maintenance (RCM) is the analytical process, formalized in SAE JA1011, that assigns one of these strategies to each failure mode based on its consequence.
The distinction between TBM and CBM matters in practice. TBM replaces a bearing at 16,000 hours whether it needs it or not, and roughly a third of time-based tasks address failure modes that are not age-related at all, a finding that traces back to the original Nowlan and Heap reliability studies for United Airlines. CBM replaces that bearing when its vibration signature says stage-two defect, which typically recovers weeks of useful component life per intervention. The U.S. Department of Energy's Federal Energy Management Program, whose maintenance approaches guidance is published through Pacific Northwest National Laboratory, estimates that a functioning predictive maintenance program saves 8% to 12% over a preventive-only program, and that facilities heavily reliant on reactive maintenance can capture savings of 30% to 40% by moving to structured approaches.
Assigning strategies starts with a criticality analysis, and a workable version needs only three inputs. Score each asset from 1 to 5 on production impact if it trips, on safety and environmental consequence, and on spare availability, where 5 means no installed spare and long lead time. Multiply the three scores.
Assets scoring above roughly 60 justify continuous condition monitoring and a CBM program. Scores between 20 and 60 fit route-based monitoring plus time-based PM. Scores below 20 get operator rounds and basic lubrication, nothing more. A formal FMEA refines this later, but this screening pass sorts a 300-asset plant in two working days.
Run-to-failure is a legitimate engineering decision, not a program failure, when three conditions hold together. The asset must have an installed or shelf spare, its failure must carry no safety or environmental consequence, and the cost of monitoring must exceed the cost of the failure itself. A $900 sump pump with a spare on the shelf should never carry a $1,400 wireless vibration sensor.
Monitoring vendors rarely say this because it caps the sale. Field O&M contracts force the honesty: every sensor and every PM hour must earn back more availability than it costs. The strategy decision framework above is how that arithmetic gets applied asset by asset, and it feeds directly into the predictive maintenance economics covered later in this guide.
The P-F curve describes the interval between the point where a developing failure first becomes detectable (P, the potential failure) and the point where the equipment can no longer perform its function (F, the functional failure). That interval, the P-F interval, is the entire technical basis for condition monitoring. Every inspection frequency in a sound maintenance program is set shorter than the P-F interval of the failure mode it targets.
Different detection technologies sit at different points on the curve. Oil analysis and ultrasound detect bearing degradation earliest, often months before failure, because they catch sub-surface fatigue and microscopic wear debris. Vibration analysis detects the same defect weeks to months out, once spalling produces measurable defect frequencies. Infrared heat appears later, audible noise later still, and visible smoke or seizure marks the far end of the curve, where F is hours away.
The field implication is unforgiving. If a bearing failure mode has a P-F interval of six weeks and the vibration route runs every eight weeks, the program will miss failures while every reading looks compliant. Interval-setting against the P-F window, not against what the schedule can conveniently accommodate, separates monitoring programs that prevent trips from monitoring programs that document them.
This is also why the common claim that most rotating failures give two to eight weeks of warning is only half useful. The warning exists only for teams measuring the right parameter at the right frequency. A machine that fails "without warning" almost always gave warning in a spectrum nobody collected. This guide treats the P-F curve at working depth; for the full treatment of detection windows, technology selection, and program economics, see our dedicated analysis of predictive maintenance and the P-F window for power generation assets.
Effective pump maintenance concentrates on the two components that dominate repair statistics and the one operating condition that destroys both. Mechanical seal failures account for the majority of centrifugal pump repairs across industry reliability surveys, with bearing failures the second-largest category. Both failure modes trace back, more often than not, to how the pump is operated rather than how it was built.
Mechanical seals fail early for predictable reasons: dry running, seal flush interruption, operation far from the pump's design point, and misalignment that forces the seal faces out of parallel. API 682, the governing standard for seal systems in petroleum and heavy-duty service, exists largely because seal reliability depends on the flush plan and auxiliary piping as much as on the seal itself. When a seal fails twice in a year, the corrective action is rarely a better seal. It is usually a review of the flush plan, the suction conditions, or the operating point.
Bearing failures follow lubrication problems and load problems in roughly equal measure. Contaminated oil, wrong viscosity, over-greasing, and water ingress cover the first group. Misalignment, unbalance, and pipe strain transmitted through the casing cover the second. A bearing that fails at a fraction of its L10 design life did not die of old age, and replacing it without a root cause review guarantees a repeat.
Cavitation occurs when absolute pressure at the pump suction falls below the fluid's vapor pressure, forming vapor bubbles that collapse violently against the impeller. The prevention rule is fixed: net positive suction head available (NPSHa) must exceed net positive suction head required (NPSHr) by a safety margin, with 1 meter or 3 feet a common minimum in hydrocarbon service and more for hot water.
In the field, cavitation announces itself before the impeller shows damage. The pump sounds like it is passing gravel, discharge pressure flutters, and vibration rises broadband rather than at discrete frequencies. Operators who log those three symptoms on daily rounds catch cavitation weeks before an impeller inspection would, which is why the symptom belongs on the operator checklist, not just in the engineering manual.
Spare parts discipline decides how fast this schedule recovers a failed pump. Seals, bearings, and gaskets for critical pumps belong on the shelf before the work order drops, because parts lead time, not wrench time, drives most repair duration. Structured industrial supply chain programs exist precisely to keep those long-lead spares stocked against the asset register rather than ordered after the failure.
Every centrifugal pump has a best efficiency point (BEP), and reliability degrades measurably as operation moves away from it. The Hydraulic Institute's guidance places the preferred operating region between roughly 70% and 120% of BEP flow. Below that band, internal recirculation loads the shaft radially, seals see deflection, and bearings see forces the L10 calculation never assumed.
The practical consequence: a pump throttled to 40% of BEP flow for months will consume seals regardless of maintenance quality. Fixing that pump is an operations conversation about trimming the impeller, adding a recirculation line, or resizing the pump, not a maintenance conversation about seal brands. The maintenance team that brings the pump curve to that conversation saves more seals than any PM task on the schedule.
Compressor maintenance requirements divide sharply by machine type. Reciprocating compressors fail through valves, rings, and packing on run-hour patterns that suit time-based intervals. Centrifugal compressors fail through seals, fouling, and aerodynamic instability, which demand continuous monitoring. Screw compressors live or die on oil condition. A single PM template applied across all three wastes money on some machines and misses failures on others.
Compressor valves are the highest-frequency maintenance item on any reciprocating machine, and valve condition is readable without opening the cylinder. A leaking discharge valve shows up as elevated valve cover temperature and a change in the pressure-time card where analyzers are fitted. Rod packing leaks show as measurable leak-off flow, and rod drop monitoring, a proximity measurement of piston rider band wear, tells the crew when the cylinder needs opening before the piston contacts the liner.
Typical intervals in process gas service run 4,000 to 8,000 hours for valve inspection or replacement, with packing renewed on measured leakage rather than on the calendar. API 618 governs reciprocating compressors in petroleum service and is the specification reference for allowable rod loads and monitoring provisions. Machines in hydrogen or sour service sit at the short end of every interval because consequence, not just wear rate, sets the schedule.
Surge is a flow reversal through a centrifugal compressor that occurs when discharge system resistance exceeds what the machine can develop at its current speed. The reversal happens in fractions of a second, produces an audible bang, and drives axial loads that can wreck thrust bearings and seals in a handful of events. Anti-surge control systems recycle gas to keep the operating point away from the surge line, and no maintenance or operations shortcut ever justifies bypassing them. The opposite limit, stonewall or choke, caps flow at high volume and is less destructive but equally a sign the machine is mismatched to its current process conditions.
Dry gas seals (DGS) replaced oil bushing seals as the standard on process centrifugal compressors because they eliminate seal oil systems and cut leakage to a fraction. They demand clean, dry seal gas, and most DGS failures trace to seal gas conditioning, not the seal cartridge. Owners still running oil bushing seals on older machines face a genuine upgrade decision: the DGS conversion requires engineering study of rotor dynamics and seal gas supply, but it typically removes an entire auxiliary system's worth of failure modes. That evaluation sits comfortably within OEM-agnostic operations and maintenance scopes, where the recommendation is not tied to selling either outcome.
Performance trending deserves more attention than it gets on centrifugal machines. A polytropic efficiency trend falling 3 to 4 points usually means fouling, and an online wash or planned cleaning recovers energy cost that dwarfs the maintenance cost. Crews that trend only vibration catch mechanical failures and miss the slow economic ones.
Gas turbine maintenance follows a structured inspection hierarchy with intervals set by the manufacturer in fired hours and starts, escalating from combustion inspections through hot gas path inspections to major overhauls. Steam turbine maintenance follows outage-based inspection cycles driven by blade condition and steam purity. Both machine classes reward disciplined interval tracking more than any other rotating asset, because a missed inspection window on a turbine risks components whose replacement cost runs into millions.
The gas turbine inspection sequence has three levels. A combustion inspection (CI) addresses combustor liners, fuel nozzles, and transition pieces, the components that see the highest gas temperatures and degrade first. A hot gas path (HGP) inspection extends into the turbine section, covering first-stage nozzles and buckets, and includes the CI scope. A major overhaul opens the full machine, including the compressor and rotor, and includes everything below it.
Intervals are counted in equivalent operating hours or factored fired hours, not raw clock hours, because starts, trips, peak firing, and fuel type all consume component life faster than steady baseload running. On heavy-frame machines, a common OEM pattern places combustion inspections in the range of 8,000 to 12,000 factored hours, HGP inspections around 24,000, and major overhauls around 48,000, with each start typically counted as equivalent to several fired hours. Peaking units therefore reach inspections on starts long before they reach them on hours, a distinction that catches out operators who track only the hour meter.
Aeroderivative machines run a different philosophy entirely. On TM2500 and LM2500 class units, the maintenance strategy centers on engine exchange: a degraded gas generator is swapped for a lease or spare engine in days, and the removed engine goes to a depot for overhaul. On a 110 MW TM2500 fleet operated under a multi-year O&M contract in Duqm, Oman, this swap-based approach is what keeps grid-commitment availability achievable in a coastal desert environment, where airborne salt and dust would punish any plan built around long in-situ repairs. The trade-off is structural: aeroderivative operators buy availability with engine pool access, while heavy-frame operators buy it with longer intervals and on-crank repairs.
A borescope inspection examines internal turbine components through access ports without disassembly, and it is the single highest-value diagnostic on any gas turbine. Standard scope covers combustor liners for cracking and cooling hole blockage, transition pieces for wear at the aft frame, first-stage nozzles for trailing edge cracks, and first-stage buckets for coating condition.
Coating condition is the finding that most often changes an outage plan. Thermal barrier coating (TBC) spallation on a bucket leading edge exposes base metal to gas temperatures it cannot survive for long, and spallation found at a borescope typically pulls the HGP inspection forward rather than waiting for the interval. Crews photograph and trend every finding against the previous inspection, because the rate of change, not the snapshot, drives the intervention decision.
Steam turbine reliability rests on steam purity and blade condition. Solid particle erosion from oxide scale attacks control stage blading, water induction bends rotors, and stress corrosion cracking targets low-pressure blade roots and discs, particularly in the phase transition zone where moisture first forms. Gland seal condition and steam chemistry logs deserve the same trending discipline as vibration, because most catastrophic steam turbine failures announce themselves first in chemistry excursions rather than in mechanical data.
A long-term service agreement (LTSA) with the OEM buys guaranteed parts access, fleet-wide engineering data, and interval coverage at a premium price and with limited flexibility. Third-party, OEM-agnostic maintenance buys lower cost, multi-vendor parts sourcing, and contract flexibility, in exchange for the owner carrying more responsibility for parts strategy and engineering oversight. Neither answer is universally right, and the full decision framework, including when to split scopes or exit an LTSA at renewal, deserves its own analysis: see our dedicated guide on choosing between OEM and third-party turbine O&M.
Motors, gearboxes, fans, and blowers make up the numerical majority of rotating assets in most plants, yet they receive the least structured maintenance attention because no single failure feels catastrophic. The failure data says otherwise: these balance-of-plant machines generate most reactive work orders by count, and their maintenance follows the same criticality logic as everything above, applied with lighter tools.
Electric motor reliability turns on two practices. Greasing discipline comes first: over-greasing forces grease past the bearing shields into windings, and it kills as many motor bearings as neglect does, so quantity and interval come from the lubrication chart, not habit. Insulation condition comes second, trended through periodic insulation resistance testing, because heat and moisture degrade windings years before failure. On VFD-driven motors, add bearing currents to the watch list: without proper grounding or shaft grounding rings, current discharge through the bearings produces the characteristic fluting damage that no lubrication program prevents.
Gearboxes are oil-analysis machines above all else. Iron and particle-count trends in the oil reveal gear tooth wear months before vibration does, which makes a quarterly sample the single highest-value PM task on any critical gearbox. Breather condition and seal integrity decide how fast water and dirt contaminate that oil, so both belong on the same PM.
Fans and blowers fail through imbalance more than anything else, because process buildup on blades shifts mass distribution continuously. A rising 1X vibration trend on a fan almost always means cleaning before it means bearings, and blade inspection for erosion, corrosion, and cracking at every planned stop closes the loop. All three asset classes slot into the troubleshooting table and checklist in this guide; the techniques transfer directly, only the criticality scores differ.
Rotating equipment troubleshooting follows a fixed logic: match the symptom to its most probable causes, verify with the cheapest checks first, and correct the cause rather than the damaged component. Vibration frequency content carries most of the diagnostic information, because each mechanical fault produces energy at characteristic multiples of running speed. A technician who can read that pattern resolves most problems without opening the machine.
ISO 20816, which superseded ISO 10816, provides the vibration severity zones that turn an absolute reading into a judgment: Zone A and B machines run, Zone C machines get scheduled intervention, Zone D machines come down. The zones are the common language between the analyst, the planner, and the operations manager, which is why every reading in the CMMS should carry its zone, not just its number.
A worked example shows the sequence. A gas compressor's monthly route shows 2X vibration trending up over three readings, now at the Zone B/C boundary. The order of checks matters because it runs cheapest to most invasive.
First, inspect the coupling with the machine down under lockout: element cracking or wear debris confirms misalignment is doing damage. Second, perform a laser shaft alignment check and record as-found values against the tolerance for the machine's speed. Third, test for soft foot at each hold-down bolt, because alignment corrections made on a soft-footed machine do not hold. Realign, torque to specification, and take a verification reading at the next run. If 2X persists after a verified alignment, the search widens to casing distortion from pipe strain, which no amount of realignment fixes.
Root cause analysis is the discipline of asking why the failed component failed, repeated until the answer is a process, practice, or design condition that can be corrected. A bearing that failed at 4,000 hours against a 40,000-hour design life is evidence, not a conclusion. Why did it fail? Vibration history shows sustained 2X. Why misaligned? Alignment was done at commissioning in winter and never rechecked. Why never rechecked? No PM task existed for thermal-growth realignment.
The corrective action that ends the sequence is a new PM task and a revised commissioning procedure, not a bearing brand change. Plants that skip the last two whys replace the same components on the same machines every year and call it maintenance.
A condition monitoring stack earns its cost only when each technology is deployed against assets whose criticality and failure modes justify it. Continuous vibration monitoring belongs on assets where the criticality score is high and the P-F interval is short, because a monthly route would miss the detection window. Route-based monitoring covers the mid-tier at a fraction of the cost. Everything below that tier is covered adequately by trained operator rounds.
The decision keys directly to the criticality ranking built earlier in this guide. Assets protected by permanent machinery protection systems, specified under API 670 for critical turbomachinery, already justify continuous monitoring by consequence alone. For balance-of-plant machines, the honest comparison is the cost of continuous sensors against the cost of the failures a monthly route would statistically miss. On a spared cooling water pump, that comparison rarely favors sensors. On an unspared boiler feed pump, it almost always does.
Wireless sensor pricing has moved the breakeven point down the criticality list in recent years, but it has not eliminated the tier where route-based collection wins. A defensible program states, asset by asset, which tier applies and why. A program that cannot state that has been sold, not designed.
An oil analysis report answers three questions: is the machine wearing, is the oil degrading, and is the oil contaminated. Wear metals answer the first, and the elements identify the source: iron indicates gears or cylinder surfaces, copper and lead point to bearing babbitt or bushings, chromium to rings or rolling elements. Ferrography and particle counting, reported against ISO 4406 cleanliness codes, size and count the debris, and a rising code trend matters more than any single result.
Viscosity change answers the second question: oxidation thickens oil, fuel dilution thins it, and either outside roughly 10% of the grade nominal warrants action. Water content answers the third, and in gearbox or turbine lube service, results above a few hundred ppm mean the machine is making water somewhere, usually through a cooler leak or breather failure. A maintenance planner who reads these three lines can act on a lab report without waiting for anyone's interpretation.
Infrared thermography finds heat where it should not be: a coupling running hot from misalignment, one bearing warmer than its twin, a motor connection resisting its way toward failure. It surveys fast and wide, which makes it the ideal route-day companion to vibration collection. Ultrasound earns its place in two roles: earliest-stage bearing detection, and finding compressed air and gas leaks, where the energy savings alone typically fund the instrument in its first year.
A rotating equipment maintenance checklist works only when every line pairs a task with an acceptance criterion, so the person executing it records a judgment, not just a tick. The master checklist below is organized in four tiers and tagged by equipment class: [P] pumps, [C] compressors, [T] turbines, [M] motors, [G] gearboxes, [F] fans and blowers, [All] every rotating asset. Adapt intervals to criticality; the structure holds across industries.
A printable PDF version of this checklist, formatted for clipboard use on rounds, is available for download from this page. Print it, mark it up against your own asset list, and treat every crossed-out or added line as data about your plant that no generic checklist can supply.
No maintenance task on rotating equipment begins until the energy that turns it is isolated, locked, tagged, and verified at zero. That sequence is codified in OSHA 29 CFR 1910.147, the control of hazardous energy standard, and its steps are fixed: identify all energy sources, isolate them, apply personal locks and tags, release or restrain stored energy, and verify isolation with a try-start before any hand enters the machine. The verification step is the one that saves lives, because breakers get mislabeled and valves get passed.
Rotating machinery holds energy after isolation. A large rotor coasts for many minutes after trip, discharge piping stays pressurized behind a check valve, and hot casings hold burn-injury temperatures for hours. Stored energy release belongs in every isolation plan, not just electrical lockout, and turning gear on turbines adds a specific hazard: a shaft that appears stopped may be commanded to rotate.
Guarding is the permanent control between people and couplings. A coupling guard removed for an alignment check goes back before the motor is unlocked, without exception, and no PM task, schedule pressure, or troubleshooting shortcut ever justifies bypassing an interlock or trip function. In process plants, the mechanical integrity element of OSHA 29 CFR 1910.119 makes documented inspection of covered compressors and pumps a regulatory obligation, which means the maintenance records this guide describes are also the audit evidence.
Confined space entry and hot work carry their own permit systems, and both intersect rotating equipment routinely: a compressor cylinder is a confined space, and a weld repair on a base frame is hot work beside lube oil. The permit is not paperwork around the job. It is the job's first task.
A maintenance program's health is measurable through six indicators: MTBF, MTTR, availability, PM compliance, the planned-to-reactive work ratio, and schedule compliance. Together they answer the only questions that matter: is equipment failing less often, is it recovered faster when it fails, and is the team working to plan rather than to the loudest alarm. Programs that track none of these cannot demonstrate improvement, and on O&M contracts where availability carries liquidated damages, these numbers are the contract.
Mean time between failures (MTBF) is total operating time divided by the number of failures over the period, tracked per asset and per asset class. Its value is in the trend and in bad-actor identification: the five machines with the worst MTBF typically consume a third of the reactive budget. Mean time to repair (MTTR) is total repair downtime divided by the number of repairs, and it exposes logistics problems, because most long MTTR traces to parts and permits, not wrench time.
Availability is the percentage of required time the asset was capable of running, and it is the number operations and commercial teams actually feel. Well-run generation assets sustain availability in the mid-to-high 90s; the gap to 100 decomposes into planned outage, unplanned outage, and that decomposition is where the improvement targets live. The remaining three KPIs measure the program rather than the assets. PM compliance is the percentage of scheduled PM completed within its window, with 90%-plus the working standard. The planned-to-reactive ratio should sit near 80/20 or better in a mature program, per benchmarking bodies such as SMRP, and schedule compliance measures whether the weekly plan survived contact with the plant.
One caution from contract experience: PM compliance is the easiest KPI to game. A schedule padded with low-value tasks reports 98% compliance while the bad actors keep failing. The cross-check is the planned-to-reactive ratio, which cannot be gamed, because unplanned work orders write themselves.
Annual maintenance spend on rotating equipment is best benchmarked as a percentage of replacement asset value (RAV), and industry bodies including SMRP place a healthy range at roughly 2% to 5% of RAV per year depending on industry and asset intensity. A plant spending under 2% is usually deferring maintenance rather than outperforming, and the deferral surfaces later as unplanned outage. A plant spending over 5% is usually paying for reactive work at reactive prices.
The reactive premium is the central cost fact of this field. Industry studies consistently place the cost of an unplanned repair at three to five times the cost of the same work performed planned, because emergency work carries expedited parts freight, overtime and call-out labor, collateral damage from run-to-destruction, and production loss on top of the repair itself. That multiplier, not the maintenance budget line, is what a maintenance strategy actually manages.
Downtime cost is context-dependent, and the arithmetic should be done per asset rather than quoted as a universal figure. A 30 MW gas turbine selling into a market at $60 per MWh forgoes $1,800 in revenue per hour offline, or over $43,000 per day, before any repair cost. A fully spared transfer pump in the same plant costs nothing per hour of downtime beyond its repair, which is exactly why the criticality framework earlier in this guide assigns them different strategies and different spend.
For the maintenance manager defending a budget, the cost-avoidance case writes itself from these numbers. Take the plant's reactive work orders from last year, price each at the planned-work equivalent, and the difference is the annual cost of the current strategy mix. A proposal that shifts even a fifth of that reactive work to planned execution typically pays for the condition monitoring and planning resources it requests. Where the constraint is capital rather than conviction, structured financing solutions for equipment and program investment change the timing of the spend without changing the arithmetic.
Rotating equipment maintenance operates inside a framework of API and ISO standards, and knowing which standard governs which question saves hours of argument in procurement, audits, and warranty disputes. The table below maps the standards a maintenance or procurement engineer will actually encounter, with the scope each one owns.
Standards matter in practice for one reason: they convert opinion into specification language. A purchase order that requires seal systems per API 682 and installation per API 686 is enforceable in a way that "good quality installation" never is, which is why EPCM specifications cite them by number. In an audit or a failure dispute, the standard is the neutral referee, and the party whose records align with it usually prevails.
The maintenance-facing standards deserve equal attention to the design-facing ones. ISO 14224 failure coding is what turns ten years of work orders into analyzable reliability data, and plants that adopt its taxonomy on day one of a CMMS implementation avoid the unusable free-text history that plagues most failure analysis efforts.
The repair-or-replace decision on a damaged machine turns on four factors: repair cost against replacement cost, remaining service life of the repaired asset, lead time for each path, and whether an overhaul offers an efficiency upgrade worth capturing. A widely used screening threshold puts the repair ceiling at roughly 50% to 60% of replacement cost, but lead time overrides that arithmetic more often than any other factor.
Lead time is the honest driver in most real decisions. A new OEM impeller may quote at months; a reverse-engineered impeller, manufactured from a single forging by 5-axis milling, can arrive in weeks and often exceeds the original's dimensional consistency. The modern repair toolkit is broad: localized weld repair, HVOF hard-face coatings such as tungsten carbide for erosion protection and dimensional restoration, and full reverse engineering of rotors, blades, and diaphragms. Each option makes sense when it recovers the machine faster or cheaper than replacement without compromising remaining life, and an owner's engineer should price all three paths, not just the OEM quote.
Overhaul windows are also upgrade windows. A compressor already open for a damaged impeller can receive improved vane geometry and gain gas path efficiency; a steam turbine repair after a blade failure should include a finite element reevaluation of the blade design rather than a like-for-like copy of a part that already failed. Obsolescence weighs on the same scale: when the OEM no longer supports the frame, every repair buys time on a machine whose parts problem is worsening, and replacement planning should start before the next failure forces it. Where replacement lead time is itself the crisis, ready-to-ship equipment inventory exists precisely to compress a months-long procurement into weeks.
When replacement wins, the displaced machine is not scrap. Surplus rotating equipment holds real market value to operators running the same frame, and structured equipment marketing and disposition recovers capital that offsets the replacement cost. The full decision, properly made, prices all of it: repair, replace, upgrade, and recovery value, against the downtime clock.
Most rotating equipment reliability is decided before the first PM task ever runs, at installation. API 686 exists because grouting quality, baseplate flatness, pipe strain, and precision alignment at commissioning set the vibration baseline a machine carries for life. A pump pulled into its piping during construction will consume seals and bearings for a decade, and no maintenance program downstream fully compensates for it. The cheapest reliability investment any project makes is enforcing the installation standard while the contractor is still on site.
The second execution factor is craft skill, and lubrication is where the gap shows first. Over-greasing kills as many motor bearings as under-greasing, wrong-viscosity top-ups quietly degrade gearboxes, and none of it appears in any spectrum until damage is underway. Field observation across fleets is consistent: the plant's bad actor list traces to installation defects and lubrication practice far more often than to design or component quality. Training the people who touch the machines outperforms upgrading the components they touch.
The third factor is the execution system. A CMMS earns its keep by doing three things without exception: triggering PM from run hours and condition rather than only the calendar, capturing a failure code on every corrective work order, and holding history per asset so trends are visible. Which software does this matters far less than whether the discipline is enforced, because a CMMS full of "pump broke, fixed pump" entries is a filing cabinet, not a reliability tool. Programs fail on execution discipline long before they fail on technical knowledge, and the fix is management attention, not another technology purchase.
The three core types are corrective maintenance, performed after a failure; preventive maintenance, performed at fixed time or run-hour intervals; and predictive maintenance, performed when condition monitoring shows a developing fault. Most modern frameworks add reliability-centered maintenance (RCM) as the analytical method that assigns one of the three to each asset based on failure consequence. Mature plants run all three deliberately, matched to asset criticality.
No single strategy is most effective across a whole plant. Predictive, condition-based maintenance delivers the lowest total cost on critical, unspared machines because it catches faults inside the P-F interval. Time-based preventive maintenance suits mid-criticality assets with age-related wear, and run-to-failure is economically correct for cheap, spared, low-consequence equipment. Effectiveness comes from matching strategy to criticality, not from picking one strategy.
TBM is time-based maintenance: intervention at fixed calendar or operating-hour intervals regardless of machine condition. CBM is condition-based maintenance: intervention triggered when a monitored parameter, such as vibration, oil condition, or temperature, crosses a defined threshold. CBM generally recovers component life that TBM discards, but it requires monitoring infrastructure and analysis capability that TBM does not.
Rotating equipment covers industrial machinery whose function depends on rotating components: pumps, compressors, gas and steam turbines, electric motors, generators, gearboxes, fans, and blowers. The category is defined by shared failure mechanisms, including bearing wear, misalignment, unbalance, and lubrication degradation, which is why the same maintenance techniques apply across all of it.
Inspection frequency follows a tiered pattern: daily operator rounds for leaks, noise, levels, and pressures; monthly technician PM for vibration readings against ISO 20816 zones, thermography, and seal checks; and annual or outage-window inspections for alignment verification, internal clearances, and performance testing. Critical machines add continuous monitoring, and every interval must be shorter than the P-F interval of the failure mode it targets.
Across industry reliability data, the leading causes are mechanical seal failures on pumps, bearing failures driven by lubrication problems or misalignment, and operational causes such as running far from the design point, cavitation, and frequent starts. The common thread is that most failures are induced by installation, lubrication, or operating practice rather than by component quality, which is why root cause analysis matters more than component replacement.
One decision test carries this entire guide into Monday morning. For any rotating asset in the plant, can the team answer three questions: what strategy is this machine assigned and why, what is the next warning its dominant failure mode will give, and who acts on that warning within what window? Where all three have answers, uptime is being built. Where any of them draws a blank, that asset is running on luck, and luck has an MTBF too.
We wrote this guide from the operating side of the fence. Prismecs runs OEM-agnostic O&M crews under multi-year contracts on gas turbine and balance-of-plant fleets worldwide, with CMMS-driven programs held to contractual availability numbers. If you want that experience applied to your own assets, from a single problem machine to a full-plant maintenance program, talk to our team and get a custom maintenance plan.
Partner with Prismecs to maximize uptime and extend the life of your rotating assets. To avail of our services, call us at +1 (888) 774-7632 or email us at sales@prismecs.com.
Tags: rotating equipment reliability predictive maintenance condition monitoring pump and compressor maintenance gas turbine maintenance
O&M Services
36 minutes read
Rotating Equipment Maintenance: A Field Guide to Uptime for Pumps, Compressors, and Turbines
Get rotating equipment maintenance right: daily to annual PM checklist, cost benchmarks, and troubleshooting tables built by O&M crews. Download the P...
EPCM Services
21 minutes read
EPC vs EPCM: Which Model Fits Your Project's Risk and Control Needs?
EPC transfers risk for a lump-sum premium; EPCM keeps you in control but on the hook. See the numbers, the failure modes, and score which model fits y...
O&M Services
16 minutes read
Predictive Maintenance for Power Assets: How It Works, Challenges, and Real Applications
Predictive maintenance cuts power plant downtime, but only inside the P-F window. See fault signatures, honest ROI math and false alarm traps. Read th...
Renewables
17 minutes read
Renewable Energy Integration: Challenges, Technologies, and How the Grid Adapts
Renewable energy integration breaks grids built for baseload. See how inertia loss, duck curves, and storage decide reliability, with field-proven fix...