The global knowledge network for professionals in the energy and industry

Why do assets fail? How to implement operational reliability

Failure causes, criticality, RCA, and RCM to define reliability strategies, reduce recurrences, and improve the performance of industrial assets.
operational reliability

When an asset fails repeatedly, increasing the frequency of maintenance does not necessarily eliminate the problem. The recurrence may be associated with a misidentified failure mode, an operating condition outside the expected limits, an inadequate strategy, or a repair that restored function without addressing the root cause.

Operational reliability requires reviewing the asset’s behavior within its service context, determining the consequences of a loss of function, and selecting actions commensurate with the risk. Its implementation depends both on the quality of the analysis and on the organization’s ability to incorporate its results into maintenance, inspection, operation, and life-cycle management.

Why assets fail and how to identify the failure mechanism

To understand why assets fail, it is necessary to distinguish between concepts that are often conflated in maintenance records. A functional failure refers to the inability to perform a required function; the failure mode identifies the event or condition that can cause that loss of function, while the cause and mechanism explain why it occurred and how the deterioration progressed.

In a pumping system, failing to achieve the required flow rate or pressure can constitute a functional failure. High vibration, increased temperature, or abnormal noise may indicate a degraded condition, but it is still necessary to determine whether cavitation, misalignment, bearing damage, erosion, lubrication problems, or operation outside the intended range is the cause. The strategy varies depending on the actual mechanism driving the failure.

For this reason, the asset hierarchy and its functional boundaries must be properly defined. Failure modes, causes, effects, and corrective actions must be associated with the equipment and the corresponding function. Otherwise, different events may end up being recorded under the same description and lose their usefulness for subsequent analysis.

Something similar occurs with damage mechanisms in process equipment. Corrosion, erosion, fatigue, creep, environmental cracking, or loss of thickness cannot be controlled simply by adjusting the frequency of maintenance. Materials, operating conditions, inspection history, and operational changes are all part of the assessment.

Operational reliability based on criticality

Operational reliability does not mean preventing all failures or applying the same level of maintenance to all equipment. Efforts should focus on those functions whose loss could have significant consequences for safety, the environment, production, quality, or cost.

The criticality assessment provides that distinction. A piece of equipment with numerous low-impact events may have a lower priority than another with few historical failures, but whose loss of function compromises a process unit or a safety barrier. The frequency of the event is important, but it does not replace an assessment of its consequences.

Criticality also helps identify assets that are prone to recurring issues, downtime, or high maintenance costs. So-called “bad actors” should be reviewed with sufficient context: number of failures, downtime, production losses, type of repair, and failure mode behavior. A list based solely on the number of work orders can lead to misplaced priorities.

The quality of the historical data influences the failure analysis

Equipment failure analysis depends on the quality of the records. Descriptions such as “repaired,” “mechanical failure,” “equipment shut down,” or “component replacement” allow a work order to be closed, but they contribute little to identifying reliability patterns.

ISO 14224:2016 provides a standardized framework for collecting data on equipment, failures, and maintenance in the oil, natural gas, and petrochemical industries. Its structure allows for the differentiation of data on equipment, failure modes and causes, consequences, maintenance actions, and downtime.

Metrics such as MTBF and MTTR help assess trends, but they must be interpreted within the context of the asset. MTBF provides insight into the average time between failures of repairable equipment, while MTTR provides information on the average time associated with repair or restoration, depending on the definition used by the organization. Two pieces of equipment with similar values may require completely different decisions if their functions and consequences are not comparable.

Before using these indicators to justify improvements, it is advisable to review the consistency of the historical data. Changes in nomenclature, duplicate orders, uncoded failures, repairs with no identified cause, or interventions associated with the wrong equipment can skew the conclusions.

Root cause analysis to prevent recurring failures

Root cause analysis is particularly useful for recurring failures, events with significant consequences, or situations in which the cause cannot be determined through a routine review. Its purpose is to explain the event with sufficient evidence to define a technically justifiable corrective action.

Methods such as Why-Why, cause-and-effect diagrams, fault trees, or causal analysis can be used depending on the complexity of the case. In industrial facilities, multiple factors are often involved: a mechanical condition may coincide with an operational deviation, a procedural deficiency, a maintenance practice, or a design limitation.

The investigation does not end once a probable cause has been identified. Recommendations must be incorporated into the work system, assigned, and verified after implementation to ensure that the condition that caused the failure has been eliminated or controlled. If the results are merely documented without modifying the strategy, procedure, design, or operational condition involved, recurrence remains possible.

operational reliability
Figure 1. Root cause analysis to identify causes, implement corrective actions, and prevent recurring failures.

Reliability-based maintenance and task selection

Reliability-based maintenance must address the behavior of failure modes and their consequences. The key question is not merely how often to service a piece of equipment, but rather what functional loss one aims to prevent, what mechanism could cause it, and what action offers a reasonable chance of controlling it.

RCM formalizes this analysis through functions, functional failures, failure modes, effects, and consequences. The current edition of SAE JA1011, revised in November 2024, establishes criteria for evaluating processes that are presented as Reliability-Centered Maintenance.

Based on the analysis, tasks can be defined, such as condition-based maintenance, scheduled repairs or replacements, troubleshooting hidden faults, or modifications when preventive maintenance does not provide an adequate solution. The selection depends on the failure mode behavior and the associated consequences, not solely on the frequency history.

A scheduled maintenance task does not become more effective simply because it is performed more frequently. When the failure mode is not meaningfully related to age, periodically replacing a component may add cost without significantly changing the probability of failure. In other cases, a condition-based maintenance task may be justified if there is a detectable potential failure condition and a P-F interval between the detection of a potential failure and functional failure.

Lifecycle asset management

Asset management broadens the decision-making process beyond the reliability of a single component. For static equipment, inspection results, damage mechanisms, corrosion rates, remaining life, and risk are all part of the assessment. For rotating equipment, failure history, condition, maintainability, and consequences also guide the definition of maintenance strategies.

Strategies established during commissioning should not be considered permanent. Changes in raw materials, capacity, temperature, pressure, chemical composition, operating conditions, or process modifications can alter degradation mechanisms and the criticality of the asset. A strategy that is valid under the original conditions may no longer be sufficient after several years of operation.

The necessary records must also be retained to reconstruct technical decisions. Information on design, materials, modifications, repairs, inspections, and operational changes is relevant when investigating a failure, evaluating a repair, or studying a life extension. The loss of that history forces one to work with assumptions and increases uncertainty.

Reliability strategies should be reviewed throughout the asset lifecycle, especially when there are changes in operating conditions, degradation, criticality, or failure history. This review helps determine whether existing tasks continue to control the mechanisms for which they were established.

How to implement operational reliability in an industrial facility

Implementation must be based on a consistent equipment hierarchy, defined functional boundaries, a criticality assessment, and a sufficiently reliable history to identify recurring issues and high-risk conditions. With this foundation, it is possible to determine which assets require a more in-depth analysis and which methodology is appropriate based on the identified technical problem.

The resulting actions may include changes to maintenance tasks or frequencies, condition-based monitoring, inspections, maintainability improvements, design modifications, or adjustments to operating conditions. Their definition must take into account the associated risk, available resources, and intervention windows, in addition to establishing how their effectiveness will be verified afterward. If failures associated with the same failure mode continue to recur, unavailability increases, or the risk remains above acceptable criteria, the hypothesis used, the selected strategy, or the quality of its execution must be reviewed. Operational reliability is built on this periodic review of the asset’s behavior, not on the isolated application of a single methodology.

From reliability recommendations to their implementation

Criticality, RCM, or RCA studies can generate numerous recommendations that must subsequently be incorporated into maintenance planning, maintained in relation to the asset and the evaluated failure mode, and tracked until they are implemented. This process becomes more complex in facilities with large numbers of pieces of equipment and multiple disciplines involved.

To manage this volume of recommendations, AsInt offers its Asset Risk & Reliability Suite, which includes applications for criticality analysis, RCM, and RCA. The solution can be supplemented with Recommendation Workbench+ (RWB+), which is used to manage the recommendations generated during assessments and track them through to implementation. This ensures that the actions defined in the reliability analyses remain linked to the assets and strategies that gave rise to them.

Implementation and continuity in asset management

In an interview with Inspenet TV, Michael Warren, founder and COO of AsInt, discusses the challenges that can arise during long-term implementations when personnel change or some of the accumulated knowledge is lost. His experience highlights the importance of maintaining continuity and auditability during the execution of integrity and reliability programs.

Conclusions

Recurring failures cannot be reduced simply by increasing inspections or maintenance. First, it is necessary to understand which function the asset has lost, how the failure occurred, and what consequences it has for operations. Criticality allows for the focused allocation of resources, while RCA, RCM, condition monitoring, and lifecycle management provide different tools depending on the problem.

Operational reliability is strengthened when strategies are based on evidence, incorporated into the work program, and reviewed in light of the equipment’s subsequent performance. In this way, maintenance is no longer limited to responding to breakdowns and can instead address the factors that truly influence the asset’s performance.

References

  1. International Organization for Standardization (ISO). ISO 14224:2016 – Petroleum, petrochemical and natural gas industries — Collection and exchange of reliability and maintenance data for equipment. 
  2. SAE International. SAE JA1011_202411 – Evaluation Criteria for Reliability-Centered Maintenance (RCM) Processes. 
  3. AsInt. Asset Risk & Reliability Suite. 
  4. AsInt. Recommendation Workbench+ (RWB+). 

Frequently Asked Questions (FAQs)

What is operational reliability, and how is it applied in an industrial plant?

Operational reliability aims to maintain the expected performance of assets by analyzing functions, criticality, failure modes, operational history, and risk-based maintenance strategies.

How does root cause analysis help reduce recurring failures?

Root cause analysis makes it possible to identify the technical, operational, human, or organizational factors that cause a failure and to define corrective actions that reduce the likelihood of recurrence.

When should reliability-based maintenance be used instead of preventive maintenance?

Preventive maintenance is typically performed at defined frequencies or intervals. Reliability-based maintenance is more appropriate when tasks must be selected based on the asset’s functions, its failure modes, consequences, and the behavior of each mechanism.

What metrics are used to assess the reliability of assets?

Among the most commonly used indicators are MTBF, MTTR, availability, failure recurrence, downtime, and production losses. When interpreting these indicators, one must always consider the asset’s function and criticality.

Verified Author

Mechanical Engineer with experience in the oil and gas sector, has technical skills in static equipment inspection, project control, development of work scopes and quality assurance. Contributes to the exchange of knowledge and best practices by writing technical articles related to the energy sector.