Analytics choice

SyncAI uses optional analytics and advertising measurement only after you allow it. Necessary site functions work without these trackers. Privacy details

Back to Insights
Reliability Engineering

FRACAS Is Not a Decision System

Closing the loop from failure code to approved action — why codes alone do not change reliability outcomes.

Mining reliability teams do not lack failure language. A haul truck trips, a crusher bearing runs hot, a conveyor tears, and someone enters a code. The computerized maintenance management system (CMMS) records a work order. A root-cause ticket may open. The week continues.

The code is not the decision. The work order is not the proof. And a Failure Reporting, Analysis, and Corrective Action System (FRACAS) that stops at either one is a reporting loop, not a decision system.

That distinction matters in asset-intensive operations because reliability outcomes do not move when events are classified. They move when a named person authorizes a change on the basis of evidence, then verifies whether the change did what it was supposed to do. Codes help you talk about failures. They do not, by themselves, change the failure.

What FRACAS actually is

FRACAS is a closed-loop reliability practice: report the failure, analyze it, take corrective action, and keep enough record that the organization can see whether the action prevented recurrence. DoD reliability program practice treated FRACAS as a requirement of MIL-STD-785. Uniform criteria were then written in MIL-STD-2155 (1985) and later issued as MIL-HDBK-2155, Failure Reporting, Analysis and Corrective Action Taken. The public ASSIST listing for the handbook is ident_number 207200.

The handbook is explicit about purpose. FRACAS exists to give management visibility and control for reliability and maintainability improvement by using failure and maintenance data to generate and implement effective corrective actions — and to reduce or simplify the maintenance task.

The intended sequence is not mysterious:

  • Failure reporting — what happened, on which item, under what conditions
  • Failure analysis — what the evidence supports as cause, and what it does not
  • Failure verification — confirm the reported failure is real and repeatable enough to act on
  • Corrective action — the change intended to prevent recurrence
  • Close-out — a record that the action was implemented and checked

MIL-HDBK-2155 also assumes a Failure Review Board: a human forum with authority to accept, reject, or return analysis and action. That is not a software feature. It is an accountability design. The loop is closed only when someone can answer, later, what was believed, what was authorized, and what was verified.

How mining reliability inherited a reporting loop

Few mines run a labeled “FRACAS program” with a standing Failure Review Board. Most run the fragments. Failure codes live in the CMMS. Analysis lives in a reliability spreadsheet, a contractor report, or a conversation on the radio. Corrective action lives as a work order, a setpoint change request, or a capital request. Verification — if it happens — lives in a later production meeting that no longer has the original evidence in view.

The result is a FRACAS-shaped workflow without the object FRACAS was built to produce: a reviewable decision. The mine can show that a code was entered and a task was completed. It often cannot show that the organization decided anything in a way that would survive a shift change, an audit, or a repeat failure six weeks later.

SAE reliability-program practice makes the same point without being mining-specific. SAE GEIA-STD-0009, the Reliability Program Standard, includes closed-loop feedback for corrective actions and field reliability monitoring. Its companion handbook, SAE TAHB0009A, describes that feedback method. Those documents are standards and practice context — not a SyncAI certification, and not a requirement to run a defense-style FRACAS office at a mine. They describe the industrial rule: a report that never returns as a checked action is not a closed loop.

Failure codes classify. They do not decide.

A failure code is a classification. At best it is a consistent label for a symptom, a mode, or a suspected cause. At worst it is the fastest pick-list item that lets the work order close before the next dispatch. Either way, the code answers a different question than the one the plant actually has to answer.

The code asks: How do we file this event?

The plant asks: What are we allowed to change, on what evidence, and who is accountable if we are wrong?

Those questions diverge under production pressure. A repeated “bearing failure” code on a pump family can mean a true bearing problem, a lubrication problem, a misalignment problem, a process-induced load problem, or a historian scaling problem that made a healthy machine look failed. The code collapses those possibilities into one string. A decision cannot.

If the record does not separate observed evidence from hypothesis, the code is a story the system will treat as fact.

That is why more complete coding schemes disappoint reliability leaders. Adding modes, mechanisms, and causes to the pick list increases the resolution of the archive. It does not add a governor. Nobody is forced to say what is proven, what is assumed, and what is still missing before the next action is approved.

Corrective action without a decision record

In the original FRACAS loop, corrective action is the change intended to prevent recurrence — not the repair that returns the asset to service. Replacing a failed component can be necessary work. It is not automatically corrective. If the condition that produced the failure is still in place, the organization has restored function and left the reliability problem intact.

Work orders are good at the restore-function half. They assign a craft, a duration, a part, and a completion stamp. They are weak at the prevent-recurrence half because completion is a task state, not a decision state. A closed work order proves that someone did the job as written. It does not prove that the job was the right intervention, that the diagnosis was evidenced, or that the outcome was checked against a defined signal.

A useful decision record, by contrast, has to hold at least four things:

  • Evidence — what the work history, condition data, inspection, or procedure actually shows
  • Hypothesis — competing explanations that remain unproven, stated as such
  • Named human approval — who accepted, rejected, escalated, or returned the recommended action
  • Verification — the signal that will show, later, whether the action worked

Without those four, “corrective action” is a label applied to activity. The FRACAS loop looks closed in the CMMS and remains open in the operation.

Recommend ≠ authorize

The gap gets wider when analysis is fast. A reliability engineer, a contractor, or a model can draft a recommendation in minutes: change the PM interval, replace the assembly, lower the trip setpoint, add a vibration route, park the unit. Drafting is cheap. Authorization is not — not on a haul fleet, a processing plant, or any asset whose failure has safety, environmental, or production consequence.

Recommend ≠ authorize. A recommendation is an argument. Authorization is an act by a person who can be named.

That is the industrial version of the Failure Review Board, whether or not the mine uses the term. Someone with operating authority has to accept the action, reject it, send it back for evidence, or escalate it. If the system cannot show who did that, the organization has a suggestion trail, not a governed decision.

This is also why unsupervised plant execute is the wrong default for reliability software. Writing a recommendation into a historian, a CMMS, or a control-system workflow is not the same as being allowed to change the plant. Approval, escalation, and accountability have to remain explicit. A tool that blurs recommend and authorize will look efficient in a demo and unaccountable on the night shift.

Closing the loop from failure code to approved action

A practical test for any mining reliability workflow — FRACAS software, CMMS module, or spreadsheet — is whether a later reviewer can reconstruct the decision without calling the original engineer.

If the record cannot answer these, the loop is still open

  • 1What failure was reported, and on which asset configuration?
  • 2What evidence was observed, versus hypothesized?
  • 3What competing causes remain unproven?
  • 4What action was recommended, and what was actually authorized?
  • 5Who approved it, and under what boundary?
  • 6What verification would show the action worked — and was it checked?

Those questions do not require a new standard. They are the FRACAS close-out discipline applied to the object mines already have: codes, analyses, and work. The missing piece is usually not another failure mode in the pick list. It is a decision record that can carry evidence grade, approval state, and verification criteria across shifts and systems of record.

When that record exists, failure codes become inputs instead of conclusions. Work orders become the authorized work, not the argument. Corrective action becomes a change that can be falsified. Mining reliability starts to look like what FRACAS was designed to be: a closed loop from event to approved action to checked outcome.

What this article is not claiming

This is an educational argument about a common reliability practice, not a customer case study. It does not report named plants, testimonials, or savings percentages. It does not treat a recommendation engine as live plant execution. Direct plant execute is not a capability being marketed here. Self-guided onboarding is not claimed as a live product path.

The public sources behind the FRACAS description are MIL-HDBK-2155 (historically implementing the MIL-STD-785 FRACAS requirement; earlier uniform criteria in MIL-STD-2155) and SAE closed-loop feedback practice in GEIA-STD-0009 and TAHB0009A. They are cited as reliability-program context, not as SyncAI certifications.

Bring a real reliability question

If you want to see how a decision record keeps evidence, hypothesis, and named human approval distinct, start in the Reliability Engineer workspace or request a Reliability Assessment. A bounded Strategic Pilot is the path when a specific workflow is ready to be operationalized with verification.