01
What is changing or failing?
Refineries are being offered assistants that can search procedures, summarise reports, draft work instructions and surface patterns across maintenance history. The capability is real. The operating evidence beneath it is often less ready than the demonstration suggests.
Equipment may have several names across the historian, CMMS and document archive. Procedures can remain searchable after they have been superseded. Work-order closure text may describe symptoms rather than verified causes. Technical decisions may sit in email or in the memory of an experienced employee.
An AI layer makes this information easier to reach. It does not automatically make it complete, current or authoritative.
The control requirement is not unique to AI. OSHA process-safety guidance expects changes to be reflected in process information, procedures, training and accessible documentation. NIST’s Generative AI Profile recommends documenting source provenance and testing outputs against defined risk tolerances. AI exposes the cost of weak information governance because it can distribute an old or conflicting record more persuasively and at greater speed.
02
Why does it matter commercially?
The first risk is not necessarily an obviously wrong answer. It is a convincing answer that removes the user’s reason to check the source. That can accelerate rework, inappropriate maintenance preparation or a decision based on the wrong equipment context.
The second cost is investment without adoption. If users repeatedly find conflicting answers, trust falls quickly. The organisation then carries the cost of data preparation, cyber review and implementation without changing the operating outcome.
The value case therefore depends on evidence quality and workflow fit—not model capability alone.
03
What must management decide?
Management must decide where AI may assist, which sources it may use and where accountable human authority remains mandatory. The starting point should be one valuable, bounded decision environment rather than unrestricted access to operational information.
The use case should also be separated into three levels:
- Retrieval: finding the relevant approved information and showing its source
- Synthesis: comparing records and making conflicts or missing evidence visible
- Recommendation: suggesting an action that still requires the named authority to decide
Each level needs a different evidence threshold, control boundary and benefit measure.
Score evidence readiness before model performance
For the selected corpus, sample records against six conditions: identifiable owner, current approved version, equipment or process context, effective date, superseded-record control and traceable change history. Report the proportion that passes each condition. That baseline tells management whether the next investment belongs in AI configuration or evidence remediation.
04
What evidence is required?
- A named operating decision, user and economic consequence
- The approved source hierarchy for procedures, equipment and prior decisions
- Version, ownership and review status for the selected document set
- Known conflicts across asset names, records and systems
- A visible citation path from every generated statement to its source
- Defined actions the system may draft, recommend or never perform
- Baseline time, rework or decision delay that the use case is expected to reduce
- Sample-based evidence-readiness score and the remediation owner for each failed condition
05
What should happen next?
- Choose one recurring decision or document set with a visible cost of delay or rework.
- Test the evidence before testing the model: ownership, version status, conflicts and missing context.
- Begin read-only and keep source references visible to the user.
- Run representative difficult cases, including contradictory and incomplete records.
- Scale only if the use case improves the named outcome without weakening technical accountability.
