See It Work
See It Work
SYSTEM: OPERATIONAL OT/IT CONNECTORS: 150+ AUTONOMOUS OPERATION: 15+ DAYS GOVERNED AUTONOMY: ENFORCED AUDIT TRAIL: IMMUTABLE INDUSTRIES: ASSET-INTENSIVE & MISSION-CRITICAL DEPLOYMENT: 3-6 MONTHS VIA APEX CONTROL LOOPS: 3,400+ SYSTEM: OPERATIONAL OT/IT CONNECTORS: 150+ AUTONOMOUS OPERATION: 15+ DAYS GOVERNED AUTONOMY: ENFORCED AUDIT TRAIL: IMMUTABLE INDUSTRIES: ASSET-INTENSIVE & MISSION-CRITICAL DEPLOYMENT: 3-6 MONTHS VIA APEX CONTROL LOOPS: 3,400+

How Much Rigor Does Your Operational AI Need?

How Much Rigor Does Your Operational AI Need?

During workshop sessions with a large oil and gas company, we discussed how different artificial intelligence (AI) models could be used across operational use cases. Some applications allowed room to experiment with frontier models and change models regularly. Others needed managed versions, testing, and approval. For critical applications, we discussed applying the management-of-change principles used in control systems: once a model is approved and deployed, its version stays fixed until a formal process authorizes a change.

That conversation reminded me of a rigor model we had used years earlier. The required discipline varied with the consequences of the work. I think the same approach can help us make more explicit decisions about how we introduce and maintain AI in operations.

A model update changes something the operation depends on. Whether that calls for a few checks or a formal engineering review depends on what the model contributes and what could happen if its behavior changes.

Consider two illustrative uses. An engineer uses an AI model to explore alternative ways to present training material. The engineer can inspect the output, correct it, and try another model. Experimentation may be part of doing that work well.

Now consider a system that helps operations choose a maintenance window while preserving production requirements. Its recommendation depends on equipment condition, operating constraints, available resources, and the consequences of waiting. A different interpretation of any of those inputs could change the recommendation. People may still approve the final plan, but they need a sound basis for that approval.

Human approval does not automatically make a use case low consequence.

The person reviewing a recommendation must have enough information, time, and expertise to detect a problem. If that review cannot reliably catch an error before it affects the operation, the approval button provides less protection than the design assumes.

I would start with three proposed rigor levels: Basic, Enhanced, and High.

At Basic rigor, there is room to experiment with models and upgrade them with lightweight testing. The use has limited consequences, and mistakes are readily detected and corrected. Someone still owns the application. Its permitted use is clear. Changes are recorded, and checks are sufficient to establish that the revised version remains suitable for its task. Basic rigor gives experimentation a defined scope. It also establishes when a change in purpose requires a fresh assessment. Enhanced rigor applies where errors could cause material but bounded disruption. Models and versions are managed. An update has an identified owner, an assessment of what could change, and tests against the behavior the application requires. Material changes receive approval before use. The team can adopt a new model, provided it has the evidence needed to accept that change. At High rigor, the deployed system uses approved model versions. A new release from the provider does not become an approved release for the operation. The proposed change goes through formal management of change: assess its effects, test the affected behavior, confirm the controls, obtain authorization, and retain the evidence. The same discipline applies when the organization initiates the update itself.

The approved version remains in place until the change is approved.

These are proposed levels, with boundaries to agree for the operation concerned. Existing engineering and governance requirements continue to apply. The assessment must consider the whole decision process, including people, data, software, and downstream actions.

Freezing the model version addresses only one source of change. The system can also behave differently because someone changes its instructions, the tools it can call, the information it retrieves, or the constraints it applies. A rigor assessment needs to identify which of those changes affect the approved use and what review each requires.

This connects directly to the argument in my recent book, Software-Defined Operations: Change How Operations Work. The book describes how operational decision and coordination logic can be defined and maintained in software, with explicit connections to the systems that supply information and carry out actions.

Making that logic easier to change increases the value of knowing how each change must be governed. A revised requirement, a new model, and a fresh measurement have different implications. The decision definition helps identify what depends on each. Rigor determines the discipline needed to assess and accept a change.

The pumping-station example in the book illustrates the connection. Software helps compare pumping plans and maintenance windows while people retain approval. If the reserve requirement changes, the team must identify the affected evaluation, test the revised definition, and reconsider pending recommendations. A model update needs its own assessment of which parts of that process could be affected. The fact that the application remains advisory does not settle the required rigor.

I am developing this as one part of a broader classification structure for agentic operations: type, class, maturity, and rigor.

Type describes the system arrangement. The Digital Twin Consortium’s published types distinguish static automation, conversational agents, procedural workflow agents, cognitive autonomous agents, and multi-agent generative systems. The use case determines which arrangement is appropriate. Class describes the role. Creator and Operator distinguish producing an artifact, such as an engineering design or simulation model, from supporting or executing an operational decision. I have written about this distinction before; in this assessment, an Operator can remain advisory, with its authority recorded separately. Maturity describes the evidence. For the intended use, what is specified, what has been demonstrated, what has been validated, and what has been sustained in operation? The assessment compares that evidence with the requirements for the selected type, class, and rigor. Rigor describes the required discipline. It establishes the depth of validation, oversight, evidence, and change control needed for the decision’s consequences.

The type definitions come from the Digital Twin Consortium. The combined assessment structure is my proposal, and I will develop its application further in subsequent articles.

A more complex system still has to provide evidence that it is fit for its intended use. A Creator producing an engineering design may need High rigor. An Operator making recommendations may have no authority to execute them. Calling either system advanced tells the reviewer little about whether its requirements have been met.

For one use case you are considering, record the model version in use, the decision it contributes to, the consequences of an unsuitable recommendation, and the process required to change it. Include the person accountable for accepting the change. If the model provider could alter that behavior without your team applying the required checks, that dependency needs attention before you rely on the application.

My book explains how to make the decision definition and its dependencies explicit. You can download Software-Defined Operations here.

Related: Software-Defined Operations: Change How Operations Work

The next design decision is how much freedom to change that application should have. Some uses benefit from frequent experimentation. Others require an approved configuration to remain stable. Defining the required rigor gives the team a basis for making that choice before accepting the next model update.


Pieter van Schalkwyk is the CEO of XMPro, specializing in industrial AI agent orchestration and governance. Drawing on 30+ years of experience in industrial automation, he helps organizations implement practical AI solutions that deliver measurable business outcomes while ensuring responsible AI deployment at scale.

He authored the Industrial AI Agent Manifesto, published by the Digital Twin Consortium in February 2026, a governance framework for trustworthy autonomous operations

Our GitHub Repo has more technical information. You can also contact myself or Gavin Green for more information.

Read more on MAGS at The Digital Engineer


How this was made: I worked with Codex to check sources, draft from my direction and edit. I do not publish anything I have not rewritten myself.

This article originally appeared on XMPro CEO’s LinkedIn blog, The Digital Engineer, on 7 October 2026.