Classify an AI system by consequence
Draft
Classify an AI system by consequence, at or above materiality
Layer 4. Control CLSClassification and thresholds
Classify an AI system by consequence, below materiality
Layer 4. Control CLSClassification and thresholds
Allocation
| L4-CLS-04 | L4-CLS-05 | |
|---|---|---|
| Decides | AI Risk Committee | Business Accountable Executive |
| Consulted | Business Accountable Executive, Model Owner, CISO and General Counsel | Model Owner and CAIO |
| Executes | Model Owner | Model Owner |
| Evidence | Classification record with rationale and review trigger | Classification record proportionate to materiality |
| Delegated band | Delegated | Delegated |
In plain terms
Decide how badly this system’s failure would land, and therefore how much control it must carry. Materiality decides that the system gets classified. Consequence decides how much the classification requires of it.
What is being judged
Two questions in sequence, and they are frequently collapsed into one.
First: is this system material at all? Eleven dimensions, published at Layer 4 §5A, any one sufficient. Legal and regulatory, individual position, safety, employee, fairness, fundamental rights, privacy, autonomy, financial, operational, reputational.
Materiality is a gate, not a scale. A system that trips one dimension is material. It does not become more material by tripping four, and nothing in the framework averages the dimensions or weighs them against each other. This matters because the instinct is to score, and scoring produces a system that fails one dimension badly and passes the average.
Second: how severe is the consequence? Given that the system is material, how bad is the worst credible failure? This is what sets the control depth, through the mandatory control catalog.
The judgment is about the consequence of being wrong, not the probability of being wrong. A model that fails rarely and catastrophically is high-consequence. Reliability belongs to the risk acceptance decisions at L4-RSK-02 and L4-RSK-03, not here. Classifying on expected harm rather than potential harm produces a system that carries light controls precisely because it has not failed yet.
Three practical tests, in order of usefulness:
Can the affected person tell? A system whose error is invisible to the person affected is more severe than one they can see and dispute. Invisibility removes the correction path.
Is the outcome reversible? A wrong recommendation someone declines is not a wrong decision someone acts on. Irreversibility raises consequence independently of value, which is the same principle the agent transaction limits use.
Does anything else catch it? A model whose output passes through a human who could reasonably notice the error is lower consequence than one whose output is acted on directly. Note the word reasonably. A human who rubber-stamps 400 decisions a day is not a control.
The class scheme
Layer 4 §5B publishes the four-factor mechanism as normative and leaves the scheme to each organization. The following is an illustration only and carries no normative weight.
| Class | Condition |
|---|---|
| Contained | Material system where failure is reversible, observable by the affected party, and independently checked. |
| Consequential | Failure causes material harm and at least one of reversibility, observability or independent check is absent. |
| Critical | Failure could cause serious harm to a person, bears on safety or fundamental rights, or is irreversible with no independent check. |
Three classes is a starting point. A wide estate wants four or five; three systems want two. State the scheme in the control catalog, per L4-POL-03.
What this decision does not cover
It does not decide whether the system may run. That is L4-AUT-01.
It does not decide whether the residual risk is acceptable. That is L4-RSK-02 and L4-RSK-03.
It does not decide which specific controls apply. Classification selects a class; the control catalog at L4-POL-03 states what that class carries.
When it fires
On event. Before a system enters production. Before an existing system changes purpose, data scope, or the population it affects. On the review trigger recorded in the classification itself. When a system crosses a materiality dimension it previously did not trip, which happens most often through the autonomy and employee dimensions as scope expands.
On cycle. Existing classifications are reviewed on the cycle stated in the record.
A classification with no review trigger is a classification of a system as it was when someone last looked.
What you need before deciding
- The system’s purpose, and the decision its output influences
- Who is affected, and whether they can observe the output
- The data it consumes, and its classification
- Whether it acts, or only informs
- The Layer 4 materiality mechanism, current version
- The control catalog, so the consequence of a class is visible before it is assigned
How this goes wrong
Classifying by technology rather than by consequence. Large language model systems get classified as high-consequence because the technology is unfamiliar, while a rules-driven scoring system making the same decision about the same people gets classified low because it is old. Consequence follows the decision, not the implementation.
Classifying on current scope while scope expands. The system was classified when it served one team. It now serves four. Nothing re-triggered because nobody defined an expansion as a change of purpose.
Averaging the dimensions. See above. Materiality is a gate.
Collapsing both sides of the band. Where the Business Accountable Executive classifies a system as below materiality, that determination is itself the decision that avoided committee review. It carries a rationale for the same reason.
Related decisions
Upstream L4-CLS-02 defines the materiality mechanism. L4-CLS-03 determines whether formal governance applies at all. L4-CLS-06 names the accountable executive.
Downstream L4-POL-03 control catalog. L4-AUT-01 production authorization. L4-RSK-02 and L4-RSK-03 risk acceptance. L5-05 production readiness gate, which cannot pass without the authorization this classification shapes.
Related L4-REG-01 regulatory applicability, determined separately and frequently confused with classification.
Mental model
system → potential consequences → assess against eleven dimensions → material? → severity of worst credible failure → consequence class → required controls
The two arrows that get skipped are the third and the fifth. Skipping the third produces classification of everything. Skipping the fifth produces classification with no control consequence, which is a label rather than a decision.
Instrument references
EU AI Act Annex III defines high-risk categories, which is a classification of use case rather than of consequence. A system outside Annex III can be material under this mechanism.
ISO/IEC 42001 Clause 6.1 risk and opportunity, Annex A impact assessment controls.
NIST AI RMF MAP 1 and MAP 5, context and impact characterization.
No instrument publishes a consequence class scheme. AI9GM publishes the determination mechanism at Layer 4 §5B and deliberately publishes no scheme.
Correction
The maintainer answers corrections. There is no service level. Responses are best-effort and opportunistic within a reasonable time: a correction raised on a Monday is answered that week or sooner.