AI9GM
Type to search documentation.

Set technical acceptance thresholds for a model

Draft

Allocation

L3-02
DecidesModel Owner
ConsultedBusiness Owner and CAIO
ExecutesML Engineering
EvidenceThreshold record with rationale, bound to the model version. Includes drift bounds, which define the delegated band at L3-10 and L3-11.

In plain terms

Decide what accuracy, fairness and drift figures count as good enough. Set before validation runs, or the threshold is chosen to fit the result.

What is being judged

What performance the decision downstream actually requires, expressed as numbers someone can test against.

Three properties.

Derived from consequence, not from capability. The threshold follows what the business decision needs and what failure costs the affected person, not what the model happens to achieve. A model that reaches 94 percent and a threshold set at 94 percent is a threshold set backwards.

Distributional, not aggregate. An overall accuracy figure conceals systematic error in a subgroup. Where the fairness dimension of materiality applies, thresholds are stated per relevant group or the number tells you nothing about the failure the framework is trying to prevent.

Including drift. The threshold record defines the delegated band at L3-10 and L3-11. Drift inside the recorded thresholds is operational and stays with the Model Owner; drift breaching them becomes a risk decision and leaves the layer. A threshold record without drift bounds leaves that band undefined.

What this decision does not cover

It does not accept the residual risk of the model performing at the threshold, which is L4-RSK-02. It does not authorize production use, which is L4-AUT-01. The Model Owner sets what good looks like technically and never signs off their own deployment.

When it fires

On event. Before validation of a model version. On a new version. On a change of purpose under L4-AUT-04, since thresholds set for one purpose do not transfer. On a classification change that alters consequence.

On cycle. None.

What you need before deciding

The decision the output influences and what an error costs the affected party. The consequence class. Which subgroups matter for the fairness dimension. Current performance of whatever performs the task today, including a human baseline where one exists.

How this goes wrong

Thresholds set after validation: the figures are recorded once results are known, which makes the record a description rather than a standard. The rationale field is the tell; a threshold with no rationale was usually reverse-engineered. Aggregate only: a single accuracy number for a system whose failures fall unevenly. No drift bounds: the band at L3-10 and L3-11 is undefined, so every drift observation becomes an ad hoc judgment.

Downstream L3-01 release to validation, L3-10 and L3-11 drift response, L4-AUT-01 authorization, L4-RSK-02 acceptance.

Upstream L4-CLS-04 classification.

Convention separation of build and acceptance, §0.1.

Instrument references

NIST AI RMF MEASURE 2 addresses performance and fairness measurement. ISO/IEC 42001 Annex A covers performance criteria. EU AI Act Article 15 requires accuracy levels be declared. None requires thresholds to be recorded before validation, which is the property that makes them a standard.

Correction

Correct L3-02

The maintainer answers corrections. There is no service level. Responses are best-effort and opportunistic within a reasonable time: a correction raised on a Monday is answered that week or sooner.

Attribution