AI9GM
Type to search documentation.

Maturity

Scoring

Maturity is scored per layer. Composite organizational scores are not produced under this specification.

The five levels

These levels are a reference model. No conformance regime exists at v0.9, and no organization, individual or engagement is described here as certified, accredited, assessed or compliant (AI9GM-v0_9.md section 7).

LevelDescriptor
1 · InitialGovernance occurs per project, if at all. No inventory of AI systems.
2 · ManagedGovernance exists where someone insisted on it. Inconsistent between teams. Partial inventory.
3 · DefinedDocumented decision rights and required artifacts across all six layers, applied consistently.
4 · Quantitatively ManagedGovernance effectiveness is measured per layer. Deviations are detected, not discovered.
5 · OptimizingGovernance adapts from measured outcomes. Controls automated. The pipeline generates the evidence.

Level 2 is where the minimum viable set applies. See the minimum viable set for the artifact types a small organization actually maintains.

Layer 1

Foundation

LevelDescriptor
1 · InitialAI workloads run wherever capacity was available when the project started. No capacity model. Models and datasets absent from the CMDB. Recovery for AI systems untested and probably undefined. Infrastructure learns about a new production model when it saturates something.
2 · ManagedCapacity is tracked and forecast for known workloads. AI services appear in the service catalog, though objectives are inherited from general IT rather than set from consequence. Some models are configuration items. Recovery plans exist and cover data but not model artifacts. Placement decisions are made and recorded, inconsistently.
3 · DefinedPublished capacity model covering accelerator, storage and network, with stated lead times. All production models and datasets are configuration items. Service level objectives are set per AI service from its consequence classification. Recovery plans include model artifacts and feature stores, tested annually. Placement, expansion and allocation decisions follow the register in section 5 with retained evidence.
4 · Quantitatively ManagedCapacity headroom, availability and restore time are measured per service and reported to Layer 4 on a fixed cycle. Deviation from objective triggers a defined response before it becomes an incident. Cost is attributed per model. Recovery objectives are validated by test, not asserted. Capacity constraints are published to Layer 5 in a form the portfolio schedule consumes.
5 · OptimizingCapacity is provisioned from forecast rather than from request. Configuration state is generated by the platform rather than maintained by hand, so CMDB coverage is a property of deployment rather than a compliance exercise. Recovery is exercised continuously rather than annually. Cost and carbon per inference are measured and fed to Layer 6 as inputs to strategy.

The distance between levels 2 and 3 is the largest in this layer and the one most organizations underestimate. It requires treating model artifacts as first-class configuration items, which usually means changing the deployment pipeline rather than changing the CMDB.

Layer 2

Structural

LevelDescriptor
1 · InitialIntegration happens per project, in whatever pattern the delivery team preferred. No standards register or one that nobody consults. Interface specifications are generated from code where they exist at all. Nobody can list which interfaces an AI system calls. Third-party model APIs enter through a corporate card.
2 · ManagedAn architecture function exists and reviews significant changes. A standards register exists and is mostly current. Interface specifications are published for external consumers and inconsistent internally. Exceptions are granted and recorded, with expiry dates that are not enforced. Some AI data access runs through catalogued interfaces, some directly against stores.
3 · DefinedThe standards register is authoritative and consulted before build. Every interface an AI system consumes is catalogued and carries a fitness record. The interface contract standard is published and applied, so schemas carry descriptions, examples and documented errors. Exceptions carry expiry and a named remediation owner, and those at or above the risk threshold route to Layer 4. Third-party model admission runs through the Architecture Review Board.
4 · Quantitatively ManagedSchema and error documentation coverage are measured on AI-consumed interfaces and reported to Layer 4. Direct store access by AI systems is measured and trending toward zero. Exception age is tracked and rising counts trigger review before expiry. Breaking changes carry measured consumer migration, including autonomous consumers that cannot report failure.
5 · OptimizingInterface specifications are generated and validated by the build pipeline, so description and error coverage are properties of release rather than compliance work. Fitness records expire automatically and block AI consumption until renewed. Rationalization is continuous rather than periodic. The exception register trends toward empty because remediation is scheduled rather than deferred.

The distance between levels 2 and 3 is dominated by one thing: producing the catalog of AI-consumed interfaces. Most organizations discover at that point that they cannot enumerate what their models read.

Layer 3

Intelligence

LevelDescriptor
1 · InitialModels are built by whoever needed one. No registry, so no complete list of what runs in production. Validation happens against criteria the builder chose and recorded informally. Data fitness is assumed from the fact that a pipeline exists. Security controls cover infrastructure and treat model endpoints as ordinary services.
2 · ManagedA model registry exists and is mostly current. Model cards are written at launch. Validation is documented and the thresholds are the builder's. Some data domains have stewards. Drift is monitored on the models someone remembered to instrument. Access grants are recorded and rarely reviewed.
3 · DefinedEvery production model is registered, carries a model card bound to its running version and was validated against thresholds recorded before validation ran. Every data domain and interface feeding a production model carries a current fitness declaration. Access is granted with justification and reviewed on cycle. Thresholds for drift escalation are published, and crossing them routes the decision to Layer 4. Model Owners do not authorize their own deployments.
4 · Quantitatively ManagedCard currency, drift coverage, fitness coverage and remediation time are measured and reported to Layer 4 on cycle. Deviation triggers a defined response rather than a discovery. Evidence completeness is measured before audit rather than during it. Feature reuse is tracked to its originating lawfulness basis.
5 · OptimizingModel cards, lineage and validation evidence are generated by the pipeline, so evidence completeness is a property of deployment. Fitness declarations expire automatically and block consumption until renewed. Drift detection triggers retraining within band and escalation beyond it without human initiation. Controls are automated to the point where Layer 4 verifies a stream rather than a submission.

The distance between levels 2 and 3 is dominated by one change that is organizational rather than technical: removing deployment authorization from the Model Owner. Every other level 3 requirement is achievable by a competent team. That one requires someone else to hold the authority.

Layer 4

Control

LevelDescriptor
1 · InitialPolicy exists as a document nobody applies, or does not exist. No register of AI systems, so policy applies to nothing enumerable. Risk is discussed when something goes wrong. Accountability is assumed to sit with whoever built the system. Audit has not looked at AI.
2 · ManagedAI policy is approved and circulated. Some systems are registered. Risk acceptance happens for the systems someone escalated, usually verbally or in a meeting record. Classification exists as a concept without a published mechanism. Cost is visible in aggregate. Vendor contracts are signed without AI-specific provenance terms.
3 · DefinedEvery production AI system is registered, classified against a published mechanism and carries one named Business Accountable Executive. The mandatory control catalog is published per consequence class. All ten delegated band thresholds are published. Risk acceptances are signed, time-bounded and held by domain owners, with a composite record above them. Regulatory applicability is determined per system rather than assumed. An aggregate exposure assessment is produced on the board reporting cycle. Layer 4 does not produce the evidence it verifies.
4 · Quantitatively ManagedRegister completeness, classification currency, acceptance expiry and finding remediation are measured and reported on cycle. Threshold review is scheduled rather than triggered by failure. Control catalog coverage against the source instruments is measured. Cost is attributed per model and governed against envelope rather than reported after the fact. Audit tests the decision records, not only the controls.
5 · OptimizingEvidence arrives as a stream from Layer 3 rather than as a submission, so verification is continuous. Acceptances expire and block continued operation until renewed. Classification re-evaluates automatically when a system's configuration or data scope changes. Aggregate exposure across the AI estate is computed and governed alongside per-system risk. Thresholds are revised from measured outcomes rather than from incident.

The distance between levels 2 and 3 is dominated by two artifacts that do not exist at level 2: the AI system register and the published threshold table. Both are modest documents. Neither is technically difficult. Both require someone to make decisions that have been comfortable to leave open.

Layer 5

Execution

LevelDescriptor
1 · InitialAI initiatives start where someone had budget and enthusiasm. No portfolio view. Gates exist for large projects and not for AI work, which is treated as experimental. Pilots run indefinitely. Nobody can list which AI initiatives are underway. AI literacy is assumed from job title.
2 · ManagedA portfolio register exists and covers funded initiatives. Stage gates are applied to AI work using the general project template. Some pilots have end dates. Capacity is planned for the initiatives that asked. AI literacy training is offered and completion is tracked. Business Accountable Executives are named at go-live.
3 · DefinedEvery initiative delivering an AI system names its Business Accountable Executive before the first gate. The production readiness gate carries the Layer 4 authorization as a mandatory entry condition. Pilots are registered with transition review dates and a named owner for the transition decision. AI literacy requirements are set per role and competence is assessed, not just completion. Benefit realization is assessed at a stated interval. Workforce AI tooling is authorized with a stated data boundary.
4 · Quantitatively ManagedGate compliance, pilot age, literacy competence and benefit realization are measured and reported to Layer 4 and Layer 6. Capacity demand forecasts are compared against Layer 1 actuals and the forecast method is corrected from the variance. Deferred initiatives are tracked, so the cost of sequencing decisions is visible rather than absorbed.
5 · OptimizingGate entry conditions are enforced by the delivery toolchain rather than checked by a person, so an initiative cannot reach a production gate without its authorization attached. Pilot transition triggers automatically on usage thresholds rather than waiting for a review date. Capability planning is driven by portfolio composition rather than by requisition. Benefit outcomes feed Layer 6 as evidence in the next funding cycle rather than as a retrospective.

The distance between levels 2 and 3 is dominated by one change that costs nothing and is resisted anyway: naming the Business Accountable Executive before the build rather than at go-live. Naming at go-live means the entire build ran without an accountable person, and the name selected is whoever was available rather than whoever should hold it.

Layer 6

Strategic

LevelDescriptor
1 · InitialAI activity happens where budget and enthusiasm coincided. No stated AI strategy, or one that names technologies rather than outcomes. Decisions not to use AI are not decisions, because nobody was asked. Governance capability is not considered when commitments are made.
2 · ManagedAn AI strategy document exists and is broadly accurate. It was largely assembled from initiatives already underway. Outcomes are stated in general terms without measures or intervals. Emerging technology evaluations run without kill criteria. Sourcing happens by procurement convenience rather than by posture.
3 · DefinedThe strategy states outcomes with measures and intervals, and funding at Layer 5 requires traceability to one of them. A target maturity level is set per AI9GM layer. Capability decisions are recorded, including decisions not to proceed, documented proportionately to materiality. A model sourcing posture with a concentration limit is published. Evaluations carry kill criteria and decision dates. Where a strategic initiative outruns governance capability, the gap is named in the approval and a remediation plan is bound to it.
4 · Quantitatively ManagedStrategic outcomes are measured at the stated intervals and the results inform the next cycle. Maturity is assessed per layer against target and the trajectory is tracked rather than the position. Aggregate exposure and aggregate reallocation are reviewed as strategic positions rather than as risk reports. Benefit realization from Layer 5 is used as evidence in funding decisions.
5 · OptimizingStrategy is revised from measured outcomes on a defined cycle rather than annually by convention. Maturity targets are adjusted from what the estate actually requires rather than from ambition. Sourcing posture responds to measured concentration before it becomes exposure. The capability decision register is consulted before new proposals, so the organization stops relearning conclusions it already reached.

The distance between levels 2 and 3 is the largest in the framework, and it is almost entirely about sequence. At level 2 the strategy describes what was funded. At level 3 the funding requires a strategy to point at. No new capability is needed. The order of two existing activities has to be reversed, and that is harder than it sounds because it means the Portfolio Board has to decline a good proposal that serves no stated outcome.

Target setting

L6-02 Set the target maturity level per AI9GM layer

Decide what governance capability the organization should hold, per layer. This converts the framework from a description into a roadmap and is what makes the assessment a gap analysis.

Correction

Correct this page

The maintainer answers corrections. There is no service level. Responses are best-effort and opportunistic within a reasonable time: a correction raised on a Monday is answered that week or sooner.

Attribution