C4AIL AI Maturity Diagnostic

The C4AIL AI governance canon

56 defined terms describing how organisations govern AI, and how that governance fails.

These are not general-purpose definitions. They are the vocabulary the AI Maturity Diagnostic scores against: each one names a specific, observable failure or capability that the instrument can detect and measure. They are published here so that a result can be read without a translator, and so that anyone quoting them is quoting the same version the engine uses.

Free to cite with attribution to the Centre for AI Leadership.

agent-as-principal
each agent a distinct, revocable identity, never a human's full rights
An agent is a principal: a distinct, revocable non-human identity with scoped, least-privilege credentials - never sharing a human's full access.
#agent-as-principal
AI Theatre
a full policy binder with no capability behind it
Governance that exists on paper but does not bite - accountability closes with no human able to catch the error. The failure the whole instrument is built to detect: high Apparatus, low Capability.
#ai-theatre
AI-exposed vs AI-enabled
high adoption at L1-2 is exposure, not capability
An organisation where most users sit at L1-L2 is AI-exposed, not AI-enabled; adoption metrics disguise the gap.
#ai-exposed-vs-enabled
ARCH
the agent run-time control loop; the gate lives in the harness, not the model
Action (tiered) - Reasoning (captured) - Contextual Check (pre-commit gate: authority/compliance/presence) - Horizon (mandatory human hand-off for irreversible or out-of-scope actions).
#arch
ARGS
the four adoption pillars: Agency, Architecture, Governance, Scaling
Agency (the decision to interrogate rather than accept), Architecture (Logic Pipes + clean data), Governance (the accelerator, not the brake), Scaling (decoupling output from headcount). The path from AI Theatre to Sovereign Command.
#args
autonomy tiers
actions classified by autonomy x reversibility; irreversible = always-ask
Every agent action pre-classified; irreversible / high-blast-radius actions require mandatory human approval. Raising an agent's autonomy is an explicit governance decision, not a default.
#autonomy-tiers
Bright Lines
which actions AI may take alone vs need human sign-off
Clear boundaries: the line sits where accountability begins, not where AI capability ends. Irreversible / high-blast-radius actions are always-ask.
#bright-lines
CAGE
the front half of directing AI: Context, Align, Goals, Examples
Context (all situational/domain context), Align (the AI plays back its plan before running), Goals (granular, incl. Must NOT / Must FLAG), Examples (a growing data layer). Constrains input to domain-valid ranges. Pairs with ARCH.
#cage
CAIO
the Translator made institutional, measured by dispensability
A Direct-seat Orchestrator built on Build substrate who owns the AI storyline and succeeds by manufacturing Translators org-wide, not by being the bottleneck (dispensability = the metric).
#caio
compliance theatre
policies filed that satisfy the auditor without changing the outcome
The governance equivalent of the Eloquence Trap: it looks like governance without being governance (e.g. 75% have policies, 7% embedded).
#compliance-theatre
Comprehension Debt
running on AI systems no one fully understands or has verified
The invisibly accumulating gap where decisions are made by people who no longer understand the logic behind their own work.
#comprehension-debt
concentration / monoculture
dependence on a few providers; shared models cause correlated failure
Reliance on a small set of chip/cloud/foundation-model providers, and model/data monoculture, drive correlated 'herding' failure across firms (FSB/BoE/IMF).
#concentration
Concretisation
making tacit substrate explicit and scalable
AI's core economic act. Concretise your OWN substrate = a moat; depend on generic concretised substrate = erosion. Without the Forge, concretisation becomes corporate amnesia.
#concretisation
Decision Survivability
can you defend the process after it goes wrong, from the trail that exists
The test of whether an AI-assisted decision can be reconstructed and defended after it fails - about the trail and the rational documented process, not whether the output was right.
#decision-survivability
Epistemic Credit
unearned trust granted to fluent AI output you cannot verify
Trust granted to fluent AI output by someone who lacks the substrate to verify it - reaching accountability while skipping the knowing. The opposite of Epistemic Labour (the verification work of knowing).
#epistemic-credit
Floor / Ceiling
the 90-95% who work through AI vs the 5-10% who design and govern it
Floor = the majority who work through AI-structured interfaces (mass literacy, months). Ceiling = the few who become Architects, Orchestrators, Trainers (years). The Diagnostic Map's structural echo.
#floor-ceiling
Intellectual vs Accountability Labour
AI does the intellectual labour; the human keeps the accountability labour
A boundary, not a spectrum: give the intellectual labour to AI, keep the accountability labour (sign-off, ownership under uncertainty) human. Accountability never transfers.
#intellectual-vs-accountability-labour
Legibility Debt
the gap between what you know and what you've made machine-actionable
The structural distance between an organisation's knowledge and what it has made legible in a form AI can act on; the agent can only work with what has been made legible.
#legibility-debt
Leverage Leaks
Architecture, Infrastructure, Talent - three ways value is lost
Where organisations lose even the value their best people create: Architecture Leak (judgment work without verification), Infrastructure Leak (AI on dirty/illegible data), Talent Leak (tools with no capability pipeline).
#leverage-leaks
Living Material
AI output is provisional and version-controlled, not final
AI-assisted decisions get review dates, not just approval dates; outputs are treated as version-controlled material that can expire, not a finished product.
#living-material
Logic Pipes
engineered deterministic workflows that constrain and route AI
End-to-end structured workflows (generate -> verify -> triage -> review -> approve -> audit) that replace narrative chatting; each step is traceable.
#logic-pipes
shadow AI
AI used without approval, bypassing governance
Staff using unapproved AI tools, creating data-privacy and compliance blind spots. A supply problem, not a demand problem: provide, do not just punish.
#shadow-ai
Sovereign Command
the org owns, can defend, and can scale its AI decisions
The state where an organisation owns its AI-informed decisions, can defend them, and scales them without losing control - human judgment kept above machine fluency. The goal state ARGS produces. Not a destination, a discipline.
#sovereign-command
Spec Loop vs Chat Loop
fix the spec so corrections compound, vs fix the conversation (ephemeral)
Spec Loop = encode the correction into the template/substrate so it persists and compounds. Chat Loop = fix only the conversation, so it repeats next session. Stop editing outputs, start editing specifications.
#spec-loop
Substrate vs Surface
earned tacit judgment vs tools/policies/vocabulary
Substrate = slow-built tacit domain judgment and accountability experience (defensible). Surface = tools, interfaces, policies, vocabulary (fast, hollow alone). The Map is Apparatus (Surface) vs Capability (Substrate).
#substrate-surface
taste / phronesis
the felt sense of quality that separates competent from accountable
Practical wisdom (phronesis) - the internal standard of quality that precedes articulation. AI has episteme (knowing) and techne (doing) but no phronesis.
#taste-phronesis
the 98/2 Principle
trust lives in the deterministic harness you own, not the model
98% deterministic software you own and engineer; 2% swappable probabilistic model at the edge. Upgrading the model changes neither who is accountable nor where the gate sits.
#ninety-eight-two
the accountability sink
a consequential action taken under delegated authority with no human in the moment
Accountability quietly closing with no one inside it. Accountability never transfers to the agent - a named human always owns the outcome.
#accountability-sink
the Adjacency Pathway
mid-career experts porting existing substrate in 9-18 months
Porting an existing expert's substrate into the AI-augmented work (Architect in 3-6 months, Orchestrator in 9-18) - the primary near-term source of Ceiling capacity, vs the 3-5yr Novice Pathway.
#adjacency-pathway
the Atrophy Trap
offloading -> loss of evaluative power -> structural sclerosis
Delegating judgment to AI makes the organisation progressively unable to do or judge the work it depends on: cognitive offloading, then loss of evaluative power, then structural sclerosis.
#atrophy-trap
the Building Code (NIST/ISO vs ARGS)
frameworks say what must be true; ARGS builds it
NIST AI RMF / ISO 42001 are the building code (what must be true of anything you build); ARGS is the builder - the teaching and implementation framework.
#building-code
the C4AIL Awareness Model
the 0-6 maturity spine: Unaware -> User -> Amplifier -> Orchestrator
AI Unaware (L0) -> AI User (L1-2) -> AI Amplifier (L3-4) -> AI Orchestrator (L5-6). The Knee at L3-4 is where the Eloquence Trap breaks and returns start to compound.
#awareness-model
the capability spend ratio
human-capability spend vs infrastructure spend (aim >= 1:3-4)
The Infrastructure-Capability Imbalance: the great majority of AI spend goes to infrastructure and a fraction to the humans who must make it work - the pattern most consistently found in pilots that never reach measurable impact.
#capability-spend-ratio
the co-creation model
AI does the volume, the junior does judgment under a senior
Redesigns the junior role so AI does the volume and the junior co-creates under a senior who develops their taste; 'would you sign this?' replaces 'did you catch the errors?'
#co-creation-model
the Confidence Plateau / Dunning-Kruger Peak
more confident while less capable
AI removes the visible failure that normally corrects overconfidence, so users grow more confident while growing less capable, mistaking the tool's competence for their own.
#confidence-plateau
the Cost of Not Acting
the quantifiable price of inaction / falling behind
The inaction side of two-sided risk: lost margin, talent flight, strategic obsolescence from moving too slowly while competitors move.
#cost-of-not-acting
the Eloquence Trap
trusting fluent-but-wrong AI output without checking it
Fluent output that looks right (right terminology, internal coherence) masking that it probably is not. Experts are most at risk.
#eloquence-trap
the Five Roles
Floor User, Translator, Architect, Orchestrator, Trainer
Defined by labour function, not job title: Floor User (works through AI), Translator (bridges domain + AI), Architect/Amplifier (builds Logic Pipes/verification engines), Orchestrator (designs and governs the system), Trainer (maintains the human pipeline).
#five-roles
the Four Labours
Intellectual, Physical, Accountability, Architectural
Work decomposed by the human it requires: Intellectual (commoditised by AI), Physical (lags), Accountability (the durable monopoly), Architectural (the growth category where new jobs live).
#four-labours
the Institutional Vault
the org's codified, executable knowledge store; it executes but never judges
The codified knowledge store (processes, prompts, pipelines) - surface made durable. It may execute codified rules but must never make the un-codified call. The Brain (human judgment/accountability) stays human.
#institutional-vault
the Intern Test
would you give an intern this access and let them act for you?
If you would not hand an intern this access and let them act on your behalf, apply the same limits to the AI agent (least privilege).
#intern-test
the Inverted Stack
probabilistic logic where deterministic rules belong
The failure of letting the model make the un-codified call (or compute a score) - probabilistic logic in a place that needs deterministic rules. The 98/2 failure.
#inverted-stack
the Knowledge Paradox
those who most need Sovereign Command can least initialise CAGE
The organisations that most need structured AI use are the ones whose knowledge was never made legible, so they cannot fill CAGE's Context - resolved by the Spec Loop.
#knowledge-paradox
the Missing Middle
cutting juniors faster than building verification capacity
Automating junior work severs the pipeline that produces tomorrow's senior judgment: no juniors today, no seniors in five years.
#missing-middle
the Moat Inversion
value moves from what you can codify to what you cannot
AI commoditises the explicit, so the defensible moat shifts to the tacit - judgment, taste, the ability to feel a fluent answer is wrong.
#moat-inversion
the Pre-AI Permission Audit
tighten file permissions before deploying any AI search tool
AI search inherits existing permissions and instantly surfaces mis-permissioned data; auditing permissions first is the highest-return security action.
#pre-ai-permission-audit
the Reliability Trap
errors compound across multi-step / agentic chains
Chained probabilistic steps compound error (0.95^5 = 77%); each step looks sound in isolation, so the failure is invisible until the end.
#reliability-trap
the Six Numbers
governance measured, not attested
First-time-right rate, acceptance-without-verification rate, rework rate, error-classification, correction-encoding (Spec Loop vs Chat Loop), and Decision-Survivability score. Governance you measure rather than attest to.
#six-numbers
the Trail
the record that makes invisible AI degradation diagnosable
What was generated, verified, changed, and approved - the record that lets you catch AI degradation that otherwise looks stable, and defend a decision later.
#the-trail
the Trainer Paradox
you need L4+ practitioners to develop L4+ practitioners
The bootstrap problem: the system that needs Trainers to produce accountable practitioners cannot produce Trainers without already having accountable practitioners.
#trainer-paradox
the Translator
bridges domain expertise and AI capability; a trait, not a rung
Domain-fluency + AI-fluency + judgment + legibility, meant to distribute across every seat rather than be a maturity level. Failure modes: fluent mistranslator, credulous bridge, silent expert.
#translator
the Verification Bottleneck
an expert drowning in manual review of AI output
What domain expertise without AI architecture produces: a brilliant expert reviewing AI output line by line, which does not scale. Fixed by architecture, not more reviewers.
#verification-bottleneck
the Wrong Scoreboard
measuring adoption instead of capability
When you measure adoption, you optimise for adoption; capability, verification quality, and pipeline health are what actually matter.
#wrong-scoreboard
two-sided risk
risk of adopting badly AND risk of falling behind
AI risk is two-sided: the risk of adopting AI badly and the competitive risk of not adopting (or adopting too slowly). Governance must weigh both, not only say no.
#two-sided-risk
Verification Capacity
the human ability to catch fluent-but-wrong output and own the call
The competency AI adoption needs more of, not less. The leading indicator of whether governance will keep biting.
#verification-capacity
workslop
AI content that looks professional but lacks substance
Plausible, well-structured AI output missing the actual business logic; the cost lands on the recipient, who must supply the verification the sender skipped.
#workslop

Using these

Every term has a stable anchor, so a specific definition can be linked directly - /glossary#ai-theatre, for instance. The same definitions are served to any AI connected to the assessment, so an assistant explaining your result is explaining it in these terms and not its own.

Take the governance screen ยท Run the assessment with your own AI