AI IaC Utopia
A description of somewhere nobody works
Everything in this piece is buildable with technology that shipped years ago.
None of it requires a breakthrough, a vendor, or a budget line that doesn't already exist somewhere in your organisation under a worse name. It is not a prediction and it is not a roadmap. It is a description of what "good" looks like when nobody has been allowed to negotiate it downwards.
It is written in the present tense on purpose. Target states written in the conditional — we would, the organisation should aim to — are hedged before they are read, and everyone in the room knows it. Written flat, in the indicative, a target state becomes a thing you can be measured against.
That is uncomfortable. That is the point.
The morning
An engineer arrives at nine. She wants to add a queue to a service, expand its network reach slightly, and grant it access to a data store it has not previously used.
She writes a declaration. Not a ticket, not a request, not a diagram that someone else will translate into a ticket — a declaration, composed from curated modules that already encode the organisation's decisions about encryption, logging, network position and lifecycle. The composition takes eleven minutes, because the hard thinking was done once, centrally, and has been reused four hundred times since.
She opens a change. Within ninety seconds, automated evaluation tells her three things: the queue is compliant, the network expansion is compliant, and the data store grant conflicts with a residency rule. Not "flagged for review by the security team." Conflicts. With a rule, by name, with a link to that rule as it is written in code, and a suggested alternative that satisfies both her intent and the rule.
She takes the alternative. A colleague reviews the change for whether it is a good idea — the only question a human review is actually good at — because every question about whether it is permitted has already been answered by a machine, deterministically, in under two minutes.
It merges. It applies through a pipeline. She never touches production. She has never had credentials that could touch production. Neither has anyone else.
At no point in this morning did anyone request evidence, produce a screenshot, complete a control questionnaire, or attend a meeting. The evidence was emitted as exhaust. It is already in the system of record, already queryable, already attributable to her, to the artefact version, to the policy version that evaluated it, and to the reviewer who approved it.
That is the whole thing, really. Everything below is the machinery that makes that morning possible — and makes it identical whether the author was an engineer or an agent.
One control plane
There is exactly one way for the estate to change.
Not one way for humans and a different way for automation. Not a governed path for the platform team and an ungoverned path for people in a hurry. Not a clean process for new build and a shrug for legacy. One path, and technical prevention of every alternative — not policy language forbidding alternatives, but the absence of any credential capable of taking one.
This matters more than it used to. When the only authors were people, a second path was a governance weakness. When the authors include agents operating at machine cadence, a second path is an unmonitored production surface changing faster than anyone can read it.
The organising principle is authorship-indifference:
Every change to the estate is a declared artefact, evaluated by policy, evidenced automatically, and reversible — and none of those four properties depend on who or what authored it.
Provenance is a field on the artefact. It is not a different process.
Everything is an artefact
Five classes of thing are governed. Most organisations govern the first two competently, the third inconsistently, and the last two not at all.
Infrastructure. Declarative configuration, state, modules, provider versions. The well-trodden case.
Application. Source, dependencies, images, tests. Also well-trodden.
Policy. Guardrails, detections, entitlements, classification rules. Policy is code, versioned and reviewed exactly like the things it governs. A control that exists only as a paragraph in a standard is not a control; it is an aspiration with a document number.
Cognition. Models, system prompts, tool definitions, retrieval corpora, agent scopes. This is the class that gets missed, and the omission is not small. A system prompt is a control surface. A tool definition is an entitlement grant. A retrieval corpus is a body of text the model will treat as authoritative. All three change production behaviour, and in most organisations all three can be edited by one person, in a web console, with no review, no version history, and no way to answer the question what was it yesterday?
Data. Schemas, lineage, retention, residency. Governed as declaration rather than discovered by archaeology.
Five classes, one lifecycle, one set of gates. The uniformity is the feature. A model that governs infrastructure rigorously and prompts not at all has simply moved the risk to where nobody is looking.
Four gates
Every artefact, in every class, from every author. No exemptions by seniority, urgency, or authorship.
Provenance. What produced this, from what inputs, at what version, under what authority. For agent-authored change this extends to the model version, the prompt version, and the context set that informed it. An action that cannot be traced to those three is not auditable, and is therefore not permitted.
Policy. Automated evaluation before merge and again before apply. Fast enough that the engineer waits for it rather than context-switching away — because a gate that takes an hour is a gate people learn to route around.
Evidence. Assurance output is a by-product of the pipeline, not a task performed by humans afterwards. The test is blunt: if anyone in an assurance function ever has to ask an engineer for proof, the model has failed. The auditor does not send an email. The auditor writes a query.
Reversibility. A declared path back, proportionate to blast radius, tested rather than assumed. This gate quietly does more work than the other three combined once autonomous change is real, because approval does not scale and containment does.
Machines that build, machines that are built, machines that reason
Keep these three separate. Collapsing them is why most conversations about AI in operations produce nothing.
Machines that build
Agents author configuration, remediate findings, triage alerts, write tests, open changes. Through the same control plane as everyone else.
Each agent holds its own identity — scoped, time-bound, attributable. Never a shared service principal. Never a borrowed human credential. Never, under any circumstances, the credentials of the engineer who invoked it, because the moment that happens the audit trail stops meaning anything and every action from then on is deniable.
Authority is graded by consequence rather than by task. The same agent may act autonomously in a low-consequence blast radius and be restricted to proposing in a high one. The grading is a property of the environment, not of the agent's cleverness.
Autonomy is earned per capability, against a measured reversal rate, and revoked the same way. Nothing is granted autonomy because a vendor demonstration was impressive. Things are granted autonomy because they have demonstrated, over a meaningful sample, in the actual estate, that their proposals are accepted and their actions are not reversed.
The non-human identity population outnumbers the human one and grows faster. It inherits every privileged access problem the industry spent two decades learning, plus one that is genuinely new: the credential holder can be persuaded by its inputs. No human privileged account has ever had that property. It changes the threat model.
Machines that are built
Models, agents and retrieval systems are production systems with an unusual supply chain. Weights, prompts, tool definitions and corpora each carry versioning, provenance, integrity verification and a rollback path. Evaluation suites gate promotion the way tests gate application code.
The exposure with no conventional equivalent is the context supply chain. Retrieval corpora and tool outputs are untrusted input that the model treats as authoritative. The nearest analogue is injection, but injection assumes a parser you control rather than a reasoner you persuade.
Machines that reason
Deterministic systems assemble the context. Dependency models, telemetry, configuration state, relationship traversal — resolved by systems that are correct rather than plausible. Inference runs as the final step, over structured material that has already been reasoned about.
The graph does the thinking. Inference is the last step, and it is last precisely because it is the least trustworthy component in the chain.
The inverse arrangement — model first, evidence retrieved afterwards to support whatever it said — produces confident narrative rather than analysis. It demonstrates beautifully and fails in exactly the situations you bought it for.
Drift, properly understood
Conventional drift is deployed state diverging from declared state. The same taxonomy runs across all five artefact classes:
Configuration — declared vs actual
Dependency — pinned vs current vs vulnerable
Policy — intent vs enforced
Model — evaluation performance vs baseline
Context — corpus vs source of truth
Entitlement — granted vs exercised vs required
One taxonomy, one detection loop, one remediation queue, one set of severities.
This matters most for the cognition class, because AI systems mostly do not break. They degrade. Nothing goes red. Nothing pages. Quality drifts downward across weeks, and without continuous evaluation against a baseline, the first reliable signal that something is wrong arrives from a customer, a regulator, or a journalist.
Identity as the spine
Human, workload and agent identities on one model with one lifecycle. Credentials short-lived and federated; static secrets an exception with an owner and an expiry rather than a fact of life.
Entitlements are declared artefacts, subject to the same four gates as everything else. Granted access is reconciled continuously against exercised access, and the difference between the two is understood for what it is: the standing attack surface, quantified, trending, owned.
Emergency access is designed against the circular dependency — the path to restore identity does not itself require identity. And it has been rehearsed recently enough that the rehearsal date is a number on a dashboard rather than a matter of recollection.
Proof by reconstruction
The strongest claim here is not that everything is under code. It is that this has been demonstrated.
A foundational component is complete when it has four properties: a code artefact, a position in a reconstruction sequence, a proven import path, and a dated reconstruction test. Three out of four is not complete. It is a component that will be discovered to be incomplete at the worst available moment.
This is the honest test of the entire model. Everything else can be asserted. Reconstruction cannot. Either the estate rebuilds from declaration or the declarations were decorative, and the only way to know which is to have done it.
Anything that can be reconstructed from its declaration is genuinely under code. Anything that cannot, isn't — regardless of how many repositories it appears in.
What it costs
Utopias are cheap to describe and expensive to inhabit, and a piece that omitted the price would be advertising.
Curated modules require a funded team. Not a rota, not twenty percent of somebody's time. A paved road that is not maintained becomes a paved road that is out of date, and an out-of-date paved road is worse than none, because people fork it and the forks become permanent.
Policy-as-code requires people who can write both. The population that understands regulatory intent and the population that can express it as an evaluable rule overlap less than anyone plans for.
The import campaign is the real programme. The proportion of foundational infrastructure that exists only as running infrastructure determines whether adoption is a governance exercise or a multi-quarter excavation. Measure it before committing to a date. It will be worse than the estimate.
Someone must be able to say no. Every mechanism here can be dissolved by a sufficiently senior person in a sufficiently urgent moment. The real precondition is not technical.
Two honest gaps
Autonomy thresholds have no industry baseline. Nobody can currently tell you what reversal rate justifies allowing an agent to apply rather than propose within a given blast radius. The mechanism is sound; the numbers have to be earned per environment, from measurement, over time. Anyone quoting you a figure is quoting their marketing department.
Evaluation of agentic systems is immature. Evaluation is reasonable for single-turn model behaviour and weak for multi-step agents holding tools — which is precisely the configuration everyone is deploying. The gate exists. It is not as strong as the diagram implies.
Both gaps close with time and measurement. Neither closes by being ignored, and neither justifies waiting.
The distance
Nothing above is speculative. Every mechanism described is running somewhere today, in some organisation, at some scale. What does not exist anywhere is all of it, together, uniformly, without exception.
That is what makes it a utopia rather than a specification. The individual parts are ordinary. The absence of exceptions is the impossible bit.
Which gives you a diagnostic that takes about a minute to run. Ask for the exception register. For each entry, ask two questions: who owns this, and when does it expire.
Most organisations cannot answer either.
That is the distance. It has never once been measured in technology.
Threat intelligence every morning — new victims, new groups, what matters, in plain English. Free, with receipts.
Subscribe to the Daily →
Scott Gardner ·