Part II — The Universal Operating Framework

Chapter 5

The AI Tool Register

Governs how AI enters the organization. Establishes the difference between sanctioned and shadow AI, the record the organization keeps of what it has sanctioned, the two gates every AI system passes, the evaluation criteria and contractual requirements applied before approval, the response when unsanctioned AI is found, and the retirement framework that closes the lifecycle. The chapter's trace to accountability runs through the second corollary: an organization that cannot say what it has deployed must still answer for what is deployed.

Published September 24, 2026 24 min read

Chapter 4 required that every AI deployment be classified into a Trust and Harness tier before it is approved for use, and insists that the tier determines where the responsibility sits for verifying output. This chapter answers the three questions Chapter 4 left unaddressed: what approval consists of, who performs it, and where the classification is recorded once someone has made it.

The second corollary of the Accountability Principle is why this work is necessary. The organization cannot disclaim responsibility for the behavior of AI tools it has deployed, regardless of how those tools were marketed, acquired, or configured. The organization has to be able to say what AI it runs, where it runs, and who answers for its output.

Sanctioned and Shadow AI Systems

An AI system is sanctioned when the organization has evaluated it, classified it into a Trust and Harness tier, attached it to a named verification owner, and documented all of it. An AI system is shadow when it is in active use without those things.

Shadow AI is usually discussed as a security exposure, and the exposure is real. It is also the smaller part of the problem. Accountability does not wait for approval: when an unregistered deployment produces output that reaches a customer or a system of record, the Accountability Principle attaches that output to the operator who initiated it, the custodian of the domain it crossed, and the organization who owns the work. The absence of an approval changes nothing about who’s responsible.

What the absence changes is every process the organization would have used to catch the problem. Without a tier assignment, no one has assessed how much of the checking the operator was performing. Without a named verification owner, the checking itself was never assigned to anyone. The audit trail that Chapter 9 requires does not exist either, so the work cannot be reconstructed after the fact.

We hold that no AI system may be in use in the organization that the organization cannot describe, and this position does not soften for tools that cost nothing.

Cost is the wrong boundary for a sanctioned ruling. Several legitimate deployments carry no incremental price: an open-weight model running on corporate infrastructure, a new capability released into software the organization already licenses, or a free tier used under an enterprise agreement. A rule that prohibits what is free prohibits those consequently. A rule that governs only what appears on an invoice leaves an opening for anything an employee acquires through personal means. The boundary that works is account and control. Account and control is any AI system the organization uses through its own identity, under its own contractual control, and may be sanctioned regardless of cost. AI acquired through a personal account, or otherwise outside organizational control, is prohibited.

The prohibition creates an obligation for the organization. An employee reaching for a commercial or free AI tool is trying to complete work with no sanctioned way to do it, and prohibiting the tool without supplying an alternative does not change that. It only moves the work into the shadows the organization cannot see.

Everything in this chapter, and an organization’s adherence to the Accountability Principle depends on the organization keeping record of what it has sanctioned. That record is called the AI Tool Register.

The AI Tool Register

The AI Tool Register (Register) is the organization’s account of every governed AI system it operates: what the system is, where it is deployed, which Trust and Harness tier each deployment occupies, and who owns verifying its output.

The Register holds one entry per product. The AI features and workflows that product contains or have been created from, are enumerated beneath it as deployments. A deployment is one feature, used by one business unit, for one business purpose.

Two levels are created by these entries and answer to different audiences. Product-level identification lets an auditor or an executive see how much AI is in the business and where it exists, which is the question those readers will ask. Deployment-level detail is where the tier and the verification owner live, which informs change management, SRE’s, application support, NOC staff, the Policy Owner, and the AI Governance Council. These two levels are required because Chapter 4 established the tier as a property of the deployment rather than of the model in use. This reasoning can be illustrated by an example. A summarization feature used by a marketing team to identify impactful metrics from internal research, and also used by a claims team to compress case files before a decision, are two deployments at two different tiers; a single tier recorded against the product would mislabel both.

Additionally, a harness the organization built for itself exists as a deployment beneath the entry for the model it runs on.

Each entry has a named owner inside the business unit that owns the AI tool (product or deployment-level), accountable for keeping that entry accurate, while the Policy Owner is accountable for the Register as a whole. The split is deliberate. A central management function cannot track feature-level change across an entire software estate, and organizations that have tried to maintain distributed inventories from the center have produced records that were accurate only on the day they were built, and decayed from that day on. The business unit running the deployment is the party positioned to notice when the AI tool changes in a way that warrants an entry’s update.

Appendix A contains the AI Tool Register Template which includes the full set of fields. Four fields need comment here, because each of them holds a requirement this chapter places elsewhere. The tier and the verification owner are what Chapter 4’s control resolves to. The tier records how much checking the deployment’s output needs before the organization can rely on it. The verification owner records who is responsible for doing that checking. An entry missing either describes a deployment with unfinished governance in place.

The halt usage method records how a particular deployment is stopped, which differs by tier and is specified later in this chapter. Having the halt usage method on file at registration time is what separates a response measured in minutes from one that begins with an hour of research, text messaging, and confusion.

Vendor contract term (as in policy, not time-bound) status sits in the Register rather than in the intial vendor evaluation files, often part of a bake-off. The reason is practical. Evaluation artifacts are produced once, filed under whatever discipline was applied by the owner or team at the time, and are frequently unlocatable years later when questions arise. The Register survives this, and it carries the status of each term, a pointer to the evidence, and a reference to any exception the organization approved where a term was not obtained.

The Register carries no review cadence of its own. Currency is maintained through the review and discovery processes the organization already runs: internal technology audit, vendor and contract renewal timings, access recertification, the change advisory process, and annual policy reviews. Access recertification is the most useful of these for AI discovery, since it already enumerates which people hold access to which applications. A separate AI review calendar would create a recurring obligation competing with obligations that already have owners, budgets, and audit evidence. A new review process would lose that competition. Each Register entry records when it was last reviewed and which process performed the review, so the review leaves evidence an auditor can follow.

Reconciliation is what makes the Register a control. Inside those existing processes, the Register is compared against procurement and accounts payable records, the identity provider’s application inventory, the software renewal calendar, and the business unit’s own declaration of what it is running. Each source catches a different control failure. Payment records expose what was bought and never entered into the Register, while the identity provider exposes what employees sign into whether or not anyone paid for it. Renewal review is where a vendor’s shipped capability clearly surfaces. Self-declaration is the only one of the four that finds an internally built AI system that reached production without a prior entry.

This is also where Chapter 4’s classification requirement becomes provable. That chapter required every deployment to be classified before approval. The classification becomes the control point, without specifying how the organization demonstrates that classification happened or that it remains accurate. Reconciliation is that demonstration and proof. A Register with entries never reconciled, records what someone believed only at the time of its recording.

The Approval Pipeline

Every AI system must pass through policy registration, and where money is exchanged, it passes commercial approval first. Registration into the AI Tool Register determines what tier the deployment occupies and who answers for its output. Buying an AI system decides nothing about either.

Commercial approval is the organization’s existing process for approving a technology purchase, and we do not alter it. Technology purchases are made to support the execution of a strategy or to create a specific capability, and both exist because of a business justification and an approver holding the budget. Mature organizations run this process well. Building a second approval alongside it for AI purchases only would put two authorities over one decision and the new process would almost certainly fail to survive.

AI policy registration is the creation of the Register entry: the product, the AI features it contains, the tier each deployment occupies, its intended use, the data classifications it is approved for, and the role-holder who owns verifying its output. No AI system may be used in the organization without that entry. Registration activity attaches to the purchase process where a purchase exists and stands on its own where none does, and completing one gate does not complete the other. A manager with discretionary budget who signs for an AI product has finished with procurement and has registered nothing.

There are four paths that bring AI into the organization we should be aware of. A purchased AI product travels the familiar, first path through procurement and carries the full evaluation. An open-weight or self-hosted model usually involves no purchase path, and its evaluation rests on provenance, hosting choices, and the harness assembled around it. A deployment built internally on a model already approved has no procurement event of its own and is examined on its permissions, its validation logic, and its named owner. The fourth path is AI capability that a vendor adds to software the organization already runs.

That fourth path delivers AI without a procurement event of any kind. No purchase order exists to find, no contract amendment prompts a review, and no request for approval reaches anyone, because the organization made no decision. The vendor made it. New capabilities are frequently enabled by default, inside a product that was correctly absent from the Register on grounds that it previously contained no AI. Enterprise systems are constantly being enhanced with AI-based generative content assistance, task automation, insights, and customer support. Under the scope Chapter 2 fixed, each of these becomes a governed AI system on the day the vendor ships it, and nothing in the organization’s ordinary operations prompts anyone to record it.

How much of an organization’s AI exposure arrives this way varies with the size and composition of its software estate. What can be said without qualification is that a Register built on AI purchases cannot see any of it, and that the gap is invisible from within the Register. An organization that has not gone looking for this pathway of AI introduction, should not assume the amount is small.

Closing that gap is the Policy Owner’s first task in establishing the Register. A baseline sweep of the existing software estate for AI capability the organization already owns is the largest single body of work in establishing the Register, and it is also finite. Three inventories the organization already maintains or has access to produce the candidate list: the identity provider’s record of every application employees sign into, accounts payable’s software renewal calendar, and the published release notes for the products those two sources name. The time goes into the determination, product by product, of whether a capability exists, whether it is enabled, what it does, and what tier it occupies.

Internally built harnesses need a line drawn for a registration event, and production reach is that line. Sandbox use of a procured model by an engineering team sits under the approval that brought the model into the organization and needs no entry of its own. Once a harness built on that model can act on production systems, it is registered, with its permissions, its validation logic, its named owner, and its tier recorded as any other deployment would be.

The technical approval of that harness stays in the owning department’s internal documentation and change management processes. The approvers would be the same people either way, and this policy does not improve an engineering organization’s release governance by adding a signature to it. Registration is a separate act. A change record establishes that a change was safe to make; a Register entry establishes what the organization runs, at what tier, and who owns verifying it. A department that documents its harness thoroughly under change management has satisfied its own governance and has still left that harness invisible to the Policy Owner and the AI Governance Council.

The depth of review scales with the tier the deployment will occupy. A Tier 1 deployment carries a light evaluation and a complete Register entry. A Tier 4 deployment carries the full criteria, the security review, the contractual conditions, and a named role-holder who has accepted its outcomes before it runs.

Where registration is refused, the requester follows the operator-initiated escalation pathway Chapter 3 established, terminating at the AI Governance Council. This policy creates no separate appeal for tool requests, because a contested decision under it is already provided for.

Arrival path Commercial gate Registration What the evaluation asks
AI product purchased Existing purchase process Required The full vendor criteria
AI capability added by a vendor to a product already in use None; no purchase occurs Required The vendor criteria, applied to the new capability
Open-weight or self-hosted model Usually none Required Provenance, hosting, and the harness around it
Deployment built internally on a procured model None; the model was approved earlier Required at production reach Permissions, validation logic, named owner

Evaluating What the Organization Is About to Trust

An evaluation gathers the answers the organization will need later, when a customer asks what happened to its data, a regulator asks what the tool was permitted to do, or an executive asks why nobody knew this AI existed. Every criterion below exists because one of those questions is coming.

The first question concerns what’s available in the product. Which model or models does it run on, whose are they, and does the vendor change them without notice? AI output behavior moves when the model moves. A deployment whose underlying model is swapped quietly has had its prior testing invalidated without anyone being informed, and the organization usually discovers this after the fact by noticing that the output has changed.

The second group of criteria concerns what happens to the organization’s content once it enters the AI tool. Whether the vendor trains on the organization’s inputs, what it retains and for how long, and whether the behavior can be disabled, together determine which data classifications the deployment may hold (Chapter 6: Data Classification and AI Use). The same answers govern what the organization can recover when the deployment is retired, since content a vendor has absorbed into a training corpus does not come back when an account is deleted. Data residency belongs to this group as well, covering where inference happens and where content rests. So does sub-processor disclosure: the organization’s commitments to its own customers pass through whichever other parties touch the content, and a vendor unwilling to enumerate them is asking the organization to create policy based on something it cannot substantiate.

The third group of criteria concerns what the organization can check rather than accept. The security review of the tool belongs to the security function under existing security policy, and this chapter neither specifies it nor duplicates it; it requires that the review happened before approval and that the Register records where the result lives. Audit rights or vendor-provided audit results determine whether the organization is able to verify the vendor’s claims at all, which matters most for the claims the organization will repeat to its own regulators, business partners, or potential customers.

The evaluation ends by assigning the Trust and Harness tier the deployment will occupy, and that assignment governs what work the deployment may perform, which data it may access or hold, and where its verification responsibility rests. Assigning the tier during evaluation puts it in place before anyone has relied on the output, which is what makes it a control.

Evaluation runs against the deployment the organization intends to operate rather than against the vendor’s product literature. As we’ve seen, the same product, configured two ways for two purposes, can occupy two tiers, and the tier follows its configuration.

A harness the organization built for itself on a procured model answers a different set of questions, because most of the criteria above address a supplier’s AI capability. What can the harness touch, what validates its actions before those actions execute, and which named role-holder owns its outcomes? Those are the Tier 4 qualifying conditions from Chapter 4, levied against the deployment here at the time of approval.

This Policy states that no AI system is approved until it has been evaluated against the criteria that apply to it and the result has been written down. An evaluation whose conclusions were never documented leaves no trail to show what was examined.

The Contract Is the Backstop

The organization must begin asking its AI vendors for six things: a data processing addendum, indemnification, audit rights, clauses governing the use and retention of the organization’s content for training, the right to disable specific product behaviors, and termination assistance.

Most organizations will not obtain all six from an AI vendor today. The list stays at full length regardless. A large organization runs a standing set of terms through procurement and legal as ordinary practice, and what survives into the vendor agreement reflects the relationship the purchase creates. A smaller organization rarely has such a list at all, and benefits from seeing what a complete one looks like before signing. Where enough buyers put the same term in front of the same vendors, that term tends to find its way into the standard agreement given time. As with each major shift in technological capability, vendors first determine terms of usage that aligns with their technological oligarchy. Over time however, customer segments with significant holdings over vendor contract revenue, can push vendors to adjust terms that better meet their customers’ acceptable risk exposure.

An aspirational list still produces a prescriptive requirement, and we separate the six accordingly. Terms that govern what happens to the organization’s owned and generated content are conditions of approval wherever the deployment will handle restricted or regulated data above Tier 1: the data processing addendum, the training-data clause, and sub-processor disclosure. A deployment that will hold regulated content under a vendor unwilling to commit to how that content is handled is not approved for that content. Terms that govern the organization’s remedies rather than its content, such as indemnification, audit rights, and termination assistance, are negotiated as far as the relationship allows and recorded at the level reached.

The right to disable specific behaviors deserves separate attention because it governs capability the organization never asked for. Vendors add features to products already in service, and this term is what lets the organization decide whether a newly delivered AI capability enters its estate at all. Where a product cannot be configured to disable a capability, the vendor has made that deployment decision on the organization’s behalf.

A term the organization does not obtain becomes an exception under the mechanism Chapter 3 already provides, approved by the Policy Owner and referenced in the Register entry rather than left to the recollection of whoever ran the negotiation. Recording it there produces something the organization should use deliberately: because the AI Governance Council sees the Register, it can also see which terms the organization consistently fails to obtain and from which vendors. That view supports a procurement position the organization takes against each vendor, at the enterprise level.

When Discovering Shadow AI in Use

The pipeline describes how a deployment is supposed to arrive. Some arrive anyway, and the organization finds them through reconciliation, through an incident, or because someone mentions in a meeting that their team found an online tool that helps them complete tasks in less time, with fewer resources, or with less cost.

The correct action is to stop the AI first, and what stopping means depends on what the deployment is. Access to a Tier 1 open-ended chat application gets revoked from the operator. A Tier 2 retrieval-augmented system is stopped by disconnecting its corpus or disabling the endpoint that serves it. For a Tier 3 constrained operational agent, the credentials are removed and permissions denied. A Tier 4 purpose-built skill runs on an event trigger rather than on a person’s prompt, so the trigger pipeline is what must be disabled; instructing someone to stop using it accomplishes nothing, because no one is using it in the sense the instruction assumes.

Next, a conversation with the operator and the operator’s management follows. This conversation covers the business context, the need the deployment was meeting, and the outcome it was meant to produce. This is the part of the response to a shadow AI deployment that returns something to the organization. Shadow use is evidence of a gap between what people are asked to do and what traditionally they have been given to do it with. Where the need is legitimate and the deployment can be brought within this policy, the answer to a discovery is approval rather than prohibition.

The deployment’s work product needs its own assessment, sized to consequence through the organization’s existing change management review. Stopping the system does not retire what the system may have already produced, and that output can sit in systems of record, in customer correspondence, or behind decisions still in force. Much of it may prove unapproved and harmless as most shadow AI usage likely falls within Tier 1’s open-ended chat interactions. Where the effects persist and require remediation, the matter goes to the domain custodian for the affected domain under the escalation pathway Chapter 3 established.

A discovered deployment does not enter the Register on discovery. It is registered only if the organization subsequently approves it. The Register exists to answer one question unambiguously: what is sanctioned.

When Retiring an AI Deployment

A deployment that entered through the pipeline eventually exits, by decision or by circumstance, and an approval process with no defined and documented exit governs only half the tool’s lifecycle.

Organizations with a mature technology retirement process apply it here, and again, this policy does not attempt to build a second and separate process. What follows serves the organizations that have no such process, and identifies the obligations an ordinary retirement process does not anticipate. Mature organizations may benefit from the suggested processes, further tightening their existing retirement process.

Not surprisingly, the decision to retire a deployment is made and owned at the level that approved the deployment or above it. The contract is read for termination terms, notice periods, transition assistance, data retention, and the disposition of the organization’s content. Human resources are addressed, including the people whose work changes when the deployment is withdrawn. A risk assessment establishes what depends on the deployment, which is where the surprises usually live. Given enough time and awareness, dependencies accumulate around a useful tool regardless of the existence of proper documentation. The retirement is documented, then tested in whatever form the deployment permits, with the test looking specifically for consequences no one anticipated. Execution of scripts, migrations, file and permissions changes, etc. follow to cause the deployment’s effectual retirement. Change management records are updated and the Register entry moves to “retired” with its retirement record attached.

Four obligations distinguish this from retiring other software.

Work product does not retire when the deployment does. Retiring a sanctioned deployment predicates obligation of preservation: code that shipped stays in the repository, analyses that informed decisions stay in the record, generated content have disclaimers about retired tooling used to produce the content. What the organization loses when access is lost is the ability to explain or regenerate the work, and it loses the audit trail entirely where that trail lived only in the vendor’s console. An audit trail is therefore exported before access ends.

Removal is the right answer for a different case, and the two must be kept apart. Work product discovered at retirement that was never documented, produced by a deployment operating outside what it was registered for, is made ineffectual by whatever standard the organization considers sufficient: deletion from systems and repositories, removal from production, or halting the process that continues to generate it. An organization that applies this treatment to legitimate work will destroy the output of years of engineering work. Complete removal must be a last resort as it destroys any audit trail. Organizations must apply industry standard practices of care regarding retention when considering removal of AI output.

Vendor-held content is another obligation deserving consideration and divides into two halves with different remedies. Content the organization submitted: prompts, uploaded documents, and conversation history; is generally deletable through the interface the vendor already provides, and that deletion is part of executing the retirement. Content the vendor retained for training is beyond any action the organization can take at retirement. Whether it exists at all was settled by the training-data term in the agreement, which is why that clause is a condition of approval for regulated and restricted data.

Migration across a tier boundary is a governance event and the final obligation. When operators move from a retired deployment to its replacement, the replacement carries its own entries in the Register, tiers, and the tier governs what work each is approved to perform. A Tier 2 retrieval-augmented deployment replaced by a Tier 1 open-ended chat assistant is a reduction in harness affecting everyone who used it, and work that was correctly tiered last week is mis-tiered this week without anyone having been informed. Register entries are updated before the migration, and the operators are told what changed about the work they are permitted to do.

Retirement is not always a deliberate choice. A vendor discontinues a product, is acquired by a competitor of the organization, or changes terms the organization now finds unsuitable. A general retirement process assumes a planned decision on a timeline the organization controls. Deployments at Tier 3 and Tier 4 therefore warrant a note in the Register on how replaceable they are, because a harness tuned to the behavior of one model does not port on short notice, and the week the notice arrives is a poor time to find that out.

The second corollary of the Accountability Principle asks the organization to answer for what it deployed. The Register as discussed here and presented in Appendix A, is where the answer is kept.