InsightsGuideJuly 202611 min read

How to select AI governance tools without buying a filing cabinet.

Most AI governance tools are register software with a compliance vocabulary. This is the evaluation we run with clients: what to test, which questions separate enforcement from documentation, and how to map a platform onto the EU AI Act, the NIST AI RMF, and ISO 42001 before you sign anything.

The AI governance category assembled itself quickly, and the labels have not settled. A model registry, a risk-assessment questionnaire, an evaluation harness, a data-lineage scanner, and a full policy-as-code control plane are all sold as "AI governance platforms". They solve genuinely different problems, and buying the wrong one is expensive in a specific way: you discover the gap during an audit or an incident, when replacing the tool is the least available option.

What follows is the evaluation we use. It is deliberately unromantic about vendor categories and deliberately strict about one distinction — whether the tool enforces a control or merely records that someone said the control exists.

First, decide which problem you are solving

Three needs get conflated, and a tool that is excellent at one is usually mediocre at the others.

Visibility. You do not know how many models, agents, and third-party AI features are in production, who owns them, or what data they touch. Your first requirement is discovery, not policy — an inventory built from what your infrastructure actually runs, not from a form people are asked to fill in.

Control. You know what is running and need to stop specific things from happening: an unevaluated model reaching production, personal data leaving a jurisdiction, a high-risk use case deploying without human oversight. This requires enforcement inside the delivery pipeline.

Evidence. You have controls and need to demonstrate them repeatedly to regulators, auditors, and enterprise customers at acceptable cost. This requires a tamper-evident, queryable record produced as a by-product of the work.

Ask a vendor which of the three they are strongest at. The ones worth working with will answer honestly.

Seven criteria that actually separate platforms

1. Controls execute as code, in your pipeline. The decisive question is what happens when a rule is violated. If the answer is "a ticket is created" or "it appears on a dashboard", you have reporting. If the answer is "the deployment fails", you have a control. Ask to see a failing check block a real pipeline step during the evaluation, not in a recorded demonstration.

2. Evidence is append-only and queryable. Records that the governed team can edit are a courtesy. Test three things: can a record be altered after the fact, is the ordering verifiable, and can you answer an arbitrary question — "which decisions did model X influence last quarter, and who approved each change?" — in minutes rather than weeks.

3. Inventory is discovered, not declared. A registry populated by self-reporting decays from the day it is created. Prefer platforms that read from your cloud accounts, CI/CD, and model-serving infrastructure and flag what they find that nobody registered.

4. Framework mappings are maintained, not marketed. Every vendor claims EU AI Act, NIST AI RMF, and ISO 42001 coverage. The useful question is mechanical: when an obligation changes, who updates the mapping, how quickly, and does the platform re-evaluate historical decisions against the rules in force at the time rather than today's rules? Retrospective evaluation against current rules is unfair to your teams and useless to an auditor.

5. Privacy and residency are first-class. AI governance and data governance separate badly. Ask where evaluation data, prompts, and evidence are stored; whether personal data can be pinned to a jurisdiction; whether the vendor's own models or staff can read your content; and what the deletion path looks like. Vague answers here are the most reliable negative signal in the entire process.

6. It fits the stack you already have. Governance that requires engineers to leave their tools is governance that gets routed around. Native integration with your CI/CD, identity provider, ticketing, and data catalogue matters more than feature breadth.

7. Your evidence can leave. Audit records outlive vendor relationships. Confirm you can export the complete evidence store in an open, self-describing format — and that the export is legible to someone who has never used the platform.

Reading the frameworks as requirements, not reading material

The EU AI Act is the sharpest source of concrete tooling requirements for high-risk systems: risk management across the lifecycle, data and data-governance quality measures, technical documentation, automatic logging, human oversight, accuracy and robustness, and post-market monitoring. Translate each into a question about the tool — where does this artefact live, who generates it, and how is it produced without manual assembly?

The NIST AI RMF is voluntary and organised around Govern, Map, Measure, Manage. Its value in a procurement is structural: Measure and Manage are where tooling earns its place, and a platform that helps only with Govern and Map is a documentation product.

ISO/IEC 42001 describes a management system rather than controls. It is the right reference if you need certification, and the wrong one to evaluate whether a platform can stop a bad deployment.

No tool makes an organisation compliant. Tools make compliance cheap enough to be continuous.

A four-week evaluation that produces a real answer

Week one — pick one high-consequence decision your organisation already makes with model assistance. Write down what you would need to prove about the last hundred instances of it: inputs, model version, approver, data location, outcome.

Week two — attempt that proof with what you have today. The gap is your requirement list, ordered by materiality, and it will be more honest than any maturity assessment.

Week three — run two vendors against the same gap on your own infrastructure with your own pipeline. Insist on connecting a real repository. A platform that cannot be trialled against real delivery is telling you something.

Week four — attempt the proof again and measure elapsed time and manual steps. That number, not the feature matrix, is the product.

Build, buy, or both

Buy the commodity: registries, evidence stores, framework mappings, monitoring plumbing. Build the parts that encode your own risk appetite and domain-specific controls, because those are the ones a vendor cannot know and cannot maintain for you. The failure mode of building everything is a governance platform nobody owns eighteen months later. The failure mode of buying everything is a set of controls that describe a delivery process you do not have.

This is the thesis behind Lumiaxiom, our compliance-as-a-service and code-compliance platform: controls expressed as executable checks, evidence emitted by the work rather than written about it, and framework mappings maintained as obligations change. We built it because the evaluation above kept ending the same way — strong documentation products, and very little that would actually refuse a bad deployment at two in the morning.

Whichever way your own evaluation lands, hold the line on one question: can this system answer for itself? Everything else is a promise that someone will remember.

Further reading

Written by the senior practitioners at AITW Authentica. To discuss how this applies to your organisation, start a conversation.