An AI governance checklist is a short set of questions that tests whether an AI vendor gives you control, visibility, and proof: can you see every action, require approval before anything reaches a customer, limit what data and tools the AI touches, cap spend, and export the record when compliance asks. Eight questions cover it. A vendor with real governance answers all eight in minutes; a vendor without it changes the subject.
This post is that checklist. It’s deliberately vendor-neutral — every question works against any AI platform you’re evaluating, including ours. We build Turtle AI Coworker around exactly these controls, so yes, we have a horse in this race. But the honest version of this pitch is: ask everyone the same questions and let the answers sort the field. If a vendor stumbles on a question, believe the stumble.
One framing note before the list. Governance is not the enemy of getting value from AI. The teams that skip these questions are usually the ones whose pilots quietly die — not because the AI was bad, but because nobody could answer “what did it do last Tuesday?” and trust never formed. Governance is what lets a cautious owner say yes.
The 8 questions
1. Can I see every action the AI took?
Why it matters. The first time something goes wrong — a strange email, an unexpected charge, an angry customer — your first question will be “what exactly happened?” If the vendor can’t answer at the level of individual actions, you’re debugging a black box, and every incident becomes a trust crisis instead of a five-minute lookup.
What a good answer looks like. A per-run activity history: every task the AI executed, and within each task, every tool call it made — what it read, what it wrote, what it sent, with timestamps, in order. Not a monthly usage report. Not “logs are available on request.” A screen you can open yourself, today, and drill from “the AI handled 40 tickets yesterday” down to “here is the exact draft it produced for ticket #38 and where the source text came from.”
2. Can I require approval before anything reaches a customer?
Why it matters. The highest-consequence AI failure mode is not a wrong answer — it’s a wrong answer sent. A draft that’s 90% right is a time-saver; the same draft auto-sent is a customer incident. Whoever controls the send controls the risk.
What a good answer looks like. Draft-and-approve as a first-class mode, not a workaround: the AI prepares the reply, the invoice reminder, the candidate email — and a human reviews and sends. Autonomy should be settable per AI worker, so your invoice-chaser can run looser than your customer-facing support drafter. Ask specifically: “can I configure this AI so nothing it writes reaches a customer without a named human approving it?” The mechanics of doing this well are their own topic — we’ve written up how approval workflows actually work in practice.
3. Can I control what data and tools each AI worker touches?
Why it matters. An AI with access to everything is an incident with access to everything. Your support AI has no business reading payroll; your recruiting AI doesn’t need the invoice table. Blast radius is a design decision, and it’s made here.
What a good answer looks like. Per-worker resource permissions: this AI employee can read these tables, search these knowledge bases, and use these specific tool connections — and nothing else. Grants should be explicit and reviewable, not “the AI has an API key to your CRM.” If the vendor’s answer is one shared credential for the whole platform, that’s your answer.
4. Can I cap spend?
Why it matters. AI usage costs scale with activity, and autonomous activity is the point. A retry loop, a runaway task, or simple month-two enthusiasm can turn a $300 line item into a $3,000 one. Finance will ask; you should be able to answer before they do.
What a good answer looks like. Hard budgets and limits you set — per period, ideally per AI worker — with alerts before the cap and a stop at it. “You can monitor your usage dashboard” is not a cap. A cap is: it stops.
5. What always escalates to a human?
Why it matters. Some situations should never be handled by software at full autonomy, no matter how good the drafts are: a furious customer, a refund over threshold, anything with legal or privacy language, a safety hazard. If escalation depends on the AI deciding to escalate case-by-case with no rules behind it, you’re trusting judgment exactly where judgment is least reliable.
What a good answer looks like. Named, built-in escalation categories — not just a configurable option you have to think to set up. In Turtle’s packs this is baked into the domain: the support pack always escalates anger, churn risk, refunds, account deletion, and legal or privacy issues; the e-commerce pack always escalates denials, over-threshold refunds, and damaged items; the field-trades pack flags safety hazards for immediate human attention and never calls a hazard resolved without a technician on site. Whatever vendor you pick, make them list their equivalents out loud.
6. Are domain guardrails built in?
Why it matters. Every regulated or liability-heavy domain has lines the AI must never cross — and “we told it not to in the prompt” is not a control. If you’re a law firm, an insurance agency, or a staffing firm, a single crossed line is a malpractice claim, an E&O exposure (errors and omissions — professional liability), or a discrimination complaint.
What a good answer looks like. Refusals designed into the product for your domain, stated in writing. Concretely, the kind of thing to listen for: a legal-ops AI that never gives legal advice, assesses merits, or clears a conflict check; recruiting screening that considers only skills and qualifications and never rejects a candidate or extends an offer; an insurance assistant that never quotes a premium or binds coverage. Those are the actual guardrails in Turtle’s legal, recruiting, and insurance packs — and any serious vendor in your domain should be able to recite theirs just as specifically.
7. Who can change the rules, and is that logged?
Why it matters. Approval settings, permissions, and budgets are only as strong as the controls on changing them. If any user can quietly flip an AI worker from draft-and-approve to auto-send, questions 2 through 6 are decorative.
What a good answer looks like. Role-based access — a defined set of people who can change autonomy, permissions, and budgets — plus a record of configuration changes: who changed what, when. When an auditor asks “who authorized this AI to send emails?”, the answer should be a lookup, not an investigation.
8. Can I export the record if compliance or legal asks?
Why it matters. Eventually someone outside your company will ask what your AI did: an auditor, a regulator, opposing counsel, a big customer’s security questionnaire. An audit trail you can see but not produce is only half a control.
What a good answer looks like. History and evidence you can get out — the run and action records from question 1, exportable in a usable format, covering the period in question. Ask the vendor to actually show you an export, not describe one.
Does governance slow the AI down?
Less than you’d think, and the trade is usually worth making explicit. Draft-and-approve adds a human minute or two per outbound item — that’s the cost. What it buys is the ability to run the AI on real customer-facing work at all, instead of restricting it to internal chores because nobody quite trusts it. Permissions and budgets cost nothing at runtime; they’re constraints, not steps. The one place governance genuinely adds friction is approvals piling up when volume grows — which is a routing and delegation problem, and a much better problem than an unsupervised AI.
Frequently asked questions
Do small businesses really need AI governance?
Yes — a lighter version of it. A 5-person firm doesn’t need a governance committee, but it absolutely needs questions 1, 2, and 3: see what the AI did, approve what goes out, and limit what it touches. Small firms have less slack to absorb an AI-caused customer incident, not more.
Does governance slow down the AI’s usefulness?
Approval adds minutes per outbound item; permissions and budgets add nothing at runtime. In practice governance speeds up adoption, because owners extend autonomy when they can verify behavior. The slowest AI rollout is the one that gets paused indefinitely after an incident nobody can explain.
What’s the minimum governance for a pilot?
Three things: full activity history on (question 1), draft-and-approve on everything customer-facing (question 2), and access limited to the pilot’s own data and tools (question 3). Add a spend cap if usage pricing is involved. That’s enough to run a 30-day pilot you can defend to anyone.
How should I use this checklist in an actual vendor evaluation?
Send all eight questions to every vendor in writing, before the demo. In the demo, ask them to show the audit screen, the approval flow, and an export — not slides about them. Score answers as shown / claimed / absent. Vendors with real controls love this exercise; that reaction is itself a signal.
If you want to see what “yes” to all eight looks like in a working product, install any Turtle Solution Pack — the AR Collections, Customer Support, or Legal Intake packs all ship with draft-and-approve, built-in escalations, and per-worker permissions wired in — and run it as a 30-day pilot with the audit history open. Bring this checklist to the review meeting.