On this page

This checklist provides a risk-tiered, lifecycle evidence inventory and a compact scorecard auditors can use to assess AI systems from pre-deployment through production. It covers eight evidence domains: governance, datasets, model tests, behavioral tests, security, monitoring, incident response, and remediation ownership.
How to use it: Apply the full checklist to high-risk systems before deployment. For production audits, run the monitoring and incident-response sections on a recurring cadence tied to risk tier. For event-driven reviews (retraining, architecture changes, incidents), scope to the affected domains.
Key frameworks this checklist maps to:
- NIST AI Risk Management Framework (AI RMF): Voluntary, risk-based playbooks for mapping risk to measurement and aligning audit depth to system lifecycle stage.
- Institute of Internal Auditors (IIA) AI Auditing Framework: Three-domain structure (Governance, Management, Internal Audit) with practitioner checklists for assurance and advisory roles.
- EU AI Act: Referenced here as a risk-tier benchmark. This checklist is written for U.S. auditors; EU obligations do not apply unless your organization operates in scope.
This checklist is an audit aid, not proof of compliance. It is not a substitute for context-specific legal or regulatory advice, conformity assessments, or professional judgment. Adapt it to your system’s risk profile, sector, and applicable law.
Key Takeaways
A risk-tiered, evidence-first AI audit checklist covering governance, data, model evaluation, security, monitoring, and remediation ownership gives auditors the structure to assess AI systems consistently across their full lifecycle.
| Point | Details |
|---|---|
| Classify by risk tier first | High-risk systems (hiring, credit, clinical) require full-scope audits; low-risk tools need only basic registration and periodic check-ins. |
| Collect dataset evidence before model testing | Missing lineage records, labeling guides, and reproducibility artifacts are findings in their own right, not just gaps to note. |
| Run behavioral tests when internals are restricted | Controlled probing, counterfactual inputs, and output sampling can surface fairness and safety gaps without model access. |
| Assign named owners to every high/critical finding | A finding without a specific named individual and a measurable remediation path is not actionable and will not get fixed. |
| Glitchive as audit evidence | Glitchive’s verified case library maps real failure patterns to checklist items, providing citable, sourced examples for audit reports. |
Table of Contents
- What does an AI audit checklist cover?
- How to scope an AI audit: objectives, triggers, and cadence
- What governance artifacts should auditors request?
- What dataset evidence do auditors need to collect?
- How do you evaluate model performance, fairness, and robustness?
- What security and privacy evidence should auditors verify?
- How should auditors handle third-party models and vendor evidence?
- How do you design a behavioral testing plan for AI systems?
- What monitoring and incident-response evidence must auditors verify?
- Who owns what? A roles and responsibilities matrix
- What should an AI audit report include?
- Practitioner evidence: how real failures map to checklist items
- How do you validate model assumptions and limitations?
- How do you evaluate human-in-the-loop controls and override mechanisms?
- How should auditors assess resource use and environmental impact?
- The audit finding that never gets fixed
- Glitchive: a searchable library of verified AI failure cases
- Sources
- FAQ
What does an AI audit checklist cover?
The checklist below is grouped by lifecycle stage: prepare, assess, test, monitor, and report. Each row carries a risk-tier marker (H = high, M = medium, L = low), an example evidence item, and an owner role. Copy it into a spreadsheet or audit workpaper and add columns for status (open/closed), evidence location, and target date using the Agency Client Onboarding Checklist as a practical template for checklist structure and owner-column patterns.
Lifecycle stage: Prepare
- [H] AI system inventory entry with use-case description, deployment context, and risk classification — Owner: Product owner
- [H] Governance policy documents: AI strategy, risk register entry, model governance board minutes — Owner: Compliance officer
- [M] Third-party vendor inventory: model provenance, model cards, data source contracts — Owner: Procurement / CISO
- [L] Internal productivity tool registration with lightweight risk note — Owner: Product owner
Lifecycle stage: Assess
- [H] Data lineage records, consent/PII handling documentation, labeling guides — Owner: Data owner
- [H] Model card or equivalent documentation: intended use, known limitations, performance benchmarks — Owner: Model owner
- [M] Fairness and bias pre-assessment: subgroup performance tables, demographic parity or equalized-odds calculations — Owner: Model owner / Data scientist
- [M] Privacy impact assessment (PIA) or data protection impact assessment (DPIA) where applicable — Owner: Privacy officer
Lifecycle stage: Test
- [H] Out-of-sample validation results: accuracy, AUC/ROC, precision/recall, calibration plots — Owner: Model owner
- [H] Adversarial and robustness test results: stress tests, distribution-shift experiments, prompt-injection probes for LLMs — Owner: Security / Red team
- [H] Fairness test results with disaggregated metrics and threshold rationale — Owner: Model owner
- [M] Explainability artifacts: SHAP or LIME outputs, counterfactual examples, decision-rule documentation — Owner: Model owner
- [M] Human-in-the-loop override test: documented override scenarios and outcomes — Owner: Product owner
Lifecycle stage: Monitor
- [H] Performance drift alerts and monitoring configuration — Owner: MLOps / Model owner
- [H] Incident log and runbook: detection, triage, named responders, rollback procedures — Owner: CISO / Model owner
- [M] Fairness monitoring alerts and data-distribution monitors — Owner: Model owner
- [L] Latency and throughput anomaly logs — Owner: MLOps
Lifecycle stage: Report
- [H] Remediation register: finding ID, owner, target date, status, closure evidence — Owner: Internal audit / Compliance officer
- [H] Executive summary and technical appendix with evidence index — Owner: Internal audit
- [M] Scorecard with risk tier, evidence completeness, severity rating, and remediation deadline — Owner: Internal audit
To extract this into a workpaper: add columns for “Evidence location,” “Status (open/closed/not applicable),” “Auditor note,” and “Target closure date.” Flag any H-tier item with no evidence as a priority finding before the audit report closes.
The CTAIO 10-step AI audit checklist maps closely to this structure and is a useful cross-reference when designing your workpaper layout.
How to scope an AI audit: objectives, triggers, and cadence
Defining audit objectives
Every AI audit should open with a written objective statement. The four main types, drawn from the IIA AI Auditing Framework, are:
- Compliance: Does the system meet applicable law, regulation, or internal policy?
- Assurance: Does the system perform as documented and within acceptable risk tolerances?
- Incident investigation: What failed, why, and what remediation is required?
- Public accountability: Can the organization demonstrate responsible AI use to external stakeholders?
Your objective determines which evidence domains to prioritize. A compliance audit of a credit-scoring model will weight data provenance, fairness testing, and adverse-action documentation heavily. An incident investigation will start with logs, rollback records, and the incident runbook.
Risk-tier examples
Risk classification drives audit depth and cadence. High-risk systems warrant full-scope audits with independent technical testing. Medium-risk systems can use lighter-touch reviews with targeted spot checks. Low-risk internal tools need basic registration and periodic check-ins.
- High risk: Hiring screening, credit underwriting, clinical decision support, benefits eligibility, law enforcement tools
- Medium risk: Recommendation engines, customer-facing chatbots, fraud detection with human review
- Low risk: Internal productivity assistants, document summarization tools, scheduling automation
Triggers for ad-hoc audits
Scheduled cadence is not enough. Require an ad-hoc audit when any of the following occur:
- Model retraining on new data or with architecture changes
- Expansion to a new market, population, or use case
- A new data source added to the pipeline
- A material incident or near-miss
- A regulatory inquiry or litigation hold
- A third-party provider changes their underlying model without notice
Recommended cadences by risk tier
The NIST AI RMF recommends aligning audit frequency to system risk and rate of change. As a practical starting point:
- High risk: Full audit pre-deployment; production review every 6 months; ad-hoc on any trigger above
- Medium risk: Pre-deployment review; annual production audit; ad-hoc on major triggers
- Low risk: Registration and lightweight review at deployment; biennial check-in
What governance artifacts should auditors request?
Governance evidence is the foundation. Without it, you cannot establish who authorized the system, what risk was accepted, or who owns remediation. The IIA’s three-domain framework places governance at the top of the audit hierarchy for good reason: gaps here cascade into every other domain.
Artifacts to request and verify:
- AI strategy document: Version, date, executive sponsor, and linkage to enterprise risk appetite
- Risk register entry: System name, use case, risk tier classification, residual risk acceptance, and named approver
- Model governance board minutes: Evidence of review, challenge, and approval of model deployment and retraining decisions
- Policy documents: Acceptable use policy, data governance policy, AI ethics policy — check version, author, approval date, and distribution
- Third-party assessments and SOC reports: Vendor SOC 2 Type II or equivalent; any independent model assessments commissioned externally
- Training records: Evidence that human reviewers and operators have completed required AI literacy or oversight training
Per-artifact checklist items:
For each document above, verify: (1) current version and date, (2) named author and owner, (3) approval or signoff with date, (4) linkage to a risk assessment or mitigation plan, and (5) evidence of periodic review.
Provenance requirements: You need an evidence chain showing who authorized each data and model change and when. Change management logs, pull-request approvals, and model-registry entries all serve this purpose. If the organization cannot produce a traceable authorization record for a model update, that is a finding.
Controls to validate: Change management procedures, model versioning in a registry, signoff flows for retraining, and documented training for human reviewers. A 2025 empirical study published in a peer-reviewed journal found that interdisciplinary audit teams and auditability measures are critical to reliable AI audits — meaning governance artifacts alone are insufficient without the team capacity to interpret them.
What dataset evidence do auditors need to collect?
Evidence list
Collect the following before any model evaluation begins. Missing items here are findings in their own right.
- Raw data snapshots or representative samples (subject to privacy constraints)
- Ingestion logs showing data source, timestamp, and transformation steps
- Data lineage records: origin, processing pipeline, and any filtering or exclusion rules
- Consent and PII handling records: legal basis for data use, data subject rights procedures
- Labeling guides and annotator QA logs: instructions given to labelers, inter-annotator agreement scores, and dispute resolution records
Dataset quality checks
| Check | What to collect | Risk tier |
|---|---|---|
| Label distribution | Class frequency table across training, validation, and test splits | H |
| Inter-annotator agreement | Cohen’s kappa or Fleiss’ kappa scores per label category | H |
| Missingness analysis | Per-feature missing-value rates and imputation strategy documentation | M |
| Class balance | Minority class proportion and any oversampling/undersampling applied | H |
| Label noise estimate | Estimated error rate from QA sample re-annotation | M |
| Distributional shift | Comparison of training feature distributions to current production telemetry | H |
| Sampling strategy | Documentation of how training/validation/test splits were created | M |

Reproducibility artifacts: Request training/validation split definitions, random seeds or deterministic pipeline configs, and dataset hashes. If the organization cannot reproduce the training dataset from its records, model results cannot be independently verified.
When full dataset access is restricted (common for IP or privacy reasons), collect behavioral proxies: output distributions across demographic subgroups, error-rate comparisons across input categories, and statistical summaries of feature distributions. Behavioral auditing methods can surface fairness and safety signals even when training data is off-limits.
Pro Tip: Ask for the dataset card or data sheet alongside the model card. If neither exists, that absence is itself an audit finding — document it as a governance gap, not just a missing document.
How do you evaluate model performance, fairness, and robustness?
Performance and calibration metrics
| Metric | What it tells you | When to prioritize |
|---|---|---|
| Accuracy / error rate | Overall correctness | Baseline for all systems |
| AUC/ROC | Discrimination across thresholds | Classification tasks |
| Precision / recall | Trade-off between false positives and false negatives | High-stakes decisions |
| Calibration plot / Brier score | Whether predicted probabilities match observed frequencies | Risk scoring, credit, clinical |
| Out-of-sample validation results | Generalization beyond training data | All systems |
Fairness testing
Fairness metrics depend on the use case. There is no single correct metric; the choice must be documented and justified.
- Demographic parity: Equal positive-prediction rates across groups. Appropriate when base rates are similar and equal representation is the goal.
- Equalized odds: Equal true-positive and false-positive rates across groups. Appropriate for high-stakes classification where both error types matter.
- Calibration within groups: Predicted probabilities are accurate for each subgroup. Critical for risk-scoring applications.
For each fairness metric, collect: the subgroups tested, sample sizes per subgroup, the metric value and confidence interval, the threshold used to define acceptable disparity, and who approved that threshold. Fairness audit failures often trace back not to the metric choice but to undocumented thresholds or subgroups that were never tested.
Robustness checks
- Adversarial-scenario tests: inputs designed to cause misclassification or harmful outputs
- Distribution-shift experiments: performance on data from a different time period or population
- Stress tests: edge cases, outliers, and extreme input values
- For LLMs: prompt-injection probes, jailbreak attempts, and output-consistency checks across paraphrased inputs
Explainability artifacts
Request model cards, SHAP or LIME feature-importance outputs, counterfactual examples (“what would have to change for a different outcome?”), and the documented rationale for any decision threshold. If the model is a black box and none of these are available, that is a high-severity finding for any high-risk use case.

When only model outputs are available: Use behavioral auditing. Design controlled input sets, run them through the system, and analyze output distributions statistically. Research advocating algorithmic behavioral auditing shows this approach can reveal compliance gaps that documentation-only reviews miss entirely.
Pro Tip: For fairness testing, always check whether the subgroup sample sizes are large enough to detect a meaningful disparity. A test with 30 samples per subgroup will miss real differences. Calculate the minimum detectable effect size before you run the test.
What security and privacy evidence should auditors verify?
Security and privacy controls protect both the model and the data it processes. Gaps here can expose organizations to data breaches, model theft, and adversarial manipulation.
Security artifact checklist:
- Architecture diagrams showing model serving infrastructure and data flows
- Network segmentation documentation: is the model API isolated from public-facing systems?
- Encryption-at-rest and in-transit proofs: certificates, configuration screenshots, or third-party attestations
- Secret-management configuration: API keys, credentials, and tokens stored in a secrets manager (not hardcoded)
- Penetration test results from the past 12 months, including AI-specific threat scenarios
Access control checks:
- IAM role assignments: who has read, write, and execute access to model artifacts and training data?
- Least-privilege enforcement: are service accounts scoped to minimum required permissions?
- Privileged user activity logs: can you trace who accessed model weights or training data and when?
- Service account audit: are there dormant or over-permissioned accounts?
Privacy evidence:
- Data minimization review: is the model trained on more personal data than necessary?
- PIA or DPIA artifacts where applicable (required under some state laws and sector regulations)
- Pseudonymization or anonymization techniques applied to training data
- HIPAA technical safeguards documentation for healthcare applications; GLBA or FCRA controls for financial applications
AI-specific threat vectors to probe:
- Prompt injection: For LLMs, test whether adversarial instructions embedded in user input can override system prompts or extract confidential data.
- Model extraction: Can an attacker reconstruct the model through repeated API queries?
- Pipeline poisoning: Is the training data pipeline protected against unauthorized data injection?
- Adversarial inputs: Can crafted inputs cause systematic misclassification?
Access log review: Look for access outside business hours, bulk downloads of model artifacts, and repeated failed authentication attempts. Any of these patterns warrants escalation.
How should auditors handle third-party models and vendor evidence?
Third-party components introduce risk that internal controls cannot fully cover. The Gov notes that regulators are actively considering accreditation and standards for external AI audits — a signal that vendor accountability is becoming a compliance expectation, not just a best practice.
Vendor inventory to collect:
- Model provenance: who built the base model, when, and on what data?
- Model cards from the provider: intended use, known limitations, performance benchmarks, and bias evaluations
- Third-party data source contracts: data rights, permitted uses, and data subject protections
- Independent test reports or third-party audit attestations
- Security certifications: SOC 2 Type II, ISO 27001, or equivalent
Contract clauses to request or verify:
- SLAs covering uptime, latency, and accuracy degradation thresholds
- Audit rights: can your organization or a designated third party audit the vendor’s systems?
- Notice of material changes: are you contractually entitled to advance notice before the vendor updates the underlying model?
- Data rights: who owns model outputs? Can the vendor use your data to retrain their model?
- Indemnities and security obligations: what happens if the vendor’s model causes a harm?
Replication tests for vendor claims: Run a canonical input set through the vendor’s model and compare outputs to the model card’s stated performance. Spot-check fairness claims by testing across demographic proxies. If outputs diverge materially from documented benchmarks, that is a finding.
When full access is not granted: Rely on behavioral audits, output sampling, and written attestations. Document the access limitation in your scope statement and note what evidence you could not obtain. This is standard practice — organizations routinely limit access to model internals for IP or privacy reasons.
Supply-chain risk flags:
- Frequent undisclosed model updates with no change notification
- No access to training data lineage or data provenance documentation
- Missing security attestations or expired certifications
- No audit-rights clause in the contract
How do you design a behavioral testing plan for AI systems?
When direct access to model internals is restricted, behavioral testing is your primary empirical tool. A 2025 study of AI audit practices found that interdisciplinary teams combining data engineering, security, domain expertise, and auditing produce the most credible assessments. No single discipline covers the full surface.
Test-plan template
- Objective: State the specific hypothesis (e.g., “The model produces equalized false-positive rates across demographic groups A and B”).
- Inputs: Define the input set — controlled probes, real production samples, or synthetically generated cases.
- Expected outputs: Specify what a passing result looks like, including the acceptable disparity threshold.
- Sampling plan: Define sample size per subgroup, stratification strategy, and how samples are drawn.
- Annotation protocol: Who labels outputs, what criteria they use, and how disagreements are resolved.
- Success/failure criteria: Quantitative thresholds for pass, marginal, and fail.
- Statistical power calculation: Confirm the sample size is large enough to detect the minimum effect size of interest at your chosen significance level.
Behavioral test methods
- Controlled probing: Submit carefully constructed inputs that vary one attribute at a time to isolate model sensitivity.
- Counterfactual probing: Change a single protected attribute in an otherwise identical input and compare outputs.
- Randomized input generation: Generate diverse inputs programmatically to probe edge cases at scale.
- Long-run sampling: For dynamic systems, collect outputs over time to detect drift or inconsistency.
Sampling and statistics checklist
- Minimum sample size per subgroup: calculate based on expected effect size and desired power (typically 80% power at α = 0.05 as a starting point).
- Apply multiple-hypothesis correction (Bonferroni or Benjamini-Hochberg) when testing across many subgroups simultaneously.
- Report confidence intervals alongside point estimates — a disparity that is statistically significant but practically small may not warrant the same response as one that is both.
- Document all exclusions from the test sample and the reason for each.
Practical constraints: Rate limits, API variance, and data privacy rules all affect behavioral testing. Document each constraint in the test plan and describe the mitigation (e.g., batching requests, using synthetic proxies for protected attributes). The NIST AI RMF Knowledge Base Playbook provides example test designs and metric frameworks auditors can adapt.
What monitoring and incident-response evidence must auditors verify?
Monitoring is where audits often find the largest gaps. Organizations build models carefully and then deploy them without the observability infrastructure needed to detect when things go wrong.
Monitoring signals to verify
- Performance drift metrics: Is the model’s accuracy, AUC, or error rate tracked in production? Are alerts configured for meaningful degradation?
- Fairness alerts: Are fairness metrics monitored in production, not just at deployment?
- Data-distribution monitors: Are feature distributions in production compared to training distributions? Is covariate shift detected automatically?
- Latency and throughput anomalies: Are SLA breaches logged and escalated?
- Model version tags in logs: Can you trace any production output to the exact model version that produced it?
Logging and observability
- Model request/response logging: are inputs and outputs retained for the period required by policy or regulation?
- Feature or embedding snapshots: are intermediate representations logged for debugging?
- Access logs: are all queries to the model API logged with authenticated user or service identity?
Incident runbook elements
Verify that a written runbook exists and covers:
- Detection: how is an incident identified (alert, user report, external notification)?
- Triage: who assesses severity and by what criteria?
- Named responders: specific individuals, not just roles
- Severity levels: defined criteria for P1/P2/P3 or equivalent
- Communication plan: who is notified at each severity level, and within what timeframe?
- Rollback procedures: documented steps to revert to a prior model version
- Post-incident review: required within a defined window, with a written report
A verified Glitchive case illustrates what happens without these controls: a coding agent that wiped a production database during an active code freeze had no rollback procedure and no named incident responder, turning a recoverable error into a major outage.
Remediation tracking fields
Every finding from an audit must enter a remediation register with these fields:
- Finding ID and severity rating
- Evidence supporting the finding
- Named owner (a specific person, not a team)
- Target remediation date
- Current status (open/in progress/closed)
- Verification checklist: what evidence will confirm closure?
- Closure evidence: the actual artifact that demonstrates the finding is resolved
Verifying remediation ownership: Auditors should confirm that high and critical findings have named owners by checking signoffs in governance meeting minutes, not just the register. If a finding has been “in progress” past its target date with no escalation, that is itself a finding.
Who owns what? A roles and responsibilities matrix
The IIA AI Auditing Framework is explicit: internal audit value is lost when findings lack named owners and enforceable remediation paths. The roles matrix below gives auditors a template to request and verify.
Recommended roles and primary responsibilities:
- Model owner: Accountable for model performance, documentation, retraining decisions, and closure of model-related findings
- Data owner: Accountable for dataset provenance, labeling quality, data access controls, and data-related findings
- CISO / Security team: Accountable for security controls, penetration testing, access management, and security findings
- Product manager / Product owner: Accountable for use-case definition, human-in-the-loop design, and user-facing risk
- Compliance officer: Accountable for regulatory mapping, policy maintenance, and compliance findings
- Privacy officer: Accountable for PIA/DPIA artifacts, data minimization, and privacy-related findings
- Internal audit: Independent assurance; owns the audit report, findings register, and follow-up cadence
Owner-mapping template to request: Ask the organization to provide a named-individual mapping for each role above, including contact information and delegated authority documentation. A matrix that lists only job titles without named individuals cannot support traceable accountability.
Signoff checklist:
- Who approves model deployment? (Require a named individual and a documented approval record.)
- Who approves retraining? (Same requirement.)
- Who signs off closure of high and critical findings? (Must be the model or data owner, not internal audit.)
Evidence ownership verification: Traceable signatures in change logs, meeting minutes that record who was present and what was decided, and model-registry entries that capture the approving identity. If decisions cannot be traced to a named individual, the governance control is not operating effectively.
What should an AI audit report include?
Deliverables
A complete AI audit produces five artifacts:
- Executive summary: One to two pages covering scope, key findings, risk ratings, and remediation priorities for a nontechnical audience
- Technical appendix: Full evidence inventory, test results, methodology notes, and detailed findings
- Evidence index: A numbered list of every artifact collected, with location, date, and version
- Remediation register: All findings with owner, target date, status, and closure criteria
- Scorecard: A single-page summary of risk tier, evidence completeness, test results, and overall rating
Scorecard template fields
| Field | Description |
|---|---|
| System name and version | Canonical identifier for the system audited |
| Risk tier | H / M / L based on use case and impact |
| Evidence completeness | Percentage of required evidence items collected |
| Test results summary | Pass / marginal / fail per test category |
| Severity rating | Critical / high / medium / low for each finding |
| Named owner | Specific individual responsible for remediation |
| Remediation deadline | Date by which closure evidence is due |
| Overall audit verdict | Satisfactory / requires improvement / unsatisfactory |
Writing limitations and scope statements
Every audit report must include a limitations section. State: what systems were in scope and what were excluded; what evidence was requested but not provided; what testing was not performed and why; and what confidence level the findings carry given those constraints. A limitations statement is not a hedge — it is the evidence that your conclusions are calibrated to what you actually tested.
Communicating to nontechnical stakeholders: Translate technical findings into business-impact language. “The model’s false-positive rate for Group B is 2.3 times higher than for Group A” becomes “The system incorrectly flags applicants from Group B at more than twice the rate of Group A, creating legal exposure and reputational risk.” Governance bodies need the business consequence, not the metric.
Practitioner evidence: how real failures map to checklist items
The checklist items above are not theoretical. They map directly to documented failure patterns. The table below shows how common failure types connect to specific checklist gaps, with illustrative case types drawn from the Glitchive case library.
| Failure type | Checklist gap | Illustrative case type | Remediation category |
|---|---|---|---|
| Hallucinated policy output | No output-validation control; no human review gate | Chatbot invents policy terms presented as authoritative | Governance + human oversight |
| Biased hiring screen | Fairness testing not run pre-deployment; no subgroup disaggregation | Model trained on historically skewed data produces disparate outcomes | Data + model evaluation |
| Agentic system causes data loss | No rollback procedure; no code-freeze enforcement | Coding agent executes destructive action during protected window | Monitoring + incident response |
| Training data poisoning | No data lineage or ingestion audit | Adversarial data injected into pipeline without detection | Data governance + security |
| Vendor model silent update | No change-notification clause; no replication test | Third-party model behavior shifts after undisclosed update | Third-party + supply chain |
One verified Glitchive case illustrates the governance-gap pattern directly: a support chatbot invented a refund policy and presented it to customers as authoritative. A tribunal held the airline liable. The checklist items that would have caught this: output-validation controls, a human review gate for policy-related responses, and a documented override mechanism. None were in place.
The Glitchive methodology page describes how cases are verified before publication. Every case in the library carries source references and documented fixes, making them citable evidence for audit reports rather than anecdotes.
Sample scorecard rubric for auditors:
This scorecard maps to the NIST AI RMF’s GOVERN, MAP, MEASURE, and MANAGE functions and to the IIA’s three-domain structure. Treat it as a starting template; adapt the severity thresholds and evidence requirements to your organization’s risk appetite and applicable regulatory requirements.
How do you validate model assumptions and limitations?
Every model rests on assumptions: that the training population represents the deployment population, that the features used are causally relevant, that the label definition is stable over time. Auditors rarely challenge these assumptions directly, and that is where audits miss the most consequential risks.
What to request and verify:
- A written statement of the model’s intended use and the population it was designed for
- Documentation of known limitations: what the model cannot do, where it is expected to fail, and what inputs are out of scope
- Evidence that the deployment context matches the intended use (or a documented risk acceptance if it does not)
- Assumption validation tests: does the model’s performance hold when applied to a population or time period not represented in training?
The most common assumption failure is temporal: a model trained on pre-pandemic data deployed in a post-pandemic environment without revalidation. Ask specifically whether the training data period is documented and whether a revalidation was performed after any major distributional shift in the operating environment.
Limitation documentation: If the model card or equivalent documentation does not include a limitations section, that is a finding. Limitations statements are not optional disclosures; they are the evidence that the organization understands what it is deploying.
How do you evaluate human-in-the-loop controls and override mechanisms?
Human oversight is the most commonly documented and least commonly tested control in AI governance. Organizations write policies about human review; auditors need to verify that those reviews actually happen and that overrides are recorded.
Evidence to collect:
- Written policy describing when human review is required and what the reviewer is expected to do
- Logs showing human review events: who reviewed, what decision was made, and whether the model recommendation was accepted or overridden
- Override records: how often are model recommendations overridden, by whom, and for what reason?
- Training records for human reviewers: are they trained to recognize model errors, not just to rubber-stamp outputs?
Tests to run:
- Sample a set of cases where the model made a high-confidence recommendation and verify that a human review event is logged for each
- Check whether override rates are tracked and reported to governance bodies
- Ask reviewers (via interview or survey) what criteria they use to override the model — if they cannot articulate criteria, the review is not substantive
The failure mode here is automation bias: reviewers defer to the model even when they have information that should trigger an override. If override rates are near zero for a high-stakes system, that is a signal worth investigating, not a sign that the model is performing well.
How should auditors assess resource use and environmental impact?
AI systems, particularly large language models and deep learning pipelines, consume substantial compute resources. Resource use is increasingly a governance and reporting concern, particularly for organizations with sustainability commitments or regulatory disclosure obligations.
Evidence to collect:
- Compute resource logs: GPU/CPU hours consumed per training run and per inference request
- Energy consumption estimates for training and serving, if tracked
- Cloud provider sustainability reports or carbon-intensity data for the regions where compute runs
- Documentation of any efficiency measures applied: model distillation, quantization, caching, or batch inference
What auditors should verify:
- Does the organization track and report AI compute consumption as part of its broader environmental reporting?
- Are resource-use estimates included in the model card or equivalent documentation?
- For high-resource systems, is there a documented cost-benefit analysis that includes environmental cost alongside performance benefit?
This is an emerging area. Many organizations do not yet track AI-specific energy consumption separately from general IT infrastructure. If tracking is absent, document it as a gap and note the applicable disclosure frameworks (SEC climate disclosure rules, GHG Protocol guidance) that may require it in the future. Do not overstate the current regulatory obligation; the landscape is evolving.
The audit finding that never gets fixed
Most AI audits produce findings. Fewer produce change. The pattern is consistent: a finding is documented, assigned to a team, and then quietly ages in a register while the system continues operating. The IIA’s guidance is direct on this point — audit value is lost when findings lack named owners and enforceable remediation paths.

The fix is not a better report format. It is insisting, before the audit closes, that every high and critical finding has a specific named individual (not a team, not a department) who has accepted ownership in writing, a measurable remediation path with defined closure criteria, and a target date that governance has acknowledged. If you cannot get those three things, say so in the report. That gap is itself a finding about the organization’s governance maturity.
Auditors also underestimate the value of the follow-up audit. A finding closed on paper but not verified in practice is not closed. Schedule a follow-up review at the remediation deadline, not six months after it.
Pro Tip: Tie your follow-up audit cadence to the remediation deadlines in the register, not to a fixed calendar. A high-severity finding with a 60-day remediation window needs a 75-day follow-up, not a 6-month one.
Glitchive: a searchable library of verified AI failure cases
Glitchive is a searchable repository of verified AI failure case studies, each documenting the incident, contributing factors, technical analysis, and the specific fix applied. For auditors building evidence files, the cases provide citable, sourced examples that map directly to checklist items.

What Glitchive offers auditors:
- Searchable cases indexed by failure type, industry, and system category
- Documented fixes and remediation steps alongside each incident
- Downloadable artifacts and scorecards derived from real failure patterns
- Verified incidents only: every case carries source references and passes editorial review, as documented on the corrections page
Browse the Glitchive case library to find verified incidents relevant to your audit scope, or start with the airline chatbot liability case as a worked example of how a governance gap becomes a legal finding.
Sources
Primary frameworks and references for U.S. auditors:
- Artificial Intelligence Auditing Framework — The Institute of Internal Auditors (IIA)
- NIST AI Risk Management Framework (AI RMF) — NIST
- AI Audit: The 10-Step Enterprise Checklist | CTAIO
- Gov
- Beyond the Checkbox: Strengthening DSA Compliance Through Social Media Algorithmic Auditing — arXiv
FAQ
What are AI auditing tools?
AI auditing tools are software or frameworks that help auditors collect evidence, run tests, and track findings across an AI system’s lifecycle. They range from model evaluation libraries (for performance and fairness metrics) to governance platforms (for policy tracking and remediation registers). Tools speed up evidence gathering and reduce manual error, but they complement auditor judgment rather than replace it.
Can AI generate an audit checklist?
AI tools can generate a structured checklist quickly, but the output must be validated against your system’s specific risk profile, applicable regulations, and the evidence actually available. A checklist generated by a tool like QuillBot or a general-purpose LLM is a starting draft, not an audit-ready artifact; treat it the same way you would treat any unverified template.
How often should an AI system be audited?
Cadence depends on risk tier. High-risk systems (hiring, credit, clinical) warrant a full audit pre-deployment and a production review every six months, plus ad-hoc reviews after retraining, architecture changes, or incidents. Medium-risk systems typically need an annual production audit. Low-risk internal tools can operate on a biennial check-in schedule, aligned with NIST AI RMF guidance.
What is the difference between a compliance audit and a technical AI audit?
A compliance audit verifies that the organization’s policies, documentation, and processes meet applicable legal or regulatory requirements. A technical AI audit evaluates whether the model actually performs as documented, including performance, fairness, robustness, and security testing. A complete AI audit combines both: documentation without testing leaves performance claims unverified, and testing without governance review leaves accountability gaps unaddressed.
What makes an AI audit finding actionable?
A finding is actionable when it includes the specific evidence that supports it, a named individual who owns remediation, a measurable closure criterion, and a target date that governance has acknowledged. Findings that assign ownership to a team or department without a named individual consistently age in registers without resolution, as the IIA’s guidance on remediation accountability makes clear.