Back to The Signal

    The Signal

    AI Compliance Checklist: 12 Steps Auditors Can Verify

    Build an AI compliance checklist of 12 steps that connects security controls to audit evidence, ISO/IEC 42001, NIST AI RMF, and EU AI Act duties.

    11 October 2026 · 11 min read

    AI Compliance Checklist: 12 Steps Auditors Can Verify

    An AI compliance checklist must do five things before any auditor accepts it: discover every AI system and agent in use, map the data each one touches, enforce access and approval controls, monitor behavior continuously, and produce evidence mapped to recognized frameworks. Standards bodies including ISO/IEC 42001 and the NIST AI Risk Management Framework give structure to this work. Start with inventory and evidence capture today, not after the next audit notice arrives.


    TL;DR:

    • Start discovery with browser extension inventories, SaaS connector audits, endpoint scans, and API traffic; record each agent’s owner, provider, version, permissions, and approval status.
    • Map each agent’s data path and classify sensitive information, then prioritize agents with regulated data or sensitive records ahead of those handling routine support tickets.
    • Give every agent a separate identity, short lived credentials, and least privilege access; require human approval before writes, deletions, bulk exports, or high risk decisions.
    • Store tamper resistant logs with the agent, action, data touched, request and response, timestamp, and correlation ID; retain them for applicable regulatory periods.
    • Map controls to ISO/IEC 42001, NIST AI RMF, and EU AI Act requirements; transparency rules begin August 2, 2026, and Annex III high risk duties start December 2, 2027.

    Lexic AI
    lexic.ai
    Audit AI Conversations With Clarity
    Lexic Compass independently audits AI agents from any vendor, surfacing blind spots in customer experience, security, and technical issues.
    Explore Lexic Compass

    Table of Contents

    The 12-step checklist at a glance

    A compliance checklist only works when it turns into tasks someone owns. These twelve steps give security and compliance teams a build order they can translate directly into tickets.

    1. Discover and inventory every AI tool and agent in use, including shadow AI.
    2. Map the data each AI system can access.
    3. Classify data sensitivity across every mapped connection.
    4. Scope access to each agent using least-privilege rules.
    5. Require human approval for high-risk or irreversible actions.
    6. Run third-party risk assessments on every AI vendor.
    7. Apply in-flight remediation, such as redaction and masking, before data reaches a model.
    8. Extend coverage to every surface: browser extensions, endpoints, and MCP servers.
    9. Monitor agent behavior and log every action continuously.
    10. Retain logs and evidence according to regulatory windows.
    11. Map each control to ISO/IEC 42001, NIST AI RMF, and applicable regulation.
    12. Operationalize the checklist as a continuous control cycle, with a fixed review cadence.

    Each step below expands on the mechanics, the artifacts to produce, and the common points where teams stall.

    Step 1: discover and inventory every AI tool and agent

    You cannot govern what you cannot see, and most organizations underestimate how many AI tools are already running inside their environment. Discovery work combines several data sources: SIEM alerts, endpoint scans, SaaS connector audits, MCP and agent server logs, and direct surveys of teams that build or buy automation.

    Shadow AI is the most common blind spot practitioners report, often arriving through browser extensions, personal accounts, or embedded chat widgets that nobody formally approved, according to KPMG’s analysis of ISO/IEC 42001 adoption. Finding these tools takes the same discipline as any shadow IT hunt: checking SSO logs for unapproved sign-ins, scanning browser extension inventories across managed devices, and reviewing API gateway traffic for unrecognized model endpoints.

    Once discovered, each agent needs a consistent inventory record:

    • Agent name and the owning team or business unit.
    • Model, provider, and version in use.
    • Data scopes the agent can reach and the access tokens tied to it.
    • Approval status and the date of its last risk test.

    Pro Tip: Run the browser extension and SaaS connector scan first. It typically surfaces more unsanctioned AI than any formal survey.

    Step 2: map and classify the data each AI can reach

    Inventory tells you an agent exists. Data mapping tells you what it can actually touch, and that distinction drives every prioritization decision that follows.

    Start by linking each inventoried agent to the data sources it connects to, then diagram the flow so a reviewer can trace a straight line from input to model to output. This exercise routinely reveals agents with broader reach than anyone intended, a chatbot plugin with database read access far beyond its stated purpose, for example.

    Classify what you find into clear buckets:

    • Personally identifiable information (PII): names, addresses, contact details.
    • Protected health information (PHI) and payment card data (PCI).
    • Secrets and credentials: API keys, passwords, tokens.
    • Source code and regulated records: financial filings, legal documents, audit logs.

    Prioritize remediation by volume and sensitivity. An agent touching a small volume of low-sensitivity support tickets waits behind one that can query a customer PHI database. The EDPB’s AI auditing checklist treats this kind of traceability as a baseline expectation, not an advanced practice.

    Step 3: scope and enforce identity-based access controls and approval gates

    Access control for AI agents works the same way it does for human users, except agents tend to accumulate permissions faster and get reviewed less often. Each agent should carry its own identity, scoped to the minimum access it needs, with short-lived tokens rather than standing credentials.

    Role-based policies keep this manageable at scale. A customer support agent gets read access to tickets and write access to a response field; it does not get database export rights by default.

    Approval gates matter most for actions that are hard to reverse:

    • Require human sign-off before any agent executes a write, delete, or bulk export.
    • Flag high-risk agents, those touching regulated data or customer-facing decisions, for mandatory human review of their outputs.
    • Set alerts for permission drift, when an agent’s access expands beyond its original scope.

    Pro Tip: Review agent permissions on the same cadence as employee offboarding reviews. Drift accumulates fastest right after integration changes.

    A short remediation playbook helps here too: revoke the excess scope, log the correction, and note the root cause so the same drift does not recur with the next agent deployment.

    Step 4: manage vendor and third-party AI risk through the full lifecycle

    Every AI vendor you add is a new party with access to your data, and third-party risk management for AI needs to go further than a one-time security questionnaire.

    Assessment should capture model and provider metadata, named subprocessors, the vendor’s security posture, any SOC or HIPAA claims, and a record of past incidents. Contracts need to lock in specific commitments:

    • Logging and audit access provisions that let you pull evidence on demand.
    • Breach notification timelines specific to AI-related incidents.
    • Explicit data handling and retention commitments tied to your regulatory obligations.

    Monitoring does not stop once the contract is signed. Compare the data flows you actually observe against what the vendor claims it processes, and flag mismatches early. Maintain an offboarding checklist too, so that when a vendor relationship ends, access revocation and data deletion are confirmed rather than assumed.

    Step 5: monitor, log, and produce auditor-ready evidence

    Policy documents do not satisfy auditors. Evidence does, and evidence has to be structured before the audit request arrives, not assembled afterward.

    Each log entry needs enough detail to reconstruct what happened: the acting agent, the action taken, the data touched, a snapshot of the request and response, a timestamp, and a correlation ID that ties the event to related records.

    Retention periods should map to the regulatory windows that apply to your data, and the storage itself needs to resist tampering. Write-once-read-many (WORM) storage and hashed indices are standard techniques for keeping logs defensible under scrutiny.

    Checklists built for AI auditing consistently emphasize traceability and documentation as the core requirement, according to the EDPB’s AI auditing checklist, which means an evidence package needs more than raw logs.

    • Indexed event records tied to specific agents and time windows.
    • Remediation tickets showing how past issues were resolved.
    • Policy-change records documenting when and why controls shifted.
    • A set of representative conversation or transaction samples.

    Step 6: map controls to ISO/IEC 42001, NIST AI RMF, and EU AI Act timelines

    Mapping the same control to multiple frameworks saves real work, and it starts with knowing what each framework actually asks for.

    ISO/IEC 42001 is a certifiable AI management system standard built on a Plan-Do-Check-Act cycle: you document a policy, implement it, check its effectiveness, and adjust. The NIST AI RMF organizes the same kind of work into four functions, GOVERN, MAP, MEASURE, and MANAGE, with MEASURE covering the testing and metrics work that produces hard evidence.

    The EU AI Act’s implementation timeline adds binding dates on top of these voluntary frameworks. Prohibited practices and AI literacy obligations applied from February 2, 2025. Governance rules and general-purpose AI model obligations followed on August 2, 2025. Transparency obligations under Article 50 apply from August 2, 2026, and high-risk system obligations under Annex III begin December 2, 2027.

    EU AI Act obligations across four dates

    A single access-control log can satisfy an ISO 42001 audit clause, a NIST MEASURE outcome, and an EU AI Act documentation requirement at once, provided it is written clearly and labeled against each framework it supports.

    Step 7: assign governance roles, classify risk, and set a review cadence

    Governance only holds up when responsibility is assigned before an incident forces the question. A three-lines-of-defense model, adapted from traditional risk management, works well for AI: project teams own day-to-day use and first-line checks, a governance or quality committee sets policy and reviews exceptions, and an independent assurance function audits both.

    Risk classification should combine probability and severity rather than relying on gut judgment. An agent with a low chance of failure but severe consequences, one handling medical appointment scheduling, for instance, still warrants escalation.

    Classification triggers matter in practice:

    • Flag any agent making or materially influencing decisions about people for a Fundamental Rights Impact Assessment (FRIA) or equivalent review.
    • Escalate agents handling regulated data categories automatically, regardless of stated risk level.
    • Require human oversight sign-off for any agent classified as high-risk under your framework.

    A recommended cadence keeps this from going stale: refresh the inventory monthly, review evidence retention quarterly, and revalidate model behavior at least every six months or after any material vendor update.

    How independent audits close the evidence gap

    Internal documentation tells you what a system is supposed to do. It rarely proves what it actually did across thousands of real conversations, and that gap is where most audit evidence falls apart.

    Independent, vendor-neutral auditing addresses this by reviewing agent behavior directly, across any underlying model or provider, rather than relying on vendor-supplied summaries. Coverage across the full population of customer conversations, not a sample, matters here because escalation failures tend to hide in edge cases that spot-checks miss.

    Auditors generally value a specific evidence set:

    • Sampled transcripts that show how an agent handled sensitive or ambiguous requests.
    • Escalation success and failure rates, tracked over time rather than as a single snapshot.
    • Continuous alert logs that flag anomalies as they happen, not after a complaint arrives.

    Audited findings have shown that a large majority of AI agents reviewed failed to properly escalate issues to a human representative when they should have, underscoring why continuous, independent oversight matters more than periodic self-reporting.

    Moving from snapshots to continuous evidence

    Most teams stall at the same three points: discovery takes longer than expected, data remediation gets deprioritized once the inventory is built, and evidence collection turns into a scramble right before an audit. A checklist treated as a one-time project will always run into this pattern.

    Treat discovery and evidence capture as ongoing operations, not milestones. Build your next 90-day plan around continuous controls rather than a single compliance sprint, and the audit stops being an event you prepare for and becomes a byproduct of how you already operate.

    — Sergio Llorens

    Another option worth considering: vendor-independent continuous audits

    Building every piece of this checklist in-house takes real engineering time, and plenty of teams reach a point where continuous evidence collection outpaces what internal tooling can handle. Our audit and governance platform is designed to address that gap.

    Lexic AI

    Our Lexic Compass platform audits conversational AI agents from any vendor, independent of who built or operates them, and verifies real behavior rather than relying on vendor-reported summaries. It reviews regulatory compliance, security vulnerabilities, escalation handling, and customer disclosure across the full population of conversations, which gives compliance teams a defensible evidence trail they can hand to a regulator without depending on the vendor’s own data.

    For teams that need a faster starting point, our Flash Preview and Audit Sprint offerings on the compliance audit page produce a focused evidence package without a long engagement. If your checklist work has stalled at the evidence stage, that is the exact problem our platform solves. Explore the audit approach and see where your current coverage has gaps.

    FAQ

    What is the 30% rule for AI?

    Such wording shows up informally in different contexts, often referring to a cap on AI-generated content or automation in a given workflow, so treat any specific percentage you encounter as organization-specific guidance rather than a regulatory standard.

    What is an AI checklist?

    An AI checklist is a structured set of steps an organization follows to inventory its AI systems, assess risks, and verify that controls like access limits, monitoring, and human oversight are in place. A strong checklist also ties each control to evidence auditors can review, rather than existing only as a policy document.

    How do I prepare a compliance checklist?

    Start by discovering every AI tool and agent in use, including shadow AI that was never formally approved, then map the data each one can access and classify its sensitivity. From there, build in access controls, continuous monitoring, and logging, and map each control to applicable frameworks such as ISO/IEC 42001 or the NIST AI RMF.

    What is the 10/20/70 rule for AI?

    This is not a standard defined by major regulatory or standards bodies covering AI compliance. Where it appears, it is typically used informally to describe resource allocation in AI projects (such as splitting effort across data, algorithms, and infrastructure), so it should not be treated as a compliance or governance benchmark.

    Does an AI compliance checklist need to cover voice and call-based AI separately?

    Yes, voice and call-based AI agents carry obligations distinct from text-based systems, including call recording disclosure and consent requirements that vary by jurisdiction. A dedicated checklist for call compliance covers these voice-specific obligations in more depth than a general AI framework typically does.

    Sources

    Want an independent verdict on your AI agent?

    Back to The Signal