Building AI without auditing readiness is like deploying smart contracts without reviewing assumptions: you’ll ship something impressive… and brittle. An AI readiness audit is a short, disciplined assessment that tells you whether you’re prepared to build (or buy) AI systems, what risks you’re taking on, and what to fix first.
Below is a field-tested checklist we use in AI consulting—biased toward shipping real systems (not slide decks) and avoiding expensive rewrites.
1) Business clarity: the “decision” you’re improving
Most AI projects fail because they start with a model, not a business decision.
Audit questions:
- What decision will AI improve (approve/deny, route, recommend, forecast, detect)?
- Who owns that decision today, and how is success measured?
- What’s the baseline performance (time, cost, accuracy, revenue impact)?
- What’s the action path—what happens when AI is confident vs uncertain?
Practical insight: if you can’t write a one-page “decision spec” (inputs, outputs, constraints, KPIs, failure modes), you’re not ready to build.
2) Use-case selection: feasibility vs value
Not all “good ideas” are AI-ready. You want high business value and near-term feasibility.
Audit questions:
- Is there enough historical signal to learn from or rules to automate?
- Can you tolerate occasional errors? If not, can you design human review?
- Will the model face distribution shift (new products, new markets, fraud adapts)?
- Is the task better solved with analytics or rules?
Example: A support chatbot is often pitched first, but a better early win is “agent assist” that drafts replies and cites sources—lower risk, faster adoption.
3) Data inventory: what you have, where it lives, who owns it
AI is downstream of data plumbing. If you don’t know your data assets, you can’t estimate cost or timeline.
Audit questions:
- What are the authoritative systems of record (CRM, ticketing, ERP, on-chain indexer, data warehouse)?
- What data is structured vs unstructured (PDFs, chats, calls, docs)?
- Who is the data owner (the person who can approve access and definitions)?
- What joins are required to create training/evaluation datasets?
Deliverable to aim for: a simple data map (tables/sources, update frequency, access method, PII fields).
4) Data quality: completeness, correctness, and label reality
Teams routinely underestimate label problems. “We have data” is not the same as “we have learnable targets.”
Audit questions:
- Are key fields missing or inconsistently populated across teams/time?
- Are labels objective (e.g., chargeback) or subjective (e.g., “good lead”)?
- Are there feedback loops (model influences labels)?
- How noisy is the ground truth and what’s the inter-annotator agreement?
Practical insight: for many business workflows, you should start with an evaluation dataset before training anything—100–1,000 carefully curated examples can reveal whether the task is even well-defined.
5) Privacy & data handling: PII, retention, and consent
If you’re moving fast with AI, you’re also moving fast toward privacy incidents unless you audit controls.
Audit questions:
- What PII/PHI/PCI exists in your corpora (tickets, emails, call transcripts)?
- Are you allowed to use it for model training or vendor processing?
- What are retention rules and deletion obligations?
- Do you need redaction, tokenization, or field-level access controls?
Concrete check: verify you can produce a “data processing register” for your AI pipeline (sources → transformations → storage → vendors).
6) Security posture: threat model your AI system
AI expands your attack surface: prompt injection, data exfiltration, model misuse, and insecure tool access.
Audit questions:
- Will the model have tool access (email, database, wallet, admin actions)?
- How do you prevent prompt injection from untrusted inputs (web pages, user uploads)?
- Do you log prompts/responses—and are logs sensitive?
- Do you have secrets management and least-privilege for connectors?
Opinionated take: if your LLM can call tools, you need an explicit policy layer (allowlists, schema validation, rate limits) and human approval for high-impact actions.
7) Governance & accountability: who signs off when AI is wrong
If AI makes or recommends decisions, someone must own outcomes.
Audit questions:
- Who is the accountable owner (product, operations, compliance)?
- What decisions require human-in-the-loop vs human-on-the-loop?
- What is your incident process (rollback, disable features, notify users)?
- What documentation is required (model cards, system cards, audit trails)?
In regulated contexts (finance, healthcare), plan for auditable rationales, not just outputs.
8) Evaluation: metrics that reflect reality (not vanity)
Accuracy is rarely the metric that matters. You need offline evaluation + online monitoring tied to business KPIs.
Audit questions:
- What are your offline metrics (precision/recall, calibration, hallucination rate, retrieval hit rate)?
- What are your online metrics (conversion lift, handle time reduction, defect rate)?
- What’s the cost metric (tokens, latency, human review time)?
- What’s the acceptance threshold to launch—and the rollback threshold?
Example: For RAG (retrieval-augmented generation), track: citation coverage, groundedness, refusal correctness, and “answer usefulness” scored by reviewers.
9) Architecture readiness: build vs buy vs blend
Most companies should not start by training models. Start with capabilities: retrieval, workflow automation, and strong evaluation.
Audit questions:
- Can you meet latency and uptime requirements with API-based models?
- Do you need on-prem/VPC deployment due to data constraints?
- Do you have a plan for vendor lock-in (abstraction layer, model fallback)?
- What is the integration surface (Slack, CRM, dashboards, internal tools)?
Practical approach: design for model pluralism—swap models as quality/cost changes.
10) Team & operating model: who runs this after launch
AI is not “ship and forget.” Models drift, costs creep, prompts regress, and data changes.
Audit questions:
- Who owns prompt/version changes and approvals?
- Do you have MLOps/LLMOps capabilities (CI for prompts, eval gates, monitoring)?
- Can support/ops teams report issues with reproducible traces?
- Do you have annotation/review capacity to continuously improve?
If nobody owns post-launch operations, your AI feature becomes a permanent fire drill.
11) Cost & ROI: unit economics before ambition
Compute is only part of the bill. Human review, data work, security, and tooling can dominate.
Audit questions:
- What’s your expected cost per task (tokens + retrieval + infra + review)?
- How many tasks/month, and what’s the marginal value per task?
- Where is the budget held (central AI, product team, ops)?
- What’s the cheapest experiment that can prove value in 2–4 weeks?
Rule of thumb: require a path to a measurable KPI delta within one quarter, or treat it as R&D—not product.
12) Legal & compliance: contracts, IP, and content risk
Using vendors and user data introduces IP and compliance pitfalls.
Audit questions:
- Do vendor terms allow training on your data? Are there opt-outs?
- How do you handle copyrighted content in training or retrieval?
- What are disclosure requirements (users informed they’re interacting with AI)?
- Do you need to meet SOC 2, ISO 27001, GDPR, or industry-specific rules?
If you’re in Web3: also consider wallet risk, transaction signing flows, and whether AI outputs could be construed as financial advice.
What a good audit produces (deliverables)
A readiness audit should end with concrete artifacts:
- A ranked backlog of AI use cases with effort/value estimates
- A data map and access plan (including privacy classification)
- A target architecture sketch (RAG, agents, analytics, or hybrids)
- An evaluation plan (datasets, metrics, launch/rollback criteria)
- A security and governance plan (tool permissions, logging, incident response)
- A 30/60/90-day execution plan with owners
Conclusion: audit first, build second
AI rewards teams that are systematic. The fastest way to ship is to reduce unknowns: define the decision, verify data reality, design evaluation, and lock down security and ownership. Do the audit, and your build becomes engineering—not gambling.
If you’re deciding where to start: pick one high-value workflow, run a two-week readiness audit, and only then commit to a prototype. You’ll save months, avoid compliance surprises, and end up with an AI system that survives contact with production.