Blog · Essay · AI-ready data
AI-ready data: The audit checklist every mid‑market CEO must run before funding an AI project

When I was consulting for a Singapore‑based fintech that wanted to embed a predictive fraud model, the first thing that broke was not the algorithm but the data pipeline. The team had assumed that “we have data, so we can train” and spent six months chasing missing timestamps, duplicate customer IDs, and inconsistent currency fields. The project stalled, the budget ballooned, and the board started questioning the AI ambition altogether. This is the kind of story that is now being discussed in the media – see the recent piece “AI Transformation Is A CEO Decision” on *Chief Executive* (source). The headline captures the strategic weight of AI, but the real work starts with the data foundation.
In this essay I walk through exactly what I inspect when a client says “we want to fund an AI project tomorrow”. I break the audit into three layers – source integrity, schema consistency, and governance readiness – and give you the concrete questions you can ask on a Monday morning. The goal is not to produce a buzz‑word checklist, but to surface the hidden gaps that cause budget overruns and missed deadlines.
1. What broke in the real world?
The fintech client I mentioned had three data‑related failures that cascaded into a full‑scale project halt:
- Missing timestamps – Transaction logs from two legacy systems used different time zones and formats. The model needed a single, monotonic timeline to calculate velocity features, but the data engineer spent weeks writing conversion scripts that still produced gaps.
- Duplicate identifiers – Customer IDs were reused after account closure. The deduplication logic was applied after model training, so the model learned spurious patterns that later manifested as false positives.
- Inconsistent currency fields – Some records stored amounts in SGD, others in USD, without a flag. The feature engineering step assumed a single currency, leading to wildly inaccurate risk scores.
The root cause was not a lack of talent or compute; it was an assumption that the data lake was “AI‑ready”. In reality, the data was a patchwork of legacy exports, manual uploads, and ad‑hoc APIs. The audit I performed uncovered these issues in a single day, allowing the team to re‑scope the project and allocate budget for data remediation instead of endless model re‑training.
2. The three layers of AI‑ready data
2.1 Source integrity
Source integrity is about trusting the raw feed. If the upstream system drops records, corrupts files, or delivers data late, any downstream AI will inherit that unreliability.
2.2 Schema consistency
Even if the source is reliable, the shape of the data must be stable. Column names, data types, and enumerated values should not change without version control. A single schema drift can break feature pipelines that have been in production for weeks.
2.3 Governance readiness
Finally, governance covers data lineage, access controls, and compliance tags. CEOs need to know who can edit the data, how it is audited, and whether it meets regulatory requirements (e.g., PDPA in Singapore).
3. The CEO audit checklist – questions to ask this week
Below are the exact questions I put on a one‑page spreadsheet and walk through with the executive team. Answering “yes” to all of them means you have a solid foundation; a “no” signals a risk that must be mitigated before you sign a contract.
3.1 Source integrity questions
- Do we have end‑to‑end SLAs for each data feed?
- What is the maximum tolerated latency?
- Who is accountable if the SLA is breached?
- Are there automated health checks that alert us to missing or corrupted files?
- Frequency of checks (hourly, daily)?
- Alert channel (Slack, email, ticketing)?
- Do we retain raw source files for at least the model’s training horizon?
- Retention period aligned with regulatory requirements?
- Is there a documented fallback source if the primary feed fails?
- Example: batch dump from the ERP system.
3.2 Schema consistency questions
- Is the data schema version‑controlled in a repository (Git, Bitbucket)?
- Do we tag releases when a schema changes?
- Do we run schema validation as part of the ingestion pipeline?
- Reject or quarantine records that violate the contract?
- Are enumerated values (e.g., transaction types) centrally managed?
- A change in an enum should trigger a change request.
- Do downstream feature pipelines have automated tests that fail on schema drift?
- Unit tests, integration tests, and data‑drift monitors.
3.3 Governance readiness questions
- Do we have a data lineage map that shows origin → transformation → consumption?
- Can we trace a model prediction back to the exact source row?
- Are access permissions reviewed quarterly?
- Principle of least privilege enforced on data stores.
- Is there a documented process for data privacy impact assessments (DPIA)?
- Required for any personal data used in AI.
- Do we have a data retention and deletion policy that aligns with PDPA and ISO 27001?
- Automated purging for data older than the policy horizon.
If you can answer “yes” to at least 90 % of these questions, you are in a position to fund an AI project with confidence. Anything less is a red flag that should trigger a remediation sprint before any dollars are committed.
4. Monday action plan – what to do right now
- Gather the data owners – Schedule a 30‑minute stand‑up with the heads of the three systems that feed your AI use case. Bring the checklist and ask them to fill it out live.
- Run a quick health‑check script – Use a simple Python script that verifies record counts, timestamp continuity, and duplicate keys for the last 30 days. Document any anomalies.
- Create a schema snapshot – Export the current schema to a JSON file and commit it to a version‑control repo. Tag it as `baseline‑AI‑ready`.
- Assign a data‑governance champion – This could be a senior analyst or the CDO. Their first task is to map lineage for the data set you intend to use.
- Report back to the board – Prepare a one‑page risk register that lists the audit findings, remediation actions, and estimated effort. This turns a vague “AI budget” request into a concrete, accountable plan.
By the end of the week you will have a clear picture of whether the data can support a model, and you will have a realistic budget line for any data‑cleaning work that is required.
5. Beyond the checklist – building a culture of AI‑ready data
The audit is a one‑time gate, but the real advantage comes from embedding data readiness into your operating model. Here are two practices that have worked for the companies I’ve helped:
- Data‑first sprint cycles – Treat data remediation as a deliverable in every agile sprint, not an after‑thought. This keeps the backlog visible and ensures that new data sources are vetted before they land in production.
- Cross‑functional data guilds – Form a lightweight guild that meets monthly, composed of engineers, product managers, and compliance officers. The guild reviews schema changes, shares lessons learned, and updates the checklist as the business evolves.
When data becomes a first‑class citizen, AI projects stop being “big bets” and become incremental improvements that deliver ROI on a predictable cadence.
FAQ
What if my organization already has a data lake?
A data lake is a storage layer, not a guarantee of quality. Apply the same audit questions to the lake’s ingestion pipelines and metadata catalog. If the lake lacks version‑controlled schemas, you need to add that layer before proceeding.
How much budget should I allocate for data remediation?
It varies, but a rule of thumb is 10‑20 % of the total AI project budget for data work. The audit helps you quantify the exact effort, turning a vague estimate into a line item.
Can I outsource the audit to a vendor?
You can, but the CEO should remain the final sign‑off. Vendors may miss organization‑specific compliance nuances, especially around PDPA and ISO 27001. Use the audit as a negotiation tool, not a delegation.
How often should I repeat the audit?
At a minimum quarterly for high‑impact data sets, and anytime you add a new source or change a critical schema.
Does AI‑ready data mean I need a data lake?
No. Mid‑market companies often succeed with a well‑governed relational warehouse or a set of curated CSV feeds. The key is trustworthy, versioned, and governed data, not the technology stack.
If you’re ready to run this audit with a seasoned operator by your side, I’m happy to walk through the checklist in a brief discovery call. You can book a slot at my calendar or explore more of my work on the homepage.
*Related reading:*
In the market
Headlines this post is responding to — not invented stats.