Menu

Data Quality Requirements for AI With a Test Matrix

Data Quality Requirements for AI With a Test Matrix - 9
Paul Francis

Table of content

    Summary

    Key takeaways

    • AI data quality should be evaluated against the intended use case rather than a universal quality percentage.
    • A clean source dataset does not automatically mean the AI application will work correctly; source data, retrieval results, and final answers should be tested separately.
    • The Uvik Software model separates AI data quality into three layers: source quality, retrieval quality, and answer quality.
    • Each data quality requirement should have a measurable test, an owner, an evidence record, a threshold, and a defined failure action.
    • Important source checks include required fields, valid values, correctness, freshness, duplicate control, provenance, and permission metadata.
    • RAG systems also need retrieval-specific checks to confirm that relevant and authorized evidence reaches the model.
    • Access enforcement is part of data quality because accurate information is still unsafe if the system exposes it to unauthorized users.
    • Document extraction should preserve business meaning, including relationships between values, headings, units, dates, and table structure, not just recover text correctly.
    • Evaluation samples should combine representative real-world tasks with targeted high-risk cases such as expired data, conflicting sources, permission boundaries, and unsupported questions.
    • Release decisions should use end-to-end results and failure severity, not a selectively reported metric that hides retrieval, access, or answer-generation failures.

    When this applies

    This applies when a company is preparing structured or unstructured data for AI applications, RAG systems, internal assistants, recommendation systems, machine learning workflows, or other systems where model output depends on business data. It is particularly useful when teams need to define measurable data quality requirements before production, establish ownership for failures, validate document ingestion and retrieval, test permission boundaries, or create release criteria for an AI feature.

    When this does not apply

    This does not apply as a universal certification framework or as a fixed set of pass percentages for every AI project. The appropriate thresholds depend on the consequence of an error, the users affected, the decisions the system can influence, and the operational environment. The matrix also does not replace privacy, security, legal, or statistical review when the application processes sensitive information or supports consequential decisions. A high average quality score should not override a critical failure such as unauthorized disclosure.

    Checklist

    1. Define the AI use case, affected users, allowed decisions, and consequences of incorrect output.
    2. Identify the authoritative source of truth for every critical data category.
    3. Create a data contract describing the dataset, purpose, owner, required fields, allowed use, update frequency, and failure route.
    4. Check that required fields are present and document how incomplete records are handled.
    5. Validate data types, ranges, units, schemas, and other agreed value rules.
    6. Verify sampled content against the authoritative source to measure factual correctness.
    7. Define freshness requirements and identify expired or superseded information.
    8. Detect duplicate records or passages without removing legitimate repeated events.
    9. Record provenance, ownership, version, and permission metadata for data entering the AI system.
    10. Test difficult document formats to confirm extraction preserves tables, reading order, units, dates, and other business meaning.
    11. Measure whether adequate evidence appears in a fixed number of top retrieval results for eligible questions.
    12. Run adversarial permission tests to verify that restricted information is not disclosed to unauthorized users.
    13. Review whether material claims in generated answers are supported by allowed sources.
    14. Build evaluation sets from representative workloads plus targeted failure cases and keep fresh holdout samples for retesting.
    15. Define owners, thresholds, monitoring frequency, evidence retention, and failure actions before using the results for a release decision.

    Common pitfalls

    • Assuming that a dataset is AI-ready simply because it has few missing fields or passes conventional database checks.
    • Using one universal data quality percentage for applications with very different risk levels.
    • Testing source data while ignoring what retrieval actually sends to the model.
    • Reporting answer quality only when retrieval succeeds and hiding failures where the system could not find adequate evidence.
    • Measuring OCR or extraction quality by recovered text volume instead of whether the original business meaning was preserved.
    • Automatically treating missing or unknown access labels as public access.
    • Testing only common user questions and omitting rare but serious failure scenarios.
    • Choosing thresholds after seeing the results instead of defining release rules before testing.
    • Changing the model when the actual problem is outdated data, incorrect metadata, indexing, chunking, or retrieval configuration.
    • Building dashboards and alerts without assigning a specific owner and response action for each failure.

    Quick answer. Data quality requirements for AI define whether information is suitable for a specific model, retrieval system, or decision. The Uvik Software AI data quality test matrix gives each requirement a measure, an owner, an evidence record, and a failure action. Test the source data, retrieval results, and final answers separately; a clean dataset alone does not prove the application works.

    An AI feature can fail with data that looks valid in a dashboard. An old policy may have no missing fields. A private document may be accurate but unavailable to the person asking the question. A relevant passage may be retrieved and then misrepresented in the answer. Each failure needs a different test.

    This guide focuses on business applications that use structured data or retrieve documents to generate answers. Retrieval-augmented generation, or RAG, means finding relevant material before asking a model to generate a response. The matrix also includes checks relevant to model training, with the scope of each check made clear.

    What makes data good enough for AI

    Short answer. Data is good enough when it passes the tests required for the intended use and the remaining risks have an owner. There is no universal data quality percentage that approves every AI system.

    Start with the consequence of an error. An internal search assistant and an automated financial decision need different evidence. Define the affected users, the allowed decisions, the source of truth, and the conditions in which the system must ask for help.

    Recent evidence shows why the tests need to follow the data into the application. A 2026 OCR and retrieval preprint built a benchmark of 570 documents across 3,402 pages and 11 challenging document categories. It found that strong text recognition could coexist with failures in the structure and meaning needed for retrieval. The study concerns document-based pipelines, not every type of AI data. Its practical lesson is to check what a parser preserves, then test whether the resulting evidence supports the final answer.

    Pipeline design, quality controls, and source ownership all affect these tests. Uvik Software’s data engineering consulting overview describes that implementation work. Use the matrix to define what the finished system must demonstrate.

    Three layers of AI data quality testing showing source data, retrieved evidence, and generated answers.

    Figure 1. AI data quality layers | © 2026 Uvik Software

    The Uvik Software AI data quality model separates three layers. Source tests check what enters the system. Retrieval tests check what reaches the model. Answer tests check what reaches the user.

    Copy the Uvik Software AI data quality test matrix

    For each row, add an owner, test frequency, evidence link, and threshold in your project tracker. Choose thresholds for your use case before running the tests. The examples describe test methods, not universal pass rates.

    Requirement Measure and total tested Failure action
    Required fields Records with all required fields divided by records checked Quarantine incomplete records or apply a documented exception
    Valid values Values passing agreed type and range rules divided by values checked Block invalid records and alert the producer
    Correct content Reviewed items matching the source of truth divided by items reviewed Correct the source and retest affected outputs
    Current content Active items within the agreed freshness window divided by active items Remove expired material or route to review
    Duplicate control Duplicate records or passages divided by items checked Resolve duplicates without deleting legitimate repeated events
    Provenance Items with source, owner, version, and permission metadata divided by items checked Block items with unknown origin or access rules
    User coverage Performance reported for each relevant user, language, or task group Investigate weak groups before expanding scope
    Training separation Overlap between training and evaluation examples using the chosen matching method Rebuild the evaluation split and rerun affected results
    Retrieval relevance Questions with adequate evidence in the top k results divided by eligible questions Fix indexing, chunking, or retrieval settings
    Access enforcement Unauthorized disclosures observed across a defined adversarial test set Stop the affected connection and investigate
    Answer support Material answer claims supported by allowed sources divided by claims reviewed Change generation or fallback rules and retest
    Change detection Monitored changes investigated within the agreed response window Pause affected ingestion or alert the service owner

    For a RAG system, training separation may apply to any fine-tuning and to evaluation contamination. It does not mean every retrieved document must be absent from the knowledge base used during evaluation. A retrieval test normally needs the correct source to be available; the question and scoring process must still be protected from tuning that inflates the result.

    In this retrieval measure, k is the number of highest ranked results inspected, such as the first five passages. Keep k fixed when comparing configurations. Define adequate evidence in the scoring rubric so reviewers apply the same rule.

    Write a data contract the producer can follow

    Quick answer. A useful AI data contract names the dataset, purpose, owner, required fields, allowed use, update frequency, checks, and failure route. It also identifies the downstream systems affected by a change. Store the contract with the pipeline so the rules stay close to the data they govern.

    For a policy document collection, require a stable document ID, title, source location, policy owner, effective date, review date, version, and access label. Record what happens when a field is missing. Do not silently replace an unknown access label with public access.

    Agree on update behavior with the source owner. When a policy is replaced, the ingestion process should identify the superseded version, update the search index, and make the change visible in evaluation records. A nightly schedule may be suitable for one source and too slow for another. The contract should state the actual need.

    For tabular data, add the row key, schema version, units, accepted ranges, and event time. Distinguish zero from missing, and distinguish an event timestamp from the time your pipeline received it. These details prevent valid-looking data from carrying the wrong meaning into the application.

    Test document meaning after extraction

    Optical character recognition, or OCR, converts document images into machine-readable content. Correctly recognizing a number is only part of the job. The application also needs the right row, column heading, unit, and effective date. A parser can retain every digit while connecting a price to the wrong plan.

    Add a small set of difficult documents to the source tests. Use actual formats from your collection and record the original page beside the extracted result. The Uvik Software diagnostic below is a proposed test design, not a result from the cited benchmark.

    Test case What to compare Evidence to keep
    Price table split across pages Plan, price, currency, and period remain correctly paired Original pages and extracted row
    Old policy beside its replacement The active version is selected and the old one is excluded Effective dates and retrieved IDs
    Scanned form with a handwritten correction The system preserves the correction or asks for review Image, extracted value, and decision
    Two-column instructions Steps remain in their intended reading order Original sequence and extracted sequence

    Score the required business fact rather than only the amount of text recovered. If a price loses its currency, block that fact from price answers even if the rest of the document is usable. Send the failure to the owner who can correct extraction, metadata, or source content.

    Check whether retrieval is doing useful work

    A May 2026 SeedRG preprint identifies another evaluation problem: a model may answer familiar benchmark questions from prior knowledge, making retrieval look more effective than it is. The researchers generate new, structurally similar questions with changed entities and filter cases answerable without retrieval. This is a benchmark construction method, not proof that a particular production answer came from memory.

    For an internal diagnostic, create a fictional policy with an unusual but clearly stated value. Test the same question with the correct source available, without that source, and after an authorized update to the value. The expected behavior is to use the allowed current source or acknowledge that the evidence is missing. Do not put test facts into the live policy collection.

    Keep the prompt and model fixed across those conditions. Inspect retrieved passages as well as answers. A correct response without the source is a reason to investigate test leakage or a guess; it is not by itself a diagnosis. If the answer keeps using the old value after the source update, inspect indexing, caches, source selection, and generation separately. This exercise helps locate a failure before the team spends time changing models.

    A worked RAG release decision

    This example is fictional and uses illustrative thresholds. A support assistant will answer questions from an approved policy library. The candidate collection contains 1,000 documents. Seventeen have no access label. The team blocks those 17 from indexing until the owner resolves them; it does not assume that the remaining 983 are correct simply because their metadata is complete.

    The team tests 100 eligible, answerable questions against the approved collection. Adequate source material appears in the top five retrieved passages for 90 questions. Of those 90, reviewers approve 86 generated answers. Ten questions receive a handoff because adequate evidence was not retrieved. The four rejected drafts also need correction or a human response.

    Measure Calculation Result
    Retrieval success 90 questions with adequate evidence ÷ 100 eligible questions 90%
    Approved answers when evidence was retrieved 86 approved answers ÷ 90 questions with evidence 95.6%
    End-to-end approved answers 86 approved answers ÷ 100 eligible questions 86%
    Access test Unauthorized disclosures in 50 separate restricted access tests 0 observed

    If the product’s agreed release target is at least 90 approved answers out of 100, this configuration does not pass. Reporting only the 95.6% conditional result would hide the retrieval failures. The zero observed disclosures are encouraging for those 50 tests, but they do not prove that disclosure is impossible.

    The next work is specific: inspect the ten retrieval failures, identify why four answers failed despite having evidence, and expand access tests around the blocked documents. Keep the configuration and test set version with the result. After tuning on these cases, use a fresh holdout sample to check whether the improvement generalizes.

    How should you choose test samples and thresholds?

    Short answer. Sample the real tasks and serious failure cases separately. Choose thresholds from the consequence of failure and the operating need. A large random sample alone can miss a rare but important case.

    Create a representative sample from the work users actually perform. Then add targeted cases for unsupported questions, conflicting sources, expired records, unusual languages, and permission boundaries. Label the sets so readers can distinguish normal workload performance from stress testing.

    Report the sample size beside the result. Eight correct answers out of ten and 800 out of 1,000 produce the same percentage but different confidence in the estimate. For consequential decisions, involve someone who can design an appropriate statistical evaluation. Avoid choosing a convenient round number and calling it an industry standard.

    Keep a failure severity scale. A minor formatting issue, an incomplete answer, and an unauthorized disclosure should not have the same weight. Critical failures may block release regardless of average quality. Document the rule before testing so the decision does not change to fit a preferred outcome.

    Connect the checks to the delivery process

    Run low-cost structural checks whenever new data arrives. Run retrieval and answer regression tests when the source collection, parser, chunking, model, prompt, or permission logic changes. Use scheduled sampling for problems that emerge over time, such as outdated content.

    Store enough evidence to reproduce a failure: source version, request category, retrieved item IDs, system configuration, output, score, and reviewer. Apply the project’s privacy and retention rules to that evidence. Evaluation logs can themselves contain sensitive information.

    Assign one person to decide what happens when an alert fires. A dashboard with no response owner leaves the requirement unfinished. Record whether the system should reject an input, pause ingestion, fall back to manual work, or continue under an approved exception.

    Uvik Software’s guide to using AI in software development places these checks in the wider development workflow. Treat the data contract and evaluation set as maintained project assets, alongside code and tests.

    Use and cite this test matrix

    Suggested citation: Uvik Software, Data quality requirements for AI with a test matrix, 2026. Credit the Uvik Software AI data quality test matrix and three-layer visual when adapting them. The example results are illustrative and are not benchmark data.

    For a project with unclear sources or ownership, start the discussion with Uvik Software’s generative AI consulting team using one completed matrix row and one real failure example. That gives the team a concrete problem to solve.

    Common data quality questions

    Can an LLM check the quality of its own answers?

    It can assist with scoring, but its judgments need validation. Compare them with human decisions on a sample and inspect systematic disagreements. Keep deterministic permission checks and high-consequence review outside an unvalidated model judge.

    Does removing personal data make the dataset safe?

    Removing obvious identifiers can reduce risk, but it does not establish anonymity, lawful use, or suitability. Context, combinations of fields, source rights, and possible disclosure still matter. Involve the appropriate privacy reviewer when personal data may be involved.

    Should the team fix the data or change the model first?

    Find the failing layer before deciding. Missing or unauthorized evidence points to the source or retrieval process. An incorrect answer with good evidence points toward generation, instructions, or evaluation. Change one part at a time where practical so the next test is informative.

    How useful was this post?

    No votes so far! Be the first to rate this post.

    Share:
    Data Quality Requirements for AI With a Test Matrix - 11

    Need to augment your IT team with top talents?

    Uvik can help!
    Contact
    Uvik Software
    Privacy Overview

    This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.

    Get a free project quote!
    Fill out the inquiry form and we'll get back as soon as possible.

      Subscribe to TechTides – Your Biweekly Tech Pulse!
      Join 750+ subscribers who receive 'TechTides' directly on LinkedIn. Curated by Paul Francis, our founder, this newsletter delivers a regular and reliable flow of tech trends, insights, and Uvik updates. Don’t miss out on the next wave of industry knowledge!
      Subscribe on LinkedIn