Menu

GDPR and AI engineering evidence checklist

GDPR and AI engineering evidence checklist - 9
Paul Francis

Table of content

    Summary

    Key takeaways

    • GDPR applies to AI engineering when the system processes personal data within the regulation’s scope, so teams need evidence of what data is used, why it is used, where it goes, who can access it, and how long it remains.
    • A privacy checklist is an engineering evidence tool, not proof of compliance and not a replacement for legal review, a DPO assessment, or a DPIA where one is required.
    • Legal decisions and engineering controls must stay connected: the approved purpose, legal basis, data categories, processing roles, and restrictions should be reflected in the actual system design.
    • AI data mapping should include the original source, preparation steps, search indexes, prompts, model providers, tools, outputs, logs, caches, and recovery copies.
    • RAG does not avoid GDPR obligations automatically because retrieved passages, prompts, outputs, embeddings, and logs may still contain or reveal personal data.
    • Removing a record from the source system does not prove deletion from indexes, caches, logs, backups, or models that were trained or fine-tuned on that data.
    • Access controls should be tested using authorized and unauthorized users so the team can prove that restricted information cannot be retrieved or inferred.
    • Vendor claims about retention, data location, subprocessors, deletion, training, and support access should be backed by current contracts or documentation rather than marketing statements alone.
    • Privacy evidence must stay current when the purpose, model, vendor, connector, user group, data source, or access model changes.
    • A strong evidence register links each privacy claim to an owner, supporting document or configuration, test result, review date, known limitations, and remediation status.

    When this applies

    This applies when an AI feature processes personal data and falls within the scope of EU GDPR. It is especially relevant for RAG assistants, internal copilots, AI agents, support tools, AI-enabled search, model integrations, and other systems that retrieve, transform, log, transmit, or generate content involving identifiable people. The framework is useful for engineering teams that need to prepare technical evidence for a DPO, privacy reviewer, legal team, controller, or processor assessment before release or during an existing compliance review. :chatgpt-content-reference{index=”0″}

    When this does not apply

    This does not establish GDPR compliance by itself and does not replace a use-case-specific legal assessment, a formal DPIA, or advice from the organization’s privacy and legal specialists. It also does not automatically cover UK GDPR or other privacy regimes, which require their own review. If an AI experiment genuinely uses no personal data and creates no realistic path for personal data to enter prompts, logs, retrieval, tools, or outputs, the full evidence process may be unnecessary, although that assumption should itself be verified rather than guessed. :chatgpt-content-reference{index=”1″}

    Checklist

    1. Document the approved use case, intended users, exclusions, and accountable decision owner.
    2. Record controller, processor, and other relevant processing-role decisions.
    3. Inventory personal data sources, categories, identifiers, and derived data used by the AI system.
    4. Map the full data path through connectors, parsers, indexes, prompts, model calls, tools, outputs, logs, caches, and backups.
    5. Record the approved legal basis and any special-category or DPIA decision supplied by the responsible privacy reviewer.
    6. Collect current vendor agreements, subprocessor details, data-location information, retention terms, and deletion behavior.
    7. Document roles, service accounts, permission filters, and the process for removing access.
    8. Test that an authorized user can retrieve permitted information and an unauthorized user cannot.
    9. Define retention and deletion rules separately for every relevant store, including indexes, caches, logs, and recovery copies.
    10. Run a synthetic deletion drill and verify what disappears, what remains, and why.
    11. Document how data-subject requests are received, verified, searched, answered, and escalated.
    12. Confirm that the user interface shows the reviewed notices and explanations for the AI feature.
    13. Document incident detection, containment, evidence preservation, escalation, and the ability to disable the feature.
    14. Define change triggers that require a new technical or privacy review, including model, vendor, connector, purpose, and user changes.
    15. Keep every checklist row linked to an owner, evidence location, status, review date, known limitation, remediation action, and retest date.

    Common pitfalls

    • Assuming a RAG system is automatically GDPR-compliant because it does not train a foundation model.
    • Treating embeddings as anonymous data simply because they are not directly human-readable.
    • Deleting a source document and assuming all copies have disappeared from indexes, caches, logs, and backups.
    • Relying on a privacy policy or architecture diagram without producing testable implementation evidence.
    • Using a successful technical test without checking whether it proves the policy or legal claim that actually matters.
    • Accepting vendor statements such as no model training or EU region as complete answers to retention, access, subprocessor, or transfer questions.
    • Failing to distinguish source-data lawfulness, system handling of the data, and whether a trained model can reveal personal information.
    • Testing access removal only in the source system and not verifying existing sessions, cached results, or downstream stores.
    • Hiding failed deletion or access tests instead of recording them as evidence with remediation and a retest date.
    • Leaving the original privacy review unchanged after adding new models, connectors, data sources, user groups, or processing purposes.

    By Uvik Software | Research checked 4 October 2026

    Quick answer. GDPR applies when an AI project processes personal data within the regulation’s scope. A useful engineering checklist records where that data goes, why it is used, who can access it, how long it remains, and how the team verifies its controls. The Uvik Software AI privacy evidence checklist helps engineers prepare that record for the responsible privacy reviewer.

    This guide focuses on EU GDPR engineering evidence for an AI feature, including a retrieval-based assistant. It is an implementation aid to use with your data protection officer or legal reviewer. Completing it does not establish compliance or replace a use case specific legal assessment. UK GDPR and other jurisdictions need their own review.

    The reviewer needs evidence for each important claim about the system: the relevant document, configuration, test result, and owner. A diagram or vendor statement alone rarely provides that whole chain.

    What should the team establish before building?

    Short answer. Define the purpose, identify the personal data and processing roles, and have the controller document the legal basis and applicable obligations with appropriate privacy and legal advice. Then translate those decisions into system behavior and evidence.

    The GDPR text sets out principles including purpose limitation, data minimization, accuracy, storage limitation, security, and accountability in Article 5. Article 6 addresses lawful bases. Article 9 adds conditions for special category data, and Article 35 concerns impact assessments for processing likely to create high risk. Which provisions apply depends on the processing, not simply on whether a vendor calls a product enterprise-grade.

    Keep legal decisions and engineering checks connected. The controller remains accountable for the processing. A legal adviser helps assess the proposed basis, while a data protection officer advises and monitors independently, as described in the European Commission guidance on DPO responsibilities. Engineering records how the approved purpose limits the inputs, connections, retention, and user actions. If the design later expands, reopen the decision rather than assuming the earlier review covers it.

    Use Uvik Software’s generative AI consulting overview to scope the technical work around systems, data, and access. Keep legal judgments with the responsible organization and its advisers.

    What has recent regulator guidance clarified?

    The European Data Protection Board’s AI model opinion summary explains that whether a model is anonymous requires a case-by-case assessment. It also sets out a three-stage analysis for legitimate interests: identify a legitimate interest, assess necessity, and balance the interests and rights involved. Public availability of training data does not by itself settle whether its use is lawful.

    CNIL’s AI development recommendations, available in English with updates in 2026, organize practical guidance around purpose, roles, legal basis, minimization, retention, rights, security, and model assessment. CNIL also provides a checklist. The additional tool in this article is a software team’s evidence register and verification drill, intended to support that official guidance.

    Do not collapse three separate questions into one: whether the source data may be used, whether the system handles it appropriately, and whether a trained model can disclose personal information. Each needs the right evidence and reviewer.

    Flowchart showing the AI privacy evidence chain from purpose and data flow through controls and tests to review.

    Figure 1. AI privacy evidence chain | © 2026 Uvik Software

    The Uvik Software AI privacy evidence chain links purpose, data flow, controls, tests, and review. Keep the evidence version with the approval so a later change is visible.

    Copy the Uvik Software AI privacy evidence checklist

    The following rows are proposed engineering handoff records. They are not an exhaustive statement of legal duties. Assign an owner, evidence link, review date, and status to each applicable row. Use not applicable only with a reason.

    Evidence area Record to prepare Verification question
    Purpose and scope Approved use case, users, exclusions, and decision owner Does the implemented feature stay inside that scope
    Processing roles Controller, processor, and other relevant role decisions Do contracts and actual behavior match the decision
    Data inventory Inputs, sources, categories, identifiers, and derived data Can the team find each category in the system
    Data flow Collection, retrieval, model calls, tools, logs, and backups Are all destinations and copies accounted for
    Legal review Basis, special category review, and impact assessment decision where relevant Has the reviewer seen the current design
    Vendor evidence Agreement, instructions, subprocessors, retention, and locations Are material vendor claims supported by current documents
    Access Roles, service accounts, permission filters, and removal process Can an unauthorized user retrieve or infer restricted content
    Retention and deletion Rules for each store, expiry behavior, exceptions, and recovery copies Does a test show the rule working as designed
    Rights handling Request intake, identity check, search, response, and escalation route Can the team locate relevant data and explain limitations
    User information Approved notices and explanations of the feature Does the interface show the reviewed version
    Incident response Detection, containment, evidence, and escalation process Can the team stop the feature and preserve relevant records
    Change control Triggers that require a new technical or privacy review Are model, connector, and purpose changes captured

    Keep this register in the same project space as the architecture and release record. A row marked complete should point to evidence that another person can inspect. A policy statement without a test may describe intent, while a test without the policy may prove the wrong behavior.

    Map the full data path

    Quick answer. An AI data map should include the original source, any preparation step, the search index, prompts, model provider, tools, outputs, logs, and recovery copies. Record the purpose and retention rule at each destination. A document can leave copies behind after it disappears from the user interface.

    For a retrieval assistant, start at the source system. Record which fields the connector reads, what the parser extracts, and what metadata the index stores. Include embeddings where used. Embeddings are numerical representations used for search; do not assume they are anonymous merely because they are difficult for a person to read.

    Follow the request through the application. A prompt may contain a user’s question, retrieved passages, and conversation history. A tool call may send that information to another service. An evaluation or error log may retain content that the main application deletes. Draw these as distinct destinations because the same retention rule may not apply to every one.

    Use Uvik Software’s data engineering consulting overview for the wider pipeline context. The privacy evidence record should identify the actual stores and owners in your implementation, rather than a generic architecture.

    A worked evidence record for a support assistant

    This example is fictional. The team proposes an assistant that drafts responses from approved help articles. It does not need customer account records for its initial scope. A privacy reviewer has still asked the team to assess user questions and logs, because employees may paste personal information into the input.

    The engineering team records four design choices: only the approved article collection is connected; account records are excluded; support employees review drafts before sending them; and the input interface explains what information should not be entered. Those choices reduce scope but do not guarantee that personal data never enters the system.

    The team adds input handling tests, access tests, and a log review. It documents which services receive the request, which logs contain content, and who can inspect them. The reviewer can then assess the real processing rather than relying on the label internal assistant.

    Register field Example entry
    Control ID PRIV 007
    Claim to verify Removing a source document removes it from active retrieval
    Test object Synthetic document with a unique test identifier
    Locations Source collection, parser output, index, cache, and relevant logs
    Test evidence Before and after query results, deletion event, and store checks
    Known limit Backup expiry follows the reviewed recovery and retention policy
    Owner and decision Named engineer supplies results; privacy reviewer assesses the evidence

    The use of synthetic test data avoids introducing a real person’s record just to demonstrate the control. The evidence still needs to reflect the paths used in production.

    Run a deletion and access removal drill

    Short answer. Test removal across the application path, record what remains, and explain why. Deleting one source row does not demonstrate deletion from indexes, caches, logs, backups, or a model trained on that data.

    Create a synthetic test item with a unique identifier and known access rules. Confirm that an authorized test user can retrieve it and an unauthorized user cannot. Record the configuration and source version.

    Trigger the normal removal process. Check the original store, parsed copies, search index, and application cache. Repeat the user queries. If relevant, remove the user’s access and verify that an existing session or cached result does not bypass the new permissions.

    Inspect logs and recovery copies under their separate rules. Record their retention, access restrictions, expiry process, and any approved exception. Do not claim immediate deletion from a backup if the actual design uses a controlled expiry cycle. Ask the reviewer whether that design is appropriate for the processing.

    If the item was used for model training or fine-tuning, a successful retrieval deletion test does not answer the model question. Document that path separately and escalate the technical and rights handling implications. Keep the response to the individual accurate about what has and has not changed.

    Finish the drill with a result, owner, remediation action, and retest date. A failed drill is useful evidence when it produces a tracked fix; hiding the exception makes the register less reliable.

    What should you ask an AI vendor?

    Ask which entity contracts with you and which role it takes for each service. Request the relevant processing agreement, subprocessor information, data locations, retention options, support access rules, and the behavior of deletion requests. Confirm the precise product edition covered by the answer.

    Ask whether prompts, outputs, uploaded files, and feedback are handled differently. Ask what changes when a connector or third-party tool is enabled. Get evidence for material claims in the agreement or current documentation, and identify questions that still require a legal judgment.

    No model training by default is useful information, but it does not settle every issue about processing, access, retention, or transfers. Similarly, a region setting does not describe every service path. Test what you can and send the remaining questions to the responsible reviewer.

    Keep evidence current as the system changes

    Review the register when the purpose expands, a new data source is connected, access changes, a vendor or model changes, or a new group of users joins. Set a review cadence appropriate to the system, and record the reason for that cadence.

    Uvik Software’s AI software development risk guide can help engineering teams maintain a broader risk record alongside the privacy evidence. Keep the two linked so an incident or design change reaches the right owners.

    Use and cite this checklist

    Suggested citation: Uvik Software, GDPR and AI engineering evidence checklist, 2026. Credit the Uvik Software AI privacy evidence checklist and original visual when adapting them. Legal statements rely on the linked official sources; the register, example, and drill are proposed engineering aids.

    Common GDPR and AI questions

    Is a RAG system automatically compliant because it does not train a model?

    No. Retrieval, prompts, outputs, and logs may still process personal data. The team needs to assess the actual processing and applicable obligations.

    Does removing names make the data anonymous?

    Not necessarily. Context and other fields may still identify someone, and a model may retain information. An anonymity conclusion needs an appropriate assessment rather than a naming convention.

    Can this checklist replace a data protection impact assessment?

    No. It can supply engineering evidence for a review or assessment. The controller determines whether an impact assessment is required, with appropriate advice, and remains responsible for the assessment. The checklist supplies technical evidence; it does not replace that process.

    How useful was this post?

    No votes so far! Be the first to rate this post.

    Share:
    GDPR and AI engineering evidence checklist - 11

    Need to augment your IT team with top talents?

    Uvik can help!
    Contact
    Uvik Software
    Privacy Overview

    This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.

    Get a free project quote!
    Fill out the inquiry form and we'll get back as soon as possible.

      Subscribe to TechTides – Your Biweekly Tech Pulse!
      Join 750+ subscribers who receive 'TechTides' directly on LinkedIn. Curated by Paul Francis, our founder, this newsletter delivers a regular and reliable flow of tech trends, insights, and Uvik updates. Don’t miss out on the next wave of industry knowledge!
      Subscribe on LinkedIn