EDPB's New Anonymisation and AI Scraping Drafts: What to Do Now

EDPB's New Anonymisation and AI Scraping Drafts: What to Do Now
August 8, 2026

Two Draft Guidelines Just Reset the GDPR Anonymisation Baseline

On 7 July 2026, the EDPB adopted two draft guidelines that privacy counsel must read this quarter: draft Guidelines 02/2026 on Anonymisation and draft Guidelines 03/2026 on Web Scraping in the Context of Generative AI. Both are open for public consultation until 30 October 2026, as IAPP reported.

The headline for practitioners: the anonymisation draft, once finalised, will replace the Article 29 Working Party's Opinion 05/2014 on Anonymisation Techniques (WP216), which had stood since 10 April 2014. That is a twelve-year-old benchmark being retired.

Do not treat these as academic. The consultation window is your window to comment, and the draft framework will shape how supervisory authorities test your anonymisation claims and your generative-AI training pipelines. Start mapping your data flows against the new criteria now, not after finalisation.

The Three-Criteria Test: No Record Isolation, No Linkage, No Inference

The core of Guidelines 02/2026 is a cumulative three-part test. Data can be considered anonymous only if all three criteria are satisfied: No Record Isolation, No Linkage, and No Inference. If any single criterion is not met, the data is not anonymous and further analysis is required.

This is the practitioner trap. Teams often strip direct identifiers, declare victory, and treat the dataset as out of scope. That will not fly with a regulator under this framework. Residual re-identification risk through linkage or inference keeps the data inside the GDPR.

  • No Record Isolation — you cannot single out an individual's record within the dataset.
  • No Linkage — you cannot connect records relating to the same individual across datasets.
  • No Inference — you cannot deduce, with significant probability, the value of an attribute about an individual.

The practical test is not what your vendor's marketing deck says. It is whether residual risk remains after you strip direct identifiers, and whether that risk survives all three prongs.

Data Provenance and the CJEU EDPS v SRB Signal

Why provenance drives the anonymisation analysis

The anonymisation draft lands against the CJEU's judgment in EDPS v SRB (Case C-413/23 P), issued 4 September 2025. The Court held that pseudonymised data transferred to a third party is not automatically personal data from that recipient's perspective if the recipient lacks reasonable means to re-identify individuals.

That holding is context-relative, and context is a provenance question. Whether a dataset is personal data can differ between the party that holds the identification key and the party that does not. Counsel must therefore document, per recipient and per transfer, what re-identification means each party actually possesses.

What provenance documentation must capture

  • The source of each dataset and the transformation steps applied to it.
  • The legal basis for each processing step, especially for generative-AI training corpora assembled by web scraping.
  • The recipient-relative re-identification analysis — who holds keys, auxiliary data, or realistic linkage capacity.

For generative AI, Guidelines 03/2026 put the scraping stage under scrutiny. If you cannot show where training data came from and on what legal basis you collected it, you cannot govern it. Build the provenance record first; then argue anonymisation.

The Digital Omnibus Backdrop and Why Definitions Still Matter

These drafts arrive while the definition of personal data itself is contested. On 10 February 2026, the EDPB and EDPS adopted Joint Opinion 2/2026 on the Digital Omnibus, expressing significant reservations about the proposed narrowing of the GDPR's definition of personal data in Article 4(1).

A Council of the EU compromise text dated 20 February 2026, circulated by the Cypriot presidency, then eliminated the Digital Omnibus's proposed new definition of personal data, removing the language that would have codified a narrower, entity-relative definition, according to IAPP reporting.

Do not build compliance on a definition that may not survive trilogue. The legislative status remains unsettled. Design your data maps to the current Article 4(1) definition as interpreted through the three-criteria anonymisation test, and treat any statutory narrowing as a contingency, not a foundation.

Action Items for Counsel This Quarter

What to operationalize before 30 October 2026

  1. Re-test every dataset you call anonymous. Run it through No Record Isolation, No Linkage, and No Inference. Flag any dataset that fails one prong; it is personal data.
  2. Build a recipient-relative register. Following EDPS v SRB, document for each transfer whether the recipient holds reasonable means to re-identify. Do not assume a uniform answer across parties.
  3. Inventory your generative-AI training corpora. Identify scraped sources, collection dates, and legal basis. You cannot govern what you have not mapped.
  4. Draft consultation comments. The window closes 30 October 2026. If the three-criteria test creates operational friction for your pipelines, comment on the record before finalisation.
  5. Refresh your DPAs and vendor diligence. Vendor anonymisation claims must be tested against the new framework, not accepted on the strength of a certification logo.

Sequence matters. Map first, classify second, then draft notices and DPAs against that classification — not the other way around.

Key Takeaways and How FinTech Law Helps

Key takeaways

  • The three-criteria test raises the anonymisation bar. Guidelines 02/2026 require No Record Isolation, No Linkage, and No Inference — all three, cumulatively — before data is anonymous.
  • Anonymisation is now recipient-relative. EDPS v SRB (Case C-413/23 P, 4 September 2025) confirms the analysis can differ between a party holding the key and one that does not.
  • Web scraping for generative AI is squarely in scope. Guidelines 03/2026 put the collection stage under GDPR scrutiny; provenance and legal basis must be documented.
  • A twelve-year benchmark is being retired. The drafts will replace WP216 from 10 April 2014 once finalised.
  • The consultation closes 30 October 2026. That is your window to comment before the framework hardens.

How FinTech Law helps

FinTech Law helps companies operationalize privacy and AI governance — from data maps and retention schedules to anonymisation testing and generative-AI training-data provenance. The firm works where product, legal, and security teams actually meet, and builds the audit trails supervisory authorities expect to see on day one of an inquiry.

If you train or fine-tune models on scraped data, or you rely on anonymisation to move data out of GDPR scope, contact FinTech Law to stress-test your position before finalisation. Learn more at fintechlaw.ai.

This post is for general information only and does not constitute legal advice. No attorney-client relationship is formed by reading it. Consult qualified counsel about your specific facts.