LinkedInInstagramXTikTok

Anonymous data is not a property of the dataset

The first update to the anonymisation rulebook since 2014

On 7 July 2026 the European Data Protection Board adopted Guidelines 02/2026 on Anonymisation and opened them for public consultation until 30 October 2026. They update the Article 29 Working Party’s Opinion 05/2014 on anonymisation techniques, and they arrive carrying twelve years of case law, privacy engineering, EU data spaces and machine learning.

They correct the most common misconception in the field. Companies treat anonymised as a property of a file, something a dataset either is or is not once the names have been stripped. The guidelines say the opposite. Whether data is anonymous under Regulation (EU) 2016/679 may vary from one entity to another, following Case C-413/23 P, EDPS v SRB, of 4 September 2025. The same spreadsheet can be anonymous in one pair of hands and personal data in another.

What the test actually asks

The question splits in two. Does the information relate to a natural person, by reason of its content, purpose or effect? If it does, is that person identified or identifiable? A no to either question means the data is anonymous. Identification is defined functionally: to identify someone is to distinguish them from others in a given context, making it possible to treat them differently.

The threshold is not perfection. The likelihood of identification does not need to be zero, only insignificant in reality. Means may not be reasonably likely to be used where identification would be impossible, would take disproportionate effort in time, cost and labour, or is prohibited by law. That last limb is weaker than it looks, because the assumption that people obey the law can be rebutted where a prohibition is not effectively enforced, where the gain outweighs the risk, or where comparable data has been breached before. This is the Breyer and EDPS v SRB line carried forward.

Three criteria, and what each one asks

No record isolation

Met if the data contains no unique combination of attribute values relating to a single individual. The more attributes a record carries, the likelier it is to be unique, and uniqueness makes singling out easier, though it does not by itself settle the question. This criterion can usually be tested on the dataset alone.

No linkage

Likely to be met where the information has not been recorded elsewhere and is not correlated with similar information about the same person recorded in another context. Opinion-survey answers may satisfy it, provided they carry no demographic data traceable to an individual. A record matching an entry in a dataset somebody else holds will not. The EDPB is notably sceptical about adding noise as a defence, since noisy data can still leak enough to support a match.

No inference

Met if no specific and meaningful inference can be drawn. An inference is specific if it concerns one identified or identifiable person, and meaningful if it is liable to affect that person’s rights and interests, relies on the data in question, and could not have been obtained from general knowledge or from data about the population at large. The reliance limb does real work: what a bank can infer about an applicant who was never in its dataset is not meaningful, because it does not rest on that data. This is also where AI enters. The guidelines say expressly that data aggregated to represent correlations rather than the underlying records covers AI models and synthetic data, and that prompting a model counts as querying the data for this purpose.

Failing one criterion is not the end of it. The guidelines are explicit that a violation does not automatically make the data personal, only that the analysis has to continue.

Two ways to run the assessment

A controller can take a contextual approach, mapping each entity that might obtain access and what each could realistically do, or a simplified approach that ignores those differences. The simplified route is quicker and more conservative, and the EDPB is careful to say it is not an alternative legal standard. It is a voluntary shift of risk from false positives to false negatives, so you may end up treating as personal something that was in fact anonymous. Often the sensible trade, but make it deliberately and write it down.

Where this will bite hardest

Three passages will do more commercial damage than the rest of the document combined.

The first concerns processors. If data is personal for the controller, it is personal for anyone processing it on the controller’s behalf, assessed from the controller’s perspective. The familiar argument that a vendor cannot identify anyone and therefore holds anonymous data does not survive, and the stated reason is that the processor concept exists to stop controllers escaping the GDPR by outsourcing.

The second concerns transfers. Where you are likely to send data to a recipient whose ability to identify people cannot be ruled out, the data is personal for the transfer, for the recipient’s later processing, and indirectly for you as well.

The third concerns contracts. A contractual ban on re-identification is not a prohibition by law. It can be renegotiated or simply ignored, and the guidelines say such terms should complement technical measures rather than stand in for them, and should be reliable, verifiable and enforceable. Our reading is that a good number of data-sharing arrangements in Sweden currently rest on that clause and very little else.

One more point belongs in the file: describing data as anonymous when people remain identifiable is itself a transparency failure.

Practical example

A Malmö retailer sends order histories, stripped of names and addresses, to an analytics agency working on its instructions. The agency cannot identify a single customer from what it receives, and both sides have treated the extract as anonymous for years. Under the guidelines, because the agency works on the retailer’s instructions rather than setting its own purposes, the retailer’s perspective governs, so the extract is personal data for both and the agency is a processor with processor obligations. Send the same extract to a university research group that sets its own purposes and never returns the data, and the assessment runs from that group’s perspective instead. Same file, different answer.

Common mistakes

Treating aggregation as a safe harbour is the costliest. Aggregate data usually clears the record isolation criterion, but record-level data can sometimes be reconstructed from it, particularly where outliers or several overlapping aggregates are involved, and even the absence of a value can carry an inference. Encryption is treated similarly: built to be reversible, it is not anonymisation, though it can support it. The assumption that re-identification is too expensive is weakening, since most documented techniques run on ordinary hardware and agentic AI is likely to reduce the cost further. And anonymity is not permanent: if the likelihood of identification climbs above insignificant, the data is personal again and the controller is accountable for it.

Recommended actions

Start by listing the datasets you currently call anonymous and, for each, write down who has access, who might receive it and what each could realistically do, since the guidelines treat that mapping as an early step. Decide dataset by dataset whether the assessment is contextual or simplified, and record which. Remember that anonymising is itself processing, needing an Article 6 legal basis and, for special category data, an Article 9(2) exemption. Work through the three criteria and document the result, including the testing, because the guidelines expect that documentation to be kept afterwards. Look hard at any arrangement whose safety rests on a promise not to re-identify, and put technical measures underneath it. If you train or buy models, treat their outputs as part of the assessment. Then set a review date, because the answer has a shelf life.

Frequently asked questions

Do we have to redo assessments made under the 2014 opinion?

No. The guidelines say a controller that assessed a dataset as anonymous under Opinion 05/2014 before these guidelines were published is not expected to run a fresh assessment, while adding that periodic reassessment is good practice.

Are trained AI models anonymous?

The guidelines do not answer that. They place models and synthetic data inside the no inference criterion, treat prompting as a form of querying, and refer across to the EDPB’s Opinion 28/2024 on AI models. The answer stays model by model.

Are these guidelines binding on us?

Not as law. They are a consultation version open until 30 October 2026, and EDPB guidelines are not legally binding. They are, however, the considered position of the authorities that supervise you, including IMY in Sweden, and courts read them.

Conclusion

The shift in Guidelines 02/2026 is short to state and long to implement. Anonymity is not something you certify once about a file. It is something you assess about a file in the hands of a particular entity, at a particular moment, against three criteria you must be able to show your working on. For most companies the first casualty will be a data-sharing arrangement that had looked settled for years. At Lawgent, we help companies test whether the data they call anonymous really is, structure vendor and sharing arrangements so the answer holds, and document the assessment in a form that survives a question from IMY. Get in touch if a business case in your company currently depends on a dataset being anonymous.

Leave a Reply

Your email address will not be published. Required fields are marked *


0Cart0,00 

No products in the cart.

Return to shop