LinkedInInstagramXTikTok

Training AI on personal data: what the GDPR actually requires

Why this matters now

Almost every company that builds, fine-tunes or buys an AI model is processing personal data somewhere in the chain – in the training set, in the prompts, or in the outputs. The EU AI Act sets the rules for how AI systems must behave. The GDPR still decides whether you were allowed to use the data in the first place, and it is the GDPR that supervisory authorities have been enforcing hardest.

The European Data Protection Board’s Opinion 28/2024, adopted in December 2024, is the closest thing we have to a rulebook for AI models and personal data. It answers three questions that come up in almost every AI project we advise on: is the model itself personal data, can you rely on legitimate interest, and what happens if the training data was collected unlawfully.

Is an AI model “anonymous”?

Many vendors claim that once a model is trained, the personal data is gone – only weights and parameters remain. The EDPB does not accept that as a general rule. A model trained on personal data can only be considered anonymous if, taking all means reasonably likely to be used into account, the likelihood of extracting personal data from the model and the likelihood of obtaining it through queries are both insignificant.

That is a case-by-case assessment, and the burden of proof sits with the controller. In practice it means documenting what you did to reduce memorisation, how you tested the model for regurgitation of training data, and what safeguards prevent extraction attacks. A vendor’s marketing statement is not documentation.

Can you rely on legitimate interest?

Yes – but not by default. Legitimate interest under Article 6(1)(f) GDPR remains available for both developing and deploying AI models, and the EDPB confirms it is not excluded. What it requires is the classic three-step test, done properly and in writing:

1. A real, present and lawful interest

“Improving our product” is thin. Interests such as developing a conversational assistant, detecting fraud or improving cybersecurity are far easier to defend because they are concrete and can be evidenced.

2. Necessity

Could you achieve the same result with less data, synthetic data, or aggregated data? If you could, legitimate interest fails at this step. Scraping everything “because more data is better” will not survive scrutiny.

3. Balancing against the individual’s rights

The decisive factor is reasonable expectations. Was the data publicly available? Did the individual have any relationship with you? Was the source a professional network, a public register, or a private forum? Could the data be sensitive, or reveal location, health or political views? Where expectations tilt against you, mitigating measures matter – opt-outs that actually work, filtering of special-category data, limits on output, and transparency that reaches the people concerned.

What if the training data was collected unlawfully?

This is the risk that most boards underestimate. The EDPB is clear that unlawful processing at the development stage can taint the deployment stage – unless the model has been genuinely anonymised before it is deployed. If you buy or license a model, you inherit that question. Supervisory authorities can, in principle, order the deletion of data or restrict processing, and the commercial consequence of losing the right to use a model you have built a product on is far larger than the fine.

Practical example: a Swedish SaaS company fine-tunes a support bot

A B2B software company fine-tunes an open-weight model on five years of customer support tickets. The tickets contain names, email addresses, order numbers and occasional health references from sick-leave-related cancellations. The company has a lawful basis for holding the tickets – but not necessarily for reusing them to train a model.

The workable path: run a compatibility assessment for the new purpose, pseudonymise or strip identifiers before training, exclude tickets containing special-category data, document the three-step test, update the privacy notice, and test the fine-tuned model for memorisation before it goes anywhere near production. None of this is exotic. All of it needs to be written down.

Common mistakes companies make

Treating the AI Act as the whole picture. The AI Act does not give you a legal basis to process personal data. The GDPR still applies in full, in parallel.

Assuming “publicly available” means free to use. Public data is still personal data. Web-scraped data carries a heavier balancing burden, not a lighter one.

Relying on consent where it cannot be withdrawn. If you cannot remove an individual’s influence from a trained model, consent is a fragile basis. Legitimate interest, properly documented, is often the more honest choice.

No record of the assessment. The GDPR is an accountability regime. An assessment you cannot produce did not happen.

Recommended actions

Map where personal data enters your AI lifecycle: training, fine-tuning, RAG sources, prompts, logs and outputs. Decide and document a legal basis for each stage. Run a DPIA where the processing is large-scale, involves special-category data, or uses scraped data. Test for memorisation and log the results. Put AI-specific clauses into vendor contracts – provenance of training data, indemnities, and the right to audit. And make sure the people whose data you use can actually exercise their rights.

Frequently asked questions

Does Opinion 28/2024 create new obligations?

No. It interprets existing GDPR obligations for the AI context. But it is what national supervisory authorities look at, so in practice it sets the standard you will be measured against.

Can we train on customer data if our terms allow it?

Contractual permission is not automatically a GDPR legal basis, and terms buried in a service agreement rarely create reasonable expectations. Assess the purpose separately.

Does this apply if we only use a third-party model such as a hosted LLM?

Yes. Deployment is processing. You need a legal basis for the data you send to the model, a data processing agreement, clarity on whether your inputs are used for training, and a view on where the data is transferred.

Conclusion

AI governance is not a separate discipline from data protection – it is data protection applied to a new technology, with higher stakes and less room for improvisation. Companies that document their legal basis, test their models and control their vendors can move fast. Companies that skip the paperwork are building products on a foundation a regulator can remove.

Lawgent helps companies build AI and data governance that holds up in practice – legal basis assessments, DPIAs, AI policies and vendor contracts, delivered as working documents rather than theory. Book a free first hour and we will tell you where you actually stand.

Leave a Reply

Your email address will not be published. Required fields are marked *


0Cart0,00 

No products in the cart.

Return to shop