Why training data is now a legal question, not just a technical one
For most of the last decade, the quality of the data a company used to build a model was an engineering concern. It affected accuracy and was argued about internally, but no regulator could ask to see it. Under the EU AI Act that has changed. Article 10 turns data governance into a legal obligation with a documented standard attached, and it does so for the systems businesses care most about — the ones used in hiring, credit, insurance, education and access to essential services.
The Digital Omnibus on AI, Regulation (EU) 2026/1744, made a change here that has been widely misreported. It did not weaken the data quality standard. It pulled the derogation allowing sensitive personal data to be used for bias testing out of Article 10 altogether and re-sited it, in broadened form, as a new Article 4a near the front of the Regulation — a relocation that matters more than it sounds, because the new provision reaches operators Article 10 never touched.
Which businesses Article 10 actually binds
Article 10 sits in Chapter III, Section 2 — the requirements for high-risk AI systems. It binds providers: the businesses that develop a high-risk system and place it on the market under their own name or trade mark. Deployers owe no Article 10 duty, though they owe a narrower one of their own discussed below.
Two pieces of scoping are easy to miss. Article 2(2), as rewritten by the Omnibus, provides that for high-risk systems under Article 6(1) relating to products covered by Section B of Annex I, only Article 6(1), Article 60a and Articles 102 to 112 apply — Article 10 is not among them. And a new Article 2(13) allows specific requirements in Articles 9 to 15 and 17 to 25 to be limited where Section A legislation already provides equivalent or higher protection, with delegated acts due by 2 August 2027. Article 10 falls within that range, so manufacturers of regulated products should expect their obligations to look different once that act lands.
What Article 10 requires
The article now runs 1, 2, 3, 4, 6 — a visible gap where paragraph 5 used to be, and a useful signal that you are reading current text rather than a stale copy.
The governance and management practices
Article 10(2) requires training, validation and testing data sets to be subject to data governance practices appropriate for the intended purpose, and says those practices shall concern, in particular, a specified set of matters — “in particular” making that a floor rather than a ceiling. Among the matters named are the design choices; the data collection processes and the origin of the data, and for personal data its original purpose of collection; preparation operations such as annotation, labelling, cleaning, enrichment and aggregation; the assumptions about what the data is supposed to measure; an assessment of the availability, quantity and suitability of the data sets needed; examination for biases likely to affect health and safety, fundamental rights or prohibited discrimination; measures to detect, prevent and mitigate those biases; and the identification of data gaps and how they can be addressed. This is a documentation regime as much as a technical one — almost every item is something a business must show it considered, not merely something it must get right.
The quality standard, which was not relaxed
Article 10(3) is the sentence quoted most and misquoted nearly as often. Data sets must be relevant, sufficiently representative, and, to the best extent possible, free of errors and complete in view of the intended purpose. They must also have appropriate statistical properties, including as regards the persons or groups the system is intended to be used on. It is worth being clear, because the contrary has circulated widely: this paragraph was not amended. The qualifier “to the best extent possible” attaches to freedom from errors and completeness, and it was in the original 2024 text. Relevance and sufficient representativeness carry no such qualifier.
Context, not just content
Article 10(4) requires data sets to take into account, to the extent required by the intended purpose, the characteristics particular to the geographical, contextual, behavioural or functional setting in which the system is intended to be used. For a Nordic business this is the paragraph with teeth. A model trained on North American or continental European populations, labour markets or claims histories may perform acceptably in aggregate and still be indefensible in Sweden, and the burden of showing the setting was considered sits with the provider. Systems built without training are not exempt either: under Article 10(6) as replaced, paragraphs 2, 3 and 4 and Article 4a(1) apply to them, but only to the testing data sets.
The new Article 4a, and why it reaches further
Testing a model for bias against protected characteristics generally requires processing data about those characteristics — exactly the category the GDPR restricts most tightly. The AI Act’s answer used to be Article 10(5). It is now Article 4a, sitting in Chapter I among the general provisions rather than in the high-risk chapter.
Article 4a(1) permits providers of high-risk systems, to the extent strictly necessary for bias detection and correction under Article 10(2), points (f) and (g), to process special categories of personal data subject to appropriate safeguards. Six cumulative conditions must be met, and they apply in addition to the GDPR rather than instead of it. In outline, the bias work must not be achievable using other data including synthetic or anonymised data; the data must be subject to technical limitations on re-use, to state-of-the-art security and privacy-preserving measures including pseudonymisation, and to strict documented access controls; it must not be transmitted or otherwise accessed by other parties; it must be deleted once the bias has been corrected or the retention period ends; and the records of processing must explain why it was strictly necessary and why no other data would do.
Article 4a(2) is the genuinely new reach. It extends the same permission, on the same conditions, to providers and deployers of other AI systems and models, and to deployers of high-risk systems, where the processing is strictly necessary to detect and correct biases likely to affect health and safety, fundamental rights or prohibited discrimination. It closes with a sentence businesses should read carefully: this paragraph does not create any obligation to conduct bias detection and correction. It is a permission, not a mandate.
A small amendment elsewhere makes this work. Article 2(7) previously stated flatly that the AI Act does not affect the GDPR; as amended it reads “without prejudice to Articles 4a and 59 of this Regulation”. That carve-out is what allows Article 4a to derogate rather than merely restate. One practical warning: the Commission’s own AI Act Service Desk still serves the pre-Omnibus Article 10 with a live paragraph 5, and has no page for Article 4a at all. Any memo citing “Article 10(5)” for the bias derogation is citing repealed text.
Praktiskt exempel
A Gothenburg insurance technology company builds a model that prices life and health policies and licenses it to insurers across the Nordics. Risk assessment and pricing in life and health insurance is Annex III point 5(c), so the system is high-risk, and in Sweden market surveillance for that point falls to Finansinspektionen under the interim designations made in June 2026.
The company’s data is drawn largely from one insurer’s historical Swedish book. Article 10(3) requires it to be sufficiently representative in view of the intended purpose, and that purpose is Nordic-wide. Article 10(4) requires the geographical and behavioural setting to be taken into account. A book reflecting one insurer’s underwriting appetite in one country is a weak foundation for a product sold in four, and the obligation is not only to fix that but to have documented the suitability assessment that identified it. The company also wants to test whether its pricing produces disparate outcomes by ethnicity or health status. That is permitted under Article 4a(1), but only if it can show the test cannot be run on synthetic or anonymised data, that the data is pseudonymised and access-controlled, that it goes nowhere outside the company, and that it is deleted once the bias is corrected. If the insurers deploying the model want to run their own bias testing, they are deployers rather than providers — and before the Omnibus they had no equivalent permission at all.
Vanliga misstag som företag gör
The first is believing the standard was relaxed. It was not, and a programme built on the assumption that “free of errors and complete” no longer means much rests on a misreading of the Digital Omnibus.
The second is citing repealed law. Article 10(5) no longer exists, and the Commission’s own per-article reference pages have not caught up. Any policy, template or supplier questionnaire citing Article 10(5) needs updating.
The third is treating Article 4a as an obligation to run bias testing. It is a permission to use sensitive data when you do; the duty to test comes from Article 10(2) and from discrimination law.
The fourth is deployers assuming Article 10 is somebody else’s problem. Article 26(4) provides that, to the extent the deployer exercises control over the input data, it must ensure that input data is relevant and sufficiently representative in view of the intended purpose. That is narrower than Article 10, but it is a real duty owed by the buyer, not the builder.
Rekommenderade åtgärder
Providers of systems likely to be high-risk should treat Article 10 as a documentation project starting now, because the evidence it calls for is retrospective and cannot be manufactured later on. Work through the matters listed in Article 10(2) and ask, for each, whether a regulator reading your files cold could see that you considered it. Pay particular attention to Article 10(4): Nordic deployment of models trained elsewhere is one of the most common and least documented exposures in the market.
On bias testing, rewrite any internal guidance still pointing at Article 10(5), and build the Article 4a conditions into the design of the test rather than bolting them on. That means deciding up front whether synthetic or anonymised data would do, specifying pseudonymisation and access controls, fixing a deletion trigger tied to correction of the bias, and writing the necessity reasoning into your GDPR records at the time. Deployers should ask suppliers what Article 10 documentation exists, and scope their own Article 26(4) duty. And nobody should wait for standards: none has been cited in the Official Journal, so no presumption of conformity is available to anyone.
Vanliga frågor
Did the Digital Omnibus weaken the data quality requirements?
No. Article 10(2), (3) and (4) are unchanged from the original 2024 text, including the requirement that data sets be relevant, sufficiently representative and, to the best extent possible, free of errors and complete in view of the intended purpose. The Omnibus amended Article 10 in three places only: it adjusted the cross-references in paragraphs 1 and 6, and it deleted paragraph 5, whose substance moved to the new Article 4a.
Can we use ethnicity or health data to test our model for bias?
Article 4a permits it where strictly necessary and where six cumulative conditions are met, and those conditions apply in addition to the GDPR rather than replacing it. The threshold question is whether the testing could be done effectively on other data, including synthetic or anonymised data — if it could, the derogation is not available.
When does Article 10 actually start to apply?
The Digital Omnibus deferred the high-risk obligations for Annex III systems to 2 December 2027, and for Annex I systems under Article 6(1) to 2 August 2028, and those dates are no longer contingent on standards. The GDPR, however, applies to the personal data in your training sets today, and always has.
Slutsats
Article 10 turns a machine learning team’s internal practices into evidence a regulator can demand. The Digital Omnibus left its substance intact and moved the sensitive-data derogation somewhere more useful, extending it to deployers and to systems outside the high-risk category for the first time. Businesses that read the change as a relaxation will build the wrong programme. Those that read it accurately will find a clearer legal basis for the bias testing they should be doing anyway, and rather less time than they think to document how their data was assembled.
Lawgent helps businesses build Article 10 data governance documentation, structure bias testing under Article 4a, and align AI Act data duties with their existing GDPR obligations.