What are the data governance requirements under the EU AI Act (Article 10)?
Updated 9 August 2026 for Regulation (EU) 2026/1744.
If a high-risk AI system learns from data, the law sets rules for that data: what it has to cover, how it must be looked after, and what you are allowed to keep in it. High-risk AI systems that train models with data must be developed on training, validation and testing data sets meeting the quality criteria in Article 10(2), (3) and (4) and Article 4a(1): documented governance practices, a relevance and representativeness standard, contextual fit, and tightly conditioned bias-data processing. Governance practices here means the written steps by which the data is collected, checked and prepared — the Act's own term is "data governance and management practices". The requirement applies from 2 December 2027 for Annex III systems and from 2 August 2028 for Annex I systems (Article 113, third paragraph, point (c), as amended by Regulation (EU) 2026/1744).
Quick answer
- Scope: systems "which make use of techniques involving the training of AI models with data" — training, validation and testing sets alike (Article 10(1)). Systems developed without training techniques: only the testing data sets (Article 10(6)). Both paragraphs were re-pointed at Article 4a(1) in 2026; Article 10(2), (3) and (4) are unamended.
- Practices: eight areas your data governance must cover, from design choices and data origin to bias examination and gap identification (Article 10(2)).
- The standard: data sets must be relevant, sufficiently representative, and — only "to the best extent possible" — free of errors and complete (Article 10(3)).
- Bias data: special categories of personal data may be processed for bias detection and correction only exceptionally, under six cumulative conditions — moved out of Article 10(5), which was deleted in 2026, into the new Article 4a(1).
- Who: Article 10 binds providers via Article 16, point (a); deployers carry a parallel input-data duty under Article 26(4).
Which data sets does Article 10 cover?
Article 10(1) covers "training, validation and testing data sets" for high-risk AI systems "which make use of techniques involving the training of AI models with data", and requires them to meet the quality criteria "referred to in paragraphs 2, 3 and 4 of this Article and in Article 4a(1)" whenever such data sets are used. All three set types carry the same criteria — a clean training set with an unexamined test set fails the text.
If your high-risk system is built without training techniques — a rule-based scoring engine, for example — Article 10(6) narrows the requirement: paragraphs 2, 3 and 4 and Article 4a(1) then "shall apply only to the testing data sets". Narrows, not removes.
Article 10 is one of the Section 2 requirements every high-risk system must comply with under Article 8(1). For Annex III systems it applies from 2 December 2027; for systems that are high-risk under Article 6(1) — AI systems that are safety components of, or are themselves, products covered by the Union harmonisation legislation listed in Annex I — the date is 2 August 2028 (Article 113, third paragraph, point (c), as amended by Regulation (EU) 2026/1744). Neither date has passed. The full high-risk regime covers where Article 10 sits among Articles 8–15. For providers, non-compliance is an Article 16 breach, fined up to €15 million or, if the offender is an undertaking, 3% of total worldwide annual turnover for the preceding financial year, whichever is higher (Article 99(4), point (a)).
Which practices must your data governance actually cover?
Article 10(2) requires "data governance and management practices appropriate for the intended purpose" — appropriate to your system, not a fixed template — concerning in particular:
- (a) the relevant design choices;
- (b) data collection processes and the origin of data — and, for personal data, the original purpose of its collection;
- (c) data-preparation operations "such as annotation, labelling, cleaning, updating, enrichment and aggregation";
- (d) the formulation of assumptions, in particular about what the data are supposed to measure and represent;
- (e) an assessment of the availability, quantity and suitability of the data sets needed;
- (f) examination for possible biases likely to affect health and safety, negatively impact fundamental rights, or lead to discrimination prohibited under Union law — "especially where data outputs influence inputs for future operations", the feedback-loop case;
- (g) appropriate measures to detect, prevent and mitigate the biases identified under point (f);
- (h) identification of data gaps or shortcomings that prevent compliance, and how they can be addressed.
Treat the eight points as a file structure: one documented answer per point, per system. That file is not an end in itself — it becomes source material for the technical documentation under Article 11, and the biases you examine under point (f) and address under point (g) are exactly the kind of risk your Article 9 risk management system must evaluate and treat. Write it once, in a form both can cite.
How good does the data have to be — really "error-free"?
No. Article 10(3) sets the standard precisely: data sets "shall be relevant, sufficiently representative, and to the best extent possible, free of errors and complete in view of the intended purpose". The qualifier sits where it sits: relevance and sufficient representativeness are unqualified requirements, while error-freeness and completeness are best-effort obligations. A vendor claiming Article 10 demands perfect data is misreading it; a provider with no error-handling evidence at all is breaching it.
Two refinements matter in practice. Data sets must have "the appropriate statistical properties, including, where applicable, as regards the persons or groups of persons in relation to whom the high-risk AI system is intended to be used" — representativeness is measured against the people the system will judge. And the characteristics "may be met at the level of individual data sets or at the level of a combination thereof" — a known skew in one set can be compensated by another, if you can show the combination works.
Article 10(4) adds contextual fit: data sets must reflect, "to the extent required by the intended purpose", the characteristics particular to the "specific geographical, contextual, behavioural or functional setting" of intended use. A credit model trained solely on one market's data needs an answer for why it fits another. Getting this right buys something concrete: under Article 42(1), systems trained and tested on data reflecting that setting "shall be presumed to comply with the relevant requirements laid down in Article 10(4)".
Can you process special categories of personal data to detect bias?
Exceptionally, yes — but the provision has moved. Regulation (EU) 2026/1744 deleted Article 10(5) and inserted a new Article 4a, "Processing of special categories of personal data for bias detection and correction". Article 4a(1) carries the same six conditions, and Article 10(1) and (6) now cite it in place of the deleted paragraph.
Article 4a(1) permits providers of high-risk AI systems to process special categories of personal data "[t]o the extent strictly necessary" for bias detection and correction under Article 10(2), points (f) and (g), "subject to appropriate safeguards for the fundamental rights and freedoms of natural persons". It applies "[i]n addition to the provisions set out in" Regulations (EU) 2016/679 and (EU) 2018/1725 and Directive (EU) 2016/680, "as applicable" — a pathway on top of data-protection law, not an exemption from it; how the two regimes stack is covered in the AI Act vs GDPR. All of the following conditions shall be met:
- (a) bias detection and correction "cannot be effectively fulfilled by processing other data, including synthetic or anonymised data" — try those first, and document why they failed;
- (b) technical limitations on re-use, plus state-of-the-art security and privacy-preserving measures, including pseudonymisation;
- (c) measures ensuring the data is secured and access-controlled — "strict controls and documentation of the access", authorised persons only, under appropriate confidentiality obligations;
- (d) the data is "not transmitted, transferred or otherwise accessed by other parties";
- (e) deletion "once the bias has been corrected or the personal data has reached the end of its retention period, whichever comes first" — bias fixed early ends the processing early; an expiring retention period ends it even if the bias work is unfinished;
- (f) the records of processing activities under those data-protection instruments must state why the processing was strictly necessary and why the objective "could not be achieved by processing other data".
Article 4a(2) has no counterpart in the deleted Article 10(5). It extends the same pathway, on all the paragraph 1 conditions and safeguards, to providers and deployers of other AI systems and models and to deployers of high-risk AI systems, where processing is strictly necessary against biases likely to affect health and safety, harm fundamental rights or lead to discrimination prohibited under Union law — "especially where data outputs influence inputs for future operations". It "does not create any obligation to conduct such bias detection and correction".
What this means for you
If you're a provider: Article 16, point (a) makes Article 10 your obligation for every high-risk system that touches data. Build the file the article dictates: eight answers under Article 10(2), the Article 10(3) statistical evidence, the Article 10(4) setting analysis — and, if you touch special-category data for bias work, a record against each of the six Article 4a(1) conditions. Complipath's requirements checklist breaks these into per-system items with an evidence link on each, so the file exists before anyone asks for it.
If you're a deployer: Article 10 does not bind you — but don't buy blind. Before procurement, ask the provider how the training data matches your setting (their Article 10(4) analysis is the honest answer to that question). And you carry your own version of the standard: under Article 26(4), to the extent you control the input data, you must ensure it is "relevant and sufficiently representative in view of the intended purpose" — the same vocabulary Article 10(3) applies to the provider's data sets. Unsure which side you're on? See provider vs deployer.
Which of your systems answer to Article 10?
Classify your system now — 7 questions on the main line, plus follow-ups where they apply, no account, and the classification runs in your browser: answers stay there unless you choose to keep the result.
FAQ
Does the EU AI Act require training data to be error-free? No. Article 10(3) requires data sets to be relevant and sufficiently representative without qualification, but "free of errors and complete" only "to the best extent possible" in view of the intended purpose. You need documented error-handling and completeness measures, not perfection.
Does Article 10 apply if our system doesn't train a model? Partly. Under Article 10(6), for high-risk AI systems developed without techniques involving the training of AI models, paragraphs 2, 3 and 4 and Article 4a(1) apply only to the testing data sets. You still need governed, quality-checked test data — the full training-data regime falls away.
Can we use health or ethnicity data to test our model for bias? Exceptionally. Article 4a(1) — which replaced the deleted Article 10(5) in 2026 — allows providers to process special categories of personal data where strictly necessary for bias detection and correction, but only where all six of its conditions are met, on top of GDPR, and with deletion once the bias is corrected or retention ends.
Is Article 10 a provider or deployer obligation? A provider obligation: Article 16, point (a) requires providers to ensure their high-risk systems comply with the Section 2 requirements, which include Article 10. Deployers instead have Article 26(4): where they control input data, it must be relevant and sufficiently representative for the intended purpose.
Sources: Regulation (EU) 2024/1689, Articles 8, 10, 16, 26, 40, 42, 99 and 113 (EUR-Lex), as amended by Regulation (EU) 2026/1744 (EUR-Lex), which inserted Article 4a, deleted Article 10(5), replaced Article 10(1) and (6), and replaced Article 113, third paragraph, point (c). The Article 40 harmonised standards that would give Article 10 a presumption-of-conformity route were still incomplete at the time of writing; the obligations apply on the Article 113 dates regardless — 2 December 2027 for Annex III systems, 2 August 2028 for Annex I systems.