COMPLIPATHDOC complipath.io/guides/ai-act-data-governanceRENDERED 2026-08-23ENGINE 2026-08-09.1CORPUS 2024/1689 + 2026/1744 + Commission guidelines
Guides/Requirements ·By Yobel Tzegai ·Updated 23 August 2026

What are the data governance requirements under the EU AI Act (Article 10)?

Updated 9 August 2026 for Regulation (EU) 2026/1744.

If a high-risk AI system learns from data, the law sets rules for that data: what it has to cover, how it must be looked after, and what you are allowed to keep in it. High-risk AI systems that train models with data must be developed on training, validation and testing data sets meeting the quality criteria in Article 10(2), (3) and (4) and Article 4a(1): documented governance practices, a relevance and representativeness standard, contextual fit, and tightly conditioned bias-data processing. Governance practices here means the written steps by which the data is collected, checked and prepared — the Act's own term is "data governance and management practices". The requirement applies from 2 December 2027 for Annex III systems and from 2 August 2028 for Annex I systems (Article 113, third paragraph, point (c), as amended by Regulation (EU) 2026/1744).

Quick answer

Which data sets does Article 10 cover?

Article 10(1) covers "training, validation and testing data sets" for high-risk AI systems "which make use of techniques involving the training of AI models with data", and requires them to meet the quality criteria "referred to in paragraphs 2, 3 and 4 of this Article and in Article 4a(1)" whenever such data sets are used. All three set types carry the same criteria — a clean training set with an unexamined test set fails the text.

If your high-risk system is built without training techniques — a rule-based scoring engine, for example — Article 10(6) narrows the requirement: paragraphs 2, 3 and 4 and Article 4a(1) then "shall apply only to the testing data sets". Narrows, not removes.

Article 10 is one of the Section 2 requirements every high-risk system must comply with under Article 8(1). For Annex III systems it applies from 2 December 2027; for systems that are high-risk under Article 6(1) — AI systems that are safety components of, or are themselves, products covered by the Union harmonisation legislation listed in Annex I — the date is 2 August 2028 (Article 113, third paragraph, point (c), as amended by Regulation (EU) 2026/1744). Neither date has passed. The full high-risk regime covers where Article 10 sits among Articles 8–15. For providers, non-compliance is an Article 16 breach, fined up to €15 million or, if the offender is an undertaking, 3% of total worldwide annual turnover for the preceding financial year, whichever is higher (Article 99(4), point (a)).

Which practices must your data governance actually cover?

Article 10(2) requires "data governance and management practices appropriate for the intended purpose" — appropriate to your system, not a fixed template — concerning in particular:

Treat the eight points as a file structure: one documented answer per point, per system. That file is not an end in itself — it becomes source material for the technical documentation under Article 11, and the biases you examine under point (f) and address under point (g) are exactly the kind of risk your Article 9 risk management system must evaluate and treat. Write it once, in a form both can cite.

How good does the data have to be — really "error-free"?

No. Article 10(3) sets the standard precisely: data sets "shall be relevant, sufficiently representative, and to the best extent possible, free of errors and complete in view of the intended purpose". The qualifier sits where it sits: relevance and sufficient representativeness are unqualified requirements, while error-freeness and completeness are best-effort obligations. A vendor claiming Article 10 demands perfect data is misreading it; a provider with no error-handling evidence at all is breaching it.

Two refinements matter in practice. Data sets must have "the appropriate statistical properties, including, where applicable, as regards the persons or groups of persons in relation to whom the high-risk AI system is intended to be used" — representativeness is measured against the people the system will judge. And the characteristics "may be met at the level of individual data sets or at the level of a combination thereof" — a known skew in one set can be compensated by another, if you can show the combination works.

Article 10(4) adds contextual fit: data sets must reflect, "to the extent required by the intended purpose", the characteristics particular to the "specific geographical, contextual, behavioural or functional setting" of intended use. A credit model trained solely on one market's data needs an answer for why it fits another. Getting this right buys something concrete: under Article 42(1), systems trained and tested on data reflecting that setting "shall be presumed to comply with the relevant requirements laid down in Article 10(4)".

Can you process special categories of personal data to detect bias?

Exceptionally, yes — but the provision has moved. Regulation (EU) 2026/1744 deleted Article 10(5) and inserted a new Article 4a, "Processing of special categories of personal data for bias detection and correction". Article 4a(1) carries the same six conditions, and Article 10(1) and (6) now cite it in place of the deleted paragraph.

Article 4a(1) permits providers of high-risk AI systems to process special categories of personal data "[t]o the extent strictly necessary" for bias detection and correction under Article 10(2), points (f) and (g), "subject to appropriate safeguards for the fundamental rights and freedoms of natural persons". It applies "[i]n addition to the provisions set out in" Regulations (EU) 2016/679 and (EU) 2018/1725 and Directive (EU) 2016/680, "as applicable" — a pathway on top of data-protection law, not an exemption from it; how the two regimes stack is covered in the AI Act vs GDPR. All of the following conditions shall be met:

Article 4a(2) has no counterpart in the deleted Article 10(5). It extends the same pathway, on all the paragraph 1 conditions and safeguards, to providers and deployers of other AI systems and models and to deployers of high-risk AI systems, where processing is strictly necessary against biases likely to affect health and safety, harm fundamental rights or lead to discrimination prohibited under Union law — "especially where data outputs influence inputs for future operations". It "does not create any obligation to conduct such bias detection and correction".

What this means for you

If you're a provider: Article 16, point (a) makes Article 10 your obligation for every high-risk system that touches data. Build the file the article dictates: eight answers under Article 10(2), the Article 10(3) statistical evidence, the Article 10(4) setting analysis — and, if you touch special-category data for bias work, a record against each of the six Article 4a(1) conditions. Complipath's requirements checklist breaks these into per-system items with an evidence link on each, so the file exists before anyone asks for it.

If you're a deployer: Article 10 does not bind you — but don't buy blind. Before procurement, ask the provider how the training data matches your setting (their Article 10(4) analysis is the honest answer to that question). And you carry your own version of the standard: under Article 26(4), to the extent you control the input data, you must ensure it is "relevant and sufficiently representative in view of the intended purpose" — the same vocabulary Article 10(3) applies to the provider's data sets. Unsure which side you're on? See provider vs deployer.

Which of your systems answer to Article 10?

Classify your system now — 7 questions on the main line, plus follow-ups where they apply, no account, and the classification runs in your browser: answers stay there unless you choose to keep the result.

FAQ

Does the EU AI Act require training data to be error-free? No. Article 10(3) requires data sets to be relevant and sufficiently representative without qualification, but "free of errors and complete" only "to the best extent possible" in view of the intended purpose. You need documented error-handling and completeness measures, not perfection.

Does Article 10 apply if our system doesn't train a model? Partly. Under Article 10(6), for high-risk AI systems developed without techniques involving the training of AI models, paragraphs 2, 3 and 4 and Article 4a(1) apply only to the testing data sets. You still need governed, quality-checked test data — the full training-data regime falls away.

Can we use health or ethnicity data to test our model for bias? Exceptionally. Article 4a(1) — which replaced the deleted Article 10(5) in 2026 — allows providers to process special categories of personal data where strictly necessary for bias detection and correction, but only where all six of its conditions are met, on top of GDPR, and with deletion once the bias is corrected or retention ends.

Is Article 10 a provider or deployer obligation? A provider obligation: Article 16, point (a) requires providers to ensure their high-risk systems comply with the Section 2 requirements, which include Article 10. Deployers instead have Article 26(4): where they control input data, it must be relevant and sufficiently representative for the intended purpose.


Sources: Regulation (EU) 2024/1689, Articles 8, 10, 16, 26, 40, 42, 99 and 113 (EUR-Lex), as amended by Regulation (EU) 2026/1744 (EUR-Lex), which inserted Article 4a, deleted Article 10(5), replaced Article 10(1) and (6), and replaced Article 113, third paragraph, point (c). The Article 40 harmonised standards that would give Article 10 a presumption-of-conformity route were still incomplete at the time of writing; the obligations apply on the Article 113 dates regardless — 2 December 2027 for Annex III systems, 2 August 2028 for Annex I systems.

← All guides
Complipath

Complipath is EU AI Act compliance software for AI-heavy software companies without a compliance team — an AI system register, deterministic risk classification, the obligations that follow, and the evidence behind every decision.

Complipath is built by Yobel Tzegai in Gothenburg, Sweden.

Complipath provides legal information, not legal advice. Every guide cites its source on EUR-Lex — Regulation (EU) 2024/1689, and Regulation (EU) 2026/1744 where that has amended it; where the law is still settling, the guide says so.

We measure page views with Vercel Web Analytics. It uses no third-party cookies. Visitors are identified by a hash derived from the incoming request, which is discarded after 24 hours, and no identifier is stored that could follow a visitor to another site. What is collected: the time of the visit, the URL, the referring page, filtered query parameters, city-level location, operating system, browser and device type.