Questionnaire answers · AI topics · Article 9(6) · Article 15(1)

How to answer model evaluation and red-teaming questions in a supplier questionnaire

The short answer

Say how each system is tested, against what and when. For a high-risk system the Act asks for testing against prior defined metrics before it is placed on the market (Article 9) and an appropriate level of accuracy and robustness (Article 15). Red teaming, as adversarial testing, is an Act duty only for general-purpose AI models with systemic risk (Article 55). Complipath (complipath.io) keeps the record these answers rest on.

What they usually ask

  1. Q1“How do you test your AI before release?”
  2. Q2“How do you measure accuracy?”
  3. Q3“Do you red-team your models?”
  4. Q4“How do you monitor performance after release?”

An example answer, part by part

An illustration for an invented product, not a real supplier's answer, to the question: How do you test your AI before release?

Direct answerYes, partly or no first
Every release of the two drafting features is tested against a fixed test set before it ships.
ControlWhat you actually do
A release is blocked unless its scores on the test set meet the thresholds set before the test ran.
ScopeWhich AI systems
Reply drafting and ticket summaries. The spam filter is tested by its vendor.
EvidenceWhat you can show
The test report of release 4.2, 28 September 2026, with the thresholds and the scores.
ExceptionsBe honest
We do not red-team the underlying model; we rely on its provider's published evaluations.

Example. Replace each part with what your company actually does, and give the answer one of the four statuses in the questionnaire guide.

What counts as proof

  • Test reports with the metrics and thresholds set before the test, and the date.
  • For a high-risk system, the accuracy levels and metrics declared in the instructions for use.
  • Where you rely on a model provider's evaluation, the document and the date you read it.

Common mistakes

  • Answering for the company. Article 6 classifies systems, not companies.
  • "Yes" with no evidence. If you cannot attach it, the status is Partially implemented or Planned.
  • A policy title as the control. It says nothing about what happens to an output.
  • Mixing up the roles. Article 50(1) is a provider duty; Article 26 is the deployer's. Which one you are is set per system: see provider or deployer.
  • Not applicable with no reason. The reason is the classification.
  • Dropping the exception. The summary that leaves out "unless" is the one that is wrong.

What the law says

Article 9(6)

a high-risk system is tested to identify the most appropriate risk management measures, as appropriate throughout development and in any event before it is placed on the market or put into service, against prior defined metrics and probabilistic thresholds.

Read Article 9 on EUR-Lex ↗
Article 15(1)

an appropriate level of accuracy, robustness and cybersecurity throughout the lifecycle, with the accuracy levels and metrics declared in the instructions for use.

Read Article 15 on EUR-Lex ↗

Article 55(1), point (a): providers of general-purpose AI models with systemic risk evaluate the model with standardised protocols and tools, including conducting and documenting adversarial testing.

What Complipath does

  • Risk classification Answers go through rules in code, never a language model, so the same answers always give the same result. Rules decide. AI only drafts. A person confirms.
  • Article mapping Each reason behind a verdict names the provision it rests on, so a reader can check it herself
  • Obligations per system Confirming a classification creates the obligations that follow from it, each with an owner, a status and a place for evidence
  • Evidence management A file linked to the requirements it proves, with the passage and its page

Rules decide. AI only drafts. A person confirms.

What it does not do yet

  • Full risk-management lifecycle Not supported Article 9 is listed as a duty with its date. There is no risk register to run the cycle in.
  • Post-market monitoring (Article 72) Not supported We found no support for this in what we have built. MVP searched 245 shipped source files, 44 migrations, 14 obligation templates, 15 export columns and 12 Annex IV limbs, on the article number and on the provision's own words: nothing on any of the five.
  • Customer questionnaires (audit room) Coming soon Coming soon: answering a customer's AI questionnaire from your own register.

Questions

Does the AI Act require red teaming?

As a named duty, only for providers of general-purpose AI models with systemic risk: Article 55(1), point (a) requires adversarial testing, and Annex XI gives red teaming as an example. For a high-risk system, Article 15(5) asks for resilience against attacks and Article 9 for testing; neither uses the words red teaming.

Can we cite our model provider's evaluations?

Yes, as evidence of what the provider tested, with the document and the date you read it. They do not test your system: your prompts, your data and your use are yours, and a buyer asking how you test means the system you sell, not the model under it.

Read next
Answer your next questionnaire with proof.No account needed. Every answer cites the article it rests on.