COMPLIPATHDOC complipath.io/guides/we-ran-our-engine-on-itselfRENDERED 2026-08-23ENGINE 2026-08-09.1CORPUS 2024/1689 + 2026/1744 + Commission guidelines
Guides/Risk classification ·By Yobel Tzegai ·Updated 20 August 2026

What happened when we ran our classification engine on itself?

Complipath's classification engine, run on Complipath itself, returned minimal risk for both systems in the run, answered one sentence that is false about itself, and could not ask the question that matters most. We publish all three outcomes verbatim, because a self-assessment you can check is worth more than a clean one you cannot.

Runs of 11 August 2026 through the real classify(), app repository commit 8bebc71. Inputs and outputs recorded in full in this page's claims file.

Quick answer

What exactly did we run?

The engine behind the product's guided classification is deterministic: fixed rules mapping confirmed answers to a classification with citations. We answered its six questions twice (the engine as it stood that day; it serves 7 today) — once describing the engine itself, once describing the draft generator — and ran the same classify() the workspace runs. No demo mode, no edited output; the claims file carries every answer we gave and every field that came back.

What did it say about the engine?

Risk level minimal. The reasoning began: "You answered that the system does not make, support or inform a decision about a person" — which is right: the engine classifies systems, not people.

Then it said something false: "because the model is trained in-house, also as provider". There is no model. The engine is hand-written rules over encoded legal knowledge, and its own vocabulary for "built in-house" assumes a trained model where none exists. We sent that finding into the product as a defect report instead of editing it out of this page — a self-run that only surfaced flattering output would not be evidence of anything.

What did it say about the draft generator?

Risk level minimal, with the role factor "Using a third-party model via API leaves you a deployer under Article 26." The role question for API-wrapped systems is one the engine's own coverage page lists as something it cannot decide for you where you have not said so — what this check can and cannot decide keeps that list, written once by a person, never generated per visitor.

What could it not decide?

Whether its subject is an AI system at all. Question zero is not among the engine's 7 questions — the decision tree says the same thing to every reader: every branch presumes the Article 3(1) definition is already met. Run on itself, the engine presupposes the answer to the one question this exercise was meant to illuminate. That is not a flaw discovered in embarrassment; it is the recorded boundary of the tool, and the run is evidence of the boundary, not an answer to the question. The definitional analysis — against Article 3, point (1), recital 12 and the Commission's guidelines — is its own page, and it reaches its conclusion by reading, not by running.

What this means for you

If you are a provider: this is what an honest self-assessment looks like at minimum — the real tool, the real inputs, the outputs kept even where they are wrong, and the boundary of the method stated. The same discipline applies to the assessment you record for your own systems: the reasoning is the evidence.

If you are a deployer: when a vendor hands you a clean self-assessment, ask for the inputs and for what the method cannot decide. Ours are published; the boundary is named; and the verdicts carry citations you can check against the articles they rest on.

What would it leave undetermined about yours?

Classify your system now — 7 questions on the main line, plus follow-ups where they apply, no account, and the classification runs in your browser: answers stay there unless you choose to keep the result.

FAQ

What did the engine conclude about itself? Minimal risk, on both runs — the engine itself and the draft generator. The verdicts, reasoning strings and citations are published verbatim in the claims file, together with every answer we gave, so the runs can be checked rather than taken on trust.

What did the engine get wrong? Its reasoning said "because the model is trained in-house" about a system that has no model — hand-written rules only. The vocabulary behind its build-type answer assumes a trained model where none exists. The finding went into the product as a defect report; the output stays published unedited.

What can a self-run not tell us? Whether the engine is an AI system at all. The engine's 7 questions never ask it — every run presupposes the Article 3(1) definition is met. That threshold question is answered by reading the Regulation and the Commission's guidelines, on a separate page, not by running the tool.

Why publish an unflattering result? Because a self-assessment you can verify beats a claim you cannot. A company selling classification that edits its own engine's errors out of its self-run has published marketing, not evidence. The uncomfortable output is the proof that the rest is real.


Sources: engine runs of 11 August 2026 against the app repository at commit 8bebc71 — full inputs and outputs in this page's claims file and in the claims file of is Complipath itself an AI system; Regulation (EU) 2024/1689 (EUR-Lex), Article 3, point (1), for the definition the runs presuppose.

← All guides
Complipath

Complipath is EU AI Act compliance software for AI-heavy software companies without a compliance team — an AI system register, deterministic risk classification, the obligations that follow, and the evidence behind every decision.

Complipath is built by Yobel Tzegai in Gothenburg, Sweden.

Complipath provides legal information, not legal advice. Every guide cites its source on EUR-Lex — Regulation (EU) 2024/1689, and Regulation (EU) 2026/1744 where that has amended it; where the law is still settling, the guide says so.

We measure page views with Vercel Web Analytics. It uses no third-party cookies. Visitors are identified by a hash derived from the incoming request, which is discarded after 24 hours, and no identifier is stored that could follow a visitor to another site. What is collected: the time of the visit, the URL, the referring page, filtered query parameters, city-level location, operating system, browser and device type.