Checklist · Certification Software
A Buyer's Checklist for Evaluating AI-Assisted Certification Software
A vendor-neutral evaluation framework for Notified Bodies and Certification Bodies assessing AI-assisted review software: traceability, human oversight, data sovereignty, auditability, validation evidence and exit terms.
By Conformo Editorial Team · Published
Overview
Certification Bodies and Notified Bodies evaluating AI-assisted review software are buying under unusual conditions. The category is new, no accreditation scheme yet covers hybrid digital-human assessment models, vendor claims are difficult to verify from a demonstration, and the consequences of a poor choice fall on the organisation's designation rather than on its IT budget.
This page is a vendor-neutral evaluation framework. It is organised as questions to put to any vendor, with an explanation of what a substantive answer looks like and what a deflection looks like. We build in this category; the questions below are ones we think any serious vendor should be able to answer, including us.
Use it alongside essential features of certification body software, which covers the general functional requirements this page assumes.
Before you evaluate any vendor: fix your own position
Three internal decisions determine which answers are acceptable. Making them before the first demonstration prevents vendor framing from setting your requirements.
1. Which decisions will never be automated? Write this down. At minimum it includes the conformity decision itself and the final review under Section 4.7 of Annex VII. It probably also includes classification determinations, sufficiency judgements on clinical or performance evidence, and the acceptance of a manufacturer's justification for a non-applicable requirement. This list is the specification against which you evaluate every vendor claim about autonomy.
2. What will you tell your designating authority? You will have to describe this system in your quality management documentation, and you may be asked about it at a designation audit. If you cannot construct that description from the vendor's documentation, that is a finding about the vendor.
3. Where may manufacturers' technical documentation be processed? Decide this as a policy question before a vendor tells you what is possible. Your confidentiality obligations to manufacturers are yours regardless of your supply chain.
Section 1 — Traceability
Traceability is the foundational requirement, because it is what makes every other claim verifiable. A system whose outputs cannot be traced to source cannot support a defensible assessment record regardless of its accuracy.
Q1.1 — For any statement the system produces about a technical file, can a reviewer see the exact source?
A substantive answer means: the specific document, the version of that document, the page, and the passage. Not the file as a whole, and not a document title.
Deflection to watch for: a demonstration that shows a summary with a document name beside it. Ask to click through to the passage. Ask what happens when the passage is a table, a figure, or in an annex to an annex.
Q1.2 — Can the system produce a statement that is not grounded in a source passage?
This is an architectural question, not a quality question. Some systems constrain generation to retrieved passages; others generate freely and attach citations afterwards. The distinction determines whether unsupported statements are rare or structurally impossible.
Ask directly: is it architecturally possible for this system to state something about a file that is not in the file? A vendor who says no should be asked to explain the mechanism. A vendor who says it is possible but rare is being honest, and you should then ask how such cases are surfaced to the reviewer.
Q1.3 — Is every finding mapped to a specific regulatory provision?
Section-level, not theme-level. "MDR Annex II, Section 6.1(b)" is a mapping. "Clinical evidence" is a category. Reviewers work against provisions, and deficiency letters cite provisions.
Q1.4 — When the underlying document is updated, what happens to findings derived from the previous version?
Files change during assessment. A system that silently re-bases findings against new versions destroys the audit trail; one that ignores the update produces stale findings. Correct behaviour is explicit versioning with the reviewer notified of what changed and which findings are affected.
Section 2 — Human oversight
The TIC Council framework published in February 2026 requires human approval, input or supervision at defined decision points, with responsibility allocated to identified persons. These questions test whether that requirement is genuinely met or nominally satisfied.
Q2.1 — Which decision points require human action, and can the system be configured to bypass them?
If oversight can be turned off by a configuration setting, it is a default rather than a control. Ask who can change that setting and whether the change is logged.
Q2.2 — Is the reviewer shown evidence before conclusions, or conclusions before evidence?
This is the automation bias question, and it is the one most often left unasked. A reviewer presented with a confident machine-generated conclusion and an approve button is in a materially weaker position to exercise independent judgement than one presented with the evidence and asked to reach a conclusion.
Ask to see the actual reviewer interface, not a summary dashboard. Look at the relative cost of the accept path and the disagree path. If accepting is one click and disagreeing requires opening a form and typing a justification, the interface has a thumb on the scale.
Q2.3 — Is disagreement captured, and what is done with it?
Reviewer disagreement is the single most valuable quality signal a system of this kind produces. A vendor that captures it, reports it back to you, and can show you disagreement rates by finding type is measuring itself. One that discards it cannot tell you how well the system is performing in your environment.
Ask: what is your disagreement rate at comparable organisations, and can I see mine after deployment?
Q2.4 — Can an individual reviewer be identified for every accepted finding?
The responsibility-allocation dimension requires a named person, not a team or a role. Check that the record survives staff changes and that it is exportable.
Q2.5 — What competence does the system assume of the reviewer?
A tool designed on the assumption that a less-qualified reviewer can be supervised by the system has a different risk profile from one designed to make a qualified reviewer faster. Vendors are rarely explicit about this. Ask whether any customer uses the system to extend work to staff who would not otherwise be qualified to perform it, and consider carefully what the answer implies about your designation.
Section 3 — Data sovereignty and confidentiality
Technical documentation contains manufacturers' most commercially sensitive design, manufacturing and clinical data. Your confidentiality obligations under Annex VII do not transfer to a vendor.
Q3.1 — Where is data processed and stored, physically?
Name the country and the legal entity. "EU region" is a configuration label, not an answer — ask which Member State, which operator, and under whose jurisdiction the operating entity sits.
Q3.2 — Is any data transmitted to third-party model providers, and under what terms?
Most systems in this category use third-party foundation models. That is not disqualifying, but it must be disclosed. Ask: which providers, under what data processing terms, with what retention period, and in what jurisdiction. Ask whether the provider's terms permit human review of submitted data for abuse monitoring, and whether that has been contractually excluded.
Q3.3 — Is customer data used for model training or improvement, under any circumstances?
Accept only an unambiguous contractual answer. "We do not train on customer data by default" is a materially different statement from "we do not train on customer data."
Q3.4 — Can a specific manufacturer's data be isolated, exported and deleted on request?
You may need to demonstrate this to a manufacturer or a competent authority. Test it in a proof of concept rather than accepting a description.
Q3.5 — How is data segregated between manufacturers within your tenant?
A reviewer working on Manufacturer A's file must not see content retrieved from Manufacturer B's file. Ask how this is enforced technically, and whether it holds for the retrieval layer as well as the user interface — this is a real failure mode in systems that index across a shared corpus.
Section 4 — Auditability
The test: your designating authority asks how a certification decision taken eighteen months ago was reached. Can you reconstruct it completely?
Q4.1 — Is the assessment record complete and immutable?
Every AI-assisted output, every reviewer action, every source version, timestamped and tamper-evident. Ask whether records can be edited after the fact and whether edits are distinguishable from originals.
Q4.2 — Is the record exportable in a form that outlives the vendor relationship?
This is the question most often skipped. Certificates have multi-year validity, retention obligations extend beyond that, and vendors fail or get acquired. An export that is a proprietary archive restorable only into the vendor's platform is not an audit record.
Ask for a sample export from a real assessment and open it without vendor assistance.
Q4.3 — Can you reconstruct which version of the system produced a given output?
Models and prompts change. If a finding from March 2027 is questioned in 2029, you need to know what produced it. Ask whether system versions are recorded against findings, and whether the vendor retains the ability to explain the behaviour of superseded versions.
Q4.4 — What happens to the record when a reviewer overrides a finding?
Both the original output and the override should persist with reasoning. A record showing only the final state cannot demonstrate that oversight occurred.
Section 5 — Validation evidence
This is where the category is weakest, and where you should expect the most discomfort. There is no settled industry definition of what validating a probabilistic system for conformity assessment means. Ask anyway — the quality of a vendor's engagement with the question is informative even when the answer is incomplete.
Q5.1 — What performance evidence exists, measured how, on what data?
Ask for the evaluation methodology, the composition of the test set, and whether it was drawn from real technical files or constructed. Ask what "accuracy" denotes in their reporting — retrieval accuracy, finding precision and recall, and end-to-end assessment agreement are entirely different metrics and are frequently conflated.
Q5.2 — Where does performance degrade?
Every system has weak conditions: scanned documents, non-English files, very large files, unusual structures, tables, handwritten annotations, poorly OCR'd legacy submissions. A vendor who cannot name their failure modes either has not measured them or will not tell you. Both are informative.
Ask specifically about the document conditions in your actual intake, not the vendor's idealised corpus.
Q5.3 — Can you validate it yourself, on your own files, before committing?
Insist on a proof of concept using your files, scored by your reviewers, against assessments you have already completed. Retrospective testing against known outcomes is the only evaluation that reflects your document population and your reviewers' standards.
Define the success criteria before you start, and include a false-negative measure — a system that misses a real deficiency is far more dangerous to you than one that raises a spurious one.
Q5.4 — How are changes to the system managed and communicated?
Under a quality management system, a change to a tool used in conformity assessment is a change to your process. Ask about release cadence, advance notification, whether you can defer updates, and whether you can re-run validation against a new version before it reaches production.
Q5.5 — How does the vendor handle regulatory change?
When an implementing regulation changes requirements — as EU 2026/977 does from 25 February 2027 — how quickly does the system reflect it, who is accountable for the mapping being correct, and how are you notified?
Section 6 — Fit with your quality management system
Q6.1 — What documentation does the vendor provide for your QMS?
You need to describe this system in your own procedures. Ask what the vendor supplies: process descriptions, control descriptions, validation summaries, a statement of intended use and limitations. A vendor that has been through designation audits with customers will have this. One that has not will offer marketing material.
Q6.2 — Has the vendor been through a designation or accreditation audit with a customer, and what was raised?
Ask directly. A vendor with real deployments in designated bodies will have encountered assessor questions and should be willing to describe them, at least generically. This is one of the most efficient credibility filters available.
Q6.3 — How does the system handle the timeline and interruption requirements of EU 2026/977?
From 25 February 2027, Article 2 sets binding maximum timelines and Article 3 caps interruptions per phase. Article 4 requires monitoring, as part of your QMS, of the percentage of assessments completed within those timelines, median duration and median cost — published annually from 2028.
Ask whether the system records phase transitions as discrete timestamped events, tracks clock-stops against their per-phase caps, distinguishes Article 3(3) external-opinion interruptions (which are not counted), and produces the Article 4(2) metrics directly. A system that models assessment as opened-and-closed cannot produce these figures.
Section 7 — Commercial and exit terms
Q7.1 — What happens to your data and your assessment records if you terminate?
Format, timeframe, cost, and whether records remain readable without the vendor's software. Establish this before signature; leverage disappears afterwards.
Q7.2 — What are the dependencies on third parties, and what if one becomes unavailable?
If a foundation model provider changes terms, raises prices or withdraws a model, what is the vendor's contingency, and does it require you to re-validate?
Q7.3 — Does the pricing model penalise the behaviour you want?
Per-assessment pricing can discourage thorough use; per-seat pricing can discourage broad adoption. Check that the commercial structure does not conflict with your quality objectives.
Summary checklist
| # | Question | Acceptable answer |
|---|---|---|
| 1.1 | Source for every statement? | Document, version, page, passage |
| 1.2 | Ungrounded statements possible? | Architectural explanation, not reassurance |
| 1.3 | Findings mapped to provisions? | Section-level citation |
| 1.4 | Document version changes? | Explicit versioning, reviewer notified |
| 2.1 | Oversight bypassable? | Not by configuration; changes logged |
| 2.2 | Evidence before conclusions? | Verified in the real reviewer interface |
| 2.3 | Disagreement captured? | Captured, reported, measurable by you |
| 2.4 | Named reviewer per finding? | Individual, persistent, exportable |
| 2.5 | Assumed reviewer competence? | Explicit, and consistent with designation |
| 3.1 | Processing location? | Named country, entity, jurisdiction |
| 3.2 | Third-party model providers? | Named, with terms and retention |
| 3.3 | Training on your data? | Unambiguous contractual prohibition |
| 3.4 | Isolate, export, delete? | Demonstrated, not described |
| 3.5 | Manufacturer segregation? | Enforced at retrieval layer |
| 4.1 | Complete immutable record? | Tamper-evident, edits distinguishable |
| 4.2 | Vendor-independent export? | Opened without vendor assistance |
| 4.3 | System version per finding? | Recorded and explicable |
| 4.4 | Overrides preserved? | Original and override both retained |
| 5.1 | Performance evidence? | Methodology, test set, metric definitions |
| 5.2 | Known failure modes? | Named, matched to your intake |
| 5.3 | Validate on your files? | Retrospective PoC, criteria set in advance |
| 5.4 | Change management? | Notice, deferral, re-validation |
| 5.5 | Regulatory change? | Named accountability and timeline |
| 6.1 | QMS documentation? | Process, control and limitation statements |
| 6.2 | Survived a designation audit? | Specific, describable experience |
| 6.3 | EU 2026/977 support? | Phase events, capped clock-stops, Art. 4 metrics |
| 7.1 | Exit terms? | Readable records without vendor software |
| 7.2 | Third-party dependencies? | Contingency, re-validation implications |
| 7.3 | Pricing incentives? | Aligned with thorough use |
The two questions that matter most
If you take nothing else from this page: ask to click through from a finding to the exact passage that produced it, and run a retrospective proof of concept on your own completed assessments with success criteria you set in advance.
The first tests whether traceability is real. The second tests whether performance is real in your environment rather than the vendor's. Almost every serious problem in this category is detectable by one or the other, and neither can be satisfied with a slide.