Skip to content
AEI Labs

Technical evaluation

Bring an answer you must defend.

A useful evaluation starts with a real decision, an approved source set and a failure condition. We demonstrate what the system answers, what it refuses and what evidence remains afterwards.

Evaluation path

Five questions before a pilot.

01 · Outcome

What decision will the answer support?

Define the user, task and consequence rather than beginning with a general chatbot brief.

02 · Corpus

Which sources are approved?

Identify ownership, versions, access boundaries and the conditions under which a source becomes current.

03 · Refusal

When must the system stay silent?

Agree the unsupported, contradictory and out-of-scope cases before measuring answer quality.

04 · Evidence

What must be inspectable later?

Specify the claim, citation, version, verdict and record needed for review.

05 · Success

How will we know it worked?

Use a test set and explicit thresholds for support, refusal, citation fidelity and reproducibility.