Grounded AI evaluation scorecard
Freeze the candidate, questions, expected answers and scoring rules before running a comparison. Use the same model settings where the product permits.
| Question | Method | Expected and selected record | Identity exact? | Supported claims | Contradicted claims | Unsupported claims | Citation usable? | Context bytes or tokens | Latency and cost | Notes |
|---|
Report serious failures as well as totals. Separate retrieval, compactness, attribution, faithfulness, domain correctness, safety and operational behaviour. Do not claim that results generalise beyond the frozen bundle, questions, model, prompt, date and access conditions actually tested.