Domain 2 — Output Evaluation and Validation
CCA Associate Foundations course · Page 3 of 10 · ← Back to all courses · Weight: 21% — the heaviest domain. Master this one first. Mark this page complete at the bottom to advance your course progress.
2.1 — Measuring Accuracy: The Right Metric for the Right Failure · Core
You cannot measure accuracy without a labeled dataset — inputs with known-correct answers. Run this eval at two moments: before first launch, and after every model update (a model swap can change behavior even with nothing else touched).
Match the metric to the failure — never substitute one for another:
| Failure | Metric |
|---|---|
| Wrong or ungrounded answers | Accuracy against a labeled dataset |
| Slow responses | Latency percentiles (p95/p99) — never the average |
| Budget overruns | Cost per request / tokens per task |
| Harmful output | Safety violation rate, measured on adversarial test sets (a benign test set proves nothing) |
Operational metrics ≠ quality metrics. A system can be fast, cheap, and error-free at the infrastructure level while still returning factually wrong content — uptime and latency don't check correctness; that has to be tracked separately.
⚠ Often-missed — Aggregate Accuracy Hides Segment Failure · Gap
A 96% overall accuracy figure can conceal 60% accuracy on one specific document type or field. Always validate accuracy by segment (document type, language, field) before reducing human review — reducing review based on the aggregate number alone is the single most common exam trap in this domain. Coverage annotations (WELL_SUPPORTED / PARTIAL / COVERAGE_GAP) exist precisely to expose this at the finding level rather than the report level.
2.2 — Hallucination vs. Retrieval Failure vs. Prompt Failure · Core
A hallucination is confident, fluent content with no basis in any source — and it sounds exactly as convincing as a correct answer. The diagnostic that the exam tests hardest:
| Symptom | Cause | Fix |
|---|---|---|
| Correct documents retrieved, answer still wrong | Hallucination / weak grounding | Tighten grounding instructions; require citations to the retrieved text |
| Wrong documents retrieved | Retrieval failure | Fix chunking, embeddings, or query construction |
| Fails only on complex, multi-step tasks | Model mismatch | Test on a higher-capability tier |
| Fails uniformly across every input type | Prompt failure | Rewrite criteria to be explicit and testable |
"Correct docs retrieved, wrong answer" is not a retrieval problem — the retrieval succeeded. This is the most common wrong-answer distractor on the exam: don't reach for chunking or embedding fixes when the actual documents were fine.
The fix for a hallucination is grounding — never a stronger instruction. "Only report figures in the database" does not stop fabrication; requiring the model to retrieve and cite a specific record does.
2.3 — Fact-Checking at Scale: The Five-Stage Eval Workflow · Core
- Define the task — specific, measurable pass criteria.
- Build the golden dataset — real inputs including edge cases and adversarial examples; a dataset of only clean inputs predicts nothing about production.
- Run automated (code-based) checks — schema validity, required fields, exact matches. Cheapest, fastest, never drifts.
- Score with an LLM-as-judge — for behaviors needing interpretation (tone, helpfulness). The judge must be a different model than the one being evaluated, or self-preference bias inflates scores. Requires a written rubric.
- Interpret and act — aggregate scores and break them down by category.
Never tune your prompt against the same set you use to measure it — this is eval contamination: scores climb (91%) while production accuracy collapses (68%) because the prompt learned the test cases, not the task. Thresholds come from the business requirement, not from whatever the first prototype happened to score.
Judge calibration is not optional. Judge scores drift over time; periodically sample against human ratings and rewrite the rubric where they diverge. An uncalibrated judge is worse than no automated grade — it looks authoritative while being wrong.
Two credible sources disagree on a number? Include both values with their sources. Never pick one, never average them — an average is a number neither source actually reported.
2.4 — When Human Review Is Required · Core
Valid escalation triggers: low-confidence extraction (below a calibrated threshold), an ambiguous or contradictory source document, the customer explicitly asks for a human, or the request falls into a genuine policy gap.
Invalid triggers — never route on these: customer sentiment/frustration level, or the model's own self-reported confidence score. Both are documented as unreliable — an agent is frequently confidently wrong precisely on the hardest cases.
Confidence scores must be calibrated against a labeled validation set before they're used for routing at all — an uncalibrated score may route everything to review, or nothing. Even after calibration, apply stratified random sampling to the auto-accepted (high-confidence) items on an ongoing basis — this is what catches a new document format quietly producing high-confidence wrong answers.
2.5 — Choosing the Right Output Format · Core
| Content type | Format |
|---|---|
| Financial data, comparisons | Tables |
| News, narrative | Prose |
| Technical findings, steps | Structured lists |
| Output feeding another system/agent | Structured format with metadata (source, date, confidence) |
When synthesizing multiple sources, every claim needs a claim → source mapping: URL/document name, excerpt, and publication date. A handoff payload to a human reviewer must include the input, Claude's output, the confidence score, and the specific routing reason — the reviewer should never have to reconstruct context from scratch.
Exam reflexes for Domain 2
- "Correct documents retrieved, answer still wrong" → hallucination/grounding failure, not retrieval.
- "Fix a hallucination" → grounding (retrieval, citations, verification) — never a stronger instruction.
- "97% aggregate accuracy, team removes review" → wrong — validate by segment first.
- "Judge and system-under-test are the same model" → self-preference bias; use a separate judge model.
- "Prompt tuned against the eval set, scores drop in production" → eval contamination.
- "Customer is frustrated / agent reports low confidence" → not valid escalation triggers.
- "Two sources give different numbers" → show both with attribution; never average or pick one.
- "Financial comparison output" → table. "News summary" → prose.
Test yourself on this domain. Take the Domain 2 practice quiz — 38 questions, instant scoring, an explanation for every answer.