AI Evaluation & Eval-Driven Development · 1 of 3
Eval Terms & Metrics
AtomicReps30″Navigation
Evaluation as a first-class engineering discipline: trace review, error analysis, golden datasets, LLM-as-judge with judge alignment, CI eval gates, and eval validity. The canonical methodology home for LLM evaluation (answer-level RAG eval lives here; component retrieval metrics live in rag_systems).
AI Evaluation & Eval-Driven Development · 1 of 3
Eval Terms & Metrics
AI Evaluation & Eval-Driven Development · 2 of 3
Offline vs Online Evaluation
AI Evaluation & Eval-Driven Development · 3 of 3
Offline vs Online Evaluation
function evalGate(results) { const passRate = results.filter(r => r.score >= 0.7).length / results.length; const anyUnsafe = results.some(r => r.unsafe); if (anyUnsafe) return "BLOCK"; return passRate >= 0.9 ? "SHIP" : "HOLD";}Three of the 100 AI Evaluation & Eval-Driven Development questions in the bank.
Keep going with AI Evaluation & Eval-Driven Development, free