Skip to content
Read the original: Artificial Analysis Articles· Published 50/100AI score50/100

Harvey LAB-AA v1.1 adds hallucination checks to legal AI benchmark

Original titleAnnouncing Harvey LAB-AA v1.1: adding hallucination checks to raise the bar for agentic legal work

AISummary

Harvey LAB-AA v1.1 adds hallucination checks that audit every model deliverable against task source documents, with material hallucinations zeroing a task's score. GPT-6 Astra averaged 0.03 material hallucinations per task across 120 tasks, while Gemini 3.8 Flash averaged 13.96. Harvey uses GPT-6 Sol (high) as the hallucination checker, separate from its three-judge rubric panel.

Read the original artificialanalysis.ai

Source: Artificial Analysis Articles · artificialanalysis.aiPublished · added here