Skip to content
Read the original: Sherwin Wu· Published 62/100AI score62/100

Harvey LAB-AA v1.1 adds hallucination gate, reshaping legal benchmark rankings

Original titleExcited to see this updated version of LAB! The original LAB results left us scratching our heads.

AISummary

Artificial Analysis and Harvey released LAB-AA v1.1, which credits a legal task only when deliverables pass every rubric criterion with no material hallucinations.

Grok 4.7 (xhigh) leads at 9.4%, ahead of Muse Spark 1.3 (max) at 8.9% and GPT-6 Astra (max) at 8.6%, while over 60% of otherwise passing results contained a material hallucination.

The sharper reordering appears in the hallucination counts, where GPT-6 Astra averages 0.03 material hallucinations per task against 13.96 for Gemini 3.8 Flash (high).

Read the original x.com

Source: Sherwin Wu · x.comPublished · added here