Skip to content

Standard of Review

Identify the standard of review a court applied (reasonableness, correctness, or palpable and overriding error).

near ceilingNear ceiling: 4 of the 5 rows score 90% or more here (96.8%–98.9%; the other row scores 68.3%), so this column separates those rows little.

verified by an AI council (GPT-6 Astra, Gemini 3.8 Flash, Claude Opus 5.5, with a Claude Opus 5.5 adjudicator) against primary sourcesAnswer keys verified on 2026-10-05 by an AI council against primary sources: GPT-6 Astra, Gemini 3.8 Flash and Claude Opus 5.5 on every item (Grok 4.7 also on 43 items — 34 Regulation QA, 5 Bijural Equivalence, 4 Temporal Reasoning — until its cost guard stopped it), each member given web access and the stored source text where the benchmark holds one; a key counts as confirmed only when at least two members back it with a machine-verified quote, and disputes were ruled from the quoted texts by a Claude Opus 5.5 adjudicator. Of 210 keys: 190 confirmed, 0 changed, 20 quarantined. On 2026-10-07 the two-backer rule was applied to the keys the adjudicator had kept with fewer than two verified backers (missing members re-asked once): 4 quarantined (186 keys).

186 records|5 models evaluated|Best: 98.9%|Worst: 68.3%
Model Performance (avg 92.0%)
1
Claude Fable 5
186 records · 184 correct
98.9%
model →
2
Grok 4.5
186 records · 183 correct
98.4%
model →
3
Gemini 3.1 Pro
186 records · 182 correct
97.8%
model →
4
GPT-5.6 Sol
186 records · 180 correct
96.8%
model →
5
Cohere Command A+
186 records · 127 correct
68.3%
model →