Skip to content

Case Treatment

Classify how a later court treated an earlier decision, from a passage of the later decision (followed, considered, distinguished, criticized, questioned or overruled).

near ceilingNear ceiling: 4 of the 5 rows score 90% or more here (94.8%–96.7%; the other row scores 83.1%), so this column separates those rows little.

verified by an AI council (GPT-6 Astra, Gemini 3.8 Flash, Claude Opus 5.5, with a Claude Opus 5.5 adjudicator) against primary sourcesAnswer keys verified on 2026-10-05 by an AI council against primary sources: GPT-6 Astra, Gemini 3.8 Flash and Claude Opus 5.5 on every item (Grok 4.7 also on 43 items — 34 Regulation QA, 5 Bijural Equivalence, 4 Temporal Reasoning — until its cost guard stopped it), each member given web access and the stored source text where the benchmark holds one; a key counts as confirmed only when at least two members back it with a machine-verified quote, and disputes were ruled from the quoted texts by a Claude Opus 5.5 adjudicator. Of 389 keys: 338 confirmed, 2 changed, 49 quarantined. On 2026-10-07 the two-backer rule was applied to the keys the adjudicator had kept with fewer than two verified backers (missing members re-asked once): 7 quarantined; with the 26 passages that carry the Supreme Court Reports' own treatment label for the earlier case, 307 keys remain.

307 records|5 models evaluated|Best: 96.7%|Worst: 83.1%
Model Performance (avg 93.2%)
1
GPT-5.6 Sol
307 records · 297 correct
96.7%
model →
2
Grok 4.5
307 records · 296 correct
96.4%
model →
3
Claude Fable 5
307 records · 291 correct
94.8%
model →
4
Gemini 3.1 Pro
307 records · 291 correct
94.8%
model →
5
Cohere Command A+
307 records · 255 correct
83.1%
model →