Skip to content

Citation Holding

Given a case citation, state what the court decided — the holding — free-text answer scored by an LLM judge against the gold answer.

Judge-scored column: scores are strict accuracy — a record counts only when the judge gives it full credit (verdict 1.0) — since 2026-10-05. The lenient pass rate (verdict ≥ 0.5, half credit counted) is shown beside each model.

1395 records|5 models evaluated|Best: 56.0%|Worst: 1.1%
Model Performance (avg 36.6%)
1
Claude Fable 5
1395 records · 781 correct
56.0%
strict · lenient 56.0%
model →
2
Gemini 3.1 Pro
1395 records · 695 correct
49.8%
strict · lenient 49.8%
model →
3
Grok 4.5
1395 records · 535 correct
38.4%
strict · lenient 38.4%
model →
4
GPT-5.6 Sol
1395 records · 525 correct
37.6%†
strict · lenient 37.6%
model →
5
Cohere Command A+
1395 records · 16 correct
1.1%†
strict · lenient 1.1%
model →

† the model's shared thinking and answer token budget ran out before it wrote any answer.These answers stay in the denominator and are scored 0; the count is disclosed per cell. Empty-output policy: an answer that came back empty is never removed from its cell; it counts as not correct against the frozen record count, and the number of such answers and their cause are published with the cell.

  • GPT-5.6 Sol: † 16 answers hit the token budget before any answer token; scored 0
  • Cohere Command A+: † 18 answers hit the token budget before any answer token; scored 0