Skip to content

Cite Verify

Classify Canadian legal citations as real, fabricated, or ambiguous — testing citation knowledge across 12 jurisdictions.

near ceilingNear ceiling: 4 of the 5 rows score 90% or more here (93.3%–97.8%; the other row scores 54.4%), so this column separates those rows little.

495 records|5 models evaluated|Best: 97.8%|Worst: 54.4%
Model Performance (avg 87.7%)
1
Claude Fable 5
495 records · 493 scored (2 ambiguous keys excluded) · 482 correct
97.8%
model →
2
GPT-5.6 Sol
495 records · 493 scored (2 ambiguous keys excluded) · 480 correct
97.4%
model →
3
Gemini 3.1 Pro
495 records · 493 scored (2 ambiguous keys excluded) · 471 correct
95.5%
model →
4
Grok 4.5
495 records · 493 scored (2 ambiguous keys excluded) · 460 correct
93.3%
model →
5
Cohere Command A+
495 records · 493 scored (2 ambiguous keys excluded) · 268 correct
54.4%†
model →

† the model's shared thinking and answer token budget ran out before it wrote any answer.These answers stay in the denominator and are scored 0; the count is disclosed per cell. Empty-output policy: an answer that came back empty is never removed from its cell; it counts as not correct against the frozen record count, and the number of such answers and their cause are published with the cell.

  • Cohere Command A+: † 3 answers hit the token budget before any answer token; scored 0