Skip to content

Citation Lookup

Verify whether a Canadian legal citation refers to a real court or tribunal decision — a hallucination probe across courts and years.

scoring policySince scorer 2026-10-08 an explicit answer line ("Answer:", "Final answer:", "**Answer**", « Réponse : ») that leads with a verdict (yes / no / oui / non / real / fabricated among its first three words: "No — …", "Likely no — …") decides; otherwise the first verdict word of the answer decides, as before (no negation handling: "not a real decision" still reads as "real"). Before, the first verdict word always decided, so an answer that reasoned first ("the court token is real") and then answered "**Answer:** Likely no" was read as "real". Every row was re-scored from its stored answers: claude-fable-5 93.45% → 93.95% (40 records 0 → 1, 1 record 1 → 0 whose answer line says "Likely yes"); the other four rows unchanged. (since 2026-10-08)

7911 records|5 models evaluated|Best: 99.2%|Worst: 63.4%
Model Performance (avg 85.9%)
1
GPT-5.6 Sol
7911 records · 7849 correct
99.2%
model →
2
Gemini 3.1 Pro
7911 records · 7690 correct
97.2%
model →
3
Claude Fable 5
7911 records · 7432 correct
93.9%
model →
4
Grok 4.5
7911 records · 6008 correct
75.9%
model →
5
Cohere Command A+
7911 records · 5014 correct
63.4%†
model →

† the model's shared thinking and answer token budget ran out before it wrote any answer.These answers stay in the denominator and are scored 0; the count is disclosed per cell. Empty-output policy: an answer that came back empty is never removed from its cell; it counts as not correct against the frozen record count, and the number of such answers and their cause are published with the cell.

  • Cohere Command A+: † 18 answers hit the token budget before any answer token; scored 0