Skip to content

Citation Resolution

Name the Supreme Court of Canada decision a citation refers to (style of cause + year) — deterministic scoring; a declined answer earns no credit, like a wrong one.

scoring policyHeadline = accuracy; a decline earns no credit; the net score (+1/0/−1) is a secondary hallucination-aware diagnostic. The accuracy cell counts a declined answer as not correct, exactly like a wrong one; the net score (correct +1; a decline, or a record that returned no answer, 0; a wrong case or none named −1) is computed beside it only to show which rows guess rather than decline. (since 2026-10-04)

known scorer limitA distinctive party name translated or abbreviated other than by its initials, a given name omitted or spelled out, a dropped union-local number and a generic-only short form still score wrong case — 11 records over the four API rows in a hand read of 2026-10-05 (at most 0.33 pp per row) — and a few evasive non-answers without decline wording still score wrong case. (since 2026-10-05)

scoring policySince scorer 2026-10-07 an answer that says it has no access to the decision and names no case except placeholder parties from a fixed list ("Appellant v. Respondent", "Plaintiff v. Defendant", "Smith v. Jones", "A v. B", "X v. Y", "Party A v. Party B", "[Name] v. [Name]", « Appelant c. Intimé », « X c. Y », …) or a name that could be real ("R. v. Smith") introduced as a format example is a decline (0 in the net score), not a wrong case (−1); any other case named keeps it an attempt. Every row was re-scored from its stored answers and accuracy did not change: command-a-plus 67 rows reclassified (net −0.325 → −0.281; its archived temperature-0 cell 30, −0.153 → −0.133), gemini-3.1-pro-preview 2 (−0.212 → −0.211), claude-fable-5, gpt-5.6-sol and grok-4.5 none. (since 2026-10-07)

scoring policySince scorer 2026-10-08 the bare abstention words that also describe a decision — "ambiguous", "insufficient evidence", "insufficient information" — make an answer a decline only when it names no case: an answer that names the case and says the Court "held that an ambiguous clause was unenforceable" or found "insufficient evidence of discrimination" is an answer, not a decline. The identity-directed decline patterns and the other abstention words are unchanged, so a hedge still scores as a decline. Every row was re-scored from its stored answers: 3 records move decline → correct (claude-fable-5 44.27% → 44.40%, 2; gemini-3.1-pro-preview 33.20% → 33.27%, 1); the other rows unchanged. (since 2026-10-08)

1500 records|5 models evaluated|Best: 44.4%|Worst: 0.0%
Model Performance (avg 24.7%)
1
Claude Fable 5
1500 records · 666 correct
44.4%†
model →
2
GPT-5.6 Sol
1500 records · 540 correct
36.0%†
model →
3
Gemini 3.1 Pro
1500 records · 499 correct
33.3%
model →
4
Grok 4.5
1500 records · 150 correct
10.0%
model →
5
Cohere Command A+
1500 records · 0 correct
0.0%†
model →

† the model's shared thinking and answer token budget ran out before it wrote any answer.These answers stay in the denominator and are scored 0; the count is disclosed per cell. Empty-output policy: an answer that came back empty is never removed from its cell; it counts as not correct against the frozen record count, and the number of such answers and their cause are published with the cell.

  • Claude Fable 5: † 202 answers hit the token budget before any answer token; scored 0
  • GPT-5.6 Sol: † 211 answers hit the token budget before any answer token; scored 0
  • Cohere Command A+: † 143 answers hit the token budget before any answer token; scored 0