Citation Resolution
Name the Supreme Court of Canada decision a citation refers to (style of cause + year) — deterministic scoring; a declined answer earns no credit, like a wrong one.
scoring policyHeadline = accuracy; a decline earns no credit; the net score (+1/0/−1) is a secondary hallucination-aware diagnostic. The accuracy cell counts a declined answer as not correct, exactly like a wrong one; the net score (correct +1; a decline, or a record that returned no answer, 0; a wrong case or none named −1) is computed beside it only to show which rows guess rather than decline. (since 2026-10-04)
known scorer limitA distinctive party name translated or abbreviated other than by its initials, a given name omitted or spelled out, a dropped union-local number and a generic-only short form still score wrong case — 11 records over the four API rows in a hand read of 2026-10-05 (at most 0.33 pp per row) — and a few evasive non-answers without decline wording still score wrong case. (since 2026-10-05)
scoring policySince scorer 2026-10-07 an answer that says it has no access to the decision and names no case except placeholder parties from a fixed list ("Appellant v. Respondent", "Plaintiff v. Defendant", "Smith v. Jones", "A v. B", "X v. Y", "Party A v. Party B", "[Name] v. [Name]", « Appelant c. Intimé », « X c. Y », …) or a name that could be real ("R. v. Smith") introduced as a format example is a decline (0 in the net score), not a wrong case (−1); any other case named keeps it an attempt. Every row was re-scored from its stored answers and accuracy did not change: command-a-plus 67 rows reclassified (net −0.325 → −0.281; its archived temperature-0 cell 30, −0.153 → −0.133), gemini-3.1-pro-preview 2 (−0.212 → −0.211), claude-fable-5, gpt-5.6-sol and grok-4.5 none. (since 2026-10-07)
scoring policySince scorer 2026-10-08 the bare abstention words that also describe a decision — "ambiguous", "insufficient evidence", "insufficient information" — make an answer a decline only when it names no case: an answer that names the case and says the Court "held that an ambiguous clause was unenforceable" or found "insufficient evidence of discrimination" is an answer, not a decline. The identity-directed decline patterns and the other abstention words are unchanged, so a hedge still scores as a decline. Every row was re-scored from its stored answers: 3 records move decline → correct (claude-fable-5 44.27% → 44.40%, 2; gemini-3.1-pro-preview 33.20% → 33.27%, 1); the other rows unchanged. (since 2026-10-08)
† the model's shared thinking and answer token budget ran out before it wrote any answer.These answers stay in the denominator and are scored 0; the count is disclosed per cell. Empty-output policy: an answer that came back empty is never removed from its cell; it counts as not correct against the frozen record count, and the number of such answers and their cause are published with the cell.
- Claude Fable 5: † 202 answers hit the token budget before any answer token; scored 0
- GPT-5.6 Sol: † 211 answers hit the token budget before any answer token; scored 0
- Cohere Command A+: † 143 answers hit the token budget before any answer token; scored 0