Skip to content

Issue Spotting

Identify the legal issues a Canadian appellate court had to decide from the facts of the decision — testing issue recognition.

Judge-scored column, strict since 2026-10-05: a record counts when every key issue is identified; additional issues are not penalized. The judge matches the model's issues to the key issue by issue; the score counts the records where every key issue was found. Two diagnostics are shown beside each model: the exact-set-match rate (every key issue found and no extra items) and the lenient pass rate (set F1 ≥ 0.5).

438 records|5 models evaluated|Best: 58.7%|Worst: 39.3%
Model Performance (avg 47.6%)
1
Claude Fable 5
438 records · 257 correct
58.7%
every key issue · exact set 0.7% · lenient 55.3%
model →
2
GPT-5.6 Sol
438 records · 224 correct
51.1%
every key issue · exact set 7.5% · lenient 66.9%
model →
3
Grok 4.5
438 records · 209 correct
47.7%
every key issue · exact set 5.5% · lenient 64.4%
model →
4
Cohere Command A+
438 records · 181 correct
41.3%†
every key issue · exact set 1.6% · lenient 49.1%
model →
5
Gemini 3.1 Pro
438 records · 172 correct
39.3%‡
every key issue · exact set 9.6% · lenient 68.0%
model →

† the model's shared thinking and answer token budget ran out before it wrote any answer.‡ the provider's safety filter blocked the answer, or the answer came back empty with no recorded finish reason.These answers stay in the denominator and are scored 0; the count is disclosed per cell. Empty-output policy: an answer that came back empty is never removed from its cell; it counts as not correct against the frozen record count, and the number of such answers and their cause are published with the cell.

  • Cohere Command A+: † 28 answers hit the token budget before any answer token; scored 0
  • Gemini 3.1 Pro: ‡ 8 answers blocked by the provider's safety filter (empty output); scored 0