Skip to content

Case Outcome

Classify the outcome of an appeal or judicial review (allowed, allowed in part, dismissed) from the court's reasons, with its conclusion removed.

composition199 items: allowed 50, allowed in part 72, dismissed 77; 133 English and 66 French; 35 civil-law items. Courts: Ontario, Quebec, Alberta, British Columbia, Manitoba and Newfoundland and Labrador courts of appeal, the Federal Court of Appeal and the Federal Court; 172 appeals, 26 judicial reviews, 1 application. Each input is the facts and reasons of the decision with the conclusion, the order and every sentence stating the outcome removed. Keys confirmed by an AI council against the decisions (two members with verified quotes); 8 items the council found not answerable from their input were quarantined after the runs. (since 2026-10-07)

known cuesNo removal mark, paragraph number or outcome word is left in any input, and nothing about how the reasons were cut predicts the class: the best such feature, the input's length, reaches 39.7% against a 38.7% majority (each at or below the majority + 5 pp). Cues that remain are part of the reading, not leaks: how often the reasons use words of contrast or name an error in the decision below (46.2% and 43.2% alone), and whether the case is criminal (42.5%). (since 2026-10-07)

contaminationThe decisions are real, so a model may know some of them. In a 25-item memorization probe (the opening of the facts, answered from memory), each API row named the decision for 1 to 3 items (all four for one well-known Ontario Court of Appeal decision) and Command A+ for none; no row reproduced a decision's text (8-gram recall 0). (since 2026-10-07)

verified by an AI council (GPT-6 Astra, Gemini 3.8 Flash, Claude Opus 5.5, with a Claude Opus 5.5 adjudicator) against primary sourcesv3 answer keys verified on 2026-10-07 by an AI council (Gemini 3.8 Flash and Claude Opus 5.5 on every item; GPT-6 Astra on the 9 items sent to a second round, where members without a verified quote were re-asked once), each member given web search and the benchmark's stored copy of the decision; a key counts as confirmed only when at least two members back it with a machine-verified quote of the decision's outcome wording, and a key changes only on a verified quote, never by vote. Of 218 candidate keys: 207 confirmed, 0 changed, 11 quarantined on a verified contrary quote. A Claude Opus 5.5 adjudicator, without tools and from the item and the quoted texts only, then found 8 of the 207 not answerable from the input; they were quarantined (199 keys).

199 records|5 models evaluated|Best: 91.0%|Worst: 79.4%
Model Performance (avg 87.8%)
1
GPT-5.6 Sol
199 records · 181 correct
91.0%
model →
2
Grok 4.5
199 records · 180 correct
90.5%
model →
3
Claude Fable 5
199 records · 178 correct
89.4%†
model →
4
Gemini 3.1 Pro
199 records · 177 correct
88.9%
model →
5
Cohere Command A+
199 records · 158 correct
79.4%†
model →

† the model's shared thinking and answer token budget ran out before it wrote any answer.These answers stay in the denominator and are scored 0; the count is disclosed per cell. Empty-output policy: an answer that came back empty is never removed from its cell; it counts as not correct against the frozen record count, and the number of such answers and their cause are published with the cell.

  • Claude Fable 5: † 3 answers hit the token budget before any answer token; scored 0
  • Cohere Command A+: † 1 answer hit the token budget before any answer token; scored 0