Skip to content

Multi-Doc Reasoning

Read a passage of an earlier decision and a passage of a later decision that cites it (treatment words removed) and classify how the later court treated the earlier one: followed, distinguished, not followed or cited only — testing multi-document legal reasoning.

known cuesPassage B carries no treatment word, removal mark, ellipsis or paragraph number; in the shortcut audit of the 154-item set (2026-10-06) no feature of how it was cut predicted the label better than the majority + 5 pp. Two kinds of cue remain and are part of the task, not leaks: how often the later decision names or cites the earlier one (the evidence the item is about), and the court mix of the selected pairs (on the current 130 items, 58 of the 66 cited-only items come from the Supreme Court and 29 of the 42 Federal Court items are followed; the majority class, cited only, is 50.8%). (since 2026-10-07)

composition130 items: cited only 66, followed 44, distinguished 17, not followed 3; 92 English and 38 French; 19 civil-law items. Keys verified by an AI council against primary sources (2026-10-05); on 2026-10-07, 23 items whose passage B still stated the treatment and 1 item the council found not answerable from its input were quarantined. (since 2026-10-07)

verified by an AI council (GPT-6 Astra, Gemini 3.8 Flash, Claude Opus 5.5, with a Claude Opus 5.5 adjudicator) against primary sourcesAnswer keys verified on 2026-10-05 by an AI council against primary sources: GPT-6 Astra, Gemini 3.8 Flash and Claude Opus 5.5 on every item (Grok 4.7 also on 43 items — 34 Regulation QA, 5 Bijural Equivalence, 4 Temporal Reasoning — until its cost guard stopped it), each member given web access and the stored source text where the benchmark holds one; a key counts as confirmed only when at least two members back it with a machine-verified quote, and disputes were ruled from the quoted texts by a Claude Opus 5.5 adjudicator. Of 166 keys: 157 confirmed, 0 changed, 9 quarantined. On 2026-10-06 the v2.1 rebuild set aside 3 thin items (154), and on 2026-10-07 the 23 items whose passage B still stated the treatment were quarantined (131 keys). A council check the same day confirmed one more item's key but found the item not answerable from its input, so it was quarantined too (130 keys).

130 records|5 models evaluated|Best: 90.0%|Worst: 63.1%
Model Performance (avg 77.5%)
1
Gemini 3.1 Pro
130 records · 117 correct
90.0%
model →
2
Claude Fable 5
130 records · 104 correct
80.0%†
model →
3
GPT-5.6 Sol
130 records · 103 correct
79.2%
model →
4
Grok 4.5
130 records · 98 correct
75.4%
model →
5
Cohere Command A+
130 records · 82 correct
63.1%
model →

† the model's shared thinking and answer token budget ran out before it wrote any answer.These answers stay in the denominator and are scored 0; the count is disclosed per cell. Empty-output policy: an answer that came back empty is never removed from its cell; it counts as not correct against the frozen record count, and the number of such answers and their cause are published with the cell.

  • Claude Fable 5: † 1 answer hit the token budget before any answer token; scored 0