Skip to content
about

CanLegal Bench

A closed-book benchmark evaluating frontier AI models on Canadian legal reasoning.

14,547 bilingual gold records across 17 task types, spanning English and French common-law/civil-law domains. Every record is drawn from Canadian legal materials and validated for structural and semantic correctness before inclusion.

version
v0.8.3
records
14,547
task types
17
french
30.8%
software tests
7,945 pytest

Two different counts, not to be confused: 14,547 is the evaluated gold corpus (17 tasks, per-task n=57–7,911); 7,945 is our internal automated test suite validating the scoring code — models are evaluated on the gold records, not on the software tests.

tasks

Cite Verify

Real vs. fabricated legal citations
495 records

Statute QA

Statutory provision recall (judge-scored free text)
557 records

Jurisdiction Router

Which order of government has jurisdiction (federal, provincial, shared, federal paramountcy)
57 records

Bijural Equivalence

Civil law ↔ common law concept mapping
175 records

Case Treatment

Judicial treatment classification (six labels, from followed to overruled)
307 records

Case Outcome

Appeal and judicial-review outcome from the reasons, conclusion removed (allowed, allowed in part, dismissed)
199 records

Standard of Review

Standard of review identification (reasonableness, correctness, palpable and overriding error)
186 records

Doctrine Shift

4-class doctrine-shift detection between two authorities (expanded, narrowed, reversed, unchanged)
70 records

Regulation QA

Federal regulation provision recall (key elements; judge-scored)
232 records

Tribunal QA

Tribunal outcome from the reasons, conclusion and order removed (judge-scored)
77 records

Multi-Doc Reasoning

Two-passage treatment classification (earlier and later decision)
130 records

Temporal Reasoning

Date-aware statute and SCC decision questions (assent/in-force/amendment years, decision dates, before/after ordering)
218 records

Citation Lookup

Citation existence verification (hallucination probe)
7911 records

Citation Holding

SCC holding recall (judge-scored free text)
1395 records

Limitation Periods

Limitation-period date calculation
600 records

Issue Spotting

Appellate issue identification (judge-scored recall)
438 records

Citation Resolution

SCC case-name resolution from a citation (deterministic, abstention-aware)
1500 records

per-task record counts

TaskRecords (n)Note
Cite Verify495
Statute QA557
Jurisdiction Router57small-n
Bijural Equivalence175
Case Treatment307
Case Outcome199
Standard of Review186
Doctrine Shift70small-n
Regulation QA232
Tribunal QA77small-n
Multi-Doc Reasoning130
Temporal Reasoning218
Citation Lookup7,911
Citation Holding1,395
Limitation Periods600
Issue Spotting438
Citation Resolution1,500

Task sizes are deliberately heterogeneous. citation_lookup alone is 7,911 of the frozen 14,547 records (54%) — a deterministic existence check — but the suite mean is an unweighted macro mean over all 17 tasks, so that task contributes 1/17 of the score, not 54%. Rows flagged small-n (n<100) have wide Wilson 95% confidence intervals; per-cell CIs are published in the API (ci95).

scoring

Structured tasks

Exact-match with alias-aware scoring. Gold records carry bilingual aliases for cross-language synonym matching (e.g. granted ↔ allowed, raisonnabilité ↔ reasonableness). Word-overlap fallback at threshold 0.45 for multi-word answers.

Free-text tasks (statute, tribunal, regulation, bijural, citation holding, issue spotting)

A semantic judge scores each answer by whether it matches the gold answer in meaning — not by token overlap, so a concise correct answer scores the same as a verbose one. Since 2026-10-05 the published score is strict accuracy: a record counts as correct only when the judge gives it full credit (verdict 1.0). Half credit (0.5) is not counted. For Issue Spotting the judge matches the model's issues to the court's key issues one by one: a record counts when every key issue is identified; additional issues are not penalized. The exact-set-match rate (every key issue found and no extra items) is published beside it as a diagnostic. The lenient pass rate (verdict ≥ 0.5, half credit counted; for Issue Spotting, a set F1 ≥ 0.5), the headline until 2026-10-04, is shown beside each judged cell. Strict is harder. Citation-anchor grounding is computed internally as a supplementary signal and does not affect the published score.

Suite mean composition — all 17 tasks

The headline suite mean is an unweighted macro-average over all 17 launched tasks; no column is excluded. Citation Resolution, launched on 2026-10-04, joined it on 2026-10-05 once every public row had a cell (Command A+ last). Its cell is accuracy: a declined answer earns no credit, like a wrong one. Since 2026-10-07 Command A+ is scored like its other columns: a 65,536-token pass, then one 131,072-token re-run of every answer that ran out of budget (the last answer counts); the 143 answers that ran out of budget both times while reasoning count as not correct (marked †). Citation Resolution is not part of the legacy 11-task mean. Tribunal QA, Multi-Doc Reasoning and Standard of Review returned to the headline mean on 2026-10-05 after their rebuilds, once Command A+ had a cell on each rebuilt key. Tribunal QA asks how a tribunal decided from reasons whose concluding paragraphs and order are hidden (77 items after the AI council; about a third are partial outcomes, scored together with their claims) — a small set with a wide confidence interval; all five rows were judged in one pass by one judge. Multi-Doc Reasoning was rebuilt from independent passages (two passages from two decisions, keyed by the reporters' treatment designation); on 2026-10-06 its second passage was re-cut as whole sentences so that no mark shows where a treatment sentence was taken out, and on 2026-10-07 the items whose second passage still stated the treatment were quarantined (130 items). Standard of Review was rebuilt with balanced classes (210 items keyed by the court's own stated standard; 186 after the AI council); four rows score near the ceiling, so it separates them weakly. Regulation QA returned to the headline mean on 2026-10-05: its earlier answers had been scored on truncated text, so it was rebuilt as a 232-item key and all five public rows were re-run on it. The earlier 11-task mean, which includes Regulation QA and Tribunal QA and leaves out Case Treatment, Case Outcome, Standard of Review, Multi-Doc Reasoning, Issue Spotting and Citation Resolution, stays available as a secondary legacy figure frozen at its 2026-09-29 composition and definition, so it reads judge-scored columns at their lenient pass rate (suiteMeanLegacy11 in the API); it does not set the ranking. Case Outcome and Case Treatment returned to the headline mean on 2026-10-04 after two independent AI readers checked their answer keys; on 2026-10-07 Case Outcome was rebuilt as v3 (199 items: the facts and reasons of appellate and judicial-review decisions with the conclusion removed; allowed, allowed in part or dismissed), and Case Treatment has 307 items after the inputs that stated their own answer were removed. On 2026-10-05 an AI council (GPT-6 Astra, Gemini 3.8 Flash and Claude Opus 5.5, with a Claude Opus 5.5 adjudicator; Grok 4.7 also on 43 items) verified the answer keys of these six columns and Jurisdiction Router against primary sources, replacing lawyer review: 4 keys were changed on verified quotes and 154 items were quarantined (14,738 → 14,584 records); on 2026-10-06, 6 of the keys it had set aside were re-drafted, re-verified and re-admitted; on 2026-10-07, 167 more items were quarantined (keys without two verified backers, and inputs that stated their own answer) and Case Outcome v3 replaced that column's key: 14,547 records today. Issue Spotting returned to the headline mean on 2026-10-04: every public row is now judged on the rebuilt answer lists with the set-level rubric under one judge. A counted task with no data scores 0.0 — every row is divided by the same 17 tasks, never by whichever tasks happen to have data.

Row status — all public rows complete

A row counts as complete only when every launched task carries at least 90% of its frozen record count; otherwise it is flagged partial. Every public row is complete and has a cell on all 17 launched tasks. A column a row had not been run on would read pending run (a rebuilt column it was not re-run on, pending re-run), never 0%, and would stay out of the headline mean until every public row had a cell. Until 2026-09-30 four rows were partial — claude-opus-4-8, gemini-3.1-pro-preview, gpt-5.5 and grok-4.3 — because their Case Treatment cells were scored on the earlier pre-rebuild version of that column (478 scored records against the rebuilt 683-record denominator). On 2026-09-30 Gemini 3.1 Pro was re-run on its five older-gold tasks on the current record sets, and the other three rows were retired from the public board (see the model inclusion policy below). A row below 90% completion on any launched task is never shown on this site as a final score.

Cross-jurisdiction probes

64 probes across 22 legal dimensions (aboriginal rights, administrative law, class actions, contract law, corporate law, court hierarchy, criminal procedure, environmental law, evidence, family law, good faith, labour law, liability framework, precedent weight, prescription, privacy law, property, securities regulation, source of law, succession law, tax law, trusts) testing civil-law vs common-law reasoning correctness. An LLM judge evaluates PASS/FAIL/PARTIAL against red-flag patterns of wrong-system reasoning. This is an auxiliary probe set, not one of the 17 scored tasks; its results are not published in this release.

Scorer identities, per-cell uncertainty, and known limitations are disclosed on the Transparency page.

policies

Model inclusion policy + roadmap

The public frontier board shows the newest row per lab — Claude Fable 5, GPT-5.6 Sol, Gemini 3.1 Pro and Grok 4.5 — plus Cohere Command A+ (re-added 2026-09-19). New models are added as new rows between phases, never swapped mid-phase, so rankings stay comparable within a phase. On 2026-09-30 the owner retired Claude Opus 4.8, GPT-5.5, Grok 4.3 from the public board — retired 2026-09-30: superseded models whose cells were scored on older gold / stale generations; not re-run; their data is kept, labelled retired. Further models may be added as new rows in a later phase.

Abstention policy

Gold records are constructed to be answerable — every record has a verifiable correct answer, so a refusal on an answerable item scores 0 by design (the suite measures task competence, not calibration). Behaviour on genuinely unanswerable inputs is not part of the suite score.

Contamination screening

The retired 3,901-record v0.5.x manifest was screened on 2026-08-20 (verbatim-recall probe: LOW memorization signal for all 7 board models), the full 14,470-record v0.8.0 corpus on 2026-08-28 (gold-memorization probe + provenance audit + web-verbatim spot checks: LOW signal across all three methods), and the 2,927 records added after it on 2026-10-06 (the same three methods: LOW signal), so every record id of the 14,584-record frozen manifest of that date has been screened (as of 2026-10-06); the 6 items re-admitted later that day after their answer keys were re-drafted and re-verified were in quarantine during that screen. The Multi-Doc Reasoning items re-cut that day (130 remain) carry the screen of the items they rebuild. The 199 Case Outcome v3 records that replaced that column's key on 2026-10-07 were screened the same day (provenance of all 207 built ids; a 25-item memorization probe on all five public rows: LOW; no row reproduced a decision's text). The 2026-10-06 gold-memorization probe covered all five public rows, Command A+ in its own serving window the same day: LOW for every row. The screening dates and results are summarized on the Transparency page.

infrastructure

corpus: Canadian legal documents across 14 jurisdictions · federal + provincial + territorial
protocol: closed-book evaluation · no tool use · no RAG injection
validation: automated schema validation · 7,945 regression tests · audited gold (2026-09-08 and 2026-09-29 audits; defective records quarantined, not deleted) · answer keys verified by an AI council against primary sources
CanLegal Bench· scoring version 2026-10-08 · latest score change: v083-open-release-keys
Victoria, BC · Canada