CanLegal Bench
A closed-book benchmark evaluating frontier AI models on Canadian legal reasoning.
14,547 bilingual gold records across 17 task types, spanning English and French common-law/civil-law domains. Every record is drawn from Canadian legal materials and validated for structural and semantic correctness before inclusion.
Two different counts, not to be confused: 14,547 is the evaluated gold corpus (17 tasks, per-task n=57–7,911); 7,945 is our internal automated test suite validating the scoring code — models are evaluated on the gold records, not on the software tests.
tasks
Cite Verify
Statute QA
Jurisdiction Router
Bijural Equivalence
Case Treatment
Case Outcome
Standard of Review
Doctrine Shift
Regulation QA
Tribunal QA
Multi-Doc Reasoning
Temporal Reasoning
Citation Lookup
Citation Holding
Limitation Periods
Issue Spotting
Citation Resolution
per-task record counts
| Task | Records (n) | Note |
|---|---|---|
| Cite Verify | 495 | |
| Statute QA | 557 | |
| Jurisdiction Router | 57 | small-n |
| Bijural Equivalence | 175 | |
| Case Treatment | 307 | |
| Case Outcome | 199 | |
| Standard of Review | 186 | |
| Doctrine Shift | 70 | small-n |
| Regulation QA | 232 | |
| Tribunal QA | 77 | small-n |
| Multi-Doc Reasoning | 130 | |
| Temporal Reasoning | 218 | |
| Citation Lookup | 7,911 | |
| Citation Holding | 1,395 | |
| Limitation Periods | 600 | |
| Issue Spotting | 438 | |
| Citation Resolution | 1,500 |
Task sizes are deliberately heterogeneous. citation_lookup alone is 7,911 of the frozen 14,547 records (54%) — a deterministic existence check — but the suite mean is an unweighted macro mean over all 17 tasks, so that task contributes 1/17 of the score, not 54%. Rows flagged small-n (n<100) have wide Wilson 95% confidence intervals; per-cell CIs are published in the API (ci95).
scoring
Structured tasks
Free-text tasks (statute, tribunal, regulation, bijural, citation holding, issue spotting)
Suite mean composition — all 17 tasks
Row status — all public rows complete
Cross-jurisdiction probes
Scorer identities, per-cell uncertainty, and known limitations are disclosed on the Transparency page.