A Canadian legal reasoning benchmark
CanLegal Bench is an open evaluation benchmark measuring the legal-reasoning competence of large language models across Canadian case law, statutes, and regulations — bilingual, bijural, and expert-curated.
We'll only email you about the CanLegal Bench release. Unsubscribe anytime. See our Privacy Policy.
Leaderboard
Each model answers every gold record closed-book — no tools, no retrieval. Suite mean is the unweighted average across all sixteen task types. Live results from the canonical results files.
What it measures
CanLegal Bench tests the skills a legal professional actually relies on — statutory interpretation, outcome prediction, citation verification, and jurisdictional routing — across both of Canada's legal traditions and both official languages.
Every task carries explicit English/French and common-law/civil-law dimensions — the legal duality that defines Canadian law and that no other benchmark captures.
From statute and regulation interpretation to case-outcome prediction, citation verification and lookup, issue spotting, and jurisdiction routing — a broad surface of legal-reasoning skills.
Authoritative reference answers so scores reflect genuine legal competence, not annotation noise.
Perturbation and contamination safeguards expose whether a model truly reasons or merely pattern-matches against memorized text.
Leading frontier models evaluated head-to-head on identical, closed-book prompts — a clean, comparable measure of Canadian legal reasoning.
No tools, no retrieval, no web access during evaluation — each model answers from its own learned knowledge, making results auditable and reproducible.
By the numbers
The leaderboard above is live — the open dataset, paper, and public-repo launch are next. Straight to your inbox.
We'll only email you about the CanLegal Bench release. Unsubscribe anytime. See our Privacy Policy.