CanLegal Bench logoCanLegal Bench
CanLegal Bench logoCanLegal Bench

A Canadian legal reasoning benchmark

How well do AI models reason about Canadian law?

CanLegal Bench is an open evaluation benchmark measuring the legal-reasoning competence of large language models across Canadian case law, statutes, and regulations — bilingual, bijural, and expert-curated.

We'll only email you about the CanLegal Bench release. Unsubscribe anytime. See our Privacy Policy.

Leaderboard

Frontier models, head-to-head

Each model answers every gold record closed-book — no tools, no retrieval. Suite mean is the unweighted average across all sixteen task types. Live results from the canonical results files.

Loading the leaderboard…

What it measures

A rigorous, Canada-first evaluation suite

CanLegal Bench tests the skills a legal professional actually relies on — statutory interpretation, outcome prediction, citation verification, and jurisdictional routing — across both of Canada's legal traditions and both official languages.

Bilingual & Bijural

Every task carries explicit English/French and common-law/civil-law dimensions — the legal duality that defines Canadian law and that no other benchmark captures.

Sixteen Task Types

From statute and regulation interpretation to case-outcome prediction, citation verification and lookup, issue spotting, and jurisdiction routing — a broad surface of legal-reasoning skills.

Expert-Curated Answers

Authoritative reference answers so scores reflect genuine legal competence, not annotation noise.

Adversarial Robustness

Perturbation and contamination safeguards expose whether a model truly reasons or merely pattern-matches against memorized text.

Frontier-Model Leaderboard

Leading frontier models evaluated head-to-head on identical, closed-book prompts — a clean, comparable measure of Canadian legal reasoning.

Closed-Book Protocol

No tools, no retrieval, no web access during evaluation — each model answers from its own learned knowledge, making results auditable and reproducible.

By the numbers

Built at scale

14,672
Expert-curated records
16
Task types
14
Jurisdictions
30%
French-language

Get CanLegal Bench release news first

The leaderboard above is live — the open dataset, paper, and public-repo launch are next. Straight to your inbox.

We'll only email you about the CanLegal Bench release. Unsubscribe anytime. See our Privacy Policy.