Skip to content

Statute QA

Answer questions about specific sections of Canadian statutes — testing precise statutory recall across 51+ Acts.

Judge-scored column: scores are strict accuracy — a record counts only when the judge gives it full credit (verdict 1.0) — since 2026-10-05. The lenient pass rate (verdict ≥ 0.5, half credit counted) is shown beside each model.

557 records|5 models evaluated|Best: 81.0%|Worst: 9.9%
Model Performance (avg 54.2%)
1
Claude Fable 5
557 records · 451 correct
81.0%
strict · lenient 93.4%
model →
2
GPT-5.6 Sol
557 records · 359 correct
64.5%
strict · lenient 95.5%
model →
3
Grok 4.5
557 records · 328 correct
58.9%
strict · lenient 95.5%
model →
4
Gemini 3.1 Pro
557 records · 317 correct
56.9%‡
strict · lenient 89.0%
model →
5
Cohere Command A+
557 records · 55 correct
9.9%†
strict · lenient 56.6%
model →

† the model's shared thinking and answer token budget ran out before it wrote any answer.‡ the provider's safety filter blocked the answer, or the answer came back empty with no recorded finish reason.These answers stay in the denominator and are scored 0; the count is disclosed per cell. Empty-output policy: an answer that came back empty is never removed from its cell; it counts as not correct against the frozen record count, and the number of such answers and their cause are published with the cell.

  • Gemini 3.1 Pro: ‡ 26 answers blocked by the provider's safety filter (empty output); scored 0
  • Cohere Command A+: † 19 answers hit the token budget before any answer token; scored 0