Skip to content

Regulation QA

Recall what a named provision of a Canadian federal regulation requires — the actor, number, period, condition, permission or prohibition — scored by an LLM judge.

Judge-scored column: scores are strict accuracy — a record counts only when the judge gives it full credit (verdict 1.0) — since 2026-10-05. The lenient pass rate (verdict ≥ 0.5, half credit counted) is shown beside each model.

verified by an AI council (GPT-6 Astra, Gemini 3.8 Flash, Claude Opus 5.5, Grok 4.7 on 34 items, with a Claude Opus 5.5 adjudicator) against primary sourcesAnswer keys verified on 2026-10-05 by an AI council against primary sources: GPT-6 Astra, Gemini 3.8 Flash and Claude Opus 5.5 on every item (Grok 4.7 also on 43 items — 34 Regulation QA, 5 Bijural Equivalence, 4 Temporal Reasoning — until its cost guard stopped it), each member given web access and the stored source text where the benchmark holds one; a key counts as confirmed only when at least two members back it with a machine-verified quote, and disputes were ruled from the quoted texts by a Claude Opus 5.5 adjudicator. Of 232 keys: 227 confirmed, 0 changed, 5 found wrong or incomplete and set aside; on 2026-10-06 those 5 were re-drafted through the column's build gates, re-verified by the council with verified quotes and re-admitted (232 keys).

232 records|5 models evaluated|Best: 62.9%|Worst: 12.9%
Model Performance (avg 48.3%)
1
Claude Fable 5
232 records · 146 correct
62.9%†‡
strict · lenient 62.9%
model →
2
GPT-5.6 Sol
232 records · 142 correct
61.2%
strict · lenient 61.2%
model →
3
Gemini 3.1 Pro
232 records · 133 correct
57.3%
strict · lenient 57.3%
model →
4
Grok 4.5
232 records · 109 correct
47.0%
strict · lenient 47.0%
model →
5
Cohere Command A+
232 records · 30 correct
12.9%†
strict · lenient 12.9%
model →

† the model's shared thinking and answer token budget ran out before it wrote any answer.‡ the provider's safety filter blocked the answer, or the answer came back empty with no recorded finish reason.These answers stay in the denominator and are scored 0; the count is disclosed per cell. Empty-output policy: an answer that came back empty is never removed from its cell; it counts as not correct against the frozen record count, and the number of such answers and their cause are published with the cell.

  • Claude Fable 5: † 1 answer hit the token budget before any answer token; scored 0 · ‡ 1 answer blocked by the provider's safety filter (empty output); scored 0
  • Cohere Command A+: † 4 answers hit the token budget before any answer token; scored 0