Skip to content

Tribunal QA

From a Canadian administrative tribunal's reasons, with its conclusion and order removed, state how the tribunal disposed of the matter — scored by an LLM judge.

Judge-scored column: scores are strict accuracy — a record counts only when the judge gives it full credit (verdict 1.0) — since 2026-10-05. The lenient pass rate (verdict ≥ 0.5, half credit counted) is shown beside each model.

verified by an AI council (GPT-6 Astra, Gemini 3.8 Flash, Claude Opus 5.5, with a Claude Opus 5.5 adjudicator) against primary sourcesAnswer keys verified on 2026-10-05 by an AI council against primary sources: GPT-6 Astra, Gemini 3.8 Flash and Claude Opus 5.5 on every item (Grok 4.7 also on 43 items — 34 Regulation QA, 5 Bijural Equivalence, 4 Temporal Reasoning — until its cost guard stopped it), each member given web access and the stored source text where the benchmark holds one; a key counts as confirmed only when at least two members back it with a machine-verified quote, and disputes were ruled from the quoted texts by a Claude Opus 5.5 adjudicator. Of 91 keys: 76 confirmed, 0 changed, 15 quarantined (3 of them keys the council found wrong, set aside for a re-draft); on 2026-10-06 1 of those 3 was re-drafted, re-verified and re-admitted (77 keys), and the other 2 stay quarantined.

77 records|5 models evaluated|Best: 90.9%|Worst: 71.4%
Model Performance (avg 85.2%)
1
Gemini 3.1 Pro
77 records · 70 correct
90.9%
strict · lenient 90.9%
model →
2
GPT-5.6 Sol
77 records · 69 correct
89.6%
strict · lenient 89.6%
model →
3
Claude Fable 5
77 records · 68 correct
88.3%
strict · lenient 88.3%
model →
4
Grok 4.5
77 records · 66 correct
85.7%
strict · lenient 85.7%
model →
5
Cohere Command A+
77 records · 55 correct
71.4%
strict · lenient 71.4%
model →