Tribunal QA
From a Canadian administrative tribunal's reasons, with its conclusion and order removed, state how the tribunal disposed of the matter — scored by an LLM judge.
Judge-scored column: scores are strict accuracy — a record counts only when the judge gives it full credit (verdict 1.0) — since 2026-10-05. The lenient pass rate (verdict ≥ 0.5, half credit counted) is shown beside each model.
verified by an AI council (GPT-6 Astra, Gemini 3.8 Flash, Claude Opus 5.5, with a Claude Opus 5.5 adjudicator) against primary sourcesAnswer keys verified on 2026-10-05 by an AI council against primary sources: GPT-6 Astra, Gemini 3.8 Flash and Claude Opus 5.5 on every item (Grok 4.7 also on 43 items — 34 Regulation QA, 5 Bijural Equivalence, 4 Temporal Reasoning — until its cost guard stopped it), each member given web access and the stored source text where the benchmark holds one; a key counts as confirmed only when at least two members back it with a machine-verified quote, and disputes were ruled from the quoted texts by a Claude Opus 5.5 adjudicator. Of 91 keys: 76 confirmed, 0 changed, 15 quarantined (3 of them keys the council found wrong, set aside for a re-draft); on 2026-10-06 1 of those 3 was re-drafted, re-verified and re-admitted (77 keys), and the other 2 stay quarantined.