Skip to content
transparency

Scoring Transparency

Everything that determines a published number, in one place: the scoring version, the exact scorer that produced each task's cells, the uncertainty on suite means and per-task ranks, and the failure/capping accounting. Nothing on this page changes a score — it discloses how the scores on the leaderboard are produced and where their limits are.

scoring version

version date: 2026-10-08
latest changelog entry: v083-open-release-keys

The score changelog is append-only — score-affecting changes are appended with a date and id, never edited in place. Methodology: About →

judge & scorer identities

TaskScorerType
Cite Verifyneutral_cite_verify_score@2026-09-17deterministic
Statute QAqwen3.8-27b-i53f6qqg-semantic-full-v2LLM judge
Jurisdiction Routerletter_answer_process@2026-07-20deterministic
Bijural Equivalenceqwen3.8-27b-i53f6qqg-semantic-full-v2LLM judge
Case Treatmentcase_treatment_score@2026-09-17deterministic
Case Outcomecase_outcome_score@2026-10-07deterministic
Standard of Reviewstandard_of_review_score@2026-10-05deterministic
Doctrine Shiftdoctrine_shift_v2_score@2026-08-28deterministic
Regulation QAqwen3.8-27b-i53f6qqg-regv2-element-v1LLM judge
Tribunal QAqwen3.8-27b-tribv32-disposition-v1LLM judge
Multi-Doc Reasoningmulti_doc_v2_score@2026-10-04deterministic
Temporal Reasoningtemporal_reasoning_score@2026-10-07deterministic
Citation Lookupcitation_lookup_score@2026-10-08deterministic
Citation Holdingclaude-sonnet-5-semantic-v2LLM judge
Limitation Periodslimitation_period_calc_score@2026-10-08deterministic
Issue Spottingclaude-sonnet-5-issue-set-v1LLM judge
Citation Resolutioncitation_resolution_score@2026-10-08deterministic

statusLLM-judge human validation: PENDING— judge outputs have not yet been human-audited; treat judge-scored cells with appropriate caution.

Scorer identities are read from the per-cell scorer fields of the published leaderboard data; every model's cell for a task shares one scorer.

suite-mean uncertainty

ModelSuite MeanBootstrap 95% CI± half-width
Claude Fable 580.1%[78.9%–81.3%]±1.2pp
GPT-5.6 Sol78.1%[76.9%–79.3%]±1.2pp
Gemini 3.1 Pro77.5%[76.4%–78.6%]±1.1pp
Grok 4.571.6%[70.3%–72.9%]±1.3pp
Cohere Command A+49.5%[47.9%–51.1%]±1.6pp

Method: parametric-binomial-bootstrap — 100,000 replicates, fixed seed; exact for binary-scored cells (every cell is exact-match or an LLM-judge cell binarized at its strict verdict 1.0). A closed-form normal-approximation cross-check is computed with it (generated 2026-10-08T17:06:45Z).

Reading: overlapping CIs mean the suite-mean gap between those models is not statistically separable at 95% confidence — the ranking within an overlapping band is provisional, even though the point estimates are exact.

per-task n and confidence intervals

Taskn (frozen)Model accuracy rangeWidest CI ±Ranks
Cite Verify49554.4%–97.8%±4.4ppresolvable
Statute QA5579.9%–81.0%±4.1ppresolvable
Jurisdiction Router5775.4%–94.7%±10.9pplow-n · ranks not resolvable
Bijural Equivalence17533.1%–57.1%±7.3ppresolvable
Case Treatment30783.1%–96.7%±4.2ppresolvable
Case Outcome19979.4%–91.0%±5.6ppresolvable
Standard of Review18668.3%–98.9%±6.6ppresolvable
Doctrine Shift7055.7%–87.1%±11.3pplow-n · ranks not resolvable
Regulation QA23212.9%–62.9%±6.4ppresolvable
Tribunal QA7771.4%–90.9%±9.9ppresolvable
Multi-Doc Reasoning13063.1%–90.0%±8.2ppresolvable
Temporal Reasoning21871.1%–91.7%±6.0ppresolvable
Citation Lookup7,91163.4%–99.2%±1.1ppresolvable
Citation Holding1,3951.1%–56.0%±2.6ppresolvable
Limitation Periods60058.2%–100.0%±3.9ppresolvable
Issue Spotting43839.3%–58.7%±4.7ppresolvable
Citation Resolution1,5000.0%–44.4%±2.5ppresolvable

A task's ranks are resolvable when its widest per-cell Wilson 95% CI half-width is at most ±10pp; above that, the ordering is within noise and rank badges are de-emphasized across the site (scores always stay visible). Per-cell CIs are published in the API (ci95).

network failures

ModelNetwork failures (all tasks)
Claude Fable 50
GPT-5.6 Sol0
Gemini 3.1 Pro0
Grok 4.50
Cohere Command A+0

Policy: a network failure (empty output from an API timeout or infra error) counts as WRONG against the frozen denominator — failures are never silently dropped, so no model benefits from infra luck. Per-cell failure counts are disclosed via the API (networkFailures).

capped cells

No cells are currently clamped by the score-invariants guard.

The score-invariants guard clamps every served cell to [0, 1] and to n ≤ the frozen denominator, catching stale pre-quarantine record sets and corrupt upstream values before they reach the board. When a clamp fires, the cell is flagged (capped + capReason in the API) — flags are disclosed, never hidden.

known limitations

  • Small-n tasks. Jurisdiction Router n=57 (±10.9pp), Doctrine Shift n=70 (±11.3pp), Tribunal QA n=77 (±9.9pp), Multi-Doc Reasoning n=130 (±8.2pp), Bijural Equivalence n=175 (±7.3pp), Standard of Review n=186 (±6.6pp), Case Outcome n=199 (±5.6pp) — per-task ranks on these tasks carry wide uncertainty. Currently over the ±10pp rank-resolvability threshold (rank badges de-emphasized site-wide): Jurisdiction Router (±10.9pp), Doctrine Shift (±11.3pp).
  • Doctrine Shift is small. Its v2 key is 70 records (English and French versions of each item, 4 classes — the retired v1 was 105 English-only records; 6 records with wrong shift-direction gold were quarantined 2026-09-09 and 6 more by the AI council on 2026-10-05), and the four API rows sit in a 80.0%–87.1% band with intervals up to ±11.3pp; treat differences there as low-signal.
  • Contamination screening: LOW for all five public rows. The 3,901-record v0.5.x manifest screened LOW for gold memorization (2026-08-20 verbatim-recall probe, all 7 board models), the full 14,470-record v0.8.0 corpus screened LOW on 2026-08-28, and the 2,927 records added after it screened LOW on 2026-10-06 (gold-memorization probe on the four API rows + provenance audit + web-verbatim spot checks), so every record id of the 14,584-record frozen manifest of that date has been screened (the 6 items re-admitted later that day after their answer keys were re-drafted and re-verified were in quarantine during that screen; the Multi-Doc Reasoning items re-cut on 2026-10-06 carry the screen of the items they rebuild). Command A+ (self-hosted) was probed the same day in its own serving window, on the same probe set. The 199 Case Outcome v3 records that replaced that column's key on 2026-10-07 were screened the same day (provenance of all 207 built ids; a 25-item memorization probe on all five public rows: LOW; no row reproduced a decision's text).
  • LLM-judge human validation is pending (see judge & scorer identities above) — judge-scored free-text cells have not yet been human-audited.
  • The suite mean is an unweighted macro-average — small-n tasks carry equal weight (1/17 of the headline suite mean) by design, so a noisy small task moves the mean as much as the 7,911-record one. Rationale and weighting discussion: About →.
CanLegal Bench· scoring version 2026-10-08 · latest score change: v083-open-release-keys