Scoring Transparency
Everything that determines a published number, in one place: the scoring version, the exact scorer that produced each task's cells, the uncertainty on suite means and per-task ranks, and the failure/capping accounting. Nothing on this page changes a score — it discloses how the scores on the leaderboard are produced and where their limits are.
scoring version
The score changelog is append-only — score-affecting changes are appended with a date and id, never edited in place. Methodology: About →
judge & scorer identities
| Task | Scorer | Type |
|---|---|---|
| Cite Verify | neutral_cite_verify_score@2026-09-17 | deterministic |
| Statute QA | qwen3.8-27b-i53f6qqg-semantic-full-v2 | LLM judge |
| Jurisdiction Router | letter_answer_process@2026-07-20 | deterministic |
| Bijural Equivalence | qwen3.8-27b-i53f6qqg-semantic-full-v2 | LLM judge |
| Case Treatment | case_treatment_score@2026-09-17 | deterministic |
| Case Outcome | case_outcome_score@2026-10-07 | deterministic |
| Standard of Review | standard_of_review_score@2026-10-05 | deterministic |
| Doctrine Shift | doctrine_shift_v2_score@2026-08-28 | deterministic |
| Regulation QA | qwen3.8-27b-i53f6qqg-regv2-element-v1 | LLM judge |
| Tribunal QA | qwen3.8-27b-tribv32-disposition-v1 | LLM judge |
| Multi-Doc Reasoning | multi_doc_v2_score@2026-10-04 | deterministic |
| Temporal Reasoning | temporal_reasoning_score@2026-10-07 | deterministic |
| Citation Lookup | citation_lookup_score@2026-10-08 | deterministic |
| Citation Holding | claude-sonnet-5-semantic-v2 | LLM judge |
| Limitation Periods | limitation_period_calc_score@2026-10-08 | deterministic |
| Issue Spotting | claude-sonnet-5-issue-set-v1 | LLM judge |
| Citation Resolution | citation_resolution_score@2026-10-08 | deterministic |
statusLLM-judge human validation: PENDING— judge outputs have not yet been human-audited; treat judge-scored cells with appropriate caution.
Scorer identities are read from the per-cell scorer fields of the published leaderboard data; every model's cell for a task shares one scorer.
suite-mean uncertainty
| Model | Suite Mean | Bootstrap 95% CI | ± half-width |
|---|---|---|---|
| Claude Fable 5 | 80.1% | [78.9%–81.3%] | ±1.2pp |
| GPT-5.6 Sol | 78.1% | [76.9%–79.3%] | ±1.2pp |
| Gemini 3.1 Pro | 77.5% | [76.4%–78.6%] | ±1.1pp |
| Grok 4.5 | 71.6% | [70.3%–72.9%] | ±1.3pp |
| Cohere Command A+ | 49.5% | [47.9%–51.1%] | ±1.6pp |
Method: parametric-binomial-bootstrap — 100,000 replicates, fixed seed; exact for binary-scored cells (every cell is exact-match or an LLM-judge cell binarized at its strict verdict 1.0). A closed-form normal-approximation cross-check is computed with it (generated 2026-10-08T17:06:45Z).
Reading: overlapping CIs mean the suite-mean gap between those models is not statistically separable at 95% confidence — the ranking within an overlapping band is provisional, even though the point estimates are exact.
per-task n and confidence intervals
| Task | n (frozen) | Model accuracy range | Widest CI ± | Ranks |
|---|---|---|---|---|
| Cite Verify | 495 | 54.4%–97.8% | ±4.4pp | resolvable |
| Statute QA | 557 | 9.9%–81.0% | ±4.1pp | resolvable |
| Jurisdiction Router | 57 | 75.4%–94.7% | ±10.9pp | low-n · ranks not resolvable |
| Bijural Equivalence | 175 | 33.1%–57.1% | ±7.3pp | resolvable |
| Case Treatment | 307 | 83.1%–96.7% | ±4.2pp | resolvable |
| Case Outcome | 199 | 79.4%–91.0% | ±5.6pp | resolvable |
| Standard of Review | 186 | 68.3%–98.9% | ±6.6pp | resolvable |
| Doctrine Shift | 70 | 55.7%–87.1% | ±11.3pp | low-n · ranks not resolvable |
| Regulation QA | 232 | 12.9%–62.9% | ±6.4pp | resolvable |
| Tribunal QA | 77 | 71.4%–90.9% | ±9.9pp | resolvable |
| Multi-Doc Reasoning | 130 | 63.1%–90.0% | ±8.2pp | resolvable |
| Temporal Reasoning | 218 | 71.1%–91.7% | ±6.0pp | resolvable |
| Citation Lookup | 7,911 | 63.4%–99.2% | ±1.1pp | resolvable |
| Citation Holding | 1,395 | 1.1%–56.0% | ±2.6pp | resolvable |
| Limitation Periods | 600 | 58.2%–100.0% | ±3.9pp | resolvable |
| Issue Spotting | 438 | 39.3%–58.7% | ±4.7pp | resolvable |
| Citation Resolution | 1,500 | 0.0%–44.4% | ±2.5pp | resolvable |
A task's ranks are resolvable when its widest per-cell Wilson 95% CI half-width is at most ±10pp; above that, the ordering is within noise and rank badges are de-emphasized across the site (scores always stay visible). Per-cell CIs are published in the API (ci95).
network failures
| Model | Network failures (all tasks) |
|---|---|
| Claude Fable 5 | 0 |
| GPT-5.6 Sol | 0 |
| Gemini 3.1 Pro | 0 |
| Grok 4.5 | 0 |
| Cohere Command A+ | 0 |
Policy: a network failure (empty output from an API timeout or infra error) counts as WRONG against the frozen denominator — failures are never silently dropped, so no model benefits from infra luck. Per-cell failure counts are disclosed via the API (networkFailures).
capped cells
No cells are currently clamped by the score-invariants guard.
The score-invariants guard clamps every served cell to [0, 1] and to n ≤ the frozen denominator, catching stale pre-quarantine record sets and corrupt upstream values before they reach the board. When a clamp fires, the cell is flagged (capped + capReason in the API) — flags are disclosed, never hidden.
known limitations
- Small-n tasks. Jurisdiction Router n=57 (±10.9pp), Doctrine Shift n=70 (±11.3pp), Tribunal QA n=77 (±9.9pp), Multi-Doc Reasoning n=130 (±8.2pp), Bijural Equivalence n=175 (±7.3pp), Standard of Review n=186 (±6.6pp), Case Outcome n=199 (±5.6pp) — per-task ranks on these tasks carry wide uncertainty. Currently over the ±10pp rank-resolvability threshold (rank badges de-emphasized site-wide): Jurisdiction Router (±10.9pp), Doctrine Shift (±11.3pp).
- Doctrine Shift is small. Its v2 key is 70 records (English and French versions of each item, 4 classes — the retired v1 was 105 English-only records; 6 records with wrong shift-direction gold were quarantined 2026-09-09 and 6 more by the AI council on 2026-10-05), and the four API rows sit in a 80.0%–87.1% band with intervals up to ±11.3pp; treat differences there as low-signal.
- Contamination screening: LOW for all five public rows. The 3,901-record v0.5.x manifest screened LOW for gold memorization (2026-08-20 verbatim-recall probe, all 7 board models), the full 14,470-record v0.8.0 corpus screened LOW on 2026-08-28, and the 2,927 records added after it screened LOW on 2026-10-06 (gold-memorization probe on the four API rows + provenance audit + web-verbatim spot checks), so every record id of the 14,584-record frozen manifest of that date has been screened (the 6 items re-admitted later that day after their answer keys were re-drafted and re-verified were in quarantine during that screen; the Multi-Doc Reasoning items re-cut on 2026-10-06 carry the screen of the items they rebuild). Command A+ (self-hosted) was probed the same day in its own serving window, on the same probe set. The 199 Case Outcome v3 records that replaced that column's key on 2026-10-07 were screened the same day (provenance of all 207 built ids; a 25-item memorization probe on all five public rows: LOW; no row reproduced a decision's text).
- LLM-judge human validation is pending (see judge & scorer identities above) — judge-scored free-text cells have not yet been human-audited.
- The suite mean is an unweighted macro-average — small-n tasks carry equal weight (1/17 of the headline suite mean) by design, so a noisy small task moves the mean as much as the 7,911-record one. Rationale and weighting discussion: About →.