Skip to content

GPT-5.6 Sol

17 of 17 task types · 14,547 records · 11,957 correct
78.1%
Suite Mean
unweighted avg of the headline tasks
legacy 11-task 84.5%
Evaluation Configuration
Reasoning
Effort-based — Reasoning effort set to xhigh — the maximum supported by this model. Thinking and answer share the output token budget.
Max output tokens32,768
TemperatureAPI default
NotesTemperature is not sent (model rejects it). API default applies.
Task Performance
TaskAccuracy
Cite Verify97.4%
Statute QA64.5%
strict · lenient 95.5%
Jurisdiction Router94.7%
Bijural Equivalence57.1%
strict · lenient 78.9%
Case Treatment96.7%
Case Outcome91.0%
Standard of Review96.8%
Doctrine Shift85.7%
Regulation QA61.2%
strict · lenient 61.2%
Tribunal QA89.6%
strict · lenient 89.6%
Multi-Doc Reasoning79.2%
Temporal Reasoning90.8%
Citation Lookup99.2%
Citation Holding37.6%†
strict · lenient 37.6%
Limitation Periods99.3%
Issue Spotting51.1%
every key issue · exact set 7.5% · lenient 66.9%
Citation Resolution36.0%†

† the model's shared thinking and answer token budget ran out before it wrote any answer.These answers stay in the denominator and are scored 0; the count is disclosed per cell. Empty-output policy: an answer that came back empty is never removed from its cell; it counts as not correct against the frozen record count, and the number of such answers and their cause are published with the cell.

  • Citation Holding: † 16 answers hit the token budget before any answer token; scored 0
  • Citation Resolution: † 211 answers hit the token budget before any answer token; scored 0