GPT-5.6 Sol
17 of 17 task types · 14,547 records · 11,957 correct
78.1%
Suite Mean
unweighted avg of the headline tasks
legacy 11-task 84.5%
Evaluation Configuration
Reasoning
Effort-based — Reasoning effort set to xhigh — the maximum supported by this model. Thinking and answer share the output token budget.
Max output tokens32,768
TemperatureAPI default
NotesTemperature is not sent (model rejects it). API default applies.
Task Performance
| Task | Accuracy | |
|---|---|---|
| Cite Verify | 97.4% | |
| Statute QA | 64.5% strict · lenient 95.5% | |
| Jurisdiction Router | 94.7% | |
| Bijural Equivalence | 57.1% strict · lenient 78.9% | |
| Case Treatment | 96.7% | |
| Case Outcome | 91.0% | |
| Standard of Review | 96.8% | |
| Doctrine Shift | 85.7% | |
| Regulation QA | 61.2% strict · lenient 61.2% | |
| Tribunal QA | 89.6% strict · lenient 89.6% | |
| Multi-Doc Reasoning | 79.2% | |
| Temporal Reasoning | 90.8% | |
| Citation Lookup | 99.2% | |
| Citation Holding | 37.6%† strict · lenient 37.6% | |
| Limitation Periods | 99.3% | |
| Issue Spotting | 51.1% every key issue · exact set 7.5% · lenient 66.9% | |
| Citation Resolution | 36.0%† |
† the model's shared thinking and answer token budget ran out before it wrote any answer.These answers stay in the denominator and are scored 0; the count is disclosed per cell. Empty-output policy: an answer that came back empty is never removed from its cell; it counts as not correct against the frozen record count, and the number of such answers and their cause are published with the cell.
- Citation Holding: † 16 answers hit the token budget before any answer token; scored 0
- Citation Resolution: † 211 answers hit the token budget before any answer token; scored 0