Gemini 3.1 Pro
17 of 17 task types · 14,547 records · 11,817 correct
77.5%
Suite Mean
unweighted avg of the headline tasks
legacy 11-task 84.5%
Evaluation Configuration
Reasoning
Always-on — Built-in thinking is always active. No toggle — the model reasons on every query.
Max output tokens32,768
TemperatureAPI default
NotesTemperature is not sent. API default (1.0) applies.
Task Performance
| Task | Accuracy | |
|---|---|---|
| Cite Verify | 95.5% | |
| Statute QA | 56.9%‡ strict · lenient 89.0% | |
| Jurisdiction Router | 94.7% | |
| Bijural Equivalence | 56.0% strict · lenient 79.4% | |
| Case Treatment | 94.8% | |
| Case Outcome | 88.9% | |
| Standard of Review | 97.8% | |
| Doctrine Shift | 87.1% | |
| Regulation QA | 57.3% strict · lenient 57.3% | |
| Tribunal QA | 90.9% strict · lenient 90.9% | |
| Multi-Doc Reasoning | 90.0% | |
| Temporal Reasoning | 88.5% | |
| Citation Lookup | 97.2% | |
| Citation Holding | 49.8% strict · lenient 49.8% | |
| Limitation Periods | 99.5% | |
| Issue Spotting | 39.3%‡ every key issue · exact set 9.6% · lenient 68.0% | |
| Citation Resolution | 33.3% |
‡ the provider's safety filter blocked the answer, or the answer came back empty with no recorded finish reason.These answers stay in the denominator and are scored 0; the count is disclosed per cell. Empty-output policy: an answer that came back empty is never removed from its cell; it counts as not correct against the frozen record count, and the number of such answers and their cause are published with the cell.
- Statute QA: ‡ 26 answers blocked by the provider's safety filter (empty output); scored 0
- Issue Spotting: ‡ 8 answers blocked by the provider's safety filter (empty output); scored 0