Skip to content

Gemini 3.1 Pro

17 of 17 task types · 14,547 records · 11,817 correct
77.5%
Suite Mean
unweighted avg of the headline tasks
legacy 11-task 84.5%
Evaluation Configuration
Reasoning
Always-on — Built-in thinking is always active. No toggle — the model reasons on every query.
Max output tokens32,768
TemperatureAPI default
NotesTemperature is not sent. API default (1.0) applies.
Task Performance
TaskAccuracy
Cite Verify95.5%
Statute QA56.9%‡
strict · lenient 89.0%
Jurisdiction Router94.7%
Bijural Equivalence56.0%
strict · lenient 79.4%
Case Treatment94.8%
Case Outcome88.9%
Standard of Review97.8%
Doctrine Shift87.1%
Regulation QA57.3%
strict · lenient 57.3%
Tribunal QA90.9%
strict · lenient 90.9%
Multi-Doc Reasoning90.0%
Temporal Reasoning88.5%
Citation Lookup97.2%
Citation Holding49.8%
strict · lenient 49.8%
Limitation Periods99.5%
Issue Spotting39.3%‡
every key issue · exact set 9.6% · lenient 68.0%
Citation Resolution33.3%

‡ the provider's safety filter blocked the answer, or the answer came back empty with no recorded finish reason.These answers stay in the denominator and are scored 0; the count is disclosed per cell. Empty-output policy: an answer that came back empty is never removed from its cell; it counts as not correct against the frozen record count, and the number of such answers and their cause are published with the cell.

  • Statute QA: ‡ 26 answers blocked by the provider's safety filter (empty output); scored 0
  • Issue Spotting: ‡ 8 answers blocked by the provider's safety filter (empty output); scored 0