Skip to content

Grok 4.5

17 of 17 task types · 14,547 records · 9,578 correct
71.6%
Suite Mean
unweighted avg of the headline tasks
legacy 11-task 78.4%
Evaluation Configuration
Reasoning
Effort-based — Reasoning effort set to high — the maximum supported by this model. Thinking and answer share the output token budget.
Max output tokens32,768
TemperatureAPI default
NotesTemperature is not sent. API default (1.0) applies.
Task Performance
TaskAccuracy
Cite Verify93.3%
Statute QA58.9%
strict · lenient 95.5%
Jurisdiction Router82.5%
Bijural Equivalence49.7%
strict · lenient 75.4%
Case Treatment96.4%
Case Outcome90.5%
Standard of Review98.4%
Doctrine Shift85.7%
Regulation QA47.0%
strict · lenient 47.0%
Tribunal QA85.7%
strict · lenient 85.7%
Multi-Doc Reasoning75.4%
Temporal Reasoning87.2%
Citation Lookup75.9%
Citation Holding38.4%
strict · lenient 38.4%
Limitation Periods95.3%
Issue Spotting47.7%
every key issue · exact set 5.5% · lenient 64.4%
Citation Resolution10.0%