Grok 4.5
17 of 17 task types · 14,547 records · 9,578 correct
71.6%
Suite Mean
unweighted avg of the headline tasks
legacy 11-task 78.4%
Evaluation Configuration
Reasoning
Effort-based — Reasoning effort set to high — the maximum supported by this model. Thinking and answer share the output token budget.
Max output tokens32,768
TemperatureAPI default
NotesTemperature is not sent. API default (1.0) applies.
Task Performance
| Task | Accuracy | |
|---|---|---|
| Cite Verify | 93.3% | |
| Statute QA | 58.9% strict · lenient 95.5% | |
| Jurisdiction Router | 82.5% | |
| Bijural Equivalence | 49.7% strict · lenient 75.4% | |
| Case Treatment | 96.4% | |
| Case Outcome | 90.5% | |
| Standard of Review | 98.4% | |
| Doctrine Shift | 85.7% | |
| Regulation QA | 47.0% strict · lenient 47.0% | |
| Tribunal QA | 85.7% strict · lenient 85.7% | |
| Multi-Doc Reasoning | 75.4% | |
| Temporal Reasoning | 87.2% | |
| Citation Lookup | 75.9% | |
| Citation Holding | 38.4% strict · lenient 38.4% | |
| Limitation Periods | 95.3% | |
| Issue Spotting | 47.7% every key issue · exact set 5.5% · lenient 64.4% | |
| Citation Resolution | 10.0% |