Cohere Command A+
17 of 17 task types · 14,547 records · 6,885 correct
49.5%
Suite Mean
unweighted avg of the headline tasks
legacy 11-task 52.6%
Evaluation Configuration
Reasoning
Always-on — Thinking runs server-side on every query (the vLLM serve's reasoning parser separates it from the answer). Thinking and answer share the output token budget.
Max output tokens65,536
Temperature0.6
NotesSince 2026-10-07: the vendor's recommended decoding from the Command A+ model card (temperature 0.6, top_p 0.95; seed 20261007 recorded), as the API rows use their providers' defaults. max_tokens=65536 (thinking + answer share it), then one re-run of every budget-exhausted record at 131,072; records that exhaust both budgets are scored 0 and marked †. The earlier temperature-0 cells are archived.
Task Performance
| Task | Accuracy | |
|---|---|---|
| Cite Verify | 54.4%† | |
| Statute QA | 9.9%† strict · lenient 56.6% | |
| Jurisdiction Router | 75.4% | |
| Bijural Equivalence | 33.1% strict · lenient 58.9% | |
| Case Treatment | 83.1% | |
| Case Outcome | 79.4%† | |
| Standard of Review | 68.3% | |
| Doctrine Shift | 55.7%† | |
| Regulation QA | 12.9%† strict · lenient 12.9% | |
| Tribunal QA | 71.4% strict · lenient 71.4% | |
| Multi-Doc Reasoning | 63.1% | |
| Temporal Reasoning | 71.1%† | |
| Citation Lookup | 63.4%† | |
| Citation Holding | 1.1%† strict · lenient 1.1% | |
| Limitation Periods | 58.2%† | |
| Issue Spotting | 41.3%† every key issue · exact set 1.6% · lenient 49.1% | |
| Citation Resolution | 0.0%† |
† the model's shared thinking and answer token budget ran out before it wrote any answer.These answers stay in the denominator and are scored 0; the count is disclosed per cell. Empty-output policy: an answer that came back empty is never removed from its cell; it counts as not correct against the frozen record count, and the number of such answers and their cause are published with the cell.
- Cite Verify: † 3 answers hit the token budget before any answer token; scored 0
- Statute QA: † 19 answers hit the token budget before any answer token; scored 0
- Case Outcome: † 1 answer hit the token budget before any answer token; scored 0
- Doctrine Shift: † 2 answers hit the token budget before any answer token; scored 0
- Regulation QA: † 4 answers hit the token budget before any answer token; scored 0
- Temporal Reasoning: † 1 answer hit the token budget before any answer token; scored 0
- Citation Lookup: † 18 answers hit the token budget before any answer token; scored 0
- Citation Holding: † 18 answers hit the token budget before any answer token; scored 0
- Limitation Periods: † 36 answers hit the token budget before any answer token; scored 0
- Issue Spotting: † 28 answers hit the token budget before any answer token; scored 0
- Citation Resolution: † 143 answers hit the token budget before any answer token; scored 0