Skip to content

Cohere Command A+

17 of 17 task types · 14,547 records · 6,885 correct
49.5%
Suite Mean
unweighted avg of the headline tasks
legacy 11-task 52.6%
Evaluation Configuration
Reasoning
Always-on — Thinking runs server-side on every query (the vLLM serve's reasoning parser separates it from the answer). Thinking and answer share the output token budget.
Max output tokens65,536
Temperature0.6
NotesSince 2026-10-07: the vendor's recommended decoding from the Command A+ model card (temperature 0.6, top_p 0.95; seed 20261007 recorded), as the API rows use their providers' defaults. max_tokens=65536 (thinking + answer share it), then one re-run of every budget-exhausted record at 131,072; records that exhaust both budgets are scored 0 and marked †. The earlier temperature-0 cells are archived.
Task Performance
TaskAccuracy
Cite Verify54.4%†
Statute QA9.9%†
strict · lenient 56.6%
Jurisdiction Router75.4%
Bijural Equivalence33.1%
strict · lenient 58.9%
Case Treatment83.1%
Case Outcome79.4%†
Standard of Review68.3%
Doctrine Shift55.7%†
Regulation QA12.9%†
strict · lenient 12.9%
Tribunal QA71.4%
strict · lenient 71.4%
Multi-Doc Reasoning63.1%
Temporal Reasoning71.1%†
Citation Lookup63.4%†
Citation Holding1.1%†
strict · lenient 1.1%
Limitation Periods58.2%†
Issue Spotting41.3%†
every key issue · exact set 1.6% · lenient 49.1%
Citation Resolution0.0%†

† the model's shared thinking and answer token budget ran out before it wrote any answer.These answers stay in the denominator and are scored 0; the count is disclosed per cell. Empty-output policy: an answer that came back empty is never removed from its cell; it counts as not correct against the frozen record count, and the number of such answers and their cause are published with the cell.

  • Cite Verify: † 3 answers hit the token budget before any answer token; scored 0
  • Statute QA: † 19 answers hit the token budget before any answer token; scored 0
  • Case Outcome: † 1 answer hit the token budget before any answer token; scored 0
  • Doctrine Shift: † 2 answers hit the token budget before any answer token; scored 0
  • Regulation QA: † 4 answers hit the token budget before any answer token; scored 0
  • Temporal Reasoning: † 1 answer hit the token budget before any answer token; scored 0
  • Citation Lookup: † 18 answers hit the token budget before any answer token; scored 0
  • Citation Holding: † 18 answers hit the token budget before any answer token; scored 0
  • Limitation Periods: † 36 answers hit the token budget before any answer token; scored 0
  • Issue Spotting: † 28 answers hit the token budget before any answer token; scored 0
  • Citation Resolution: † 143 answers hit the token budget before any answer token; scored 0