Limitation Periods
Compute the last day to commence proceedings under Canadian limitation-period statutes — testing statutory date arithmetic.
scoring policySince scorer 2026-10-08 the answer's FINAL date is scored: the date on its last answer line ("Answer:", "Final answer:", "**Answer**", « Réponse : », « Réponse finale : »), else its opening line when that is a date and nothing else, else the last date it states, compared as a date with the key and its listed forms (so 2022-03-01 counts where the key lists "March 1, 2022"). An answer whose final date is wrong scores 0 even when the key's date appears earlier, and a hedge between two dates with no answer line is read as its last date. Before, the whole answer was matched (the first date mentioned, or the leading span of a long answer), so an answer that showed its working scored 0 even when its final date was right. Every row was re-scored from its stored answers (182 records, all 0 → 1): command-a-plus 37.8% → 58.2% (122), gemini-3.1-pro-preview 94.5% → 99.5% (30), grok-4.5 90.5% → 95.3% (29), claude-fable-5 99.8% → 100.0% (1), gpt-5.6-sol unchanged. (since 2026-10-08)
near ceilingNear ceiling: 4 of the 5 rows score 90% or more here (95.3%–100.0%; the other row scores 58.2%), so this column separates those rows little.
† the model's shared thinking and answer token budget ran out before it wrote any answer.These answers stay in the denominator and are scored 0; the count is disclosed per cell. Empty-output policy: an answer that came back empty is never removed from its cell; it counts as not correct against the frozen record count, and the number of such answers and their cause are published with the cell.
- Cohere Command A+: † 36 answers hit the token budget before any answer token; scored 0