Skip to content

Limitation Periods

Compute the last day to commence proceedings under Canadian limitation-period statutes — testing statutory date arithmetic.

scoring policySince scorer 2026-10-08 the answer's FINAL date is scored: the date on its last answer line ("Answer:", "Final answer:", "**Answer**", « Réponse : », « Réponse finale : »), else its opening line when that is a date and nothing else, else the last date it states, compared as a date with the key and its listed forms (so 2022-03-01 counts where the key lists "March 1, 2022"). An answer whose final date is wrong scores 0 even when the key's date appears earlier, and a hedge between two dates with no answer line is read as its last date. Before, the whole answer was matched (the first date mentioned, or the leading span of a long answer), so an answer that showed its working scored 0 even when its final date was right. Every row was re-scored from its stored answers (182 records, all 0 → 1): command-a-plus 37.8% → 58.2% (122), gemini-3.1-pro-preview 94.5% → 99.5% (30), grok-4.5 90.5% → 95.3% (29), claude-fable-5 99.8% → 100.0% (1), gpt-5.6-sol unchanged. (since 2026-10-08)

near ceilingNear ceiling: 4 of the 5 rows score 90% or more here (95.3%–100.0%; the other row scores 58.2%), so this column separates those rows little.

600 records|5 models evaluated|Best: 100.0%|Worst: 58.2%
Model Performance (avg 90.5%)
1
Claude Fable 5
600 records · 600 correct
100.0%
model →
2
Gemini 3.1 Pro
600 records · 597 correct
99.5%
model →
3
GPT-5.6 Sol
600 records · 596 correct
99.3%
model →
4
Grok 4.5
600 records · 572 correct
95.3%
model →
5
Cohere Command A+
600 records · 349 correct
58.2%†
model →

† the model's shared thinking and answer token budget ran out before it wrote any answer.These answers stay in the denominator and are scored 0; the count is disclosed per cell. Empty-output policy: an answer that came back empty is never removed from its cell; it counts as not correct against the frozen record count, and the number of such answers and their cause are published with the cell.

  • Cohere Command A+: † 36 answers hit the token budget before any answer token; scored 0