Model Accuracy
Does the forecast work when it has never seen the answer? The headline numbers, federal and provincial, are out-of-sample.
Out-of-Sample Backtest — the honest number
How the model does when it has never seen the 2025 result — predicting it from the 2021 election and the pre-election polls alone.
Ridings Called Correctly
88.0%
301 of 342 ridings ⓘ343 ridings total. Terrebonne (QC) is excluded; its result was nullified by the courts after the election.
vs. Naive Baseline
+4.7pp
over “2021 winner holds” (83.3%)
National Vote — Average Error
0.87pp
how far each party’s national share was off, on average
Chance of a Liberal Majority
73%
seats 182 (158–205); actual 168
What “out-of-sample” means. This is the test that matters. The model is given only what a forecaster would have had before the 2025 vote: each riding's 2021 result as a starting point, plus the final pre-election national polling average. It applies the poll-implied swing and projects a winner in every riding. No 2025 data ever touches a prediction — the certified result is used only to mark the answers.
Because the model beats a naive “last winner holds” baseline (83.3% → 88.0%), the extra accuracy is genuine skill, not the churn of a stable map. The projection also correctly treated an LPC majority as likely-but-not-certain; the Liberals landed at 168, four short, well inside the projected seat range.
Method: GE44 (2021) riding baseline · Apr 27, 2025 national poll average · uniform swing · 10,000 Monte-Carlo iterations. Source: simulation/backtest_ge45_oos.py.
Out-of-Sample Accuracy by Region
Where the 2021→2025 backtest was strong and where it was hard: Quebec and the Prairies held up best; three-way races and the territories are hardest.
British Columbia — the same test, provincially
The BC model run forward from the 2020 result on polls published before the 2024 vote, with each firm’s correction calibrated on 2017 and 2020 alone so it could not be fitted to the election being called.
Districts Called Correctly
86.0%
80 of 93 districts
vs. Naive Baseline
+6.4pp
over “2020 winner holds” (79.6%)
Provincial Vote — Average Error
1.23pp
how far each party’s share was off, on average
Seat Totals In Range
3 of 3
parties whose result landed inside the 80% band
| Party | Projected | 80% range | Actual | Error |
|---|---|---|---|---|
| BC NDP | 49.1 | 43–55 | 47 | +2.1 |
| Conservative | 42.6 | 37–48 | 44 | −1.4 |
| BC Green | 1.3 | 1–2 | 2 | −0.7 |
| Other | 0.0 | 0–0 | 0 | +0.0 |
The backtest is one auditable script, simulation/bc/backtest_2024.py. It does not cover the OneBC and CentreBC columns, which had nothing to stand on in 2024, or the floor-crossing transfer, whose table describes 2026.
Nova Scotia: the same test, provincially
The Nova Scotia model run forward from the 2021 result on polls published before the 2024 vote, with each firm's correction calibrated on 2017 and 2021 alone so it could not be fitted to the election being called.
Run at the election-eve setting: the full house-effect correction, as a campaign uses it, rather than the relative correction applied between elections.
Districts called correctly
96.4%
54 of 56 districts
vs. naive baseline
+26.8pp
over "2021 winner holds" (69.6%)
Provincial vote: average error
1.67pp
how far each party's share was off, on average · Brier 0.090
Seat totals in range
3 of 3
parties whose result landed inside the 80% band
| Party | Projected | 80% range | Actual | Error |
|---|---|---|---|---|
| Progressive Conservative | 42.7 | 39–46 | 44 | −1.3 |
| NS NDP | 8.2 | 7–10 | 9 | −0.8 |
| Liberal | 4.1 | 2–7 | 2 | +2.1 |
| Independent | 1.0 | – | 1 | +0.0 |
The backtest is one auditable script, simulation/ns/backtest_2024.py. Its two misses were Sackville-Cobequid and Halifax Atlantic, whose 2021 Liberal MLA crossed to the PCs in January 2024. The PCs were understated in the polls' final week in all three elections we calibrate on, which the 3.5-point floor allows for but does not correct.
Prince Edward Island: the same test, provincially
The PEI model run forward from the 2019 result on polls published before the 2023 vote, with polling error measured on 2015 and 2019 alone so it could not be fitted to the election being called.
Run as the live model runs: VoteLab's averaging rules and no house-effect correction, which one regular pollster leaves nothing to measure against.
Districts called correctly
80.8%
21 of 26 districts
vs. naive baseline
+15.4pp
over "2019 winner holds" (65.4%)
Provincial vote: average error
2.01pp
how far each party's share was off, on average · Brier 0.293
Seat totals in range
4 of 4
parties whose result landed inside the 80% band
| Party | Projected | 80% range | Actual | Error |
|---|---|---|---|---|
| Progressive Conservative | 22.2 | 18–26 | 22 | +0.2 |
| Liberal | 2.5 | 0–5 | 3 | −0.5 |
| Green | 1.8 | 0–4 | 2 | −0.2 |
| PEI NDP | 0.5 | 0–1 | 0 | +0.5 |
The backtest is one auditable script, simulation/pe/backtest_2023.py. District 9 is left out of the district score, because its 2019 vote was a deferred election held in July after a candidate died; the seat totals count all 27. Two of the five misses were called at about 90%, which is where the Brier score comes from, and Others are held at their 2023 share, the one figure the test borrows from the result it is scored against.
By-Election Track Record
Top-two margin predicted within 9.5pp on average
Chicoutimi—Le Fjord
QC · 2026-08-31
| Party | Predicted | Actual | Error |
|---|---|---|---|
| Liberal | 29.9% | 51.3% | +21.4pp |
| Conservative | 25.2% | 12.5% | -12.7pp |
| NDP | 7.6% | 2.0% | -5.6pp |
| Bloc Québécois | 30.9% | 32.9% | +1.9pp |
| Green | 3.3% | 0.7% | -2.6pp |
Predicted = 2025 election result for this riding + national polling swing as of 2026-08-30.
Beaches—East York
ON · 2026-08-31
| Party | Predicted | Actual | Error |
|---|---|---|---|
| Liberal | 66.5% | 55.7% | -10.8pp |
| Conservative | 14.6% | 15.5% | +0.9pp |
| NDP | 12.6% | 25.8% | +13.2pp |
| Green | 3.7% | 1.7% | -2.0pp |
Predicted = 2025 election result for this riding + national polling swing as of 2026-08-30.
North Vancouver—Capilano
BC · 2026-08-31
| Party | Predicted | Actual | Error |
|---|---|---|---|
| Liberal | 58.5% | 58.6% | +0.1pp |
| Conservative | 24.8% | 29.2% | +4.4pp |
| NDP | 9.9% | 3.1% | -6.8pp |
| Green | 4.1% | 8.7% | +4.6pp |
Predicted = 2025 election result for this riding + national polling swing as of 2026-08-30.
University—Rosedale
ON · 2026-04-13
| Party | Predicted | Actual | Error |
|---|---|---|---|
| Liberal | 64.0% | 64.4% | +0.5pp |
| Conservative | 16.4% | 12.4% | -4.0pp |
| NDP | 12.8% | 18.9% | +6.1pp |
| Green | 3.6% | 2.9% | -0.7pp |
Predicted = 2025 election result for this riding + national polling swing as of 2026-04-12.
Scarborough Southwest
ON · 2026-04-13
| Party | Predicted | Actual | Error |
|---|---|---|---|
| Liberal | 61.5% | 69.6% | +8.1pp |
| Conservative | 23.5% | 18.8% | -4.7pp |
| NDP | 7.9% | 5.9% | -2.0pp |
| Green | 3.3% | 2.5% | -0.8pp |
Predicted = 2025 election result for this riding + national polling swing as of 2026-04-12.
Terrebonne
QC · 2026-04-13
| Party | Predicted | Actual | Error |
|---|---|---|---|
| Liberal | 38.6% | 48.3% | +9.6pp |
| Conservative | 11.1% | 3.3% | -7.8pp |
| NDP | 5.5% | 0.5% | -5.0pp |
| Bloc Québécois | 38.5% | 46.9% | +8.4pp |
| Green | 2.9% | 0.4% | -2.5pp |
Predicted = 2025 election result for this riding + national polling swing as of 2026-04-12.
Battle River—Crowfoot
AB · 2025-08-18
| Party | Predicted | Actual | Error |
|---|---|---|---|
| Liberal | 11.8% | 4.3% | -7.5pp |
| Conservative | 76.2% | 80.9% | +4.8pp |
| NDP | 6.8% | 2.1% | -4.7pp |
| Green | 2.0% | 0.2% | -1.8pp |
Predicted = 2025 election result for this riding + national polling swing as of 2025-08-17.
Why margins, not vote shares? By-election turnout and local candidate effects shift absolute vote-share levels in ways that national polling can't anticipate. But the relative positioning of parties, which is what determines seat outcomes, is captured by the national swing prior. Margin accuracy is therefore the better signal: it tests exactly what the model is optimised for. Four margin errors averaging single-digit accuracy is a more informative validation sample than four binary winner calls.
Reproduce these numbers yourself
The scoring harness is public; you don't have to take our word for it. Clone the repo and run:
python scoring/evaluate.py