Model Accuracy
Does the forecast work when it has never seen the answer? The headline number is out-of-sample.
Out-of-Sample Backtest — the honest number
How the model does when it has never seen the 2025 result — predicting it from the 2021 election and the pre-election polls alone.
Ridings Called Correctly
88.0%
301 of 342 ridings ⓘ343 ridings total. Terrebonne (QC) is excluded; its result was nullified by the courts after the election.
vs. Naive Baseline
+4.7pp
over “2021 winner holds” (83.3%)
National Vote — Average Error
0.87pp
how far each party’s national share was off, on average
Chance of a Liberal Majority
73%
seats 182 (158–205); actual 168
What “out-of-sample” means. This is the test that matters. The model is given only what a forecaster would have had before the 2025 vote: each riding's 2021 result as a starting point, plus the final pre-election national polling average. It applies the poll-implied swing and projects a winner in every riding. No 2025 data ever touches a prediction — the certified result is used only to mark the answers.
Because the model beats a naive “last winner holds” baseline (83.3% → 88.0%), the extra accuracy is genuine skill, not the churn of a stable map. The projection also correctly treated an LPC majority as likely-but-not-certain; the Liberals landed at 168, four short, well inside the projected seat range.
Method: GE44 (2021) riding baseline · Apr 27, 2025 national poll average · uniform swing · 10,000 Monte-Carlo iterations. Source: simulation/backtest_ge45_oos.py.
Out-of-Sample Accuracy by Region
Where the 2021→2025 backtest was strong and where it was hard: Quebec and the Prairies held up best; three-way races and the territories are hardest.
By-Election Track Record
Top-two margin predicted within 7.7pp on average
University—Rosedale
ON · 2026-04-13
| Party | Predicted | Actual | Error |
|---|---|---|---|
| Liberal | 64.0% | 64.4% | +0.5pp |
| Conservative | 16.4% | 12.4% | -4.0pp |
| NDP | 12.8% | 18.9% | +6.1pp |
| Green | 3.6% | 2.9% | -0.7pp |
Predicted = 2025 election result for this riding + national polling swing as of 2026-04-12.
Scarborough Southwest
ON · 2026-04-13
| Party | Predicted | Actual | Error |
|---|---|---|---|
| Liberal | 61.5% | 69.6% | +8.1pp |
| Conservative | 23.5% | 18.8% | -4.7pp |
| NDP | 7.9% | 5.9% | -2.0pp |
| Green | 3.3% | 2.5% | -0.8pp |
Predicted = 2025 election result for this riding + national polling swing as of 2026-04-12.
Terrebonne
QC · 2026-04-13
| Party | Predicted | Actual | Error |
|---|---|---|---|
| Liberal | 38.6% | 48.3% | +9.6pp |
| Conservative | 11.1% | 3.3% | -7.8pp |
| NDP | 5.5% | 0.5% | -5.0pp |
| Bloc Québécois | 38.5% | 46.9% | +8.4pp |
| Green | 2.9% | 0.4% | -2.5pp |
Predicted = 2025 election result for this riding + national polling swing as of 2026-04-12.
Battle River—Crowfoot
AB · 2025-08-18
| Party | Predicted | Actual | Error |
|---|---|---|---|
| Liberal | 11.8% | 4.3% | -7.5pp |
| Conservative | 76.2% | 80.9% | +4.8pp |
| NDP | 6.8% | 2.1% | -4.7pp |
| Green | 2.0% | 0.2% | -1.8pp |
Predicted = 2025 election result for this riding + national polling swing as of 2025-08-17.
Why margins, not vote shares? By-election turnout and local candidate effects shift absolute vote-share levels in ways that national polling can't anticipate. But the relative positioning of parties, which is what determines seat outcomes, is captured by the national swing prior. Margin accuracy is therefore the better signal: it tests exactly what the model is optimised for. Four margin errors averaging single-digit accuracy is a more informative validation sample than four binary winner calls.
In-Sample Reproduction & Poll Diagnostics
Detailed 2025 breakdown — vote-share error, seat bands, and the missed ridings. The riding-call figure here is in-sample; read it as a consistency check, not the headline.
Reproduction check, not a headline. The figures below start each riding from its actual 2025 result and apply the final-day poll swing — a calibration/consistency check, not an out-of-sample forecast. It runs hot (a very small swing over the true result) and is shown for full transparency. For the honest measure of predictive skill, see the out-of-sample backtest above.
Ridings Called Correctly
97.7%
334 of 342 ridings ⓘ343 ridings total. Terrebonne (QC) is excluded; its result was nullified by the courts after the election.
Vote Share — Average Error
0.9pp
Mean absolute error across 6 parties
80% Bands Hit
5/5
Parties whose actual seats fell in projected range
CPC / LPC Seat Swing
CPC +15 · LPC -14
Seats vs. median projection
Why 80% intervals? We report 80% bands (not 95%) because they're easier to validate against real results. A 95% interval almost never misses, making it hard to test whether the model is actually calibrated. With a single election, a model can pass a 95% CI test by accident; an 80% CI is genuinely informative. See methodology →
National Vote Share
Predicted = model simulation output using the poll average as of 2025-04-27 (not the raw poll average). Actual = certified 2025 election result.
| Party | Predicted | Actual | Error |
|---|---|---|---|
| Liberal | 43.4% | 44.0% | +0.6pp |
| Conservative | 39.1% | 41.5% | +2.4pp |
| NDP | 7.7% | 6.3% | -1.3pp |
| Bloc Québécois | 6.2% | 6.3% | +0.1pp |
| Green | 1.4% | 1.2% | -0.2pp |
| People's | 1.3% | 0.7% | -0.6pp |
Error = actual − predicted. Positive = party outperformed the forecast; negative = underperformed.
Model Forecast
Actual Result
LPC: 168 seats · CPC: 144 seats
Majority threshold: 172 seats · LPC fell 4 short
Seat Projection
Monte Carlo distribution from 10,000 simulated elections using the pre-election poll average (validation runs use 10,000; live scenario runs on the map use 2,000). Brackets show the 80% range (p10–p90).
| Party | Median | 80% range | Actual | Error |
|---|---|---|---|---|
| Liberal | 182 | 156–212 | 168 | -14 |
| Conservative | 129 | 101–155 | 144 | +15 |
| NDP | 8 | 5–12 | 7 | -1 |
| Bloc Québécois | 22 | 17–27 | 22 | 0 |
| Green | 1 | 0–2 | 1 | 0 |
| People's | 0 | 0–0 | — | 0 |
Actual seat counts in green fell within the 80% projected range; red fell outside. Error = actual − median. Positive = party won more seats than projected.
Why vote share accuracy and seat accuracy can differ
A small miss on vote share can produce a large miss on seats, or the reverse. Both are normal.
Comparing the vote share table above to the seat projection table can be confusing. The model might predict a party's national vote share within half a point but still miss several seats, or hit the seat count almost exactly while overestimating the vote by a full percentage point. Both outcomes are common under first-past-the-post.
Three structural reasons account for the gap. First, vote share is a national number but seats are decided riding by riding: a party can be off by a point nationally with most of that error landing in regions where the seat count doesn't change. Second, three-way races are sensitive to where second- and third-place support sits, so a small shift in NDP support can flip several Liberal versus Conservative marginals without moving the headline LPC and CPC numbers. Third, the model applies regional adjustments (Alberta CPC floor, BC NDP baseline, Atlantic variance multiplier) that reflect historical patterns and partially decouple national swing from local outcomes.
This is why we publish vote share error and riding call accuracy separately. They measure different things, and a forecast is only credible if it gets both roughly right. See methodology for the full mechanism →
Regional Accuracy (in-sample)
In-sample riding-call accuracy broken down by region.
Analysis & Commentary
A written breakdown of where the model was right, where it was wrong, and why.
Going into election night, the model gave the Liberals a 68.4% chance of a majority. They fell 4 seats short, landing at 168, just below the 172-seat majority threshold and near the lower edge of the model's 80% range (156–212). That was the headline miss: the model leaned majority, but a narrow minority was the outcome.
The CPC miss is the key residual error. Using the final day's polls, the model showed them at 39.1%; they ended up at 41.5%, a 2.4pp gap that is entirely attributable to polling error. The polls hadn't fully closed on the late CPC surge, even by April 27. Critically, the model itself didn't amplify this: the simulation reproduced the poll average faithfully and added no further CPC underestimation. CPC's actual 144 seats landed comfortably inside the 80% band (101–155).
Everywhere else, the model was sharp. LPC vote share landed within 0.6pp of the forecast. BQ, NDP, GPC, and PPC were all inside their calibrated error bands. Starting from the certified 2025 result, this reproduction check called 334 of 342 ridings correctly (97.7%) — but that figure is in-sample: it begins from the actual riding outcomes and applies only a small poll swing, so it flatters the model. The honest measure is the out-of-sample backtest above, which predicts 2025 from the 2021 map and pre-election polls alone and calls 88.0% of ridings.
This reproduction check uses the final-day poll average from April 27, 2025, one day before voting, applied on top of the certified 2025 riding results. The 0.87pp MAE reflects residual polling error even with fully up-to-date inputs; the simulation layer added no noise beyond what the polls contained. It is a consistency check — not a forecast made in advance.
Missed Ridings (8)
Ridings where the model called the wrong winner, sorted by actual certified margin — closest calls first.
| Riding | Prédiction du modèle (prob. de victoire) | Vainqueur réel | Margin |
|---|---|---|---|
| Windsor—Tecumseh—LakeshoreON | LPC51.9% | CPC | 0.1pp |
| NunavutNU | LPC60.4% | NDP | 0.6pp |
| Acadie—AnnapolisNS | LPC71.8% | CPC | 1.1pp |
| Markham—UnionvilleON | LPC62.8% | CPC | 3.6pp |
| New Westminster—Burnaby—MaillardvilleBC | NDP50.8% | LPC | 3.6pp |
| Cowichan—Malahat—LangfordBC | NDP55% | CPC | 4.6pp |
| Edmonton RiverbendAB | LPC55% | CPC | 5.4pp |
| North Island—Powell RiverBC | NDP49.7% | CPC | 6.1pp |
% = model's win probability for the party it (wrongly) called. Margin = certified vote share gap between 1st and 2nd place; red = under 1pp, amber = under 3pp.
Reproduce these numbers yourself
The scoring harness is public; you don't have to take our word for it. Clone the repo and run:
python scoring/evaluate.py
github.com/Northern-Vibe/scoring → · Neither 338 nor Poliwave publishes an equivalent harness.