xJawn NHL expected goals

← Blog

Model v3: the cascade

v2 ended on a registered question: would a nested backfit win outright? It did — sixteen seasons out of sixteen, the first mechanical bar-pass in this project's history — and not for the reason we predicted. Three questions instead of one, and blocking becomes a number.

Model v2 ended on a registered question. The promotion that made its joint fit the canonical model lost on architecture while winning on the idea — the decomposition said a nested version of the same fit should be the first configuration to win outright, and that model did not exist yet. This post is the answer: it was built, it won — sixteen seasons out of sixteen, the first mechanical bar-pass in this project's history — and the decomposition of why it won inverted the thesis we built it on. Both halves of that sentence are the story.

Three questions instead of one

A shot attempt is not one event; it is a gauntlet. First it has to get past the defenders in the lane — a quarter of all attempts don't. Then it has to hit the net — nearly a third of the survivors miss. Only then does the goalie get a say. Model v1 scored that gauntlet as a cascade of three chained probabilities:

xg = P(through) · P(on net | through) · P(goal | on net)

The v2 joint fit flattened it: one booster, one target, did it go in? That was a simplification bought for the backfit machinery, and the promotion paid for it in the open — the adjudication measured the flat chassis losing 5.76×10-4 of log-loss to v1's cascade while decontaminating the structure won 5.20×10-4. Nearly equal and opposite. So the registered prediction wrote itself: rebuild the backfit on the cascade, take both effects, and the first outright win should follow.

It won — and not for the reason we predicted

The bar was the same pre-registered one everything in v2 was held to: beat the incumbent's held-out log-loss on the headline season by 5×10-4, improve in a majority of seasons. The nested fit, people layer included, cleared it without an over-rule: +6.51×10-4 on the headline season, better in 16 of 16, bootstrap interval excluding zero (+4.41 to +8.59×10-4, ten thousand resamples), pooled improvement +1.09×10-3. Every season individually helps. After a summer of adoptions by owner over-rule, the mechanical verdict finally just said yes.

Then the decomposition took the win apart, and the two effects did not contribute what the prediction said they would:

Bar chart isolating two effects on overall log-loss: the chassis effect at +7.79e-4 and the decontamination effect at −0.36e-4, against the track's predicted +5.20e-4.
The win, isolated on 2,344,674 identical shots. The chassis — three specialized boosters on the cascade instead of one general booster on the flat problem — is worth +7.79×10-4, larger than predicted and enough to account for the entire outright win on its own. The decontamination — the backfit loop this track was named for — measures −0.36×10-4: slightly negative, smaller than the loop's own retrain noise.

The honest headline, recorded the way the failed gates in v2 were: the win is architectural, not from decontamination. The track's stated reason for existing — that re-estimating structure net of the people would recover real value — is not what the data shows on this chassis. The mechanism underneath is real and it replicated: as the loop converges, the signal the people layer finds grows by 13–14% in every one of the three stages, right inside the flat fit's own 11.5–16% range. The talent absorption v2 measured is there, in triplicate. It just no longer shows up as a measurably better structure once the cascade is doing the work.

One more thing came back with the cascade, and it settles an open debt. The flat promotion carried a cost stated as permanent at the time: 0.0022 of overall AUC, "structural, and no future calibrator will recover it." The nested structural model recovers it and more — +0.0030 of overall AUC against the flat canonical, positive against even the old v1 cascade — because the sentence was true about calibrators and said nothing about architectures. Danger AUC moves +0.0002, thin but non-negative. The per-stage feature breakdown also returns to the methods page, which renders it live from the shipped model.

So on 25 July 2026 the cascade's identity-blind structural sub-model — same talent-blind contract as ever, no player terms in it at all — became the canonical model behind every xG number on this site, on the chassis margin and the AUC recovery above. The people-inclusive fit that actually passed the bar went where player-aware models go here: into the game engine's audition, where that engine's own pre-registered scoreboard decides what ships. You can watch the cascade behave in the canonical model's own held-out predictions:

What actually happenedshots mean xg
blocked602,587 0.0134
missed503,017 0.0490
saved1,126,835 0.0565
goal112,235 0.1410

Held-out mean xg by eventual outcome, recomputed from the shipped model's stored out-of-fold predictions on every build — live numbers, not a pasted table. The gradient is the model working: attempts that ended blocked were, on average, longer-odds propositions at the moment of release than attempts that got through, and goals were the best chances of all. (A perfect model would not put blocked shots at zero — the shooter doesn't know the block is coming, and neither does the model.)

Five people terms, one per question

The reason to nest was architecture; the reason it matters is measurement. The backfit fits a shrunk, cross-fitted people layer per stage, and each stage's universe gives its people terms a better target than did it go in? ever was. Goals are ~5% of attempts; getting blocked is a quarter of them — a far higher-entropy target, so the people terms that read it are far better powered per shot. Five components ship, in cascade order:

  • Getting through (shooter, stage 1) — whose attempts evade the lane.
  • Blocking (defender, stage 1) — who takes attempts away. New in this generation, and the first defensive people component in the model.
  • Accuracy (shooter, stage 2) — whose unblocked attempts find the net.
  • Finishing (shooter, stage 3) — whose on-net shots beat the goalie.
  • Saving (goalie, stage 3) — who turns away more than the chances predict. Keyed, in the cascade, to the goalie actually in net for every on-net shot — the row-support wedge that nearly fooled v2 dissolves by construction, because stage 3's universe has no blocked shots in it.

Every component cleared its pre-registered split-half persistence gate: getting through 0.54, blocking 0.50, accuracy 0.51, finishing 0.41, saving 0.50 — blocking repeats about as well as goaltending, which is the finding the blocked-shots series turns into a product. And blocking carries the validation that makes it science rather than a fit artifact: an instrument the model never sees.

Scatter of the model's blocking rating against observed opponent blocks per 100 on-ice attempts, 2,668 players, with a clear positive trend.
The check the fit cannot game: the play-by-play records who blocked each shot, and that column is never a model input — it is held out as a measurement instrument. The blocking rating correlates 0.31 with a player's observed blocks per 100 on-ice attempts against (and 0.32 with blocks per 60, an independent second instrument). Not large, but positive, twice, on a component that has no other way to be checked — a stable rating that predicts an outcome-blind behavioral proxy did not exist for defense in this codebase before this model.

And the shipped ratings themselves, drawn live from the artifact the site actually reads (joint__baseline_xg__f3e3111e, fit on shots strictly before 2026-06-14, 3,125 players):

One violin per component, every fitted player included. The mass pinned at zero is the prior doing its job: a player the model has little evidence on is held at league average, so a rating near zero means "not enough shots to say," never "confirmed average." The goal-stage components (orange) are the same finishing and saving v2 shipped; the three to their left are what nesting bought.

What reaches the site

The players table carries the shooter and defender components as the Sh and Blk rating columns, and puts them on each player's current-season shots as goal impacts. Blocking's converted value — about a goal a season at the top of the league — lands as its own term in the value accounting, beside finishing, saving and the RAPM rate terms. Two of v2's closing caveats have since been retired the honest way, by building the missing thing rather than rewording the caveat: "no power play, no penalty kill" ended when special-teams value got its own estimator, and "a rating is a piece of player value, never the whole of it" is answered by the value accounting that now composes the pieces — rate, conversion, blocking, special teams — into one goals ledger, with goaltending measured as its own object beside it.

Honest limitations

  • The decontamination null is a finding, not a footnote. The backfit loop did not measurably improve the cascade's structure over its margin-less start. The loop's formal convergence gate also reads FAIL at the pre-registered tolerance — the fit reaches a stable fixed point two iterations in, and the tolerances stay where they were registered rather than being tuned until the gate passes.
  • Stage 1 leans on imputed geometry. Blocked shots are recorded where they died, so the model's block-stage discrimination is measured partly on reconstructed origins. Its stage-1 AUC should be read as an upper bound until a registered sensitivity read (imputed rows vs observed) lands.
  • The danger-AUC margin is thin. +0.0002 over the flat canonical, within season-to-season noise. The decisive reads were log-loss, overall AUC, and the architecture decomposition — not danger ranking.
  • The displayed ratings are a final refit on all shots, labeled in-sample. The gates and validations quoted here are the held-out evidence; the numbers on the players table are the best single estimate, not a held-out score. (Same discipline, and same wording, as v2.)
  • The adjudication figures on this page are the July record. The shipped artifacts were rebuilt in August on corrected blocked-shot origin imputation — a data fix, not a design change — and the live elements above read from the current pins. The archived adjudication numbers are quoted as what was measured when the decision was made.

v2's registered question is answered, and this post registers none — the cascade is the canonical model, the five components are shipping, and what happens next is product work: where the ratings surface, and what the value accounting does with them. The model, for once, is allowed to just run.