What a hot finish is worth in October
Methods and further results
The protocol, the data, every prespecified comparison, and the checks behind the essay. Numbers here come from the same saved results file as the essay.
What was fixed, and when
The research protocol (version 1.0) was frozen at Oct. 2, 2026, 23:44 UTC, before any model was scored on postseason outcomes. Its SHA-256 hash begins edc2517981e19677. Amendments were only appended, and runs/protocol-at-freeze.md, the file with them removed, reproduces that hash; the validation script checks it on every run. At that point I had looked at source coverage, field layouts, series counts, home-site patterns, the 2026 Padres game log, standings, and the 2026 postseason schedule. No relationship between a heat measure and a postseason result had been computed.
Two amendments followed, both logged in the protocol file:
- A1, written after the core results and before any roster result: one recency reference date per season; the roster coefficient’s prior; FIP without its league constant; and the 2026 roster inputs, since neither Division Series roster had been announced.
- A2, written after the core results: the protocol described the 3-point threshold as roughly home-field advantage in a five-game series. Home field is worth about 1.6 points there. A 3-point change equals about 0.14 runs per game of team quality, about two wins over a season. The threshold itself didn’t change.
The Division Series forecasts were frozen at Oct. 2, 2026, 23:53 UTC and the roster-aware addendum at Oct. 2, 2026, 23:55 UTC. Game 1 was scheduled for Oct. 4, 2026, 00:30 UTC. The freeze files record each forecast’s hash. The timestamps are written by the scripts themselves; a commit would make them independently checkable. The Wild Card Series against the Cubs had already been played, so its model probability (about 49% for San Diego under the recency-weighted model) is retrospective.
Data and reconciliation
Regular-season and postseason games from 1995 to 2025 come from Retrosheet game logs, which include both teams’ batting lines and starting pitchers. The 2026 season, standings, rosters, player lines, box scores, transactions, and probable pitchers come from the MLB Stats API, retrieved October 2 and 3, 2026. The full source ledger, with retrieval times, grain, transformations, and limits, is research/padres-october/sources.csv.
A second parser, written separately from the pipeline, rebuilt every team’s wins, losses, runs scored, and runs allowed from the raw Retrosheet files. All 924 team-seasons from 1995 to 2025 matched MLB’s standings exactly. The 2026 games matched the standings for all 30 teams. The Padres’ box scores sum to their official 722 runs scored and 681 allowed.
| Seasons | Format | Matchups each season |
|---|---|---|
| 1995 to 2011 | Division Series, LCS, World Series | 7 |
| 2012 to 2019, 2021 | Adds two single wild-card games | 9 |
| 2022 to 2025 | Adds four best-of-three Wild Card Series; top two seeds get byes | 11 |
| 2020 | Excluded: 60-game season, 16-team bracket | 0 |
Each matchup is one row. Team A is the team that hosted game 1. The Division Series used a 2-3 site pattern in 1995 to 1997 and 2012 and 2-2-1 otherwise; the pipeline checks every series’ actual sites against its declared pattern. I traced 17 series by hand across each format change, including all six Padres postseasons in the period:
| Series | Season | Sites played | Matches |
|---|---|---|---|
| Mariners over Yankees, 2-3 site format | 1995 | AABBB | Yes |
| Cardinals over Padres | 1996 | AAB | Yes |
| Padres over Astros | 1998 | AABB | Yes |
| Padres over Braves | 1998 | AABBBA | Yes |
| Yankees over Padres | 1998 | AABB | Yes |
| Cardinals over Padres | 2005 | AAB | Yes |
| Cardinals over Padres | 2006 | AABB | Yes |
| Cardinals over Braves, first single-game wild card | 2012 | A | Yes |
| Giants over Reds, 2-3 site format | 2012 | AABBB | Yes |
| Nationals over Astros | 2019 | AABBBAA | Yes |
| Braves over Astros | 2021 | AABBBA | Yes |
| Padres over Mets, first best-of-three wild card | 2022 | AAA | Yes |
| Padres over Dodgers | 2022 | AABB | Yes |
| Phillies over Padres | 2022 | AABBB | Yes |
| Padres over Braves | 2024 | AA | Yes |
| Dodgers over Padres | 2024 | AABBA | Yes |
| Dodgers over Blue Jays | 2025 | AABBBAA | Yes |
Definitions
- Rating. A regularized fit of every regular-season game’s run margin: home runs minus away runs equals home rating minus away rating plus a home constant. Ratings are runs per game against an average opponent, shrunk toward zero by 40 games of average play.
- Recency weighting. Each game counts half as much for every 120 days before the season’s last day. The half-life and the shrinkage were chosen once, on 1995 to 2014 regular seasons, before any postseason scoring.
- Final-30 record. Winning percentage over a team’s last 30 regular-season games, ties excluded. Frozen at the end of the regular season; playoff games never count.
- Underlying runs. BaseRuns: A = H + BB + HBP − HR − 0.5 IBB; B = 1.1 (1.4 TB − 0.6 H − 3 HR + 0.1 (BB + HBP − IBB) + 0.9 (SB − CS − GIDP)); C = AB − H + CS + GIDP; runs = A B / (B + C) + HR. Computed for a team’s batting line and its opponents’.
- FIP. (13 HR + 3 (BB − IBB + HBP) − 2 K) / innings + a league constant that puts it on the ERA scale.
- After the break. Games after the All-Star Game date: July 14 in 2026.
- Hot group. The top third of playoff teams by final-30 record, with cutoffs set on 1995 to 2014 qualifiers (.567 and .633).
Models and scoring
Every model predicts a game: the log-odds that team A wins equal a slope times the rating difference, plus a home edge, plus (for heat models) a coefficient times the standardized heat difference. Series probabilities follow exactly from the game probabilities, the format, and the site pattern, treating games as independent given the inputs. Heat coefficients are fit by penalized maximum likelihood with a normal prior (standard deviation 0.2 per standard deviation of heat).
One choice was made on 2015 to 2019: take the slope and home edge from the postseason fit, or fix them at the regular-season values. The fixed version scored better (Brier 0.2307 against 0.2350 on 45 matchups), so every model uses a slope of 0.475 per run and a home edge of 0.168 in log-odds. Only heat and roster coefficients are fit on postseason games.
The Brier score is the mean squared difference between the probability given to team A and the outcome. Log loss penalizes confident misses more heavily. The calibration slope is 1 when probabilities are spread correctly and below 1 when they’re too spread out.
Held-out results
Fit on 1995 to 2019 (191 matchups), scored on 2021 to 2025 (53 matchups). No model beat always saying 50% on Brier score over this stretch.
| Model | Brier | Log loss | Accuracy | Calibration slope |
|---|---|---|---|---|
| C0: 50% every time | 0.2500 | 0.6931 | 47% | 0.00 |
| C1: Log5 from full-season records | 0.2662 | 0.7281 | 47% | -1.05 |
| C2: Full-season rating | 0.2566 | 0.7080 | 58% | 0.03 |
| C3: Recency-weighted rating | 0.2591 | 0.7137 | 55% | -0.23 |
| C4: C3 plus final-30 record | 0.2574 | 0.7097 | 58% | -0.20 |
| C5: C3 plus final-30 underlying runs | 0.2588 | 0.7125 | 55% | -0.36 |
| C2W: C2 plus final-30 record | 0.2558 | 0.7061 | 60% | 0.06 |
| Comparison | Brier change | Within-season resampling, 95% | Season resampling, 95% | Better / worse |
|---|---|---|---|---|
| Add final-30 record (C4 vs C3) | −0.002 | −0.008 to +0.005 | −0.007 to +0.004 | 25 / 22 |
| Add final-30 underlying runs (C5 vs C3) | 0.000 | −0.004 to +0.003 | −0.006 to +0.004 | 24 / 29 |
| Weight recent games more (C3 vs C2) | +0.003 | −0.005 to +0.009 | −0.003 to +0.011 | 21 / 32 |
| Add final-30 record without recency (C2W vs C2) | −0.001 | −0.004 to +0.003 | −0.004 to +0.002 | 25 / 22 |
| Rating model vs log5 (C2 vs C1) | −0.010 | −0.023 to +0.004 | −0.031 to +0.009 | 26 / 27 |
With only five held-out seasons, the season-resampling interval is rough. Dropping one season at a time, the C4-versus-C3 change ranged from −0.004 to +0.001.
Expanding window, 2005 to 2025
Each season forecast from all earlier normal seasons: 174 matchups. Adding the final-30 record changed the Brier score by +0.002 (95% −0.002 to +0.006); adding underlying runs, by +0.002 (−0.001 to +0.005); recency weighting, by +0.001 (−0.003 to +0.004).
| Model | Brier | Log loss | Accuracy | Calibration slope |
|---|---|---|---|---|
| C2: Full-season rating | 0.2495 | 0.6926 | 54% | 0.27 |
| C3: Recency-weighted rating | 0.2502 | 0.6941 | 54% | 0.21 |
| C4: C3 plus final-30 record | 0.2522 | 0.6978 | 52% | 0.09 |
| C5: C3 plus final-30 underlying runs | 0.2523 | 0.6984 | 53% | 0.01 |
| Group | Team-matchups | Won | Forecast expected |
|---|---|---|---|
| cold | 113 | 55 | 53.7 |
| middle | 138 | 68 | 69.3 |
| hot | 97 | 51 | 50.9 |
Size of the effect
The heat coefficient fit on all 244 matchups (1,044 games, 30 postseasons), converted to the change in a 2-2-1 best-of-five probability for an otherwise equal team one standard deviation hotter. Intervals resample whole postseasons (1,000 draws for the final-30 record, 300 for the others). This is estimation on all seasons, not a held-out test.
| Heat measure | Estimate | 95% interval |
|---|---|---|
| Final 30 games, record | −2.0 | −6.4 to +2.4 |
| Final 30 games, run differential | −2.1 | −6.0 to +2.6 |
| Final 30 games, underlying runs | −1.0 | −5.2 to +3.0 |
| Final 15 games, record | −0.3 | −3.6 to +3.8 |
| Final 15 games, run differential | −0.2 | −3.2 to +5.0 |
| After the All-Star break, record | −3.9 | −8.1 to +1.3 |
| After the All-Star break, run differential | −3.9 | −7.3 to 0.0 |
| Sample | Points per SD | Matchups |
|---|---|---|
| 1995 to 2011 | −1.2 | 119 |
| 2012 to 2025 | −2.8 | 125 |
| All seasons, Padres matchups removed | −1.2 |
Power, shuffle, and robustness
Power: I simulated postseason games from a known model with a planted heat effect, ran the full fit-and-score pipeline at four effect sizes, and counted how often each test called it.
| Planted effect, series points | Held-out test detects | All-season estimate detects | Mean estimated coefficient |
|---|---|---|---|
| 0.0 | 4% | 2% | 0.004 (true 0.000) |
| +3.0 | 7% | 26% | 0.060 (true 0.064) |
| +6.0 | 14% | 78% | 0.123 (true 0.129) |
| +11.8 | 33% | 100% | 0.247 (true 0.258) |
With no planted effect, the tests called one 4% and 2% of the time, and the estimated coefficients tracked the planted ones. These check the pipeline, not baseball.
Shuffle: permuting final-30 records among playoff teams within each training season (200 times) gave a mean coefficient of 0.009, and 42% of the shuffled models improved the held-out score, as expected when the input carries no signal.
| Check | Brier change | 95% interval |
|---|---|---|
| Window: win15 | +0.003 | −0.002 to +0.007 |
| Window: rd30 | −0.001 | −0.007 to +0.005 |
| Window: comp30 | 0.000 | −0.004 to +0.003 |
| Window: post_asb_win | −0.009 | −0.022 to +0.005 |
| Window: post_asb_rd | −0.007 | −0.021 to +0.008 |
| Window: excess30 | −0.001 | −0.007 to +0.005 |
| Prior scale ×0.5 | −0.001 | −0.007 to +0.004 |
| Prior scale ×2.0 | −0.002 | −0.009 to +0.005 |
| without padres (47) | −0.001 | −0.008 to +0.006 |
| without extreme streaks (44) | −0.002 | −0.007 to +0.003 |
| first entry only (20) | −0.001 | −0.011 to +0.010 |
| Series-level logistic model | 0.000 | −0.006 to +0.006 |
Roster study
MLB’s dated active rosters are complete from 2009, so this study covers 146 matchups from 2009 to 2025 (292 team rosters, none shorter than 24 players). Roster quality uses the active roster on each matchup’s first day: the nine hitters with the most regular-season plate appearances, the four pitchers with the most starts, and the six other pitchers with the most relief appearances, valued by shrunk wOBA and FIP and converted to runs per game. It correlates 0.58 with the recency-weighted rating.
| Comparison | Brier change | 95% interval |
|---|---|---|
| Add roster quality to C3 (R1 vs C3) | +0.001 | −0.009 to +0.011 |
| Add final-30 record to R1 (R2 vs R1) | +0.001 | −0.001 to +0.003 |
For 2026, the usage rules picked these players. San Diego’s list is its Wild Card Series roster. Milwaukee’s comes from its 36-man active list, because its Division Series roster wasn’t announced; the rules missed Game 2 starter Logan Henderson, who made fewer starts than the four selected.
| Team | Hitters | Starters | Relievers | Runs per game vs. average |
|---|---|---|---|---|
| Padres | Fernando Tatis Jr., Manny Machado, Jackson Merrill, Xander Bogaerts, Ty France, Jake Cronenworth, Gavin Sheets, Freddy Fermin, Miguel Andujar | Michael King, Walker Buehler, Robbie Ray, Nick Pivetta | Adrian Morejon, Mason Miller, Bradgley Rodriguez, Yuki Matsui, Jason Adam, David Morgan | +0.38 |
| Brewers | Brice Turang, William Contreras, Jake Bauers, Jackson Chourio, Garrett Mitchell, Christian Yelich, Joey Ortiz, David Hamilton, Sal Frelick | Dustin May, Jacob Misiorowski, Kyle Harrison, Robert Gasser | Aaron Ashby, Trevor Megill, Abner Uribe, Antonio Senzatela, JoJo Romero, Grant Anderson | +0.88 |
Regular-season study
For each season from 1995 to 2014, ratings were fit on the first seven-eighths of the schedule and used to forecast the remaining 6,106 games, holding out one season at a time. A 120-day half-life scored best, narrowly. Adding each team’s last-30 record moved the Brier score from 0.24250 to 0.24256 with the recency-weighted rating, and from 0.24265 to 0.24266 with the full-season rating.
San Diego in detail
| Measure | Before break | After break | Final 30 | Season |
|---|---|---|---|---|
| Record | 48-48 | 43-23 | 20-10 | 91-71 |
| Run differential per game | −0.45 | +1.27 | +0.87 | +0.25 |
| Pythagorean win pct. | .451 | .626 | .583 | .527 |
| One-run games | 13-12 | 13-10 | 7-4 | 26-22 |
| Mean opponent rating | +0.03 | −0.01 | +0.01 | +0.01 |
| Runs per game | 3.95 | 5.20 | 5.20 | 4.46 |
| OPS | .673 | .781 | .768 | .718 |
| Strikeout rate | 22.8% | 19.8% | 20.7% | 21.6% |
| Starters ERA / FIP | 4.78 / 4.62 | 3.92 / 4.88 | 4.16 / 4.80 | 4.42 / 4.73 |
| Starter innings per game | 4.61 | 4.70 | 4.90 | 4.65 |
| Bullpen ERA / FIP | 3.68 / 3.84 | 3.31 / 3.26 | 3.74 / 3.63 | 3.53 / 3.60 |
| Window | Padres | Brewers | MLB leader |
|---|---|---|---|
| Final 15 games | 12-3, +22, #2 | 12-3, +42, #1 | Brewers |
| Final 18 games | 15-3, +32, #2 | 15-3, +46, #1 | Brewers |
| Final 20 games | 17-3, +34, #1 | 15-5, +40, #3 | Padres |
| Final 30 games | 20-10, +26, #4 | 22-8, +54, #1 | Brewers |
Material roster moves: Robbie Ray and Casey Mize acquired Aug. 3; Nick Pivetta activated from the 60-day injured list Sept. 7; Mize placed on the injured list September 26. Source: MLB Stats API transactions.
Comparison teams
Candidates: all 274 playoff teams from 1995 to 2025. Strength, final-30 record, improvement after the break, and offense are standardized within season across all teams; bullpen, rotation, and top-three hitter share are measured against the league and standardized across playoff teams. Distance is weighted Euclidean (weights: rating 1, win30 1, asb_improvement 1, bullpen 1, rotation 1, offense 1, concentration 0.5). Outcomes were joined after the ranking was fixed.
| Trait | Value | Standard deviations |
|---|---|---|
| Full-season strength (runs per game) | +0.21 | +0.41 |
| Final-30 record | 20-10 | +1.46 |
| Improvement after the break (runs per game) | +1.72 | +2.54 |
| Bullpen FIP vs. league | +0.57 | +0.57 |
| Rotation FIP vs. league | −0.55 | −2.04 |
| Runs per game vs. league | 100% of average | −0.08 |
| Top-three hitter share (Fernando Tatis Jr., Ty France, Luis Campusano) | 77% | +0.67 |
| Team | Distance | Record | After break | Final 30 | Bullpen | Rotation | Biggest difference | Entered | Result |
|---|---|---|---|---|---|---|---|---|---|
| 2021 Cardinals | 1.69 | 90-72 | 46-26 | 22-8 | +0.25 | −0.22 | Bullpen (−1.0 SD) | Wild Card | Lost Wild Card |
| 2012 Orioles | 1.77 | 93-69 | 48-29 | 20-10 | +0.38 | −0.51 | Improvement after the break (−1.4 SD) | Wild Card | Lost Division Series |
| 2014 Orioles | 2.02 | 96-66 | 44-24 | 20-10 | +0.24 | −0.49 | Full-season strength (+1.1 SD) | Division Series | Lost LCS |
| 2007 Rockies | 2.03 | 90-73 | 46-29 | 22-8 | +0.33 | −0.21 | Offense (+1.2 SD) | Division Series | Lost World Series |
| 2023 Brewers | 2.03 | 92-70 | 43-28 | 18-12 | +0.32 | −0.04 | Rotation (+1.4 SD) | Wild Card | Lost Wild Card |
| 2025 Padres | 2.09 | 90-72 | 38-28 | 16-14 | +0.67 | −0.24 | Improvement after the break (−1.3 SD) | Wild Card | Lost Wild Card |
| 2006 Athletics | 2.18 | 93-69 | 48-26 | 17-13 | +0.64 | −0.15 | Improvement after the break (−1.3 SD) | Division Series | Lost LCS |
| 1996 Cardinals | 2.21 | 88-74 | 42-33 | 19-11 | +0.70 | −0.10 | Improvement after the break (−1.5 SD) | Division Series | Lost LCS |
| Weighting | Shared with main set | Teams |
|---|---|---|
| equal weights | 7 of 8 | 2021 Cardinals, 2012 Orioles, 2007 Rockies, 2014 Orioles, 2023 Brewers, 2025 Padres, 2022 Guardians, 1996 Cardinals |
| strength double | 7 of 8 | 2021 Cardinals, 2012 Orioles, 2023 Brewers, 2007 Rockies, 2025 Padres, 1996 Cardinals, 2006 Athletics, 2022 Guardians |
| without improvement | 3 of 8 | 2012 Orioles, 2024 Guardians, 2018 Brewers, 2016 Orioles, 2007 Cubs, 2021 Cardinals, 2008 Phillies, 1996 Cardinals |
| without final 30 | 7 of 8 | 2021 Cardinals, 2025 Padres, 2012 Orioles, 2007 Rockies, 2023 Brewers, 2014 Orioles, 2022 Guardians, 2006 Athletics |
The 40 nearest teams, with every feature, are available as a download.
Forecast details
| Model | San Diego | 90% model range |
|---|---|---|
| C2: full-season rating | 32.1% | 21.7% to 45.7% |
| C3: recency-weighted rating | 34.7% | 23.0% to 49.7% |
| C4: C3 plus final-30 record | 36.1% | 23.8% to 51.4% |
| C5: C3 plus final-30 underlying runs | 35.3% | 23.7% to 50.5% |
| R1: C3 plus roster quality | 31.0% | Not computed |
| R2: R1 plus final-30 record | 31.4% | Not computed |
Inputs: 2026 regular-season games through September 27. Ratings: Milwaukee +0.94 and San Diego +0.31 runs per game (recency-weighted). Ranges come from 500 refits that resample 2026 games and, for the heat models, whole postseasons. Context, labeled exploratory: in the expanding window, favorites given 60% to 70% won 31 of 54 against an average forecast of 64%.
Media ledger
The opening uses one video and three written lines. Each was checked on the original page; the spoken words come from MLB’s own caption file, not automatic captions. The video embed loads only when a reader asks for it. This isn’t an exhaustive survey of coverage, and syndicated copies count once.
| Item | Kind | Date | Wording |
|---|---|---|---|
| MLB.com video, Pat Murphy | Spoken, from publisher captions | Oct. 2 | I mean, this is a super talented, confident team right now with, we all know the best bullpen in baseball, um, and with some real megastars in the, in the lineup, um, and they're rolling. |
| NBC Sports, D.J. Short | author statement | Sept. 14 | Winners of seven straight, the Padres are the hottest team in baseball. |
| MLB.com, Will Leitch | author statement | Sept. 20 | The top of the NL is tough this year, but no one’s hotter than the Padres. |
| CBS Sports, Matt Snyder | headline | Oct. 2 | Fernando Tatis Jr. is the engine driving the Padres, baseball's hottest team, into NLDS showdown vs. Brewers |
| ESPN, Bradford Doolittle | author statement | Oct. 1 | As hot as San Diego was down the stretch, in terms of team temperature, the Brewers were the hottest team in baseball when the postseason opened. |
| Lead | Why not |
|---|---|
| Baseball Bar-B-Cast: White-hot Padres are 14-4 in their last 18 games | Published Aug. 12; described the 18 games through Aug. 11, not September. |
| AP: a 20-3 surge | Doesn't match the game log. The closest runs are 19-3 over 22 games and 20-4 over 24, both counting the Wild Card games. |
| MLB.com Wild Card picks quoting 'Murphy' on the hottest team | The quote is from Brian Murphy, an MLB.com reporter, not Brewers manager Pat Murphy. |
| Yahoo Sports Daily on-air 'hottest team' line | Headline verified; speaker not verifiable without captions, and the embed doesn't load outside Yahoo. |
Reproducing the results
Everything lives in apps/blog/research/padres-october. From that folder, with the pinned environment in requirements.txt:
python3 -m venv .venv && .venv/bin/pip install -r requirements.txt.venv/bin/python scripts/run_all.pyrebuilds every result from saved inputs, checks the frozen forecasts, runs validation, and writes a run manifest with hashes and seeds..venv/bin/python scripts/refresh_sources.py allrefreshes sources; thenscripts/normalize.pyrebuilds the inputs. A refresh never changes a published number on its own.
The website reads exhibits/essay-data.json. It doesn’t refit anything. Raw MLB Stats API responses stay local; the factual fields the analysis uses are saved in inputs/.
Limitations
- This studies playoff teams in matchups they reached. It says nothing about whether a surge helps a team reach October.
- The forecasts for past matchups are retrospective and conditional on the actual pairings.
- The held-out test has five seasons and 53 matchups. It can’t resolve effects of a few points either way.
- The all-season estimate is observational. It bounds an association; it doesn’t identify a cause.
- Ratings ignore which pitchers start each game. Game outcomes within a series are treated as independent.
- Roster quality omits defense, baserunning, and park, and the usage rules can miss a pitcher, as they did with Henderson.
- The 2026 forecasts were frozen before the Division Series rosters were announced.