Baylee LaneOctober 1, 2026

What a hot finish is worth in October

Methods and further results

The protocol, the data, every prespecified comparison, and the checks behind the essay. Numbers here come from the same saved results file as the essay.

What was fixed, and when

The research protocol (version 1.0) was frozen at Oct. 2, 2026, 23:44 UTC, before any model was scored on postseason outcomes. Its SHA-256 hash begins edc2517981e19677. Amendments were only appended, and runs/protocol-at-freeze.md, the file with them removed, reproduces that hash; the validation script checks it on every run. At that point I had looked at source coverage, field layouts, series counts, home-site patterns, the 2026 Padres game log, standings, and the 2026 postseason schedule. No relationship between a heat measure and a postseason result had been computed.

Two amendments followed, both logged in the protocol file:

  • A1, written after the core results and before any roster result: one recency reference date per season; the roster coefficient’s prior; FIP without its league constant; and the 2026 roster inputs, since neither Division Series roster had been announced.
  • A2, written after the core results: the protocol described the 3-point threshold as roughly home-field advantage in a five-game series. Home field is worth about 1.6 points there. A 3-point change equals about 0.14 runs per game of team quality, about two wins over a season. The threshold itself didn’t change.

The Division Series forecasts were frozen at Oct. 2, 2026, 23:53 UTC and the roster-aware addendum at Oct. 2, 2026, 23:55 UTC. Game 1 was scheduled for Oct. 4, 2026, 00:30 UTC. The freeze files record each forecast’s hash. The timestamps are written by the scripts themselves; a commit would make them independently checkable. The Wild Card Series against the Cubs had already been played, so its model probability (about 49% for San Diego under the recency-weighted model) is retrospective.

Data and reconciliation

Regular-season and postseason games from 1995 to 2025 come from Retrosheet game logs, which include both teams’ batting lines and starting pitchers. The 2026 season, standings, rosters, player lines, box scores, transactions, and probable pitchers come from the MLB Stats API, retrieved October 2 and 3, 2026. The full source ledger, with retrieval times, grain, transformations, and limits, is research/padres-october/sources.csv.

A second parser, written separately from the pipeline, rebuilt every team’s wins, losses, runs scored, and runs allowed from the raw Retrosheet files. All 924 team-seasons from 1995 to 2025 matched MLB’s standings exactly. The 2026 games matched the standings for all 30 teams. The Padres’ box scores sum to their official 722 runs scored and 681 allowed.

Matchups per season, by format
SeasonsFormatMatchups each season
1995 to 2011Division Series, LCS, World Series7
2012 to 2019, 2021Adds two single wild-card games9
2022 to 2025Adds four best-of-three Wild Card Series; top two seeds get byes11
2020Excluded: 60-game season, 16-team bracket0

Each matchup is one row. Team A is the team that hosted game 1. The Division Series used a 2-3 site pattern in 1995 to 1997 and 2012 and 2-2-1 otherwise; the pipeline checks every series’ actual sites against its declared pattern. I traced 17 series by hand across each format change, including all six Padres postseasons in the period:

Manual traces against known results
SeriesSeasonSites playedMatches
Mariners over Yankees, 2-3 site format1995AABBBYes
Cardinals over Padres1996AABYes
Padres over Astros1998AABBYes
Padres over Braves1998AABBBAYes
Yankees over Padres1998AABBYes
Cardinals over Padres2005AABYes
Cardinals over Padres2006AABBYes
Cardinals over Braves, first single-game wild card2012AYes
Giants over Reds, 2-3 site format2012AABBBYes
Nationals over Astros2019AABBBAAYes
Braves over Astros2021AABBBAYes
Padres over Mets, first best-of-three wild card2022AAAYes
Padres over Dodgers2022AABBYes
Phillies over Padres2022AABBBYes
Padres over Braves2024AAYes
Dodgers over Padres2024AABBAYes
Dodgers over Blue Jays2025AABBBAAYes

Definitions

  • Rating. A regularized fit of every regular-season game’s run margin: home runs minus away runs equals home rating minus away rating plus a home constant. Ratings are runs per game against an average opponent, shrunk toward zero by 40 games of average play.
  • Recency weighting. Each game counts half as much for every 120 days before the season’s last day. The half-life and the shrinkage were chosen once, on 1995 to 2014 regular seasons, before any postseason scoring.
  • Final-30 record. Winning percentage over a team’s last 30 regular-season games, ties excluded. Frozen at the end of the regular season; playoff games never count.
  • Underlying runs. BaseRuns: A = H + BB + HBP − HR − 0.5 IBB; B = 1.1 (1.4 TB − 0.6 H − 3 HR + 0.1 (BB + HBP − IBB) + 0.9 (SB − CS − GIDP)); C = AB − H + CS + GIDP; runs = A B / (B + C) + HR. Computed for a team’s batting line and its opponents’.
  • FIP. (13 HR + 3 (BB − IBB + HBP) − 2 K) / innings + a league constant that puts it on the ERA scale.
  • After the break. Games after the All-Star Game date: July 14 in 2026.
  • Hot group. The top third of playoff teams by final-30 record, with cutoffs set on 1995 to 2014 qualifiers (.567 and .633).

Models and scoring

Every model predicts a game: the log-odds that team A wins equal a slope times the rating difference, plus a home edge, plus (for heat models) a coefficient times the standardized heat difference. Series probabilities follow exactly from the game probabilities, the format, and the site pattern, treating games as independent given the inputs. Heat coefficients are fit by penalized maximum likelihood with a normal prior (standard deviation 0.2 per standard deviation of heat).

One choice was made on 2015 to 2019: take the slope and home edge from the postseason fit, or fix them at the regular-season values. The fixed version scored better (Brier 0.2307 against 0.2350 on 45 matchups), so every model uses a slope of 0.475 per run and a home edge of 0.168 in log-odds. Only heat and roster coefficients are fit on postseason games.

The Brier score is the mean squared difference between the probability given to team A and the outcome. Log loss penalizes confident misses more heavily. The calibration slope is 1 when probabilities are spread correctly and below 1 when they’re too spread out.

Held-out results

Fit on 1995 to 2019 (191 matchups), scored on 2021 to 2025 (53 matchups). No model beat always saying 50% on Brier score over this stretch.

Held-out scores, 2021 to 2025
ModelBrierLog lossAccuracyCalibration slope
C0: 50% every time0.25000.693147%0.00
C1: Log5 from full-season records0.26620.728147%-1.05
C2: Full-season rating0.25660.708058%0.03
C3: Recency-weighted rating0.25910.713755%-0.23
C4: C3 plus final-30 record0.25740.709758%-0.20
C5: C3 plus final-30 underlying runs0.25880.712555%-0.36
C2W: C2 plus final-30 record0.25580.706160%0.06
Paired comparisons on the same matchups (negative favors the addition)
ComparisonBrier changeWithin-season resampling, 95%Season resampling, 95%Better / worse
Add final-30 record (C4 vs C3)−0.002−0.008 to +0.005−0.007 to +0.00425 / 22
Add final-30 underlying runs (C5 vs C3)0.000−0.004 to +0.003−0.006 to +0.00424 / 29
Weight recent games more (C3 vs C2)+0.003−0.005 to +0.009−0.003 to +0.01121 / 32
Add final-30 record without recency (C2W vs C2)−0.001−0.004 to +0.003−0.004 to +0.00225 / 22
Rating model vs log5 (C2 vs C1)−0.010−0.023 to +0.004−0.031 to +0.00926 / 27

With only five held-out seasons, the season-resampling interval is rough. Dropping one season at a time, the C4-versus-C3 change ranged from −0.004 to +0.001.

Expanding window, 2005 to 2025

Each season forecast from all earlier normal seasons: 174 matchups. Adding the final-30 record changed the Brier score by +0.002 (95% −0.002 to +0.006); adding underlying runs, by +0.002 (−0.001 to +0.005); recency weighting, by +0.001 (−0.003 to +0.004).

Expanding-window scores, 2005 to 2025
ModelBrierLog lossAccuracyCalibration slope
C2: Full-season rating0.24950.692654%0.27
C3: Recency-weighted rating0.25020.694154%0.21
C4: C3 plus final-30 record0.25220.697852%0.09
C5: C3 plus final-30 underlying runs0.25230.698453%0.01
Out-of-sample matchups won by heat group, expanding window
GroupTeam-matchupsWonForecast expected
cold1135553.7
middle1386869.3
hot975150.9

Size of the effect

The heat coefficient fit on all 244 matchups (1,044 games, 30 postseasons), converted to the change in a 2-2-1 best-of-five probability for an otherwise equal team one standard deviation hotter. Intervals resample whole postseasons (1,000 draws for the final-30 record, 300 for the others). This is estimation on all seasons, not a held-out test.

Points of five-game series probability per standard deviation of heat
Heat measureEstimate95% interval
Final 30 games, record−2.0−6.4 to +2.4
Final 30 games, run differential−2.1−6.0 to +2.6
Final 30 games, underlying runs−1.0−5.2 to +3.0
Final 15 games, record−0.3−3.6 to +3.8
Final 15 games, run differential−0.2−3.2 to +5.0
After the All-Star break, record−3.9−8.1 to +1.3
After the All-Star break, run differential−3.9−7.3 to 0.0
Final-30 record by era and without the Padres
SamplePoints per SDMatchups
1995 to 2011−1.2119
2012 to 2025−2.8125
All seasons, Padres matchups removed−1.2

Power, shuffle, and robustness

Power: I simulated postseason games from a known model with a planted heat effect, ran the full fit-and-score pipeline at four effect sizes, and counted how often each test called it.

Detection rates for planted effects (400 simulations each)
Planted effect, series pointsHeld-out test detectsAll-season estimate detectsMean estimated coefficient
0.04%2%0.004 (true 0.000)
+3.07%26%0.060 (true 0.064)
+6.014%78%0.123 (true 0.129)
+11.833%100%0.247 (true 0.258)

With no planted effect, the tests called one 4% and 2% of the time, and the estimated coefficients tracked the planted ones. These check the pipeline, not baseball.

Shuffle: permuting final-30 records among playoff teams within each training season (200 times) gave a mean coefficient of 0.009, and 42% of the shuffled models improved the held-out score, as expected when the input carries no signal.

Prespecified robustness checks, held-out Brier change for the added input
CheckBrier change95% interval
Window: win15+0.003−0.002 to +0.007
Window: rd30−0.001−0.007 to +0.005
Window: comp300.000−0.004 to +0.003
Window: post_asb_win−0.009−0.022 to +0.005
Window: post_asb_rd−0.007−0.021 to +0.008
Window: excess30−0.001−0.007 to +0.005
Prior scale ×0.5−0.001−0.007 to +0.004
Prior scale ×2.0−0.002−0.009 to +0.005
without padres (47)−0.001−0.008 to +0.006
without extreme streaks (44)−0.002−0.007 to +0.003
first entry only (20)−0.001−0.011 to +0.010
Series-level logistic model0.000−0.006 to +0.006

Roster study

MLB’s dated active rosters are complete from 2009, so this study covers 146 matchups from 2009 to 2025 (292 team rosters, none shorter than 24 players). Roster quality uses the active roster on each matchup’s first day: the nine hitters with the most regular-season plate appearances, the four pitchers with the most starts, and the six other pitchers with the most relief appearances, valued by shrunk wOBA and FIP and converted to runs per game. It correlates 0.58 with the recency-weighted rating.

Roster study, fit 2009 to 2019, scored 2021 to 2025
ComparisonBrier change95% interval
Add roster quality to C3 (R1 vs C3)+0.001−0.009 to +0.011
Add final-30 record to R1 (R2 vs R1)+0.001−0.001 to +0.003

For 2026, the usage rules picked these players. San Diego’s list is its Wild Card Series roster. Milwaukee’s comes from its 36-man active list, because its Division Series roster wasn’t announced; the rules missed Game 2 starter Logan Henderson, who made fewer starts than the four selected.

2026 roster inputs
TeamHittersStartersRelieversRuns per game vs. average
PadresFernando Tatis Jr., Manny Machado, Jackson Merrill, Xander Bogaerts, Ty France, Jake Cronenworth, Gavin Sheets, Freddy Fermin, Miguel AndujarMichael King, Walker Buehler, Robbie Ray, Nick PivettaAdrian Morejon, Mason Miller, Bradgley Rodriguez, Yuki Matsui, Jason Adam, David Morgan+0.38
BrewersBrice Turang, William Contreras, Jake Bauers, Jackson Chourio, Garrett Mitchell, Christian Yelich, Joey Ortiz, David Hamilton, Sal FrelickDustin May, Jacob Misiorowski, Kyle Harrison, Robert GasserAaron Ashby, Trevor Megill, Abner Uribe, Antonio Senzatela, JoJo Romero, Grant Anderson+0.88

Regular-season study

For each season from 1995 to 2014, ratings were fit on the first seven-eighths of the schedule and used to forecast the remaining 6,106 games, holding out one season at a time. A 120-day half-life scored best, narrowly. Adding each team’s last-30 record moved the Brier score from 0.24250 to 0.24256 with the recency-weighted rating, and from 0.24265 to 0.24266 with the full-season rating.

San Diego in detail

2026 Padres by period
MeasureBefore breakAfter breakFinal 30Season
Record48-4843-2320-1091-71
Run differential per game−0.45+1.27+0.87+0.25
Pythagorean win pct..451.626.583.527
One-run games13-1213-107-426-22
Mean opponent rating+0.03−0.01+0.01+0.01
Runs per game3.955.205.204.46
OPS.673.781.768.718
Strikeout rate22.8%19.8%20.7%21.6%
Starters ERA / FIP4.78 / 4.623.92 / 4.884.16 / 4.804.42 / 4.73
Starter innings per game4.614.704.904.65
Bullpen ERA / FIP3.68 / 3.843.31 / 3.263.74 / 3.633.53 / 3.60
Who was hottest depends on the window (wins, run differential, rank by wins)
WindowPadresBrewersMLB leader
Final 15 games12-3, +22, #212-3, +42, #1Brewers
Final 18 games15-3, +32, #215-3, +46, #1Brewers
Final 20 games17-3, +34, #115-5, +40, #3Padres
Final 30 games20-10, +26, #422-8, +54, #1Brewers

Material roster moves: Robbie Ray and Casey Mize acquired Aug. 3; Nick Pivetta activated from the 60-day injured list Sept. 7; Mize placed on the injured list September 26. Source: MLB Stats API transactions.

Comparison teams

Candidates: all 274 playoff teams from 1995 to 2025. Strength, final-30 record, improvement after the break, and offense are standardized within season across all teams; bullpen, rotation, and top-three hitter share are measured against the league and standardized across playoff teams. Distance is weighted Euclidean (weights: rating 1, win30 1, asb_improvement 1, bullpen 1, rotation 1, offense 1, concentration 0.5). Outcomes were joined after the ranking was fixed.

San Diego's profile in standard deviations
TraitValueStandard deviations
Full-season strength (runs per game)+0.21+0.41
Final-30 record20-10+1.46
Improvement after the break (runs per game)+1.72+2.54
Bullpen FIP vs. league+0.57+0.57
Rotation FIP vs. league−0.55−2.04
Runs per game vs. league100% of average−0.08
Top-three hitter share (Fernando Tatis Jr., Ty France, Luis Campusano)77%+0.67
The eight nearest teams
TeamDistanceRecordAfter breakFinal 30BullpenRotationBiggest differenceEnteredResult
2021 Cardinals1.6990-7246-2622-8+0.25−0.22Bullpen (−1.0 SD)Wild CardLost Wild Card
2012 Orioles1.7793-6948-2920-10+0.38−0.51Improvement after the break (−1.4 SD)Wild CardLost Division Series
2014 Orioles2.0296-6644-2420-10+0.24−0.49Full-season strength (+1.1 SD)Division SeriesLost LCS
2007 Rockies2.0390-7346-2922-8+0.33−0.21Offense (+1.2 SD)Division SeriesLost World Series
2023 Brewers2.0392-7043-2818-12+0.32−0.04Rotation (+1.4 SD)Wild CardLost Wild Card
2025 Padres2.0990-7238-2816-14+0.67−0.24Improvement after the break (−1.3 SD)Wild CardLost Wild Card
2006 Athletics2.1893-6948-2617-13+0.64−0.15Improvement after the break (−1.3 SD)Division SeriesLost LCS
1996 Cardinals2.2188-7442-3319-11+0.70−0.10Improvement after the break (−1.5 SD)Division SeriesLost LCS
Stability: how many of the eight survive other weightings
WeightingShared with main setTeams
equal weights7 of 82021 Cardinals, 2012 Orioles, 2007 Rockies, 2014 Orioles, 2023 Brewers, 2025 Padres, 2022 Guardians, 1996 Cardinals
strength double7 of 82021 Cardinals, 2012 Orioles, 2023 Brewers, 2007 Rockies, 2025 Padres, 1996 Cardinals, 2006 Athletics, 2022 Guardians
without improvement3 of 82012 Orioles, 2024 Guardians, 2018 Brewers, 2016 Orioles, 2007 Cubs, 2021 Cardinals, 2008 Phillies, 1996 Cardinals
without final 307 of 82021 Cardinals, 2025 Padres, 2012 Orioles, 2007 Rockies, 2023 Brewers, 2014 Orioles, 2022 Guardians, 2006 Athletics

The 40 nearest teams, with every feature, are available as a download.

Forecast details

Frozen probabilities that San Diego wins the Division Series
ModelSan Diego90% model range
C2: full-season rating32.1%21.7% to 45.7%
C3: recency-weighted rating34.7%23.0% to 49.7%
C4: C3 plus final-30 record36.1%23.8% to 51.4%
C5: C3 plus final-30 underlying runs35.3%23.7% to 50.5%
R1: C3 plus roster quality31.0%Not computed
R2: R1 plus final-30 record31.4%Not computed

Inputs: 2026 regular-season games through September 27. Ratings: Milwaukee +0.94 and San Diego +0.31 runs per game (recency-weighted). Ranges come from 500 refits that resample 2026 games and, for the heat models, whole postseasons. Context, labeled exploratory: in the expanding window, favorites given 60% to 70% won 31 of 54 against an average forecast of 64%.

Media ledger

The opening uses one video and three written lines. Each was checked on the original page; the spoken words come from MLB’s own caption file, not automatic captions. The video embed loads only when a reader asks for it. This isn’t an exhaustive survey of coverage, and syndicated copies count once.

Items used
ItemKindDateWording
MLB.com video, Pat MurphySpoken, from publisher captionsOct. 2I mean, this is a super talented, confident team right now with, we all know the best bullpen in baseball, um, and with some real megastars in the, in the lineup, um, and they're rolling.
NBC Sports, D.J. Shortauthor statementSept. 14Winners of seven straight, the Padres are the hottest team in baseball.
MLB.com, Will Leitchauthor statementSept. 20The top of the NL is tough this year, but no one’s hotter than the Padres.
CBS Sports, Matt SnyderheadlineOct. 2Fernando Tatis Jr. is the engine driving the Padres, baseball's hottest team, into NLDS showdown vs. Brewers
ESPN, Bradford Doolittleauthor statementOct. 1As hot as San Diego was down the stretch, in terms of team temperature, the Brewers were the hottest team in baseball when the postseason opened.
Leads checked and not used
LeadWhy not
Baseball Bar-B-Cast: White-hot Padres are 14-4 in their last 18 gamesPublished Aug. 12; described the 18 games through Aug. 11, not September.
AP: a 20-3 surgeDoesn't match the game log. The closest runs are 19-3 over 22 games and 20-4 over 24, both counting the Wild Card games.
MLB.com Wild Card picks quoting 'Murphy' on the hottest teamThe quote is from Brian Murphy, an MLB.com reporter, not Brewers manager Pat Murphy.
Yahoo Sports Daily on-air 'hottest team' lineHeadline verified; speaker not verifiable without captions, and the embed doesn't load outside Yahoo.

Reproducing the results

Everything lives in apps/blog/research/padres-october. From that folder, with the pinned environment in requirements.txt:

  • python3 -m venv .venv && .venv/bin/pip install -r requirements.txt
  • .venv/bin/python scripts/run_all.py rebuilds every result from saved inputs, checks the frozen forecasts, runs validation, and writes a run manifest with hashes and seeds.
  • .venv/bin/python scripts/refresh_sources.py all refreshes sources; then scripts/normalize.py rebuilds the inputs. A refresh never changes a published number on its own.

The website reads exhibits/essay-data.json. It doesn’t refit anything. Raw MLB Stats API responses stay local; the factual fields the analysis uses are saved in inputs/.

Limitations

  • This studies playoff teams in matchups they reached. It says nothing about whether a surge helps a team reach October.
  • The forecasts for past matchups are retrospective and conditional on the actual pairings.
  • The held-out test has five seasons and 53 matchups. It can’t resolve effects of a few points either way.
  • The all-season estimate is observational. It bounds an association; it doesn’t identify a cause.
  • Ratings ignore which pitchers start each game. Game outcomes within a series are treated as independent.
  • Roster quality omits defense, baserunning, and park, and the usage rules can miss a pitcher, as they did with Henderson.
  • The 2026 forecasts were frozen before the Division Series rosters were announced.