October Terms 2005 to 2025
The Chief Justice counted questions. He was asking the wrong one.
The shortcuts we use on this Court work. They just answer a smaller question than the one we asked. Counting questions tells you where the bench is leaning. People use it to call the winner. Who appointed a justice tells you how a vote gets labeled. It doesn’t tell you how that justice voted. An hour of argument can read one person. It can’t call the case.
The rule holds up, and it still loses
In 2005, John Roberts counted the questions in 28 Supreme Court arguments. The losing side, he found, was almost always asked more. He published it with a joke: the secret to winning is getting the Court to ask your opponent more questions. Months later he was Chief Justice. I reran the count on every argued case since October Term 2005. The pattern is real. It still loses to a guess you can make before anyone speaks: the Court reverses.
Across 1,110 cases where we know who won, the side the bench spent more words on lost more often than it won. I counted it six different ways: turns instead of words, dropping the amicus sections. Every version still beats a coin flip.
So Roberts saw something true. The bench really does lean on the eventual loser. Can you use that to pick winners? That takes a number his paper never mentioned.
Any prediction has to beat the guess you’d make knowing nothing. That free guess is called the base rate, and here it’s lopsided. A petitioner is at the Court because a lower court ruled against them and the justices agreed to take another look. They could have said no. Saying yes isn’t neutral: by the first minute of argument, the petitioner is already the two-to-one favorite.
Always guessing the petitioner calls ~66% of these same cases. Counting questions calls ~61%. Run the two head to head, case by case, and the free guess wins by a margin too big to be luck.
Exhibit 1. Every version of the question-counting rule loses to the free guess
Read this as a table
| Version | Cases | Correct | Accuracy |
|---|---|---|---|
| Words, all speakers | 1110 | 672 | 60.5% |
| Words, principal advocates only | 1112 | 700 | 62.9% |
| Words, excluding amicus sections | 743 | 461 | 62.0% |
| Turns rather than words | 1081 | 625 | 57.8% |
| Unanimous decisions only | 475 | 309 | 65.1% |
| Divided decisions only | 635 | 363 | 57.2% |
| Always predict the petitioner | 1110 | 737 | 66.4% |
That’s the bar I use for everything here, including my own model. A rule can detect something real and still be worth less than the free guess. “Almost always” sounds like plenty until you know what guessing gets you.
“Who wins” is one number per case. The Court doesn’t give you one number. It gives you nine.
Every vote answers two questions
I scored every argued case since October Term 2005. Transcripts come from Oyez, which publishes the audio and the speaker-labeled text. Outcomes and votes come from the Supreme Court Database.
Any single vote can be scored two ways.
The first is the one people argue about. Was this a conservative vote or a liberal one? Somebody has to decide that. Spaeth, Epstein, Nelson, Martin, Segal, Ruger, and Benesh code each decision against a published rulebook: a vote for the criminal defendant is liberal, a vote for the business against the regulator is conservative, and so on. It’s a judgment call. Every ideology number below rests on it.
The second question sounds like paperwork. Which of the two lawyers in the room did the justice side with? No interpretation is needed, because the Court’s own judgment says who won.
Throwing out a lower-court ruling isn’t conservative or liberal by itself. It depends entirely on what that court did. The same data answers the two questions in opposite ways.
The exhibit below is every vote I have, one box per justice per argued case, twenty years wide. Rows are justices and columns are cases in the order they were heard. Use the control to switch between the two questions.
Score by ideology and the bench sorts into two horizontal blocks you could have drawn from memory. Switch to which lawyer each justice sided with and the blocks dissolve into vertical stripes. A vertical stripe means a whole column voted the same way, which is what near-unanimity looks like. The only thing that changed is the question.
Exhibit 2. Change the question and the party blocs disappear
Who appointed a justice predicts the label, not the vote
If you want to guess how a justice votes, the free answer is the president who appointed them. It’s in every news story about the Court, usually as the whole explanation. If you’re asking how the vote gets labeled conservative or liberal, it belongs there. Across 7,956 votes in cases with at least one dissent, party calls the ideological direction ~69% of the time, against 50% for a coin flip.
That’s a count of how the Database coded the votes. When I get to the argument scores, those come from a model.
Every ideology figure here uses only those divided cases, because a unanimous decision gives nobody a chance to break from a label.
People treat that label as if it got stronger over twenty years. It didn’t. Two justices left. The people who stayed didn’t get more predictable. Take John Paul Stevens and David Souter out of the average and the climb is gone. No sitting justice drifted toward their label. The bench was replaced.
Exhibit 3. Two justices voted against their own label the whole time they sat
0%coin flip100%
Read this as a table
| Justice | Appointed by | Terms | Divided votes | Label correct |
|---|---|---|---|---|
| Stevens | Gerald Ford | OT2005 to OT2009 | 258 | 19.4% |
| Souter | George H. W. Bush | OT2005 to OT2008 | 208 | 27.4% |
| Kennedy | Ronald Reagan | OT2005 to OT2017 | 602 | 59.5% |
| Roberts | George W. Bush | OT2005 to OT2024 | 893 | 63.2% |
| Gorsuch | Donald J. Trump | OT2016 to OT2024 | 343 | 64.4% |
| Barrett | Donald J. Trump | OT2020 to OT2024 | 201 | 65.7% |
| Kavanaugh | Donald J. Trump | OT2018 to OT2024 | 294 | 66.3% |
| Breyer | Bill Clinton | OT2005 to OT2021 | 776 | 68.3% |
| Scalia | Ronald Reagan | OT2005 to OT2015 | 493 | 71.0% |
| O'Connor | Ronald Reagan | OT2005 to OT2005 | 7 | 71.4% |
| Kagan | Barack Obama | OT2009 to OT2024 | 619 | 74.0% |
| Ginsburg | Bill Clinton | OT2005 to OT2019 | 696 | 75.7% |
| Alito | George W. Bush | OT2005 to OT2024 | 873 | 77.8% |
| Sotomayor | Barack Obama | OT2009 to OT2024 | 683 | 78.5% |
| Thomas | George H. W. Bush | OT2005 to OT2024 | 896 | 79.0% |
| Jackson | Joe Biden | OT2022 to OT2024 | 114 | 81.6% |
Stevens was appointed by Gerald Ford, and in divided cases the label built from that fact was right ~19% of the time. A coin flip is 50%. Souter was appointed by George H. W. Bush and is in the same place. The label pointed the wrong way.
This only asks how well the appointing president predicts the direction of a vote. On that question, the rate is flat. It says nothing about whether the Court’s decisions got more consequential, whether the cases it agrees to hear have changed, or whether a six to three majority behaves differently from a five to four. Those have changed. This can’t see them.
The argument reads one justice
Now the other question: which lawyer will each justice side with? I let four things compete: nothing at all, the appointing party, the justice’s own voting record, and the argument.
“The argument” means four measurements taken off the transcript, and no names, parties, or votes. Which lawyer the justice spent words on, how far that lean sat from the rest of the bench that day, how much the justice talked relative to their own habit, and how fast. The model doesn’t see the Court’s own decision. It only sees the transcript.
Every score below is out of sample: each vote is called by a model trained only on earlier terms, so the model has never seen the vote it’s predicting. A model graded on cases it studied would just be memorizing.
Ask which side a justice backs. Party, the justice’s own record, and a guess that knows nothing all land on ~62%. That’s the base rate again. Knowing who appointed you says nothing about whether you’ll throw out a Ninth Circuit tax ruling. The argument gets ~65%.
That’s only a few points, so I wanted to know how firm it was. I scored them with a stricter measure, one that punishes confident mistakes hardest, and reshuffled the same 1,201 cases. The argument beats the justice’s own record every time.
Ask the ideology question and it goes the other way. The president who appointed a justice and the entire oral argument land on the same number. An hour of listening tells you nothing a Wikipedia page wouldn’t. Adam Feldman reached the same split independently on a single term while I was fitting this. His numbers are in Sources.
That’s the average across the whole bench. Look at justices one at a time and some of them break it. Test 15 justices one at a time and roughly one will look remarkable by pure chance. So I set the bar higher, high enough that the whole set of tests still has the usual odds of a false alarm. Only Stevens, Gorsuch and Thomas clear it. The other bars in the exhibit are noise.
Exhibit 4. For one justice, the room reads him better than his own voting record
-35 pointsno difference+35 points
Read this as a table
| Justice | Appointing party | Own record | The argument | Argument minus record | Survives correction | Votes |
|---|---|---|---|---|---|---|
| Stevens | 19.7% | 80.3% | 59.8% | -20.5 points | yes | 122 |
| Souter | 27.5% | 72.5% | 61.5% | -11.0 points | no | 91 |
| Kennedy | 56.9% | 56.9% | 54.1% | -2.8 points | no | 399 |
| Roberts | 61.5% | 61.5% | 61.1% | -0.3 points | no | 646 |
| Gorsuch | 62.9% | 60.4% | 75.5% | 15.0 points | yes | 245 |
| Barrett | 65.8% | 59.9% | 55.9% | -3.8 points | no | 152 |
| Kavanaugh | 65.9% | 59.6% | 63.2% | 3.6 points | no | 223 |
| Scalia | 67.1% | 67.1% | 68.3% | 1.2 points | no | 319 |
| Breyer | 67.4% | 67.4% | 69.0% | 1.7 points | no | 533 |
| Kagan | 73.5% | 70.8% | 73.1% | 2.2 points | no | 490 |
| Ginsburg | 73.7% | 73.7% | 70.4% | -3.4 points | no | 476 |
| Thomas | 77.2% | 77.2% | 59.9% | -17.3 points | yes | 167 |
| Alito | 77.3% | 77.3% | 74.4% | -2.8 points | no | 594 |
| Sotomayor | 77.8% | 76.3% | 72.7% | -3.7 points | no | 545 |
| Jackson | 79.5% | 79.5% | 79.5% | 0.1 points | no | 88 |
Gorsuch’s questions predict his vote at ~76%, against ~60% for his own voting history.
That margin isn’t one strange term. The room reads him better than his record in seven of nine terms.
Clarence Thomas goes the other way. Once he speaks, his questions point away from his vote.
The third one who clears the bar is John Paul Stevens, and the gap looks bigger than it is. His own record is close to unbeatable because he voted one way almost every time. Stevens wasn’t hard to read. He was easy to predict for anyone who watched him instead of reading his résumé.
I can’t tell you why the room works for Gorsuch and not the other 14. My reading is that he argues his way toward a position in public, so a measure of where his attention goes tracks something real. But that’s me interpreting a number. I can’t test it.
Add the nine votes together and the edge disappears. Across 1,201 cases, the model and the one-word guess “the Court reverses” land within a point of each other. Run both rules on the same cases and you can’t tell them apart.
If you want the likely winner, the best clue arrives before the argument begins. If you want to know how a named justice is leaning, the argument is the only thing I tested that helps.
I published predictions for 22 undecided cases from October Term 2025, timestamped before the outcomes, in the forecast file. If a case sits near 50%, the model has nothing to say. That’s not the same as the case being a toss-up.
The pattern is real. It starts too late.
Roberts saw a real pattern. You still can’t use it to pick winners. By the time anyone asks a question, the Court has already tipped the odds by agreeing to hear the case.
We do the same thing with the party label. We talk about justices as if the president who appointed them is the whole story. Stevens voted liberal for thirty years in public and the label survived anyway.
Method
Vote direction, majority membership, area of law, and lower-court direction are coded by the Supreme Court Database, not by me. Speaking counts, timestamps, and quoted turns are observed from Oyez transcripts. Model scores are modeled. The reading of Gorsuch is my interpretation and is marked as such where it appears.
Two questions, two answers
This piece scores two different outcomes on one data set and they behave in opposite ways. Vote direction (conservative or liberal) is predicted well by appointing party and not improved by the argument. Vote for the petitioner (which side the justice backed) isn’t predicted by party at all and is predicted by the argument. Both are reported, and both run through the same fitting code, so a difference between them can’t be an artifact of the routine.
Two windows, and why they differ
Direction measures cover October Terms 2005 to 2024, the end of Supreme Court Database release 2025 Release 01. Argument measures and the petitioner-vote model run through October Term 2025 using Oyez outcomes, because the database hasn’t coded that term. The two aren’t comparable to each other and are never placed on one axis.
Why an earlier draft said the Court reverses 77% of its cases
An earlier version reported the petitioner win rate as 77%. That number was a property of a join, not of the Court: it counted only the cases where Oyez’s free-text winner field resolves by name matching, and whether it resolves is correlated with who won. Adding the cases the Database settles brings the rate to 66%, in line with Exhibit 1’s 66.4%. The per-term selection table is in the ledger.
Why the same measure carries two numbers
The party label appears at two rates: 68.7% described across all 7,956 divided-case votes, and 67.6% scored as a prediction on the narrower set a model can be tested against. The difference is the denominator. The petitioner base rate also appears twice, at 66.4% on the Exhibit 1 case set and 67.4% on the narrower rolling-origin set; the sets differ and the two figures never share an axis.
Held-out design
Rolling origin throughout: each term is scored by a model trained only on the terms before it, with feature means, standardization, and every justice-level prior computed inside each fold’s training rows. The argument model never sees a justice’s voting record. It isn’t blind to identity: one feature compares a justice’s word share to their own typical share, which measures departure from habit and carries no information about which way they vote.
October Term 2025 is additionally held out end to end and scored against a label source the training data never used. Oyez and the Database agree on 96.0% of the 9,771 votes where both exist, so that residual is a ceiling on the held-out accuracy, and the held-out gap is smaller than the residual. Same shape as the rolling test: 59% for party and for own record, 62% for the argument. One term is too small to prove the gap.
Calibrating the forecast
The published forecast probabilities are Platt-calibrated: a logistic map fitted on the rolling-origin predictions only, never on October Term 2025, applied at two grains. The per-justice probabilities are mildly over-extreme and the case-level aggregates badly so. The difference between those two corrections is a measurement of how wrong the independence assumption is.
Intervals and multiple comparisons
Single rates carry 95% Wilson intervals. Differences between models carry a bootstrap that resamples whole cases rather than individual votes, because justices hearing one case mostly share an answer. Two rules scored on the same cases are compared with a paired interval on the difference, since two separate intervals mostly share sampling noise. With 15 justices tested, about one apparently significant result is expected by chance, so claims about a named justice rest on the family-wise corrected interval.
Known limits
Oyez leaves decision records null for a large share of some terms, the largest single exclusion from every outcome test here, and its winning-party field is free text resolved by name matching. The argument measures are word counts and timings; they don’t represent an understanding of what was said. On the direction grain, for most justices the party and own-record baselines collapse to the same constant, so beating the record there is a weaker comparison than it sounds.
None of this is cause and effect. A justice probably presses the side they already doubt, so the questioning and the vote most likely share a cause rather than one producing the other. Nothing here shows an argument changing a mind, and nothing here shows that being the petitioner makes a side more likely to win.
Sources
- Vote direction, majority membership, issue area, party winning, and lower-court direction from the Supreme Court Database, release 2025 Release 01, justice-centered data organized by docket (Spaeth, Epstein, Nelson, Martin, Segal, Ruger & Benesh). 14,792 justice-votes across 1,655 cases.
- Case and transcript payloads from the Oyez API, October Terms 2005 through 2025, 1,408 arguments. Oyez publishes no API documentation, so field meanings are taken from observed responses. Oyez material is licensed CC BY-NC.
- John G. Roberts Jr., “Oral Advocacy and the Re-emergence of a Supreme Court Bar”, Journal of Supreme Court History 30 (2005), 68 to 81, for the 28-case question count: the first and last case of each argument sitting in the 1980 and 2003 terms. Secondary accounts report his finding as “the losing side was almost always asked more questions” without publishing a hit rate, so no specific success count is attributed to him here.
- Sarah Levien Shullman, “The Illusion of Devil’s Advocacy”, Journal of Appellate Practice and Process 6 (2004), ten arguments scored by hostility, which preceded Roberts by a year. Lee Epstein, William M. Landes and Richard A. Posner, “Inferring the Winning Party in the Supreme Court from the Pattern of Questioning at Oral Argument,” Journal of Legal Studies 39 (2010), on 1979 to 1995.
- Adam Feldman, “How predictable is the Supreme Court from oral argument?” SCOTUSblog, July 27, 2026. Reaches the vote-versus-case distinction independently, working from one term rather than twenty: 56 decisions and 495 votes, a word-imbalance rule calling 64.2% of individual votes and 58.9% of case outcomes, and a justice-specific version at 66.5% and 66.1%. He finds it strongest for Gorsuch, Kagan and Jackson and near chance for Roberts and Barrett. Two people who didn’t coordinate landing in the same place is worth more than either result alone.
- Argument format history from the Supreme Court’s press release of April 28, 2020 and its Guide for Counsel, October Terms 2021 through 2024.
- Code and data: github.com/baylee1lane/BayleeLaneBlog. Forecast file: forecast-v1.json. Measures version direction-v2, argument measures argument-measures-v1, direction gate argument-gate-v1, outcome holdout outcome-holdout-v1, forecast forecast-v1.