Supreme Court · October Terms 2005 to 2025
What the Supreme Court’s questions reveal
The questions helped predict individual votes. They didn’t make my model better at picking the winner.
I spent the summer listening to Supreme Court oral arguments on walks and in the car. I debated in high school, and I still enjoy hearing someone defend an argument under pressure. I wanted to know whether the justices’ questions told me who would win.
In Bowe v. United States, Neil Gorsuch presses a lawyer to explain why two groups of prisoners should be treated differently.
Neil Gorsuch · Bowe v. United States
But can --but can -- can you think of -- and maybe -- maybe the answer is you can't on your feet, and I get it -- any reason why pre-conviction prisoners would be barred, but post-conviction prisoners wouldn't be?
I matched twenty years of transcripts to the justices’ votes and tested whether exchanges like this helped predict the outcome.
The questions improved my predictions of individual votes. Most of those gains came in cases where I was already picking the same winner. Adding the questions didn’t give me a better case prediction than always choosing the petitioner.
Picking the petitioner set the bar
The petitioner is the side asking the Court to review a lower court’s ruling; the respondent defends it. Petitioners won about two-thirds of the cases I tested. That was the prediction my model needed to beat.
John Roberts tested a similar idea before becoming Chief Justice. In a 2005 article, he described counting questions in 28 cases. The side asked more usually lost. Sarah Levien Shullman had studied questioning tone the year before. Lee Epstein, William Landes, and Richard Posner later tested the relationship on larger samples.
I used Oyez’s transcripts, which identify each speaker, to count how many words the justices addressed to each side. The collection covers 1,408 arguments and 128,223 substantive speaking turns by justices, from October Terms 2005 through 2025. An October Term is the Court’s working year, beginning in October.
The rule picked the side that received fewer words. Counting words gives a long exchange more weight than a short question, so I also checked speaking turns and exchanges with the main lawyers.
My first count overstated how often petitioners won. Oyez records winners by name, and my code dropped cases when it couldn’t match the name to a side. That removed too many petitioner losses. Filling the gaps with Supreme Court Database outcomes brought the win rate from about 77% to about 66%.
On the same 1,110 cases, counting words got about 61% right. Always choosing the petitioner got about 66% right.
Exhibit 1. A simple benchmark beats counting words
Read this as a table
| Rule | Cases | Correct | Accuracy |
|---|---|---|---|
| Petitioner rule | 1110 | 737 | 66.4% |
| Word-counting rule | 1110 | 672 | 60.5% |
Changing which source supplied the outcome changed the word-counting result by just one correct prediction. The side questioned more did tend to lose. That pattern still gave me worse predictions than picking the petitioner without listening.
Predicting each justice’s vote
Counting words across the Court loses who asked them. One justice can press the petitioner while another presses the respondent. Their questions may point to different votes, even when the total looks balanced.
For each justice, I wanted to predict whether they would support the petitioner or respondent. Knowing whether a justice usually votes liberal or conservative doesn’t answer that on its own.
A defendant can ask the Court to overturn a conviction. A prosecutor can ask it to reverse a ruling favoring the defendant. A justice who favors defendants in both cases would support the petitioner in one and the respondent in the other.
The map shows the same votes both ways. The Supreme Court Database’s ideological coding shows which justices usually vote liberal or conservative. Switching to case side shows who agreed within each case. The ideological labels describe outcomes under a published rulebook, not the justices’ motives.
Exhibit 2. The same votes look different when the question changes
My starting model used each justice’s earlier votes and appointing president’s party, along with the court the case came from, how it reached the Supreme Court, and whether the United States was named as a party. That let it account for differences in a justice’s record across types of cases.
The model didn’t read the briefs or assess the legal merits. It used selected background facts reconstructed from historical records.
What the questions added
Adding the argument gave the model each justice’s word balance between the two sides, their speaking share relative to the other justices and their own history, and the length and pace of their turns. These were counts and timing. The model didn’t interpret what anyone said.
I trained on earlier terms and tested on the next one, repeating the comparison from October Term 2007 through 2024. Every model was scored on the same 8,878 votes in 1,136 cases. The first two terms supplied the initial training data.
Adding the questions raised vote accuracy from about 61% to 65%. After subtracting the predictions it turned from right to wrong, the model called 333 more votes correctly.
Exhibit 3. The questions improve individual vote predictions
Read this as a table
| Model | Votes | Correct | Accuracy | Log loss | Brier score |
|---|---|---|---|---|---|
| Always petitioner | 8878 | 5544 | 62.4% | 0.6623 | 0.2347 |
| Before argument | 8878 | 5429 | 61.2% | 0.6779 | 0.2387 |
| Argument alone | 8878 | 5767 | 65.0% | 0.6306 | 0.2192 |
| Before + argument | 8878 | 5762 | 64.9% | 0.6509 | 0.2241 |
The extra background detail didn’t earn its place. Always choosing a petitioner vote got about 62% right, ahead of the background model. The argument alone reached about 65%, as good as the combined model. The questions carried the useful information in this comparison.
The questions also improved a score that penalizes confident mistakes. That gain held when I resampled whole decisions or whole terms. The vote improvement appeared in both halves of the testing period and with different limits on model complexity.
The counts still missed what made an exchange interesting to listen to. In Cox Communications, Gorsuch asks whether the Court can resolve the dispute more narrowly.
Neil Gorsuch · Cox Communications, Inc. v. Sony Music Entertainment
We don't have to go that far, though, to recognize that the jury instructions here were improper, do we?
He sounds to me like he’s working through an answer aloud. My model counted how much he spoke. It didn’t understand the distinction he was asking the lawyer to make.
Most of the improvement left the winner unchanged
To pick the winner, I combined the individual vote probabilities into a prediction of the majority. Then I compared the calls before and after adding the questions.
In 964 cases, the predicted winner stayed the same. Those cases accounted for 316 of the 333 additional votes called correctly. Almost all the improvement was in who I expected to support each side, without changing which side I picked.
Exhibit 4. Most of the additional correct votes left the winner unchanged
Read this as a table
| Winner prediction | Cases | Votes corrected | Votes spoiled | Net vote gain | Cases corrected | Cases spoiled |
|---|---|---|---|---|---|---|
| Winner stayed the same | 964 | 889 | 573 | 316 | 0 | 0 |
| Winner changed | 172 | 348 | 331 | 17 | 80 | 92 |
The winner changed in 172 cases. Those changes fixed 80 calls and turned 92 from right to wrong. Even in those cases, the model gained 17 correct individual votes. It got more votes right and more winners wrong.
Combining the votes assumed they were independent once the model had accounted for each justice’s measured behavior. Many cases also lacked measurements for the full bench. Limiting the test to cases with complete recorded coverage didn’t establish an improvement in picking winners.
Predicting the winner directly didn’t solve it
I then built a model that predicted the winner directly, using the case background and summaries of the justices’ questioning. This removed the need to combine separate vote probabilities into a majority.
On the same 1,136 cases, always picking the petitioner got about 68% right. The direct model with questions got about 66% right. Adding the questions didn’t establish an overall improvement over the background model in either winner accuracy or the probability score.
Exhibit 5. The direct model doesn’t beat picking the petitioner
-5.0 pointsno improvement3.0 points
Read this as a table
| Model or comparison | Cases | Accuracy or difference | Paired 95% interval |
|---|---|---|---|
| Always petitioner | 1136 | 67.6% | |
| Before argument | 1136 | 66.5% | |
| Before + argument | 1136 | 65.8% | |
| Over background model | 1136 | -0.8 points | -2.6 to 0.9 percentage points |
| Over petitioner rule | 1136 | -1.8 points | -3.8 to 0.1 percentage points |
When this model disagreed with the petitioner rule, it fixed 51 calls and turned 72 from right to wrong. Those were the cases where I needed it to add value. It didn’t.
So do I trust this model enough to put money behind it on Kalshi or Polymarket? I’ll stick to listening on walks and in the car. I still enjoy a good argument, even when I can call the winner two-thirds of the time.
Methods, further results, and sources
The companion notes contain the model definitions, exact results, uncertainty estimates, source audit, transcript explorer, corrections, and reproducible code.
Correction: the earlier Gorsuch prediction used a field derived from the Court’s eventual decision. I’ve withdrawn that result. The tests above exclude the field. Details are in the notes.
- Oyez: argument recordings and transcripts.
- Supreme Court Database: vote and case coding, release 2025_01.
- John G. Roberts Jr., “Oral Advocacy and the Re-emergence of a Supreme Court Bar” (2005).
- Sarah Levien Shullman, “The Illusion of Devil’s Advocacy” (2004).
- Lee Epstein, William M. Landes, and Richard A. Posner, “Inferring the Winning Party in the Supreme Court from the Pattern of Questioning at Oral Argument” (2009 working paper; published 2010).
- Code and saved results.