Here is the new sound of a big-league ballpark: a hitter taps his helmet, the crowd leans in, a light blinks, and a strike three is quietly turned into ball four — or isn’t. The ABS challenge is a tidy piece of theater, and like everything in baseball it immediately became a leaderboard. ESPN ranks the best challenging teams; broadcasts flash a hitter’s challenge record; the word “challenge IQ” is now a thing people say. We wanted to know whether any of it is a skill — whether some players and teams are genuinely good at picking their moment, or whether we’re just watching a coin land heads a few times in a row and calling it a hot hand.
So we pulled every challenge in the majors this year — 5,843 of them, through the All-Star break — with the one thing the leaderboards leave out: exactly where each challenged pitch crossed, measured against the strike zone’s edge. That last detail is what lets you separate a real read from a lucky one.
1. Winning a challenge is close to a coin flip
Start with the number that should cool everyone’s jets: league-wide, 53% of challenges are overturned. Just over half. That is roughly what you’d expect from a system where players, being sensible, only challenge pitches they think were close — and close pitches are, by definition, the ones the machine and the umpire are most likely to disagree about. You don’t challenge the pitch down the middle. You challenge the one on the black, and the one on the black is a coin flip.
Which means an individual’s challenge record is built on a tiny pile of coin flips. The median player has made a handful of challenges all year. When we ask whether the catchers who’ve challenged the most are reliably better than each other — the real test of a skill — the answer is: barely. Plot every catcher against the range pure luck would produce, and almost all of them sit inside it.
Each dot is a catcher; the shaded band is where chance alone would scatter them around the league average, given how many challenges each one made. If “good challenger” were a real individual trait, dots would spill well outside the band. They mostly don’t. Only about 24% of the spread between the “best” and “worst” catcher-challengers is real; for hitters it’s 8%. The rest is the small-sample bounce. This is the same lesson as our Coin-Flip Line: a rate measured over a dozen events is mostly noise, so the leaderboard it feeds is mostly noise too.
2. But who challenges is the whole game
Here is where the coin turns out to be weighted. Sort those same challenges not by the person but by the position — who actually tapped for the review — and a gap opens up that dwarfs anything between individuals.
The catcher overturns 59% of the calls he challenges. The hitter, 48% — a touch worse than a coin flip. The pitcher, a dismal 36%, and he should mostly stop. These aren’t close: on 5,843 challenges the intervals don’t come near overlapping. The single biggest lever in the entire challenge game isn’t who your best challenger is. It’s making sure the catcher is the one holding the tap.
The reason isn’t mysterious, and that’s what makes it convincing. The catcher is crouched eighteen inches behind the plate, tracking the ball into his own glove, watching it cross the exact plane the zone is drawn on. The hitter is standing off to the side, six feet away, watching a 95-mph pitch move away from where his eyes started. One of them has the best view of the strike zone in the building. The other has one of the worst. The scoreboard just made that difference legible.
3. Catchers have judgment. Hitters have conviction.
If the catcher’s edge is really about seeing, it should show up in a specific way: he should be better at telling an obvious miss from a genuine coin-flip. And this is the part of the data we didn’t expect to be so clean. For every challenge, we know how far the pitch actually was from the edge of the zone — a pitch two inches off is a much clearer miss than one grazing the line. So: as the pitches get more obviously wrong, does the challenger cash in?
The catcher’s line climbs — from a coin flip on the tightest pitches up to nearly 70% on the clear ones. He knows an obvious miss when he sees it, and he wins those. The hitter’s line is flat, and if anything droops: his “obvious” challenges succeed no more often than his desperate ones. That’s the whole story of the ABS challenge in two lines. The catcher has judgment — his confidence tracks reality. The hitter has conviction — the batter who stalks back to the box certain he got robbed is, more than half the time, simply wrong.
It also quietly sharpens a finding of our own. We once showed that the best challengers are catchers. True — but now we can say why, and what it does and doesn’t mean. It isn’t that certain catchers have a gift. It’s that the catcher’s chair has a gift, and any of them sitting in it inherits most of it.
What this means for tonight’s game
The next time a broadcast crowns someone the team’s challenge ace on the strength of a 7-for-9 record, remember what nine coin flips look like. Individual challenge records are almost all bounce; a hot one tells you little about the next one. What actually repeats is duller and more useful: a team’s habits — how often it challenges, how close to the edge it’s willing to go — carry from month to month far more than its win rate does (a stable style, a coin-flip result). And the one durable, sizable edge on offer is free and structural: let the catcher decide. He has the best seat in the house. Everyone else is guessing, and the hitter who’s surest is guessing worst.
Methodology
How we built and stress-tested this
Data. Every ABS challenge in the 2026 MLB regular season from opening day through July 8 (the All-Star break) — 5,843 challenges — pulled per-game from Baseball Savant’s gamefeed, each with the challenger’s role, the initial and final call, and the pitch’s distance from the strike-zone edge. Overturn = the umpire’s original call was reversed by ABS.
Is winning a skill? For catchers and hitters with at least eight challenges, we compared the observed spread in individual overturn rates to the spread pure binomial chance would produce given each player’s challenge count. Reliability is the share of the observed variance that survives after subtracting that sampling noise: 0.24 for catchers (n=77), 0.08 for hitters (n=135). Low reliability means the leaderboard mostly reflects small-sample luck, not a durable individual skill — the events-per-measurement lens from The Coin-Flip Line.
Position and judgment. Overturn rates by challenger role carry Wilson 95% intervals; catcher (58.8%, [57.0, 60.6]) and hitter (47.8%, [45.9, 49.7]) do not overlap on n=3,013 and n=2,727. The obviousness curves bin challenges by the pitch’s distance from the zone edge and hold within role, so the catcher’s rising line isn’t a matter of which pitches he happens to challenge. One honest caveat: catchers challenge on defense (a called ball they want as a strike) and hitters on offense (a called strike they want as a ball), so the two face different pitch populations; the catcher edge holds within every edge-distance bin, but we can’t fully separate “better view” from “easier calls.”
Habits vs. results. Splitting each team’s challenges into odd- and even-numbered attempts, a team’s selection style (average distance-from-edge of the pitches it challenges) repeats across the two halves at r = 0.49; its overturn rate repeats at only 0.22 (30 teams). Style is more repeatable than winning. Limitations: a half-season; overturn is our only outcome and it treats a game-changing reversal and a trivial one alike; individual-reliability estimates are themselves noisy at these sample sizes and are reported as directional, not precise.
Cite this analysis
CalledThird. "The Best Seat in the House." CalledThird.com, July 16, 2026. https://calledthird.com/analysis/the-best-seat-in-the-house
All CalledThird analysis is original research. If you reference our findings, data, or charts in your work, please link back to the original article. For data inquiries: [email protected]