Every ABS challenge is a small bet with a strange payout. If you’re right, the call flips and you keep the challenge. If you’re wrong, nothing happens — and you’ve lost one of the two bets you get. So the decision is never “is this call worth fixing?” It’s “is fixing this call worth the risk of not having a challenge for the next one?” That’s a question with an answer, and this season’s data can finally give it.
The reason it can is a detail buried in Baseball Savant’s game feed. Alongside every pitch, the feed carries a flag for whether the ball crossed the ABS strike zone — the same zone the challenge system rules against. We checked it against every challenge of the first half: on 5,843 of 5,843, the flag matched the ABS verdict. Which means we don’t just know which challenged calls were wrong. We know which unchallenged calls were wrong — all ~14,000 of them, 5.1% of the called pitches nobody contested (9.6% of the uncontested called strikes, 3.2% of the uncontested called balls). That is the population an optimizer needs: not the bets teams made, but the bets they could have.
Two independent analyses — one a backward-induction model of the game, one a replay of real game remainders — priced the same asset and, after three rounds of adversarial review, landed within 4% of each other on the headline (what two challenges are worth at first pitch) and with a typical per-state gap of about 14% across a shared grid of 2,000 game states. This is what they found.
Three pitches
Why bother calculating? Because the stakes on a single challenge run from nothing to a third of a win, and teams get exactly two. Three pitches from this season, all scored by the optimizer below.
June 27, Reds at Pirates. Top of the ninth, Cincinnati down 7–6, nobody out, runner on first, 3-2. Elly De La Cruz is called out on strikes, challenges, and ball four is restored. That flip was worth 26 points of win probability — and the inning it kept alive produced a three-run Eugenio Suárez homer two batters later. Reds 9–7. (The season’s single biggest successful flip, for the record, was a catcher’s: Edgar Quero, bases loaded in a tied ninth, a 3-2 ball four turned into strike three, 36 points. The White Sox walked the next two hitters anyway and lost.)
August 2, Rangers at Astros. Texas lost a challenge in the first inning on a 3-2 pitch 1.8 inches off the edge, and another in the third on a 2-0 pitch a sixth of an inch off the edge — a coin flip they didn’t need to take. Top of the seventh, tied 3–3, two out, runners on first and second, 3-2: Nicky Lopez is called out on a slider 2.7 inches below the zone. Ball four would have loaded the bases; the flip was worth 12 points. Texas had nothing left to challenge with. Astros 7–3.
July 17, Cardinals at Diamondbacks. Bottom of the ninth, Arizona down 5–4, two out, tying run on second, 2-2. Ketel Marte takes a 100-mph sinker 2.6 inches above the zone and is called out on strikes. Game over. Arizona had challenged four times that night and won all four, so both challenges were still in its pocket — and the last pitch of the game, worth 15 points, went unchallenged. Nobody will ever use those two.
The second story is rare: teams burned both challenges inside the first three innings in just 72 of 3,706 team-games this season (10 inside the first inning). The third story is the epidemic, and the rest of this piece is about it.
The arithmetic of a challenge
Three numbers decide it. G is what flipping this call is worth — the swing in win probability between the count as called and the count as it should have been, including the cases where the flip ends the at-bat (strike three, ball four). C is what a challenge in your pocket is worth for the rest of the game — the expected value of having one more than you otherwise would. And q is how sure you are the umpire got it wrong. Because a successful challenge is retained, the only thing you risk is C, and only when you lose. So:
Challenge if q ≥ q* = C / (G + C).
That threshold is the whole optimizer. When a lot is at stake on this pitch (big G) or little is left in the game (small C), the bar falls and you should challenge on a hunch. When the pitch is trivial and the game is long, the bar rises. Here it is, for a tied game with nobody on and nobody out:
How sure you need to be, in percent, before challenging a called strike against your hitter — by inning and count, tied, bases empty, nobody out. Left: two challenges in hand. Right: one. Values from the reference perception model under optimal continuation; the independent replay method reproduces the two anchor situations discussed below within about a point and a half.
Two situations, to fix ideas. First inning, 0-0, two in hand: you need to be about 55% sure (54–55 across the two methods) — better than a coin flip, because a 0-0 call is worth 0.7 points and a challenge in hand is worth 0.85. Bottom of the eighth, down one, one out, tying run on second, 3-2, one challenge left (off the map above, which keeps the bases empty): about 5% (4.7 and 6.0). Flipping strike three into ball four there is worth 14.5 points; the challenge in your pocket is worth 0.7. You should challenge almost anything that looks close. Across competitive states (within three runs, innings 1–9) the required confidence spans about 60–65 points from the 5th to the 95th percentile — the rule is not “always challenge when you’re fairly sure.” It moves.
What a challenge in your pocket is worth
The interesting number turned out to be C. Two challenges at the first pitch of a game are worth 2.38 points of win probability [2.29, 2.46] by the model-based method and 2.29 [2.24, 2.35] by the replay method — about two and a half typical wrong calls’ worth, since flipping a median call moves win probability by roughly a point. The first challenge in hand is worth about 1.4 points in the first inning and about 0.75 in the ninth; the second is worth 0.85 falling to 0.3–0.4. The second challenge is worth about 60% of the first (0.59 and 0.63 in the two methods). That ratio is the tell: a game usually contains more challengeable calls than one challenge can cover. If wrong calls were rare, the second challenge would be nearly worthless — the first would already cover what little there is; instead it carries real weight all game, because teams do lose challenges (roughly 46% of the time), there is usually something left to spend the second one on, and a team with one left plays a different game than a team with two.
The shadow price of a challenge by inning, first challenge (red) and second (green), at a tied/empty/nobody-out state. Solid line and band: dynamic program with 500-game bootstrap intervals. Open dots: the independent real-remainder replay at the 1st and 9th — the convergence check. Dashed: a sharper-radar sensitivity (higher early, crossing below the reference late).
One caveat carries through everything below, so we’ll state it once. All of these prices depend on how well the people making the decision can see a wrong call, and that isn’t observable directly. We calibrated a “radar” — the decision-maker sees the pitch’s location with about 2.5–2.6 inches of noise — to reproduce exactly what the league does: it challenges 24.9% of the wrong calls it could and 1.5% of the correct ones, which is why 53.7% of challenges succeed. A sharper radar raises the headline prices (two challenges at first pitch go from 2.38 to 2.60 points under the most defensible alternative, and to 3.15 under an aggressively sharp one; ninth-inning prices barely move); a blurrier one lowers them (to 1.79 under an aggressively blurry one). The reference numbers sit toward the lower-middle of that band.
The bar drops. Not enough, and not soon enough.
Here is what teams actually do. They challenge 2.14 times per team-game and win 53.7%. Their success rate falls by inning — around 60% in the first two innings, 36–46% in the ninth depending on how many they have left. That decline is correct. A challenge is a wasting asset; as the game runs out, C collapses toward zero, the bar falls, and you should be spending your challenges on closer and closer calls. A falling success rate is what good decisions look like.
But it doesn’t fall enough. An optimal team with the same radar would challenge 2.7 times per game and win only 39% — and it would still end with 0.46 challenges unused, because sometimes there’s simply nothing worth contesting. In the ninth inning with two in hand it would challenge 15% of all the called pitches that go against it; the league challenges 6%. With one left, 10% against the league’s 5%. Its ninth-inning success rate would be around 23–28%, not the league’s 36–46%. Teams treat a ninth-inning challenge like a fifth-inning one.
Challenge success rate by inning, split by challenges remaining so the comparison isn’t an artifact of who still has two. Dots: what teams did (95% intervals; early-inning one-left samples are small). Red: what an optimal team using the same radar would run. The gap between them, late, is the finding.
Both methods reached the same verdict on this one, and it’s the piece’s cleanest sentence: the direction of the league’s adjustment is right, the magnitude is too small, and it arrives too late.
Hoarding, not panic
Which of the two mistakes costs more — challenging when you shouldn’t, or holding when you shouldn’t? It isn’t close. Against the optimal policy, the league leaves about 0.62 points of win probability per team-game on the table. 0.57 of that [0.54, 0.61] is hoarding — challenges that should have been made and weren’t. 0.05 [0.05, 0.06] is panic. Eleven to one. And the hoarding is concentrated exactly where the theory says it should be: the eighth and ninth innings, with two challenges still in hand.
Win probability left on the table per team-game, by inning, split into hoarding (red; whiskers are 95% intervals) and panic (gold). The late-inning spike is teams carrying two challenges into the ninth and not using them.
The plainest version of the same fact needs no model at all. 75% of team-games [74, 77] end with at least one challenge unused; 29% end with both. Teams average 1.04 unused challenges per game against 2.14 used. Priced by the optimizer, the inventory that dies unused was worth 0.58 points per team-game [0.55, 0.61] on entering the seventh (0.54 by the replay method) — about a quarter of what two challenges were worth at first pitch. If you’d rather not trust the model at all: with perfect hindsight, using the ABS flag, teams captured about a third (33%) of the win probability that was sitting in wrong calls against them, and about 21% of team-games (20.7% and 21.9% in the two analyses) ended with a challenge unused and a wrong call worth at least a full point of win probability having gone uncontested from the seventh inning on. We had guessed a quarter before we looked; it’s a fifth. The hindsight upper bound — every wrong call, perfectly seen — is 3.7–3.9 points per team-game, and no human radar gets there. The optimizer’s 0.6 is the model’s estimate of what a team with the league’s radar could reach by deciding better; if it keeps deciding like the league from each state onward, the unused inventory is worth 0.35–0.39.
Which teams do this well? Mostly, you can’t tell.
Teams differ in how they use challenges — that part is just counting. The Yankees challenge 2.5 times a game and end with one unused 66% of the time; the Red Sox challenge 1.9 times and end with one unused 84% of the time. But whether those habits translate into win probability is a different question, and the answer is: not detectably, yet. We pre-registered a split-half reliability bar of .30 for a team-level “challenge efficiency” leaderboard. One method’s point estimates cleared it (.37 and .40 on two metrics) with intervals that ran to negative; the other’s did not (.03 and .23). Both analyses recommended the same thing, and we’re taking it: no ranked leaderboard. The between-team spread in realized value is under 0.2 points per game after shrinkage, which is smaller than the noise in measuring it. So the table below is descriptive — how teams behave, not how well.
| Team | Games | Challenges per game | Success rate | Ends with one unused | Efficiency vs optimal (pp / game ± se) |
|---|---|---|---|---|---|
| MIN | 123 | 2.77 | 55% | 62% | +0.08 ± 0.14 |
| CWS | 122 | 2.59 | 47% | 56% | +0.10 ± 0.13 |
| NYY | 122 | 2.51 | 55% | 66% | +0.25 ± 0.12 |
| COL | 123 | 2.42 | 54% | 69% | +0.05 ± 0.13 |
| MIA | 124 | 2.42 | 57% | 69% | +0.04 ± 0.14 |
| NYM | 124 | 2.40 | 51% | 65% | -0.15 ± 0.15 |
| HOU | 124 | 2.31 | 54% | 75% | +0.09 ± 0.14 |
| LAA | 124 | 2.31 | 53% | 72% | -0.07 ± 0.13 |
| ATH | 123 | 2.31 | 57% | 76% | -0.08 ± 0.15 |
| MIL | 124 | 2.30 | 50% | 66% | -0.17 ± 0.13 |
| CIN | 122 | 2.27 | 64% | 84% | +0.04 ± 0.14 |
| KC | 124 | 2.13 | 56% | 77% | +0.09 ± 0.12 |
| CLE | 124 | 2.08 | 47% | 73% | -0.01 ± 0.13 |
| PIT | 125 | 2.08 | 45% | 67% | -0.23 ± 0.15 |
| WSH | 125 | 2.05 | 50% | 74% | -0.03 ± 0.13 |
| PHI | 123 | 2.04 | 57% | 85% | -0.13 ± 0.15 |
| TOR | 125 | 2.02 | 53% | 75% | +0.02 ± 0.13 |
| TB | 122 | 2.02 | 52% | 80% | -0.12 ± 0.15 |
| SD | 122 | 2.02 | 53% | 78% | -0.20 ± 0.15 |
| SF | 122 | 2.01 | 50% | 76% | -0.07 ± 0.14 |
| DET | 123 | 2.00 | 60% | 81% | -0.06 ± 0.13 |
| BAL | 123 | 1.99 | 50% | 73% | -0.14 ± 0.14 |
| SEA | 124 | 1.97 | 52% | 79% | -0.19 ± 0.15 |
| CHC | 124 | 1.94 | 59% | 85% | +0.00 ± 0.15 |
| ATL | 123 | 1.93 | 52% | 77% | -0.01 ± 0.13 |
| BOS | 123 | 1.93 | 55% | 84% | -0.22 ± 0.13 |
| LAD | 124 | 1.92 | 56% | 82% | -0.21 ± 0.14 |
| TEX | 124 | 1.81 | 57% | 85% | -0.20 ± 0.15 |
| AZ | 122 | 1.80 | 61% | 85% | -0.11 ± 0.14 |
| STL | 123 | 1.76 | 52% | 86% | -0.05 ± 0.14 |
Sorted by challenges per game. The last column is a shrunken estimate of each team’s realized challenge value minus what the optimal policy would have earned on the same opportunities under the reference radar, shrunk toward the league’s grand mean (about −0.06 pp per game); note that almost every interval covers zero. This is not a leaderboard.
Try it
The tool below runs the full policy table — every inning, half, out state, base state, score within four runs, count, call, and inventory. Set the situation, tell it how sure you are, and it tells you whether to burn one. There’s a standalone version under Tools.
Bottom of the 7th, 1 out, runner on second, down 1, 3-2 — strike called on my hitter, one challenge in hand.
Loading the policy table (about half a megabyte)…
q* = C / (G + C). Values from the reference perception model under optimal continuation; a sharper radar raises early-game challenge values by up to ~10% (q* by a point or two) and barely moves late-game ones; values are rounded to a thousandth of a point. Extras (X) use the “a team at zero regains one each extra inning” rule. Score is clipped at ±4.
What this isn’t
It isn’t a claim about who challenges well. We showed in July that individual challenge skill is mostly noise, and that catchers succeed more often than hitters (Best Seat in the House); this piece takes those as given and asks only about the state of the game. It isn’t a postseason model — the postseason has its own leverage and no data yet, and we won’t guess. And it isn’t a claim that any team is being stupid: the league’s adjustment runs the right way. It’s a claim that a challenge is a wasting asset priced too high in the ninth inning, and that the adjustment is simple: with two in hand and three outs to go, challenge the ones that look close. The one you’d be protecting is worth about a third of a point by then.
One takeaway for tonight: watch the eighth inning of a close game and count the challenges each side is still holding. If it’s two, the optimizer says the bar for that team is now somewhere between “probably wrong” and “might be wrong” — and the broadcast will still call the challenge they don’t make “disciplined.”
Methodology
Data, the decision model, the two methods, the perception calibration, and what changed under review
Data. Every called ball and called strike in every 2026 regular-season game with pitch data from March 26 through August 15 — 283,512 pitches in 1,853 of 1,859 completed games — from the Baseball Savant game feed, joined to Statcast pitch-level state (score, runners, count-aware win expectancy) at a 100% match rate. 7,908 challenges, 53.7% overturned. The game feed’s in-zone flag equals the ABS verdict on 7,907 of 7,908 challenged pitches once the umpire’s original call is reconstructed (the feed records the post-challenge call). Rules were verified from the replay and against MLB’s published guidelines: two challenges per team, retained on success, lost on failure; at the start of each extra inning a team at zero regains one; zero inventory exceptions in regulation.
The decision model. A challenge succeeds with the decision-maker’s belief probability q and is retained; it fails with probability 1 − q and is lost. G is the absolute win-probability swing between the called and correct counts, including PA-ending flips, from Statcast’s count-aware win expectancy (an exact lookup on states seen this season, a fitted fill on the few unseen). C(state, k) is the marginal value of holding k rather than k − 1 challenges for the rest of the game. The rule challenge-iff q ≥ C/(G+C) follows. Transitions after a successful flip (a changed count, an ended or revived at-bat, a run) shift q* by a median of about 0.03 percentage points (the transition term is a median 2.6% of G) and are reported as a sensitivity; the replay method’s transition-aware variant did not converge to its tolerance, so its primary numbers are the no-transition ones.
Two methods. Method A solves C by backward-induction dynamic programming over a discretized game state (inning, half, outs, bases, score within four, count, call, k) with an empirical arrival model of wrong calls by count, half, and inning bucket. Method B never fits an arrival model: for each state it samples real remainders of real 2026 games against the challenging team, matched on state, and values a held challenge under a state-based threshold policy over those remainders. Both score a shared grid of 2,000 real pitch states; after convergence the shadow prices correlate at .95 (Spearman), with a typical per-state gap of 14% (the pre-registered gate statistic, median absolute gap over the median price, is 11%), and G at .97 against an independent Statcast-only anchor. Bootstrap intervals are from 500 game resamples in both methods (Method B holds its fitted perception and win-probability models fixed across resamples).
Perception. Players don’t see the ABS zone, so a model of belief is required, and it is the load-bearing assumption. The shared reference model gives the decision-maker the pitch location with isotropic Gaussian noise (σ = 2.61 inches [2.55, 2.70] in Method A, 2.51 in Method B), fitted so a single-threshold policy reproduces the league’s realized operating point: 24.9% of wrong calls challenged, 1.5% of correct ones (belief AUC .908). The two methods share that operating point and the resulting belief accuracy (AUC .908), not an identical belief construction — a disclosure the cross-review required. Each method also ran its own alternative (a latent Beta radar; a coarsened-feature classifier) calibrated to the same point; the headline numbers move by a few percent. A stratum-consistent sharper radar (σ = 2.29 in) raises the two-challenge value from 2.38 to 2.60 points; a radar 1.5× blurrier lowers it to 1.79, one 1.5× sharper raises it to 3.15. The tool prints reference values, rounded to a thousandth of a point; in near-decided states its threshold can differ from the research table by up to a point. The model treats every called pitch as challengeable (a slight simplification of the rules, which carve out a few cases), so false alarms consume inventory. “Optimal continuation” prices assume the team plays optimally from here; “league-typical” prices (about 25–35% lower) assume it keeps behaving like the league.
Pre-registration and verdicts. H1 (the rule is state-dependent, ≥15-point q* range across competitive states): supported, 63/59 points within one/two in hand (Method B: 66/61). H2 (unused inventory is worth ≥0.5 points per team-game): supported at optimal prices (0.58, 0.54), not at league-typical prices (0.39, 0.35); the pre-registered hindsight guess of 25% of team-games came in at 21% and is reported as below the guess. H3 (timing of the late-game bar drop): both methods, “too late.” H4 (team differences are noise): descriptive table only. Estimands are stated in the research brief; one — the exact game state at which unused inventory is priced for H2 — was made precise during reconciliation so both methods computed the same thing, and we say so.
What changed under review. Round 1 failed the pre-registered cross-method agreement gate. Review found three coding errors in Method B (a count double-advance, a decoding error on the shared grid — partly the orchestrator’s, which shipped ambiguous columns — and post-plate-appearance out counts) and one design flaw both reviewers caught: a “fair” belief model built on exact pitch location that was, in effect, the truth flag, so challenges never failed and the second challenge priced at zero. Round 2 fixed those and imposed the shared perception operating point; Round 3 removed a remaining hindsight leak in Method B’s valuation (it had optimized against each game’s realized future) and delivered the mandated 500 bootstraps in Method A. Every number above traces to a produced artifact.
Limitations. The perception model is calibrated, not observed; the article prints reference numbers and states the band. Extra innings are handled by rule, not modeled richly. Score is clipped at four runs. Nothing here estimates individual or positional challenge skill, and no team ranking is claimed.
Cite this analysis
CalledThird. "Should You Burn One? A Challenge Is a Wasting Asset. Teams Hoard It Anyway." CalledThird.com, August 19, 2026. https://calledthird.com/analysis/should-you-burn-one
All CalledThird analysis is original research. If you reference our findings, data, or charts in your work, please link back to the original article. For data inquiries: [email protected]