Few analytical findings have shaped modern starter usage as visibly as the times-through-the-order penalty. The raw fact is easy to see: against a starter facing them for the third time in a game, hitters produced a .338 wOBA in 2025, against .317 the first time. Across 2022–25 that face-value gap is about 25 points — and if you follow one batter against one starter within a game, 27 points.

We set out to take that number apart: how much is the pitcher declining, how much is the hitter adjusting, how much is an accounting artifact — and how much is simply the pitch count wearing a costume. Two independent analysis pipelines — different models, different code, forbidden from talking to each other — ran the same pre-registered brief, adversarially cross-reviewed each other twice, reproduced each other’s anchor rows exactly, and converged on a signed ledger of publishable claims — every remaining disagreement in this piece is reported as a range. What they also converged on, and what we didn’t expect to be the finding: the last question on that list is unanswerable with the data baseball has. We’ll show you exactly why, because the reason is itself the best explanation of how the third-time penalty actually works.

Every trip through the order, 2022–25423,861 plate appearances vs. starters · descriptive · whiskers: 95% game-cluster intervals.290.300.310.320.330.340.350wOBA.3171st trip172,533 PA34% slots 1–3~pitch 16.3262nd trip163,041 PA34% slots 1–3~pitch 51.3423rd trip86,316 PA53% slots 1–3~pitch 78.3234th+ trip1,971 PA89% slots 1–3~pitch 94.321.331.338wOBA allowedseason wOBA of thehitters at the plateThe survivors’ club:4th trips come from ahighly selected set ofstarts — long outings bybetter-than-average arms —the line falls while thehitters are the best yet.

The effect, with nothing hidden: wOBA allowed by trip through the order, 2022–25, with the season-long quality of the hitters actually at the plate each trip drawn beneath it (red, dashed). Three things to notice before reading on: the raw line climbs; the quality line climbs with it — by the third trip, more than half the plate appearances belong to lineup slots 1–3; and on the rare fourth trip the raw line falls while facing the best hitters yet — because fourth trips come from a highly selected set of starts and starters. Composition and selection, both visible before any model. All descriptive; the corrected penalty comes next.

The lineup card does part of the work

Start with the accounting artifact, because it’s the piece nobody quotes. The third-look sample is not the first-look sample with more mileage: it is disproportionately the top of the order, because in a modern start the bottom of the lineup often never gets a third look at all. Hitters who reach a third plate appearance against a starter carry a season wOBA about 5 to 10 points better than the average first-look hitter, depending on which pre-registered quality measure you use. The pitchers who are still out there pull the other way — the ones allowed to face a lineup a third time are better than average, which gives back a couple of points. And there’s a subtler leak: the starters who get to a third trip are partly the ones who got lucky early, and luck doesn’t persist. A simulation with the penalty set to exactly zero — same lineups, same realistic manager pull behavior — still produces a spurious +7-point “penalty” from that selection alone.

Where the face-value penalty goeswOBA points, third trip vs first, 2022–25 · ranges span two pre-registered quality accountings010203024.6−5 to −10+1.5 to +2.617–21Face value(raw third-trip gap)− who is batting(better hitters reacha third look)+ who is pitching(survivors arebetter pitchers)What is left(the penalty)

The face-value third-trip gap (raw, 2022–25) and what remains after accounting for who is batting and who is pitching. Ranges span two pre-registered ways of measuring hitter quality; the paired estimate (27 points, same batter vs. same starter) minus its own selection correction lands in the same 17–21 window. One sentence we are contractually obligated to include, because it’s true: the paired estimate controls who is batting and who is pitching; it does not control the manager’s decision to leave him in.

After the accounting, a real penalty of about 17 to 21 points remains — roughly the gap between a league-average hitter and a good one. The third time through is real. Then we went looking for what carries it, and this is where the story turns.

The arm: real, replicated, and worth almost nothing

The oldest explanation is the pitcher’s. By the third trip his fastball has lost 0.36 mph — in every one of the four seasons, in every fastball type (four-seam −0.36, sinker −0.44, cutter −0.36), measured within the same start so no cross-game comparison can fake it. Spin is down 6 rpm, the arm slot has sagged a quarter of a degree, vertical break has flattened a fifth of an inch. Stuff declines with pitch count. That much is settled.

The arm on the third tripfastball velocity, 3rd trip vs 1st, demeaned within start and pitch type · 2022–25, every season-0.5-0.4-0.3-0.2-0.10mphFour-seam-0.36Sinker-0.44Cutter-0.36Also on the third trip:spin -6 rpmarm angle -0.24°vertical break -0.018 ftextension +0.03 ftmix shift adds −0.11 mphEstimated wOBA cost: ≈ zero.Cross-lane uncertainty is on the order of±5 wOBA points: the decline replicates;harm from it is not established.

Fastball velocity on the third trip vs. the first, demeaned within each start and pitch type, 2022–25 pooled (whiskers: 95% game-cluster intervals). The decline replicated in all four seasons in both pipelines. The box is the punchline.

What that decline is worth is another matter. Price the run value of a third of a mile per hour any defensible way we tried — and both pipelines tried several — and the estimated wOBA contribution of the observed stuff decline is approximately zero, with cross-lane uncertainty on the order of five wOBA points either way. The decline replicates; harm from it is not established. The pitcher’s plan changes far more than his arm does: by the third look, starters throw 20.6 percentage points fewer first-pitch fastballs, land 4.1 points fewer first pitches in the zone, and get 6.9 points fewer first-pitch called strikes. That is the pitcher’s plan retreating on the third look — a change in choices far larger than the change in his arm.

And the decoupling holds pitcher by pitcher. Who loses velocity late in starts is a repeatable trait among qualified starter-seasons — it repeats year over year at r ≈ .8. Some starters shed nearly two miles per hour from their mid-start peak; a few finish throwing harder than they peaked. But who gets hit on the third trip is not a stable fact about a pitcher: the same-season split-half correlation of the per-pitcher paired penalty is .05, and across 496 pitcher-seasons the correlation between losing velocity and getting hit is .02. Here are all 141 qualified 2026 starters — find yours:

hover or tap a dot to pin; 141 starters, 2026 through Aug 15
Who loses velocity late — and who gets hit on the third tripeach dot a 2026 starter (≥ 10 qualified starts) · the absence of a pattern is the finding: r ≈ .02 across 2022–25-1.5-1.0-0.50.0+0.5−.2−.10+.1+.2+.3leaguefastball velocity, pitches 76–90 vs. 16–35 peak (mph, within start & pitch type)a real trait: year-over-year r ≈ .83rd-trip − 1st-trip wOBA, same battersmostly noise per pitcher: split-half r ≈ .05

Each dot is a 2026 starter with at least 10 qualified starts, through August 15 — descriptive, one season, no grades. Horizontal position (late-start velocity change) is a repeatable trait; vertical position (paired third-trip penalty, at least 25 same-batter comparisons per pitcher — median 97, range 28–206) is mostly sampling noise, which is why the cloud has no shape — and that shapelessness, replicated on 2022–25 (r = .02, n = 496 pitcher-seasons), is the point: the radar gun’s decline does not tell you who pays the third-trip penalty, and the penalty itself is not a stable per-pitcher skill. No dot on this chart licenses a “fades late” or “third-time-proof” label.

The hitter: the ambush that wasn’t, and the contact that is

Flip to the batter’s box and the third look appears transformed. Same batter, same game, same starter: on the third look he swings at 10.2 percentage points more in-zone first pitches, whiffs on 1.2 fewer pitches per hundred, makes contact on 2.8 points more two-strike pitches in the zone, and strikes out 3.7 points less often. Contact quality holds up too: expected wOBA on contact is about 19 points higher. At face value, those rows resemble a hitter ambushing the first strike he sees.

Most of that picture is the same accounting trap wearing a batting helmet. The third-look plate appearances happen 60 or 70 pitches into a start; the first-look ones happen at pitch 10. To compare like with like, you have to hold the pitch count still — a second look and a third look, inside the same start, both in the pre-specified pitch 55–80 window. In that comparison — a descriptive one, not a causal estimate of the look — the famous first-pitch ambush doesn’t just shrink. It reverses.

The third look: what survives the pitch countpercentage points vs the earlier look · hollow = as a fan sees it (3rd vs 1st trip) · solid = same start, same pitch count (3rd vs 2nd)0-15-10-5+5+10Whiff per pitch-1.1-2.1 to -1.3Contact on 2-strike pitches in the zone+2.8+3.6 to +4.5First-pitch swing, in-zone+10.2-11.5 to -7.1the “first-pitch ambush” does not survive the local comparison

Three hitter signatures, two comparisons. Hollow dots: third look vs. first look, paired within batter and game — what a fan watching the broadcast sees. Solid bars: third look vs. second look at the same pitch count inside the same start (the range spans the two independent pipelines). The contact rows survive the local comparison; the first-pitch swing rise does not.

The contact signature is the one that survives every in-sample design. Fewer whiffs per pitch (1 to 2 per hundred), more contact on two-strike pitches in the zone (+3.6 to +4.5 points) — in the paired comparison, in the local same-pitch-count comparison, in every season, under every model variant, in both pipelines — and its 2026 sign held out of sample, though its magnitude beat our frozen ceiling (more below). That pattern is consistent with familiarity: whatever else the third look gives a hitter, it gives him the ball. What we could not show is that it buys runs at a fixed pitch count — expected damage on contact changes sign across our two pipelines there, and strikeout and walk rates reorganize under the local controls — we won’t quote a number we can move by changing an assumption. (The first-pitch reversal itself is stable in both pipelines.)

The question the data refuses to answer

Which brings us to the question this piece was supposed to be about: how much of those 17–21 points is the trip — the hitter’s third look — and how much is just the pitch count? We pre-registered that split as our headline question. The data refused to answer it. Twice, in two independent pipelines. Understanding why is the most useful thing we learned.

Three facts gang up on the question. First: a first look and a third look almost never happen at the same pitch count — trip one spans pitches 1–36 at the 5th to 95th percentiles, trip three 62–96 — so separating them statistically requires assumptions the data can’t check. Second: within a start, timing is not random. Starts end right after damage — the last two batters a pulled starter faces hit .42 and .39 — so any model that tries to read “the clock” off within-start timing inhales the manager’s pull decision and hallucinates. Our pre-registered primary model produced a nonsensical answer for exactly this reason, and a zero-penalty simulation reproduced the nonsense. Third, and most stubborn:

Who is at the plate at pitch 55–80?trip 1: pitch 1–36trip 2: pitch 33–72trip 3: pitch 62–96the pre-specified local window →2nd look3rd look128%224%319%413%59%8%615%721%824%925%avg hitter quality (wOBA): .302.337 — a 35-point gap

Every plate appearance at pitches 55–80 of a start, 2022–25 — the pre-specified local window where second and third looks coexist. The second looks are the bottom of the order; the third looks are the top. The average hitter-quality gap between the two groups is 35 points of wOBA — twice the size of the effect being hunted.

At the same pitch count, in the same start, the third look is the top of the order. The thing you must adjust for is the thing that defines the comparison, and the local estimate moves point-for-point with how you measure hitter quality — our two pipelines, using two defensible quality models, got local answers on opposite sides of zero, and a zero-effect simulation showed the design itself can wander by about 13 wOBA points in either direction. The honest statement, which both pipelines signed: at the same pitch count, a third-look effect on wOBA cannot be resolved on these data — anything smaller than about 15 points of wOBA, in either direction, is invisible here. If someone tells you they’ve measured the pure times-through-order effect net of pitch count, ask them very politely what they did about the lineup card.

What this looks like in a box score

Population estimates are how you know; box scores are where you look. Four moments from 2026 — each picked by a pre-registered rule (the largest qualifying instance of its pattern this season; exact criteria in the methodology), each one finding in miniature, and none of them proof of anything on its own.

The collapse that was the lineup card. April 10 in Atlanta, Slade Cecconi’s first two trips through the order went .200 and .178. His third: five batters, starting at pitch 69 — homer, single, homer, single, out — a 1.160 wOBA. That line is the face-value penalty at its most extreme, and it is also exactly what the composition finding describes: the third look opened with the Braves’ 1-2-3 hitters, at pitch 69, for five plate appearances. Both things are true at once — he got hit, and the sample that hit him was the best of the order arriving on schedule. When tonight’s broadcast shows a third-trip meltdown, ask the Cecconi questions first: who was batting, at what pitch count, over how many hitters?

The radar gun, wrong in both directions. April 26 in San Francisco, Landen Roupp’s fastball ran about 1.4 mph above his start-wide norm early and 0.9 below it by the third trip — a within-start slide of more than 2 mph, the kind of graphic that gets a bullpen moving. Miami’s third look went nine up: eight outs and a walk, .078. A week earlier in Boston, Garrett Crochet’s velocity held (−0.2 early, +0.2 late) — and Detroit’s third look, beginning at the top of the order, hit 1.083 over six batters: homer, walk, single, homer, single, out. One start each, and neither says anything about either pitcher beyond that night — but together they are the aggregate result in miniature: the velocity decline is real, and the results don’t follow the gun.

The signature, in one at-bat. August 9 in Milwaukee, Kody Clemens faced Jacob Misiorowski three times and saw nothing but four-seam fastballs — fourteen of them. First look, starting at pitch 10: a seven-pitch strikeout with three swings and misses at 101–102 mph. Second look, pitch 52: three pitches, two more whiffs, strike three called at 103. Third look, pitch 86: four fastballs, now 99–100 — and the fourth one left the park. Five swinging strikes against the four-seamer, then the four-seamer in the seats. An anecdote, not evidence — but it is the exact shape of the one finding that survived everything we threw at it: on the third look, the misses turn into contact.

Season by season — and a test we couldn’t rig

Everything above was fit on 2022–25. Before we opened the 2026 data, both pipelines wrote down predictions — the paired penalty, the velocity decline, the first-pitch swing rise, the whiff drop — froze them with cryptographic hashes, and then scored the season once.

The penalty, season by season — and the sealed testpaired third-trip penalty (same batter, same starter), wOBA points · 2026 was predicted, hashed, then scored once010203040wOBA pointsfrozen interval352022242023302024192025182026sealed holdout

The paired third-trip penalty by season (95% game-cluster intervals), and sealed 2026: predicted before the holdout was opened, scored once. The 2026 value (+18 points) landed inside the frozen interval; so did the velocity decline (−0.43 mph) and the first-pitch swing rise (+8.8 points).

Three of the four sealed predictions landed inside their intervals. The fourth missed — honorably, but a miss: the 2026 whiff-per-pitch drop came in at −1.8 per hundred, stronger than the frozen interval allowed. The sign held; the magnitude beat our ceiling; by our own pre-registered rule that scores as a failure, so treat the size of the contact signature as provisional. Worth watching alongside it: the paired penalty itself has drifted from +35 points in 2022 to +19 in 2025 and +18 in 2026 — descriptive only; if the third-time penalty is quietly shrinking, this chart is us on the record early. It’s a monitoring question, not a finding.

The 2026 leash, descriptively

The stretch-run read: the selection-corrected population penalty is 17–21 points — real, but smaller than the face-value number commonly quoted in quick-hook discussions — and this study does not produce a pull rule. The velocity decline is real; harm from it is not established. The leash itself has stayed in a narrow 4.3–4.5 third-look batters per start since 2022 (4.3 in the 2026 data through August 15), with Seattle (5.4 per start) running the longest leash in baseball this year.

Two 2026 descriptive examples, because their by-trip splits will get quoted down the stretch. Tarik Skubal’s 2026 by trip: .241 the first time through, .256 the second, .371 the third (93 PA) — while his type-demeaned fastball velocity runs higher at pitches 76–90 than at 1–15. Before anyone quotes that .371 in a pregame segment: the population results above say lineup composition and selection inflate every starter’s face-value third-trip line, and 93 plate appearances decide nothing. And Cam Schlittler, whose April profile in these pages left “the second time through” as an open question: in 2026 he’s at .249 / .240 / .332 by trip across 26 starts — the by-trip numbers the April profile lacked. These rows are descriptive snapshots, not estimates of either pitcher’s times-through-order skill; the population composition and selection caveats apply to both.

One more piece of housekeeping. In April, before we had run any of this, an early piece on this site asserted that late-game decline is “the batter adjusting, not the pitcher deteriorating” and endorsed pulling starters “based on times through the order regardless of how sharp he looks.” That paragraph was half right at best: the deterioration is real but worth almost nothing, the adjustment that survives scrutiny is contact-only, and the trip-versus-pitch-count distinction that rule relies on turns out to be unanswerable. We’ve appended a correction note to that article. The takeaway for tonight’s game: when the broadcast flashes a third-time-through split, remember that lineup composition and selection inflate the face-value number — and don’t apply our population bridge as a fixed correction to any one pitcher. In this study, the contact signature is clearer than any established harm from the velocity decline — the hitter who stopped missing is the real story of the third look.

Methodology

Data, the paired design, selection corrections, the identification failure, and the sealed 2026 test

Data. Every regular-season plate appearance against a game’s starting pitcher, 2022–2025 (426,206 PAs; 86,177 paired batter-games with a first and third look), built from Statcast pitch-level data with a rebuilt times-through-order count (Statcast’s own n_thruorder_pitcher is a lineup-pass counter, not a per-batter count; phantom at-bats from mid-inning runner outs and mid-PA pitching changes are excluded; automatic balls/strikes don’t count as pitches). 2026 (through August 15) was held out untouched behind a hash-sealed prediction file.

The headline penalty. The paired estimate compares the same batter against the same starter within a game (third look minus first), +.027 [.022, .032]. The corrected range +.017 to +.021 subtracts the regression-to-the-mean component of the selected-low first-look baseline (starters who reach a third trip were getting good outcomes earlier, partly luck), with the range spanning two pre-registered hitter-quality models; a composition-adjusted unpaired model lands in the same window. A synthetic-null simulation — zero true penalty, realistic manager pull rule — produces a spurious +.007 [.003, .011] paired gap and is printed as the scale of what selection alone can do. These estimates control who is batting and who is pitching; they do not control the manager’s decision to leave a pitcher in.

Two pipelines, two cross-review stages. The anchor rows and the signed ledger were computed independently by two pipelines from a shared data build (one interpretability-first, one ML-first, different model families), each blind to the other until filing, then cross-reviewed with required recomputation — once after the main round and once after the local-estimand round. Implementations agreed exactly where specifications matched (to the fourth decimal on the anchor rows); every disagreement traced to a documented assumption and is reported as a range. Shared-data chart payloads and the post-holdout Schlittler descriptive row were source-checked rather than duplicated. The pre-launch brief was itself adversarially reviewed by a 12-agent panel, which restructured the identification strategy before any analysis ran.

The identification failure, precisely. Trip and pitch count share almost no common support (trip 1 spans pitches 1–36 and trip 3 spans 62–96 at the 5th–95th percentiles, with almost no overlap); within-start timing is endogenous to the pull decision (the last two PAs before a pull average .42/.39 wOBA), which poisons any start-fixed-effects “clock” model — our pre-registered primary specification returned design-determined values that a zero-effect placebo reproduced; and in the pre-specified local window (pitches 55–80) the third-look PAs are the top of the order (average quality gap .035 wOBA, twice the effect size sought), so the local contrast moves point-for-point with the hitter-quality model. Two defensible quality models put the local effect on opposite sides of zero; placebo simulations show the local design itself can produce ±.013 under zero truth. We therefore report the local third-look effect as unresolvable at a resolution of about ±.015, in either direction.

The contact signature. Whiff per pitch and two-strike in-zone contact are computed conditional on pitch type, location, count, and handedness, within batter-game (paired) and within start at pitches 55–80 (local); both survive every specification in both pipelines (local whiff −1 to −2 per hundred; local two-strike in-zone contact +3.6 to +4.5 points). Expected wOBA on contact and strikeout/walk rates are not stable in the local design and are quoted only in the paired form.

The trip chart. The opening chart is descriptive: pooled 2022–25 real PAs vs. starters (substitutes and mid-PA replacements excluded), game-cluster bootstrap intervals, with trip 4+ (1,971 PAs) shown for completeness — fourth trips were excluded from all primary analyses, and their .323 against the best hitters (89% slots 1–3; mean final start length ≈ 99–100 pitches depending on PA- vs. start-weighting; PA-weighted pitcher quality .316 vs. .324 at the first trip) reflects aggregate survivor selection, not late-game improvement.

The scatter. Per-pitcher late-start velocity change = mean type-demeaned fastball velocity at pitches 76–90 minus the 16–35 peak, per start, averaged within pitcher-season (2026 dots: ≥ 10 such starts, through August 15, descriptive). Its stability on 2022–25: split-half r = .74 (halves by game-id parity, ≥ 8 qualified starts per half, n = 374 pitcher-seasons), year-over-year r = .75–.85 (≥ 15 qualified starts in both seasons; n = 74/72/78). The per-pitcher paired penalty’s split-half r = .05 (same parity halves, ≥ 20 pairs per half, n = 603 pitcher-seasons); the decline–penalty correlation is r = .02 (n = 496 pitcher-seasons at ≥ 15 starts and ≥ 40 pairs).

Box-score examples. The four 2026 examples were selected by rules written before the search, not narrative convenience. In full: (E1) starts with ≥ 5 completed third-trip PAs, third-trip wOBA allowed ≥ .500, ≥ 60% of those PAs from lineup slots 1–3, and a third look beginning at pitch ≥ 60 — largest third-trip wOBA wins; this rule produced a TIE at 1.160 (Slade Cecconi and Kyle Leahy, April 22), and because no tie-break was pre-registered, the displayed pick fell to script row order — disclosed here. (E2a/E2b) starts of ≥ 90 pitches with ≥ 15 fastballs and ≥ 5 PAs at both trips: largest type-demeaned fastball-velocity drop of at least 0.8 mph with third-trip wOBA ≤ .150 (Roupp), and largest third-trip wOBA ≥ .500 with velocity change ≥ 0 (Crochet). (E3) the highest-value trip-3 hit (single/double/triple/home run) on a pitch group the batter had whiffed on at least twice earlier in the same game, ties broken by latest date then game id (Clemens). Every stated number was recomputed from the holdout PA and pitch tables; the examples are descriptive and carry no claim about any player’s times-through-order skill.

The sealed test. Both pipelines froze 2026 predictions with SHA-256 hashes before the holdout was opened and scored once: paired penalty +.018 (inside), fastball velocity −0.43 mph (inside), in-zone first-pitch swing +8.8 points (inside), whiff per pitch −1.8 per hundred — outside its interval on the stronger side, a pre-registered failure that makes the contact signature’s magnitude provisional. A sealed pass certifies stability of the estimates in a new season, not the mechanism story. Prior art we’re building on, and departing from where noted: Lichtman’s original penalty work, Carleton’s pitch-count mediation analyses (2016, 2018), Brill, Deshpande & Wyner’s Bayesian analysis (2023, which found no discrete trip jump and is, we can now add, survivor-conditioned like everything else), and Sutton-Brown’s hitter-side work (2025).

Share Twitter/X Bluesky

Cite this analysis

CalledThird. "The Third Time Through Is Real. Here’s What’s Carrying It." CalledThird.com, August 27, 2026. https://calledthird.com/analysis/third-time-through

All CalledThird analysis is original research. If you reference our findings, data, or charts in your work, please link back to the original article. For data inquiries: [email protected]