{"id":"d034aaba-7915-4f4b-9cef-41a2d6225f3c","arxiv_id":"2411.15075","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"MLB's 2023 shift ban raised left-handed batters' BABIP and OBP by about nine points, with larger synthetic-control estimates for several high-shift players.","lead":"This paper measures how Major League Baseball's 2023 ban on infield shifts changed hitting results by comparing left-handed and right-handed batters, and by building synthetic comparison versions of high-shift players. It finds a modest boost in batting average on balls in play and on-base percentage for left-handed batters, plus large estimated gains for several individual players.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"In-time placebo estimates appear systematically negative for high-shift players, suggesting SCM player-level gains may be inflated by donor-pool bias rather than the shift ban.","rationale":"The reader's weakest assumption identified the general SCM concern: weights must remain stable over time, and post-2022 divergence could bias the player-specific estimates. My concern is a specific, observable manifestation of that instability: the in-time placebo analysis, which the paper itself conducts, produces negative estimates in a year without the intervention. If those negative estimates are non-trivial and systematic, they indicate the synthetic control overpredicts high-shift players' outcomes, which would directly inflate the 2023 effect estimates. This is more load-bearing than the broader 'divergence' concern because it is grounded in a reported diagnostic result, not a hypothetical scenario. The league-wide DID result (nine-point BABIP/OBP increase) is less threatened by this SCM bias, so the paper retains a credible core finding; however, the headline player-level claims (e.g., Seager's 271-point OPS gain) may be overstated. This does not change the reader's CONDITIONAL verdict, but it adds a concrete condition: the in-time placebo bias must be formally evaluated and, if present, corrected or the player-level results must be reframed as hypothesis-generating. I disagree with a fully unchecked reading that would accept the player-level point estimates at face value, but I also do not see grounds for REJECT given the DID result and the acknowledged limitations.","tokens_in":18011,"tokens_out":11171,"duration_ms":105661,"concrete_test":"Compute the mean (or median) in-time placebo effect estimate across the 30 target players from the 2022 pseudo-intervention re-fit (Appendix A, Figure A.4). Test whether this mean is significantly different from zero using the placebo distribution as a null. If the mean is significantly negative, apply it as a bias correction to each 2023 estimate; then determine whether any player (including Seager) still shows a substantial, 'significant' gain (e.g., OPS > 100 points). This directly tests whether the 2023 SCM estimates are an artifact of the systematic overprediction seen in 2022.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Appendix A Figure A.4 and the Results section report that the in-time placebo analysis (using 2022 as a dummy intervention year) yields 'moderate negative effect estimates for many players.' This is a critical diagnostic: if the synthetic control predicts outcomes that are systematically too high for high-shift players in a year without the ban, then the method has a predictive bias for exactly the population studied. The paper interprets these negative estimates as reflecting a real pre-existing decline in high-shift batters' performance, but this interpretation is not independently verified. Under the stated SCM assumption (Appendix B) that weights must be stable over time, a systematic in-time placebo effect is evidence against that assumption. If the negative bias persists in 2023, the positive 2023 estimates (e.g., Seager's 85-point OBP, 271-point OPS gains) would be overstated by that amount. Seager's synthetic OPS of 0.742 is well below his 2021–2022 level and heavily weights donors (Trea Turner, Carlos Correa) whose 2023 seasons declined sharply, raising the possibility that the counterfactual is an artifact of donor-specific declines rather than a valid prediction. Because these player-level estimates are a headline result of the abstract, this bias threatens the central claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper uses difference-in-differences (DID) and synthetic control methods (SCM) to estimate the effect of Major League Baseball's 2023 ban on infield shifts. The league-wide DID compares left-handed batters to right-handed batters in bases-empty plate appearances before and after the ban, reporting a 9-point increase in both BABIP and OBP for left-handed batters in 2023, with the effect persisting in 2024. The player-level SCM analysis focuses on 30 high-shift batters, using 58 low-shift batters as donors, and reports substantial estimated gains for several players, including Corey Seager (OBP +85 points, OPS +271 points, wOBA +115 points). The paper also discusses robustness checks, including placebo pre-trends, an in-unit placebo analysis, and an in-time placebo analysis using 2022 as a dummy intervention year, and provides all data and code in a GitHub repository.","tokens_in":18274,"tokens_out":5158,"duration_ms":51593,"significance":"If the league-wide estimate is credible, it provides a rigorous, reproducible quantification of a real policy change in professional sports and demonstrates the utility of quasi-experimental methods for sports analytics. The analysis is genuinely transparent: the synthetic control weights are shown in detail, placebo distributions are constructed, and all code and data are provided. The league-wide result is plausible and consistent with prior batted-ball analyses, and the 2024 persistence check is a valuable addition. However, the player-specific claims are substantially weaker than the abstract and discussion imply: the in-time placebo indicates a systematic divergence between targets and synthetic controls in a pre-intervention year, and the permutation p-values are not adjusted for the 90 tests performed. These issues do not invalidate the league-wide finding, but they do require a significant reframing and additional analysis of the player-level results.","major_comments":[{"comment":"The league-wide DID estimates of 0.009 for BABIP and OBP are presented without any measure of statistical uncertainty. No standard errors, confidence intervals, or p-values are reported for the 2023 comparison, so the reader cannot judge whether the 9-point effect is distinguishable from year-to-year noise. The placebo pre-trends and the 2024 persistence check are helpful, but they do not provide a formal test of the 2023 estimate. Please report a variance estimate (e.g., a cluster-robust or bootstrap standard error) and a confidence interval for the main DID estimates, and ideally for the placebo-year estimates as well.","section":"Results, Analysis 1; Table 1"},{"comment":"The in-time placebo analysis using 2022 as a dummy intervention year produces 'moderate negative effect estimates for many players.' This is a diagnostic failure of the synthetic control: under the stable-weights assumption stated in Appendix B, the synthetic control should track the target player's outcome in the absence of the intervention, but it systematically over-predicts the 2022 outcomes for high-shift players. The paper interprets this as a real pre-existing decline in high-shift batters' performance, but that interpretation is not independently verified. If the same over-prediction continues in 2023, the 2023 effect estimates are contaminated, although the direction of the bias is unclear (over-prediction would understate the positive effects, not overstate them). Please provide a sensitivity analysis that quantifies how the 2023 player-specific estimates change when adjusted for the mean (or player-specific) 2022 placebo effect, or otherwise demonstrate that the in-time placebo does not affect the substantive conclusions.","section":"Results, Analysis 2; Appendix A, Figure A.4; Appendix B, Assumptions"},{"comment":"The player-level p-values in Table 3 are unadjusted for multiple testing. With 30 target players and 3 outcomes, 90 hypotheses are effectively tested, and the smallest possible permutation p-value with 58 placebo players is 1/59 ≈ 0.017, which is exactly the value reported for Seager, Olson, Alvarez, and Ohtani. After any standard multiple-comparison correction (e.g., Bonferroni or Benjamini-Hochberg), none of the individual player effects would be statistically significant. The authors acknowledge that the p-values are 'not strict hypothesis tests,' but the abstract and discussion highlight individual players with 'substantial' gains as if they were individually confirmed. Please re-analyze the player-level results using a family-wise error control procedure, or clearly reframe the player-specific estimates as exploratory findings that are not individually confirmatory.","section":"Results, Analysis 2; Table 3"}],"minor_comments":[{"comment":"Table 1 does not report the number of plate appearances behind each rate, so the reader cannot assess the precision of the group-specific averages. Please add PA counts for each cell.","section":"Table 1"},{"comment":"Figure 1 would benefit from confidence bands or at least a statement about the variability of the annual DID estimates. Currently the pre-trend panel (C,D) shows point estimates only, making it hard to judge how 'centered around 0' the placebo estimates are.","section":"Figure 1"},{"comment":"The description of the placebo p-value in the text could be more precise: it should state that the p-value is the proportion of placebo estimates, including the target estimate, that are at least as extreme in absolute value, and that the minimum attainable value is 1/(number of placebos+1).","section":"Materials and Methods, Analysis 2"},{"comment":"The 'Weight Ranking' column in Table 2 is inconsistent across outcomes: for OBP only two weights are shown, while for OPS and wOBA five weights are shown. Use the same format for all three columns.","section":"Table 2"},{"comment":"The term 'average treatment effect on the treated' (ATT) is applied to the player-specific SCM estimates, but the estimand is a unit-specific effect for a single player, not an average over treated units. Consider using 'unit-specific treatment effect' or 'synthetic control effect' to avoid confusion.","section":"Discussion"},{"comment":"The paper states that the donor pool consists of 58 low-shift players, but Appendix B notes that the donor pool size changes for each target player because of the 250-PA requirement. Please report the distribution of donor pool sizes across the 30 target players, since this affects the resolution of the placebo p-values and the stability of the synthetic controls.","section":"Materials and Methods, Analysis 2"}],"recommendation":"major_revision","confidential_remarks":"The paper is a useful and reproducible methods demonstration, and the league-wide DID result is likely of interest to a sports-analytics audience. However, the player-level headline claims are currently oversold given the in-time placebo issue and the lack of multiple-testing control. I would be comfortable with a revised version that (1) adds uncertainty quantification to the DID estimates, (2) provides a sensitivity analysis for the in-time placebo bias, and (3) reframes the player-specific results as hypothesis-generating rather than confirmatory. The extensive appendices and code sharing are commendable and should be preserved."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The league-wide difference-in-differences result is the most solid part of this paper: a 9-point increase in BABIP and OBP for left-handed batters, supported by placebo pre-trends and a near-zero 2024 estimate that suggests persistence. The simple LHB/RHB comparison appeared in public analyst work before, but the formal panel design, the explicit statement of assumptions, and the full 2015-2023 series are a genuine step up. The synthetic control application is also new in this setting and is transparently presented, with code, data, and placebo checks. A reader who wants to know how quasi-experimental methods can be used in sports will get real value here.\n\nThe soft spots are what you would expect. Table 1 reports no confidence intervals or standard errors for the DID estimates, which is easy to fix and should be fixed. The player-level SCM results involve 30 targets times three outcomes, the p-values are not multiplicity-adjusted, and the donor pools are small. The paper is honest about this, explicitly framing the placebo p-values as screening indicators rather than strict tests. That is the right framing, and the reader should keep it.\n\nOn the stress-test note: the concern that negative in-time placebo estimates imply the player-level gains are inflated has the sign backwards. If the synthetic control over-predicts for high-shift players in a pre-ban year (which is what a negative placebo means), and that bias persists into 2023, then the estimated effect (observed minus synthetic) would be understated, not overstated. The author's interpretation that the negative placebo reflects a pre-existing decline in high-shift batters' performance is plausible, and it is consistent with the negative DID pre-trends for 2021-2022. So that specific worry does not land as stated. What does land is the more general point that the stable-weights assumption is untestable and the in-time placebo provides some evidence against it. The paper acknowledges this directly in Appendix B and the Discussion.\n\nWorth a serious referee. The right outcome is likely minor-to-moderate revision: add uncertainty quantification to Table 1, keep the screening language for the individual estimates, and maybe show the sensitivity of the Seager result to dropping Turner or Correa from the donor pool. I would not desk-reject this.","headline":"A credible league-wide DID estimate of the shift ban plus a genuinely new per-player synthetic control application; the player-level results are suggestive, and the in-time placebo concern is real but the sign argument in the stress-test note is backwards.","tokens_in":18763,"tokens_out":3443,"would_cite":false,"duration_ms":36097,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62D20"],"pacs":[],"model":"deepseek-v4-flash","headline":"MLB's 2023 ban on infield shifts raised left-handed batters' BABIP and on-base percentage by nine points each, and synthetic-control analysis attributes far larger offensive gains to individual high-shift players such as Corey Seager.","keywords":["difference-in-differences","natural experiments","panel data","quasi-experiments","sports analytics","synthetic control"],"falsifier":"Two checks would settle the player-level claims. First, re-run Seager's synthetic control with his most heavily weighted donor player (60% of the weight in the OPS fit) removed from the donor pool: if the estimated 271-point OPS gain collapses toward zero, the headline player result is an artifact of a single comparison player. Second, examine public batted-ball location data for left-handed ground balls: if the nine-point BABIP gain is really the ban working, the new hits should cluster in the vacated middle-infield zone, whereas if the gain appears instead in exit velocity or launch-angle distributions, it belongs to a different rule change or to luck.","tokens_in":17815,"feed_emoji":"⚾","tokens_out":18237,"duration_ms":150205,"temperature":0.7,"pith_summary":"This paper tries to establish whether Major League Baseball's 2023 ban on infield shifts achieved its aim of generating more offense, and it does so with two observational designs that borrow causal-inference methods from policy evaluation. It claims the ban raised left-handed batters' batting average on balls in play (BABIP) and on-base percentage by nine points each, a modest league-wide effect that persisted into 2024, while player-level synthetic control estimates attribute far larger gains to the most-shifted hitters, including a 271-point on-base-plus-slugging (OPS) gain for Corey Seager. A careful reader should care because the same tools apply to any sports rule change, injury, or tactical experiment in which a league provides no randomized control group.","feed_headline":"Nine-point gain for left-handed batters tied to MLB shift ban","feed_subtitle":"Synthetic controls attribute far larger gains, like Seager's 271-point OPS jump, to the heaviest-shift hitters.","key_machinery":"The argument runs on two quasi-experimental machines. Difference-in-differences compares left-handed batters, the group the shift primarily targeted, with right-handed batters, and subtracts the right-handers' year-to-year change in BABIP and on-base percentage from the left-handers' change across the 2022-2023 boundary, so that shared time trends cancel; its validity depends on the parallel-trends assumption. The synthetic control method builds a counterfactual for each of 30 high-shift target players as a weighted average of low-shift donor players, with weights between zero and one summing to one, chosen so that the donor composite matches the target's pre-2023 trajectory in the outcome, age, plate appearances, hits, singles, home runs, walk rate, and strikeout rate; the composite's actual 2023 value is the counterfactual, and the gap between it and the target's real outcome is the estimated effect of the ban. Placebo procedures that re-fit the same machinery for the donor players themselves and for a mid-shift group of players provide a null distribution, while an in-time placebo using 2022 as a fake intervention year checks for systematic overestimation.","core_discovery":"The central claim is that the 2023 shift ban raised offensive output for the batters it most affected, and that the size of the effect varied predictably with how often a batter had been shifted. The league-wide difference-in-differences estimate puts the effect on left-handed batters' batting average on balls in play and on-base percentage at nine points each, with near-zero estimates in prior no-policy years and a 2024 estimate of -0.001, which indicates the gain persisted. The player-level synthetic control estimates attribute the largest gains to the most-shifted hitters: Corey Seager, whose 2022 shift rate was 92.8%, is estimated to have gained 85 points of on-base percentage, 271 points of on-base-plus-slugging (OPS), and 115 points of weighted on-base average (wOBA) in 2023, and four players (Seager, Matt Olson, Yordan Alvarez, and Shohei Ohtani) show estimated wOBA gains above 80 points. The paper also finds a dose-response pattern: each additional 10 percentage points of 2022 shift rate maps to roughly 11 points of on-base percentage, 31 points of OPS, and 17 points of wOBA in estimated gain.","pith_inferences":["An unstated implication is a valuation problem: if Seager's 271-point OPS gain is real, the shift was suppressing roughly a quarter of his offensive production, and teams' pre-2023 willingness to accept that suppression implies they judged the shift's run-prevention value to exceed the batter's adjustment cost, a tradeoff the new rule has now settled.","The league-wide design cannot separate the shift ban from the other 2023 rule changes, such as the pitch clock and larger bases; a clean extension would replicate the lefty-versus-righty contrast in minor-league seasons where the shift ban was introduced in isolation.","The same synthetic-control machinery could be turned on injuries or suspensions, since a player's pre-injury donor composite is already constructed here and could estimate the counterfactual season a team lost.","A batted-ball test follows directly: if the ban caused the effect, left-handed hitters' gains should concentrate in ground balls through the vacated middle-infield gaps, and if the gains instead appear in exit velocity or launch-angle changes, the attribution to the shift ban weakens."],"forward_implications":["The league-wide effect, nine points of BABIP and on-base percentage for left-handed batters in bases-empty situations, amounts to roughly one additional on-base event per 500 plate appearances, so the ban was a modest success on its own terms.","The near-zero 2024 difference-in-differences estimate indicates the 2023 gain did not revert, meaning teams' counter-strategies such as 'strategic shades' did not erase the effect.","Because the paper finds a dose-response relationship between 2022 shift rate and 2023 gain, future rule changes intended to help a specific subset of players can be evaluated at both league and player levels with the same design.","The synthetic control estimates are average treatment effects on the treated for the 2023 season specifically, so they are not directly generalizable to other seasons, leagues, or rule environments without further assumptions.","The player-level estimates are transparent in a way batted-ball models are not: the weighted donor players are listed explicitly, so analysts can judge whether each comparison is plausible.","Most target players (over 75%) show positive estimated effects across all three outcomes, with target-player estimates averaging four to eight times the size of placebo-player estimates."],"supporting_citations":[{"why":"Supplies the synthetic control method's formal requirements, including the weight-stability assumption the player-level analysis depends on.","marker":"Abadie (2021)"},{"why":"Establishes the synthetic control estimator and the placebo-test logic used to build null distributions for each player estimate.","marker":"Abadie et al. (2010)"},{"why":"The canonical difference-in-differences application that motivates the league-wide lefty-versus-righty comparison.","marker":"Card and Krueger (1994)"},{"why":"Frames difference-in-differences and synthetic control as complementary policy-evaluation tools and states the parallel-trends assumption.","marker":"Basu et al. (2017)"},{"why":"The only prior formal causal analysis of the infield shift, whose pre-ban run-prevention estimate the results are triangulated against.","marker":"Markes et al. (2024)"},{"why":"Batted-ball modeling predicted Seager would benefit from the ban, and the synthetic-control estimate for Seager is positioned against that prediction.","marker":"Petriello (2023b)"},{"why":"Within-batter comparisons of shifted and non-shifted plate appearances whose predicted range brackets the league-wide nine-point estimate.","marker":"Carleton (2022b)"},{"why":"Supplies the league-wide BABIP and on-base percentage data by batter handedness for bases-empty plate appearances used in the difference-in-differences analysis.","marker":"Splits Leaderboard (2024)"},{"why":"Supplies the player-season batting statistics used to fit and evaluate the synthetic controls.","marker":"Statcast Custom Leaderboard (2024)"},{"why":"Supplies the shift-rate data that assigns players to high-shift target and low-shift donor groups.","marker":"Statcast Batter Positioning Leaderboard (2024)"}],"fun_headline_variants":["MLB shift ban adds nine points to lefty on-base percentage","Shift ban: nine points for lefties, 271 for Seager","Seager's 271-point OPS surge tied to shift ban","Most-shifted hitters gain most from MLB shift ban","Nine-point lefty gain from shift ban persists"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The player-level results stand on the assumption that the weighted mix of low-shift players who tracked a target's stats before 2023 would have kept tracking his stats in 2023 and 2024 if the shift ban had never happened, so a changed role, injury, or aging divergence would masquerade as an effect of the ban.","fun_headline_variants_meta":{"raw":{"variants":["MLB shift ban adds nine points to lefty on-base percentage","Shift ban: nine points for lefties, 271 for Seager","Seager's 271-point OPS surge tied to shift ban","Most-shifted hitters gain most from MLB shift ban","Nine-point lefty gain from shift ban persists"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001276,"raw_usage":{"total_tokens":5262,"prompt_tokens":1032,"completion_tokens":4230,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":648,"completion_tokens_details":{"reasoning_tokens":4144}},"tokens_in":648,"tokens_out":4230,"duration_ms":29393,"temperature":1.0,"reasoning_tokens":4144,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:31:42.252597+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Two checks would settle the player-level claims. First, re-run Seager's synthetic control with his most heavily weighted donor player (60% of the weight in the OPS fit) removed from the donor pool: if the estimated 271-point OPS gain collapses toward zero, the headline player result is an artifact of a single comparison player. Second, examine public batted-ball location data for left-handed ground balls: if the nine-point BABIP gain is really the ban working, the new hits should cluster in the vacated middle-infield zone, whereas if the gain appears instead in exit velocity or launch-angle distributions, it belongs to a different rule change or to luck.","supporting_citations":[{"cited_title":"Splits Leaderboard","cited_arxiv_id":null,"evidence_quote":"Supplies the league-wide BABIP and on-base percentage data by batter handedness for bases-empty plate appearances used in the difference-in-differences analysis."}],"review_version":1}