{"id":"3bf8d69d-5579-490a-9091-1a346eb48b63","arxiv_id":"2607.21312","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A rank-based DID test uses the historical distribution of absolute pre-treatment differential trends as the reference distribution for the post-treatment effect, allowing non-parallel trends.","lead":"This paper proposes a way to test for a treatment effect in difference-in-differences by comparing the final outcome gap between two groups with the gaps seen in earlier, pre-treatment periods. The method works without assuming parallel trends, using instead the stability of the size of the differential trend over time.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Proof is sound; the load-bearing risk is assumption (2), whose violation can cause large size distortions. Since (2) is explicit, this is an applicability caveat, not a correctness flaw.","rationale":"The paper presents a clean asymptotic conformal argument. Proposition 1 follows from (1)-(3): under H0, the test statistic converges in distribution to Unif(0,1) by the probability integral transform and uniform consistency of the empirical CDF. I checked the proof carefully; there is no hidden circularity or missing step. The DID specialization is sensible, and the contrast with parallel-trends-based conformal predictors is informative. The weakest point is condition (2), exactly as the reader stated. This condition is substantive: it requires the entire distribution of absolute differential trends to be time-invariant. The paper's claim that the method does not require parallel trends is true, but it substitutes a different strong assumption. The method does allow the mean differential trend to vary via sign dynamics, which is a genuine relaxation, but it does not allow the magnitude distribution to drift. In many empirical settings—long panels with macroeconomic fluctuations or structural breaks—that magnitude is likely to change, and the test will then over- or under-reject. However, this is an assumption explicitly stated by the author, not a flaw in the proof. I would therefore keep the ACCEPT verdict while perhaps encouraging the author to add a sensitivity analysis or at least a discussion of how practitioners might assess the plausibility of (2). The proposed simulation would demonstrate the practical stakes of the assumption. Overall, the paper is a sound theoretical contribution, and the reader's assessment is accurate.","tokens_in":5131,"tokens_out":10091,"duration_ms":105663,"concrete_test":"Simulate T=50 pre-periods and one post-period with U_t(0)=sigma_t*epsilon_t, epsilon_t iid N(0,1), sigma_t=1+delta*(t/T), and tau=0. For delta in {0, 0.25, 0.5, 1}, compute the rejection rate of the test at alpha=0.05 over 10,000 repetitions. If the rejection rate exceeds 0.08 for delta=0.25, the procedure is materially sensitive to mild violations of (2), confirming that the assumption is not merely technical but central to practical validity.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—asymptotic exactness of the DID pre-trend test—depends entirely on condition (2): the distribution of |U_t(0)| is the same for every t, including the post-treatment period under H0. If the scale of shocks to the untreated differential trend changes over time (e.g., variance grows because of macroeconomic volatility, policy changes, or composition shifts), then the pre-treatment empirical distribution is the wrong reference for the post-treatment trend. The paper offers no measure of how rejection rates or confidence-set coverage degrade as a function of deviations from (2), and no pre-test or sensitivity check. The sufficient condition for (3) via strict stationarity of the joint increment process is strong and often implausible in long panels. The theorem itself is proven correctly under (1)-(3); the concern is about the gap between the theorem and practical reliability. This is the same assumption the reader flagged, and it is genuinely load-bearing, but it is an explicit assumption rather than an internal inconsistency.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a nonparametric, rank-based test for the average treatment effect on the treated in the last period of a two-group difference-in-differences design. The test uses the empirical distribution of absolute pre-treatment differential trends as a reference distribution: for a candidate effect x, the post-treatment differential trend adjusted by x is compared with the T pre-treatment absolute differential trends. The main theoretical result, Proposition 1, shows that under an assumption that the absolute untreated differential trend has a constant continuous distribution over time (condition (2)) and the empirical CDF is uniformly consistent (condition (3)), the test is asymptotically exact. Confidence sets are obtained by inverting the test. The paper also positions the method relative to conformal inference and the parallel-trends literature, arguing that it relaxes parallel trends while imposing a different, explicit stationarity-type restriction.","tokens_in":5372,"tokens_out":6274,"duration_ms":73571,"significance":"If the result is taken at face value, the paper contributes a very simple and transparent inference procedure for DiD settings with long pre-treatment panels, with no parallel-trends requirement. The proof of Proposition 1 is short, clear, and mathematically correct. The paper is also honest about the substantive nature of condition (2). However, the practical value is not demonstrated: there are no simulations or empirical applications, and the load-bearing assumption (2)—time-invariance of the entire distribution of |U_t(0)|—is strong. The novelty relative to existing conformal methods, while real, is incremental: it is a specific predictor for DiD within the Chernozhukov, Wüthrich, and Zhu (2021) framework. The paper is better viewed as a useful theoretical note than as a fully developed inference method.","major_comments":[{"comment":"The asymptotic exactness of the test and the validity of the confidence set depend entirely on the assumption that P(|U_t(0)| ≤ u) = F(u) for every t, including the post-treatment period under H0. The paper's advertised advantage over parallel-trends-based methods is that it allows stable violations of parallel trends, but condition (2) is itself a strong restriction: it requires the entire distribution of absolute differential trends, including tails and quantiles, to be time-invariant. The manuscript provides no quantification of how rejection rates or coverage degrade if this distribution changes over time—for example, if the variance of differential trends grows or shrinks. Because this is the central identifying restriction, the paper should include a formal sensitivity analysis (e.g., a bound under a perturbation of F) or a simulation study showing the method's behavior under reali","section":"§3.2, condition (2)"},{"comment":"The test uses the empirical CDF \\widehat F that includes the post-treatment observation itself, so the p-value is the in-sample rank of the post-period statistic among T observations, not a leave-one-out conformal p-value. While the asymptotic argument is correct, for finite T the actual rejection probability can differ from α by O(1/T) and can be notably higher than α for small T (e.g., with T=50 and α=0.05, the in-sample rank can cause the level to be about 0.06). The paper contains no finite-sample analysis or simulation evidence to show how large T must be before the asymptotic approximation is reliable. This is a practical concern because the method is explicitly motivated by settings with 'many' pre-treatment periods, and the paper should provide guidance on what 'many' means.","section":"§3.3, Proposition 1"},{"comment":"The sufficient condition for condition (3)—strict stationarity and ergodicity of the joint increment process—is strong and often implausible in long panels where group composition, measurement, or the economic environment changes. The paper offers no discussion of how a researcher could assess the plausibility of condition (2) from the pre-treatment data, nor the implications of conducting a pre-test for (2) on the final inference (which would itself be a form of pre-testing with its own distortions, as in Roth 2022). A practical recommendation, even a tentative one, would help close the gap between the theorem and its application.","section":"§3.2, sufficient conditions for (3)"}],"minor_comments":[{"comment":"The indicator function in the definition of \\widehat F(u;x) is written as '1{...}', which is standard but could be typeset as \\mathbf{1}{...} for clarity, especially since '1{t=T}' appears later as the same indicator notation.","section":"§2, notation"},{"comment":"The claim that the proposed predictor 'does not imply parallel trends' should be phrased more precisely: it does not require parallel trends, but it does require a different and equally substantive distributional restriction. The current phrasing could be read as suggesting the method is assumption-free, which it is not.","section":"§4.1"},{"comment":"The discussion of Rambachan and Roth (2023) is brief. Given that the paper's method is a special case of extrapolating pre-trends via distributional stability, a fuller comparison of the restrictions imposed by condition (2) with the restrictions imposed by their sensitivity-analysis approach would be helpful.","section":"§4.2"},{"comment":"The paper has no conclusion or limitations paragraph. A brief final section summarizing the identifying assumption, the scope of applicability, and open directions (e.g., multiple treated periods, covariates, cluster dependence) would make the note more self-contained.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The theoretical result is clean and the proof of Proposition 1 is correct. However, the manuscript is extremely short for the claim it makes: it positions itself as a practical inference method for DiD, yet provides no simulation evidence, no empirical illustration, and no sensitivity analysis with respect to the load-bearing assumption (2). The concern is not internal inconsistency but rather the gap between the asymptotic theorem and the practical reliability of the procedure. A revision that adds a limited simulation study (including small-T behavior and violations of condition (2)) and a more careful discussion of the identifying restriction would substantially strengthen the contribution. I would then be willing to reconsider the paper favorably."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper does something genuinely new: it takes the absolute pre-treatment DID trends as a reference distribution for the post-treatment DID and tests a candidate effect by ranking the adjusted post DID inside that historical distribution. The identifying condition is not parallel trends but time-invariance of the distribution of |U_t(0)|. That move is simple, clearly explained, and correctly differentiates the procedure from Chernozhukov-Wüthrich-Zhu's conformal DID predictor, which does imply parallel trends. Proposition 1 is proved correctly: under (1)-(3), the empirical CDF converges uniformly to F, the probability integral transform applies, and rejection probability goes to alpha. The paper is honest about what the assumption requires and even flags its own novelty claim as provisional.\n\nCredit where due: this is a clean piece of theory. It puts a new tool on the table for applied DID with many periods. It is not a breakthrough but a solid, well-scoped contribution.\n\nSoft spots: The load-bearing condition is (2). It requires the whole distribution of absolute untreated differential trends to be stable over time, including tails. If macro volatility or composition shifts make the scale of shocks grow, the pre-period empirical distribution is the wrong benchmark and size and coverage will be off. The paper gives no sense of how bad this gets under plausible deviations, and no simulations. The stationarity sufficient condition for (3) is strong. These are applicability caveats, not internal errors. The theorem does what it claims under stated assumptions. Also, the absolute value throws away sign; that is fine for this test but means the power profile is symmetric in a way that may be worth discussing.\n\nWho this is for: applied DID researchers who have many pre-periods and want an inference method that does not rest on parallel trends. A reader who wants a robustness check alongside Rambachan-Roth will find this useful. It deserves a serious referee. I would send it out rather than desk reject. The main ask to the author for revision is some simulations showing size and power under DGPs that satisfy and mildly violate (2), and a sensitivity discussion.","headline":"New DID inference method that uses absolute pre-trends as a reference distribution, with a correct proof and an explicit but strong identifying assumption that limits its practical reach.","tokens_in":5799,"tokens_out":1326,"would_cite":true,"duration_ms":13700,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G10","62P20"],"pacs":[],"model":"deepseek-v4-flash","headline":"A difference-in-differences test that uses the historical distribution of pre-treatment differential trends as its reference distribution, without requiring parallel trends, is asymptotically exact.","keywords":["difference-in-differences","pre-trends","conformal inference","distribution-free testing","parallel trends","treatment effect","time series","rank-based inference"],"falsifier":"Simulate a DID panel in which the variance of the untreated differential trend increases monotonically over time (e.g., shocks scale with time); compute the pre-treatment empirical distribution of absolute DIDs from the early periods and apply the test to a later post-treatment period with a zero treatment effect. If the rejection frequency exceeds the nominal α substantially, the test's validity depends on the time-invariance of the absolute-trend distribution in a way that a researcher can check empirically.","tokens_in":5044,"feed_emoji":"📊","tokens_out":3602,"duration_ms":34942,"temperature":0.7,"pith_summary":"The paper proposes a nonparametric test for the treatment effect in difference-in-differences when many pre-treatment periods are available. Instead of assuming parallel trends, it assumes that the absolute value of the untreated differential trend between treated and control groups has a time-invariant distribution. Under that condition, the test compares the absolute post-treatment DID with the empirical distribution of absolute pre-treatment DIDs and rejects if it falls in the top α of the historical distribution. The test is asymptotically exact and does not require the mean differential trend to be zero or constant. A confidence set for the effect follows by inverting the test.","feed_headline":"No parallel trends needed: pre-trends give valid DID inference","feed_subtitle":"A simple rank-based test uses the historical spread of differential trends to test treatment effects.","key_machinery":"The central object is the untreated differential trend U_t(0), defined as the change in the treated group's untreated outcome minus the contemporaneous change in the control group's untreated outcome. The paper assumes its absolute value |U_t(0)| has the same continuous distribution F at every date (condition 2), and that the empirical distribution of the observed adjusted absolute trends converges uniformly to F (condition 3). The test ranks the post-treatment absolute DID among its pre-treatment counterparts; the probability integral transform converts this rank into an asymptotically uniform variable. A stationary and ergodic joint increment process is offered as a sufficient condition fo","core_discovery":"The core claim is that the distribution of the absolute pre-treatment difference-in-differences can be used as the reference distribution for the absolute post-treatment difference-in-differences, once the hypothesized treatment effect is subtracted. The test statistic is the empirical CDF of the adjusted absolute DIDs evaluated at the post-treatment value. Under the null, the probability integral transform makes this statistic asymptotically uniform, so rejecting when it exceeds 1−α gives an asymptotically exact test. This identifying restriction is distributional stability of |U_t(0)|, not parallel trends; stable nonzero, time-varying, or sign-changing differential trends are permitted.","pith_inferences":["If the assumption of a stable absolute-trend distribution is violated—for instance, if the variance of differential shocks grows over time—the test's size will not be controlled, and the pre-treatment distribution will be a misleading reference for the post-treatment period.","The same ranking logic could be extended to settings with multiple treated groups or multiple post-treatment periods, by pooling absolute differential trends across units and dates under an appropriate stability condition.","The test may have limited power when pre-treatment variability is large relative to plausible effect sizes, since the post-treatment DID must exceed a historical quantile to be flagged.","A direct, testable extension would be to check the stability of the absolute differential-trend distribution by using earlier sub-samples of the pre-period to predict the distribution of later pre-period differential trends."],"forward_implications":["Applied DID studies with many pre-periods can conduct inference on the treatment effect without invoking parallel trends, as long as the magnitude of differential trend shocks is stable over time.","Confidence sets for the treatment effect are obtained by inverting the test, which is simple to compute from the historical distribution of absolute DIDs.","The method permits a stable nonzero differential trend between groups, a common concern in observational panels.","Because only absolute differential trends are used, the distribution of the sign of pre-trends is unrestricted; trends may oscillate in sign over time.","The test is closely tied to a calibration principle already used in the literature, but the DID-specific predictor used here changes the identifying assumption from parallel trends to distributional stability of absolute differential trends."],"fun_headline_variants":["Pre-trends as reference: DID inference without parallel trends","DID validity from pre-trend spread, not parallel paths","Test DID effects with pre-period trend variations","Leverage pre-trends: valid DID tests without parallel trends","DID inference: Use pre-trends to gauge post-treatment shocks"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is condition (2): the distribution of the absolute untreated differential trend is identical at every date, so the historical spread of pre-treatment differential trends is the correct reference for the post-treatment period.","fun_headline_variants_meta":{"raw":{"variants":["Pre-trends as reference: DID inference without parallel trends","DID validity from pre-trend spread, not parallel paths","Test DID effects with pre-period trend variations","Leverage pre-trends: valid DID tests without parallel trends","DID inference: Use pre-trends to gauge post-treatment shocks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000747,"raw_usage":{"total_tokens":3090,"prompt_tokens":590,"completion_tokens":2500,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":334,"completion_tokens_details":{"reasoning_tokens":2414}},"tokens_in":334,"tokens_out":2500,"duration_ms":17149,"temperature":1.0,"reasoning_tokens":2414,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T07:49:01.638583+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a DID panel in which the variance of the untreated differential trend increases monotonically over time (e.g., shocks scale with time); compute the pre-treatment empirical distribution of absolute DIDs from the early periods and apply the test to a later post-treatment period with a zero treatment effect. If the rejection frequency exceeds the nominal α substantially, the test's validity depends on the time-invariance of the absolute-trend distribution in a way that a researcher can check empirically.","supporting_citations":[],"review_version":1}