{"id":"624de69e-f7b8-41be-91a1-fa3092aed8e1","arxiv_id":"2608.08405","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"A same-date comparison of differently sized strategy sleeves identifies only the sleeve's own response at fixed aggregate positioning, never the strategy's total capacity under crowding, because an arbitrary calendar effect can absorb the aggregate crowding exactly.","lead":"A proposed experiment to measure how much capital a trading strategy can absorb compares parallel versions of the strategy running at different deployment scales on the same dates. The paper shows that this comparison removes market shocks but also removes the crowding effect capacity is meant to measure, so it can only identify a narrower own-sleeve effect unless the experiment deliberately varies how sleeves overlap.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The non-identification result presumes additive separability of own and aggregate erosion in Assumption 3.2; if own and aggregate crowding interact, within-date contrasts can carry crowding information as aggregate deployment varies, so the stated trade-off is not a property of within-date designs.","rationale":"The reader's weakest assumption correctly located the exposure model, but the more precise load-bearing premise is additive separability of own and aggregate erosion, not only homogeneous overlap. Even with homogeneous overlap, a non-additive term lambda W_pt A_t breaks the span(iota) argument because the aggregate component is no longer a common additive shift: the within-date contrast becomes a function of the common aggregate level A_t, so an experiment that varies A_t across blocks while keeping contrasts within dates can recover one channel of crowding without calendar exposure. The paper is internally consistent under Assumption 3.2; the concern is scope and external correctness risk, because the abstract and introduction state the robustness-versus-crowding trade-off without this caveat. This does not overturn the conditional theoretical result, and the reader's CONDITIONAL verdict remains appropriate, but the condition should explicitly include additive separability, not just homogeneous overlap. The proposed analytical check would settle whether the impossibility is an artifact of the additive model or extends to interacting exposures.","tokens_in":30028,"tokens_out":26042,"duration_ms":291201,"concrete_test":"Re-derive the observation-equivalence for Y_pt = mu_t - c_own(W_pt) - c_agg(A_t) - lambda W_pt A_t + epsilon_pt, with A_t common and varying across blocks. Check whether the within-date contrast means, -[c_own(beta_1)-c_own(beta_0)] - lambda(beta_1-beta_0) A_t, identify lambda and the own-scale difference by OLS on A_t while exactly differencing out mu_t. If they do, Proposition 3.18 fails as soon as lambda is nonzero; if they do not, identify the step in the proof that requires the cross term to be absent.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Proposition 3.18's exact equivalence, (mu_t, c_agg) vs (mu_t - c_agg(gamma^T W_t), 0), holds because Assumption 3.2 restricts erosion to the additive pair c_own(W_pt) + c_agg(gamma_p^T W_t); the aggregate term then enters Y_pt as a common shift that any unrestricted date effect can absorb. The paper does not flag additive separability as a maintained premise, and Remark 3.8 concedes that the scale response is only a local linearisation. Suppose instead that Y_pt = mu_t - c_own(W_pt) - c_agg(A_t) - lambda W_pt A_t + epsilon_pt, with A_t = gamma^T W_t common to all sleeves. The within-date contrast between scales beta_1 and beta_0 has mean -[c_own(beta_1)-c_own(beta_0)] - lambda(beta_1-beta_0) A_t, which varies with A_t. An experiment that randomizes aggregate deployment A_t across blocks while still comparing sleeves within a date would identify lambda, a component of aggregate crowding, with zero exposure to the arbitrary calendar component mu_t. This contradicts the headline mutual exclusion. Thus the central result is a theorem about an additive-exposure model, not about within-date designs as such; since the abstract states the trade-off unconditionally, the central claim is broader than the proof supports.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper asks what controlled experiment would measure the capital capacity of a trading strategy. It proposes parallel sleeves of one strategy, randomly assigned to deployment scales, held for L periods, and contrasted within the same calendar dates. The main formal results are: (i) under geometric erosion, a finite hold recovers only a fraction F_L or G_L of the steady-state effect, and the choice of block summary is a choice of estimand; (ii) the within-date sleeve contrast identifies only the own-sleeve controlled direct effect at a fixed average deployment; (iii) under homogeneous overlap (Prop 3.18), aggregate crowding is absorbed by an unrestricted date effect and is not identified from within-date variation by any estimator; (iv) heterogeneous overlap or cross-date variation can restore some identification at a cost quantified by the scaling formulas; and (v) a calibration on thirteen long–short strategies prices a realistic study, compares it with impact-model capacity bounds, and is frankly labeled as a scenario rather than an estimate. The paper is unusually explicit about transported bridges, one-sided bounds, and the distinction between identified sets and confidence sets.","tokens_in":30287,"tokens_out":12295,"duration_ms":144071,"significance":"If the central identification theorem is accepted, the paper makes a substantive design contribution: it separates estimands that are often conflated, gives a sharp identified set for capacity, derives explicit minimum-variance recovery weights, quantifies replication saturation, and prices the opportunity cost of arm placement. The proof of Prop 3.18 is clean and does not assume its conclusion, and the constructive reading in Remark 3.21 is valuable. The paper also earns credit for reporting its limitations honestly: the transported persistence and publication bridges are stated as untestable, the scale response is described as a local linearization, and the model-free one-sided bound is presented alongside the model-assisted point estimate. The main caveat is that the headline mutual-exclusion claim is proved only for the additive exposure model, so the scope of the claim needs to be qualified before publication.","major_comments":[{"comment":"The non-identification result and the mutual-exclusion claim are proved under Assumption 3.2, which restricts erosion to the additive pair c_own(W_pt) + c_agg(gamma_p^T W_t). This additive separability is load-bearing and is not flagged as a maintained restriction. If instead the outcome were Y_pt = mu_t - c_own(W_pt) - c_agg(A_t) - lambda W_pt A_t + eps_pt with A_t = gamma^T W_t common to all sleeves, then a within-date contrast between scales beta_1 and beta_0 would have mean -[c_own(beta_1)-c_own(beta_0)] - lambda(beta_1-beta_0)A_t, which varies with A_t. An experiment that randomizes aggregate deployment A_t across blocks while still comparing sleeves within a date would identify lambda with zero exposure to the arbitrary calendar component mu_t. Thus 'robustness and aggregate identification are mutually exclusive' is a theorem about the additive exposure model, not about within-date designs as such. The abstract and introduction state the trade-off unconditionally; Remark 3.8's local-linearization caveat does not address the cross-partial term. Please either extend the model to allow an own-aggregate interaction and characterize when the exclusion survives, or qualify Assumption 3.2, the abstract, and the introductory claims accordingly.","section":"Assumption 3.2; Prop 3.18; abstract and §1"}],"minor_comments":[{"comment":"Numbered results are often called 'Theorem' in the text but stated as 'Proposition' (for example, 'Theorem 3.13' for Proposition 3.13, 'Theorem 3.18' for Proposition 3.18, and 'Theorem B.2' for Proposition B.2); the cross-references should be harmonized with the actual labels.","section":"Throughout"},{"comment":"The sentence 'By Theorem 3.1 the observed adjusted return is a function of the full deployment vector' refers to a Theorem 3.1 that does not appear in the paper; this appears to be a copy-editing error.","section":"Assumption 3.1"},{"comment":"The first paragraph repeats the sentence 'Three consequences follow for practice.'; one copy should be removed.","section":"Section 5"},{"comment":"The notation Delta_infty = -c(beta) is correct only for a zero control arm; for a general arm pair (beta, beta_0) the steady-state contrast is -[c(beta)-c(beta_0)], and this should be stated wherever the attenuation formula is applied to nonzero baseline arms.","section":"Prop 3.16 and surrounding text"},{"comment":"The value ||M_iota g_t|| = 0.27 for six sleeves in two disjoint blocks is stated without derivation; since it is used to quantify the gain from heterogeneous overlap, a short derivation or an explicit reference to Section G would help the reader.","section":"Remark 3.21"},{"comment":"The sentence 'the comparison that makes the experiment robust is the one that prevents it from measuring the crowding capacity is about' is ungrammatical and should be rewritten.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The paper is a methods and design contribution that is squarely within the journal's scope; the calibration section is clearly labeled a scenario and is appropriately hedged. The central proof is sound under the stated additive model, so I would not reject on that basis. The main risk is overstatement of the mutual-exclusion result: the paper should either broaden the model or narrow the claims, and the abstract and introduction need to be adjusted accordingly. The authors might also be asked to check the citation balance around the commonality criterion from their own prior work, but this is not a blocking issue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth reading: Proposition 3.18 and Corollary 3.19 give a genuinely new observation-equivalence result. Under homogeneous overlap and the paper's additive exposure model, an unrestricted calendar component can absorb aggregate crowding exactly, so within-date contrasts identify only the own-sleeve direct effect. That result is not in the switchback or interference literatures they cite, and the proof is clean. The attenuation algebra (F_j = 1 - a^j, G_L, the oracle deattenuation weights) is also solid, and the Monte Carlo tables reproduce the variance formulas to the third decimal. Credit where due: they are explicit that the persistence and scale-response calibrations sit on untestable bridges, and the pre-registration checklist is a useful disciplinary contribution.\n\nWhere I would push back: the headline mutual exclusion is stated more broadly than the proof supports. Assumption 3.2 restricts erosion to c_own(W_pt) + c_agg(avg_t). If own and aggregate crowding interact, say via lambda * W_pt * A_t, then the within-date contrast has mean that depends on A_t, and randomizing aggregate deployment across blocks while keeping the within-date comparison would identify that interaction component with zero exposure to the arbitrary calendar effect. So the impossibility is a property of the additive-exposure model, not of within-date designs as such. The paper flags homogeneous overlap as the key condition and even shows heterogeneous overlap breaks it (Remark 3.21), but it does not flag additive separability. The theorem is correct as stated; the abstract and some of the surrounding prose should carry the caveat.\n\nThe calibration section is honest but brittle. The transported a = 0.9177 and c(1) = 0.123 are the load-bearing inputs, and the paper itself reports that the panel bootstrap interval for the decay is essentially [0, 0.99]. The cost tables are a scenario, not an estimate. Fine as illustration, but readers should not walk away with 42 years or 6 years as point forecasts. Minor issues: no code or data shipped, and cross-referencing to theorems is inconsistent in places. Both are fixable.\n\nBottom line: this is a serious paper with one real identification result and a lot of careful supporting analysis. It deserves a serious referee. I would send it to peer review with a request to expand the scope discussion around additive separability, soften the unconditional framing in the abstract, and ideally ship the simulation code.","headline":"The core non-identification result is real and clean, but it is a theorem about an additive-exposure model, not about all within-date designs as the abstract implies.","tokens_in":30880,"tokens_out":4080,"would_cite":true,"duration_ms":46790,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A within-date experiment comparing parallel sleeves cannot measure aggregate crowding, because an unrestricted calendar effect absorbs it exactly.","keywords":["trading strategy capacity","crowding","experimental design","interference","causal inference","market impact","partial identification","switchback experiments"],"falsifier":"Measure the overlap matrix $\\Gamma$ of a real multi-sleeve book from daily executed positions and form a within-date contrast between two sleeves with materially different $\\Gamma$ rows while holding own-scale fixed: if the contrast shows a component that moves with measured aggregate deployment across dates, the homogeneous-overlap premise behind the non-identification result fails in the field.","tokens_in":1917,"feed_emoji":"📉","tokens_out":5837,"duration_ms":123263,"temperature":0.7,"pith_summary":"This paper asks what experiment could measure how much capital a trading strategy can absorb before its edge disappears, and concludes that the most credible experiment—running parallel implementations of the same strategy, called sleeves, on the same dates—cannot by itself answer the question. Deployed capital erodes the edge through a slowly dissipating stock, and that stock is common to every sleeve on a given date, so an unrestricted calendar-date effect can absorb it exactly. A same-date contrast therefore identifies only a sleeve's own response at the prevailing average deployment, not the strategy-level capacity that capacity statements refer to. The paper characterises what can be recovered through deliberately heterogeneous overlap between sleeves or through cross-date variation in average deployment, how much a fixed holding period understates the steady-state effect, and what a finite grid of deployment levels can and cannot reveal.","feed_headline":"Same-date experiments can't see strategy crowding","feed_subtitle":"Within-date sleeve comparisons miss aggregate crowding; heterogeneous overlap or cross-date exposure is required.","key_machinery":"The load-bearing object is the exposure mapping of Assumption 3.2 combined with the balanced-design assumption 3.4: each sleeve's interference enters through the overlap matrix $\\Gamma$, and in the homogeneous case every row of $\\Gamma$ equals a common vector $\\gamma^\\top$ (uniform overlap $\\Gamma = N^{-1}\\iota\\iota^\\top$ gives average deployment). This puts the aggregate crowding vector $g_t = c_{\\mathrm{agg}}(\\gamma^\\top W_t)\\,\\iota$ in $\\mathrm{span}(\\iota)$, making it observationally equivalent to the unrestricted calendar component $\\mu_t$. The argument then runs through contrast weights: $q^\\top\\iota = 0$ removes $\\mu_t$ for arbitrary $\\mu_t$, and because $g_t$ is in $\\mathrm{span}(\\iota)$, the same condition removes aggregate crowding. The within-block trajectory of contrasts, summarised by the accumulation fractions $F_j$, carries the attenuation and deattenuation results, while $\\|M_\\iota g_t\\|$ measures how much crowding information survives when overlap is heterogeneous.","core_discovery":"The central discovery is a non-identification result for aggregate crowding under homogeneous overlap. When every sleeve's exposure to the strategy's aggregate position enters through the same scalar $\\gamma^\\top W_t$, the aggregate crowding term $g_t = c_{\\mathrm{agg}}(\\gamma^\\top W_t)\\,\\iota$ lies in the span of the calendar intercept, and the pair $(\\mu_t, c_{\\mathrm{agg}})$ is observationally equivalent to $(\\mu_t - c_{\\mathrm{agg}}(\\gamma^\\top W_t), 0)$ at every date. Hence no estimator using within-date variation can identify aggregate crowding: the zero-sum contrast that removes arbitrary calendar shocks also removes the crowding term. A same-date design recovers only the controlled direct effect $\\mathrm{CDE}(\\beta,\\beta_0)$ of a sleeve's own scale at a fixed average deployment, scaled by the attenuation factor $R_L(u)$; reaching the strategy-level total effect $TE_{\\mathrm{strat}}$ requires either heterogeneous overlap, which moves $g_t$ out of $\\mathrm{span}(\\iota)$, or cross-date variation that accepts calendar exposure.","pith_inferences":["A practical pre-test follows from the paper's own Remark 3.21: before committing to a within-date design, compute the overlap matrix from historical positions and estimate $\\|M_\\iota g_t\\|$; the size of that component says whether the design could identify aggregate crowding at all.","The same observational-equivalence failure likely appears beyond trading: any experiment in which a shared accumulated resource—a recommendation model, a matching algorithm, a common supplier—is the treatment of interest and all units receive the same exposure will see the shared effect absorbed by time fixed effects. Testing this in a platform experiment would be a direct transfer of Proposition ","The calibration implies a division of labour in practice: single-strategy shops cannot run the proposed experiment at feasible cost, so the realistic adopters are institutions able to randomise many sleeves simultaneously; for everyone else, the paper's one-sided impact-model bound and the model-free terminal bound are the usable outputs.","Because a finite hold understates steady-state erosion, the paper's numbers imply that observational capacity estimates built from finite-hold return histories should be read as upper bounds on own-sleeve capacity rather than point estimates, a reinterpretation the authors leave implicit."],"forward_implications":["A same-date experiment with homogeneous overlap identifies only the own-sleeve controlled direct effect at the prevailing average deployment, not the effect of scaling the whole strategy.","Estimating aggregate crowding requires either deliberately heterogeneous overlap across sleeves, which restores information through the component $\\|M_\\iota g_t\\|$, or cross-date variation in average deployment, which costs calendar exposure; the paper's calibration puts that cost at a factor of 23.9 for 100 sleeves under staggered assignment.","A finite holding period understates the steady-state erosion effect by a factor $G_L$; the terminal contrast is a model-free lower bound, and deattenuation is reliable only when the accumulation kernel is transported rather than estimated from the experiment's own trajectory.","The sharp identified set for capacity on a finite grid is not a confidence set: a bracket between two estimated arm means covers the true capacity less than half the time, and sequential refinement makes coverage worse, so arm placement and the reporting rule must be fixed separately.","An impact model fitted to execution records bounds capacity from one side only, overstating capacity by an amount that grows with the share of crowding in total erosion, from 10% to 51% in the paper's calibration."],"supporting_citations":[{"why":"Supplies the exposure-mapping device that reduces a full assignment vector to a low-dimensional summary, the basis of Assumption 3.2.","marker":"Hudgens and Halloran (2008)"},{"why":"Provides the general-interference causal framework under which exposure consistency and the overlap design are stated.","marker":"Aronow and Samii (2017)"},{"why":"Gives the switchback experimental design theory that frames the block-randomised, carryover-prone structure used here.","marker":"Bojinov et al. (2023)"},{"why":"Documents the slow decay of market impact that motivates the persistent erosion stock driving the attenuation results.","marker":"Brokmann et al. (2015)"},{"why":"Establishes that impact persists beyond volatility, supporting the accumulation-kernel model of erosion.","marker":"Bucci et al. (2019)"},{"why":"Supplies the concave square-root price-impact law that the paper reconciles with its convex working scale response.","marker":"Tóth et al. (2011)"},{"why":"Provides the g-computation identity used to identify finite-hold capacity curves at the assigned arms.","marker":"Robins (1986)"},{"why":"Supplies the monotone-treatment-response partial-identification framework behind the sharp identified set for capacity.","marker":"Manski (1997)"}],"fun_headline_variants":["Same-date experiments blind to crowding","Robust design hides crowding effect","Crowding erased by date fixed effects","Capacity driver invisible in same-date tests","To see crowding, heterogeneous overlap required"],"cache_read_input_tokens":32896,"weakest_assumption_plain":"The result depends on homogeneous overlap: every sleeve's crowding enters through the same scalar $\\gamma^\\top W_t$, so the aggregate crowding vector lies in the span of the calendar intercept; if overlap is heterogeneous, within-date contrasts can separate crowding from calendar effects.","fun_headline_variants_meta":{"raw":{"variants":["Same-date experiments blind to crowding","Robust design hides crowding effect","Crowding erased by date fixed effects","Capacity driver invisible in same-date tests","To see crowding, heterogeneous overlap required"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000235,"raw_usage":{"total_tokens":1529,"prompt_tokens":1003,"completion_tokens":526,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":619,"completion_tokens_details":{"reasoning_tokens":466}},"tokens_in":619,"tokens_out":526,"duration_ms":5835,"temperature":1.0,"reasoning_tokens":466,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:37:42.399529+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the overlap matrix $\\Gamma$ of a real multi-sleeve book from daily executed positions and form a within-date contrast between two sleeves with materially different $\\Gamma$ rows while holding own-scale fixed: if the contrast shows a component that moves with measured aggregate deployment across dates, the homogeneous-overlap premise behind the non-identification result fails in the field.","supporting_citations":[],"review_version":1}