{"id":"f84eb705-9c56-4deb-a1d1-746cab3f2bde","arxiv_id":"2607.21955","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"CoW Protocol's CIP-74 reward reform concentrated large-order trading value onto incumbent solvers and dispersed small-order value, without changing average execution quality or trade-count concentration.","lead":"A reform to how CoW Protocol pays its solvers reallocated trading value toward large orders and the top incumbent solver, while small-order concentration fell and average execution quality did not visibly change. The paper is the first causal event-study of solver-reward design in intent-based crypto exchanges, with public code and data.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Cross-venue placebo tests only aggregate HHI, not the order-size gradient; a market-wide large-order shock could mimic the central result.","rationale":"The reader's weakest assumption—that no unobserved CoW-specific or coincident shock explains the post-reform pattern—is the right general concern. I sharpen it to a specific, testable implication: the paper's own placebo design is applied to aggregate HHI, not to the order-size gradient that constitutes the central claim. The aggregate HHI break is explicitly carried by one solver and is de-emphasized by the authors; the load-bearing result is the four-bucket monotone gradient. Yet the cross-venue placebo in §6.3 says only that UniswapX had no aggregate break on 8 December, and the market-wide series has a break in May 2025. Neither test addresses whether a market-wide large-order concentration event at the reform date would generate the same order-size gradient on a venue without reward changes. The data to run this test are public (UniswapX filler identities and fill prices are already used in the paper), so the check is feasible and decisive. I also note the monotonicity test's exact permutation p=0.042 is the minimum possible p-value with four buckets (1/24); while not itself fatal, it underscores how coarse the existing test is. The paper's replication package, public on-chain data, pre-specified 100k+ bucket, and explicit limitations are real strengths; the missing control-venue order-size placebo is a concrete, addressable gap rather than a fundamental flaw. Conditional acceptance on this check preserves the paper's contribution while closing the main identification hole.","tokens_in":8318,"tokens_out":5873,"duration_ms":58505,"concrete_test":"Replicate §6.4 on UniswapX: compute daily filler-HHI within the same four USD order-value buckets (0-1k, 1k-10k, 10k-100k, 100k+) around 8 December 2025, estimate the identical interrupted-time-series level break per bucket, and run the same Spearman permutation test for monotonicity. If UniswapX shows the same monotone gradient (rho=1.00) at the CIP-74 date, the CoW-specific causal claim is undermined. If the gradient is flat, non-monotone, or absent at that date, the concern is resolved.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central claim is the order-size gradient: concentration fell in small orders and rose monotonically in large orders after CIP-74. The causal interpretation rests on this gradient being specific to CoW and to the reform date. But the cross-venue placebo in §6.3 validates only aggregate HHI: it shows UniswapX had no sharp aggregate break aligned to 8 December 2025. It does not report whether UniswapX exhibited the same four-bucket order-size gradient. A market-wide event (e.g., a December cryptocurrency price move, a shift in institutional order flow, or a change in large-order composition) could produce exactly the same monotone within-bucket HHI breaks on a control venue that had no reward reform. Because the paper's mechanism test (§6.4) is run only on CoW, the possibility that the gradient is a coincident market phenomenon rather than a CIP-74 effect remains open. This is the load-bearing gap in the causal argument: the outcome dimension that defines the central claim is not placebo-tested on the one available control venue. The paper's Limitation 3 acknowledges only one firm-level control venue, but the more specific gap is that the order-size gradient—the result the paper rests on after leave-one-solver-out—was never computed for UniswapX. Without this check, the timing evidence alone cannot distinguish a CIP-74 effect from a concurrent large-order concentration event that happened to align with the reform date.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper estimates the causal effect of CoW Protocol's CIP-74 solver-reward reform (effective 8 December 2025) on market concentration and execution quality, using 395 days of public on-chain data. It reports a sharp aggregate rise in volume-weighted HHI, a monotone order-size gradient in within-bucket HHI breaks (small orders de-concentrate, large orders concentrate), no increase in trade-count HHI, and no detectable change in average execution quality. The identification strategy combines interrupted time series, changepoint detection, placebo-in-time, a cross-venue placebo (UniswapX), leave-one-solver-out, an exact permutation test on monotonicity, and a triple-difference around a February 2026 fee cut. The paper's stated central claim is that CIP-74 reallocated trading value by order size, with the order-size gradient as the result that survives dropping the largest solver.","tokens_in":8626,"tokens_out":6924,"duration_ms":63001,"significance":"If the causal interpretation holds, this is the first causal evidence on how solver-reward rules shape competitive structure in intent-based exchanges, providing an empirical counterpart to restricted-entry theory. The paper is unusually transparent: it ships a replication package, uses public data, relies on exact permutation inference for the headline monotonicity result, explicitly downgrades anti-conservative HAC p-values, and appends a candid Limitations section. The main weakness is that the causal identification for the central result is incomplete: the cross-venue placebo tests only aggregate concentration, not the order-size gradient that the paper itself designates as load-bearing, and the fixed USD bucket definitions introduce a potential price-composition confound. These gaps are fixable with additional analysis of the existing control-venue data, so the paper merits revision rather than rejection.","major_comments":[{"comment":"The cross-venue placebo is run only on aggregate HHI, whereas the paper's central claim, stated in §1 and §6.4, is the monotone order-size gradient. UniswapX is the only firm-level control venue, yet the paper never computes the same four-bucket within-segment HHI breaks, or a permutation test for a monotone gradient, on UniswapX. A market-wide large-order shock (e.g., a broad price move, a shift in institutional order flow, or a change in large-order composition) could produce the identical within-bucket pattern on a venue that had no reward reform. Because the mechanism test in §6.4 is run only on CoW, the causal claim that the gradient is specific to CIP-74 is not directly placebo-tested. Please either present the UniswapX order-size gradient and show it is flat or non-monotone, or explicitly restrict the causal claim to 'conditional on no concurrent large-order market shock' and assess the resulting threat. Limitation 3 acknowledges only the scarcity of control venues, not this specific missing check.","section":"§6.3, §6.4"},{"comment":"The four order-value buckets are fixed USD intervals, and no adjustment is made for token-price movements or for the resulting reclassification of trades across bucket boundaries over time. The paper reports that the incumbent's mean large-trade size rises from $316k to $435k over the sample, an increase of roughly 37%; if this reflects broad price appreciation rather than a change in order composition, the composition of the 100k+ bucket is not stable. A price-driven reclassification could mechanically change within-bucket HHI even absent any solver response, and the current design does not control for this or placebo-test it on the control venue. Please report bucket shares over time, deflate USD values by the relevant token price, or otherwise demonstrate that the bucket-level HHI breaks are not an artifact of nominal bucket boundaries.","section":"§4, §6.4"},{"comment":"The aggregate timing evidence is presented as 'time-locked' in the abstract, but the design-based evidence is borderline: placebo-in-time p-values are approximately 0.05 for HHI and 0.09 for top-solver share, and the dominant changepoint lands 2-4 days after the effective date. Since the paper itself argues that the HAC p-values are anti-conservative, the aggregate 'time-locked to CIP-74' claim rests on a single marginally significant permutation p-value. This is not fatal because the paper correctly leads with the order-size gradient, but the abstract and conclusion should qualify the aggregate claim as suggestive, and the discussion should address why a 2-4 day lag is consistent with the governance-effective date rather than with a delayed implementation or an unrelated event.","section":"§6.1, §6.2"}],"minor_comments":[{"comment":"The phrase 'reshapeswho captures value' in the abstract and introduction is missing a space; it should read 'reshapes who captures value.'","section":"§1 (Abstract)"},{"comment":"The manuscript retains the 'Conference’17, July 2017, Washington, DC, USA' template header, which should be removed for a journal submission.","section":"Header"},{"comment":"Equation (1) uses the reward cap as if it were the exact reward, but the text describes it as a cap. Please clarify whether the cap is binding in all auctions or only for high-revenue solutions, and how non-binding cases are handled in the model and the empirical predictions.","section":"§2.1, Eq. (1)"},{"comment":"The 10k+ triple-difference result (δ≈-0.103, clustered p≈2e-9) is dismissed because randomization-in-time p≥0.26, while the 100k+ result is reported at clustered p=0.099. The text would benefit from a more explicit statement of why randomization-in-time is the design-matched inference and why the clustered p-value is not used for the headline interpretation.","section":"§6.6"},{"comment":"The phrase 'An adversarial novelty search found no prior causal/event-study analysis' is vague; if such a search was performed, describe its scope, or remove the claim to avoid an unfalsifiable assertion.","section":"§3"}],"recommendation":"major_revision","confidential_remarks":"The paper is competently executed and unusually honest about its limitations, and the central finding is plausible. The main gap is the missing placebo test on the order-size gradient for the control venue, which is the outcome dimension the paper itself identifies as load-bearing. This is fixable with public data and does not require new theory. I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Ruiyang Zhang's paper is one of the first genuinely causal looks at how a reward-rule change in an intent-based DEX reshapes who captures value. The natural experiment is well chosen — a governance-dated change (CIP-74) with a clean effective date, public data, a control venue, and a real mechanism model. The headline monotone order-size gradient (small orders de-concentrate, large orders concentrate) survives dropping the largest solver, which is the right robustness test. The paper's design-based inference — permutation tests, placebo-in-time, changepoint location — is honest, and the authors explicitly downgrade the HAC p-values and disclose the one-solver caveat. That's real discipline.\n\nThe soft spots are the usual ones for observational work. The aggregate HHI break is carried by one solver; the authors know this and pivot to the gradient, which is what should be judged. The second-event triple-difference is a bounded null — directionally consistent, underpowered. Execution-quality null is bounded at ~7 bps. These are correctly framed.\n\nThe stress-test gap is the one thing I'd want addressed before publication: the cross-venue placebo on UniswapX validates only aggregate HHI. It does not show whether UniswapX exhibited the same four-bucket order-size gradient. That matters because the central claim is precisely the gradient. If a market-wide large-order shock coincided with the reform — a December price move, institutional flow, something affecting big orders broadly — both CoW and UniswapX would show the same within-bucket breaks even though only CoW changed its reward rule. The authors rule out a synchronized aggregate break, but not a synchronized order-size gradient. That's a proportionate concern: the timing, changepoint, and leave-one-solver-out evidence make the causal story plausible, but the one control venue available to test the load-bearing outcome dimension wasn't run on that dimension. I'd call this a fixable gap rather than a fatal flaw — run the bucket-level ITS on UniswapX and report it.\n\nCitation pattern looks honest; no self-citation inflation, and the relevant empirical and theoretical work (Bachu, Yuminaga, Chitra et al.) is engaged rather than ignored. Replication package is committed with frozen data, which is more than most.\n\nWho is this for? DeFi market design, protocol governance, and anyone working on auction-based execution venues. It's a solid empirical paper — the kind that deserves refereed journal/conference time rather than a desk reject. My recommendation: send it to peer review, with the UniswapX order-size gradient as a required addition.","headline":"An honest, carefully-identified empirical paper on how solver reward rules reshape value capture in intent-based DEXes; the central monotone order-size gradient is real and robust, but the cross-venue placebo doesn't cover the outcome dimension that carries the claim.","tokens_in":9076,"tokens_out":1951,"would_cite":true,"duration_ms":15873,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A solver-reward reform reallocated CoW Protocol trading value by order size, monotone from small to large orders.","keywords":["intent-based exchanges","solver rewards","market concentration","natural experiment","CoW Protocol","decentralized finance","execution quality","order flow auctions"],"falsifier":"The central claim would be falsified if re-estimating the four bucket-level HHI breaks while excluding the single largest solver and the largest 1% of trades no longer yields a monotone gradient with permutation p < 0.05, or if more than 5% of placebo dates chosen in the pre-reform window produce a similarly perfect monotone gradient.","tokens_in":8040,"feed_emoji":"📊","tokens_out":8462,"duration_ms":69037,"temperature":0.7,"pith_summary":"This paper tries to establish that a governance change in how solvers are paid can causally reallocate who captures trading value on an intent-based exchange. The specific claim is that CoW Protocol's CIP-74, which tied solver rewards to protocol revenue and added an ad-valorem volume fee, made concentration fall among small orders and rise among large orders, monotonically across four order-value buckets. The monotone gradient (Spearman rho = 1.00, exact permutation p = 0.042) survives dropping the largest solver, while trade-count concentration fell and average execution quality did not move by more than about 7 basis points. If this is right, the design of solver rewards is a first-order determinant of market structure in intent-based trading, even when users' average prices are unchanged.","feed_headline":"Reward reform concentrated large orders, dispersed small ones","feed_subtitle":"CoW Protocol's CIP-74 shifted value to inventory-rich solvers without moving users' average execution quality.","key_machinery":"The central object is the payment rule introduced by CIP-74: the solver reward cap becomes proportional to the protocol revenue of the winning solution, plus an unconditional volume fee of about 2 basis points. The model's Proposition 1 states that the ad-valorem fee enters every solver's payoff identically and therefore cannot change who wins; Proposition 2 states that the revenue-linked cap makes the marginal return to winning increase with order value, so solvers whose costs grow slowly with size are most advantaged on large orders. The empirical machinery that carries the argument is a within-bucket interrupted-time-series regression estimating a level break in volume-HHI at the reform date, with the monotonicity of the four bucket breaks tested by exact permutation and re-checked after deleting the largest solver.","core_discovery":"The paper reports that CoW Protocol's CIP-74, effective 8 December 2025, caused a reallocation of trading value by order size: volume-HHI level breaks were -0.086 in the 0-1k bucket, -0.052 in 1k-10k, +0.025 in 10k-100k, and +0.085 in 100k+, a perfect rank correlation between order-size bucket and break (Spearman rho = 1.00; exact permutation p = 0.042) that persists when the single largest solver is excluded. Aggregate volume-weighted HHI rose from 0.176 to 0.241, but the paper shows that aggregate movement is substantially carried by the incumbent top solver and treats the order-size gradient as the load-bearing result. By trade count the market de-concentrated (count-HHI -0.060), and average execution-quality changes are bounded below roughly 7 basis points and statistically indistinguishable from zero. The authors explain the pattern with a model in which the revenue-linked reward cap raises the marginal payoff to winning large orders for inventory-rich solvers, while the ad-valorem fee is competitively neutral.","pith_inferences":["If the revenue-linked cap mechanism generalizes, other intent-based venues that index solver rewards to revenue should show a similar monotone order-size gradient when their reward rules change; this is a testable prediction for future governance events on UniswapX or 1inch Fusion.","The count-volume divergence suggests that concentration metrics based on trade counts alone can miss value-level reallocation, so dashboards and oversight for intent markets may need to track both dimensions.","The bounded mean-level execution-quality null leaves open distributional effects: large-order users could have experienced worse execution even if the average did not move, and a size-bucketed execution-quality decomposition would test that.","The failure of a partial fee rollback to undo concentration hints at path dependence in solver-market structure, which could make later corrective governance, such as consistency rewards, harder to assess without the same event-study machinery."],"forward_implications":["Solver-reward parameters that look like accounting details are first-order determinants of which agents capture intent-market value.","A revenue-linked reward cap can entrench inventory-rich incumbents in the large-order segment without reducing the number of active solvers or raising trade-count concentration.","The cross-venue placebo places the concentration break solely at CoW's reform date, implying the effect is not a synchronized market-wide event.","The bounded execution-quality null implies the reallocation happened without a detectable change in the average price users received, so the immediate user-side cost is small even as market structure shifted.","The directionally consistent but underpowered triple-difference on the February 2026 fee cut leaves open whether a partial rollback can reverse the concentration."],"supporting_citations":[{"why":"Supplies the restricted-entry theoretical prediction that the paper's order-size gradient is offered as the empirical counterpart.","marker":"Chitra et al. 2024"},{"why":"Provides the prior cross-venue measurement of execution welfare that defines the user-side baseline and comparison.","marker":"Yuminaga et al. 2025"},{"why":"Gives the price-improvement formalism used for the execution-quality measurement.","marker":"Bachu, Wan & Moallemi 2024"},{"why":"Supplies the HAC standard-error correction used in the interrupted time-series regressions.","marker":"Newey–West (1987)"},{"why":"Provides the synthetic-control estimator used to bound the execution-quality null.","marker":"Abadie, Diamond & Hainmueller 2010"},{"why":"Provides the difference-in-differences template for the venue-by-post execution-quality comparison.","marker":"Card & Krueger 1994"}],"fun_headline_variants":["CIP-74 reallocates value: small orders disperse, large concentrate","CoW's CIP-74: large orders concentrate, small ones spread","Revenue-linked rewards favor big orders, study finds","Reward reform shifts value: small diversify, big consolidate","Solver reward reform creates order-size concentration gradient"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that, absent CIP-74, CoW Protocol's order-size concentration would have followed its pre-reform flat-to-declining trend and that no other CoW-specific shock, such as anticipatory solver behavior, token-price moves, or order-flow mix changes, coincided with the 8 December 2025 effective date.","fun_headline_variants_meta":{"raw":{"variants":["CIP-74 reallocates value: small orders disperse, large concentrate","CoW's CIP-74: large orders concentrate, small ones spread","Revenue-linked rewards favor big orders, study finds","Reward reform shifts value: small diversify, big consolidate","Solver reward reform creates order-size concentration gradient"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000874,"raw_usage":{"total_tokens":3859,"prompt_tokens":1096,"completion_tokens":2763,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":712,"completion_tokens_details":{"reasoning_tokens":2679}},"tokens_in":712,"tokens_out":2763,"duration_ms":19896,"temperature":1.0,"reasoning_tokens":2679,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:28:46.431774+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"The central claim would be falsified if re-estimating the four bucket-level HHI breaks while excluding the single largest solver and the largest 1% of trades no longer yields a monotone gradient with permutation p < 0.05, or if more than 5% of placebo dates chosen in the pre-reform window produce a similarly perfect monotone gradient.","supporting_citations":[],"review_version":2}