{"id":"eae8673e-534a-44f3-ac30-d0e07c2c5403","arxiv_id":"2607.24389","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Partial random disruption of delivery lowers experimental market efficiency by roughly 70% and raises concentration; attacks on numerical seller ratings do not.","lead":"A lab-style online market experiment finds that randomly blocking delivery of goods cuts market efficiency sharply and hurts sellers, while faking seller ratings does almost nothing. The result is aimed at policies that would deliberately degrade illicit online markets such as cybercrime forums.","discovery_kind":"new_application","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The headline efficiency result rests on a completer-only sample: 49.6% of groups were dropped, and the paper never reports whether attrition is balanced across treatments. Differential dropout in the low-earning delivery cells could manufacture part of the measured effect.","rationale":"The reader correctly identified external validity (MTurk lab → real cybercrime markets) as the load-bearing premise for the *policy* claim, and flagged completer-only sampling only as a secondary soundness note. I agree the lab-to-field leap limits the paper's policy framing, but the reader's verdict already prices that in via CONDITIONAL. The concern I elevate is internal and prior: if the completer-only filter is treatment-correlated, even the in-lab causal claim (\"delivery attacks reduce efficiency\") is not cleanly identified, because the estimand is a selected-subpopulation effect, not the treatment effect on formed markets. This is distinguishable from a generic small-sample worry: the 50% exclusion rate is unusually large, the attrition mechanism (timeout-forfeiture under falling earnings) is structurally tied to the treatment, and the paper provides no evidence against differential selection. I do not claim the result is wrong — the parametric robustness (Table A.3), the time trends, and the complete-vs-incomplete attack analysis (Table A.4) are genuinely supportive, and the rating-treatment null has a point estimate in the wrong direction for a power artifact. But the attrition audit is cheap, decisive, and missing, so the reader's CONDITIONAL verdict is right and should explicitly condition on this check alongside the external-validity caveat. Hence verdict UNCHANGED, agreement partial.","tokens_in":23990,"tokens_out":2159,"duration_ms":85743,"concrete_test":"Tabulate participant- and group-level completion rates by treatment cell and test for balance (Fisher exact / chi-square across the four cells). Then re-run the headline MW tests from §5.1 two ways: (a) including incomplete groups with per-round available data, and (b) with inverse-probability-of-completion weights. If completion is balanced and the D/C efficiency coefficients stay significant and within ~15% of reported magnitudes, the selection concern does not land; if D/C attrition is elevated and the effect attenuates materially under (a)/(b), the headline claim weakens.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim — 70%/76% lower efficiency in Delivery/Combined, losses borne by sellers — is estimated on \"14 complete groups for each treatment,\" with the sample \"restrict[ed]... to complete groups for which all participants completed the market experiment and the bonus tasks\" (§B.2.1). Only 49.6% of groups were complete, so half of all formed markets are discarded. This filter is plausibly endogenous to treatment: participants must click a progress bar every 30 seconds and are forfeited after three consecutive timeouts, and sellers in Delivery/Combined earn 57–63% less and watch their goods fail to arrive. Frustrated, low-earning participants have the weakest incentive to keep clicking, so attrition should be highest exactly where the treatment bites. If the surviving D/C groups are the unusually patient/cooperative ones (or groups where frustrated buyers dropped before they could express inactivity), the MW comparison of group averages (n=14 per cell) is no longer an intention-to-treat estimate and the effect size is biased — direction ambiguous, but with n=14 even a few selected groups move the medians. The paper reports balance tests on demographics (Table B.7) but only for the completer sample, which cannot detect differential selection; it reports no completion-rate-by-treatment table and no analysis including incomplete groups. The adjusted-efficiency robustness (§A.4) addresses the mechanical seizure effect, not this selection margin. Because the exclusion rule removes half the data on a dimension correlated with the treatment itself, this — more than the well-flagged external-validity gap — is where the internal claim is least secure.","agreement_with_reader":"partial"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The paper studies how to make a market less efficient, motivated by cybercrime and other illicit markets. In a web-based experiment (oTree, MTurk), groups of seven (3 sellers, 4 buyers) trade for 20 rounds under asymmetric information about quality, with a 2×2 between-subjects design crossing a delivery attack (20% of purchases undelivered, buyers not told why) and a rating attack (20% of ratings replaced by a random rating). The main findings, estimated on the last 10 rounds using group-level Mann-Whitney tests: the delivery and combined treatments reduce market efficiency by 70% and 76% relative to baseline, with the loss borne by sellers (earnings 63%/57% lower), an ~18% reduction in goods sold, higher buyer inactivity, and increased market concentration (dominant seller's share rising to 59%/56% vs 41% baseline) driven by asymmetric reputational exposure of small vs. large sellers. The rating attack alone has no significant efficiency effect, which the authors attribute to personal trading relationships substituting for public reputation. Results are supported by random-effects robustness checks distinguishing short- and long-run effects.","tokens_in":24263,"tokens_out":2305,"duration_ms":88221,"significance":"If the result holds, the paper opens a genuinely new line of work: applying experimental market-design methods in reverse, to measure how markets can be made less efficient, with direct relevance to cybercrime enforcement policy where takedowns have repeatedly failed. The paper's strengths are real: a pre-structured 2×2 factorial design, conservative group-level non-parametric inference that correctly treats the market (not the individual) as the unit of independent variation, a transparent surplus-based efficiency measure with an explicit adjustment for expected seizures (Table A.2), robustness via random-effects models separating short- and long-run effects, a mechanistic analysis of the concentration result (complete vs. incomplete delivery attacks), and full instructions and design parameters in the appendices, which makes the experiment replicable. The contrast between the effective delivery attack and the ineffective rating attack — with a plausible mechanism in personal trading relationships — is informative for both theory and policy. However, the policy relevance depends on an external-validity bridge from MTurk to illicit markets that the paper asserts rather than tests,½","major_comments":[{"comment":"The headline estimates (70%/76% lower efficiency; seller earnings 63%/57% lower) are computed on '14 complete groups for each treatment,' where completeness requires all seven participants to finish the market experiment and both bonus tasks (§B.2.1). Only 49.6% of groups are complete, so half of all formed markets are discarded. The exclusion rule is plausibly endogenous to treatment: participants must click a progress bar every 30 seconds on every wait page and are forfeited after three consecutive timeouts, while sellers in Delivery/Combined earn 57–63% less and watch their goods fail to arrive — precisely the participants with the weakest incentive to remain attentive. The balance tests in Table B.7 are computed only on the completer sample and therefore cannot detect differential selection. With n=14 groups per cell, a few selected groups can move the medians on which the MW tests r","section":"§B.2.1, §5.1"},{"comment":"The manuscript reports a large number of Mann-Whitney and Wilcoxon tests at α = 0.05 (Table A.1 alone contains dozens of starred comparisons across eight outcome variables and three round windows) with no multiple-testing correction and no pre-analysis plan referenced. Compounding this, the main-text decision to restrict treatment comparisons to the final 10 rounds 'owing to significant time trends' is a researcher degree of freedom; it is reassuring that Table A.1 shows similar effects over rounds 1–20 and that the parametric specification (Eq. 1) uses all rounds, but the headline percentages in the abstract and §5.1 are tied to the chosen window. The authors should either apply a correction (or at least state the number of hypotheses tested per family), and state whether the last-10-rounds window and the completer-only rule were pre-specified before data inspection.","section":"§5, Table A.1, §A.1"},{"comment":"The abstract and §6 move from the lab result to policy for cybercrime markets ('paves the way for evidence-based... policies to disrupt cybercrime and other illicit markets'), but the external-validity step is asserted rather than examined. The experimental market differs from darknet markets on stakes, anonymity infrastructure, multi-homing, violence/exit options, and the fact that a 20% seizure rate is imposed by the experimenter rather than achievable by law enforcement. This does not undermine the internal experimental result, but the policy claim as stated does not follow from it. The discussion should delineate which features of the setting drive the result (reputational spillover from non-delivery, small group size, fixed matching) and state explicitly what would need to hold in the field for the finding to transfer.","section":"§1, §6"}],"minor_comments":[{"comment":"The '70% and 76% lower efficiency' figures are relative reductions from a baseline efficiency of only ~29% (Table A.1); the absolute decline is roughly 20–22 percentage points. Stating both the relative and absolute magnitudes would prevent over-reading, especially since the abstract's framing invites the relative reading.","section":"§5.1, Abstract"},{"comment":"The treatment column labels are wrong: rows are labelled 'Delivery (B)', 'Rating (B)' etc., apparently copy-paste errors from the Baseline row. Also the checkmark/cross pattern should be double-checked against the text.","section":"Table B.6"},{"comment":"§A.5: the claim that firm size is not the driver is supported by the split in col. 4, but the big-seller complete-attack coefficient (2.228, SE 0.723) is estimated on what must be a small number of events (4% probability per round per big seller); report the number of complete-attack events by seller size so readers can judge the precision.","section":"§A.5, Table A.4"},{"comment":"The notation for the latent variable is inconsistent (f* in Eq. 2 vs. α_i vs. u_i for the random effect in Eq. 3; D vs. Z vs. X for the regressors). Also 'rating attach' should read 'rating attack'.","section":"Eqs. (2)–(3), §A.1"},{"comment":"Buyer earnings are described as 'unchanged across treatments,' but Table A.1 shows baseline buyer earnings falling from −6.6 to −11.7 while Combined shows a significant time trend (W, p<0.05); a sentence clarifying that the cross-treatment contrast (not the trend) is what is null would help.","section":"§5.1"},{"comment":"Consider reporting the Herfindahl-Hirschman results alongside the dominant-seller share in the main text rather than only in the appendix, since the concentration claim is one of the paper's three headline findings.","section":"§5.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript sits at the intersection of experimental economics and criminology; fit with a general economics journal is reasonable given the market-design framing, but the editor should be aware the policy motivation is cybercrime enforcement and the external-validity step from MTurk to darknet markets is asserted rather than argued. The novelty claim (\"first empirical study of how to disrupt a market using an online experiment\") appears accurate to my knowledge. The single fix that would most change my recommendation is the attrition analysis; the data to produce it already exist."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The one thing worth knowing: this is a competent 2×2 MTurk market experiment that flips the usual market-design objective and shows a partial delivery disruption cuts efficiency ~70% (losses on sellers, fewer trades, rising buyer inactivity) while a rating attack does nothing. Concentration rises because complete delivery hits wipe small sellers and leave a dominant one. That causal contrast is new relative to the conceptual Slander/Sybil notes in CS/criminology and to the eBay reputation literature.\n\nWhat they do well is straightforward. Between-subjects full factorial, fixed 3-seller/4-buyer groups, group-level Mann-Whitney and Wilcoxon that respect dependence, plus random-effects short-run/long-run specs that recover the same delivery effects. Adjusted efficiency nets out mechanical seizures so the volume channel is not just accounting. The complete-vs-incomplete attack split on failure-to-sell is a clean mechanism check for concentration. Citations are appropriate; no circularity in the efficiency definition.\n\nSoft spots in proportion. The stress-test on completer-only sampling is real: they keep only groups where everyone finished (49.6% of formed groups), report no completion rates by treatment, and balance only on the survivor sample. Sellers in delivery cells earn far less and face failed deliveries; if frustrated participants time out more there, the n=14 cell medians are selected. Direction of bias is ambiguous, but with small cells it matters and should have been tabled or bounded. External validity to real cybercrime (stakes, multi-homing, violence, anonymity tech) is stated as inspiration, not tested — fair for a first lab paper, load-bearing if you want policy instruments. The 20% attack rate and last-10-round window are free parameters; results are not wildly fragile in the appendix, but they are design choices.\n\nThis is for experimental market-design people and anyone working on illicit-market disruption who wants a causal benchmark rather than another takedown anecdote. The lab claim is internally credible enough to deserve referee time; the policy framing needs tightening and the attrition margin needs reporting. I would send it to peer review, not desk-reject. Engage if you care about reputation mechanisms or cybercrime enforcement design; skip if you only want field-identified policy effects.","headline":"Clean lab evidence that delivery noise tanks efficiency and concentrates the market; rating noise does nothing — solid internal design, real attrition and external-validity soft spots.","tokens_in":24256,"tokens_out":563,"would_cite":true,"duration_ms":24635,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Partial delivery disruption cuts market efficiency by about 70 percent and hits sellers hardest; rating attacks do almost nothing.","keywords":["market disruption","illicit markets","cybercrime","asymmetric information","reputation systems","delivery attack","market efficiency","experimental economics"],"falsifier":"A field test that seizes or degrades a measurable fraction of deliveries on a real illicit marketplace and finds no lasting drop in transaction volume or seller earnings relative to an untreated control market would falsify the central policy claim.","tokens_in":24029,"feed_emoji":"📉","tokens_out":847,"duration_ms":16319,"temperature":0.7,"pith_summary":"Most market-design work tries to make trade more efficient. This paper asks the reverse: which interventions make a market less efficient, with an eye toward reducing the social harm of illicit online markets such as cybercrime forums. In a controlled web experiment with fixed groups of three sellers and four buyers, a 20 percent chance that a purchased good simply never arrives lowers overall efficiency by roughly 70 percent relative to an undisturbed baseline, mainly because fewer goods are sold and sellers earn far less. Randomly corrupting numerical seller ratings has no comparable effect. The same delivery shock also concentrates sales in the hands of a dominant seller, because larger sellers are less likely to have every sale fail. The authors present the result as causal evidence that delivery-side interference can shrink market activity and as a template for later field tests against real illicit markets.","feed_headline":"Delivery shocks cut market efficiency ~70%; rating attacks fail","feed_subtitle":"Sellers lose volume and earnings; a dominant seller emerges. Lab evidence aimed at illicit markets.","key_machinery":"The delivery attack: after each purchase there is an independent 20 percent probability that the buyer receives nothing, and the buyer is not told whether the failure was caused by the attack or by the seller. The resulting private rating damage, combined with the asymmetric probability that multi-unit sellers lose every sale, is what drives both the efficiency drop and the rise in market concentration.","core_discovery":"A partial disruption to delivery is an effective way to decrease market efficiency. In the final ten rounds, market efficiency in the delivery and combined treatments is respectively 70 percent and 76 percent lower than baseline; the loss is borne by sellers, whose earnings fall 63 percent and 57 percent, while the number of goods sold falls about 18 percent. Attacks that randomly replace buyer ratings leave efficiency essentially unchanged.","pith_inferences":["If personal trading relationships already substitute for public ratings, any reputation attack that leaves those private histories intact will remain weak; interventions that also scramble identity or break repeat matching may be needed.","The concentration side-effect suggests a two-stage enforcement logic: first thin the market with delivery shocks, then focus scarce investigative resources on the remaining large seller.","The 20 percent seizure rate is a design choice; mapping the dose-response curve (how efficiency and concentration change with seizure probability) would be a direct next experiment."],"forward_implications":["Delivery-side interventions can be used as a policy lever to shrink the gains from trade in markets one wishes to disrupt.","Rating-noise or fake-review campaigns alone are unlikely to reduce market efficiency when buyers can rely on personal trading histories.","The same delivery shock tends to create a dominant seller, giving enforcement a more visible target even as overall volume falls.","Combined delivery-plus-rating attacks add little beyond delivery attacks alone under the conditions tested."],"fun_headline_variants":["Delivery disruption slashes market efficiency 70%; sellers bear the loss","Partial delivery shocks cut efficiency; rating attacks leave it intact","Delivery hits reduce goods sold and seller earnings; concentration rises","Rating sabotage fails to dent efficiency; delivery friction succeeds","Lab test: delivery disruption lowers efficiency 70% vs baseline"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That behavior in a short, low-stakes online experiment with legal goods and fixed small groups tells us how real cybercrime markets, with high stakes, anonymity tools, violence, and outside options, would respond to the same delivery shock.","fun_headline_variants_meta":{"raw":{"variants":["Delivery disruption slashes market efficiency 70%; sellers bear the loss","Partial delivery shocks cut efficiency; rating attacks leave it intact","Delivery hits reduce goods sold and seller earnings; concentration rises","Rating sabotage fails to dent efficiency; delivery friction succeeds","Lab test: delivery disruption lowers efficiency 70% vs baseline"]},"model":"grok-4.5","effort":"low","cost_usd":0.001937,"raw_usage":{"total_tokens":850,"prompt_tokens":682,"num_sources_used":0,"completion_tokens":66,"cost_in_usd_ticks":19368000,"prompt_tokens_details":{"text_tokens":682,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":102,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":682,"tokens_out":66,"duration_ms":3516,"temperature":1.0,"reasoning_tokens":102,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T16:06:36.822040+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A field test that seizes or degrades a measurable fraction of deliveries on a real illicit marketplace and finds no lasting drop in transaction volume or seller earnings relative to an untreated control market would falsify the central policy claim.","supporting_citations":[],"review_version":1}