{"id":"cb666a21-e939-424e-a2f3-f5eaa1131063","arxiv_id":"2510.12049","paper_version":6,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"GenAI deployed in seven retail workflows raised sales by up to 16.3% in randomized experiments, with an annualized incremental value near $5 per consumer, driven mainly by higher conversion rates.","lead":"Randomized experiments on a large cross-border retail platform show GenAI tools in customer service, search, and product descriptions increased sales by roughly 2–16% depending on the workflow, with no detectable effect in two advertising applications. The paper is a rare large-scale causal estimate of whether generative AI pays off in real retail operations.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Annualization multipliers in Table 6 rest on an untested assumption about how often consumers encounter each workflow per year; the $5 headline is directly proportional to these multipliers.","rationale":"The reader's weakest assumption is exactly §4.3's annualization, and that is the most load-bearing element of the paper's central claim. The individual treatment effects are generally well-identified through randomization, but the headline $5 per consumer is an extrapolation that depends on unvalidated frequency assumptions. Table 6 makes the sensitivity obvious: the 40.6 multiplier for a 9-day experiment and the 365 multiplier for a 1-day experiment dominate the aggregate. If the actual annual exposure is even a factor of two lower, the aggregate drops to ~$2.5–$3, undermining the 'roughly $5' claim. The paper's own caveats (constant effects, linear additivity) are acknowledged, but the frequency assumption is the least supported and most consequential. A concrete test using platform data can settle it. The reader's CONDITIONAL verdict is appropriate; no new concern changes it, so UNCHANGED.","tokens_in":38614,"tokens_out":4291,"duration_ms":37048,"concrete_test":"Use platform log data to compute the annual number of relevant consumer actions per user for each experimental population: pre-sale chat sessions per user per year, searches per user in the treated languages per year, product-page visits per year, and push-message exposure per year. Replace the time multipliers in Table 6 with the reciprocal of these annual frequencies and recompute the total. If the recomputed total falls below ~$3 (or varies by more than ±50% across plausible frequencies), the headline aggregation should be presented as a scenario analysis rather than a point estimate.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The aggregate claim of $4.6–$5 per consumer in §4.3 is computed by multiplying per-experiment ATEs by time multipliers: 6.0 for Pre-sale Chatbot (2-month experiment), 40.6 for Search Query (9-day experiment), 52.1 for Product Description (1-week experiment), and 365 for Marketing Push (1-day experiment). The footnote states: 'We assume that each experiment’s duration reflects the typical interval between treatment opportunities for a representative consumer.' No evidence is provided for these frequencies. If a representative consumer in the Search Query study searches in Arabic/Japanese/Polish only, say, 10 times per year, the 40.6 multiplier overstates by 4×, cutting that workflow's contribution from $2.63 to $0.65 and the total from $4.96 to ~$2.98. For Pre-sale Chatbot, if consumers contact pre-sale support once a year, the 6× multiplier overstates 6×, dropping that contribution from $1.64 to $0.27. The individual ATEs (e.g., 2.93% for Search Query, 16.3% vs. no-service for Pre-sale Chatbot) may be valid for the experimental window, but the annualization is a separate, load-bearing assumption. Without data on annual exposure frequencies, the headline $5 is unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports seven randomized field experiments conducted on a large cross-border e-commerce platform between September 2023 and June 2024, in which GenAI was introduced into consumer-facing workflows: pre-sale service chatbot, search query refinement, product description generation, marketing push messages, Google advertising title optimization, chargeback defense, and live chat translation. For five workflows the authors have granular transaction data and estimate OLS treatment effects with cohort fixed effects. They find a large positive effect for the pre-sale chatbot (16.3% sales increase), smaller effects for search query refinement (2.93%) and product descriptions (2.05%), a statistically insignificant positive effect for marketing push, and a negative insignificant effect for Google ad titles. The authors interpret the sales gains, with prices and inputs held constant, as TFP improvements, and they aggregate the four positive workflows into an annualized incremental value of approximately $4.6–$5 per consumer. Heterogeneity analyses show larger effects for less experienced consumers and smaller sellers.","tokens_in":38887,"tokens_out":7296,"duration_ms":68161,"significance":"If the individual ATEs are taken at face value, this is one of the largest collections of randomized evidence on GenAI in a live retail environment, covering millions of consumers and products across multiple workflows. The explicit focus on revenue-based outcomes, conversion margins, and demand-side frictions is a valuable complement to the worker-productivity literature. The main quantitative headline, however, is the annualized value, and it rests on assumptions about exposure frequencies, effect persistence, and linear additivity that are not supported by data in the manuscript. The underlying experimental estimates are likely sound; the aggregate claim needs substantial additional justification or careful reframing.","major_comments":[{"comment":"The annualized headline is directly proportional to the time multipliers (6.0, 40.6, 52.1, 365.0) in Table 6, and the only justification is the footnote that experiment duration equals the typical interval between treatment opportunities. No frequency data support this. The Search Query experiment covered Arabic/Japanese/Polish searches over nine days; multiplying by 40.6 assumes a representative consumer searches in one of these languages every nine days year-round. If the true frequency is 10 times per year, that workflow's contribution drops from $2.63 to ~$0.65 and the total from $4.96 to ~$2.98. Similarly, if pre-sale chatbot contact occurs once per year, its $1.64 contribution falls to $0.27. The calculation also assumes constant effects and linear additivity. Please provide platform exposure frequencies, or present the annualized numbers only as an illustrative sensitivity analysi","section":"§4.3, Table 6"},{"comment":"Table 6 includes Marketing Push Message with annualized contribution $0.15, but the underlying sales ATE (0.000402, SE 0.000812 in Table 4) is not statistically significant. The abstract's 'four GenAI applications with positive sales effects' therefore overstates the evidence: only three sales effects are statistically distinguishable from zero among the detailed-data workflows. Excluding the push contribution changes the aggregate to $4.81/$4.48, which is still close to '$5,' but the inclusion should be disclosed and the headline should not be stated as if based on four significant effects. The Google Advertising Title null/negative effect is excluded; since it is insignificant this is defensible, but the exclusion rule should be stated as 'statistically significant positive effects' or 'point estimates' consistently.","section":"§4.3, Table 6 vs Table 4"},{"comment":"For Chargeback Defense and Live Chat Translation, the paper relies on platform-internal estimates ('15% defense success rate increase', '5.2% consumer satisfaction increase') because raw data were not provided. The text acknowledges the chargeback estimate 'couldn't be verified.' These are not randomized experimental estimates in the same sense as the other five workflows, yet the abstract and introduction count them among 'seven workflows' and describe the paper as providing causal evidence across workflows. Please either obtain and analyze the underlying data, or clearly label these as platform-reported descriptive metrics and restrict causal claims to the five workflows with granular data.","section":"§3.4, §4.1, Appendix C"}],"minor_comments":[{"comment":"The phrase 'total factor productivity improvements' overstates the mapping. Holding inputs and prices fixed, a demand-side sales increase raises revenue per input, but this is not necessarily a technical efficiency change; 'revenue-based productivity' is the more precise term used elsewhere. Consider softening the TFP language.","section":"§3.2, Abstract"},{"comment":"The comparison of $4.6–$5 to '5.5–6% of per-user revenue growth' cites Statista but does not define the comparison universe or period precisely; add details.","section":"§4.3"},{"comment":"Table 4 footnote says standard errors are in brackets but displays parentheses; also Table C5 reports SE 0.000816 for the Marketing Push sales coefficient while Table 4 reports 0.000812. Harmonize.","section":"Table 4"},{"comment":"Typos: 'No Serice' in Table C1 note; 'cross-boarder' in Table 1; 'The it was conducted' in §5.2; 'Live Chat T ranslation' in a section heading. Also define the 'lower-bound' value in Table 6; it is from Table C1 but not identified in the table.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper's main experimental results are likely publishable; the annualization is the sole major obstacle. If the authors provide exposure-frequency data or reframe the annualized value as an illustrative calculation with sensitivity analysis, and clearly separate platform-reported metrics from the randomized estimates, I would support publication. The current abstract overclaims relative to what is verified."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take: the individual experiments are mostly well-run and the multi-workflow setup is genuinely valuable, but you should not trust the $5-per-consumer headline until the authors fix the annualization and clear up a direct contradiction between the abstract and the main text.\n\nThe real contribution is the size and scope: seven randomized evaluations in one large cross-border retailer, millions of consumers, and actual downstream sales and conversion rather than worker-task proxies. The individual ATEs look credible — covariate balance, cohort fixed effects, large samples. The chatbot experiment is the most careful: it includes comparisons against no-service, human-only, and hybrid conditions, and the hybrid-vs-human 11.5% effect is a plausible lower bound. I also give them credit for reporting a clear null for Google ad titles, and for the heterogeneity story: consistent gains for smaller sellers and less experienced buyers. That pattern has practical value.\n\nThe aggregate claim is the weak part. Table 6 multiplies each per-experiment ATE by 12/duration, and the footnote assumption — experiment duration equals the typical interval between treatment opportunities — is untested and likely wrong. The search experiment ran nine days; the multiplier is 40.6. If a consumer searches in those languages ten times a year instead of 40, that workflow's contribution drops from $2.63 to $0.65. The chatbot multiplier assumes six contacts per year; once a year would cut its contribution to a quarter. So the $4.96 total is about as solid as those multipliers. They also include the non-significant Marketing Push sales effect and exclude the negative Google Ad effect, with no confidence interval on the total. The right move is to present the aggregate as a scenario or bounding exercise, not as a point estimate.\n\nThere's also an internal inconsistency: the abstract says return rates and customer ratings do not deteriorate, but Section 6 says the authors lack data on product returns and retention. That has to be fixed before publication. Smaller issues: two workflows rely on the platform's internal metrics, which the authors couldn't verify; that's fine as secondary evidence, but they are treated as comparable to the randomized results. And the TFP mapping is an interpretation of the constant-input assumption, not something separately measured. None of this kills the core result — the individual ATEs are the contribution — but the headline needs to be rewritten.\n\nThis paper is for economists and applied researchers working on AI and productivity, plus retail practitioners. It deserves a serious referee. I'd send it to review, but I'd expect major revisions on the aggregate claims and the abstract.","headline":"Individual experiments are solid, but mistrust the $5 per consumer headline until annualization and abstract overreach are fixed.","tokens_in":39395,"tokens_out":2767,"would_cite":true,"duration_ms":26734,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Large-scale randomized retail experiments show GenAI raises sales in most workflows — up to 16.3% — worth roughly $5 per consumer per year.","keywords":["generative AI","field experiments","sales productivity","total factor productivity","online retail","conversion rate","friction reduction","treatment effect heterogeneity"],"falsifier":"Run the same workflows for a full year and compare treatment effects in the first versus later months: a decay toward zero falsifies the constancy assumption. Or replace the assumed time multipliers (6, 40.6, 52.1, 365) with the platform's actual average workflow encounters per consumer per year — if real encounter rates are far lower, the annualized $5 value collapses. A third check measures overlap of the treated consumer populations across the four positive workflows; substantial overlap would violate linear additivity.","tokens_in":38432,"feed_emoji":"📈","tokens_out":8345,"duration_ms":68921,"temperature":0.7,"pith_summary":"This paper tries to establish, with direct causal evidence, that firm-level adoption of generative AI raises revenue-based productivity in real online retail. It reports randomized field experiments across seven consumer-facing workflows at one large cross-border platform, involving millions of consumers and products. The central result: GenAI increased sales in most workflows, from no detectable effect to 16.3%, and the four workflows with positive effects imply an annual incremental value of roughly $4.6–5 per consumer. Gains show up as higher conversion rates rather than larger baskets, which the authors read as evidence that GenAI reduces search, information, communication, and personalization frictions. This matters because it is large-scale causal evidence on whether GenAI investment pays off in actual purchase behavior, not just task-level efficiency.","feed_headline":"GenAI lifts retail sales by up to 16.3% in field tests","feed_subtitle":"Seven randomized retail workflows show gains come from higher conversion; novices and small sellers gain most.","key_machinery":"The load-bearing design is a set of seven parallel randomized field experiments, each comparing a GenAI-integrated workflow against the platform's standard practice with prices and labor and capital inputs identical across arms; randomization at the consumer level (product level in one workflow) makes each sales difference an unbiased average treatment effect. The paper interprets results through standard Solow growth accounting: with capital and labor fixed, any measured output increase is attributed to total factor productivity. The mechanism probe is the conversion rate measured alongside sales — splitting the extensive margin (more consumers buying) from the intensive margin (cart value","core_discovery":"GenAI adoption increases sales in most of the seven workflows tested, with estimated treatment effects ranging from no detectable impact to a 16.3% sales lift for a pre-sale service chatbot; the four positive workflows together imply roughly $4.6–5 of annual incremental value per consumer, about 5.5–6% of global per-user e-commerce revenue growth in 2023–2024. Because prices and labor and capital inputs were held constant across experimental arms, the paper maps these output gains directly into total factor productivity growth. The mechanism operates on the extensive margin: conversion rates rose by 1–22% across workflows while average cart values did not change, and product return rates and","pith_inferences":["If friction reduction is the general mechanism, the same experimental design should find larger GenAI lifts where baseline frictions are worst — no service, missing descriptions, underserved languages — and the paper's untested post-purchase workflows (returns, logistics, payments) are the natural next place to look.","The $5 annual figure is a point estimate on a wide band: it likely overstates steady-state value if novelty or learning effects decay, and understates it if habit formation or cross-workflow synergies compound; the paper's own annualization assumptions make the long-run number genuinely open.","In competitive equilibrium, if rival platforms deploy the same tools, this platform's relative advantage should erode and the durable surplus would shift toward consumers in the form of better matches or lower prices — so the welfare reading of the $5 differs from the revenue reading.","The heterogeneity result yields a sharp testable prediction: the same 'less experienced gain more' pattern should appear in other marketplaces and consumer settings, and average effect sizes should shrink as the user population becomes more familiar with AI-assisted shopping."],"forward_implications":["If the effects are linearly additive as the paper assumes, every additional GenAI-deployed workflow with positive lift adds its own per-consumer value; the platform's expansion from 7 to over 60 workflows by 2025 implies aggregate gains well beyond the $5 estimate.","The conversion-not-basket result means GenAI expands the market at the extensive margin — converting hesitant, marginal consumers — rather than extracting more spending from existing buyers.","Disproportionate gains for small sellers and novice consumers imply GenAI adoption shifts some platform surplus toward the long tail and less experienced participants, narrowing outcome gaps.","The null and negative advertising results imply domain-specific fine-tuning is a prerequisite for positive returns; untuned foundation models can be value-neutral or value-destroying.","Because measured gains arise with constant inputs and unchanged prices, the paper treats the revenue lift as a conservative floor on GenAI returns that omits any future labor-cost savings."],"fun_headline_variants":["GenAI lifts sales 16.3% max in retail field tests","Retail GenAI gains via conversion, not cart size","GenAI adds $5 per shopper yearly in online retail trials","Novice shoppers see biggest GenAI sales boosts in retail","GenAI boosts conversions without raising return rates"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The $4.6–$5 per-consumer annual value rests on Section 4.3's assumptions that each short experiment's sales lift repeats every time a consumer meets that workflow all year and that the four workflows add up without overlap; if novelty fades or workflows cannibalize each other, the headline annual figure fails even though each individual experiment's effect can still be valid.","fun_headline_variants_meta":{"raw":{"variants":["GenAI lifts sales 16.3% max in retail field tests","Retail GenAI gains via conversion, not cart size","GenAI adds $5 per shopper yearly in online retail trials","Novice shoppers see biggest GenAI sales boosts in retail","GenAI boosts conversions without raising return rates"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000356,"raw_usage":{"total_tokens":1778,"prompt_tokens":762,"completion_tokens":1016,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":506,"completion_tokens_details":{"reasoning_tokens":933}},"tokens_in":506,"tokens_out":1016,"duration_ms":8936,"temperature":1.0,"reasoning_tokens":933,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T09:59:55.455315+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same workflows for a full year and compare treatment effects in the first versus later months: a decay toward zero falsifies the constancy assumption. Or replace the assumed time multipliers (6, 40.6, 52.1, 365) with the platform's actual average workflow encounters per consumer per year — if real encounter rates are far lower, the annualized $5 value collapses. A third check measures overlap of the treated consumer populations across the four positive workflows; substantial overlap would violate linear additivity.","supporting_citations":[],"review_version":1}