{"id":"9c4c2cc0-7ec5-4814-978b-d8aebfb119fb","arxiv_id":"2505.23809","paper_version":2,"verdict":"REJECT","confidence":"LOW","novelty_score":2.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"An LLM copywriting pipeline combining fine-tuning, vector search, and weighted reranking reportedly lifts CTR by 12.5% and CVR by 8.3%, but the evidence is unverifiable and internally inconsistent.","lead":"This paper describes an LLM-based system for writing e-commerce product copy and claims it raises click-through and conversion rates while keeping the text varied. The reported gains are large, but the supporting experiment is only sketched and the numbers in the abstract and the conclusion do not match.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline CTR/CVR lifts are not derivable from the paper's own Table 2; the central empirical claim is internally inconsistent.","rationale":"The reader's REJECT verdict is sound, and my stress test reinforces it, but I locate the load-bearing failure differently. The reader's weakest assumption concerns whether the randomized seven-day A/B test was actually executed with valid statistics. That is important but depends on external evidence we do not have. My concern is stronger because it is checkable from the manuscript alone: the aggregate effect sizes in the abstract and conclusion contradict the paper's own Table 2. Since the central claim is explicitly quantitative, internal arithmetic consistency is a necessary condition for accepting it. Table 2 reports per-category LLM-vs-baseline CTR and CVR at λ=0.6; every category shows a relative lift of roughly 36–40%. No weighting of the three categories can produce an aggregate lift of 10.4% or 12.5% (CTR) or 4.1% or 8.3% (CVR). The §5.2 sentence about FMCG conversion improving 'by over 35%' confirms that the authors were thinking in relative terms, making the inconsistency more direct. Unless the authors can supply raw session counts and per-arm metrics showing that Table 2 is not the source of the aggregate numbers, the central claim fails as stated. This is not a disagreement with consensus; it is an internal inconsistency. I also note the positive side: the framework itself is coherent at a high level and the creativity-conversion trade-off curve in §5.1 is a reasonable modeling exercise. But the headline empirical validation cannot be accepted in its present form. Therefore the reader's REJECT verdict stands unchanged, and the concrete ask is for the underlying A/B logs or a corrected aggregation that reconciles Table 2 with the reported lifts.","tokens_in":10277,"tokens_out":4234,"duration_ms":40019,"concrete_test":"Recompute the aggregate CTR and CVR lifts implied by Table 2 by weighting the FMCG, Apparel, and Electronics rows by any positive traffic allocation; because every category's relative lift is roughly 36–40%, no convex combination yields the abstract's +12.5%/+8.3% or the conclusion's +10.4%/+4.1%. The decisive check is to obtain the seven-day session counts and per-arm CTR/CVR from Section 4.2 and verify which, if either, of the two reported aggregate effect sizes matches the logged data; if the Table 2 rows are the actual per-category results, the headline percentages cannot be reproduced.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is empirical: the framework raises CTR and CVR over a human-reviewed rule-based baseline. For that claim to hold, the reported effect sizes must be reproducible from the reported experiment. They are not. The abstract reports +12.5% CTR / +8.3% CVR; the conclusion reports +10.4% CTR / +4.1% CVR at λ=0.6. Section 5.2's Table 2 gives the λ=0.6 category-level results. Computing relative lifts from that table: FMCG CTR (12.1−8.9)/8.9=36.0%, Apparel (9.7−7.1)/7.1=36.6%, Electronics (8.5−6.2)/6.2=37.1%; CVR lifts are 36.8%, 37.9%, and 40.0% respectively. Any weighted average over the three categories with positive traffic weights lies between these values, so the aggregate relative lift cannot be 10.4% or 12.5% for CTR, nor 4.1% or 8.3% for CVR. Interpreting the conclusion's numbers as percentage-point differences also fails: Table 2 differences are about 2–3 pp CTR and 1–1.4 pp CVR. The paper even states in §5.2 that FMCG conversion improved 'by over 35%', matching Table 2 but contradicting the +4.1% overall figure. Thus the headline numbers and the table cannot both describe the same experiment; one of them is wrong. This failure is internal and does not depend on whether the seven-day A/B test happened.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes an LLM-driven framework for e-commerce marketing copy generation that combines prompt engineering, multi-objective fine-tuning, vector retrieval, and post-processing. The creativity–conversion trade-off is controlled by a weighted reward R = λD + (1−λ)P_conv (Eq. 3), where D is an embedding-based diversity score and P_conv is a logistic regression conversion probability. The authors report offline evaluations and a seven-day online A/B test across FMCG, apparel, and electronics categories, claiming substantial CTR and CVR lifts over a human-reviewed rule-based baseline (12.5% and 8.3% in the abstract; 10.4% and 4.1% in the conclusion; category-level lifts in Table 2). They also recommend category-specific λ ranges. The central claim is the empirical improvement in CTR and CVR while maintaining diversity.","tokens_in":10616,"tokens_out":5656,"duration_ms":53677,"significance":"If the headline results were reproducible, the framework would be a useful practical contribution to automated e-commerce copywriting, and the category-specific λ guidance could help practitioners calibrate creativity against conversion. The reward formulation in Eqs. (1)–(3) is clear and easy to implement, and the paper is commendable for stating business guidelines and a review pipeline as part of the system. However, the paper provides no code, data, or machine-checked artifacts, and the empirical claim—the only genuinely novel part—is neither internally consistent nor statistically documented. The framework's components are standard (logistic regression, cosine diversity, weighted reranking), so the paper's value would hinge entirely on credible A/B evidence, which is absent.","major_comments":[{"comment":"The headline lift figures are internally inconsistent with the paper's own category-level results. At λ=0.6, Table 2 gives FMCG CTR of 12.1% vs. 8.9%, apparel 9.7% vs. 7.1%, and electronics 8.5% vs. 6.2%, so the relative CTR lifts are 36.0%, 36.6%, and 37.1%; the corresponding CVR lifts are 36.8%, 37.9%, and 40.0%. Any positive traffic-weighted average of these category lifts therefore lies between roughly 36% and 40% for both metrics. Consequently the abstract's '+12.5% CTR and +8.3% CVR' and the conclusion's 'CTR +10.4%, CVR +4.1%' cannot describe the same experiment; interpreting the conclusion's numbers as percentage-point differences also fails because the Table 2 differences are 2.3–3.2 pp for CTR and 1.0–1.4 pp for CVR. The text in §5.2 that FMCG conversion improved 'by over 35%' matches Table 2 but contradicts both headline pairs. This contradiction is load-bearing because the paper's central claim is precisely these quantitative improvements, and the reported evidence does not support any single version of the claim.","section":"Abstract; §5.2 Table 2; §7 Conclusion"},{"comment":"The statistical basis of the online A/B test is not reported. The manuscript claims that 'Z-tests and chi-square tests assess significance (p < 0.05), confirming valid performance lifts,' but it provides no sample sizes, traffic counts, test statistics, p-values, confidence intervals, or timestamps for the seven-day randomized traffic split. Without these, the reader cannot verify the randomization, the adequacy of the test power, or the claimed significance, and the central empirical claim is not inspectable.","section":"§4.2"},{"comment":"The category-specific λ recommendations are not grounded in the reported data. Figure 2 shows aggregate results for λ = 0.2, 0.4, 0.6, and 0.8, and Table 2 reports only λ = 0.6. Table 3 recommends FMCG λ = 0.7–0.8, apparel λ = 0.5–0.6, and electronics λ = 0.3–0.5, but no category-level results at λ values other than 0.6 are shown. The recommendations are also derived from the same ablation experiment that produced the headline numbers, so they cannot independently validate the claimed trade-off curve.","section":"§5.1, §5.2 Table 3"}],"minor_comments":[{"comment":"The formula for D appears garbled with placeholder symbols (e.g., '$s...!' and '#') in the rendering; the intended expression should be typeset clearly so that the sum over the |S| embeddings is unambiguous.","section":"§3.2, Eq. (1)"},{"comment":"The introduction calls the validation 'small-scale A/B tests' while the conclusion says 'small-traffic A/B tests'; the terminology should be consistent and the traffic volume should be stated.","section":"§1, §7"},{"comment":"Many references are unrelated to the claims they support (e.g., [5] on petroleum imaging logging, [8] on normal-vector-assisted mapping, and [17] on COVID-19 collective response), which makes it difficult to trace the related work and undermines the literature review; the manuscript should cite sources directly relevant to e-commerce copy generation and A/B testing.","section":"References"},{"comment":"There is a capitalization error: 'Moreover, Our LLM-driven framework' should read 'Moreover, our LLM-driven framework'; Table 3 also repeats the 'Category' column header in every row and should be cleaned up.","section":"§5.2"}],"recommendation":"reject","confidential_remarks":"For the editor: the reference list contains a large number of citations that are not substantively relevant to the e-commerce copy generation claims, which suggests the related-work section was assembled without direct engagement. This is not the basis for my recommendation; the basis is the internally inconsistent and statistically undocumented central empirical claim. The manuscript would need full re-execution and transparent reporting of the A/B experiment to be reconsidered."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: the headline CTR/CVR numbers do not survive a cross-check against the paper's own Table 2, and without those numbers the paper has no validated central claim.\n\nThe framework itself is a reasonable combination of known pieces: prompt engineering, multi-objective fine-tuning, vector retrieval, logistic conversion prediction, and weighted reranking. The description of the pipeline is clear, and the idea of tuning λ per product category is sensible. The category-specific guidelines in Table 3 are the kind of practical recommendations an e-commerce team might actually use. That is the extent of the contribution.\n\nThe soft spots are fatal. The abstract claims +12.5% CTR and +8.3% CVR; the conclusion claims +10.4% and +4.1% at λ=0.6; Table 2 shows per-category lifts of 36–40% for both metrics. These cannot all describe the same experiment. The stress-test arithmetic is correct: any weighted average of the category lifts lies in the 36–40% range, so neither pair of aggregate numbers is compatible with the table. The paper even admits FMCG conversion improved 'by over 35%', confirming the table and contradicting the overall figures. This is an internal inconsistency, not a missing p-value; it means the central empirical claim is broken on its own terms.\n\nBeyond that, there is no code, no data, no sample sizes, no confidence intervals, and only a hand-wave about Z-tests and chi-square tests. The references are largely disconnected from the claims—many are about unrelated machine learning applications—which does not inspire confidence in the literature grounding. The logistic regression and diversity equations are standard; no new formalism is offered.\n\nIf the underlying A/B test actually ran as described, the results would be commercially interesting, but the paper gives no way to verify that. A reader looking for a template for LLM-based marketing copy generation could skim the architecture and takeaway suggestions, but the validation is absent.\n\nMy recommendation: do not accept for peer review. The internal contradiction alone justifies desk rejection. The authors should be told to either provide the raw aggregated data with confidence intervals or withdraw the empirical claims. There is a kernel of a useful practical write-up here, but not as a research paper.","headline":"The paper's claimed CTR/CVR lifts contradict its own Table 2, making the central empirical claim impossible to trust.","tokens_in":11161,"tokens_out":2378,"would_cite":false,"duration_ms":23553,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An LLM copywriting pipeline with a tunable creativity–conversion knob is claimed to lift CTR and CVR over a human-reviewed template baseline while keeping copy novel.","keywords":["large language models","e-commerce marketing","content generation","creativity conversion trade-off","click-through rate","conversion rate","multi-objective optimization","A/B testing"],"falsifier":"Re-run the same seven-day, fixed-seed traffic split on the same three categories with the same $\\lambda=0.6$ recipe, publishing impression, click, order, and session counts per arm and the resulting $Z$-test or chi-square statistics. If the treatment-minus-control CTR and CVR lifts are not positive with $p<0.05$, the central claim fails; the abstract's +12.5%/+8.3% and the conclusion's +10.4%/+4.1% would not both be reproduced.","tokens_in":10059,"feed_emoji":"🛒","tokens_out":15215,"duration_ms":136651,"temperature":0.7,"pith_summary":"The paper tries to establish that a single LLM-based copywriting pipeline can outperform a human-reviewed, template-based baseline on both halves of the creativity–conversion trade-off: higher click-through and conversion rates without sacrificing output novelty. Its method combines prompt engineering, multi-objective fine-tuning that mixes sentiment adjustment, diversity enhancement, and call-to-action embedding, then runs vector retrieval, post-processing, and a rule-plus-human review stage. Over fast-moving consumer goods, apparel, and electronics, the paper reports relative gains at a creativity weight of $\\lambda=0.6$: CTR +12.5% and CVR +8.3% in the abstract, and CTR +10.4% and CVR +4.1% in the conclusion. If those lifts hold, marketing teams could automate copy at scale while dialing the creativity/conversion balance category by category.","feed_headline":"LLM ad copy beats human-reviewed templates on clicks and sales","feed_subtitle":"One tunable knob lets brands keep copy fresh without losing clicks or orders.","key_machinery":"The load-bearing mechanism is the weighted reward $R = \\lambda D + (1-\\lambda)P_{\\mathrm{conv}}$, where $D$ is one minus the average pairwise cosine similarity among candidate copy embeddings and $P_{\\mathrm{conv}}$ comes from logistic regression on copy features such as CTA density, keyword strength, and sentiment. This single scalar ranks and selects the top-$K$ candidates, so $\\lambda$ acts as both the creativity dial and the conversion guardrail; the vector-retrieval and multi-stage review modules are what make the ranked output deployable as brand-compliant copy.","core_discovery":"The paper's central claim is that a tunable reward $R = \\lambda D + (1-\\lambda)P_{\\mathrm{conv}}$ organizes generated marketing copy along a spectrum from novel to conversion-optimized, and that setting $\\lambda \\approx 0.6$ beats the human-reviewed rule/template baseline on every reported metric in all three categories. Here $D$ is a diversity score (one minus the average pairwise cosine similarity among copy embeddings) and $P_{\\mathrm{conv}}$ is a logistic-regression estimate of conversion probability from features like keyword strength, CTA density, and sentiment. Sweeping $\\lambda$ from 0.2 to 0.8 gives a trade-off curve with an elbow near 0.4–0.6, and the category analysis says FMCG can tolerate high creativity, apparel needs moderation, and electronics should emphasize factual specification.","pith_inferences":["The $\\lambda D + (1-\\lambda)P_{\\mathrm{conv}}$ form is a generic two-objective controller: the same construction could be lifted to any text-generation setting where novelty competes with a measurable objective, such as headline selection or recommendation descriptions.","As defined, $D$ only measures similarity among candidates generated under the same prompt, so it could be inflated by cheap paraphrasing; the human-rating component is what keeps perceived novelty meaningful.","Because $\\lambda$ is hand-set by elbow inspection, adding a learned scheduler that picks $\\lambda$ per campaign from historical CTR/CVR logs is a direct next step the paper points at but does not implement."],"forward_implications":["At $\\lambda=0.6$, the framework reports higher diversity, CTR, CVR, and human novelty/fluency ratings than the baseline in all three categories; for FMCG, CTR rises from 8.9% to 12.1% and CVR from 3.8% to 5.2%.","The $\\lambda$ sweep gives a practical tuning rule: use high creativity for impulse-driven flash sales, moderate creativity for apparel, and low creativity for electronics, where extra flair can undercut trust.","Because every candidate passes rule checks, sensitive-word filters, brand-guideline constraints, and human review, the scheme is designed to be brand-compliant while still automating the drafting step.","Category-specific prompt libraries and dynamic $\\lambda$ scheduling let the same fine-tuned model serve different campaigns without retraining, shifting $\\lambda$ down during clearance events and up during launches."],"supporting_citations":[{"why":"The paper cites this emotion-aware context modeling work to justify the creativity side of the framework: emotional tone and novelty drive engagement.","marker":"[4]"},{"why":"The paper cites this work for evidence that structured scarcity cues like 'Only 10 Left' improve CTR and CVR, motivating the CTA-embedding objective.","marker":"[5]"},{"why":"This citation supplies the parameter-adaptability idea that the trade-off weight lambda should be tuned per category.","marker":"[6]"},{"why":"The paper cites this scheduling work to support the scalable, high-concurrency generation assumption behind the vector-retrieval pipeline.","marker":"[7]"}],"fun_headline_variants":["LLM copy tuning lifts clicks 12.5% and orders 8.3% vs templates","One knob balances AI ad copy creativity and conversion, lifting CTR 12.5%","Category-aware LLM copy: FMCG creative, electronics factual, apparel mid","12.5% more clicks, 8.3% more orders via LLM copy creativity knob","Tunable diversity-conversion reward lifts LLM ad copy CTR and CVR"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The empirical claim rests on the seven-day randomized A/B test described in Section 4.2 having been run on a live platform with fixed-seed traffic splitting, consistent logging, and valid significance tests; no traffic counts, p-values, confidence intervals, or raw data are supplied.","fun_headline_variants_meta":{"raw":{"variants":["LLM copy tuning lifts clicks 12.5% and orders 8.3% vs templates","One knob balances AI ad copy creativity and conversion, lifting CTR 12.5%","Category-aware LLM copy: FMCG creative, electronics factual, apparel mid","12.5% more clicks, 8.3% more orders via LLM copy creativity knob","Tunable diversity-conversion reward lifts LLM ad copy CTR and CVR"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000651,"raw_usage":{"total_tokens":2919,"prompt_tokens":815,"completion_tokens":2104,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":431,"completion_tokens_details":{"reasoning_tokens":1989}},"tokens_in":431,"tokens_out":2104,"duration_ms":13204,"temperature":1.0,"reasoning_tokens":1989,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:45:19.506488+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same seven-day, fixed-seed traffic split on the same three categories with the same $\\lambda=0.6$ recipe, publishing impression, click, order, and session counts per arm and the resulting $Z$-test or chi-square statistics. If the treatment-minus-control CTR and CVR lifts are not positive with $p<0.05$, the central claim fails; the abstract's +12.5%/+8.3% and the conclusion's +10.4%/+4.1% would not both be reproduced.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The paper cites this work for evidence that structured scarcity cues like 'Only 10 Left' improve CTR and CVR, motivating the CTA-embedding objective."},{"cited_title":"Machine Learning-Based Research on the Adaptability of Adolescents to Online Education","cited_arxiv_id":"2408.16849","evidence_quote":"This citation supplies the parameter-adaptability idea that the trade-off weight lambda should be tuned per category."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The paper cites this scheduling work to support the scalable, high-concurrency generation assumption behind the vector-retrieval pipeline."}],"review_version":1}