{"id":"76eb770d-bd89-4e99-94f0-5ce94699780a","arxiv_id":"2606.10159","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Adversarial rephrasing of abstracts raises AI peer review scores by up to 1.31 points on a 10-point scale with 38% attack success rate across models and disciplines.","lead":"This paper shows that rephrasing a paper's abstract superficially can substantially improve scores from AI peer reviewers without changing the underlying science. A smart generalist should read it to understand risks when AI tools influence what research gets published.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"No independent check that rephrased abstracts preserve scientific content and communication quality","rationale":"The reader's weakest_assumption matches the load-bearing point exactly. Full-text access does not appear to add an explicit equivalence validation step, so the empirical results remain conditional on an untested premise. No other internal inconsistency (model choice, metric definition, or extrapolation to humans) rises to the same level of centrality for the reported attack rates.","tokens_in":1824,"tokens_out":343,"duration_ms":17865,"concrete_test":"Blind domain experts rate 50 random original-vs-rephrased abstract pairs on three 1-5 scales (scientific claims preserved, clarity, persuasiveness). If mean clarity or persuasiveness differs by >0.5 points while claims-preserved stays near 5, recompute attack success on the subset where equivalence holds; if success drops below 20%, the content-neutrality premise fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim (38% attack success, +1.31 / +0.88 score lifts) rests on the assertion that the adversarial rewrites are superficial and leave both scientific content and communication unchanged. The abstract states this but supplies no verification step (expert equivalence ratings, semantic entailment checks, or blinded human comparison of claims). Without that, observed score gains could arise from legitimate improvements in clarity, emphasis, or structure rather than model-specific gaming. This assumption is load-bearing because the paper contrasts the attack with 'ordinary scientific editing' and claims the effect is manipulative; if the rewrites are not content-neutral, the headline risk to AI-assisted review is not demonstrated.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper empirically demonstrates that superficial adversarial rephrasing of manuscript abstracts can substantially improve scores from AI peer reviewers (e.g., Gemini 3 Flash and GPT 5.4 Mini) without altering underlying scientific content, achieving ~38% attack success rate, score lifts of +1.31 and +0.88 on a 10-point scale, and >50% success when originals receive 'reject' recommendations. The effect is shown across disciplines, venues, and both human- and AI-written papers, at low cost (~5 min, $1), and is claimed to be hard to distinguish from ordinary editing.","tokens_in":1949,"tokens_out":544,"duration_ms":16561,"significance":"If the central empirical result holds with verified content neutrality, the work provides concrete evidence of a practical vulnerability in AI-assisted peer review that could bias editorial decisions and create perverse incentives. The low-cost, model-agnostic nature of the attack and its extension to core criteria (soundness, significance) make the finding actionable for robustness testing. The purely empirical approach with quantitative outcomes across multiple models is a strength.","major_comments":[{"comment":"Abstract: The central claim that 'adversarially rewritten abstracts' improve AI outcomes 'without changing the underlying scientific content and communication' is load-bearing for interpreting the results as evidence of gaming rather than legitimate improvement. No verification step (expert equivalence ratings, blinded human comparison, semantic entailment checks, or similarity metrics) is described to support content neutrality. Without this, observed score gains could reflect clarity or emphasis changes.","section":"Abstract"},{"comment":"Methods/Results (as referenced in abstract): The reported quantitative outcomes (38% ASR, specific score deltas of +1.31/+0.88, >50% conditional success) lack accompanying details on sample sizes, statistical tests, error bars, exact paper datasets, or how adversarial rewrites were generated and validated for equivalence. This prevents full assessment of reproducibility and effect robustness.","section":"Methods/Results"}],"minor_comments":[{"comment":"Clarify model versions (e.g., 'Gemini 3 Flash', 'GPT 5.4 Mini') with precise identifiers and whether prompts or system instructions were held constant across conditions.","section":"Abstract"},{"comment":"The abstract states the attack 'extends beyond overall score inflation' to core criteria; provide the per-criterion score tables or breakdowns to support this.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback. We respond to each major comment below and indicate the revisions planned for the next version of the manuscript.","responses":[{"response":"We agree that the absence of explicit verification steps for content neutrality is a limitation in the current manuscript. Although the adversarial rewrites were constructed to be superficial and preserve scientific meaning, we did not report expert equivalence ratings, blinded comparisons, or semantic metrics. We will add these verification procedures in the revised version to support the claim of content neutrality.","revision_made":"yes","referee_comment":"[Abstract] Abstract: The central claim that 'adversarially rewritten abstracts' improve AI outcomes 'without changing the underlying scientific content and communication' is load-bearing for interpreting the results as evidence of gaming rather than legitimate improvement. No verification step (expert equivalence ratings, blinded human comparison, semantic entailment checks, or similarity metrics) is described to support content neutrality. Without this, observed score gains could reflect clarity or emphasis changes."},{"response":"The manuscript describes the paper datasets and rewrite generation process in the Methods section, but we acknowledge that sample sizes, statistical tests, error bars, and explicit equivalence validation details are not presented with sufficient clarity. We will expand the Methods and Results sections to include these elements, along with appropriate statistical reporting, to improve reproducibility.","revision_made":"yes","referee_comment":"[Methods/Results] Methods/Results (as referenced in abstract): The reported quantitative outcomes (38% ASR, specific score deltas of +1.31/+0.88, >50% conditional success) lack accompanying details on sample sizes, statistical tests, error bars, exact paper datasets, or how adversarial rewrites were generated and validated for equivalence. This prevents full assessment of reproducibility and effect robustness."}],"tokens_in":1522,"tokens_out":397,"duration_ms":29364,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main point is that superficial abstract rephrasing lifts AI-assigned acceptance ratings by +0.88 to +1.31 on a 10-point scale across Gemini and GPT models, with an overall attack success rate near 38 percent that climbs above 50 percent on initial rejects. The effect also appears on sub-criteria such as soundness and significance.\n\nThe work is new in its direct empirical test of this manipulation on peer-review prompts, using both human-written and AI-generated papers across disciplines. It reports concrete numbers, notes the low effort required, and flags the downstream risk that inflated AI scores could shift human editorial calls.\n\nThe paper does a reasonable job documenting the attack surface and showing consistency across a few models. The practicality details help ground the claim.\n\nThe central soft spot is the missing verification that the rewrites are truly content-neutral. The abstract states the changes do not alter underlying science or communication, yet there is no blinded human rating, expert equivalence check, or semantic entailment test to support that. If the rephrased versions simply improve clarity or emphasis, the score gains are not evidence of gaming. That assumption is load-bearing for the risk narrative.\n\nMethods details on how the adversarial prompts were built and how papers were sampled would also need scrutiny in review, but the empirical framing itself is straightforward.\n\nThis is for groups building or evaluating AI review tools and for editors weighing their use. Readers focused on LLM robustness in evaluative settings will find the numbers worth seeing. The demonstration is sharp enough to merit a serious referee even if the content-preservation gap requires follow-up experiments.\n\nI would send it to peer review.","headline":"The paper shows rephrasing abstracts can raise AI review scores by about a point, but supplies no check that the rewrites leave scientific content and quality unchanged.","tokens_in":2422,"tokens_out":415,"would_cite":false,"duration_ms":16435,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Superficial rephrasing of abstracts raises AI peer review scores by up to 1.31 points without changing scientific content.","keywords":["AI peer review","adversarial manipulation","scientific publishing","review robustness","manuscript editing","AI vulnerability","peer review gaming"],"falsifier":"Running the same rephrased abstracts through multiple additional AI models or human reviewers and finding that score gains disappear or that humans consistently rate the rephrased versions lower would falsify the claim of model-agnostic, content-preserving manipulation.","tokens_in":2730,"feed_emoji":"🤖","tokens_out":705,"duration_ms":13132,"temperature":0.7,"pith_summary":"The paper shows that AI tools assisting peer review can be gamed by rewriting a manuscript's abstract in ways that leave the underlying science and arguments unchanged. These edits improve overall acceptance ratings, reviewer confidence, and scores on criteria such as soundness and significance. The effect appears across disciplines and paper types, with higher success when the original review leans toward rejection. Because the manipulation is cheap and hard to spot, it could encourage authors to write for AI scorers rather than for scientific quality and could tilt human editorial decisions that incorporate AI outputs.","feed_headline":"Abstract rephrasing lifts AI review scores 38 percent of the time","feed_subtitle":"Simple edits raise acceptance ratings by over a point on a 10-point scale while leaving the science unchanged.","key_machinery":"Adversarial abstract rephrasing attack that optimizes wording to raise AI-assigned scores on soundness, significance, and contribution while preserving original scientific claims.","core_discovery":"AI-mediated peer review is vulnerable to a simple, low-cost manipulation: superficial rephrasing of the manuscript abstract. Without changing the underlying scientific content and communication, and even without knowledge of the reviewing model, adversarially rewritten abstracts substantially improve AI review outcomes. Our strongest attack achieves an attack-success-rate of about 38%, increasing acceptance ratings by +1.31 for Gemini 3 Flash reviewers and by +0.88 for GPT 5.4 Mini reviewers on a 10-point scale. When the original AI review suggests 'reject', the success rate rises to more than 50%.","pith_inferences":["Detection tools could scan abstracts for stylistic markers that differ from the rest of the manuscript.","Journals might add a consistency check between abstract and full text as a low-cost safeguard.","The same vulnerability could appear in AI-assisted grant review or hiring evaluations that rely on short summaries.","Extending the attack from abstracts to full-text sections would test whether the effect scales with more text."],"forward_implications":["Inflated AI reviews can shift downstream human editorial recommendations from rejection toward acceptance.","Authors gain an incentive to optimize abstracts and manuscripts for AI judgment rather than scientific merit.","The attack works on both human-written and AI-generated papers and across disciplines and venues.","Review confidence and per-criterion scores rise along with the overall acceptance rating.","AI tools require systematic robustness testing and human oversight before influencing high-stakes decisions."],"fun_headline_variants":["Abstract rephrasing raises AI review scores in 38 percent of cases","Simple abstract edits improve AI review ratings 38 percent","AI peer review fooled by abstract rephrasing in 38 percent cases","Abstract tweaks lift AI acceptance ratings over one point"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The rephrasing leaves the underlying scientific content and communication unchanged and succeeds without any knowledge of the specific AI reviewer model.","fun_headline_variants_meta":{"raw":{"variants":["Abstract rephrasing raises AI review scores in 38 percent of cases","Simple abstract edits improve AI review ratings 38 percent","AI peer review fooled by abstract rephrasing in 38 percent cases","Abstract tweaks lift AI acceptance ratings over one point"]},"model":"grok-4.3","cost_usd":0.006369,"raw_usage":{"total_tokens":2973,"prompt_tokens":797,"num_sources_used":0,"completion_tokens":68,"cost_in_usd_ticks":63690500,"prompt_tokens_details":{"text_tokens":797,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2108,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":797,"tokens_out":68,"duration_ms":12806,"temperature":1.0,"reasoning_tokens":2108,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-27T16:02:10.104133+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Running the same rephrased abstracts through multiple additional AI models or human reviewers and finding that score gains disappear or that humans consistently rate the rephrased versions lower would falsify the claim of model-agnostic, content-preserving manipulation.","supporting_citations":[],"review_version":1}