{"id":"98f54788-b08b-46ec-ac87-e501accad84f","arxiv_id":"2508.13223","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Introduces a wild-AIGI benchmark and Mirage-R1, a reasoning-trained vision-language model that improves AI-image detection by 5% and 10% over prior detectors.","lead":"This paper introduces Mirage, a benchmark of AI-generated images built to resemble real, messy online photos, and Mirage-R1, a vision-language detector that reasons before judging. The authors report that Mirage-R1 beats prior detectors by 5% on their benchmark and 10% on a public benchmark, but this review was done from the abstract only.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Training/evaluation overlap on Mirage may inflate reported 5%/10% margins; fair-comparison evidence is missing.","rationale":"The abstract-only review leaves the strongest claim (5%/10% SOTA gains) without supporting experimental detail. The single most load-bearing condition is not just that Mirage is representative of the wild, but that Mirage-R1's reported gains are not an artifact of training/evaluating on the same benchmark. The phrase 'Building on this benchmark' makes train/test overlap a live possibility concrete enough to test. The reader focuses on benchmark representativeness; my concern is a more specific failure mode within that assumption (distribution leakage/fair comparison). I therefore partially agree. The proposed test—checking split disjointness and baseline parity—would settle the concern. If the margins persist under that check, the central claim is substantially supported; if not, it is overstated. No objection is raised about novelty or the use of human verification per se; those are secondary.","tokens_in":740,"tokens_out":4220,"duration_ms":49544,"concrete_test":"Run the released evaluation code/weights and check: (1) no Mirage training/validation samples occur in the Mirage test split used for the 5% comparison; (2) all baselines are retrained/tuned on the same training data with identical preprocessing, metric, and inference-time budget; (3) the public benchmark was not used in RL training/reward selection. If these conditions hold and the 5%/10% margins persist, the claim stands; otherwise the reported margins overstate generalization.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is the reported SOTA margins (5% on Mirage, 10% on the public benchmark). For that claim to hold, Mirage-R1's training procedure must be disjoint from the evaluation splits, and baselines must be compared under identical conditions. The abstract says Mirage-R1 is 'trained in two stages' and 'Building on this benchmark, we propose Mirage-R1,' which leaves open the possibility that Mirage was used for both training and testing. If the same benchmark distribution is used in SFT/RL training and evaluation, the 5% margin could reflect in-distribution memorization rather than in-the-wild generalization; if the public benchmark was used as a reward signal during RL, the 10% margin could reflect optimization against that test set. The abstract provides no dataset splits, baseline tuning protocol, metric definition, or significance testing. This is not an accusation of improper conduct; it is an unverified condition that must hold for the central claim to be accepted.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Mirage, a benchmark intended to emulate in-the-wild AI-generated images (AIGI), built from two sources: human-verified Internet-sourced AIGI and a multi-generator synthetic dataset. The authors also propose Mirage-R1, a vision-language model with heuristic-to-analytic and reflective reasoning, trained via supervised fine-tuning followed by reinforcement learning, and an inference-time adaptive thinking strategy. The abstract claims that Mirage-R1 outperforms state-of-the-art detectors by 5% on Mirage and 10% on a public benchmark. No full text is available; this review is based solely on the abstract.","tokens_in":988,"tokens_out":1683,"duration_ms":19602,"significance":"If the reported results hold, the paper would make a useful contribution by providing a challenging in-the-wild benchmark and a detector that combines reasoning with reflection, potentially improving generalization over laboratory-trained detectors. The benchmark's construction from expert-verified Internet images and multi-generator synthetic edits is promising. However, the abstract alone provides no way to verify the central quantitative claim, error bars, baselines, or methodology details, so the significance currently hinges on assertions that cannot be checked from the available material.","major_comments":[{"comment":"The central claim, 'leads state-of-the-art detectors by 5% and 10% on Mirage and the public benchmark, respectively,' is stated without any experimental detail: no baselines are named, no metric is defined, no error bars or significance tests are reported, and no dataset sizes or evaluation protocols are given. This is load-bearing because the entire contribution rests on these margins. Please provide the full evaluation setup, including the specific detectors compared, the evaluation metric (e.g., accuracy, AUC, F1), and variance or statistical significance.","section":"Abstract"},{"comment":"The abstract says Mirage-R1 is 'Building on this benchmark' and 'trained in two stages: a supervised-fine-tuning cold start, followed by a reinforcement learning stage.' It is unclear whether Mirage is used for training, validation, or both. If the same benchmark distribution is used in model training and evaluation, the reported 5% improvement on Mirage could reflect in-distribution overfitting rather than in-the-wild generalization. Please specify the exact train/validation/test splits of Mirage and confirm that no Mirage evaluation data is used in SFT or RL training.","section":"Abstract"},{"comment":"The 10% improvement on 'the public benchmark' is not accompanied by a name or protocol. If this benchmark was used as a reward signal in the reinforcement learning stage, then the model may have been optimized against that test set, making the 10% margin a measure of optimization rather than generalization. Please state whether the public benchmark was used in any way during training, and if so, which split was used for the reported evaluation.","section":"Abstract"},{"comment":"The benchmark is described as 'designed to emulate the complexity of in-the-wild AIGI' from two sources. However, the abstract offers no evidence that these sources are representative of the real-world distribution of AIGI. The human-verification protocol and the processes for selecting and editing synthetic images should be described, along with any comparison to existing in-the-wild datasets. Without this, the specialization to 'in-the-wild' conditions is not established.","section":"Abstract"}],"minor_comments":[{"comment":"The phrase 'leads state-of-the-art detectors by 5% and 10%' is grammatically awkward; consider 'outperforms state-of-the-art detectors by 5% and 10%'.","section":"Abstract"},{"comment":"The names 'Mirage' and 'Mirage-R1' are similar; clarify that Mirage is the benchmark and Mirage-R1 is the proposed model to avoid reader confusion.","section":"Abstract"},{"comment":"The term 'heuristic-to-analytic reasoning' is introduced without explanation; a brief definition would help readers understand the proposed mechanism.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"This review is based only on the abstract; the full text was unavailable. The central claim is unverified because the abstract omits all experimental details. The recommendation of major_revision reflects that the concerns are addressable by adding the missing information, rather than a fundamental flaw. I also note that the benchmark and detector share authors, so the potential train/evaluation leakage is a serious concern that must be explicitly ruled out in the full paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is an abstract-only read, so I can't verify the central claim. The benchmark and the training recipe are plausible and worth engaging; the reported 5% and 10% margins are simply unsubstantiated at this level of detail.\n\nWhat's actually new here: Mirage, a benchmark aimed at in-the-wild AIGI, built from human-verified internet images and multi-generator synthetic edits. That's a useful resource if it's real. And Mirage-R1 applies a two-stage VLM recipe — SFT cold start then RL — with a heuristic-to-analytic reasoning mechanism and an adaptive thinking/inference-time trade-off. That's a sensible extension of the VLM reasoning line to AIGI detection, not a paradigm shift but a plausible engineering contribution.\n\nThe soft spots are mostly what's missing. The abstract gives no dataset sizes, split construction, baseline tuning details, metric definitions, or error bars. The public benchmark result is the best independent grounding, but we don't know whether that benchmark or a reward derived from it influenced training. The stress-test concern about training/evaluation overlap on Mirage is legitimate: if Mirage was used in both stages of training and then evaluated on itself, a 5% margin could reflect in-distribution memorization. I'm not accusing anyone of that; it's an unverified condition that needs to be stated. Same for the benchmark's 'in-the-wild' claim: two sources are listed, but representativeness is asserted, not demonstrated.\n\nGiven that we only have the abstract, I'd treat this as a promising but unproven preprint. The right move is to send it to peer review — not desk-reject — and make sure the reviewers ask for explicit train/eval separation, baseline protocols, and error analysis. If the authors deliver that, the benchmark alone could be worth citing. For my own work, I wouldn't cite it yet, but I'd definitely put it on the reading group list once the full version is out.","headline":"Abstract-only: promising benchmark and training recipe, but the 5%/10% margins are unverified and need train/eval separation details.","tokens_in":1421,"tokens_out":2070,"would_cite":false,"duration_ms":21539,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Mirage-R1 beats top AI-image detectors by 10% on a public wild-image benchmark.","keywords":["AI-generated image detection","in-the-wild generalization","vision-language model","reinforcement learning","image forensics","benchmark construction","adaptive reasoning","synthetic media"],"falsifier":"Collect a fresh set of AI-generated images from models and platforms released after the Mirage training data, run Mirage-R1 and the compared detectors on them under identical conditions, and compare accuracies; a large drop in Mirage-R1's relative advantage would falsify the claim that the method generalizes to in-the-wild AI-generated images.","tokens_in":701,"feed_emoji":"🖼️","tokens_out":4551,"duration_ms":52987,"temperature":0.7,"pith_summary":"This paper addresses why AI-generated image detectors fail when images leave clean laboratory settings and enter the noisy, edited, multi-model reality of the internet. It introduces Mirage, a benchmark assembled from human-verified internet-sourced AI images and from images produced by several expert generators and then edited, meant to stand in for in-the-wild AI-generated images. On top of this benchmark, the authors propose Mirage-R1, a vision-language model that first makes a quick heuristic judgment and then reasons reflectively over visual evidence before deciding. Trained with a supervised cold start followed by reinforcement learning, Mirage-R1 reports accuracy gains of 5% over prior detectors on Mirage and 10% on a public benchmark, suggesting that explicit reasoning, rather than a single learned classifier, can generalize better to wild images.","feed_headline":"Mirage-R1 beats top AI-image detectors by 10%","feed_subtitle":"A new benchmark plus reflective reasoning helps spot AI-generated photos in noisy, edited real-world settings.","key_machinery":"The central mechanism is Mirage-R1's heuristic-to-analytic reasoning: the model first produces a quick heuristic judgment about whether an image is AI-generated, then engages in reflective reasoning over the visual evidence before settling on a verdict. This is trained with a supervised-fine-tuning cold start followed by a reinforcement-learning stage, and at inference time an adaptive thinking strategy lets the model either return a fast judgment or spend additional compute for a more accurate conclusion. The Mirage benchmark supplies the measuring stick: a corpus of human-verified internet AI images plus multi-generator synthetic edits that are meant to replicate the difficulty of real-wor","core_discovery":"The paper claims that a detector which reasons in two stages—forming a quick heuristic impression and then deliberately reflecting on the visual evidence—can outperform specialist classifiers on AI-generated images encountered in the wild. This claim is carried by the Mirage benchmark, which combines human-verified AI images collected from the Internet with synthetic images produced by multiple expert generators and then edited to mimic real-world quality-control pipelines. The proposed model, Mirage-R1, is a vision-language model trained in two stages: a supervised-fine-tuning cold start followed by reinforcement learning, and it uses an adaptive thinking strategy at inference time to choos","pith_inferences":["A natural next test is time-shift robustness: collect AI images after the model's training cutoff from platforms not represented in Mirage and measure how much of the 10% public-benchmark advantage survives.","The two-stage architecture suggests a practical pipeline design: a cheap heuristic filter that only invokes the slow analytic reasoning on uncertain or high-stakes images, extending adaptive thinking from per-image choice to an end-to-end moderation flow.","The benchmark's synthetic half, built from multiple expert generators plus edits, could be reused for confidence calibration or provenance tracing rather than only binary fake/real decisions.","A finer-grained evaluation that slices Mirage accuracy by edit type or generator would be a direct way to identify which real-world distortions the reasoning stage actually overcomes."],"forward_implications":["Mirage provides a harder, more realistic evaluation target: detectors must handle internet noise, mixed generator sources, and post-generation edits to score well.","Two-stage reasoning in a vision-language model can replace or augment dedicated binary classifiers for AI-generated image detection, shifting the problem from feature discrimination to deliberative judgment.","The adaptive thinking strategy lets a single model cover speed-critical moderation with a quick heuristic and accuracy-critical forensic review with analytic reasoning using the same weights.","Because the public-benchmark gain is larger than the gain on Mirage, the method's advantage is not confined to the new benchmark's particular construction choices.","If the reported accuracy holds, deployment in social-media-style pipelines becomes feasible, where images arrive noisy, cropped, and re-edited rather than pristine and generator-native."],"supporting_citations":[],"fun_headline_variants":["New benchmark plus reflective reasoning boosts AIGI detection by 10%","Mirage-R1: two-stage reasoning outdoes specialist AI-image detectors","Heuristic-to-analytic model tops AI-image detection in the wild","Adaptive thinking strategy helps Mirage-R1 balance speed and accuracy","Reinforcement-trained VLM with reflection beats existing AIGI detectors"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The load-bearing premise is that Mirage's mix of human-verified internet images and multi-generator edited images faithfully represents the full distribution of AI images people actually meet online; if the real world looks different, the reported 5% and 10% gains may not travel.","fun_headline_variants_meta":{"raw":{"variants":["New benchmark plus reflective reasoning boosts AIGI detection by 10%","Mirage-R1: two-stage reasoning outdoes specialist AI-image detectors","Heuristic-to-analytic model tops AI-image detection in the wild","Adaptive thinking strategy helps Mirage-R1 balance speed and accuracy","Reinforcement-trained VLM with reflection beats existing AIGI detectors"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000231,"raw_usage":{"total_tokens":1345,"prompt_tokens":787,"completion_tokens":558,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":531,"completion_tokens_details":{"reasoning_tokens":463}},"tokens_in":531,"tokens_out":558,"duration_ms":7167,"temperature":1.0,"reasoning_tokens":463,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T19:29:23.502023+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect a fresh set of AI-generated images from models and platforms released after the Mirage training data, run Mirage-R1 and the compared detectors on them under identical conditions, and compare accuracies; a large drop in Mirage-R1's relative advantage would falsify the claim that the method generalizes to in-the-wild AI-generated images.","supporting_citations":[],"review_version":1}