{"id":"e52df318-3a32-4c35-bf8e-e7eb85875461","arxiv_id":"2511.04080","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Adding VLM-generated, LLM-refined image captions into source text improves source visibility in generative search by about 1–2% relative, per the paper's G-Eval measurements on MRAMG.","lead":"Caption Injection adds image captions, rewritten by a language model, into web-page text so that the page is more likely to be cited by AI-generated search answers. On the MRAMG benchmark it reports 1.09–1.85% relative visibility gains over text-only baselines, though the evaluation rests on an LLM judge rather than human readers.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Multimodal GSE simulation reduces images to captions; the claimed multimodal benefit may be an artifact of a text-only pipeline.","rationale":"The reader's weakest assumption identifies the same central problem: the multimodal simulation is text-only because images are reduced to captions, and the supporting analysis for that design choice is not shown. This is the most load-bearing concern because it directly undermines the paper's core contribution—being the first multimodal G-SEO method and showing multimodal information helps. If the simulation never exercises a true multimodal GSE, then the claimed superiority in the 'multimodal' scenario is only about injecting extra text into a text-based system. The proposed concrete test—running a real multimodal GSE with original images—would settle whether the method transfers. Other concerns (e.g., LLM-judge validity, lack of significance testing) are real but secondary; the simulation issue is the one that most directly questions the central claim. Since the reader already issued CONDITIONAL based on this concern, my read does not change the verdict. I agree with the reader's identification of the weakest assumption.","tokens_in":13308,"tokens_out":4895,"duration_ms":51097,"concrete_test":"Replace the simulated multimodal GSE in Section IV-A1 with a genuine multimodal generator that receives original images alongside text (e.g., Qwen2.5-VL-7B or GPT-4o). Keep the same MRAMG data, same G-Eval 2.0 evaluation, and same comparison protocol. If Caption Injection still outperforms text-only baselines by a margin comparable to the reported +1.09%, the multimodal claim holds; if the gain disappears or reverses, the result is an artifact of the caption-only simulation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that Caption Injection improves subjective visibility in multimodal GSEs, yet the experimental simulation in Section IV-A1 does not use a multimodal GSE that processes images. The authors state that they compared two configurations—'text content + image captions + LLM' versus 'text content + original images + multimodal LLM'—and adopted the former because it 'generally yields higher response quality,' but this supporting comparison is not shown. In the deployed simulation, the GSE receives only text (with injected captions) and never sees the original images. Consequently, the reported +1.09% relative gain in the 'multimodal' condition only demonstrates that adding caption text to source content improves visibility in a text-only GSE. It does not show that the method helps when a GSE can reason over images directly, which is the realistic multimodal setting that the paper claims to address. If real multimodal GSEs use images as visual evidence, Caption Injection's advantage may vanish or reverse, because the model already has access to the visual information that the captions are trying to inject. The missing comparison is therefore a load-bearing gap: without it, the multimodal claim is unsupported, and the method risks being just a text augmentation technique.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Caption Injection, described as the first multimodal G-SEO method. A VLM generates a structural caption from each image, an LLM refines it against the source text, and the refined caption is injected into the textual content through prompt engineering. The method is evaluated on the MRAMG benchmark under two simulated GSE configurations: a text-only (unimodal) pipeline and a multimodal pipeline that feeds the GSE text plus image captions. Using G-Eval 2.0 as the evaluation metric, the paper reports relative improvements in subjective visibility of +1.85% (unimodal) and +1.09% (multimodal) over text-only baselines, and concludes that multimodal information benefits G-SEO.","tokens_in":13606,"tokens_out":2255,"duration_ms":22643,"significance":"If the claims are substantiated, the paper opens a useful direction by extending G-SEO from purely textual optimization to a setting where visual semantics are projected into text. The three-stage pipeline (structural generation, alignment refinement, semantic injection) is clearly described, and the authors make code available. However, the central empirical claims rest on an evaluation design that uses a text-only surrogate for a multimodal GSE, an LLM-based metric from the authors' own prior work, and no statistical validation. The reported gains are small (1–2% relative), so the significance currently depends on assumptions that are not fully disclosed or tested.","major_comments":[{"comment":"The multimodal GSE simulation feeds the model only 'text content + image captions + LLM', not original images or a multimodal LLM. The authors say they compared this configuration with 'text content + original images + multimodal LLM' and that the former 'generally yields higher response quality', but the supporting comparison is not shown. This is load-bearing: the claimed multimodal benefit (+1.09%) could simply reflect that adding caption text improves a text-only GSE. The paper does not demonstrate that Caption Injection helps when the GSE can reason over images directly, which is the realistic multimodal scenario claimed in the title and abstract. Please include the omitted comparison or evaluate with a genuine multimodal GSE.","section":"§IV-A1 (Generative Search Engine Simulation)"},{"comment":"The sole evaluation metric is G-Eval 2.0, which is introduced in the authors' own prior work [8]. There is no human validation, no correlation analysis with human judgments, no error bars, and no significance tests. Several dimensions—especially content volume, click-follow likelihood, and positional salience—by construction favor enriched text, so part of the measured gain may be metric-induced. Given the small effect sizes (+1.85% and +1.09%), the possibility that the differences are within LLM-judge noise is not addressed. Please add a human evaluation on a sample, significance testing, or at least a robustness analysis with multiple judge models and seeds.","section":"§IV-A4 (Evaluation Metrics)"},{"comment":"The text states that 'the refined captions consistently yield higher subjective visibility scores than both the original and VLM-generated captions'. Table V shows average scores of 1.18 for the original caption, 1.11 for the rewritten (refined) caption, and 1.06 for the VLM-generated caption. The refined caption is not the top performer; the original caption is. This contradicts the ablation claim and weakens the argument that the refinement stage is beneficial. Please correct the claim or the table, and clarify what is being compared.","section":"§IV-B2 (Ablation Study) and Table V"}],"minor_comments":[{"comment":"The method names in Table IV are abbreviated inconsistently ('tran seo', 'flue expr', 'quat addi', 'stat addi', 'capt addi') and the table has no descriptive caption. Please use full method names or define all abbreviations.","section":"§IV-B1, Table IV"},{"comment":"The formula for relative improvement uses 'imprs(r)' in the denominator and 'impression_s(r)' in the numerator, which is typographically inconsistent. Clarify notation and check whether the '+1' term is intentional.","section":"§IV-A4, improvement formula"},{"comment":"The table title says 'EVALUATION OF G-SEO METHODS ON DIFFERENT METRICS (VALUES×100)', but this is an ablation of caption types. The column header 'Sub.Posi.' and 'Sub.Volu.' are not expanded. Also 'rewriten' is misspelled in the last row.","section":"§IV-B2, Table V"},{"comment":"The prompt for caption injection in Table III is labeled 'Prompt for caption refinement', which duplicates the label of Table II. Please correct the label.","section":"§III-B, prompt tables"},{"comment":"The paper relies heavily on the authors' own unpublished prior work [8] for G-Eval 2.0 and the overall evaluation protocol; please make the dependency explicit in the methodology section so readers can assess the lineage of the metric.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The central concern is that the 'multimodal' evaluation is not actually multimodal in the sense that matters—the GSE never sees images. The omitted comparison between 'text+captions+LLM' and 'text+images+multimodal LLM' is exactly the evidence needed to support the paper's headline claim. Additionally, using a metric from the authors' own prior work without any human validation or significance testing is risky for a venue that expects rigorous empirical support. The ablation contradiction in Table V should be fixed before any revision is sent back."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the good news: this paper is the first to frame G-SEO in a multimodal setting, and the Caption Injection pipeline (VLM structural caption → LLM refinement → prompt injection) is a sensible, reproducible first attempt. The code is public, the ablation against original vs VLM vs refined captions is useful, and the authors are honest in the limitation section about the shallow fusion and the drop from unimodal to multimodal. This is worth reading for anyone working on generative search optimization.\n\nThe bad news is the load-bearing gap that the stress-test note identified, and I think it holds up. The 'multimodal GSE' in Section IV-A1 is actually a text-only LLM that receives captions instead of images. The authors say they compared that configuration against a true multimodal setup (text + original images + multimodal LLM) and found captions worked better, but they never show that comparison. So the +1.09% gain in the 'multimodal' condition only shows that adding caption text helps a text-based generator. It does not show that the method helps when a GSE can reason over images directly — and in that scenario the injected captions may be redundant or even harmful. Without that omitted comparison, the central claim of 'the first multimodal G-SEO approach' is not actually supported by the experiments.\n\nThe evaluation also has a circularity problem. The metric is G-Eval 2.0, which comes from the authors' own prior work, and several of its dimensions (content volume, click-follow likelihood) reward the kind of enrichment that caption injection is doing by construction. No human validation, no error bars, no significance testing. The effect sizes are 1-2% relative gains — plausible, but not enough to be confident in the claim of 'significant improvement.'\n\nNone of this makes the paper worthless. It is a legitimate exploratory result, and the authors already admit that the optimization effect is modest. But the claims in the abstract and conclusion are stronger than the evidence. A serious referee should ask for the missing simulation comparison, a human-rated subset of G-Eval 2.0, and either a true multimodal baseline or a rephrased claim that limits the result to caption-based augmentation.\n\nFor a researcher in G-SEO, this is worth a look as related work. I would not cite it as evidence for multimodal G-SEO until the gap is closed. But it is a clear, honest first step, and I'd prefer to see it in the literature after major revisions rather than desk-rejected.","headline":"Nice first stab at multimodal G-SEO, but the 'multimodal' experiment never actually shows the GSE an image — the headline claim is unsupported.","tokens_in":14018,"tokens_out":2261,"would_cite":false,"duration_ms":21197,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Refined image captions injected into web text beat text-only G-SEO, raising visibility by 1.85% (unimodal) and 1.09% (multimodal).","keywords":["caption injection","generative search engine optimization","multimodal retrieval-augmented generation","image captioning","prompt engineering","subjective visibility","G-Eval","MRAMG benchmark"],"falsifier":"Run the same MRAMG benchmark through a multimodal GSE that actually receives raw image pixels (e.g., a multimodal LLM fed image tokens) instead of captions; if the +1.09% relative visibility gain disappears or reverses, the paper's multimodal result is an artifact of its text simulation.","tokens_in":13223,"feed_emoji":"🖼️","tokens_out":20553,"duration_ms":138037,"temperature":0.7,"pith_summary":"This paper argues that the images accompanying a web page carry semantic information that text-only generative-search optimization leaves untapped. The authors propose Caption Injection, a three-stage pipeline that captions an image with a vision-language model, refines that caption by aligning it with the source text's objects, actions, and scene, and then instructs an LLM to insert the refined caption into the source text at the most relevant location. On the MRAMG multimodal benchmark, they report that this injection raises subjective visibility—how often users' attention lands on the content in a generative search answer—by 1.85% relative improvement in a unimodal setting and 1.09% in a multimodal setting, beating all text-only baselines. The paper positions this as the first multimodal G-SEO method and argues that visual semantics, even when projected into text, give web content a measurable edge in generative search.","feed_headline":"Inject image captions to raise AI search visibility 1.85%","feed_subtitle":"First multimodal G-SEO method; injected captions raise how often a source appears in AI answers.","key_machinery":"The central machinery is the Caption Injection pipeline, a pure prompt-engineering chain with three stages. Structural Generation: a vision-language model produces a caption covering object, action, and scene. Alignment Refinement: an LLM rewrites that caption by extracting source-text fragments matched to those three elements and expanding the caption accordingly, while keeping the original caption's syntactic skeleton intact—this is argued to preserve the LLM's learned attention distribution over semantic entities. Semantic Injection: the LLM is prompted to insert only the refined caption into the source text at a contextually optimal location, with strict instructions not to modify or del","core_discovery":"The paper introduces multimodal G-SEO and a method, Caption Injection, that converts images into refined captions and inserts them into source text. The central discovery: a three-stage pipeline—VLM generates an object-action-scene caption, an LLM rewrites it while preserving syntactic structure, then inserts it at the contextually best point—consistently raises relative subjective visibility on the MRAMG benchmark, outperforming all text-only baselines. Reported gains: +1.85% unimodal, +1.09% multimodal. Ablations show the refinement step is necessary, and all methods struggle more in multimodal settings, suggesting cross-modal G-SEO is intrinsically harder.","pith_inferences":["Editorial extension: The paper's multimodal GSE simulation replaces images with captions (text + captions + LLM), so the reported multimodal gain measures caption-mediated transfer only—not reasoning over raw pixels. A real engine that consumes image tokens directly could yield a different benefit, a possibility the paper leaves untested.","Editorial extension: The ablation shows VLM-generated captions can beat refined ones on specific sub-scores (e.g., click-follow likelihood), implying refinement trades off visibility dimensions. A content creator could tune the refinement prompt to favor a particular G-Eval sub-score rather than the average.","Editorial extension: The paper's claim that preserving syntactic structure protects LLM attention predicts that rephrasing the same semantic content with different syntax should reduce the gain; this is a testable separation between semantic alignment and structural preservation that the paper does not perform.","Editorial extension: Because the pipeline is text-based, it could generalize to other modalities (audio, video) that can be transcribed into captions; that generalization is not explored in the paper."],"forward_implications":["Content owners with images can apply this prompt-based pipeline today to improve how often their content is cited or elaborated in generative search answers, with no access to the search engine's internals.","The ablation result—refined captions outperform both raw alt-text and VLM-generated captions—implies that simply having an image is not enough; the caption must be textually aligned with the page's content.","The uniform drop in all methods' improvement from unimodal to multimodal GSE settings indicates that multimodal visibility optimization is a harder problem, motivating deeper cross-modal fusion research.","Near-zero gains on long-text sources (MRAMG-Manual) show that caption injection alone does not address long-context optimization; a separate mechanism is needed for such pages.","Because the effect is measured via G-Eval 2.0's averaged sub-scores, changes in any one visibility dimension could be optimized specifically by reweighting the refinement prompt—though the paper reports only the average."],"fun_headline_variants":["Caption injection lifts AI search visibility by 1.85%","First multimodal G-SEO: inject captions to boost source visibility","Add image captions to text to rise in AI search results","Caption injection improves generative search visibility up to 1.85%","Inject object-action-scene captions to boost GSE visibility"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The paper assumes a multimodal GSE turns images into captions before generating an answer; if a real engine reasons over image pixels directly, the measured multimodal gain may not transfer.","fun_headline_variants_meta":{"raw":{"variants":["Caption injection lifts AI search visibility by 1.85%","First multimodal G-SEO: inject captions to boost source visibility","Add image captions to text to rise in AI search results","Caption injection improves generative search visibility up to 1.85%","Inject object-action-scene captions to boost GSE visibility"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000194,"raw_usage":{"total_tokens":1225,"prompt_tokens":812,"completion_tokens":413,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":556,"completion_tokens_details":{"reasoning_tokens":325}},"tokens_in":556,"tokens_out":413,"duration_ms":3688,"temperature":1.0,"reasoning_tokens":325,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T23:44:53.092355+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same MRAMG benchmark through a multimodal GSE that actually receives raw image pixels (e.g., a multimodal LLM fed image tokens) instead of captions; if the +1.09% relative visibility gain disappears or reverses, the paper's multimodal result is an artifact of its text simulation.","supporting_citations":[],"review_version":1}