{"id":"13d941d1-18b8-4528-82a3-0dbeb42c7fe4","arxiv_id":"2605.22654","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":1,"one_line_summary":"An image-semantic guided method enhances MLLMs for detecting AI-generated modern Chinese poetry by combining poem text with visual representations of content, achieving 85.65% Macro-F1 with Gemini and outperforming text baselines and RoBERTa.","lead":"This paper proposes adding images that capture the meaning, imagery, and feeling of modern Chinese poetry to help multimodal LLMs detect AI-generated poems. A smart generalist might read it to understand a practical way multimodal models can improve detection of synthetic creative writing beyond text alone.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Performance gain may stem from MLLM detecting image-generation artifacts rather than semantic complementarity","rationale":"Reader's weakest assumption directly flags the same risk. Full-text method section would need to specify image provenance and include an artifact-control ablation for the claim to be secure; absent that, the result remains conditional on the images truly adding semantic rather than forensic information.","tokens_in":1693,"tokens_out":317,"duration_ms":19794,"concrete_test":"Reproduce the Gemini experiments on the same poem test set but replace generated images with human-curated real photographs or paintings thematically matched to each poem (sourced from public-domain collections); if Macro-F1 falls more than 8 points or loses statistical significance over the text-only baseline, the headline improvement is artifact-driven.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim requires that images supply independent meaning/imagery/feeling that complements text for better AI-poetry detection. If images are produced by text-to-image models (common for 'reflect the content'), they introduce correlated artifacts (texture inconsistencies, over-smoothing, prompt leakage) that modern MLLMs like Gemini are already tuned to notice in separate image-forensics tasks. The method description does not report image source, generation model, or any control condition that isolates semantic content from these low-level signals. Without that isolation, the 85.65% Macro-F1 and outperformance over RoBERTa could be explained by an unintended multimodal artifact detector rather than the intended image-semantic fusion.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper proposes an image-semantic guided detection method for AI-generated modern Chinese poetry that augments LLM/MLLM prompts with images reflecting poem content (meaning, imagery, feeling). It reports that this multimodal approach yields higher detection performance than text-only LLM baselines and surpasses the RoBERTa baseline, with the Gemini-based detector reaching 85.65% Macro-F1.","tokens_in":1847,"tokens_out":476,"duration_ms":23735,"significance":"If the performance gain is shown to arise from genuine semantic complementarity rather than low-level image artifacts, the work would provide a concrete multimodal technique for detecting generated creative text and could inform future detectors that exploit imagery in poetry and similar domains.","major_comments":[{"comment":"The central experimental claim (abstract and results) that images supply complementary semantic information rests on an unspecified image generation or selection process. No details are given on the text-to-image model, prompt construction, or any control condition that would isolate semantic content from correlated artifacts (texture inconsistencies, prompt leakage). This directly affects interpretability of the 85.65% Macro-F1 and the outperformance over RoBERTa.","section":"Method / Experimental Setup"},{"comment":"Dataset construction is described only at a high level; the manuscript does not report how human-written vs. LLM-generated poems were collected, balanced, or split, nor any statistical significance tests or error analysis that would substantiate the cross-model performance gains.","section":"Experiments"}],"minor_comments":[{"comment":"Notation for the image-semantic fusion step could be clarified with a short pseudocode or diagram to show exactly how image features are combined with text in the MLLM prompt.","section":"Method"},{"comment":"The abstract states that the method 'forms a complementary judgment'; a concrete example of a prompt template and the resulting MLLM output would help readers reproduce the integration step.","section":"Abstract / Method"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a timely topic in Chinese NLP and multimodal detection, but the missing implementation details on image sourcing make it difficult to assess whether the contribution is primarily methodological or artifact-driven; this should be resolved before acceptance."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback, which has helped us identify areas where the manuscript requires greater clarity and detail. We address each major comment below and have prepared revisions to strengthen the experimental description and reproducibility.","responses":[{"response":"We agree that the original manuscript described the image-augmentation process at too high a level, limiting readers' ability to assess whether gains derive from semantic content or from low-level artifacts. In the revised version we have added a dedicated subsection that specifies the text-to-image model, the exact prompt templates used to translate poem semantics into images, and a control experiment that compares performance with semantically faithful images against images generated from shuffled or artifact-only prompts. These additions directly address the interpretability concern raised.","revision_made":"yes","referee_comment":"[Method / Experimental Setup] The central experimental claim (abstract and results) that images supply complementary semantic information rests on an unspecified image generation or selection process. No details are given on the text-to-image model, prompt construction, or any control condition that would isolate semantic content from correlated artifacts (texture inconsistencies, prompt leakage). This directly affects interpretability of the 85.65% Macro-F1 and the outperformance over RoBERTa."},{"response":"We acknowledge that the dataset section was insufficiently detailed. The revised manuscript now contains an expanded data section that reports the sources of human-written poems, the specific LLMs and generation settings used to create the AI poems, the balancing and splitting procedures, and the final dataset sizes. We have also added statistical significance testing (paired t-tests and McNemar's test) for all reported improvements and a concise error-analysis subsection that categorizes the remaining misclassifications.","revision_made":"yes","referee_comment":"[Experiments] Dataset construction is described only at a high level; the manuscript does not report how human-written vs. LLM-generated poems were collected, balanced, or split, nor any statistical significance tests or error analysis that would substantiate the cross-model performance gains."}],"tokens_in":1296,"tokens_out":438,"duration_ms":27975,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main point is that this work shows you can improve AI-poetry detection for modern Chinese verse by feeding both the text and a matching image into models like Gemini. They report the multimodal version beats plain-text LLMs and even tops RoBERTa at 85.65% Macro-F1 across several generated datasets. That is the core empirical result they stand on.","headline":"The paper gets solid gains in detecting AI modern Chinese poetry by adding content-reflecting images to MLLMs, but the gains could easily come from image artifacts instead of real semantic complementarity.","tokens_in":2353,"tokens_out":156,"would_cite":false,"duration_ms":27660,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[{"relation":"unclear","rs_module":"IndisputableMonolith/Cost/FunctionalEquation.lean","rs_theorem":"washburn_uniqueness_aczel","paper_passage":"proposes an image-semantic guided poetry detection method... integrates information such as meaning, imagery, and feeling from the image"},{"relation":"unclear","rs_module":"IndisputableMonolith/Foundation/RealityFromDistinction.lean","rs_theorem":"reality_from_one_distinction","paper_passage":"Gemini detector using our method achieves a Macro-F1 score of 85.65%"}],"headline":"NLP poetry-detection pipeline has no structural overlap with RS distinction-forcing or J-cost machinery","alignment":"orthogonal","rationale":"The paper's core apparatus (MLLM prompting with paired poem+image examples, noun-preserving generation prompts, IP3/IP2 templates, RoBERTa baselines) operates entirely in the domain of multimodal classification and artifact detection. RS theorems (reality_from_one_distinction, Jcost uniqueness via Aczél, phi-ladder constants, 8-tick periodicity, AlexanderDuality D=3 forcing) derive spacetime and recognition costs from a single non-trivial distinction; none of these appear or are paralleled here. The work is therefore orthogonal.","tokens_in":54326,"confidence":"high","tokens_out":306,"duration_ms":10047,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Adding images that reflect poem content improves LLM detection of AI-generated modern Chinese poetry","keywords":["AI generated poetry","modern Chinese poetry","multimodal detection","LLM detectors","image semantic analysis","poetry authenticity"],"falsifier":"Running the image-semantic detector and a plain-text detector on a new collection of human and AI-written modern Chinese poems and finding no accuracy advantage for the image version.","tokens_in":2608,"feed_emoji":"🖼️","tokens_out":587,"duration_ms":42296,"temperature":0.7,"pith_summary":"The paper proposes using images that capture the meaning, imagery, and feeling of modern Chinese poems to help large language models detect whether the poetry was generated by AI. Earlier studies showed LLMs struggle as detectors for AI text in general, but this work focuses specifically on modern Chinese poetry and finds that combining text with semantic images creates a stronger detection signal. Experiments show this image-semantic approach lifts performance above plain-text LLM detectors and even beats the strong traditional baseline RoBERTa, with the Gemini model reaching 85.65 percent Macro-F1. A reader might care because reliable detection supports the integrity of literary traditions as generative AI spreads into poetry writing.","feed_headline":"Images boost LLM poetry detectors past RoBERTa","feed_subtitle":"Pairing poems with content-matching images gives detectors extra clues on imagery and emotion, raising Gemini to 85.65% Macro-F1 on AI verse","key_machinery":"Image-semantic guided poetry detection method that forms complementary judgments from poem text and matching images","core_discovery":"By incorporating images that reflect the content of the poetry, the image-semantic guided method allows LLMs to integrate complementary information on meaning, imagery, and feeling, leading to more accurate detection of AI-generated modern Chinese poetry compared to text-only methods.","pith_inferences":["This technique could extend to detecting AI content in other image-rich creative fields such as visual art descriptions or song lyrics.","Poetry detection might benefit from always generating an illustrative image as a first step before analysis.","Future detectors may need to account for how image generation models themselves introduce patterns that could be exploited or masked."],"forward_implications":["LLM detectors outperform text-only baselines on multiple AI-generated poetry datasets","The Gemini-based detector reaches state-of-the-art Macro-F1 of 85.65%","The method surpasses the best traditional detector RoBERTa","Performance gains are observed across different LLMs"],"fun_headline_variants":["Images aid LLM detection of AI Chinese poetry","Image semantics guide LLM AI poetry checks","Poem images support LLM detectors over text","Content images help LLM identify AI-generated verse","Images integrate semantics for LLM poetry AI detection"],"cache_read_input_tokens":64,"weakest_assumption_plain":"Images can be generated or selected to match the poetry's meaning, imagery, and feeling closely enough to aid detection without introducing their own biases or artifacts.","fun_headline_variants_meta":{"raw":{"variants":["Images aid LLM detection of AI Chinese poetry","Image semantics guide LLM AI poetry checks","Poem images support LLM detectors over text","Content images help LLM identify AI-generated verse","Images integrate semantics for LLM poetry AI detection"]},"model":"grok-4.3","cost_usd":0.010334,"raw_usage":{"total_tokens":4466,"prompt_tokens":611,"num_sources_used":0,"completion_tokens":63,"cost_in_usd_ticks":103340500,"prompt_tokens_details":{"text_tokens":611,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3792,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":611,"tokens_out":63,"duration_ms":39082,"temperature":1.0,"reasoning_tokens":3792,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-22T05:38:16.728683+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Running the image-semantic detector and a plain-text detector on a new collection of human and AI-written modern Chinese poems and finding no accuracy advantage for the image version.","supporting_citations":[],"review_version":1}