{"id":"e2fae8e6-11b9-4346-acbf-f34a4aaf129b","arxiv_id":"2508.15802","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":1,"one_line_summary":"MAC is a self-updating benchmark of scientific journal cover image-caption pairs, on which multimodal models show strong perception but weak cross-modal reasoning, improved by up to 11% with the DAD inference method.","lead":"This paper introduces MAC, a live benchmark that tests multimodal AI models on over 25,000 image-caption pairs from journal covers such as Nature and Science, plus a lightweight inference method that lifts accuracy by up to 11%. A smart generalist might read it because fixed AI benchmarks saturate and leak into training data, and MAC is designed to keep re-measuring scientific reasoning against freshly published research.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Measurement validity of journal cover-caption pairs is unestablished: if items are decorative or answerable from text alone, MAC scores and DAD gains measure a different skill.","rationale":"The reader's weakest_assumption is that journal cover-caption pairs are a valid instrument for measuring scientific understanding. I agree: this is the most load-bearing concern about the central claim. Even if the paper's body were fully available and experiments replicable, the benchmark's scientific value hinges on whether correct answers require genuine cross-modal reasoning. The abstract explicitly states that MLLMs have 'strong perceptual abilities' but 'cross-modal scientific reasoning remains limited,' which is only meaningful if the items actually demand such reasoning. The 'up to 11%' DAD gain is a relative improvement on a metric; if the metric is invalid, the gain is meaningless. My concrete test—an expert annotation study—would directly settle whether covers are decorative or informative. I do not recommend changing the reader's UNVERDICTED verdict: the supplied full text is a different document, so no independent verification of the paper's existence or content is possible, but this measurement-validity concern would remain even if the full text were recovered. Hence UNCHANGED is appropriate.","tokens_in":59360,"tokens_out":3295,"duration_ms":37825,"concrete_test":"Run an expert human annotation study on a random sample of 200 MAC-2025 pairs. For each pair, have three domain experts independently answer: (1) Does the image contribute information beyond the caption? (2) Is the correct answer uniquely determined by the image-plus-caption content? (3) Does solving require scientific reasoning (e.g., inference from visual metaphor, quantitative reasoning) rather than recognizing a famous cover or matching keywords? If fewer than 80% of items are judged to require cross-modal reasoning, the benchmark's construct validity fails, and DAD's gain cannot be attributed to improved scientific understanding.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central claim is that MAC 'challeng[es] MLLMs to reason across abstract visual and textual scientific content.' This presumes each cover-caption pair contains enough unambiguous, groundable scientific signal for a well-defined correct answer that requires cross-modal reasoning. If covers are largely decorative, symbolic, or recognizable via memorization of famous papers, then both the MAC scores and the reported 'up to 11%' DAD improvement are not evidence about scientific understanding. The supplied full text is not the MAC paper (it is a mojibake rendering of arXiv:2508.15798), so no protocol details, item-construction rules, or human-agreement statistics are available to check this premise. This measurement-validity issue is load-bearing because the entire benchmark's contribution—and the DAD inference-time gain—depends on it. Without expert validation that the items require genuine scientific reasoning across image and text, the headline results are uninterpretable.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The abstract introduces MAC (Multimodal Academic Cover benchmark), a 'live' benchmark of over 25,000 image-text pairs from journal covers of Nature, Science, and Cell, intended to measure MLLMs' scientific understanding. The abstract reports that on the MAC-2025 snapshot, MLLMs show strong perception but limited cross-modal reasoning, and proposes DAD, an inference-time extension that improves performance by up to 11%. The abstract also claims live-update experiments for benchmark refreshment. The full text supplied with this review is not the MAC paper; it is a mojibake rendering of arXiv:2508.15798, a Master's thesis on persuasiveness and bias in LLMs. Thus only the abstract can be evaluated, and the technical content of the claimed benchmark and experiments is absent from this submission.","tokens_in":59580,"tokens_out":3151,"duration_ms":33826,"significance":"If the stated results hold, MAC would be a practically useful continuously refreshable benchmark for multimodal scientific understanding, and DAD would be a lightweight inference-time gain. The dataset scale (25,000+ pairs) and the specific finding that MLLMs are perceptually strong but cross-modally limited are valuable claims. The paper also makes a resource public on GitHub, which is commendable. However, the benchmark's validity depends on item-level evidence that journal cover-caption pairs require genuine cross-modal scientific reasoning rather than text-only inference or memorization. That evidence is not present in the submission. The absence of the actual manuscript means I cannot verify any of the headline numbers, baselines, or the DAD mechanism.","major_comments":[{"comment":"The supplied full text is not the MAC paper. It is a mojibake rendering of arXiv:2508.15798, a Master's thesis on 'Persuasiveness and Bias in LLM.' None of the methods, benchmark construction rules, model details, tables, or figures for MAC/DAD are present. This is load-bearing: the claims in the abstract ('strong perceptual abilities,' 'limited cross-modal reasoning,' 'up to 11%') cannot be checked. The authors must provide the correct full manuscript before this submission can be reviewed.","section":"Full text (entire manuscript)"},{"comment":"The central premise is that journal cover-caption pairs 'challeng[e] MLLMs to reason across abstract visual and textual scientific content.' The abstract gives no protocol for item construction, no ground-truth labeling procedure, no human-agreement statistics, and no item-level examples. The stress-test concern is therefore not answerable from this submission: if covers are largely decorative or the captions alone determine the answer, MAC measures a different skill. I request (a) a sample of items with expert annotations showing the answer requires the image, (b) text-only and image-only baselines, and (c) a description of the curation/validation pipeline.","section":"Abstract / Benchmark validity"},{"comment":"The abstract reports DAD 'achieving performance improvements of up to 11%,' but provides no baselines, model list, metric definition, number of runs, or statistical tests. 'Up to 11%' could mean one favorable item or a consistent gain; there are no error bars. A full experimental section with variance estimates, model card, and hyperparameter settings for DAD is required before the result can be interpreted.","section":"Abstract / DAD experimental claim"}],"minor_comments":[{"comment":"The live-update experiments are mentioned in one sentence with no protocol. If the benchmark is meant to 'continuously evolve,' the update mechanism, contamination controls, and versioning policy need to be described.","section":"Abstract / Live-update experiments"},{"comment":"DAD is described as a 'lightweight inference-time approach,' but no algorithmic details or hyperparameters are given. Even in an abstract, naming the base MLLMs and the extension's computational cost would help.","section":"Abstract / DAD hyperparameters"}],"recommendation":"major_revision","confidential_remarks":"The full text is a different paper; I could not review the actual MAC content. Please verify the submission integrity and ask the authors to resubmit the correct manuscript. The abstract's claims are plausible but entirely unverifiable from this package."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The only thing in front of us worth reviewing is the abstract: the supplied full text is a mojibake rendering of an unrelated master's thesis. So everything below is provisional, and I'll keep my confidence low.\n\nWhat the abstract does well: it picks a novel data source (journal covers) for a self-refreshing benchmark, targets the benchmark-saturation problem directly, and the GitHub release is a concrete artifact. The MAC-2025 snapshot with the finding that MLLMs perceive well but reason poorly across modalities is a plausible, useful result if it holds, and DAD is a lightweight inference-time method with a claimed gain.\n\nThe soft spots are mostly things the abstract cannot answer. First, the measurement-validity concern is real and load-bearing: if journal covers are largely decorative, or captions alone carry the answer, then MAC scores measure something other than scientific multimodal reasoning. The abstract does not mention how items are labeled, whether there is human validation, or how they prevent text-only solvability. That is the single most important question to ask the authors. Second, 'up to 11%' is a maximum, not a central estimate; the abstract gives no baselines, error bars, or statistical tests. Third, there is no comparison with existing live benchmarks or inference-time reasoning methods, so the novelty is not yet anchored. I do not see circularity in the abstract, though the usual risk is that if labels come from LLM generation and validation, the benchmark partly measures agreement with the labeling model.\n\nOverall: this is a paper for MLLM evaluation folks. If the body matches the abstract and includes item-level validation and a fair baseline set, it deserves publication. Based on the abstract alone, I would not desk-reject it. Send it to peer review with a request for the label-construction and validation details, and flag the measurement-validity concern in the review. Also, the PDF needs fixing—the current full-text file is the wrong paper. Recommendation: accept for peer review; I'd want to see the actual body before writing anything stronger.","headline":"The abstract describes a plausible live benchmark for MLLM scientific reasoning, but the supplied full text is a different paper, so the measurement-validity question is the key thing to check.","tokens_in":60072,"tokens_out":2657,"would_cite":false,"duration_ms":29692,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Multimodal language models can perceive scientific images but cannot yet reason across them, according to a benchmark built from 25,000 journal cover-caption pairs.","keywords":["multimodal large language models","scientific understanding","live benchmark","journal covers","cross-modal reasoning","inference-time reasoning","visual question answering","benchmark evaluation"],"falsifier":"Hand a random sample of MAC-2025 cover-caption pairs to working scientists without an answer key and measure their inter-annotator agreement; if agreement is low, or if swapping a cover while keeping its caption changes model answers far more than human answers, then the benchmark is not primarily measuring stable scientific understanding.","tokens_in":59233,"feed_emoji":"🧪","tokens_out":3205,"duration_ms":39828,"temperature":0.7,"pith_summary":"The paper introduces MAC, a benchmark built from over 25,000 image-text pairs drawn from cover art and captions of journals such as Nature, Science, and Cell, and argues it measures how well multimodal large language models reason across abstract visual and textual scientific content. On the MAC-2025 snapshot, the paper reports that current models see scientific imagery well but cannot reliably integrate what they see with scientific reasoning. To close that gap, the paper proposes DAD, a lightweight inference-time method that adds language-space reasoning to visual features, reporting improvements of up to 11 percent. If right, MAC offers a continuously refreshable measure of cross-modal scientific understanding that tracks scientific progress rather than going stale.","feed_headline":"Journal covers show AI can see science but not reason about it","feed_subtitle":"Multimodal models see journal covers well but reason across them poorly, and a cheap inference-time trick helps.","key_machinery":"The benchmark unit is a journal cover paired with its caption, treated as a multimodal question whose answer demands cross-modal reasoning. DAD is the mechanism carrying the claimed improvement: it takes the MLLM's visual features and further processes them in language space before generating an answer, so the model can reason about what it sees rather than merely describe it.","core_discovery":"The central claim is that journal cover-caption pairs can serve as a live benchmark for scientific understanding: each pair is designed so that a correct answer requires connecting abstract visual content to the scientific meaning of the caption. The paper's experiments on MAC-2025 are claimed to show a perception-reasoning gap, with models recognizing elements in scientific images but struggling to reason across image and text. The paper further claims that DAD narrows this gap by extending the model's visual features with additional reasoning in language space, all at inference time and without retraining.","pith_inferences":["I would test the measurement claim directly by giving a sample of MAC questions to working scientists and checking whether their agreement and answer patterns match the ordering implied by model scores.","DAD's approach may transfer to other visual-scientific tasks such as interpreting figures, charts, and experimental diagrams, not only journal covers.","Because journal covers are public, future models could memorize them; live updates will need contamination audits to keep scores meaningful.","If cover art is often decorative metaphor, some MAC questions may be answerable by recognizing famous papers or visual conventions, which would mean the benchmark partly measures cultural recognition rather than scientific reasoning."],"forward_implications":["MAC-2025 results separate perceptual ability from cross-modal scientific reasoning, giving future evaluations a baseline for each.","If DAD's reported gains hold, inference-time language-space reasoning can improve scientific understanding without model retraining or architectural changes.","The live updating mechanism can keep benchmark content aligned with current research and expose when older fixed benchmarks have saturated.","The released benchmark allows direct comparisons across model versions and over time, with future yearly snapshots tracking progress.","The journal-cover source makes the benchmark renewable from the same material stream, so it can grow with scientific publication itself."],"supporting_citations":[],"fun_headline_variants":["AI sees journal covers, struggles to reason about science","New live benchmark exposes AI's science reasoning gap","Inference-time trick boosts AI science reasoning by 11%","Journal cover pairs reveal AI perception vs reasoning gap","Live benchmark: AI reads covers but can't connect them"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The load-bearing premise is that each journal cover and caption pair contains enough unambiguous, self-contained scientific content that the intended answer is well-defined and represents scientific understanding rather than recognition of famous images or captions.","fun_headline_variants_meta":{"raw":{"variants":["AI sees journal covers, struggles to reason about science","New live benchmark exposes AI's science reasoning gap","Inference-time trick boosts AI science reasoning by 11%","Journal cover pairs reveal AI perception vs reasoning gap","Live benchmark: AI reads covers but can't connect them"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000626,"raw_usage":{"total_tokens":2707,"prompt_tokens":695,"completion_tokens":2012,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":439,"completion_tokens_details":{"reasoning_tokens":1945}},"tokens_in":439,"tokens_out":2012,"duration_ms":11851,"temperature":1.0,"reasoning_tokens":1945,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T20:34:01.707916+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Hand a random sample of MAC-2025 cover-caption pairs to working scientists without an answer key and measure their inter-annotator agreement; if agreement is low, or if swapping a cover while keeping its caption changes model answers far more than human answers, then the benchmark is not primarily measuring stable scientific understanding.","supporting_citations":[],"review_version":1}