REVIEW 3 major objections 2 minor 2 cited by
MAC: A Live Benchmark for Multimodal Large Language Models in Scientific Understanding
T0 review · 3 major / 2 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read Multimodal language models can perceive scientific images but cannot yet reason across them, according to a benchmark built from 25,000 journal cover-caption pairs.
desk verdict The abstract describes a plausible live benchmark for MLLM scientific reasoning, but the supplied full text is a different paper, so the measurement-validity question is the key thing to check. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The benchmark unit is a journal cover paired with its caption, treated as a multimodal question whose answer demands cross-modal reasoning. DAD is the mechanism carrying the claimed improvement: it takes the MLLM's visual features and further processes them in language space before generating an answer, so the model can reason about what it sees rather than merely describe it.
What would settle it
Hand a random sample of MAC-2025 cover-caption pairs to working scientists without an answer key and measure their inter-annotator agreement; if agreement is low, or if swapping a cover while keeping its caption changes model answers far more than human answers, then the benchmark is not primarily measuring stable scientific understanding.
Extended reading notes
Core claim
The central claim is that journal cover-caption pairs can serve as a live benchmark for scientific understanding: each pair is designed so that a correct answer requires connecting abstract visual content to the scientific meaning of the caption. The paper's experiments on MAC-2025 are claimed to show a perception-reasoning gap, with models recognizing elements in scientific images but struggling to reason across image and text. The paper further claims that DAD narrows this gap by extending the model's visual features with additional reasoning in language space, all at inference time and without retraining.
Load-bearing premise
The load-bearing premise is that each journal cover and caption pair contains enough unambiguous, self-contained scientific content that the intended answer is well-defined and represents scientific understanding rather than recognition of famous images or captions.
Editorial extensions
If this is right
- MAC-2025 results separate perceptual ability from cross-modal scientific reasoning, giving future evaluations a baseline for each.
- If DAD's reported gains hold, inference-time language-space reasoning can improve scientific understanding without model retraining or architectural changes.
- The live updating mechanism can keep benchmark content aligned with current research and expose when older fixed benchmarks have saturated.
- The released benchmark allows direct comparisons across model versions and over time, with future yearly snapshots tracking progress.
- The journal-cover source makes the benchmark renewable from the same material stream, so it can grow with scientific publication itself.
Reading between the lines
- I would test the measurement claim directly by giving a sample of MAC questions to working scientists and checking whether their agreement and answer patterns match the ordering implied by model scores.
- DAD's approach may transfer to other visual-scientific tasks such as interpreting figures, charts, and experimental diagrams, not only journal covers.
- Because journal covers are public, future models could memorize them; live updates will need contamination audits to keep scores meaningful.
- If cover art is often decorative metaphor, some MAC questions may be answerable by recognizing famous papers or visual conventions, which would mean the benchmark partly measures cultural recognition rather than scientific reasoning.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The abstract introduces MAC (Multimodal Academic Cover benchmark), a 'live' benchmark of over 25,000 image-text pairs from journal covers of Nature, Science, and Cell, intended to measure MLLMs' scientific understanding. The abstract reports that on the MAC-2025 snapshot, MLLMs show strong perception but limited cross-modal reasoning, and proposes DAD, an inference-time extension that improves performance by up to 11%. The abstract also claims live-update experiments for benchmark refreshment. The full text supplied with this review is not the MAC paper; it is a mojibake rendering of arXiv:2508.15798, a Master's thesis on persuasiveness and bias in LLMs. Thus only the abstract can be evaluated, and the technical content of the claimed benchmark and experiments is absent from this submission.
Significance. If the stated results hold, MAC would be a practically useful continuously refreshable benchmark for multimodal scientific understanding, and DAD would be a lightweight inference-time gain. The dataset scale (25,000+ pairs) and the specific finding that MLLMs are perceptually strong but cross-modally limited are valuable claims. The paper also makes a resource public on GitHub, which is commendable. However, the benchmark's validity depends on item-level evidence that journal cover-caption pairs require genuine cross-modal scientific reasoning rather than text-only inference or memorization. That evidence is not present in the submission. The absence of the actual manuscript means I cannot verify any of the headline numbers, baselines, or the DAD mechanism.
major comments (3)
- [Full text (entire manuscript)] The supplied full text is not the MAC paper. It is a mojibake rendering of arXiv:2508.15798, a Master's thesis on 'Persuasiveness and Bias in LLM.' None of the methods, benchmark construction rules, model details, tables, or figures for MAC/DAD are present. This is load-bearing: the claims in the abstract ('strong perceptual abilities,' 'limited cross-modal reasoning,' 'up to 11%') cannot be checked. The authors must provide the correct full manuscript before this submission can be reviewed.
- [Abstract / Benchmark validity] The central premise is that journal cover-caption pairs 'challeng[e] MLLMs to reason across abstract visual and textual scientific content.' The abstract gives no protocol for item construction, no ground-truth labeling procedure, no human-agreement statistics, and no item-level examples. The stress-test concern is therefore not answerable from this submission: if covers are largely decorative or the captions alone determine the answer, MAC measures a different skill. I request (a) a sample of items with expert annotations showing the answer requires the image, (b) text-only and image-only baselines, and (c) a description of the curation/validation pipeline.
- [Abstract / DAD experimental claim] The abstract reports DAD 'achieving performance improvements of up to 11%,' but provides no baselines, model list, metric definition, number of runs, or statistical tests. 'Up to 11%' could mean one favorable item or a consistent gain; there are no error bars. A full experimental section with variance estimates, model card, and hyperparameter settings for DAD is required before the result can be interpreted.
minor comments (2)
- [Abstract / Live-update experiments] The live-update experiments are mentioned in one sentence with no protocol. If the benchmark is meant to 'continuously evolve,' the update mechanism, contamination controls, and versioning policy need to be described.
- [Abstract / DAD hyperparameters] DAD is described as a 'lightweight inference-time approach,' but no algorithmic details or hyperparameters are given. Even in an abstract, naming the base MLLMs and the extension's computational cost would help.
Circularity Check
No circularity evident from abstract; the supplied full text is a different paper.
full rationale
The abstract introduces a benchmark (MAC) of journal cover–caption pairs and an inference-time method (DAD). No claim in the abstract derives a result from its own definition, and no parameter is fitted and then called a prediction. The provided full text is not the MAC paper (it is a mojibake rendering of arXiv:2508.15798), so no protocol details, label-construction steps, or self-citations are available to inspect. Consequently, no circular step can be quoted or exhibited. Potential validity concerns about whether journal covers measure scientific reasoning are distinct from circularity and would require the actual paper's item-construction details to evaluate.
Assumptions & free parameters
free parameters (1)
- DAD inference-time hyperparameters (unstated)
assumptions (3)
- domain assumption Journal cover images paired with captions are a valid proxy for scientific understanding in MLLMs.
- domain assumption MLLMs evaluated on MAC-2025 have not been contaminated by the test items.
- domain assumption Yearly MAC snapshots remain comparable over time.
invented entities (2)
-
MAC (Multimodal Academic Cover benchmark)
-
DAD (inference-time reasoning extension)
Cite this review
Pith. "Pith review of MAC: A Live Benchmark for Multimodal Large Language Models in Scientific Understanding." pith.science (2026). https://pith.science/paper/DTXPYG3F
@misc{pith2026250815802,
author = {Pith},
title = {Pith review of: MAC: A Live Benchmark for Multimodal Large Language Models in Scientific Understanding},
year = {2026},
howpublished = {\url{https://pith.science/paper/DTXPYG3F}},
note = {Machine review of arXiv:2508.15802}
}
read the original abstract
As multimodal large language models (MLLMs) grow increasingly capable, fixed benchmarks are gradually losing their effectiveness in evaluating high-level scientific understanding. In this paper, we introduce the Multimodal Academic Cover benchmark (MAC), a live benchmark that could continuously evolve with scientific advancement and model progress. MAC leverages over 25,000 image-text pairs sourced from issues of top-tier scientific journals such as Nature, Science, and Cell, challenging MLLMs to reason across abstract visual and textual scientific content. Experiments on our most recent yearly snapshot, MAC-2025, reveal that while MLLMs demonstrate strong perceptual abilities, their cross-modal scientific reasoning remains limited. To bridge this gap, we propose DAD, a lightweight inference-time approach that enhances MLLMs by extending MLLM visual features with language space reasoning, achieving performance improvements of up to 11%. Finally, we highlight the live nature of MAC through experiments on updating journal covers and models for curation, illustrating its potential to remain aligned with the frontier of human knowledge. We release our benchmark at https://github.com/mhjiang0408/MAC_Bench.
Forward citations
Cited by 2 Pith papers
-
An Exam for Active Observers
On a new 17-task benchmark of active visual observation, the best frontier multimodal model solves 10.6% of items and humans solve 96.1%.
-
MatPhaseBench: A Semantics-Guided Benchmark for Materials Phase Diagrams Understanding
A 200-pair, literature-derived phase-diagram benchmark shows 13 VLMs lag expert thermodynamic interpretation, with best BERTScore Recall only 0.407.
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.