{"id":"76e25532-9037-41d5-8ca0-e22cb41b642d","arxiv_id":"2505.17613","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"MMMG introduces 49 evaluation tasks (29 new) and 937 instructions across four modality combinations, reporting 94.3% average agreement between its automated evaluators and human judgments.","lead":"A new benchmark suite called MMMG tests multimodal generation models on 49 tasks spanning image, audio, interleaved text-image, and interleaved text-audio outputs, using automated evaluators that the authors report agree with human raters 94.3% of the time on average.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 94.3% human-agreement headline is an in-sample best over per-task methods and thresholds; without held-out validation, the central reliability claim is not yet established.","rationale":"MMMG is a useful benchmark, and the human evaluation is substantial: 1886 questions, three annotators per question, and high inter-annotator agreement. The load-bearing issue is statistical: the headline agreement is the maximum over per-task methods and thresholds tuned on the same data. This is the reader's weakest assumption, and I agree with that identification. The selection is visible in Section 5.1, which describes selecting the method achieving the highest agreement per task, in Appendix C.3 thresholds, and in Section 5.2 where the most-aligned methods are used for benchmarking. Without a held-out split, 0.943 estimates the best of several candidate pipelines, not the expected agreement of the chosen pipeline on new data. This matters because the benchmark's purpose is to rank models; if the chosen evaluator overfits the 674 instructions, model rankings on the full 937 instructions and on future outputs can be biased. The 37-of-49 task validation gap is secondary but reinforces the need for out-of-sample evidence. I do not see an internal inconsistency or a reason to reject; the limitations section acknowledges coverage and proprietary-model dependence. The appropriate disposition remains CONDITIONAL: accept if the authors provide cross-validated or fixed-protocol human-alignment numbers, or qualify the headline as an in-sample best. Since the reader already reached CONDITIONAL, my verdict is UNCHANGED.","tokens_in":36649,"tokens_out":4869,"duration_ms":39271,"concrete_test":"Run a task-level or instruction-level cross-validation over the 674 human-validated instructions: in each fold, select the per-task evaluation method and thresholds using only the training instructions, then compute mean agreement on held-out instructions using the same majority-vote protocol. If the held-out average falls materially below 0.943, for example by more than 0.03, the headline overstates generalizable alignment. Independently, freeze the final per-task methods and collect new human judgments on a sample of the 12 currently unvalidated tasks; if agreement there is low, the full-benchmark reliability claim needs qualification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that MMMG's automated evaluation is highly aligned with human judgment, at 0.943 average agreement (Section 5.1, Table 8). The support is computed by selecting, for each task, the evaluation method with the highest human agreement on the same 674 instructions used to report the number, and by tuning thresholds on those same judgments (CLAP 0.68 for ESC-50, 0.62 for OpenMIC-2018, WavLM 0.86 for speaker similarity, Appendix C.3). This is an in-sample maximum, not the agreement of a pre-specified evaluation protocol. Because the final benchmark (Section 5.2) then uses these most-aligned methods, any selection noise or threshold overfitting directly transfers to the reported model rankings. The paper does not report a held-out split, confidence intervals on the 0.943, or a comparison of the selected method's agreement against the average or median method. Additionally, only 37 of 49 tasks are human-validated; for the remaining 12 tasks the reliability of the selected metric is assumed rather than demonstrated. The concern is not that the human labels are noisy (inter-annotator agreement is high, 0.971), but that the headline number is optimistically selected and may not generalize to unseen instructions, model outputs, or tasks.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces MMMG, a benchmark and automated evaluation suite for multimodal generation covering image, audio, interleaved image-text, and interleaved audio-text tasks. It comprises 49 tasks (29 new), 937 instructions, and four modality combinations, with each task assigned a program-, model-, or hybrid-based evaluation pipeline. The authors report a human-model agreement of 94.3% based on human evaluation of 674 instructions and 1886 evaluation questions, and benchmark 24 generation models, finding that GPT-Image leads image generation but lags on multimodal reasoning and interleaved generation. Code and data are released publicly.","tokens_in":36962,"tokens_out":3021,"duration_ms":23513,"significance":"If the reliability claim holds, MMMG would be a valuable contribution: it is substantially broader than existing benchmarks (Table 1), it combines verifiable program checks with model-based evaluation for hard generation tasks, and it ships a large human-annotation study (1886 questions, three annotations per item, 97.1% inter-annotator agreement). The public release and the careful design of VLM prompts (e.g., negative prompts, multiple-choice formats, chain-of-thought) are praiseworthy. However, the headline 94.3% agreement is computed by selecting, per task, the evaluator with the highest human agreement on the same validation data used to report the number, with thresholds also tuned on those judgments. This makes the central reliability claim an in-sample best fit rather than a predictive estimate, and it is the main load-bearing weakness of the paper.","major_comments":[{"comment":"The reported 0.943 agreement is the mean of per-task maximum agreement over the candidate evaluators (GPT-4O, GEMINI2.5, QWEN2.5-VL for images; CLAPScore-audio, CLAPScore-text, GEMINI2.5 for audio; WavLM, Wav2Vec for speech), where the maximum is taken on the same 674 instructions that are then used to report the agreement. This is an in-sample selection; it does not estimate the agreement of a fixed evaluation protocol. The paper should report a held-out estimate, e.g., leave-one-task-out or a split of instructions, and should also report the agreement of a pre-specified default method (e.g., the average across evaluators or the method chosen by a fixed rule). Without such an estimate, the 94.3% headline and the subsequent claim that MMMG is 'highly aligned with human evaluation' are not yet established for unseen instructions and model outputs.","section":"§5.1 and Table 8"},{"comment":"The thresholds used in the audio evaluation (CLAP similarity threshold 0.68 for ESC-50, 0.62 for OpenMIC-2018, and WavLM speaker-similarity threshold 0.86) are described as 'empirical optimal' and were tuned on the same human judgments that define the agreement numbers. This additional selection on the validation set further inflates the reported agreement. The authors should either tune thresholds on a separate development split and evaluate on a held-out split, or report the sensitivity of task-level agreement to these thresholds; otherwise the reported 92.6% audio agreement and the WavLM-based speech results may be optimistically biased.","section":"§4 and Appendix C.3"},{"comment":"Human validation covers only 37 of the 49 tasks; for the remaining 12 tasks the reliability of the selected metric is assumed rather than demonstrated. The paper justifies this by saying that verifiable instructions do not require human validation, but several of the unvalidated tasks (e.g., Border Fill, Region Fill, Music Tempo, Music Intensity) rely on programmatic checks with hand-set thresholds (e.g., the 15% color-deviation threshold and the BPM/intensity rules in Appendix C.3). The manuscript should explicitly list which tasks lack human validation and provide a rationale or a small-scale sanity check for those program-based metrics, or restrict the reliability claim to the 37 validated tasks.","section":"§3.1 and §5.1"},{"comment":"The correlation with Chatbot Arena is computed on only 7 models. While Spearman's 0.857 is suggestive, the small n makes the estimate fragile and no confidence interval or significance test is reported. The claim that MMMG 'provides reliable model rankings' would be strengthened by reporting a confidence interval or bootstrap estimate, and by acknowledging the small-sample caveat in the text.","section":"§5.3, Table 3"}],"minor_comments":[{"comment":"The phrase 'average agreement of 94.3%' should be qualified as 'average best human-model agreement after per-task method selection' to avoid implying a single pre-specified protocol achieves this value.","section":"Abstract and §5.1"},{"comment":"The table's use of symbols (⊷, /volume-down, T+⊷, T+/volume-down) is not intuitive; a legend or explicit expansion in the caption would improve readability.","section":"§2, Table 1"},{"comment":"The sentence 'The average inter-annotator agreement remains as high as 0.971 with the worst case being 0.917' is slightly ambiguous; clarify whether the worst-case is across tasks or across annotator pairs.","section":"§5.1"},{"comment":"The per-task agreement and correlation values are presented without standard errors or confidence intervals; adding a small-sample caveat would help readers interpret differences among evaluators (e.g., GPT-4O vs GEMINI2.5).","section":"Table 8"},{"comment":"The description of the solid-color-fill program mentions a 'relative deviation exceeds 15%' threshold, but the exact comparison metric (e.g., Euclidean distance in RGB or relative to the reference color) is not defined; please specify the formula.","section":"Appendix C.3"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid benchmark contribution with a large, carefully conducted human study and a public release, but the central reliability claim rests on an in-sample per-task maximum and tuned thresholds. I do not doubt the good faith of the authors, but the 94.3% figure is likely optimistic and should be re-estimated on a held-out basis before the benchmark's reliability is advertised. The revision is feasible within the paper's scope by adding a validation split or sensitivity analysis. The fit with the journal's scope is good."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a real contribution, but the headline 94.3% human-model agreement is a fitted maximum. It is computed by selecting, per task, the evaluator that agrees best with humans on the validation set, and thresholds (CLAP 0.68/0.62, WavLM 0.86) are tuned on those same judgments. That is not the agreement of a pre-specified protocol, and the paper reports no held-out split, no confidence intervals on the 0.943, and no comparison against the median method. The concern is real, and it transfers directly to the Section 5.2 rankings, which use the same selected methods.\n\nThe underlying work is solid and genuinely useful. MMMG is the first benchmark to cover image, audio, interleaved image-text, and interleaved audio-text generation with human-validated evaluation. 29 of 49 tasks are new, and the per-capability breakdown is more granular than prior benchmarks. The human validation effort is substantial: 1886 questions, three annotators each, 97.1% inter-annotator agreement. The Chatbot Arena correlation (Spearman 0.857) is suggestive even with n=7. Code and data are public.\n\nSoft spots beyond the selection issue: only 37 of 49 tasks are human-validated; the 12 remaining tasks are programmatic and assumed reliable, which is probably fine but left implicit. Evaluation relies heavily on proprietary models (GPT-4O, Gemini), acknowledged in the appendix, which limits reproducibility. The GPT-4O math agreement correlation is low (0.436), but the paper reports it honestly and picks a different evaluator for that task.\n\nThe central claim—that MMMG is more comprehensive and more human-aligned than previous suites—is defensible, but the 94.3% is optimistic. The fix is straightforward: pre-register the evaluation protocol, hold out a validation split, or at least report the agreement of a fixed method with confidence intervals.\n\nThis paper deserves a serious referee. The benchmark will likely be used regardless, and the validation gap is addressable in revision. I would recommend conditional acceptance after out-of-sample agreement or a fixed protocol is shown.","headline":"Solid benchmark with an in-sample agreement headline; deserves review but needs out-of-sample validation.","tokens_in":37526,"tokens_out":2860,"would_cite":true,"duration_ms":25685,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A benchmark claims automated grading of multimodal generation matches human judges 94.3 percent of the time.","keywords":["multimodal generation benchmark","automated evaluation","human alignment","interleaved image-text generation","interleaved audio-text generation","image generation","audio generation","instruction following"],"falsifier":"Take a fresh set of instructions and generations from models not among the original 24, score them with the paper's per-task methods and thresholds, and collect the same three-annotator majority judgments; if human-model agreement falls substantially below 0.943, the reported alignment is specific to the validation set rather than a general property of the suite.","tokens_in":36465,"feed_emoji":"🧪","tokens_out":5997,"duration_ms":42243,"temperature":0.7,"pith_summary":"MMMG is a benchmark for multimodal generation built on a specific bet: that automated evaluation of images, audio, and interleaved text-image or text-audio outputs can be made nearly as trustworthy as human evaluation, without expensive annotation, if tasks are chosen carefully. The paper claims 94.3% average agreement between the best automated method per task and human annotators across 37 tasks, based on 674 instructions and 1,886 evaluation questions. If that holds, model rankings in multimodal generation no longer need to rely on costly human studies or on unvalidated model-as-a-judge scores. The paper also benchmarks 24 models and finds the strongest image generator still weak at multimodal reasoning and interleaved generation, with audio generation showing much headroom.","feed_headline":"Automated multimodal grading matches humans 94.3%","feed_subtitle":"A 49-task suite ranks 24 image, audio, and interleaved generation models with little human review.","key_machinery":"The load-bearing mechanism is the per-task evaluation pipeline, with the method for each task chosen to maximize human agreement. For image tasks, manually designed VQA prompts push a vision-language model to reason step by step, answer as multiple choice, or reject negative prompts. For audio, CLAP similarity against reference audio, program checks for tempo and silence, and Whisper/WavLM-based transcript or speaker checks are used. Programmatic checks, such as pixel-level border verification, cover objectively verifiable tasks. The paper's methodological move is selecting one method per task based on validation-set human agreement and then applying it to all models.","core_discovery":"The central claim is that a benchmark built from verifiable tasks and tasks with a large generation-evaluation gap can be evaluated automatically with high human agreement. Across image, audio, interleaved image-text, and interleaved audio-text generation, the per-task evaluation method that best agrees with human raters averages 0.943 agreement, while inter-annotator agreement averages 0.971. The suite contains 49 tasks (29 newly developed) and 937 instructions. Benchmarking 24 models with these methods shows GPT Image leading image generation at 78.3% accuracy but scoring only 13.1% on interleaved math and code reasoning, while the best sound and music models reach 48.7% and 41.9% accuracy.","pith_inferences":["The paper does not state this, but the per-task method selection and thresholds were tuned on the same human judgments used to measure agreement, so a held-out validation split could reveal optimistic bias in the 94.3% figure.","If evaluator alignment outweighs instruction distribution, benchmark designers could generate synthetic instructions freely and focus annotation budget on evaluator validation.","The same verifiable-task design could extend to video generation, where currently no comparable human-aligned automated benchmark exists.","The dependence on proprietary graders means benchmark scores may shift as those graders update, so reproducing the benchmark requires caching model outputs."],"forward_implications":["Automated rankings on MMMG can be updated cheaply as new multimodal models appear, without rerunning human studies.","The benchmark's fine-grained 49-task breakdown lets a developer see whether a model fails at counting, spatial reasoning, text rendering, or audio-level control.","The reported gaps, such as GPT Image reaching 78.3% on image tasks but 13.1% on interleaved math and code, suggest instruction-following and multimodal reasoning are separate capabilities that need separate benchmarks.","MMMG's 0.857 Spearman correlation with a human-preference leaderboard indicates that evaluator alignment can matter more than instruction distribution matching."],"supporting_citations":[{"why":"Defines verifiable instructions, the criterion MMMG uses to admit programmatically checkable tasks.","marker":"[Zhou et al., 2023]"},{"why":"Provides the GenEval baseline whose 0.830 human agreement MMMG surpasses on image generation.","marker":"[Ghosh et al., 2023]"},{"why":"ISG-Bench is the interleaved image-text baseline whose Pearson correlation MMMG improves by 28.1%.","marker":"[Chen et al., 2025a]"},{"why":"TIFA motivates the VQA-based approach and is cited as evidence that unvalidated question-answer checks misalign with humans.","marker":"[Hu et al., 2023]"},{"why":"Chatbot Arena supplies the human-preference leaderboard against which MMMG reports 0.857 Spearman correlation.","marker":"[Chiang et al., 2024]"},{"why":"DreamSim is the perceptual similarity model used for reference-based image editing and interleaved editing evaluation.","marker":"[Fu et al., 2023]"},{"why":"CLAP embeddings provide the audio similarity score used for sound and music tasks with validation thresholds.","marker":"[Wu et al., 2023]"},{"why":"Whisper supplies transcripts for speech text constraints and editing checks.","marker":"[Radford et al., 2023]"},{"why":"WavLM provides speaker-similarity verification with the 0.86 threshold for voice replication tasks.","marker":"[Chen et al., 2022]"}],"fun_headline_variants":["Multimodal benchmark matches human judgment 94.3%","49-task suite ranks 24 multimodal generators near-human","GPT Image tops image gen but flops interleaved reasoning","Audio generation still lags in new 49-task benchmark","MMMG: 94.3% human agreement on multimodal generation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"For every task, the evaluation method and numeric thresholds that best matched human ratings on a validation sample will keep matching human ratings on the full benchmark and on outputs from models and instructions not seen during selection.","fun_headline_variants_meta":{"raw":{"variants":["Multimodal benchmark matches human judgment 94.3%","49-task suite ranks 24 multimodal generators near-human","GPT Image tops image gen but flops interleaved reasoning","Audio generation still lags in new 49-task benchmark","MMMG: 94.3% human agreement on multimodal generation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000166,"raw_usage":{"total_tokens":1236,"prompt_tokens":907,"completion_tokens":329,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":523,"completion_tokens_details":{"reasoning_tokens":245}},"tokens_in":523,"tokens_out":329,"duration_ms":3943,"temperature":1.0,"reasoning_tokens":245,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T14:43:33.550948+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a fresh set of instructions and generations from models not among the original 24, score them with the paper's per-task methods and thresholds, and collect the same three-annotator majority judgments; if human-model agreement falls substantially below 0.943, the reported alignment is specific to the validation set rather than a general property of the suite.","supporting_citations":[],"review_version":1}