{"id":"9ee072f3-9d3e-4f2f-bdfa-8716ecbd0d9a","arxiv_id":"2506.12958","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A review paper that organizes domain-specific MLLM benchmarks into an eight-discipline taxonomy, with summary tables and performance highlights.","lead":"This paper surveys domain-specific benchmarks for multimodal large language models across eight disciplines, categorizing many evaluation tools and summarizing model performance. It offers a map for researchers and practitioners who need to choose or design benchmarks for specialized applications.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Coverage is unfalsifiable as reported: §2's search strategy lacks screening, deduplication, and selection criteria, and the tables include non-benchmarks (ToT) and general benchmarks (BIG-bench) while the abstract says seven disciplines and §2 says eight.","rationale":"The reader's weakest assumption correctly identifies search strategy and coverage as the load-bearing premise, and I agree that this is where the central claim is least secure. I add a more concrete failure mode that the reader did not emphasize: the inclusion criteria, even if the search were exhaustive, are not applied consistently. The presence of general benchmarks (BIG-bench), methods (ToT), and model families (KOSMOS, LLaVA, ChatGLM) in tables that are supposed to catalogue domain-specific benchmarks means the taxonomy itself is partially miscategorized. This is not a fatal flaw for the survey's usefulness as a starting bibliography, but it is a correctness risk in the 'accessible resource' contribution. The seven-versus-eight discrepancy in the abstract versus §2 is a smaller symptom of the same issue: the cataloguing discipline is not yet reliable enough to support the word 'comprehensive.' I do not think this warrants rejection; a conditional acceptance with a required revision to the methodology and table audits is appropriate, which matches the reader's verdict. Hence verdict_should_be is UNCHANGED.","tokens_in":41583,"tokens_out":3345,"duration_ms":45552,"concrete_test":"Implement the §2 search as an auditable protocol: define explicit inclusion/exclusion criteria, record total hits screened per database and per discipline, log excluded items with reasons, and then check every row of Tables 1-8 against the criteria (benchmark, domain-specific, MLLM-relevant). If more than a small fraction of rows fail scope, or if a held-out set of well-known domain-specific MLLM benchmarks per discipline has recall below a stated threshold, the comprehensive-coverage claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the paper provides a comprehensive, domain-specific taxonomy of MLLM benchmarks across eight disciplines. The load-bearing condition is that the search and inclusion process in §2 actually captures the relevant landscape and that every catalogued item is in scope. As written, this condition is not met in a checkable way. The Search Strategy and Scope paragraph lists four sources and a few generic keywords, but gives no screening criteria, deduplication procedure, date range, number of records retrieved, or exclusion log. That makes coverage non-falsifiable: a reader cannot tell whether the taxonomy is representative or merely a convenience sample. The problem is visible internally. The abstract states a taxonomy of seven disciplines, while §2 and Figure 1 present eight. More importantly, several Table entries fall outside the paper's own stated scope of domain-specific benchmarks. Table 1 places BIG-bench under Software Engineering/Knowledge Graphs and Semantic Systems, but BIG-bench is a general LLM benchmark, not a domain-specific software-engineering benchmark. Table 8 lists ToT, a prompting method, and LongLLaVA, LLaVA-OneVision, KOSMOS-1, and ChatGLM, which are models or methods rather than benchmarks. If the scope filter is applied this loosely, the 'domain-specific benchmark' categorization is unreliable even where retrieval was thorough. The weakest link is therefore not the keyword list itself; it is the absence of a reproducible inclusion protocol and the visible scope drift in the summary tables.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper is a survey of domain-specific benchmarks for multimodal large language models (MLLMs). It proposes a taxonomy of disciplines—Engineering, Science, Technology, Mathematics, Humanities, Finance, Healthcare, and Language Understanding—and, for each, provides summary tables with scale, task type, input modality, models, performance, and key focus, along with narrative discussion of trends and limitations. The stated contribution is a comprehensive, accessible resource that maps the domain-specific evaluation landscape and highlights where current MLLMs succeed or fail in specialized fields.","tokens_in":41870,"tokens_out":3519,"duration_ms":45246,"significance":"If the survey's scope and categorization were reliable, it would be a useful entry point for researchers seeking domain-specific evaluation resources. The paper's strength is its breadth: it organizes a large number of benchmarks and studies across eight disciplines and identifies recurring gaps (e.g., in medical image reasoning and in modality-dependent performance). However, the value of the survey depends entirely on the reproducibility of its selection process and the accuracy of its categorizations; the current inconsistencies substantially weaken the central claim of comprehensiveness. The paper does not provide machine-checked proofs or reproducible code; its contribution is the structured synthesis itself.","major_comments":[{"comment":"The abstract states that the paper introduces 'a taxonomy of seven key disciplines,' while Section 2 explicitly lists eight disciplines and Figure 1 displays eight. This is a direct internal contradiction about the scope of the survey. Since the paper's central claim is that it provides a comprehensive taxonomy, the intended number of disciplines must be clarified; as written, a reader cannot tell whether one section is extraneous or the abstract is wrong.","section":"§1 (Abstract) vs §2"},{"comment":"The search strategy paragraph names databases and general keyword categories but provides no screening criteria, deduplication procedure, date range, number of records retrieved, number of records screened, or exclusion log. The paper claims to be a 'systematic review' and concludes that its resource is comprehensive, but without these details the coverage is not checkable or reproducible. The claim that the review 'examines eight key disciplines' and that the tables consolidate 'relevant benchmarks and survey papers' is therefore not falsifiable; the authors should add a principled inclusion/exclusion protocol and a flow diagram or equivalent transparency measure.","section":"§2, Search Strategy and Scope"},{"comment":"Several entries in the summary tables are not domain-specific benchmarks, which contradicts the paper's stated scope. Table 1 places BIG-bench under Software Engineering/Knowledge Graphs & Semantic Systems, but BIG-bench is a general-purpose benchmark spanning 204 tasks across many areas. Table 2 includes WeatherBench 2, a weather forecasting benchmark that is not an LLM or MLLM evaluation benchmark. Table 8 lists ToT (a prompting method), LongLLaVA, LLaVA-OneVision, KOSMOS-1, KOSMOS-2, and ChatGLM (models or architectures) as if they were benchmarks. These inclusions make the consolidated resource unreliable and undermine the central claim that the paper catalogs domain-specific benchmarks for evaluating MLLMs.","section":"Tables 1, 2, and 8"},{"comment":"The manuscript states: \"OpenAI's GPT-4o (referred to as 'o3') at 69.1%.\" This is a factual error: GPT-4o and o3 are different models, and the parenthetical misattributes a reported score. In a survey whose purpose is to summarize model performance on benchmarks, such a mislabeling is a load-bearing accuracy problem because readers may rely on the reported comparisons for model selection. The sentence should be corrected and the source of the 69.1% figure should be cited.","section":"§1.2, paragraph 3"}],"minor_comments":[{"comment":"The label 'Chain & Crypto' should read 'Blockchain & Cryptocurrency' to match the terminology used in Section 5.3 and Table 3.","section":"Figure 1"},{"comment":"The model name 'LLaVA' is repeatedly typeset as 'LLaV A' with an erroneous space; please correct this globally.","section":"Throughout"},{"comment":"The benchmark cited as [104] is called 'KnowledgeMath' in Table 4 and Section 6.3 but 'FinanceMath' in Table 6; the text in §6.3 notes the alternative name, but the tables should use one canonical name or explicitly cross-reference the alias.","section":"§6.1 vs §8.2.1"},{"comment":"The row for FinanceBench lists 'GPT-4+retriever' with performance '19% correct'; the surrounding text (Section 8.2.3) says GPT-4 provided correct responses to only 19% of questions, but the table does not indicate whether this is the best result among the evaluated models or a representative one; please clarify.","section":"Table 6"}],"recommendation":"major_revision","confidential_remarks":"The paper is best treated as a resource paper whose value hinges on accuracy and reproducibility. The internal inconsistencies (seven vs. eight disciplines, the o3 mislabeling, and the inclusion of non-benchmarks in the tables) are pervasive enough that a light copyedit will not suffice; the authors should be asked to rework the methodology description and table curation before the survey can be relied upon."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a useful compilation, not a research contribution. It organizes a large set of domain-specific MLLM benchmarks into eight discipline buckets, gives a readable narrative per domain, and the finance and healthcare sections in particular cover recent work that I had not fully tracked. For someone entering a new domain, the tables are a decent starting point. Credit where due: the discipline hierarchy is sensible and the paper brings together a lot of 2023–2025 material in one place.\n\nThe soft spots are real but mostly fixable. The abstract says seven disciplines while Section 2 and Figure 1 present eight; that looks like a leftover from a revision and needs to be reconciled. More substantively, the search strategy in Section 2 is too thin to support the word \"comprehensive\": it lists four sources and generic keywords, but gives no screening criteria, deduplication procedure, date range, record counts, or exclusion log. A reader cannot tell whether the taxonomy is representative or a convenience sample. The stress-test note is right about the tables too: BIG-bench is not a domain-specific software engineering benchmark, ToT is a prompting method rather than a benchmark, and LongLLaVA, KOSMOS-1, and ChatGLM are models or systems, not benchmarks. That scope drift undermines confidence in the categorization even where retrieval itself was thorough.\n\nThere are also smaller accuracy slips, such as referring to GPT-4o as \"o3\" in the introduction. Performance numbers are drawn from individual papers without independent verification, which is normal for a survey, but the reader should not treat them as adjudicated. None of this is fatal: the paper is a survey, not a theorem, and the errors are correctable in revision.\n\nMy take: this deserves serious peer review, not a desk reject, because a well-scoped survey of this kind is genuinely useful and the problems are addressable. The authors should be asked to fix the seven/eight discrepancy, tighten the inclusion criteria so the search is reproducible, and purge the tables of items that are not benchmarks. The paper would also be stronger if it explicitly stated that the taxonomy is a organizational convenience rather than a validated framework. For my own work, I would not cite it in its current form, but if the tables are cleaned and the method described properly, I could see it becoming a standard pointer reference. I would bring it to a reading group only to discuss what counts as a benchmark survey; the intellectual content is not deep enough for more than that.","headline":"A workmanlike survey that will be handy as a pointer resource, but its central coverage claim is not checkable as written and several table entries are off-scope.","tokens_in":42411,"tokens_out":1481,"would_cite":false,"duration_ms":20491,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"General-purpose multimodal models fall short in specialized fields, so this survey organizes the domain-specific benchmarks that measure that gap into a single eight-discipline taxonomy.","keywords":["multimodal large language models","domain-specific benchmarks","benchmark taxonomy","LLM evaluation","survey","artificial general intelligence","domain adaptation"],"falsifier":"A systematic replication that runs the paper's own search terms against the same databases and screens for domain-specific MLLM benchmarks could settle the coverage claim: finding a substantial discipline or an established benchmark family that fits none of the eight taxonomy branches would refute comprehensiveness. A cheaper check is row-level: every benchmark named in the domain tables should be traceable to a paper that actually introduces or evaluates it, and any miscategorized row would weaken the resource's reliability.","tokens_in":41432,"feed_emoji":"🧭","tokens_out":7587,"duration_ms":66749,"temperature":0.7,"pith_summary":"This paper is trying to establish that general-purpose multimodal large language models (MLLMs) are not enough: measuring and advancing progress in specialized fields requires domain-specific benchmarks. It surveys the benchmark literature across the eight disciplines listed in the methodology—engineering, science, technology, mathematics, humanities, finance, healthcare, and language understanding—and organizes the results into a taxonomy of domains, sub-domains, and application areas. The paper's contribution to a reader is a consolidated map of which benchmarks exist, what input modalities they use, and where current models succeed or fail. A sympathetic reading is that it makes the 'last mile problem' concrete by assembling repeated evidence that strong general models underperform when faced with specialized data and reasoning.","feed_headline":"Domain-specific benchmarks expose where multimodal LLMs fail","feed_subtitle":"One taxonomy groups specialized tests across eight disciplines, showing where model abilities hold up.","key_machinery":"The organizing device is the eight-branch domain hierarchy (Figure 1): disciplines → domains → sub-domains → application areas. Each branch is paired with a consolidation table that records the benchmark's scale, task type, input modality, models evaluated, and performance. That hierarchy carries the survey's argument: it converts scattered benchmark papers into a structured picture of where the MLLM evaluation ecosystem is dense, where it is thin, and where model failure is systematic.","core_discovery":"The central claim is that the field has reached the point where general benchmarks no longer tell the full story: MLLMs need domain-specific benchmarks to expose and guide their specialized capabilities. The paper's positive contribution is a taxonomy, announced as seven disciplines in the abstract but presented as eight in the methodology and Figure 1, with per-domain tables consolidating each benchmark's scale, task type, input modality, models evaluated, and reported performance. Across the domains the assembled evidence shows a recurring pattern: frontier models handle high-level reasoning and familiar text well, but stumble on fine-grained perception, specialized data formats, and domain-specific reasoning—low accuracy on financial question answering, weak geospatial localization, pathology understanding far below human experts, and poor materials property prediction. The survey frames these failures not as isolated results but as a systematic gap that domain-specific benchmarking is meant to close.","pith_inferences":["Going beyond the paper: its own tables suggest a testable hypothesis that MLLM performance on a benchmark depends less on model size than on how closely the task's input format and vocabulary match the model's training distribution.","The abstract's seven-discipline count versus the methodology's eight suggests the taxonomy's boundaries were still shifting; a natural extension is a living, community-maintained registry that updates the map as new benchmarks appear.","The recurring 'text shortcut' failure—models answering from language cues instead of analyzing images—implies that future benchmarks should include diagnostic variants that remove one modality at a time, so scores cannot be gamed by linguistic priors.","The survey's cited fine-tuning results imply that domain benchmarks can serve as training signal, not just evaluation tools; adversarially noisy or domain-specific examples measurably improve robustness and accuracy."],"forward_implications":["If the taxonomy is accurate, a researcher entering an unfamiliar domain can locate the relevant benchmarks and their reported baselines in one place, lowering the cost of designing a new evaluation.","The cross-domain pattern of failures—text shortcuts in pathology, poor counting and localization in geospatial data, low accuracy on financial QA—implies that benchmark design should include controls that isolate each input modality.","The survey's evidence supports making future benchmarks 'living' and multimodal, and scoring robustness, efficiency, and safety as well as accuracy.","Domain-specific benchmark results argue for domain-adapted fine-tuning and hybrid pipelines (LLM plus retrieval, symbolic solvers, or verification tools) rather than relying on a generalist model alone.","If the map is complete, it gives a concrete way to track whether MLLM progress toward generally capable AI is actually broadening across disciplines, the paper's stated long-term aim."],"supporting_citations":[{"why":"Supplies the opening evidence that even GPT-4 answers only 19% of straightforward financial questions.","marker":"[6]"},{"why":"Shows that strong general coding performance does not transfer to domain-specific code generation.","marker":"[7]"},{"why":"Provides robotics evidence that model strengths in high-level planning hide weak basic perception.","marker":"[8]"},{"why":"Demonstrates the text-over-image bias in scientific QA, a pattern the survey generalizes.","marker":"[40]"},{"why":"Anchors the math section's argument that models often rely on text cues rather than diagrams.","marker":"[97]"},{"why":"Anchors the healthcare section's claim that even the best model scores only 53.96% on comprehensive medical tasks.","marker":"[181]"},{"why":"Shows models below random guessing on paired medical images, the survey's reliability warning.","marker":"[188]"},{"why":"Quantifies the human-model gap in pathology (71.8% vs 49.8%) and the text-shortcut concern.","marker":"[194]"}],"fun_headline_variants":["Domain benchmarks expose MLLM blind spots","Eight-domain taxonomy maps MLLM failures","Survey: General tests miss MLLM domain gaps","Why MLLMs need specialized benchmarks","New taxonomy reveals MLLM domain limits"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the search strategy described in Section 2 captured the full landscape of domain-specific MLLM benchmarks, so that the taxonomy and summary tables are representative and complete; the paper's own inconsistency between seven disciplines (abstract) and eight (methodology) shows those coverage boundaries are not clearly pinned down.","fun_headline_variants_meta":{"raw":{"variants":["Domain benchmarks expose MLLM blind spots","Eight-domain taxonomy maps MLLM failures","Survey: General tests miss MLLM domain gaps","Why MLLMs need specialized benchmarks","New taxonomy reveals MLLM domain limits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000182,"raw_usage":{"total_tokens":1254,"prompt_tokens":835,"completion_tokens":419,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":451,"completion_tokens_details":{"reasoning_tokens":352}},"tokens_in":451,"tokens_out":419,"duration_ms":5018,"temperature":1.0,"reasoning_tokens":352,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:36:17.283937+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A systematic replication that runs the paper's own search terms against the same databases and screens for domain-specific MLLM benchmarks could settle the coverage claim: finding a substantial discipline or an established benchmark family that fits none of the eight taxonomy branches would refute comprehensiveness. A cheaper check is row-level: every benchmark named in the domain tables should be traceable to a paper that actually introduces or evaluates it, and any miscategorized row would weaken the resource's reliability.","supporting_citations":[],"review_version":1}