{"id":"79df91c3-f85a-4173-b83f-4a70f5e07df8","arxiv_id":"2501.04150","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A broad benchmark of four multimodal models shows small models can match large ones on specific recognition tasks but lag sharply on reasoning, multilingual, and complex multimodal tasks.","lead":"This paper compares two large multimodal AI models (GPT-4V and GPT-4o) with two smaller ones (LLaVA-NeXT and Phi-3-Vision) across dozens of visual tasks. It finds that small models match large ones on narrow recognition jobs but fall far behind on reasoning and complex understanding.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Small-model 'wins' are post-hoc selections from tiny samples without multiple-comparison control; the positive half of the central claim is not established.","rationale":"The reader identified sample size, annotator agreement, and prompt refinement as the weakest assumptions, leading to CONDITIONAL. I agree, but the most load-bearing issue is more specific: the paper's positive claim about 'specific scenarios' where small models are competitive is derived by scanning many tasks with tiny samples and then selecting the apparent wins, with no correction for multiple comparisons. This is a post-hoc selection problem, not just a measurement-noise problem. It directly threatens the actionability of the central claim: a practitioner cannot know whether the highlighted niche tasks are genuine strengths of small models or artifacts of chance. The negative claim, that large models dominate complex reasoning, is likely robust because the margins are large and consistent across tasks, but it is also based on small, hand-picked samples and non-transparent correctness judgments. The additional observation that Table 2.11's GPT-4V row duplicates Table 2.4 exactly, while the overalls differ, is a concrete internal inconsistency that reinforces the need for data release and recomputation. My recommended verdict remains CONDITIONAL: the paper's qualitative direction is plausible, but the quantitative support for the positive half is insufficient, and the specific scenarios should be re-verified with proper statistical controls and released data.","tokens_in":51729,"tokens_out":6759,"duration_ms":62510,"concrete_test":"Release the raw per-sample responses for all tasks and run: for each task where a small model's accuracy is greater than or equal to that of GPT-4o or GPT-4V, compute a two-sided exact binomial test (paired by sample if possible) and apply Benjamini-Hochberg correction across all tasks at FDR 0.05. Report which 'specific scenarios' survive correction, along with bootstrap 95% confidence intervals for each accuracy difference. Separately, re-run the TableInsights evaluation for GPT-4V and confirm whether its per-task scores equal the ChartInsights values; if they do, the table-reasoning comparison in Tables 2.11-2.13 is invalid and must be recomputed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim has two halves: small models match large models in 'specific scenarios', and they lag on complex reasoning. The first half is not statistically supported. The paper reports 48 tasks, many with only 5-28 samples (safety inspection n=5, celebrity recognition n=15, guided prompting n=15, logo in-the-wild n=18), then highlights tasks where a small model wins or ties, such as celebrity recognition, safety inspection, and logo recognition, as evidence of 'specific scenarios'. With such small n, a one- or two-response flip changes accuracy by 7-20 percentage points, so nominal wins are expected by chance when scanning dozens of tasks. No significance tests, confidence intervals, or multiple-comparison corrections are reported; the 'specific scenarios' are selected post hoc from the same data. This makes the positive half of the headline claim unfalsifiable from the reported evidence. The negative half (large models better on complex reasoning) is more robust because margins are larger and consistent, but even it rests on the same small samples and subjective correctness judgments. A related data-integrity issue: in Table 2.11 (TableInsights), the GPT-4V row is identical to the GPT-4V row in Table 2.4 (ChartInsights) for all ten task columns, yet the listed overall accuracy differs (50.47 vs 57.20), which is impossible unless one table was copy-pasted. This suggests the quantitative tables need verification before any ranking is accepted.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper benchmarks two large proprietary MLLMs (GPT-4V, GPT-4o) and two smaller open models (LLaVA-NeXT, Phi-3-Vision) across 48 tasks organized by input format: single image-text pairs, multiple images, multi-frame sequences, and interleaved image-text inputs. The main claims are that small models can match large models in specific recognition or narrow scenarios but lag substantially on complex reasoning, and that GPT-4o is the best overall model. The evaluation combines reused samples from external benchmarks with newly collected samples, uses three annotators to judge correctness, and includes both quantitative tables and qualitative failure-case analyses.","tokens_in":51976,"tokens_out":5594,"duration_ms":54858,"significance":"If the quantitative findings were reliable, this would be a useful practical guide: it would tell deployers where small models suffice and where only large models are adequate. The paper's strengths are its broad task taxonomy, the inclusion of multiple input formats, and the qualitative failure analysis, which highlights recurring weaknesses such as spatial localization and fine-grained counting. Credit is also due for grounding parts of the evaluation in existing datasets (MSCOCO, FSC-147, ChartInsights, BLINK) rather than relying solely on self-constructed examples. However, the quantitative evidence as reported is not strong enough to support the headline ranking claims, and at least one table contains an internal inconsistency that must be resolved before the numbers can be trusted.","major_comments":[{"comment":"The positive half of the central claim—that small MLLMs achieve comparable performance in specific scenarios—is not statistically supported. The evidence for this claim consists of post-hoc selections of tasks with very small samples, such as celebrity recognition (n=15), safety inspection (n=5), logo recognition in the wild (n=18), and guided prompting (n=15). No confidence intervals, significance tests, or multiple-comparison corrections are reported; scanning 48 tasks, a few nominal small-model wins are expected by chance. For example, in the safety-inspection task with n=5, a single response flip changes Phi-3-Vision's reported 100% accuracy to 80%, and a two-response flip changes it to 60%. The authors should either provide error bars and a pre-specified or corrected analysis of which tasks test the small-model claim, or explicitly reframe these observations as descriptive rather than as confirmatory evidence.","section":"§2.1.1, §2.1.2, §6.2.2"},{"comment":"The GPT-4V row in Table 2.11 (TableInsights) lists the exact same ten per-task values as the GPT-4V row in Table 2.4 (ChartInsights)—39.55, 60.29, 65.87, 32.03, 73.96, 41.73, 53.69, 67.75, 88.91, 73.44—while the reported overall accuracy differs (50.47 versus 57.20). Identical per-task percentages across two different datasets are extremely unlikely and strongly suggest that one row was copied from the other. The authors must regenerate and verify all quantitative tables, and the discrepancy must be explained before any model-ranking conclusion can be accepted.","section":"Tables 2.4 and 2.11"},{"comment":"The scene-understanding evaluation uses inconsistent denominators. The text states that 12 scene images were collected and three models are scored over 12, but Phi-3-Vision is reported as [2 + 2 × 0.5]/14. Either Phi-3-Vision was evaluated on a different sample size than the other models, or the printed denominator is a typo. In both cases the reported accuracy needs correction and the evaluation protocol should be stated uniformly.","section":"§2.1.2, Scene Understanding"},{"comment":"The manuscript states that three annotators judged model outputs, but no inter-annotator agreement is reported. Many tasks use 0.5 partial-credit scores (e.g., scene text recognition, multilingual translation, multimodal commonsense), so a single annotator's disagreement can shift a reported accuracy by several percentage points and can flip a claimed small-model win. The authors should report agreement statistics, release the annotation rubric and raw outputs, or at least quantify the sensitivity of the rankings to individual annotator judgments.","section":"§1.1 and throughout §2"}],"minor_comments":[{"comment":"The model name 'LLaV A-NeXT' appears with an extra space throughout; the Section 6 introduction also contains 'LLaV A-Vision', which should be 'LLaVA-NeXT'.","section":"§1.2, §6 intro"},{"comment":"The word 'illustarted' should be 'illustrated'.","section":"§2.1.2"},{"comment":"The sentence 'We also nake an interesting finding' should read 'We also make an interesting finding'.","section":"§2.1.4"},{"comment":"The phrase 'the mmetric evaluation' should read 'the metric evaluation'.","section":"§2.2.1"},{"comment":"Reproducibility details are missing: the authors should report API model versions and access dates, the exact random-sampling procedure for samples drawn from MSCOCO and FSC-147, and the annotator instructions and scoring rubric.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The identical GPT-4V row in Tables 2.4 and 2.11 is a serious data-integrity concern that goes beyond ordinary statistical weakness; I would ask for a full audit of the quantitative tables and the raw evaluation logs before sending the paper back for another round. If the discrepancy cannot be resolved, I would move toward rejection of the quantitative claims. The qualitative failure analysis and the task taxonomy are valuable, but they alone do not support the paper's central comparative claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, here's my read on arXiv:2501.04150. The paper compares GPT-4V, GPT-4o, LLaVA-NeXT, and Phi-3-Vision across 48 tasks and four input formats, with a mix of borrowed and newly collected samples. The scope is genuinely broad, and the task-to-capability mapping is helpful for practitioners deciding between on-device small models and API-based large ones. The qualitative headline—small models can match large ones on narrow recognition tasks but lag on complex reasoning—is plausible and broadly consistent with the tables.\n\nWhat is actually new is the cross-model, cross-format breadth, plus real-world applications like safety inspection, grocery checkout, and insurance reporting. Some of the newly collected samples and prompts are fresh. The authors also flag common failure modes and where even GPT-4o struggles, which is useful.\n\nThe soft spots are real and need attention. First, many tasks use only 5-28 samples. A one-response flip changes accuracy by 7-20 percentage points, and the paper highlights small-model wins in celebrity recognition (5/15) and safety inspection (5/5) as evidence of \"specific scenarios.\" With dozens of tasks and no significance tests or multiple-comparison control, those wins could easily be noise. The negative half—large models better at complex reasoning—is more robust because the margins are larger, but it still rests on small samples and subjective judging.\n\nSecond, and more concerning, there is a data-integrity red flag. In Table 2.11 (TableInsights), the GPT-4V row is identical to the GPT-4V row in Table 2.4 (ChartInsights) for all ten task columns, yet the overall accuracy differs (50.47 vs 57.20). That is impossible unless one table was copy-pasted and the overall recalculated from different data. This suggests the quantitative tables need verification before any ranking is accepted.\n\nThe authors openly state that they refined prompts when outputs were unsatisfactory and used three annotators without reporting agreement. That is honest, but it also means the evaluations are not independently reproducible. No code or data is released, which limits the paper's value as a benchmark.\n\nWho is this for? Practitioners who want a rough capability map and don't need precise scores will get some value. Researchers working on MLLM evaluation could use it as a source of task ideas, but not as a reference for exact numbers.\n\nRecommendation: send it to peer review, but with major revision required. The reviewers should demand significance tests or confidence bounds, a resolution of the TableInsights inconsistency, and a release of the benchmark assets. The core qualitative message may survive, but the paper in its current form is not a reliable quantitative reference.","headline":"A broad but statistically fragile capability map; the qualitative pattern may hold, but a copied table row and tiny samples undermine the quantitative rankings.","tokens_in":52546,"tokens_out":2833,"would_cite":false,"duration_ms":27236,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A four-model benchmark of multimodal language models finds GPT-4o the strongest overall, with small models competitive only in narrow recognition and prompt-specific tasks.","keywords":["multimodal large language models","benchmarking","GPT-4o","GPT-4V","LLaVA-NeXT","Phi-3-Vision","visual reasoning evaluation","multimodal capability boundaries"],"falsifier":"Run the same comparison on a pre-registered set of 200 chart-reasoning and 200 spatial-reasoning questions with fixed prompts and independent blind scoring by three or more annotators; if LLaVA-NeXT or Phi-3-Vision match GPT-4o within a few percentage points on those tasks, the paper's central boundary claim would collapse.","tokens_in":51517,"feed_emoji":"📊","tokens_out":6261,"duration_ms":58118,"temperature":0.7,"pith_summary":"This paper tries to map where small multimodal language models (MLLMs) can substitute for large ones and where they cannot. It benchmarks GPT-4V and GPT-4o against two small open models, LLaVA-NeXT and Phi-3-Vision, across 48 tasks organized by four input formats and by capability. The central finding is that small models match large ones on narrow recognition and prompt-specific jobs, such as logo, emotion, and celebrity identification, but fall far behind on tasks that require deeper reasoning, including visual math, Raven's progressive matrices, chart and table analysis, temporal ordering, and spatial localization. A sympathetic reader would care because the result gives a practical rule of thumb: choose small models for fast, cheap, on-device recognition and inspection, but do not expect them to replace large models on reasoning-heavy multimodal inputs. The paper also documents that even GPT-4o fails badly at precise object localization and abstract visual reasoning, so the benchmark is as much about a capability ceiling as about model size.","feed_headline":"Small MLLMs win narrow tasks, lag on deep reasoning","feed_subtitle":"GPT-4o leads on reasoning-heavy tasks; small models hold recognition niches.","key_machinery":"The mechanism that carries the argument is a multi-part evaluation protocol rather than a single theorem. It organizes 48 tasks into four input formats—single image-text pairs, multi-image with text, multi-frame with text, and interleaved image-text—and maps each task to a capability such as understanding, reasoning, assessment, or interaction. Quantitative evaluations are anchored to existing datasets (MSCOCO, FSC-147, BLINK, ChartInsights, MVTec AD) plus a new TableInsights set with 166 table questions, while correctness is judged by three annotators; when a model's output was unsatisfactory, the authors refined the prompt themselves. This protocol is what lets the paper claim a systematic comparison instead of an anecdotal one.","core_discovery":"On the paper's own terms, the discovery is that the capability boundary between large and small MLLMs is real and task-dependent. In most understanding and reasoning tasks, GPT-4o ranks first, with GPT-4V close behind; in specific recognition tasks, LLaVA-NeXT and Phi-3-Vision can reach comparable accuracy, sometimes beating the large models, as with celebrity recognition for LLaVA-NeXT and several chart and prompt-format tasks for Phi-3-Vision. When the task asks for multi-step inference or fine-grained interpretation, the small models' accuracy falls sharply, in several reasoning tasks to zero. The paper also identifies common failure cases shared by all four models, especially in spatial localization, where F1 scores stay below 0.1, and in abstract visual reasoning such as Raven's progressive matrices.","pith_inferences":["Not tested in the paper: whether the small-model gap on reasoning tasks comes from model size or from differences in training data and alignment, since the two small models differ on both axes; a size-matched, data-matched ablation would separate these causes.","A practical extension the paper only implies: small models could be embedded in constrained recognition pipelines while large models handle chart/table QA, visual math, and temporal reasoning, creating a two-tier deployment strategy.","Because the authors refined prompts when outputs were unsatisfactory, the absolute accuracy percentages should be read as prompt-dependent; a broader prompt search on the same tasks might narrow the measured gap on specific skills."],"forward_implications":["Small MLLMs can serve in narrow recognition and inspection jobs—logo, emotion, and celebrity identification, and helmet counting when an external detector supplies crops—at lower cost and faster speed.","Large models remain the default choice for reasoning-heavy multimodal tasks such as chart and table question answering, visual math, Raven's progressive matrices, and multi-frame temporal ordering.","Model choice should be task-specific rather than size-based, since Phi-3-Vision beats GPT-4V on several chart-reasoning prompt formats and LLaVA-NeXT leads on celebrity recognition.","Even the best model in the study, GPT-4o, shows a clear capability ceiling on precise object localization and abstract spatial reasoning, so these tasks remain open for future work.","Adding a reference image of a defect-free product improves large models' anomaly detection but not small models' performance, suggesting a boundary in how small models exploit additional visual context."],"supporting_citations":[{"why":"Supplies the GPT-4o reference model that ranks first in most understanding and reasoning tasks throughout the benchmark.","marker":"[OpenAI 2024]"},{"why":"Supplies the GPT-4V reference model, the second-strongest performer, with strengths in OCR, language generation, and spatial understanding.","marker":"[OpenAI 2023a]"},{"why":"Defines the LLaVA-NeXT small model, which wins narrow recognition tasks such as celebrity identification.","marker":"[Liu et al. 2024]"},{"why":"Defines the Phi-3-Vision small model, which outperforms expectations on some chart and prompt-format tasks but fails on higher-level reasoning.","marker":"[Abdin et al. 2024]"},{"why":"Provides the ChartInsights dataset and prompt methodology used for quantitative chart understanding and reasoning experiments.","marker":"[Wu et al. 2024]"},{"why":"Provides the BLINK tasks and data used for multi-view reasoning and forensic detection evaluations.","marker":"[Fu et al. 2024]"},{"why":"Supplies MSCOCO val2014 images and annotations used in counterfactual examples, object counting, object localization, and commonsense reasoning.","marker":"[Lin et al. 2014]"},{"why":"Supplies the FSC-147 dataset for object counting in dense, large-quantity scenarios.","marker":"[Ranjan et al. 2021]"},{"why":"Supplies SAM segment cut-outs used in dense captioning and part-object association applications.","marker":"[Kirillov et al. 2023]"}],"fun_headline_variants":["Small MLLMs match big on niche tasks, fail on reasoning","MLLM size: small models win recognition, lose reasoning","GPT-4o leads reasoning, small models shine in recognition","Small MLLMs hit zero on complex reasoning in benchmark","Benchmark: small MLLMs lag big on deep reasoning tasks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that small, hand-collected sample sets (sometimes 5-28 items per task) and the authors' own correctness judgments, made without reported inter-annotator agreement, are representative and fair enough to rank the four models.","fun_headline_variants_meta":{"raw":{"variants":["Small MLLMs match big on niche tasks, fail on reasoning","MLLM size: small models win recognition, lose reasoning","GPT-4o leads reasoning, small models shine in recognition","Small MLLMs hit zero on complex reasoning in benchmark","Benchmark: small MLLMs lag big on deep reasoning tasks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000165,"raw_usage":{"total_tokens":1248,"prompt_tokens":940,"completion_tokens":308,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":556,"completion_tokens_details":{"reasoning_tokens":221}},"tokens_in":556,"tokens_out":308,"duration_ms":3398,"temperature":1.0,"reasoning_tokens":221,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:21:21.680497+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same comparison on a pre-registered set of 200 chart-reasoning and 200 spatial-reasoning questions with fixed prompts and independent blind scoring by three or more annotators; if LLaVA-NeXT or Phi-3-Vision match GPT-4o within a few percentage points on those tasks, the paper's central boundary claim would collapse.","supporting_citations":[],"review_version":1}