{"id":"a9246ddc-eb51-41d8-8058-535362f050bc","arxiv_id":"2508.06585","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":3,"one_line_summary":"On CountQA, a benchmark of dense real-world counting, the best of 15 MLLMs scored 42.9% and accuracy fell as object counts rose.","lead":"CountQA is a new benchmark of over 1,500 question-answer pairs for counting objects in busy, real-world photos, and the paper reports that the best of 15 multimodal AI models scores 42.9 percent accuracy. It offers a concrete diagnostic for a basic capability gap, but the dataset and code are only promised for release after acceptance.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"CountQA's ground-truth labels and question-generation protocol are not specified, so the 42.9% top score and count-trend claim may partly reflect annotation or question artifacts rather than MLLM counting ability.","rationale":"The reader's weakest_assumption—that the benchmark's ground-truth counts and QA pairs are accurate and not auto-generated by the same class of models under evaluation—is exactly the load-bearing condition for the abstract's central claim. My concern is a concrete unpacking of that assumption: label noise and MLLM-generated questions would directly inflate the apparent counting failure and could artificially create the reported decline with count. Because the dataset and code are withheld, this cannot be checked from the abstract alone; the paper is therefore not verifiable, but it is also not shown to be wrong. I align with the reader's UNVERDICTED verdict and recommend no change. The proposed human-validation test is a single concrete check that would settle whether the 42.9% figure and the trend are measuring model counting ability or benchmark artifacts.","tokens_in":852,"tokens_out":2381,"duration_ms":29194,"concrete_test":"Release the dataset and code, then re-annotate a random 100-image subset with three independent human counters and administer the original QA pairs to human participants. If the benchmark labels disagree with majority human counts on more than ~5% of items, or if human accuracy on the QA pairs is at or below the top MLLM's 42.9%, then the benchmark's labels/question difficulty are not calibrated to isolate MLLM counting ability, and the central claim would need to be qualified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that 15 MLLMs plateau at 42.9% and worsen with count—rests entirely on CountQA being a valid probe of counting. The abstract states only that the benchmark has 'over 1,500 question-answer pairs' and 'real-world images with high object density, clutter, and occlusion'; it gives no annotation protocol, no question-authoring method, and no count-distribution statistics. This matters concretely. In dense/occluded real images, object counting is hard for humans too; if labels come from non-expert annotators or from an MLLM, label noise will depress all model scores and will grow with object count (because higher-count images are harder to label), producing exactly the reported declining trend without any statement about model capability. Likewise, if the questions were auto-generated by an MLLM from the images, the benchmark may measure self-consistency with that generator rather than perceptual counting. The 42.9% figure is therefore uninterpretable until the ground truth is independently validated. This is not an internal inconsistency; it is an unverified premise, and the paper explicitly withholds the dataset and code until acceptance.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces CountQA, a new benchmark of over 1,500 question-answer pairs built from real-world images with high object density, clutter, and occlusion. The authors evaluate 15 prominent multimodal large language models (MLLMs) on this benchmark and report that the best-performing model reaches only 42.9% accuracy, with accuracy declining as object counts increase. The stated goal is to diagnose a fundamental limitation of MLLMs in object counting and to encourage the development of more numerically grounded and spatially aware models. The paper promises to release the dataset and code upon acceptance.","tokens_in":1088,"tokens_out":2118,"duration_ms":27116,"significance":"If the benchmark's ground truth is valid and the evaluation protocol is sound, CountQA would fill an important gap: existing counting benchmarks tend to use sparse scenes or restricted domains, whereas real-world applications involve clutter and occlusion. The headline result—that 15 current MLLMs plateau below 50% accuracy and degrade with object count—would be practically significant for deployment decisions and for guiding future multimodal research. The explicit commitment to open-source the dataset and code is a strength. However, based on the abstract alone, the measurement's validity is not yet established: no annotation protocol, human baselines, question-generation method, or statistical analysis is reported, so the central numerical claims are currently unverifiable.","major_comments":[{"comment":"The central claim that the top model achieves only 42.9% accuracy, with performance declining as object counts rise, depends entirely on the validity of CountQA's ground-truth counts and question-answer pairs. The abstract does not describe how the ground truth was obtained, whether annotators were domain experts or MLLM-assisted, whether there was inter-annotator agreement, or how the questions were authored. If labels are noisy in dense scenes or if questions were auto-generated by an MLLM, the reported numbers could reflect label noise or model self-consistency rather than counting ability. The paper states it will open-source the data only upon acceptance, which further prevents independent verification of this load-bearing premise. This is not an internal inconsistency, but it is an unverified premise on which the paper's main conclusion rests.","section":"Abstract"},{"comment":"The decline in accuracy with object count is presented as a general result, but the abstract provides no per-count accuracy numbers, no distribution of object counts in the benchmark, and no confidence intervals. Higher-count images are typically harder for any estimator, human or machine; if the benchmark is skewed toward high counts, the declining trend may partly reflect image difficulty or label noise that also increases with count. The authors should provide a stratified analysis by count range, compare against human performance on the same images, and report statistical uncertainty. Without this, the trend claim is not yet established.","section":"Abstract"},{"comment":"The evaluation protocol is underspecified. The abstract names '15 prominent MLLMs' but does not list them, state the prompting scheme, decoding parameters, number of runs, or how accuracy was computed (exact match vs. tolerance, answer extraction, etc.). The 42.9% figure is a point estimate without variance; there is no way to assess whether differences among models are meaningful. For an empirical benchmark paper, the model roster and evaluation details are essential to interpreting the headline result, and they should at least be summarized in the abstract.","section":"Abstract"}],"minor_comments":[{"comment":"The abstract states 'over 1,500 question-answer pairs' but does not indicate how many unique images this corresponds to. Since multiple questions can be generated from one image, the effective sample size for image-level conclusions may be much smaller. Please clarify.","section":"Abstract"},{"comment":"The phrase 'mere 42.9%' would benefit from context: what is the chance level for the task, and what is human performance on the same benchmark? A 42.9% accuracy can be difficult or easy to interpret without such a baseline.","section":"Abstract"},{"comment":"The 'top-performing model' is not named. At least one example or a reference to the full results would help readers calibrate the claim.","section":"Abstract"},{"comment":"The final sentence about paving the way for models that are 'numerically grounded and spatially aware' is somewhat promotional; consider making the benchmark's diagnostic value more concrete.","section":"Abstract"}],"recommendation":"uncertain","confidential_remarks":"This report is based solely on the abstract; the full text was not available for review. The central empirical claims are plausible and important if the benchmark construction is sound, but the abstract does not provide enough methodological detail to assess validity. I recommend obtaining the full manuscript before making a decision. The policy of releasing the dataset and code only upon acceptance is a reproducibility concern, though it is not by itself disqualifying if the paper includes a complete protocol in the main text or supplement."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What's genuinely new: a benchmark aimed at high-density, cluttered, occluded real-world images, with a 15-model roster and a headline top accuracy of 42.9% that declines as object counts rise. If the benchmark is a valid probe, that number is a stark, citable diagnostic. The claimed gap in existing benchmarks—sparse densities or narrow domains—is plausible and worth filling.\n\nThe paper does something right even in abstract form: it states a clear gap, builds an instrument to address it, and makes an empirical claim that is falsifiable in principle. That's a legitimate contribution pattern for an evaluation paper.\n\nNow the soft spots, and they are real. The abstract gives no annotation protocol, no question-authoring method, no count-distribution statistics, and no error bars. The stress-test note is on target: if ground-truth counts were produced by non-experts or if the questions were auto-generated by an MLLM, the reported trend could reflect label noise or self-agreement rather than counting ability. Higher-count images are harder for humans to label too, so the decline with count could be an artifact. That concern is load-bearing until the benchmark artifacts are released—which the paper explicitly withholds until acceptance. This is not a fatal flaw; it is missing information.\n\nA smaller point: the abstract's tone oversells (\"paves the way for a new generation\"), but that doesn't affect the work's substance.\n\nMy take: this deserves peer review. A serious referee can examine the images, labels, and question-generation code to determine whether the 42.9% figure means anything. The empirical claim is important enough to warrant that scrutiny, even though I wouldn't cite it myself until the data and code are public and independently checkable. I'd bring it to a reading group to discuss what makes a benchmark trustworthy.\n\nRecommendation: send it to review, with a clear request for details on annotation validation and question-generation methodology.","headline":"CountQA is a plausible new benchmark for dense real-world counting, but the abstract-only evidence leaves the headline 42.9% uninterpretable until annotation and question-generation details are released.","tokens_in":1617,"tokens_out":984,"would_cite":false,"duration_ms":13867,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new benchmark shows multimodal AI models count objects in real photos correctly less than half the time.","keywords":["multimodal large language models","object counting","benchmark","real-world images","occlusion","object density","evaluation"],"falsifier":"Take a random sample of CountQA question-answer pairs and have independent human annotators verify each object count against the source image. If substantial label disagreement or evidence that the questions were generated by an MLLM emerges, the central accuracy claim would need reinterpretation. Alternatively, if human-verified labels match the benchmark and a well-known MLLM scores significantly above 42.9% on a re-split, the reported ceiling would be called into question.","tokens_in":718,"feed_emoji":"🔢","tokens_out":1368,"duration_ms":16873,"temperature":0.7,"pith_summary":"This paper introduces CountQA, a benchmark of over 1,500 question-answer pairs built from real-world images with high object density, clutter, and occlusion. Using it, the authors evaluate 15 prominent multimodal large language models and report that the best model reaches only 42.9% accuracy, with accuracy dropping further as the number of objects increases. A sympathetic reader would take the paper to establish that current MLLMs cannot reliably perform a basic cognitive skill—object counting—in realistic visual conditions, and that this weakness deserves dedicated measurement and repair.","feed_headline":"MLLMs count real-world objects at only 42.9 percent accuracy","feed_subtitle":"A 1,500-question benchmark of cluttered, occluded photos shows counting accuracy falls as object numbers rise.","key_machinery":"CountQA itself is the central mechanism: a curated benchmark of real-world images paired with counting questions, designed to test models under high object density, clutter, and occlusion. The benchmark's question-answer pairs serve as the controlled instrument that measures counting accuracy across 15 MLLMs.","core_discovery":"The central claim is that multimodal large language models are fundamentally unreliable at counting objects in real-world scenes, not just in sparse or synthetic settings. CountQA is presented as a benchmark designed to expose this gap: it contains over 1,500 question-answer pairs on real images with high object density, clutter, and occlusion. Evaluating 15 MLLMs on this benchmark, the authors report that the top-performing model achieves only 42.9% accuracy, and that accuracy declines as object counts rise. The paper argues this shows a severe limitation in numerical grounding and spatial awareness that existing benchmarks fail to capture.","pith_inferences":["The paper does not describe how ground-truth counts and question-answer pairs were annotated; if those labels were produced by an MLLM or by a noisy automated pipeline, the 42.9% figure could partly reflect annotation artifacts rather than pure counting ability.","Counting ability is a natural probe for broader numerical and spatial grounding, so CountQA's results may predict weaknesses in other relational tasks such as estimating area, ordering objects, or comparing quantities.","A testable extension is to vary image clutter and occlusion systematically while holding object count fixed, which would isolate whether the decline stems from numerosity itself or from visual difficulty.","The benchmark's reliance on real-world images makes it sensitive to domain bias; cross-domain validation would clarify whether the measured gap generalizes beyond the selected image sources."],"forward_implications":["If the reported accuracy holds, current MLLMs cannot be deployed in real-world tasks that require trustworthy object counts, such as inventory, medical imaging, or autonomous inspection.","Performance declining with object count implies the failure is not a fixed quirk but scales with task complexity, making dense scenes particularly unreliable.","A dedicated benchmark like CountQA provides a concrete target for future training and evaluation, pushing models toward numerically grounded rather than purely descriptive visual understanding.","The 42.9% top accuracy establishes a baseline that future MLLMs must surpass, making counting improvement measurable and comparable across models."],"supporting_citations":[],"fun_headline_variants":["MLLMs can't count: top accuracy 42.9% in the wild","Counting chaos: MLLMs score just 42.9% on cluttered scenes","CountQA benchmark shows MLLMs fail at high-density counting","MLLMs miscount in real photos: 42.9% accuracy max","Why MLLMs can't count: new benchmark exposes 42.9% ceiling"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The benchmark's ground-truth object counts and question-answer pairs are accurate and were not generated by the same class of multimodal models being evaluated; if the labels are noisy or the questions were written by an MLLM, the 42.9% accuracy could reflect annotation artifacts rather than genuine counting ability.","fun_headline_variants_meta":{"raw":{"variants":["MLLMs can't count: top accuracy 42.9% in the wild","Counting chaos: MLLMs score just 42.9% on cluttered scenes","CountQA benchmark shows MLLMs fail at high-density counting","MLLMs miscount in real photos: 42.9% accuracy max","Why MLLMs can't count: new benchmark exposes 42.9% ceiling"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000302,"raw_usage":{"total_tokens":1561,"prompt_tokens":716,"completion_tokens":845,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":460,"completion_tokens_details":{"reasoning_tokens":740}},"tokens_in":460,"tokens_out":845,"duration_ms":9458,"temperature":1.0,"reasoning_tokens":740,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T23:00:59.798718+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of CountQA question-answer pairs and have independent human annotators verify each object count against the source image. If substantial label disagreement or evidence that the questions were generated by an MLLM emerges, the central accuracy claim would need reinterpretation. Alternatively, if human-verified labels match the benchmark and a well-known MLLM scores significantly above 42.9% on a re-split, the reported ceiling would be called into question.","supporting_citations":[],"review_version":1}