{"id":"b68259a7-70e6-498a-87b9-d419be188b7d","arxiv_id":"2501.09672","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"CHIRP, a 104-question pairwise benchmark, exposes scaling trends in vision-language models that existing benchmarks miss, and its evaluations correlate better with training loss.","lead":"The authors built CHIRP, a benchmark of 104 open-ended image-question pairs, plus a suite of vision-language models called Robin, to test how well current evaluation methods rank models. They report that standard benchmarks fail to distinguish model scales, while CHIRP's pairwise, human-centered evaluations reveal clear scaling trends.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"CHIRP's LLM-scaling signal may be a response-fluency confound: larger Pythia models produce longer, more fluent answers, and human preference data are not controlled for length or style.","rationale":"I read the paper in good faith: the benchmark, Robin suite, and bootstrapped Elo analyses are useful, and the VE-size human trend plus hallucination criterion provide some evidence that images matter. The reader's concern about 104 questions, 5 matchups per question, and the 100-question GQA/TextVQA samples is legitimate and already supports a conditional verdict. However, the more load-bearing issue is construct validity, not sample size. If CHIRP's pairwise human preferences reward verbosity and grammatical fluency, then the headline observation that standard benchmarks miss 'the effect of language model scaling' is unsurprising and not specific to vision-language understanding. The paper acknowledges that CHIRP depends on evaluators' language proficiency (Section 7), but it never tests whether the preference signal survives when language quality is controlled or when the image is removed. Appendix C.2.7 checks only the GPT-4V(R) judge and only token log-probability, not human raters or response length. The proposed no-image control is a single experiment that would separate vision-grounded judgment from text-fluency preference. Provided the authors add such a control or a length/fluency covariate analysis, the conditional acceptance is appropriate; without it, the central scaling claim remains underdetermined.","tokens_in":19823,"tokens_out":5996,"duration_ms":62799,"concrete_test":"Run a no-image control for the LLM-size ablation: present the same 104 CHIRP questions and the same sampled response pairs to a fresh set of CloudResearch raters without showing the image, using identical pairwise instructions, and compute the resulting Elo ranking. Compare this ranking to the original image-present human ranking (e.g., Spearman correlation across the five LLM-size models and per-question preference agreement). If the no-image ranking is highly correlated (Spearman > 0.8) or the LLM-size trend still appears, the CHIRP LLM-scaling result is explained by response language quality rather than vision-grounded understanding. If the ranking collapses or reverses, the vision component is load-bearing.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that CHIRP captures LLM-scaling differences in visual response quality is threatened by a response-fluency confound. In the LLM-size ablation, only the Pythia language model changes while the vision encoder is fixed (Section 4); larger Pythia models naturally generate longer, more grammatical, and more fluent text. CHIRP scores are pairwise human preferences on open-ended questions (Section 3.2), and no analysis controls for response length, lexical diversity, grammaticality, or verbosity. Prior work cited by the authors (Wu & Aji, 2023) shows that human and LLM evaluators prefer fluent but flawed content, so the monotonic human-preference trend with LLM size (Figure 5, top row) may measure language quality rather than image-grounded understanding. The paper's only check against text-probability confounds (Appendix C.2.7) applies to GPT-4V(R) judgments, not human raters, and it does not address length or style. The VE-size ablation shows an image-related signal, but it does not isolate the LLM-size effect. If the LLM trend persists after matching responses for length and fluency, the claim stands; if not, CHIRP is partly a language-fluency benchmark.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces CHIRP, a 104-question open-ended vision-language benchmark scored by pairwise preference judgments from humans or VLMs, together with Robin, a suite of 20 LLaVA-style models that systematically vary vision-encoder and LLM sizes. The authors claim that standard benchmarks (ScienceQA, GQA, VQAv2, TextVQA, MM-Vet, LLaVA-Bench) fail to reveal scaling trends in model quality, while CHIRP shows a clear LLM-size trend. They support this with heatmaps, Elo ratings over bootstrapped pairwise preferences, an audit of GQA/TextVQA ground-truth errors, and an analysis of AI-judge agreement with human raters.","tokens_in":20027,"tokens_out":5866,"duration_ms":61223,"significance":"If the central claim holds, CHIRP would be a valuable complement to existing VLM benchmarks, and the Robin suite would provide a controlled testbed for studying scale effects in open-ended response quality. The paper's strengths include the public release of training code, model suite, benchmark, and hand-validated images; the use of bootstrap resampling for Elo scores; and the concrete audit of ground-truth errors in GQA and TextVQA. However, the current evidence base is too thin and too confounded to establish the central claim: the human-preference samples are small, no significance tests are reported, and the LLM-size ablation is not controlled for response length or fluency.","major_comments":[{"comment":"The central claim that CHIRP reveals LLM-size scaling rests on small, sparse preference samples with no inferential statistics. The LLM-size study consists of 104 questions × 5 matchups = 520 human judgments and the VE-size study of 312 judgments, with matchups randomly sampled per question; the Elo trajectories in Figure 5 are therefore averages over non-independent, sparse data. I ask for bootstrap confidence intervals on Elo differences between adjacent sizes, a permutation or sign test for monotonicity, and a statement of the minimum effect size detectable at this sample size. Without such tests, the visual contrast between the CHIRP heatmaps and the existing-benchmark heatmaps in Figure 2 does not establish that CHIRP reliably discriminates scales.","section":"§5.1, Figure 5"},{"comment":"The LLM-size trend is confounded with response length and fluency. In the LLM-size ablation the vision encoder is fixed and only the Pythia LLM varies; larger Pythia models typically produce longer, more grammatical, and more fluent text, and the paper itself cites Wu & Aji (2023) for the finding that human and LLM evaluators prefer fluent but flawed content. The only related check, Appendix C.2.7, tests whether GPT-4V (R) preferences agree with token log-probabilities (48.7%, near chance); this does not control for length, lexical diversity, or grammaticality, and it does not apply to human raters. I ask for a matched-response analysis, such as truncating responses to equal length or including length and fluency as covariates in a preference model, before interpreting the CHIRP trend as evidence of improved image-grounded understanding.","section":"§6.3, Appendix C.2.7"},{"comment":"The claim that AI evaluations can serve as proxies for human evaluations is not supported by the reported agreement statistics. Cohen's Kappa values of 0.10, 0.114, 0.204, and 0.216 fall into the 'slight' to 'fair' range, yet the text states that GPT-4V evaluations 'still exhibit very similar trends to human evaluations.' Low instance-level agreement can coincide with similar aggregate trends by construction, so I ask for a formal trend-level test: report bootstrap distributions of the Elo trajectory for each evaluator and test whether the human and AI trajectories have the same shape, rather than relying on visual similarity.","section":"§6.2.1, Table 1; §6.3"},{"comment":"The extrapolation from 9/100 incorrect GQA ground truths to '770,769 to 3,309,773 questions' uses a normal-approximation confidence interval (3.4%–14.6%) that is very wide and assumes a simple random sample. The subsequent statement that two SoTA models scoring within 3% 'could very well be equal' extrapolates from this fragile estimate and is used to motivate CHIRP. I ask for either a larger audit sample (e.g., 500 questions) or a substantially softened conclusion, since the current interval does not support a tight bound on the error rate.","section":"§5.2.3"}],"minor_comments":[{"comment":"The abstract in the paper body differs from the metadata abstract and contains typos such as 'langauge' and 'identifiying'; unify the two versions and proofread the text.","section":"Abstract and §1"},{"comment":"The number of Elo bootstrap iterations is reported as 500 in §3.4 but as '1000 samples' in the Figure 5 caption; make the numbers consistent.","section":"§3.4 and Figure 5 caption"},{"comment":"The GPT-4 evaluation prompt states that the model should respond only with 'Correct' or 'Incorrect', but the surrounding text allows 'Partially Correct'; clarify how partial-credit responses are scored.","section":"Appendix B.3"},{"comment":"The pseudocode contains several typos ('asssisstant', 'correspondig', 'exmaple') and inconsistent capitalization; clean up the prompts before publication.","section":"Appendix B.3 and Figure 20/21"},{"comment":"The caption says the first row concerns the VE-size ablation and the second row the LLM-size ablation, while Figure 5 presents the top row as LLM size and the bottom row as VE size; align the captions to avoid confusion.","section":"Figure 22 caption"},{"comment":"The displayed confidence-interval formula 'p∈p±z∗√(p∗(1−p))/n' should be written as a standard interval for a proportion, and the source of the population total 22,669,678 should be stated explicitly.","section":"§5.2.3"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a timely and important evaluation problem and the released artifacts are a genuine contribution. My main concern is that the central claim outruns the evidence: the human-preference sample is small, the LLM-size trend is confounded with fluency, and the AI-proxy claim rests on low kappa values. These issues are addressable with additional analysis rather than requiring a new paper, so I recommend major revision rather than rejection. I would also ask the editor to verify that the CHIRP dataset and evaluation code are actually accessible at the provided anonymous link before final acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"CHIRP is a useful addition to VLM evaluation, and the Robin scaling suite is a legitimate piece of engineering. The best parts are the benchmark itself—104 curated image-question pairs, eight question categories, pairwise preference protocol—and the care taken in the human study design. The ground-truth error analysis on GQA and TextVQA is a genuine contribution; finding 9/100 GQA questions with bad answers in a small sample is enough to worry about, even if the confidence interval is wide. The model suite is exactly the right tool for controlled scaling comparisons, and releasing code and data is the right call.\n\nThe main soft spot is the LLM-size ablation. With only the language model varying, larger Pythia models naturally produce longer, more fluent text. Human raters prefer fluent but flawed content, as the paper itself cites. The VE-size ablation partially rescues the story: with the LLM fixed, human preferences still scale with VE size, so an image-grounded signal is real. But the LLM-size trend may be overstating language quality rather than vision-language reasoning. The authors should run a length- or fluency-matched control on the human preference data.\n\nStatistical claims also outrun the evidence. 104 questions and a few hundred pairwise judgments do not support sharp conclusions about ranking reliability, and the Cohen's kappa between GPT-4V and humans is only 'slight' to 'fair,' yet the paper leans on AI proxy agreement. The extrapolation from 100-question samples to millions of GQA errors is an overreach, even if the direction is plausible.\n\nThe paper is honest about its limitations, and the fixes are straightforward: significance tests, matched-length controls, and clearer reporting of disagreement. This deserves a serious referee. I would cite the benchmark, but I would not rely on its LLM-scaling claim until the confound is addressed.","headline":"CHIRP is a useful benchmark and the Robin suite is a solid empirical contribution, but the LLM-scaling signal is likely confounded by response fluency and the statistical support is thinner than the claims.","tokens_in":20597,"tokens_out":3135,"would_cite":true,"duration_ms":34882,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Standard VLM benchmarks give similarly scoring models statistically indistinguishable scores; CHIRP, a 104-question open-ended pairwise benchmark, separates them and reveals clear language-model scaling that those benchmarks hide.","keywords":["vision-language models","benchmark design","open-ended evaluation","pairwise preference","Elo rating","scaling laws","LLM as judge","human evaluation"],"falsifier":"A direct check would be to run the dense version of the human study: recruit a fresh panel of raters to judge all 190 pairwise matchups of the 20 Robin models on all 104 CHIRP questions, several raters per matchup, and see whether the clear LLM-size scaling and the vision-encoder sweet spot still appear in the bootstrapped Elo rankings; if the ordering shifts materially, the headline result is an artifact of the sparse five-matchup sampling. A complementary check is to re-run the GQA and TextVQA comparisons after deleting the roughly 9% and 5% of questions flagged as faulty, since if the LLM-size trend still fails to appear on the cleaned sets, the faulty-ground-truth explanation for benchmark insensitivity would be weakened.","tokens_in":19611,"feed_emoji":"🎯","tokens_out":11960,"duration_ms":119260,"temperature":0.7,"pith_summary":"The paper argues that standard vision-language benchmarks grade short, right-or-wrong answers so coarsely that models of genuinely different quality end up with statistically indistinguishable scores, and introduces CHIRP to measure what those benchmarks miss. CHIRP is a benchmark of 104 open-ended questions (with generated images and no fixed correct answers) on which two models' responses are compared side by side by human raters or a VLM judge across five criteria, and the preferences are turned into bootstrapped Elo ratings. Evaluated on Robin, a suite of 20 models varying only in language-model size and vision-encoder size, CHIRP shows what the existing benchmarks do not: clear improvement as the language model grows, a strictly rising human-preference trend with vision-encoder size, and an optimal vision-to-language size ratio. The paper also traces why older benchmarks are blind, estimating that about 9% of sampled GQA questions and 5% of sampled TextVQA questions have faulty or ambiguous ground truths, and showing that switching to longer responses changes which questions a model gets right. The sympathetic reading is that human-centered, open-ended, pairwise evaluation is a more sensitive tool for telling models apart, and that AI judges such as GPT-4V can reproduce its overall trends at lower cost.","feed_headline":"A 104-question test exposes VLM quality gaps benchmarks miss","feed_subtitle":"Pairwise human and AI preferences show clear language-model scaling that VQA-style accuracy scores hide.","key_machinery":"The carrying object is CHIRP itself: a 104-question, eight-category benchmark of open-ended image-question pairs whose images were generated rather than scraped, so models cannot have memorized them, graded by pairwise preference, one evaluator compares two models' responses per question on overall preference, relevance and completeness, understanding and reasoning, hallucinations, and details, and the results are collapsed into Elo ratings through 500 bootstrap iterations. The pairing is essential: it converts evaluation from absolute accuracy against a fixed answer into a relative judgment, which is where size-dependent differences appear. Two supporting machines carry the argument: the Robin suite of 20 LLaVA-style models with independently varied language-model and vision-encoder sizes, which makes controlled single-parameter comparisons possible; and a set of diagnostic experiments (long-versus-short response scoring, LLM and VLM grading against human grades, and manual auditing of ground-truth quality) that attribute the insensitivity of older benchmarks to response length, grader rigidity, and faulty ground truths.","core_discovery":"On the paper's terms, the discovery is that the way a benchmark is designed, whether it asks short factual answers matched against a ground truth, is what hides real differences between vision-language models, not the underlying model itself. Using the Robin suite of 20 models in which only the language model (Pythia 410M to 12B) or the vision encoder (CLIP Base to ViT-g) changes, the authors find that standard benchmarks show no clear vision-encoder trend and only a weak language-model trend, while their own CHIRP benchmark, which asks 104 open-ended questions and records pairwise human and VLM preferences on five criteria, shows a clear language-model scaling effect in every evaluation. Human preferences on CHIRP also rise monotonically with vision-encoder size and expose a preferred ratio between vision and language scale. The paper attributes the gap to three design flaws in existing benchmarks, short responses that under-determine quality, graders too rigid to accept correct phrasings, and ground truths that are themselves wrong or ambiguous, and it measures the last of these directly: in a 100-question sample from GQA, 9% of questions had incorrect ground truths (95% confidence interval 3.4% to 14.6%), which is large relative to the typical few-percent gap between competing models.","pith_inferences":["If CHIRP's sensitivity is real, the same pairwise-preference machinery could be extended into a continuously updated public VLM leaderboard in which any two models are matched on a rotating pool of open-ended questions, since CHIRP already shows that Elo ratings from sparse matchups correlate with the model property, training loss, that scaling research cares about.","The paper's error audit method, sampling 100 questions and hand-checking ground truths, transfers directly to any other harvested VQA dataset, and applying it to newer benchmarks would immediately show whether the multi-percent noise floor is widespread.","Because near-chance agreement (48.7%) between GPT-4V's preferences and token log-probabilities indicates the judge responds to image content rather than fluency, a natural next test is whether CHIRP rankings predict which model a user prefers in a genuinely interactive setting, such as asking follow-up questions about an image.","Because every CHIRP image was generated for the benchmark, the dataset cannot be contaminated by training exposure today, but this also means its questions may not reflect the distribution of natural images; a natural-image twin of CHIRP would show whether the extra sensitivity survives outside the generated domain."],"forward_implications":["Two models that score within roughly 3% of each other on GQA, or 0.7% on TextVQA, are statistically indistinguishable on those benchmarks, so small leaderboard gaps there should not be read as real capability differences.","Long-form and short-form responses draw on different skills, since the same model answers a different set of questions correctly in each mode, so benchmarks must be deliberately designed for the response format they claim to measure.","A GPT-4V judge that reasons before choosing reproduces the overall CHIRP trends seen with human raters and correlates with training loss (distance correlation 0.96 versus 0.91 for humans), making AI-based evaluation a viable low-cost substitute for broad trend detection, though per-case agreement with humans is only slight to fair.","Human preferences on CHIRP reveal an optimal vision-encoder-to-language-model ratio for each language-model size, a pattern the standard benchmarks did not surface.","Performance on CHIRP tracks training loss more closely than performance on the other benchmarks tested, which the paper takes as evidence that CHIRP measures a distinct, human-valued skill that existing suites do not cover."],"supporting_citations":[{"why":"Supplies the Pythia language models from which the Robin suite is built, providing the controlled LLM-size scaling axis.","marker":"Biderman et al., 2023"},{"why":"Supplies the CLIP vision encoders whose size forms the second controlled axis and whose published scaling motivates the VE-size comparisons.","marker":"Cherti et al., 2023"},{"why":"GQA is the benchmark whose 9% faulty-ground-truth rate is measured and whose scores show no clear scaling, motivating CHIRP.","marker":"Hudson & Manning, 2019"},{"why":"TextVQA is the second audited benchmark (5% problematic questions) and a comparison point showing grading and question-design limits.","marker":"Singh et al., 2019"},{"why":"MM-Vet is both a comparison benchmark and the precedent for small, high-quality, AI-graded multimodal evaluation.","marker":"Yu et al., 2023"},{"why":"Provides the LLaVA visual-instruction-tuning recipe used to train Robin and the LLaVA-Bench open-ended benchmark that CHIRP argues is too narrow.","marker":"Liu et al., 2023b"},{"why":"Establishes the LLM-as-judge methodology (MT-Bench, Chatbot Arena) that justifies using GPT-4V as a proxy for human preference.","marker":"Zheng et al., 2023"},{"why":"Supplies Cohen's Kappa, the statistic used to quantify the slight-to-fair agreement between AI and human judges.","marker":"Cohen, 1960"},{"why":"Frames the scaling-law view against which CHIRP's correlation with training loss is measured.","marker":"Kaplan et al., 2020"}],"fun_headline_variants":["New benchmark exposes VLM evaluation blind spots","CHIRP: 104 questions reveal what VLM benchmarks miss","VLM quality gaps hidden by flawed benchmarks, CHIRP shows","CHIRP benchmark uncovers scaling that VQA hides","Why VLM benchmarks fail: CHIRP's fine-grained test"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim rests on the assumption that preference judgments collected from a modest pool of crowd-sourced English-speaking raters, each seeing a random subset of model pairings on 104 questions, yield Elo rankings stable enough to reveal true differences between models, and that 100-question samples are representative enough to estimate the error rates of GQA and TextVQA, while the paper's own Limitations section acknowledges the deliberately small size and the strong reliance on evaluator language proficiency.","fun_headline_variants_meta":{"raw":{"variants":["New benchmark exposes VLM evaluation blind spots","CHIRP: 104 questions reveal what VLM benchmarks miss","VLM quality gaps hidden by flawed benchmarks, CHIRP shows","CHIRP benchmark uncovers scaling that VQA hides","Why VLM benchmarks fail: CHIRP's fine-grained test"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000243,"raw_usage":{"total_tokens":1528,"prompt_tokens":944,"completion_tokens":584,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":560,"completion_tokens_details":{"reasoning_tokens":500}},"tokens_in":560,"tokens_out":584,"duration_ms":6071,"temperature":1.0,"reasoning_tokens":500,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T19:45:40.629057+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct check would be to run the dense version of the human study: recruit a fresh panel of raters to judge all 190 pairwise matchups of the 20 Robin models on all 104 CHIRP questions, several raters per matchup, and see whether the clear LLM-size scaling and the vision-encoder sweet spot still appear in the bootstrapped Elo rankings; if the ordering shifts materially, the headline result is an artifact of the sparse five-matchup sampling. A complementary check is to re-run the GQA and TextVQA comparisons after deleting the roughly 9% and 5% of questions flagged as faulty, since if the LLM-size trend still fails to appear on the cleaned sets, the faulty-ground-truth explanation for benchmark insensitivity would be weakened.","supporting_citations":[{"cited_title":"Pythia: A suite for analyzing large language models across training and scaling","cited_arxiv_id":null,"evidence_quote":"Supplies the Pythia language models from which the Robin suite is built, providing the controlled LLM-size scaling axis."},{"cited_title":"Reproducible scaling laws for contrastive language-image learning","cited_arxiv_id":null,"evidence_quote":"Supplies the CLIP vision encoders whose size forms the second controlled axis and whose published scaling motivates the VE-size comparisons."},{"cited_title":"Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei","cited_arxiv_id":null,"evidence_quote":"Frames the scaling-law view against which CHIRP's correlation with training loss is measured."}],"review_version":1}