{"id":"8eb0ae28-52ec-4189-965e-898a179ab9ea","arxiv_id":"2608.11074","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"CapProbe converts detailed caption evaluation into region-aligned multiple-choice checking with 25,650 questions and finds large coverage gaps and a competency-efficiency trade-off across 13 VLMs.","lead":"This paper introduces CapProbe, a benchmark that scores image captions by asking a language model to answer 74 region-specific multiple-choice questions per image using only the caption text. It reports that dense regional probing surfaces large coverage differences across 13 vision-language models that overlap-based caption metrics miss.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim that CapProbe reveals genuine cross-model coverage gaps depends on the QA set being neutral across model families; the conceded generator-benchmark alignment could inflate Coverage for Gemini-style captions under the caption-only judge.","rationale":"The reader's weakest assumption is exactly the benchmark-neutrality concern: the QA generators are also evaluated models. My stress-test agrees and makes the concern more concrete: the mechanism through which this bias would operate is the caption-only judge's commitment rate, which is the metric that drives the headline ranking. The paper's own fixed-judge robustness check (Table 4) is evidence that the protocol gives stable rankings under different readers, but it does not address QA-set bias because the QA set is held fixed; a Gemini-leaning QA set would inflate Gemini-style captions under any judge. The proposed cross-generator QA subset is a direct test: if the ranking flips, the CONDITIONAL verdict should be reinforced or moved to REJECT for the centrality claim; if it persists, the reviewer's concern is answered. Given the honest limitations and the strong judge-robustness check, the paper remains CONDITIONAL pending release of artifacts and this validation. No change to the reader's verdict is needed.","tokens_in":23266,"tokens_out":4711,"duration_ms":39680,"concrete_test":"Construct an independent QA set for a 50-image stratified subset using a VLM not in the evaluated 13, e.g., Claude-Opus-4.8 or Llama-3.1-405B, following the same region-mask and human-verification pipeline. Re-run the full 13-model evaluation on this subset under the same Qwen3-32B judge. If Gemini-3.1-Pro no longer ranks first, or if the coverage ordering changes substantially, then the headline coverage gap is a construction artifact. If the ranking persists, the neutrality concern is substantially mitigated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that CapProbe's dense region-aligned QA reliably exposes coverage gaps in caption quality—requires the QA set to be roughly neutral across caption model families. The paper concedes (Sec. A) that both QA generators (Gemini-3.1-Pro and GPT-5.5) are themselves evaluated models, and that wording, fact granularity, category emphasis, and distractor style may align with generator habits. This is not a cosmetic caveat. The judge answers from the caption alone; Coverage is simply the judge's willingness to select A–D. If a Gemini-generated caption shares wording or option phrasing with Gemini-influenced QA text, the judge will more readily commit, inflating Coverage without any change in underlying caption factuality. Table 3 shows Effective Accuracy clusters tightly (92.7–94.5%), so the headline separation between models comes almost entirely from Coverage (36.4–73.7%). If that Coverage gap reflects output-style overlap with the QA set rather than factual coverage of the image, the main observation—that models differ in coverage, not committed-answer accuracy—becomes an artifact of benchmark construction.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces CapProbe, a dense, region-aligned QA benchmark for evaluating detailed image captions. Images are segregated into coarse semantic regions; for each region, multiple-choice questions across ten categories are generated by Gemini-3.1-Pro and GPT-5.5, then deduplicated, balanced, and human-reviewed. A language judge answers the questions from the caption alone, with an 'Uncertain' (E) option. The authors define competency metrics (Overall Accuracy, Effective Accuracy, Uncertain Ratio, Coverage) and efficiency metrics (Tokens/Image, Mean Density, Global Density, Density CV), and report results for 13 VLMs. The headline findings are that Effective Accuracy is tightly clustered across models while Coverage varies widely, and that dense probing reveals coverage gaps invisible to sparse or overlap-based metrics. The authors also report a judge-reliability analysis on three top models showing rank stability. The paper explicitly lists limitations in Appendix A, including the generator-benchmark alignment risk, missing region-recall evaluation, and absent inter-annotator agreement metrics.","tokens_in":23488,"tokens_out":4914,"duration_ms":46042,"significance":"If the benchmark is neutral across model families and publicly released, CapProbe would be a useful contribution: it offers higher probe density than CaptionQA (74 vs. 50.3 QA pairs per image), adds explicit region-level anchoring, and proposes a transparent metric framework where Overall Accuracy is exactly the product of Effective Accuracy and Coverage. The derivations in Sec. 4 are sound, and the protocol is cost-effective by design. The authors are also commendably explicit about the judge-conditioned nature of the scores and the distinction between proxies and causal measurements. However, the central claim that the observed Coverage gaps reflect genuine caption-quality differences rather than construction bias is not yet backed by a neutral-QA experiment, and the benchmark itself is unreleased. These issues must be addressed before the paper's central conclusions are convincingly established.","major_comments":[{"comment":"The headline observation—that models differ mainly in Coverage, not Effective Accuracy—is potentially confounded by benchmark construction bias. The QA set and region metadata are generated by Gemini-3.1-Pro and GPT-5.5, both of which are among the evaluated caption models, and the judge answers from the caption alone. If a caption uses phrasing, fact granularity, or distractor style aligned with the QA generator, the judge will commit to A–D more readily, inflating Coverage and Overall Accuracy without any change in underlying caption factuality. Table 3 shows Effective Accuracy clustering in a narrow range (92.7–94.5%) while Coverage spans 36.4–73.7%, so the coverage gap is doing essentially all the work in separating models—exactly the pattern that a style-overlap artifact would produce. The paper acknowledges this in Sec. A and calls for cross-generator validation, but that validation is essential, not optional, for the paper's central claim. I request a concrete experiment: generate a QA subset with models outside the evaluated set (e.g., open-weight or non-Gemini models), run the same judge and metric pipeline, and show that the ranking and coverage gaps persist. Without this, the claim that CapProbe 'reliably' exposes coverage gaps cannot be distinguished from a construction artifact.","section":"Sec. A; Sec. 5.2; Table 3"},{"comment":"The main contribution of the paper is the CapProbe benchmark itself, yet the data, annotations, and evaluation code are only described as 'will be released soon' with no URL, repository, or supplementary material. For a benchmark paper, this is a load-bearing issue: none of the experimental results can be reproduced or independently verified, and the proposed evaluation protocol cannot be adopted by the community. To support the claims, the authors should provide either a public release of the benchmark and code or a clear, concrete release plan (e.g., repository URL and estimated date) in addition to a data card detailing the QA pairs, region masks, and human-review information. Without this, the contribution is not yet assessable as a usable artifact.","section":"Abstract; Sec. 1; Sec. 3.8"},{"comment":"The human quality assurance protocol is described as three annotators spending over 200 hours reviewing all 26,062 QA pairs, but the paper does not report inter-annotator agreement, the distribution of accept/edit/delete decisions, or the nature of edits made. The paper explicitly states this in Sec. A, but the absence of these numbers matters for the benchmark's reliability: if a large fraction of pairs were edited, or if annotators disagreed substantially, the ground truth labels may contain systematic noise that affects all subsequent metric values. I ask the authors to report at least the per-annotator accept/edit/delete counts, pairwise Cohen's kappa or similar agreement statistics, and a few examples of common edits. This is necessary to substantiate the claim that the benchmark has high-quality ground truth.","section":"Sec. 3.7; Sec. A (fifth limitation)"},{"comment":"The 'full-scene' claim is explicitly a design goal rather than a measured property, since region recall against human-enumerated salient entities is not reported. This is more than a terminological caveat: if the region set systematically misses certain kinds of content (e.g., small background objects, amorphous regions, or rare categories), then the QA probes derived from those regions will not cover those facts, and the benchmark's ability to expose coverage gaps becomes asymmetric across content types. The paper acknowledges this limitation, but it remains a load-bearing gap for the central claim that CapProbe checks 'full-scene' factual coverage. I request a human-annotated subset (e.g., 50 images) where annotators enumerate salient entities/regions, and a comparison of the YOLOv26-seg/SAM3 region set against that enumeration, reporting recall and missed-region characteristics. This would quantify the degree to which 'full-scene' holds in practice.","section":"Sec. 3.2; Sec. A (third limitation)"}],"minor_comments":[{"comment":"The judge-reliability analysis is performed only on the top-3 caption models; the claim that rankings are 'stable' across judges would be stronger if evaluated on a broader subset of the 13 models, given that the Coverage gaps among lower-ranked models are smaller.","section":"Sec. 5.4"},{"comment":"The image-balancing step is described as reducing the dataset from 664 images and 39,127 QA pairs to 346 images, 1,868 regions, and 26,062 QA pairs, while Sec. 3.7 reports the final count as 25,650 after human deletion; the intermediate number 26,062 is correct, but the text could clarify that the 1.6% reduction is from human deletion, not from the balancing step.","section":"Sec. 3.6"},{"comment":"In the 'CapProbe (Ours)' row, the 'Probes/Img' column shows 74 and the 'Region' column shows ✓, but the table does not show the number of QA pairs per region (13.7) or the fact that the region count per image is only 5.4; adding these would help readers interpret the density statistics.","section":"Table 1"},{"comment":"The definition of per-image density d_i uses a factor of 1000 to express values in per-mille; the paper states this in prose, but the units ('‰') are only introduced in Table 3, not at the equation site, which could cause initial confusion.","section":"Sec. 4.2, Eq. (6)"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: worth reading if you care about detailed caption evaluation. CapProbe gives you a genuinely denser, region-anchored QA protocol than CaptionQA, and the authors are more honest than usual about its main weakness. The headline result—models differ mostly in Coverage, not in committed-answer accuracy—is plausible, but the size of that gap may be partly an artifact of construction.\n\nWhat's new: 74 probes per image, explicit region grounding, 10 semantic categories, an Uncertain option that separates abstention from wrong answers, and density metrics. That combination doesn't exist in prior work. The fixed-judge ranking robustness check (three judges, same ranking) is real evidence for comparative use. Metric definitions are clean, and Overall = Effective × Coverage is a useful decomposition, not a gimmick.\n\nSoft spots, in order of severity. First, the QA set is generated by Gemini-3.1-Pro and GPT-5.5, both of which are among the evaluated models. The paper flags this in Sec. A and frames results as \"under the current construction pipeline,\" which is the right frame. But the concern isn't cosmetic: Effective Accuracy is tightly clustered (92.7–94.5%), so almost all cross-model separation comes from Coverage (36.4–73.7%). If Gemini-style captions share wording with Gemini-influenced QA text, the judge will commit more often and Coverage rises without any change in factual content. The paper offers no cross-generator QA subset, so this remains a live threat to the central observation. Second, human QC reporting is thin: no inter-annotator agreement, no accept/edit/delete breakdown, and \"over 200 person-hours\" is vague for a benchmark meant to be reused. Third, no region recall against human-enumerated salient entities; \"full-scene\" is a design goal, not a verified property—the authors say so, but it limits claims about coverage completeness. Fourth, the artifacts are unreleased; \"will be released soon\" can't be checked.\n\nCitation pattern looks fine; self-citations are to actual prior benchmarks with concrete differences. The math is straightforward and correct. This is not a paper with a load-bearing flaw; it's a solid benchmark with a disclosed source of bias that needs measurement.\n\nWho this is for: anyone building or evaluating caption evaluation methods, especially people comparing VLM captioners across families. A serious referee should engage. The right outcome is a revise-and-resubmit asking for data release, a non-Gemini QA subset, and inter-annotator statistics. The core idea is sound; the evidence for the headline ranking is conditional.","headline":"CapProbe is a genuine step up in dense, region-anchored caption evaluation, but the headline coverage gaps are plausibly inflated by generator-family overlap and need cross-generator validation before they're trusted.","tokens_in":24051,"tokens_out":2015,"would_cite":true,"duration_ms":19885,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Dense region-level QA shows caption models differ mainly in how much of the scene their captions cover, not in how accurate their committed facts are.","keywords":["detailed image captioning","dense question answering","region-aligned evaluation","caption evaluation benchmark","vision-language models","coverage metrics","caption information density","multiple-choice probing"],"falsifier":"Build a second question set for the same 346 images whose metadata and MCQs come from outside the two generating models—human-written questions or an open-weight VLM with no Gemini or GPT involvement—keep the judge and protocol fixed, and check whether Gemini-3.1-Pro still ranks first on Overall Accuracy and Coverage. A cheaper variant: compute Gemini-3.1-Pro's accuracy separately on its own questions versus GPT-5.5's questions; a systematic gap would show the leaderboard reflects generator alignment rather than caption quality.","tokens_in":23040,"feed_emoji":"🔍","tokens_out":11234,"duration_ms":88584,"temperature":0.7,"pith_summary":"CapProbe turns evaluation of detailed image captions from fuzzy scoring into fact-checking: each test image is split into coarse semantic regions, and roughly 74 multiple-choice questions per image probe those regions across ten fact categories, with a language judge answering each question using only the caption under evaluation. The paper's central claim is that this dense, region-aligned questioning is a reliable diagnostic that exposes differences between caption models that sparse or overlap-based metrics miss. Across 13 vision-language models, the decisive result is that all models are about equally accurate once the judge commits to an answer, so the real differentiator is coverage—how much of the probed scene the caption actually lets the judge resolve. The accompanying density metrics show that verbosity does not equal information: the longest captions are often the least dense in verifiable facts. If the method holds up, detailed caption evaluation becomes a reproducible, comparatively cheap multiple-choice reading task instead of an open-ended scoring exercise.","feed_headline":"Dense probing shows caption models differ in coverage, not accuracy","feed_subtitle":"A region-aligned checklist of 25,650 questions separates what captions get right from what they never mention.","key_machinery":"The load-bearing object is the region-aligned dense QA checklist. Each image is decomposed by YOLOv26-seg and SAM3 into coarse foreground and background regions; Gemini-3.1-Pro writes structured metadata for every region, and Gemini-3.1-Pro together with GPT-5.5 generates multiple-choice questions across ten categories (attributes, recognition, count, OCR/text, camera features, spatial relations, and others), yielding 25,650 human-checked QA pairs over 346 images. At evaluation time, a language judge reads only the caption and answers each question, choosing the Uncertain option (E) when the caption lacks the information; that option drives the analytic identity $\\text{Overall Acc} = \\text{Effective Acc} \\times \\text{Coverage}$, which separates whether a caption lets a judge answer from whether the answered facts are correct. The efficiency side counts correct answers per thousand caption tokens ($d_i = (C_i/M_i)/T_i \\times 1000$), averaged per image versus pooled over request-token mass, so that verbose but uninformative captions are penalized.","core_discovery":"The authors claim that detailed caption quality decomposes into two judge-measured components with very different behavior across today's models. Using the identity $\\text{Overall Acc} = \\text{Effective Acc} \\times \\text{Coverage}$, they report that all 13 evaluated VLMs cluster near 93 percent Effective Accuracy when the judge commits to an answer, while Coverage spans from 36.43 percent (GPT-4o) to 73.68 percent (Gemini-3.1-Pro); the primary differentiator between caption models is therefore coverage, not the correctness of committed facts. The benchmark's second finding is that caption length is a poor proxy for informativeness: Gemini-3.1-Pro writes the longest captions (561 tokens per image) yet sits near the bottom on density, while GPT-5.5's short captions (222 tokens) reach the top densities, and no model simultaneously dominates coverage and efficiency along the reported Pareto frontier.","pith_inferences":["A natural next step the authors leave implicit: since coverage is the measured bottleneck rather than fact accuracy, it could be used directly as a training reward (for example, via reinforcement learning on CapProbe-style probes) to push caption models toward fuller scene coverage.","Because absolute scores are judge-conditioned (the same captions score between 43 and 78 percent across three judges), anchoring every judge with a fixed reference caption would convert judge-dependent numbers into comparable quantities—a calibration step the fixed-judge protocol does not currently provide.","Since 'full-scene' is a design goal rather than verified exhaustive coverage, an untested sensitivity remains: models that systematically under-describe background or 'stuff' regions will be penalized most on images with many such regions, so a per-image analysis by region count could reveal whether the coverage gap is concentrated in background description."],"forward_implications":["Under a fixed judge, model rankings are reproducible even with cheap open-weight readers, so detailed caption evaluation no longer requires a strong proprietary scorer.","Because coverage—not committed-answer accuracy—separates today's models, the main bottleneck for detailed captioning is telling the full scene, not stating facts correctly.","Length does not buy informativeness: density metrics expose which models pack verifiable facts per token, making them a usable efficiency axis for model comparison.","The competency–efficiency Pareto frontier means comprehensive captioners and efficient ones are currently different models, so improving either axis without losing the other is an open gap.","Per-category results localize the shared weaknesses of all current models—Count, Camera Features, and Attributes—giving concrete targets for the next generation of caption models."],"supporting_citations":[{"why":"Supplies the LLM-as-reader paradigm and the 50.3 QA-per-image density baseline that CapProbe extends to 74 questions with explicit region anchoring.","marker":"[45]"},{"why":"The scene-graph caption benchmark whose sparse QA and limited domain coverage CapProbe is designed to surpass.","marker":"[32]"},{"why":"Stands for the semantic-overlap metric family the paper argues conflates lexical similarity with factual correctness.","marker":"[5]"},{"why":"Generates the per-region structured metadata and part of the MCQs, and is also the top-ranked caption model under the benchmark.","marker":"[18]"},{"why":"Generates the other part of the MCQs and serves as the key efficiency contrast (short captions, highest densities).","marker":"[34]"},{"why":"Carries the 'things' side of segmentation, producing masks for discrete objects in each scene.","marker":"[24]"},{"why":"Supplies 'stuff' and background segments that complete the region set used for QA anchoring.","marker":"[11]"},{"why":"Provides the default language judge (Qwen3-32B) and the smaller judges used in the reliability analysis.","marker":"[44]"},{"why":"Represents the n-gram overlap evaluation that dense QA is proposed to replace for detailed captions.","marker":"[39]"}],"fun_headline_variants":["Coverage, not accuracy, separates caption models","Caption quality is all about coverage — accuracy is a tie","Long captions don't mean informative: coverage is key","25,650 questions show caption models miss details, not facts","Caption benchmarks: accuracy is saturated, coverage is not"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the benchmark is neutral enough across model families for its rankings to reflect caption quality, even though Gemini-3.1-Pro—the model that tops the leaderboard—also generated the region metadata and co-wrote the questions, a circularity the paper itself concedes in its limitations section.","fun_headline_variants_meta":{"raw":{"variants":["Coverage, not accuracy, separates caption models","Caption quality is all about coverage — accuracy is a tie","Long captions don't mean informative: coverage is key","25,650 questions show caption models miss details, not facts","Caption benchmarks: accuracy is saturated, coverage is not"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000892,"raw_usage":{"total_tokens":3893,"prompt_tokens":1036,"completion_tokens":2857,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":652,"completion_tokens_details":{"reasoning_tokens":2778}},"tokens_in":652,"tokens_out":2857,"duration_ms":18620,"temperature":1.0,"reasoning_tokens":2778,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T10:48:08.186097+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Build a second question set for the same 346 images whose metadata and MCQs come from outside the two generating models—human-written questions or an open-weight VLM with no Gemini or GPT involvement—keep the judge and protocol fixed, and check whether Gemini-3.1-Pro still ranks first on Overall Accuracy and Coverage. A cheaper variant: compute Gemini-3.1-Pro's accuracy separately on its own questions versus GPT-5.5's questions; a systematic gap would show the leaderboard reflects generator alignment rather than caption quality.","supporting_citations":[{"cited_title":"Cap- tionqa: Is your caption as useful as the image itself? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 23741–23750, 2026","cited_arxiv_id":null,"evidence_quote":"Supplies the LLM-as-reader paradigm and the 50.3 QA-per-image density baseline that CapProbe extends to 74 questions with explicit region anchoring."},{"cited_title":"Benchmarking large vision-language models via directed scene graph for comprehensive image captioning","cited_arxiv_id":null,"evidence_quote":"The scene-graph caption benchmark whose sparse QA and limited domain coverage CapProbe is designed to surpass."},{"cited_title":"Spice: Semantic propositional image caption evaluation","cited_arxiv_id":null,"evidence_quote":"Stands for the semantic-overlap metric family the paper argues conflates lexical similarity with factual correctness."},{"cited_title":"Gemini 3.1 pro: A smarter model for your most complex tasks.https://blog.google/ innovation - and - ai / models - and - research / gemini-models/gemini-3-1-pro/, 2026","cited_arxiv_id":null,"evidence_quote":"Generates the per-region structured metadata and part of the MCQs, and is also the top-ranked caption model under the benchmark."},{"cited_title":"Introducing GPT-5.5.https://openai.com/ index/introducing-gpt-5-5/, 2026","cited_arxiv_id":null,"evidence_quote":"Generates the other part of the MCQs and serves as the key efficiency contrast (short captions, highest densities)."}],"review_version":1}