{"id":"11c436cd-6817-4475-b195-dbab83fdcc47","arxiv_id":"2508.13692","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A hierarchical benchmark for multimodal models on human-centric visual understanding finds frontier models average under 60% and miss question-uncued visual evidence, with test-time scaling helping only marginally.","lead":"HumanPCR is a new benchmark with over 6,000 multiple-choice questions and 442 open-ended video questions that tests AI vision-language models on human-centric perception, comprehension, and reasoning. Early results show top models score near 60% or lower and struggle most when questions require gathering visual clues the question itself does not mention.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Human-R measurement validity is unverified: no human baseline, no inter-annotator agreement, no released data; the proactive-evidence failure rates and error taxonomy may overstate model deficiencies.","rationale":"Reading the paper in good faith: the Human-P/C construction via existing datasets with human verification is reasonable, and the 39-model evaluation is broad. The Human-R design criteria (multi-evidence, proactive, reasoning) are well-motivated, and the QC pipeline has the right shape. However, the central claim is not just that models score low; it is that they fail specifically because they miss proactive evidence and rely on query-guided retrieval. That mechanism claim rests on three unquantified links: (i) the official Human-R answers are uniquely correct/objective, (ii) the o3-mini adjudicator agrees with human experts on open-ended answers, and (iii) the manual error taxonomy in Fig. 10 is reliable. The reader's weakest assumption covers (i); the paper itself notes LLM-based metrics as a limitation (Conclusion), but no human agreement data are given for (ii), and (iii) is not described at all. These are addressable and do not make the result implausible—many independent findings point in the same direction—so the correct verdict remains CONDITIONAL/UNCHANGED rather than ACCEPT or REJECT. My proposed audit would settle the concern in one pass.","tokens_in":20795,"tokens_out":4609,"duration_ms":52839,"concrete_test":"Release the full Human-R data and run an independent human audit: for each of the 442 questions, have ≥3 domain experts answer from the video/question alone (blind to official answers, CoTs, and model outputs), and have a separate set of annotators re-label the Fig. 10 error sample (or a fresh 200-question sample) using a written taxonomy. Compute pairwise expert agreement and expert-vs-official answer agreement. If expert-official agreement is below ~80% or expert-expert agreement (e.g., Fleiss' kappa) is below 0.6 for answers or error labels, the reported Human-R accuracy numbers and the 'missed proactive evidence' dominance are not established; if agreement is high, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central diagnostic claim—that MLLMs miss proactive visual evidence and rely on query-guided retrieval—depends on Human-R being an instrument whose ground-truth answers and evidence requirements are uncontested. The manuscript reports a multi-stage QC pipeline (§3.2.3) and a below-1/5 acceptance rate, but provides no inter-annotator agreement for the 442 Human-R items, no human expert baseline accuracy on those items, and no release of videos, questions, answers, or CoTs for independent audit. Open-ended answers are scored by o3-mini (§4.1), and the supporting error analysis in §4.4/Fig. 10 is a manual classification of 200 questions with no stated protocol or reliability measure. With only 442 questions, even a modest share of ambiguous/disputable items would materially shift the reported percentages (e.g., o4-mini 58.60% vs most models <40%) and could inflate the 'missed proactive evidence' error category, since that label presupposes a unique correct answer and evidence chain. This is a measurement-validity concern, not a dispute with the direction of the result; but the central claim is not yet independently checkable from the preprint.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes HumanPCR, a three-level benchmark for human-centric multimodal understanding: Human-P (perception, 17 tasks), Human-C (comprehension, 17 tasks), and Human-R (442 manually curated open-ended video-reasoning questions with human-annotated chain-of-thought rationales). The authors evaluate over 30 open-source and proprietary MLLMs under a common protocol, reporting that most models score below 60% on Human-P/C and below 40% on Human-R, with the best model (o4-mini) reaching 58.60%. The central interpretive claim is that MLLMs fail on Human-R because they miss proactive visual evidence not cued by the question and rely on query-guided retrieval; the paper supports this with frame-scaling experiments, context-extraction comparisons, test-time-compute scaling, and a 200-question error taxonomy.","tokens_in":21005,"tokens_out":3774,"duration_ms":45586,"significance":"If the measurement is valid, HumanPCR would be a useful and fairly comprehensive diagnostic for human-centric MLLM evaluation. The three-level taxonomy, the 34-task coverage, the human-verified multiple-choice component, the human-annotated CoT rationales, and the extensive model sweep are all strengths. The Human-R design—requiring multiple visual evidence and proactive evidence seeking—targets a real gap in current video QA benchmarks, and the frame-scaling and test-time-compute experiments are informative. However, the headline diagnostic claim about missed proactive evidence depends entirely on the validity of the Human-R ground truth and its error labels. The current preprint does not establish inter-annotator agreement, a human expert baseline, or public access to the dataset, so the core measurement is not yet independently auditable. The direction of the result is plausible, but the magnitude and the error taxonomy are not yet verifiable.","major_comments":[{"comment":"Human-R is the load-bearing instrument for the paper's central claim, but its measurement validity is not established. The manuscript reports a multi-stage QC pipeline and a below-1/5 acceptance rate, yet no inter-annotator agreement is given for the 442 Human-R items, no human expert baseline accuracy is reported, and the data are not released. With only 442 questions, a modest share of ambiguous or contestable items would materially shift the reported percentages (e.g., 22 ambiguous items would move o4-mini's 58.60% by roughly 5 points) and could inflate the 'missed proactive evidence' error category, since that label presupposes a unique gold answer and evidence chain. Please report IAA on a held-out subset, provide a human baseline, and release the videos, questions, answers, evidence chains, and CoTs for independent audit.","section":"§3.2.3, Table 2"},{"comment":"Open-ended Human-R accuracy is adjudicated by o3-mini, but there is no validation that this judge agrees with human experts. If the judge is too strict, accepts plausible but wrong paraphrases, or is itself biased toward certain reasoning styles, all Human-R scores are systematically confounded. This directly affects the headline numbers in Table 2 and the comparisons in Table 3. Provide a human-scored subset with agreement statistics (e.g., Cohen's kappa) between o3-mini and expert annotators, and report sensitivity to the adjudication prompt or to using alternative strong judges.","section":"§4.1, Appendix D"},{"comment":"The error analysis is the direct evidence for the claim that models rely on query-guided retrieval and miss proactive visual evidence. However, the classification of 200 errors into five categories plus three subcategories is manual, with no stated protocol, no dual-annotation reliability, and no release of the classified examples. Because the labels 'Missed Proactive', 'Missed Referred', and 'Irrelevant Evidence' presuppose the exact evidence chain in the gold CoT, the taxonomy is only as valid as the gold evidence annotations. Please report inter-annotator agreement on error labels and make the 200 error instances and their classifications public.","section":"§4.4, Figure 10"},{"comment":"The definitions of 'proactive visual evidence' and the scoring rubric for evidence count are described only qualitatively. Since Human-R inclusion thresholds require at least one essential proactive evidence and multiple evidence, and since Figures 2 and 5 quantify proactive vs. referred evidence, the rubric itself should be operationalized. Please release the annotation checklist, item-level evidence scores, and examples that were accepted/rejected at meta-review, so readers can assess the construct validity of the proactive-evidence measure.","section":"§3.1, Figures 2 and 5"}],"minor_comments":[{"comment":"The manuscript contains several typos and grammatical slips: 'artifical' (Abstract), 'includeing' (§3.1), 'suffient' (§3.1), 'levles' (§3.3), and 'specially' (Introduction). A careful proofread is needed.","section":"Throughout"},{"comment":"The x-axis label in Figure 10(a) appears garbled ('/glyph1197umber'), and the panel captions are hard to parse. Please fix the rendering and clarify whether the bars show counts or percentages.","section":"Figure 10"},{"comment":"The column structure of Table 3 is confusing: 'Direct', 'M=0 (CoT)', 'M=1', 'M=2', 'M=3' are not aligned consistently across model rows. Please unify the formatting and define exactly what each column measures.","section":"Table 3"},{"comment":"The sentence 'These results validate the necessity of our benchmark' is partly a restatement of the design choices (Human-R was built to require proactive multi-evidence reasoning). I recommend softening this to avoid the appearance of circularity: the results demonstrate that models fail on this constructed operationalization.","section":"§1"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of a general vision/multimodal venue. My major concern is not with the direction of the results but with the verifiability of Human-R: without IAA, a human baseline, adjudicator validation, or dataset release, the central claim cannot be independently checked from the preprint. I would encourage the editor to treat the measurement-validity requests as essential rather than cosmetic. The paper also cites several proprietary model pages as references; that is acceptable in this field but the evaluation timestamps should be explicit."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a useful benchmark paper with a genuinely new evaluation twist—Human-R questions force models to go after visual evidence the question doesn't point to. The 442-question set is small but human-annotated with CoTs, and evaluating 39 models gives the failure story some breadth. The main finding, that models lean on query-guided retrieval and miss proactive evidence, is credible as far as it goes.\n\nWhat's actually new: the 'proactive visual evidence' construction criterion is more specific than prior video reasoning benchmarks, and the comparisons against context extraction and test-time compute give the paper empirical value. The QC pipeline (annotator to reviewer to meta-reviewer, acceptance below 20%) is the right shape.\n\nWhere it's soft: the load-bearing measurement is not yet independently checkable. No human baseline, no inter-annotator agreement on the 442 Human-R items, no released data or code in the preprint. With 442 questions, even a modest share of ambiguous items shifts the percentages and could inflate the 'missed proactive evidence' category, which presupposes a unique gold answer. The 200-question error analysis is manual, with no stated protocol or reliability. Open-ended scoring via o3-mini is a reasonable but acknowledged shortcut; it adds LLM noise to the main numbers. There is also an internal inconsistency: the abstract says test-time thinking and context scaling yield 'only limited benefits,' while Section 4.4/Table 3 say BoN delivers over 5% gains and is 'generally effective.' That should be fixed. Human-P/C are built from existing datasets, so training contamination is plausible and not discussed.\n\nI don't buy the circularity charge strongly. Filtering for the construct you want to test is standard benchmark design; the bigger issue is validity, not circularity. The 'validates the necessity of our benchmark' phrasing is overreach, but it doesn't corrupt the measurement.\n\nWho it's for: anyone working on MLLM video understanding or human-centric evaluation. It deserves a serious referee—the flaws are addressable and the direction is worth checking. I'd want release and human baselines before trusting the precise failure rates, but I'd send it to review rather than desk reject.","headline":"A serious benchmark effort with a genuinely new proactive-evidence design, but the headline Human-R failure rates rest on unreleased annotation validity and need human baselines before they should be taken at face value.","tokens_in":21593,"tokens_out":3186,"would_cite":true,"duration_ms":35153,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"HumanPCR shows that today's multimodal models fail at human-centric video reasoning because they rely on question cues instead of proactively seeking visual evidence.","keywords":["multimodal large language models","human-centric video understanding","proactive visual evidence","video reasoning benchmark","chain-of-thought rationale","query-guided retrieval","perception comprehension reasoning","evaluation suite"],"falsifier":"Independent human baseline: have a panel of expert annotators answer the 442 Human-R questions without seeing the curated answers, and measure agreement with the ground truth (and with each other). If human agreement is well below ceiling or human accuracy is not near-perfect, the benchmark's answerability premise fails. A complementary check: re-word each question so the key proactive evidence is explicitly cued; if models then answer at near-ceiling accuracy, the deficit is question phrasing rather than a fundamental inability to seek evidence.","tokens_in":20596,"feed_emoji":"🧠","tokens_out":11991,"duration_ms":119008,"temperature":0.7,"pith_summary":"The paper introduces HumanPCR, a three-level benchmark for how well multimodal large language models (MLLMs) understand people in real-world scenes: perception (Human-P), comprehension (Human-C), and reasoning (Human-R). Its central claim is that current models, including the strongest proprietary ones, systematically fail at human-centric video reasoning when the question does not point to the relevant visual evidence: the best model reaches only about 59% on Human-R, and most models score below 40%. Error analysis attributes most failures to missed 'proactive visual evidence' — information present in the video but not cued by the question — rather than to missing world knowledge. The paper argues that scaling frames, adding context-extraction or retrieval modules, and increasing test-time thinking produce only limited gains, because these methods are themselves query-guided. A sympathetic reader should care because the result separates genuine video understanding from retrieval of question-cued content, which is closer to what real-world use demands.","feed_headline":"59% is the ceiling for human-scene video reasoning","feed_subtitle":"Across 30+ models, the bottleneck is not knowledge but proactively finding visual evidence the question never mentions.","key_machinery":"The load-bearing device is the Human-R protocol and its construct of 'proactive visual evidence' — video information not, or only partially, indicated by the question. Each question passed a three-part meta-review: it must require multiple visual evidence, its evidence-integration pattern must not be fully determined by the question, and at least one essential evidence must be proactive. Human-annotated Chain-of-Thought rationales itemize the key visual evidence, making the extraction step observable rather than hidden. This construct does the work: it converts a vague 'models don't understand humans' claim into a documented failure to extract evidence outside the query's cue, and lets the p","core_discovery":"Discovery: MLLMs fail at human-centric video reasoning because they retrieve evidence the question cues and miss 'proactive' evidence not cued by the question, so they never assemble the full human picture. Human-R measures this: 442 open-ended questions, each meta-reviewed to require multiple visual evidence with at least one essential piece unmentioned in the question, each with human-annotated CoT rationales. o4-mini scores 58.60%, o3 59.28%, most open-source models below 40%; error analysis shows missed proactive evidence dominates while referred evidence is rarely missed. Frame scaling, context extraction, best-of-n, and self-refinement give small or negative gains. The 6,176-question P","pith_inferences":["One testable extension not in the paper: use the Human-R CoT rationales as supervision for an explicit evidence-selection step; if accuracy rises, the correlation between missed proactive evidence and failures becomes a causal demonstration.","The proactivity criterion likely transfers beyond human scenes: any video-QA benchmark that names the relevant objects in the question invites query-guided shortcuts, so Human-R-style meta-review could sharpen general long-video and world-model evaluation.","The model-specific error patterns (one model misses less proactive evidence but adds more irrelevant evidence; another is more selective) imply a precision-recall tradeoff in evidence extraction that could be scored automatically against the annotated rationales."],"forward_implications":["If the proactive-evidence finding holds, accuracy on Human-R is a more realistic estimate of real-world video understanding than accuracy on benchmarks whose questions name the relevant objects, because real queries rarely specify where to look.","The benchmark separates model families: open-source models match proprietary ones on perception and comprehension but fall behind on reasoning, so the open-source gap is specifically a multi-evidence integration problem.","Scaling model size continues to improve Human-R while perception and comprehension plateau near 38B parameters, suggesting that reasoning over multiple visual evidence is where scale still buys performance.","Context-extraction methods validated on Video-MME — frame selection, pooling, token selection, memory retrieval — show smaller or negative gains on Human-R, indicating that query-guided compression is not a route to holistic video understanding.","Test-time compute helps only through wide sampling: Best-of-N gives more than 5% gains on three base models, while Self-Refine saturates or degrades, and the strongest thinking model still leaves a wide gap to human-level reasoning."],"supporting_citations":[{"why":"The general video benchmark used as the comparison baseline; context-extraction methods validated on it lose most of their gains on Human-R, supporting the proactive-evidence diagnosis.","marker":"[24]"},{"why":"A token-selection context-compression method whose query-guided filtering underperforms on Human-R, illustrating the failure of text-guided retrieval.","marker":"[49]"},{"why":"A long-context reasoning-oriented model whose modest Human-R scores show that scaling context and thinking does not close the proactive-evidence gap.","marker":"[30]"},{"why":"The proprietary o3/o4-mini reasoning models that reach the highest Human-R accuracy (59.28%/58.60%), defining the ceiling the paper explains.","marker":"[31]"},{"why":"The chain-of-thought technique whose prompting effects are analyzed and whose format the human-annotated rationales follow.","marker":"[45]"},{"why":"A large egocentric human-video corpus that supplies much of the real-world human-centric data and context for Human-R questions.","marker":"[10]"},{"why":"A video chain-of-thought reasoning benchmark used to contrast Human-R's higher demand for multiple visual evidence and proactive extraction.","marker":"[44]"},{"why":"An open-source model series used in the scaling and open-versus-proprietary analysis across HumanPCR levels.","marker":"[51]"}],"fun_headline_variants":["59% ceiling: AI can't spot unasked clues in human scenes","Vision AI's blind spot: unasked evidence in human scenes","MLLMs miss proactive clues: 59% ceiling on human video reasoning","Proactive evidence is the wall: AI tops at 59% in human scenes"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"Everything hinges on the 442 Human-R questions having one right answer that human experts would independently agree on; the paper reports a strict annotation-review pipeline but no inter-annotator agreement scores or human-baseline accuracy, so if many questions are ambiguous, the reported failure rates and 'missed proactive evidence' error counts partly measure benchmark noise.","fun_headline_variants_meta":{"raw":{"variants":["59% ceiling: AI can't spot unasked clues in human scenes","Vision AI's blind spot: unasked evidence in human scenes","MLLMs miss proactive clues: 59% ceiling on human video reasoning","Proactive evidence is the wall: AI tops at 59% in human scenes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000635,"raw_usage":{"total_tokens":2784,"prompt_tokens":781,"completion_tokens":2003,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":525,"completion_tokens_details":{"reasoning_tokens":1923}},"tokens_in":525,"tokens_out":2003,"duration_ms":14267,"temperature":1.0,"reasoning_tokens":1923,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T18:58:03.575471+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Independent human baseline: have a panel of expert annotators answer the 442 Human-R questions without seeing the curated answers, and measure agreement with the ground truth (and with each other). If human agreement is well below ceiling or human accuracy is not near-perfect, the benchmark's answerability premise fails. A complementary check: re-word each question so the key proactive evidence is explicitly cued; if models then answer at near-ceiling accuracy, the deficit is question phrasing rather than a fundamental inability to seek evidence.","supporting_citations":[],"review_version":1}