{"id":"f6cea882-4fc4-4854-bb0e-89add979ab49","arxiv_id":"2412.12075","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A 12,129-question benchmark for long videos that requires models to retrieve the specific video moments supporting each answer, exposing a gap between multiple-choice accuracy and genuine video understanding.","lead":"CG-Bench is a new benchmark that tests whether AI video models answer questions about long videos by finding the right moments in the video, not by guessing from the options. It adds clue-grounded scoring to standard multiple choice, and finds that even GPT-4o drops sharply when asked to show which video segment supports its answer.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Credibility metrics are confounded by frame sampling: long/white-box tasks use 128 frames over 10-80 min videos, leaving only ~1.5 frames per 19s clue, so low acc@IoU/CRR may be a resolution artifact rather than evidence of failed clue-grounded understanding.","rationale":"The reader's conditional verdict is sound, but the load-bearing weakness is slightly different from the one highlighted. The clue-completeness concern is real and would affect interpretation of CRR and acc@IoU when hidden clues exist. However, the more immediate and testable threat is the frame-count mismatch in Sec 4.1: long-video evaluation uses 128 frames spread over 10-80 minute videos, while clue-based evaluation uses 32 frames over ~19 second clue intervals. This makes the paper's central demonstration - GPT-4o long-acc 45.2 dropping to 4.38 acc@IoU or CRR 77.5 - ambiguous. The paper's own human data show that 128-frame sampling drops human long-acc from 90.3 to 59.9, and full-video human acc@IoU is only 29.8, so the reported model gaps are contaminated by sampling resolution. The proposed human-baseline experiment would settle this directly. If human acc@IoU under 128 frames is near 4-5, the benchmark's credibility conclusions need substantial revision; if humans still achieve much higher scores under the same protocol, the central claim would be supported. The dataset itself is large, diverse, and carefully annotated, and the clue-grounded idea is valuable; the issue is with the diagnostic interpretation of the metrics, not the release. Thus I keep the conditional verdict.","tokens_in":17635,"tokens_out":15253,"duration_ms":132905,"concrete_test":"Run the white-box acc@IoU task with human annotators on the existing 30-video/296-question subset under the exact model protocol: 128 uniformly sampled frames with timestamps, requiring predicted intervals scored by tIoU>0. If human acc@IoU under this protocol is close to GPT-4o's 4.38 rather than the full-video 29.8, the reported acc@IoU drop is a measurement-resolution artifact rather than evidence of failed understanding.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In Sec 3.3.2 the paper interprets low acc@IoU and CRR as evidence that models do not ground answers in the annotated clues. Both metrics are computed under a temporal-undersampling asymmetry. Sec 4.1 sets Long-video MCQ at 128 uniformly sampled frames, Clue-based MCQ at 32 frames. With average video length 1624s (Table 2) and average clue length 19.24s (Table 1), 128 frames over the full video leaves only about 1.5 frames inside a typical clue interval. For white-box acc@IoU the model must output an interval with tIoU>0 from such discrete timestamps; a correct interval containing one or two frame times cannot be localized to 19 seconds. The human full-video acc@IoU is only 29.8 (Table 3), so even with full video this is hard, and the model was not given full video. For CRR, the premise long-acc >= clue-acc fails because long-acc sees ~1.5 clue frames while clue-acc sees 32 dense clue frames. The paper's own sparse-frame human long-acc (59.9 vs 90.3 full-video) confirms 128 frames is a severe bottleneck. If human clue-acc under the 32-frame protocol is anywhere near the full-video 92.2, human CRR under the paper's protocol would be about 65, comparable to or below GPT-4o's reported 77.5, so the metric does not isolate model clue recovery from sampling resolution.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces CG-Bench, a long-video question-answering benchmark with 1,219 manually curated videos and 12,129 human-annotated question-answer-clue (QAC) triplets. It proposes two credibility evaluation protocols—white-box grounding (mIoU, rec@IoU, acc@IoU) and black-box clue recovery (CRR)—in addition to traditional MCQ and clue-aided open-ended evaluation, aiming to determine whether MLLMs base their answers on specific video evidence. Experiments across open- and closed-source MLLMs report moderate MCQ accuracy but much lower scores on credibility metrics (e.g., GPT-4o long-acc. 45.2 versus acc@IoU 4.38), which the paper interprets as evidence that current models do not genuinely understand long videos.","tokens_in":17960,"tokens_out":6506,"duration_ms":54291,"significance":"If the credibility metrics are valid, CG-Bench makes a valuable contribution: it is the largest long-video QA benchmark with question grounding annotations, uses a granular 14/171/638 category taxonomy, releases data and a leaderboard, and includes human baselines. The white-box/black-box distinction is a useful conceptual advance over pure MCQ evaluation, and the paper provides careful analyses of frame sampling, prompt/modality effects, and the open-ended evaluator, with human agreement used to validate the latter. However, the paper's central claim that models fail to ground answers in video evidence currently rests on metrics that are confounded by the evaluation protocol, so the quantitative findings and the benchmark's credibility contribution need strengthening before they can be fully accepted.","major_comments":[{"comment":"The frame-count asymmetry confounds the credibility metrics. Long-video MCQ uses 128 uniformly sampled frames over full videos (average 1624 s), while clue-based MCQ uses 32 frames over clue clips (average 19.24 s per Table 1). This yields roughly 1.5 sampled frames inside a typical clue interval for the long-video setting. Consequently, low long-acc is partly a temporal-resolution artifact, and the drop in CRR from 100 may reflect the model seeing fewer clue frames rather than a failure to retrieve clues. The paper's own human sparse-frame long-acc of 59.9 versus 90.3 full-video (Table 3) quantifies this bottleneck. To support the claim that low CRR indicates poor clue recovery, the authors should report human CRR under the same 128/32 frame protocol or match frame density across the two settings.","section":"Sec 4.1, Table 3"},{"comment":"The premise underlying CRR, that long-acc should be greater than or equal to clue-acc, is violated by the design of the frame sampling. Because 32 frames over a 19-second clue are far denser than 128 frames over a 27-minute video, clue-acc has an inherent advantage unrelated to context dilution. The metric thus conflates temporal resolution with clue-retrieval ability. Additionally, the paper acknowledges in this same section that 'hidden clues' may exist, but it provides no coverage validation that the annotated intervals are either sufficient or necessary for answering each question. Without inter-annotator agreement or a coverage study, a low CRR could reflect incomplete annotation rather than a model deficit.","section":"Sec 3.3.2, Eq. (2)"},{"comment":"The acc@IoU threshold is under-specified. The text states that the default threshold is τ=0 (so any overlap, tIoU>0, counts) but also says that acc@IoU is calculated at IoU thresholds of 0.1, 0.2, 0.3, 0.4, and 0.5. It is unclear which value is reported in Table 3 (e.g., GPT-4o's 4.38). If the table reports τ=0, the metric requires only a nonzero overlap; if it reports an average over the thresholds, the interpretation is different. The authors should state explicitly which threshold (or whether an average) is used, and report the full threshold sweep in a figure or table.","section":"Sec 3.3.2, Table 3"},{"comment":"No inter-annotator agreement or coverage validation is reported for the clue intervals. The benchmark's credibility depends on the human-annotated clue intervals being both necessary and sufficient for each question, but the annotation process in Sec 3.1 describes only a qualitative 'review iteration' process. Provide quantitative agreement (e.g., temporal IoU between annotators, or kappa on interval selection) and a check that questions cannot be answered from outside the annotated clues. Without this, the interpretation of low acc@IoU and CRR as evidence of failure to ground answers is not fully supported.","section":"Sec 3.1"}],"minor_comments":[{"comment":"There are minor typos: 'Electonic' should be 'Electronic' in Figure 2, and 'clue-grouded' should be 'clue-grounded' in the Introduction.","section":"Figure 2, Sec 1"},{"comment":"In the CG-Bench-Clue row, the columns #Video and #Dur list 12,129 and 22.8, but these appear to be the number of QAC triplets and the average clue duration, not the number of videos and video duration. Please relabel the columns or correct the values.","section":"Table 2"},{"comment":"The sentence 'use 32 frames as the for Clue-based MCQ' is grammatically incomplete; please revise it.","section":"Sec 4.1"},{"comment":"The tIoU formula sums pairwise intersections over all ground-truth and predicted intervals; if intervals within a set overlap, pairwise intersections can exceed the true union. Please clarify whether intervals are merged before computing the metric and whether the formula is the standard multi-interval IoU.","section":"Eq. (1)"},{"comment":"The human sparse-frame evaluation uses only 30 videos (296 questions), which is a small sample for a strong claim that 128 frames are insufficient. Please report confidence intervals or increase the sample size.","section":"Sec 4.2"},{"comment":"The ablations in Figure 7 and Table 4 are computed on a 1000-QAC subset; please state explicitly whether these results were verified to be representative of the full benchmark.","section":"Figure 7, Table 4"}],"recommendation":"major_revision","confidential_remarks":"The benchmark resource itself is valuable and the release appears well-executed, but the paper's main novelty—the credibility metrics—is currently compromised by the frame-sampling asymmetry between the long-video and clue-clip settings. This is fixable by re-running with matched frame density or by providing human calibration under the same protocol, so I would not reject. I also suggest the editor require a clear statement of the acc@IoU threshold and inter-annotator agreement statistics for the clue annotations, as these are essential for interpreting the reported results."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"CG-Bench is a big, careful benchmark release: 1,219 long videos, 12k human-annotated QA pairs with clue intervals, a sensible taxonomy, and broad model evaluation. The motivation is real: MCQ-only long-video benchmarks let models guess from sparse frames and eliminate options. The paper's human full-video vs sparse-frame numbers (90.3 vs 59.9) confirm that undersampling is a serious issue.\n\nThe problem is that the paper's own credibility metrics inherit exactly the same undersampling asymmetry. Long-video MCQ uses 128 frames over a 27-minute video; clue-based MCQ uses 32 frames over a ~19-second clue clip. That's about 1.5 frames per clue in the long setting versus ~1.7 fps in the clue setting. So a low acc@IoU (GPT-4o: 4.38) or a CRR well below 100 may simply mean the model never saw enough frames inside the clue interval to localize it, not that it failed to ground its answer. The paper's own human data support this: human sparse-frame long-acc is 59.9, and human full-video acc@IoU is only 29.8. Under the paper's protocol, a human CRR would be roughly 59.9/92.2 = 65, below GPT-4o's 77.5. That means the metric does not isolate clue recovery from sampling resolution.\n\nThe white-box acc@IoU definition is also under-specified: the paper sets tau=0 by default but then reports values at IoU thresholds 0.1-0.5. It needs clarification. And the clue intervals are treated as ground truth without any inter-annotator agreement or coverage check. If a question can be answered from other parts of the video, low long-acc may just mean the annotation missed a clue.\n\nOn the plus side: the dataset itself is a real resource—large, open-ended, multimodal, with human review. The open-ended evaluator is sensible and validated against human scores. The ablation on prompts and modalities is useful. The paper is honest and clearly written; there is no circular fitting or invented entities.\n\nBottom line: this is a solid benchmark release that would benefit from a serious referee. The central claim about credibility, however, does not survive the frame-sampling confound. I'd ask the authors to match frame densities (or correct for them), report inter-annotator agreement, and clarify the acc@IoU threshold before treating the credibility metrics as evidence of anything about model grounding.\n\nYes, send it to review.","headline":"Large, well-annotated long-video QA benchmark, but the credibility metrics are confounded by frame sampling and don't support the paper's central claim.","tokens_in":18545,"tokens_out":3198,"would_cite":true,"duration_ms":26962,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"CG-Bench claims that multiple-choice-only benchmarks overstate long-video understanding, and supports this with clue-grounded evaluation: GPT-4o's 45.2% MCQ accuracy falls to 4.38% when the model must also point to the interval containing…","keywords":["long video understanding","multimodal large language models","video question answering","temporal grounding","clue-grounded evaluation","multiple-choice benchmark","credibility evaluation","human-annotated clues"],"falsifier":"Re-annotate a random sample of CG-Bench questions with independent annotators and compute inter-annotator $\\mathrm{tIoU}$ on the clue intervals; if many questions have low agreement, or if a model given the full video but with annotated clues masked still beats its clue-only accuracy, then hidden clues exist and the CRR inequality is not a valid test of clue retrieval.","tokens_in":1850,"feed_emoji":"🎬","tokens_out":2653,"duration_ms":87793,"temperature":0.7,"pith_summary":"CG-Bench argues that existing multiple-choice benchmarks for long-video understanding can be passed without genuinely watching the video: models can combine short glimpses, captions, and elimination to pick the right option. To close this loophole, the paper builds a benchmark of 1,219 long videos with 12,129 human-annotated question-answer-clue triplets, where each question is tied to specific clue intervals. It then evaluates models two ways: whether the answer is correct on the full video, and whether the model can locate the clue interval or keep full-video accuracy above clip-only accuracy. On these credibility metrics, GPT-4o's long-video accuracy drops from 45.2% to 4.38% when it must also point to a clue, and no model reaches 100% Clue Recovery Rate, which the paper reads as evidence that current MLLMs are not yet trustworthy for long videos.","feed_headline":"Clue-based test cuts GPT-4o long-video score from 45.2 to 4.38","feed_subtitle":"A new benchmark shows MCQ accuracy hides whether models actually watched the right part of a long video.","key_machinery":"The load-bearing object is the question-answer-clue (QAC) triplet, in which each multiple-choice question is annotated with one or more temporal intervals that contain the evidence. On top of it sit two evaluation mechanisms: white-box evaluation, which asks the model to output timestamp intervals and measures them with $\\mathrm{tIoU}$, $\\mathrm{mIoU}$, $\\mathrm{recall@IoU}$, and $\\mathrm{acc@IoU}$; and black-box evaluation, which compares full-video accuracy to clue-clip accuracy through the Clue Recovery Rate, $\\mathrm{CRR} = \\min(\\mathrm{long\\text{-}acc.}, \\mathrm{clue\\text{-}acc.})/\\mathrm{clue\\text{-}acc.}$. The identity doing the work is the inequality $\\mathrm{long\\text{-}acc.} \\ge \\mathrm{clue\\text{-}acc.}$, which turns the benchmark into a self-check: if seeing the whole video is not better than seeing the clue, the model did not effectively retrieve the clue from the long context.","core_discovery":"The paper's central claim is that multiple-choice accuracy overstates long-video understanding, and that clue-grounded evaluation exposes this. On its benchmark, the best commercial model GPT-4o scores 45.2% on long-video multiple choice, but when the same questions are scored with $\\mathrm{acc@IoU}$—correct option plus a predicted interval with $\\mathrm{tIoU} > 0$ against the annotated clue—the score collapses to 4.38%. The black-box Clue Recovery Rate, $\\mathrm{CRR} = \\min(\\mathrm{long\\text{-}acc.}, \\mathrm{clue\\text{-}acc.})/\\mathrm{clue\\text{-}acc.}$, is below 100% for every model, meaning full-video context dilutes rather than helps clue retrieval relative to watching only the clue clip. The paper concludes that current models can answer many questions without grounding them in the relevant evidence, and that CG-Bench's clue-based metrics provide a more credible measurement.","pith_inferences":["Beyond the paper's claims, the clue-completeness assumption could be tested by asking a second annotation team to find all intervals that support each answer and recomputing CRR with the union of intervals; if CRR rises materially, the original annotations were incomplete.","As an editorial extension, the benchmark's black-box assumption also suggests a direct training signal: regularize models so full-context accuracy never falls below clip-level accuracy, which could improve long-video reliability without new annotations.","Because $\\mathrm{acc@IoU}$ requires overlap with the annotated clue, a model that answers correctly using legitimate alternative evidence is penalized; a future variant could score with soft overlap against any sufficient interval, not just the human-annotated one."],"forward_implications":["MCQ-only long-video leaderboards systematically overstate model capability; adding a clue-interval requirement changes the ranking and widens the gaps between models.","All evaluated models score below 100% CRR, so current MLLMs are not reliably retrieving the relevant moments from long contexts; improving temporal grounding is a concrete training target.","Timestamp information from frames and subtitles is what lets models point at clues: adding both raises GPT-4o's $\\mathrm{mIoU}$ from 3.39 to 9.68 and $\\mathrm{acc@IoU}$ from 10.7 to 26.7.","Human performance on the same sparse 128-frame input is 59.85%, versus 90.3% with full video, indicating the benchmark is difficult and favors models that use long-context visual information rather than sparse frames.","Open-source models such as Qwen2-VL-72B approach GPT-4o's MCQ accuracy (41.3 vs. 45.2) but remain behind on clue grounding, so the open-source gap is mostly a grounding gap."],"supporting_citations":[{"why":"Supplies the visually grounded video-QA idea (question grounding with clue intervals) that CG-Bench transfers to long, open-domain videos.","marker":"Xiao et al., 2024"},{"why":"Provides Video-MME, the long-video MCQ benchmark whose reliance on multiple choice CG-Bench argues is insufficient.","marker":"Fu et al., 2024a"},{"why":"Provides MLVU, a long-video MCQ benchmark compared against CG-Bench in Table 2 as prior work without clue grounding.","marker":"Zhou et al., 2024"},{"why":"Provides LongVideoBench, another long-video benchmark used as comparison and motivation for the clue-grounded design.","marker":"Wu et al., 2024b"},{"why":"Provides MVBench, the short-video MLLM benchmark whose evaluation setting CG-Bench extends to long videos and contrasts with in Table 2.","marker":"Li et al., 2023b"},{"why":"Defines GPT-4o, the strongest model evaluated; its long-acc to acc@IoU drop is the paper's headline evidence.","marker":"OpenAI, 2024"},{"why":"Defines Qwen2-VL, the leading open-source model, demonstrating that the open-source gap is mainly a clue-grounding gap.","marker":"Wang et al., 2024a"}],"fun_headline_variants":["Clue-graded test: GPT-4o long-video score plummets to 4.38%","MCQ overstates understanding: clue test cuts GPT-4o long-video score by 90%","CG-Bench clue test: GPT-4o long-video accuracy drops 90%","Long-video AI flunks clue test: GPT-4o accuracy falls from 45.2% to 4.38%"],"cache_read_input_tokens":20608,"weakest_assumption_plain":"The credibility metrics assume the human-annotated clue intervals are both necessary and sufficient evidence—if a question can be answered from other parts of the video, then a low full-video accuracy relative to clue-clip accuracy may reflect incomplete annotations rather than failed clue retrieval.","fun_headline_variants_meta":{"raw":{"variants":["Clue-graded test: GPT-4o long-video score plummets to 4.38%","MCQ overstates understanding: clue test cuts GPT-4o long-video score by 90%","CG-Bench clue test: GPT-4o long-video accuracy drops 90%","Long-video AI flunks clue test: GPT-4o accuracy falls from 45.2% to 4.38%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000999,"raw_usage":{"total_tokens":4285,"prompt_tokens":1055,"completion_tokens":3230,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":671,"completion_tokens_details":{"reasoning_tokens":3118}},"tokens_in":671,"tokens_out":3230,"duration_ms":22958,"temperature":1.0,"reasoning_tokens":3118,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T14:16:44.946327+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-annotate a random sample of CG-Bench questions with independent annotators and compute inter-annotator $\\mathrm{tIoU}$ on the clue intervals; if many questions have low agreement, or if a model given the full video but with annotated clues masked still beats its clue-only accuracy, then hidden clues exist and the CRR inequality is not a valid test of clue retrieval.","supporting_citations":[],"review_version":1}