{"id":"4c265587-bada-4c36-a9c9-21809e0d217e","arxiv_id":"2501.05037","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"LongViTU, an automatically generated 121k-pair video QA dataset with 4.6-minute average certificate length, provides modest SFT gains on long-video benchmarks.","lead":"LongViTU is a new 121k-question dataset for long-form video understanding, generated automatically with a hierarchical video tree and a self-revision step. The authors show that fine-tuning two open-source video models on LongViTU yields small average gains on long-video benchmarks.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 276.8s certificate length is prompt-enforced but never verified; if LongViTU questions are answerable from Ask Content alone, SFT gains may reflect instruction-following rather than long-video understanding.","rationale":"The paper is honest and provides real evidence: a large released dataset, a manually reviewed benchmark subset, and zero-shot/SFT comparisons on several external benchmarks. My stress-test narrows the reader's weakest assumption to the certificate-length claim, which is the property that differentiates LongViTU from prior datasets. The QA prompt in Section C.2 tells the LLM where to place the answer and where to pose the question, but it never verifies that the question is unanswerable from the Ask Content or from language priors. The blind GPT-4 turbo baseline score of 38.2 on LongViTU shows nontrivial language-prior answerability, and the paper's own limitation admits remaining textual bias. The reported average SFT gains are dominated by EgoSchema, an in-domain Ego4D benchmark; the out-of-domain long-video gains are small and reported without error bars. A truncation-based or text-only-Ask-Content test would settle whether the 276.8s certificate length corresponds to a genuine requirement for long-range video memory. If that test passes, the central contribution is supported; if it fails, the SFT improvements should be reinterpreted as generic instruction tuning or domain adaptation. Given that the current evidence is insufficient to reject the claim but also insufficient to move beyond a conditional acceptance, I recommend no change to the reader's verdict.","tokens_in":21158,"tokens_out":5000,"duration_ms":50830,"concrete_test":"On the 600 human-reviewed LongViTU test questions, use the provided timestamps to construct two inputs for LongVU and LLaVA-Video: (a) the full video, and (b) only the frames lying in the Ask Content span (the last two segments). Compare GPT-4 scores with bootstrap confidence intervals. If ask-only accuracy is statistically indistinguishable from full-video accuracy, the certificate-length statistic is not load-bearing. Complement this with a text-only control: give GPT-4-turbo only the Ask Content segment summaries and ask each question; if text-only ask-content scores approach the full-context model scores, the questions are answerable from textual priors. Finally, rerun the headline SFT comparison excluding EgoSchema (the in-domain Ego4D benchmark) and with multiple random seeds to confirm that the remaining OOD gains of roughly 1-4% are not noise.","verdict_should_be":"UNCHANGED","load_bearing_attack":"LongViTU's central distinguishing property is an average certificate length of 276.8s, but that property is only instructed, not measured. In Section 2.1.2, the LLM is told to put answers in 'Memory Content' (first three segments) and questions in 'Ask Content' (last two segments); the prompt in Section C.2 does not require that the question be unanswerable from Ask Content alone or from language priors. The self-revision step filters 'excessive textual bias' with a pure-text evaluation, which cannot establish that the answer requires the earlier video. The paper's own limitation statement (Appendix A) concedes that roughly 9% of QA pairs retain textual bias, and human review covers only the 600-question test subset; the 100-question rubric study is far too small to certify the 121k-pair training distribution. If a substantial fraction of QA pairs can be answered from the last two segments or from generic Ego4D narration, then the reported SFT gains on EgoSchema, VideoMME-Long, MLVU, and LVBench could reflect improved instruction following or domain adaptation rather than genuine integration of long-range temporal context. This directly threatens the central claim that the dataset improves long-video understanding.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces LongViTU, a large-scale automatically generated video QA dataset built from Ego4D videos. The construction pipeline extracts frame-, event-, and segment-level descriptions into a hierarchical video tree, then prompts GPT-4 to generate QA pairs with a sliding window that places answers in the first three segments ('Memory Content') and questions in the last two ('Ask Content'), followed by an LLM self-revision step. The paper claims an average certificate length of 276.8 seconds, explicit timestamp annotations, a condensed-reasoning taxonomy, and high quality validated by human studies. It further reports that supervised fine-tuning of LongVU and LLaVA-Video on LongViTU improves average performance by 2.5% and 3.7% respectively on EgoSchema, VideoMME-Long, MLVU, and LVBench, and it introduces a 600-question human-reviewed test benchmark on which current models score far below human annotators.","tokens_in":21411,"tokens_out":4938,"duration_ms":46591,"significance":"If the certificate-length and quality claims are reliable, LongViTU is a potentially valuable community resource: it is among the first large-scale automatically generated long-video QA datasets with explicit timestamps and a structured reasoning taxonomy, and the reported OOD gains on MLVU and LVBench support the practical utility of the data. The paper's strengths include a clearly described pipeline, reproducible fine-tuning settings, comparisons with both open and proprietary models, and human studies on a subset of the data. However, the central distinguishing property—the 276.8-second certificate length—is prompt-enforced rather than directly verified, and the SFT gains are small in several reported conditions, with selective reporting for MVBench and no significance testing. These issues are load-bearing for the paper's main claims and require revision.","major_comments":[{"comment":"The central distinguishing property, the 276.8-second average certificate length, is instructed by the prompt but never measured. The prompt asks the LLM to place answers in Memory Content and questions in Ask Content, but no check establishes that the question is unanswerable from Ask Content alone or from language priors. The self-revision step (Appendix C.3) is a pure-text evaluation, and Appendix A concedes that roughly 9% of QA pairs retain textual bias. Because human review covers only the 600-question test set and the rubric study uses only 100 questions, the 121k-pair training distribution is not certified. Please add a direct measurement of certificate length: for a random sample of QA pairs, have human raters and a strong blind LLM answer with Ask Content only versus full Memory+Ask content, and report the fraction of questions whose answers require the earlier segments. Without this, the dataset's advertised long-context property remains an assumption rather than an established characteristic.","section":"Section 2.1.2, Appendix C.2"},{"comment":"The claim of 'substantial performance improvements across nearly all' benchmarks is weakened by selective reporting and small effects. The MVBench caption states that the table 'only shows the subc-category that have shown improvement', which is a clear selection bias; the full MVBench results must be reported. In addition, several reported subsets decline, including VideoMME Short (-0.1 for LLaVA-Video, -9.5 for Video-LLaVA), LVBench Summarization (-6.3 for LongVU), MLVU Anomaly Reco. (-1.3 for LongVU), and OpenEQA (-7.1/-12.6 for Video-LLaVA). No error bars, confidence intervals, or significance tests are provided, so gains such as +0.4% on VideoMME Long for LongVU and +0.6% on LVBench average for LongVU are not distinguishable from noise. Please report complete result tables, including all MVBench categories, and add statistical significance measures or per-seed variability.","section":"Table 3, Section 3.3"},{"comment":"The human quality assessment is too small to support the strong wording that the results 'prove the quality' of LongViTU. Only 100 randomly selected questions were rubric-scored, yielding 46% 'Good', 45% 'Fair', and 9% 'Poor', and the comparative assessment against VideoMME also uses only 100 questions per dataset. These samples cannot certify the full 121k-pair corpus, especially because the 600 human-reviewed samples are limited to the test set. Please report confidence intervals for the observed proportions, scale up the human evaluation, or temper the conclusion to reflect that the quality evidence comes from a small sample.","section":"Section 2.3, Figure 4"}],"minor_comments":[{"comment":"The abstract contains a doubled closing parenthesis 'etc.)).'; please fix the punctuation.","section":"Abstract"},{"comment":"The 'Overall Avg.' for Human is 81.0, but the mean of the three category averages (84.1, 74.3, and 75.5) is approximately 78.0; please clarify how the overall average is computed.","section":"Table 2"},{"comment":"The text says 'LLaVA-Video SFT improved by 1% on VideoMME Long', but Table 3 reports +0.3% for that condition; please align the text with the table.","section":"Section 3.3"},{"comment":"The caption refers to a 'bottom horizontal axis' while the figure appears to have both top and bottom horizontal axes with different scales; please make the axis mapping explicit and improve readability.","section":"Figure 3a"},{"comment":"The MVBench caption contains the typo 'subc-category'; please correct it to 'sub-category'.","section":"Table 3"},{"comment":"The word 'significant' is used repeatedly in a non-statistical sense; please reserve it for cases where significance tests are actually reported.","section":"Section 3.3"}],"recommendation":"major_revision","confidential_remarks":"The dataset and pipeline are potentially useful to the video understanding community, and the external-benchmark validation is a positive feature. However, the unverified certificate-length claim and the selective MVBench reporting are substantial. If the authors can provide a direct answerability study for certificate length and full benchmark results with uncertainty estimates, the paper could become acceptable. The EgoSchema result should be interpreted carefully because it shares the Ego4D source with the training data."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The dataset is real, the construction pipeline is genuinely new, and the SFT results on external benchmarks are credible but modest. The main thing to know: the headline 276.8-second certificate length is a design instruction given to the LLM, not a measured property of the resulting QA pairs. The paper does not verify that questions are unanswerable from the \"Ask Content\" segments alone, and its own limitation statement concedes about 9% still carry textual bias. That doesn't sink the paper, but it means the core selling point is weaker than advertised.\n\nWhat's actually new and good: the hierarchical tree representation (frame → event → segment) is a sensible way to compress long Ego4D videos for LLM-based QA generation. The condensed-reasoning taxonomy is a useful addition, and explicit timestamp annotations are rare in this space. The dataset itself—121k QA pairs, 900 hours—is a real resource, and the authors are honest that human review covers only 600 test questions. The SFT gains on EgoSchema, MLVU, and LVBench, while small, are consistent across two base models and are measured on benchmarks that don't use GPT-4 as a judge, so the central empirical claim is independently grounded.\n\nThe soft spots are proportional. First, the certificate length is asserted but not validated; a small human study could check whether questions truly require the earlier segments, and that would strengthen the paper considerably. Second, the gains are within noise territory given no error bars or significance tests; several subsets decline (VideoMME Short, OpenEQA), and the paper's own Table 3 explicitly reports only improved MVBench subcategories, which is selective reporting, even if transparent. Third, the GPT-4-as-judge on the LongViTU test set is a known circularity risk, but the external benchmarks mitigate it.\n\nWho should read this: anyone working on long-video understanding or synthetic video instruction data. The dataset is likely to be reused, and the pipeline is worth borrowing even if the certificate-length claim needs independent verification. The paper deserves a serious referee. I'd send it to review, with the expectation that the authors add a certificate-length verification study and report confidence intervals.","headline":"LongViTU is a genuinely useful dataset with a clever construction pipeline, but the headline 4.6-minute certificate length is a design instruction, not a verified property.","tokens_in":21966,"tokens_out":3094,"would_cite":true,"duration_ms":29666,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that LongViTU, a 121k-pair automatically generated video dataset with an average certificate length of 276.8 seconds, is high-quality, and that supervised fine-tuning on it lifts long-video understanding performance by…","keywords":["long-form video understanding","instruction tuning","video question answering","dataset generation","egocentric video","certificate length","hierarchical video representation","supervised fine-tuning"],"falsifier":"A concrete check: take the LongViTU test set, give a strong model only the last two segments—the 'ask content'—or only a text transcript of the full video, and measure GPT-4 scores; if scores remain high, the claimed certificate length collapses. Alternatively, re-run the human rubric evaluation on a much larger random sample than the 100 questions used and count the share of questions judged answerable without the earlier segments.","tokens_in":20970,"feed_emoji":"🎬","tokens_out":4146,"duration_ms":36164,"temperature":0.7,"pith_summary":"This paper introduces LongViTU, a machine-generated dataset of roughly 121,000 video question-answer pairs drawn from about 900 hours of egocentric video, and argues that it is a high-quality resource for training models to understand long videos. The dataset is engineered so that each question genuinely needs a long temporal context: on average, a model must watch 276.8 seconds of video to answer, far longer than most existing benchmarks. The paper reports that fine-tuning two open video-language models on LongViTU improves their average performance by 2.5% and 3.7% across four long-video benchmarks, and that human reviewers rate the generated questions as close in quality to a manually annotated benchmark. The claim, in short, is that long-form video understanding can be advanced at scale by automatically synthesizing instruction data from video summaries rather than hand-annotating it.","feed_headline":"121k auto-made video QAs push long-video scores up 2.5-3.7%","feed_subtitle":"Questions average ~4.6 minutes of needed context, and fine-tuning lifts results across four long-video benchmarks.","key_machinery":"The central object is the hierarchical video tree, which organizes about 900 hours of egocentric video into frame-level dense captions, event-level summaries, and segment-level summaries, each anchored with explicit start and end timestamps. QA generation runs on a five-segment sliding window with a strict dependency rule: the answer must come from the first three segments and the question from the last two, guaranteeing long certificate lengths. A self-revision stage uses pure-text evaluation to discard QA pairs that can be answered without watching the video.","core_discovery":"The central discovery is that a hierarchical, tree-structured representation of a long video—dense frame captions condensed into event descriptions, then merged into segment summaries—lets an LLM generate QA pairs whose certificate length averages 276.8 seconds. The paper's key mechanism forces long dependency: the question is posed using only the last two of five segments, while the answer must be extracted from the first three, so a model must look back several minutes of video. Combined with a reasoning taxonomy that pushes questions into categories like causality, planning, and risk, and a self-revision pass that filters text-only answerable pairs, the pipeline produces data on which supervised fine-tuning yields gains on both in-distribution and out-of-distribution long-video benchmarks. The human study places the best fine-tuned model at a GPT-4 score of 55.9 against a human score of 81.0, evidence that the questions remain hard.","pith_inferences":["If long-form instruction data can be synthesized this way, the bottleneck shifts from annotation cost to the quality of base video captions, so stronger caption models should directly raise the ceiling of generated QA quality.","Certificate length could be used as a training signal: preferentially sampling QA pairs with longer certificate lengths may yield even larger gains for long-video capabilities than the uniform dataset does.","The tree-and-window recipe may transfer to other long-horizon domains, such as long audio streams or embodied trajectories, wherever hierarchical summaries of the input can be built."],"forward_implications":["Fine-tuning on LongViTU improves EgoSchema accuracy by 4.7% for LongVU and 9.6% for LLaVA-Video, with larger relative gains on longer video subsets.","LongViTU questions remain far from solved: the best fine-tuned open model scores 55.9 versus a human 81.0, and the proprietary Gemini-1.5-Pro scores only 52.3 zero-shot.","Every QA pair carries explicit timestamps for the events it refers to, enabling future work on temporal grounding and localization in long videos.","The self-revision filter and structured reasoning taxonomy reduce, though do not eliminate, textual bias; the paper reports about 9% of QA pairs still retain such bias."],"supporting_citations":[{"why":"Supplies the source videos and human event annotations the whole dataset is built on.","marker":"[19]"},{"why":"Performs per-frame dense captioning at the bottom of the hierarchical video tree.","marker":"[13]"},{"why":"Defines certificate length and serves as the in-distribution long-video benchmark for fine-tuning gains.","marker":"[36]"},{"why":"The open-source long-context model that is fine-tuned on LongViTU and evaluated across benchmarks.","marker":"[43]"},{"why":"The video instruction-tuned model that is fine-tuned on LongViTU and evaluated across benchmarks.","marker":"[64]"},{"why":"Human-annotated benchmark used both for comparative quality assessment and as an out-of-distribution evaluation set.","marker":"[16]"},{"why":"An out-of-distribution long-video benchmark used to measure the generalization of fine-tuned models.","marker":"[49]"}],"fun_headline_variants":["121k video QAs with 4.6-min context boost VLM scores 2.5-3.7%","Tree-structured video QAs force 4.6-min lookbacks, lift benchmarks","Self-revised long-video QAs make models recall 4.6 minutes earlier","LongViTU: 121k auto-built QAs that need 4.6-min context","Fine-tune on LongViTU to lift long-video scores by 3.7%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole pipeline assumes that the text summaries fed to the LLM preserve enough visual and temporal detail that the generated questions genuinely depend on watching the long video, rather than being answerable from language priors or from the later segments alone; the paper concedes around 9% of pairs retain textual bias.","fun_headline_variants_meta":{"raw":{"variants":["121k video QAs with 4.6-min context boost VLM scores 2.5-3.7%","Tree-structured video QAs force 4.6-min lookbacks, lift benchmarks","Self-revised long-video QAs make models recall 4.6 minutes earlier","LongViTU: 121k auto-built QAs that need 4.6-min context","Fine-tune on LongViTU to lift long-video scores by 3.7%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00107,"raw_usage":{"total_tokens":4520,"prompt_tokens":1019,"completion_tokens":3501,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":635,"completion_tokens_details":{"reasoning_tokens":3379}},"tokens_in":635,"tokens_out":3501,"duration_ms":24679,"temperature":1.0,"reasoning_tokens":3379,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:19:38.287591+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete check: take the LongViTU test set, give a strong model only the last two segments—the 'ask content'—or only a text transcript of the full video, and measure GPT-4 scores; if scores remain high, the claimed certificate length collapses. Alternatively, re-run the human rubric evaluation on a much larger random sample than the 100 questions used and count the share of questions judged answerable without the earlier segments.","supporting_citations":[{"cited_title":"Ego4d: Around the world in 3,000 hours of egocentric video","cited_arxiv_id":null,"evidence_quote":"Supplies the source videos and human event annotations the whole dataset is built on."},{"cited_title":"Egoschema: A diagnostic benchmark for very long- form video language understanding","cited_arxiv_id":null,"evidence_quote":"Defines certificate length and serves as the in-distribution long-video benchmark for fine-tuning gains."}],"review_version":1}