{"id":"8e1e6bfc-d345-4623-a598-7b4172686153","arxiv_id":"2607.14935","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"An open 4B video MLLM with inflated-3D ViT tokenization and adaptive streaming perception outperforms comparable open models on general, long-video, and streaming benchmarks while using fewer visual tokens.","lead":"VideoChat3 is a fully open 4B video-language model with a new 3D vision encoder and adaptive streaming resolution that claims to beat comparable open models on many video benchmarks. It also releases training code, data, and weights to make video-MLLM training reproducible.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Temporal-grounding gains may be inflated by training on the same TimeLens/Charades-STA/ActivityNet/QVHighlights data used for evaluation; one overlap check could confirm.","rationale":"The reader's weakest assumption is evaluation fairness, and the strongest concrete form of that risk is the TimeLens-100K training/evaluation overlap: the paper explicitly trains on TimeLens-100K in Stage-2 and Stage-3 and evaluates on the TimeLens splits of Charades-STA/ActivityNet/QVHighlights. This is not a subtle protocol choice but a direct data-contamination risk, and it lands exactly on the largest relative improvements in Table 2. The streaming-protocol asymmetry flagged by the reader is also real, but it is more ambiguous because the appendix discloses the coordinate-range change and the paper states that Qwen3-VL is evaluated under the same settings; the TimeLens overlap, if confirmed, is a cleaner and more serious threat. I do not think this warrants rejection: the external general-video benchmarks, the detailed architecture, and the promised open release are independent evidence, and the temporal-grounding claims can be corrected without invalidating the whole system. Conditional acceptance remains the right posture, so the reader's verdict is unchanged.","tokens_in":34815,"tokens_out":7278,"duration_ms":75833,"concrete_test":"Download the released VideoChat3 datasets and TimeLens-100K; compute the exact intersection (by video ID plus query/segment overlap) between TimeLens-100K training samples and the three TimeLens evaluation splits (Charades-STA, ActivityNet Captions, QVHighlights). If any training sample shares a video (or its rewritten supervision derives from an evaluation video), rerun the model after ablating all such samples from Stage-2/3 and compare the TL mIoU columns; a drop larger than ~3 points confirms the reported gains are partly training contamination. If the intersection is empty, the objection is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.3 and Section 4.4 list TimeLens-100K [14] in the Stage-2 and Stage-3 training mixtures, while Section 5.1 evaluates temporal grounding on \"the TimeLens suite [14]\" — i.e., Charades-STA, ActivityNet Captions, and QVHighlights. TimeLens-100K is built from these same source benchmarks, so unless the released splits are provably disjoint at the video/annotation level, the largest differentials in Table 2 — +9.7/+6.4/+8.3 over Qwen3-VL-4B and +16.4/+29.8/+34.3 over VideoChat-Flash-7B on the three TL columns — may reflect train/eval overlap rather than generalization. The paper's only gesture to this is the footnote \"Temporal-grounding columns marked with TL share the TimeLens suite\"; no disclosure or analysis of overlap is provided. Because these are the strongest and most distinctive empirical advantages cited in favor of generalist temporal understanding, the headline claim \"surpasses prior open-source models\" is load-bearing on this comparison.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents VideoChat3, a 4B-parameter fully open video-centric MLLM, with two main contributions: an Inflated 3D Vision Transformer (I3D-ViT) that performs chunk-wise spatiotemporal attention and temporal pooling to reduce visual tokens, and an Adaptive Frame Resolution mechanism for streaming video that toggles per-frame pixel budgets based on predicted response-state tokens. The authors also release three training datasets (VideoChat3-Academic2M, VideoChat3-LV116K, VideoChat3-OL617K), a four-stage training curriculum, model weights, and code. Empirically, the paper claims strong generalization across temporal perception, long-video QA, temporal grounding, and streaming/proactive benchmarks, with particular gains over Qwen3-VL-4B and Molmo2-4B, plus efficiency improvements from reduced visual-token counts.","tokens_in":35103,"tokens_out":5356,"duration_ms":56964,"significance":"If the reported results are taken at face value, this is a substantial systems-and-data contribution: it demonstrates that early spatiotemporal compression in the visual tokenizer can preserve broad video understanding while cutting inference cost, and it provides fully released assets for reproducible research. The ablation isolating the state-transition mask and the dynamic perception-window policy is well designed and informative. However, the paper's strongest empirical claims—especially in temporal grounding and streaming—rest on benchmark comparisons whose train/eval disjointness and protocol symmetry are not established. The contribution is therefore potentially significant but currently conditional on resolving these evaluation concerns.","major_comments":[{"comment":"TimeLens-100K [14] is included in both the Stage-2 and Stage-3 training mixtures, and temporal grounding is evaluated on the TimeLens suite [14] built from Charades-STA, ActivityNet Captions, and QVHighlights. No evidence is provided that the training and evaluation splits are disjoint at the video or annotation level. Because the largest differentials in Table 2 are on the three TimeLens columns (+9.7/+6.4/+8.3 over Qwen3-VL-4B, and +16.4/+29.8/+34.3 over VideoChat-Flash-7B), these gains are load-bearing for the claim of surpassing prior open models in temporal grounding. Please provide exact overlap statistics between TimeLens-100K training samples and each evaluation split, and re-report results after excluding overlapping items or on independent held-out grounding benchmarks.","section":"Sections 4.3/4.4 vs 5.1, Table 2"},{"comment":"Stage-3 training explicitly includes StreamForest [19], while the streaming evaluation includes ODVBench [19], which is the benchmark introduced in the StreamForest paper. The paper does not establish that the StreamForest training data and the ODVBench evaluation clips are disjoint. Similarly, River [87] appears to be an author-created benchmark, and the Stage-3 mixture includes StreamForest/Streamo/Seeker data; no exclusion analysis is reported. Given that Table 3 claims best results on four of six streaming metrics, this potential overlap must be quantified and the results recomputed on non-overlapping data before the streaming claims can be accepted.","section":"Section 4.4 vs 5.2, Table 3"},{"comment":"The evaluation protocol for VideoChat3 is not symmetric with the baselines. For ODVBench and OVBench, the paper modifies the bounding-box output range from [0,1] to [0,1000] when evaluating VideoChat3, while most baseline results are quoted from their original papers. Likewise, Table 12 applies custom low/high token budgets and high-budget frame counts to the proactive-response benchmarks. Since the streaming comparison is cross-paper, such protocol differences could explain part of the reported improvements. Please run at least the strongest open baselines (e.g., StreamForest, StreamingVLM, Qwen3-VL-4B) under the same evaluation harness and output-format conversions, and report all models with matched token/frame budgets.","section":"Appendix Table 11, Table 3"}],"minor_comments":[{"comment":"The PDF contains recurring ligature artifacts such as 'eﬀiciency', 'suﬀicient', and 'oﬀicial'. These should be corrected in the camera-ready version.","section":"Throughout"},{"comment":"Adding VideoChat3-OL617K substantially decreases the ProactiveVQA WEB score (39.0 to 28.4). The text says gains extend across streaming benchmarks without acknowledging this drop. Please add a sentence discussing this trade-off.","section":"Table 9 / Section 5.2"},{"comment":"The term 'fully open' should be qualified: the model weights and training assets are released, but a large fraction of the supervision is generated by proprietary models (Qwen3-VL-235B-A22B and Gemini). The assets are open, but the annotation pipeline is not fully reproducible without access to those closed APIs.","section":"Section 3 / 7"}],"recommendation":"major_revision","confidential_remarks":"I agree with the reader's conditional assessment. The temporal-grounding overlap between TimeLens-100K training data and TimeLens evaluation splits is a concrete, testable concern, and the streaming protocol asymmetries in Appendix Table 11 are also serious. If the authors can demonstrate disjointness and provide matched-protocol comparisons for the key baselines, the paper would be a strong accept candidate; otherwise the headline claims would need substantial qualification."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague — quick take: this is a serious, unusually open systems paper, and the architecture and data work are real. The trouble is that its strongest empirical claims — the temporal-grounding and streaming numbers — sit on likely evaluation leakage and protocol asymmetries. The general-video results are more trustworthy.\n\nNew and good: I3D-ViT is a clean way to get early spatiotemporal compression by inflating a pretrained image ViT to chunk-wise attention with temporal pooling; the 16x token reduction is real. The adaptive frame-resolution controller, where the model's own state tokens decide the next window's pixel budget, is a nice mechanism, and the state-transition masking loss is thoughtfully designed against both overfitting to Silence and trivial transition short-cuts. The data pipeline — re-annotating academic data, synthesizing long-video supervision from segment evidence ledgers, converting offline QA into streaming targets with clue verification — is substantial engineering. And the promised release of weights, code, and data is exactly what the field needs.\n\nThe ablations are thorough and isolate each component. Efficiency numbers are concrete and honestly acknowledge the encoder-latency tradeoff.\n\nWhere it's soft: Section 4.3/4.4 explicitly train on TimeLens-100K, and Section 5.1 evaluates on the TimeLens suite built from Charades-STA, ActivityNet, and QVHighlights. No disjointness proof or overlap analysis is given — just a footnote that the columns share the suite. That directly bears on the paper's headline gains (+9.7/+6.4/+8.3 over Qwen3-VL-4B; +16–34 over VideoChat-Flash). Without an overlap check, those numbers are not evidence of generalization. Likewise, ODVBench/OVBench evaluation changes the bounding-box output range from [0,1] to [0,1000] while baselines are quoted from their original papers — a real protocol mismatch. Minor: no variance, and a 'training throughput advantage' is claimed without reporting actual training time.\n\nThe general-video benchmarks are external and untrained-on, so those comparisons stand. The architecture and data story survive. This is not a fatal flaw; it's a fixable evaluation-hygiene problem. It deserves a serious referee, and I'd want to see the authors address these points in revision.\n\nReading group? I'd bring it. Cite? Yes — the architecture and data pipeline are useful regardless of benchmark numbers. Accept for peer review? Yes, with conditions.","headline":"Serious open systems paper with a genuine architecture/data contribution, but the temporal-grounding and streaming comparisons need a fairness check before the headline numbers can be trusted.","tokens_in":35706,"tokens_out":2564,"would_cite":true,"duration_ms":25007,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"VideoChat3 claims a 4B, fully open video model can beat larger open rivals by compressing visual tokens 16x before they reach the language model, while also handling streaming video.","keywords":["video multimodal large language models","efficient video tokenization","3D vision transformer","temporal grounding","streaming video understanding","instruction tuning","open-source datasets","long video understanding"],"falsifier":"Delete the TimeLens-100K (and any same-source grounding) samples from the Stage-2/3 mixtures, retrain, and re-measure the three TimeLens splits plus VUE-TR and MomentSeeker; if the 6-20 point grounding margins shrink to noise, the temporal-grounding claim is an artifact of train/test overlap. In parallel, re-run the streaming baselines under the same [0,1]-vs-[0,1000] bounding-box normalization and identical evaluation code; if the ODVBench and OVO-Timing gaps disappear, the streaming claim is protocol-driven.","tokens_in":34681,"feed_emoji":"🎥","tokens_out":8187,"duration_ms":71430,"temperature":0.7,"pith_summary":"This paper is trying to establish that a video multimodal language model can be both broadly capable and computationally cheap if video redundancy is removed inside the vision encoder rather than by sparse frame sampling. VideoChat3, at 4B parameters, uses an inflated 3D vision transformer that chunks frames, applies spatiotemporal attention, and temporally pools, achieving a 16x spatiotemporal compression before tokens reach the language model. The authors report that this is enough to surpass prior open models of equal or larger size across general, long-form, and streaming benchmarks: 18 of 19 direct comparisons over a leading 4B baseline, matching or beating another 4B open model on most benchmarks, with roughly half the visual tokens and substantially lower long-video latency. A second mechanism, adaptive frame resolution, lets the model spend more visual budget only after a 'standby' cue during streaming. The claim matters because it would mean efficiency and generality need not trade off, and because the full stack—weights, code, training strategy, and datasets—is released.","feed_headline":"4B video model tops larger open rivals via 16x compression","feed_subtitle":"Fully open 4B model halves visual tokens and beats bigger open rivals on most benchmarks, the paper reports.","key_machinery":"The load-bearing object is the Inflated 3D Vision Transformer (I3D-ViT), a visual tokenizer created by inflating a pretrained image ViT's 2D self-attention into chunk-wise spatiotemporal attention over T-frame windows, adding learned temporal positional embeddings, and then applying chunk-wise temporal pooling. Combined with 2x2 pixel-shuffle spatial downsampling and T=4, it gives a 16x spatiotemporal compression ratio, halving the visual tokens a baseline would produce while preserving motion cues. The second mechanism is Adaptive Frame Resolution: at each streaming step the LLM emits a response-state token (Silence, Standby, Response), and a deterministic controller sets the next window's","core_discovery":"The central discovery, stated on the paper's own terms, is that compressing video early—inside a vision tokenizer that models local space-time before the LLM—preserves motion evidence while cutting LLM context cost. I3D-ViT inflates a pretrained image ViT's 2D attention into chunk-wise 3D spatiotemporal attention with learned temporal positions, then pools over T frames; with T=4 and 2x2 spatial pixel-shuffle downsampling, the visual sequence shrinks 16x. This changes the compute profile: the vision encoder's cost grows roughly linearly in frames, while the LLM's quadratic context cost drops, so long-video inference becomes faster and lighter. The paper further shows the same tokenizer suppo","pith_inferences":["If the 16x early-compression design is the real driver, it should transfer to other base LLMs and vision encoders; a controlled swap experiment would reveal how much of the gain is architectural rather than data-specific.","The adaptive frame-resolution controller could be lifted from streaming into offline long-video pipelines: a salience estimate, rather than a state token, would decide which segments get the high 448² pixel budget, potentially cutting cost without a real-time loop.","Given the TimeLens training/evaluation overlap noted in Sections 4.3 and 5.1, the most informative independent check is a zero-shot run on temporal-grounding benchmarks whose source datasets were never in the training mixture; the reported margins may not transfer.","A cheaper-to-reproduce variant—distilling the synthetic long-video and streaming supervision into a smaller open model—would test whether the state-transition masking and adaptive budget still yield proactive behavior at lower cost."],"forward_implications":["The efficiency claim is concrete: with 256 to 2048 input frames, VideoChat3 produces exactly half the visual tokens of the compared baseline under equal ViT patches per frame, and at 2048 frames the reported end-to-end latency drops from 44.4 s to 20.4 s while FLOPs fall by more than 60%.","Because the model is fully open (weights, code, training strategy, and all three datasets), any research group can replicate the training pipeline and extend it without reverse engineering.","The streaming training recipe—state tokens with a balanced transition mask and a deterministic low/high pixel-budget controller—is the paper's answer to 'when to answer' and can be applied to other video MLLMs.","The evidence-grounded annotation pipeline increases supervision density from sparse academic labels and is reported to improve temporal perception and grounding consistently, suggesting data quality rather than raw scale drives much of the gain."],"fun_headline_variants":["Open 4B video model beats larger rivals via 16x token cut","Fully open 4B video LLM: 16x fewer tokens, beats bigger models","16x compression lets 4B open video model outgeneralize bigger ones","Open-source 4B video model: 16x token drop, top generalization","Efficient 4B video MLLM: fully open, 16x smaller context, beats larger"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The comparative claims stand or fall on the assumption that every model was evaluated under identical protocols with no training/test overlap; the paper's own tables show temporal-grounding training data derived from the same datasets as the temporal-grounding evaluation splits, and streaming baselines are mostly quoted while VideoChat3 uses a modified bounding-box output range.","fun_headline_variants_meta":{"raw":{"variants":["Open 4B video model beats larger rivals via 16x token cut","Fully open 4B video LLM: 16x fewer tokens, beats bigger models","16x compression lets 4B open video model outgeneralize bigger ones","Open-source 4B video model: 16x token drop, top generalization","Efficient 4B video MLLM: fully open, 16x smaller context, beats larger"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000672,"raw_usage":{"total_tokens":2940,"prompt_tokens":829,"completion_tokens":2111,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":573,"completion_tokens_details":{"reasoning_tokens":1999}},"tokens_in":573,"tokens_out":2111,"duration_ms":11831,"temperature":1.0,"reasoning_tokens":1999,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T00:39:13.782404+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Delete the TimeLens-100K (and any same-source grounding) samples from the Stage-2/3 mixtures, retrain, and re-measure the three TimeLens splits plus VUE-TR and MomentSeeker; if the 6-20 point grounding margins shrink to noise, the temporal-grounding claim is an artifact of train/test overlap. In parallel, re-run the streaming baselines under the same [0,1]-vs-[0,1000] bounding-box normalization and identical evaluation code; if the ODVBench and OVO-Timing gaps disappear, the streaming claim is protocol-driven.","supporting_citations":[],"review_version":1}