{"id":"49c655d2-897b-49e2-af6f-5b3a6391ebb0","arxiv_id":"2606.07433","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"This is a survey that frames video MLLM research via a human-view formulation of perceptual representations, memory states, reasoning traces, and predictions, then reviews methods, datasets, benchmarks, and open problems.","lead":"This survey organizes research on multimodal large language models for video understanding around three human-like abilities: watching, remembering, and reasoning. A smart generalist might read it to see how current AI systems handle long videos, sparse evidence, and the need for memory and inference under compute limits.","discovery_kind":"review","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"The reader's weakest_assumption matches the definitional premise of the survey but does not constitute a correctness risk because the paper makes no testable scientific claim beyond providing a taxonomy. The UNVERDICTED verdict is appropriate for a survey without new results, and the abstract supplies no internal inconsistency or unsupported derivation that would alter it.","tokens_in":1843,"tokens_out":267,"duration_ms":15137,"concrete_test":"Scan the introduction and formulation section for any language claiming the three abilities constitute the unique or empirically validated complete partition of video understanding; if absent, the survey framing holds without further test.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper is a literature survey that proposes an organizational perspective on video MLLMs structured around three functional abilities (watching, remembering, reasoning) and four system components (perceptual representations, memory states, reasoning traces, final predictions). The central claim is that this supplies a unified structure for analyzing existing methods, challenges, datasets, and benchmarks rather than introducing new empirical results or a falsifiable model. The partition is presented as a useful framing for the survey; no stronger claim of exhaustiveness, disjointness, or predictive power is required for the work to fulfill its stated purpose.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"This survey paper proposes a human-view organizational framework for video understanding with multimodal large language models (MLLMs), structured around three functional abilities (watching, remembering, reasoning) and four system components (perceptual representations, memory states, reasoning traces, final predictions). It uses this framing to categorize methods for fine-grained perception, memory modeling (offline/streaming), and reasoning (text-only or video-grounded), while surveying challenges in spatio-temporal perception and long-video processing, application domains (egocentric, sports, medical, etc.), datasets, benchmarks, and open problems.","tokens_in":1944,"tokens_out":388,"duration_ms":11020,"significance":"If the proposed partition proves useful for synthesis, the work could help researchers map existing methods onto a common structure for identifying gaps in memory-aware and evidence-grounded video intelligence, particularly as the field shifts toward long, multimodal scenarios. Its value is primarily in organization and coverage rather than new derivations or measurements.","major_comments":[],"minor_comments":[{"comment":"The formulation characterizing systems by perceptual representations, memory states, reasoning traces, and final predictions is introduced in the abstract and presumably detailed early in the manuscript; if this is presented only descriptively without a diagram or explicit mapping table to the three abilities, it risks remaining informal for readers attempting to apply the framework to new papers.","section":null},{"comment":"The abstract states that representative methods are 'organized by their roles in video MLLM systems' under the three abilities, but without an explicit cross-reference table or section that lists which cited works map to which component, the organizational claim is harder to verify.","section":null},{"comment":"The GitHub link for continuously traced related works is mentioned but not cited as a reference; adding a formal citation or footnote would improve traceability.","section":null}],"recommendation":"minor_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their positive summary of our survey and the recommendation of minor revision. No specific major comments were provided in the report, so we have no individual points requiring point-by-point rebuttal. We will incorporate minor improvements for clarity and completeness in the revised version.","responses":[],"tokens_in":1300,"tokens_out":74,"duration_ms":10646,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"This survey organizes recent work on multimodal LLMs for video around three functional abilities—watching, remembering, and reasoning—plus four system components: perceptual representations, memory states, reasoning traces, and final predictions. The abstract and structure make clear that the goal is to give a unified view of how these models handle evidence, context, and output rather than to report fresh experiments.\n\nIt does a reasonable job pulling together challenges in long-video perception, memory modeling, streaming, and faithful reasoning, then mapping representative methods, datasets, and benchmarks across domains like egocentric and medical video. The GitHub repo for ongoing updates is a practical addition that many readers will use.\n\nThe soft spots are predictable for a survey. The three-ability partition is offered as a useful lens, yet the paper does not demonstrate that the categories are exhaustive or non-overlapping; any overlaps or gaps will only become visible once the full categorization is checked against the cited papers. Soundness therefore depends on accurate mapping of existing methods, which cannot be verified from the abstract alone. No new empirical claims or formal derivations are present, so there is nothing to reproduce or falsify.\n\nThis paper is mainly for people already active in video MLLMs who want a structured overview and a list of open problems. It is not aimed at readers seeking a new model or benchmark numbers. It deserves a serious referee because the subfield moves quickly and a balanced survey can help keep the literature usable, even if the framing itself is incremental.","headline":"This is a survey that organizes video MLLM work around watching, remembering, and reasoning but introduces no new results or derivations.","tokens_in":2495,"tokens_out":374,"would_cite":false,"duration_ms":10424,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Video MLLMs acquire evidence, preserve context, and produce outputs through watching, remembering, and reasoning.","keywords":["video understanding","multimodal large language models","watching remembering reasoning","long video processing","memory modeling","egocentric video","streaming understanding","faithful reasoning"],"falsifier":"A deployed video MLLM whose accuracy, efficiency, or failure modes on long videos cannot be improved or explained by separately measuring or modifying its watching, remembering, or reasoning components.","tokens_in":2743,"feed_emoji":"🎥","tokens_out":484,"duration_ms":14165,"temperature":0.7,"pith_summary":"The paper introduces a unified human-view framework for LLM-based video understanding built around three functional abilities: watching for evidence acquisition, remembering for context preservation, and reasoning for grounded outputs. Instead of isolated task benchmarks, this structure analyzes how models handle sparse evidence, long-range dependencies, and multimodal alignment under computational limits. It formulates video systems through perceptual representations, memory states, reasoning traces, and final predictions, then maps representative methods, challenges, application domains, datasets, and benchmarks onto these elements. The work reviews current approaches across perception, memory, and reasoning while highlighting open problems for scalable video intelligence.","feed_headline":"Three abilities frame how MLLMs understand video","feed_subtitle":"Watching, remembering, and reasoning supply a single structure for analyzing evidence use, context retention, and inference in long multimod","key_machinery":"The three functional abilities—watching, remembering, and reasoning—that partition video MLLM behavior and link the four system components (perceptual representations, memory states, reasoning traces, final predictions) into a unified analysis structure.","core_discovery":"Video understanding with MLLMs is best characterized by a formulation that decomposes systems into perceptual representations, memory states, reasoning traces, and final predictions, which in turn map onto the three abilities of watching, remembering, and reasoning; this decomposition supplies a single lens for organizing methods, identifying challenges in spatio-temporal perception and memory modeling, and covering domains from egocentric to narrative videos.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["MLLMs understand video through watch remember reason","Three abilities unify video MLLM analysis","Watching remembering reasoning decompose video systems","Human view frames MLLM video understanding with three abilities"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"Every video understanding system can be usefully described by four fixed components and partitioned without overlap or remainder into the three abilities of watching, remembering, and reasoning.","fun_headline_variants_meta":{"raw":{"variants":["MLLMs understand video through watch remember reason","Three abilities unify video MLLM analysis","Watching remembering reasoning decompose video systems","Human view frames MLLM video understanding with three abilities"]},"model":"grok-4.3","cost_usd":0.004637,"raw_usage":{"total_tokens":2327,"prompt_tokens":729,"num_sources_used":0,"completion_tokens":55,"cost_in_usd_ticks":46374500,"prompt_tokens_details":{"text_tokens":729,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1543,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":729,"tokens_out":55,"duration_ms":14884,"temperature":1.0,"reasoning_tokens":1543,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-27T21:59:04.020071+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A deployed video MLLM whose accuracy, efficiency, or failure modes on long videos cannot be improved or explained by separately measuring or modifying its watching, remembering, or reasoning components.","supporting_citations":[],"review_version":1}