{"id":"14559b14-96fd-4b09-b927-9f92af6bcf02","arxiv_id":"2606.11913","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Long videos are encoded once into compact attachable neural weights via agentic distillation, enabling low-latency multi-turn understanding on frozen VLMs.","lead":"The paper proposes representing long videos as small Neural Knowledge Representations (NKR) — portions of network weights distilled from video content via an agent that generates dense descriptions and QA pairs. This allows the NKR to be mounted on a frozen VLM for fast, query-based understanding without re-encoding or reloading the original video.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"AKD's synthesized descriptions/QA pairs may omit fine-grained or unprompted video details, undermining the no-loss encapsulation needed for arbitrary queries","rationale":"The reader's weakest_assumption directly matches the load-bearing step; the abstract-only review already flags the missing verification of information completeness. No stronger internal inconsistency appears from the given text.","tokens_in":1698,"tokens_out":310,"duration_ms":13687,"concrete_test":"On a 10-video LVBench subset, run AKD, then evaluate NKR accuracy on a held-out query set constructed by (a) human annotators watching the raw video and (b) an independent VLM prompted on the full video tokens; if accuracy gap >15% relative to the paper's reported LVBench numbers, the encapsulation assumption fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that a fixed-size NKR, optimized solely via agent-generated dense descriptions and QA pairs, retains enough semantic content for the frozen VLM to answer diverse, previously unseen questions at SOTA-comparable accuracy. The abstract provides no evidence that the agentic synthesis is exhaustive (e.g., no coverage metrics, no comparison against human annotations or full-video token baselines on missed events). If synthesis systematically drops temporal ordering, rare objects, or implicit causal relations, the latency win comes at the cost of the performance claim. This is the precise point where the \"decouples video length from inference cost\" argument is least secured.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper proposes a new paradigm for long-video understanding in which videos are encoded as Neural Knowledge Representations (NKR): small, fixed portions of network weights that are optimized once via Agentic Knowledge Distillation (AKD). In AKD an agent automatically generates dense descriptions and QA pairs to distill video semantics into the NKR weights. At inference the NKR is mounted on a frozen VLM, enabling query-based understanding without reloading or re-encoding the original video. The central empirical claim is that this yields performance comparable to state-of-the-art methods on the LVBench benchmark while reducing end-to-end latency by more than two orders of magnitude.","tokens_in":1869,"tokens_out":456,"duration_ms":27343,"significance":"If the performance and latency claims are substantiated, the work would demonstrate a practical route to amortized, length-independent inference for long-video tasks. The portable, weight-based representation of video content is a clear departure from token-stream or database approaches and could support interactive multi-turn applications that current methods cannot sustain at scale.","major_comments":[{"comment":"Abstract: the claim that the method 'achieves performance comparable to state-of-the-art approaches' is presented without any reported baselines, ablation results on NKR size, error bars, or statistical tests. Because the central contribution is the performance-latency tradeoff, the absence of these details renders the claim impossible to evaluate from the manuscript.","section":"Abstract"},{"comment":"Agentic Knowledge Distillation section: no coverage metrics, human-annotation comparisons, or ablation against full-video token baselines are reported to test whether the agent-synthesized descriptions and QA pairs retain fine-grained temporal ordering, rare objects, or implicit causal relations. This directly bears on the assumption that a fixed-size NKR supports accurate answers to arbitrary unseen queries without significant information loss.","section":"Agentic Knowledge Distillation"}],"minor_comments":[{"comment":"The manuscript introduces the acronym NKR without an explicit comparison table to prior knowledge-distillation or parameter-efficient adaptation techniques, which would help situate the novelty.","section":null}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback. We address each major comment below and will revise the manuscript to provide the requested supporting details and analyses.","responses":[{"response":"We agree the abstract claim requires more explicit support. The full manuscript contains LVBench results with SOTA comparisons in Section 4 and Table 2. In revision we will add NKR-size ablations, error bars, and statistical tests to the abstract, results, and a new appendix to directly substantiate the performance-latency tradeoff.","revision_made":"yes","referee_comment":"[Abstract] Abstract: the claim that the method 'achieves performance comparable to state-of-the-art approaches' is presented without any reported baselines, ablation results on NKR size, error bars, or statistical tests. Because the central contribution is the performance-latency tradeoff, the absence of these details renders the claim impossible to evaluate from the manuscript."},{"response":"The referee is correct that the AKD section lacks these explicit validation metrics and ablations. While end-to-end LVBench performance provides indirect evidence of retention, we will add coverage metrics, human-annotation comparisons, and full-video token baseline ablations in the revised AKD section and appendix to directly address retention of temporal ordering, rare objects, and causal relations.","revision_made":"yes","referee_comment":"[Agentic Knowledge Distillation] Agentic Knowledge Distillation section: no coverage metrics, human-annotation comparisons, or ablation against full-video token baselines are reported to test whether the agent-synthesized descriptions and QA pairs retain fine-grained temporal ordering, rare objects, or implicit causal relations. This directly bears on the assumption that a fixed-size NKR supports accurate answers to arbitrary unseen queries without significant information loss."}],"tokens_in":1378,"tokens_out":386,"duration_ms":22125,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main takeaway is that this paper encodes a long video as a small attachable chunk of neural weights, distilled once via an agent that generates descriptions and QA pairs, then mounted on a frozen VLM for query-time use without reprocessing the original footage.\n\nWhat is new is the framing of video content as NKR weights rather than token streams or static databases, plus the AKD process that automates the one-time encoding. The approach directly targets amortized cost for repeated queries on the same video.\n\nIt does a solid job stating the practical problem and sketching a solution that keeps the backbone frozen while claiming over two orders of magnitude latency reduction on LVBench with comparable accuracy. That direction is worth attention if the numbers hold.\n\nThe soft spot is the lack of visible checks on whether the agent synthesis actually retains everything needed. The abstract offers no coverage metrics, ablations on NKR size, or comparisons against full-video baselines for missed events or temporal relations. If the generated descriptions systematically drop unprompted details, the latency win could trade off against the performance claim. Experimental details on baselines and error bars are also absent from what is shown.\n\nThis is for people working on deployable long-video VLMs who need lower per-query cost. A reader focused on representation tricks would get value from the paradigm even before the results are fully vetted.\n\nI would send it to peer review so the full experiments and implementation can be examined.","headline":"The NKR idea is a clean conceptual shift for long-video efficiency but the abstract leaves the distillation reliability unproven.","tokens_in":2343,"tokens_out":362,"would_cite":false,"duration_ms":27373,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"A long video can be encoded as a small set of network weights that a frozen VLM uses for accurate query answering without reprocessing the video.","keywords":["long video understanding","neural knowledge representation","vision-language models","knowledge distillation","agentic distillation","low-latency inference","LVBench","frozen VLM"],"falsifier":"A controlled experiment on videos containing subtle temporal or visual details where accuracy on targeted questions drops sharply for the NKR method compared with full-video baselines.","tokens_in":2612,"feed_emoji":"⚡","tokens_out":648,"duration_ms":25557,"temperature":0.7,"pith_summary":"The paper proposes representing each long video as a Neural Knowledge Representation consisting of a small portion of network weights rather than tokens or databases. These weights are optimized once through Agentic Knowledge Distillation, in which an agent automatically generates dense descriptions and question-answer pairs to transfer the video's semantic content. At inference time the weights attach to a frozen Vision-Language Model, enabling direct query-based answers while the original video stays unloaded. This design decouples video length from per-query cost and yields high efficiency when the same video is queried multiple times. On the LVBench benchmark the method reaches accuracy levels comparable to current state-of-the-art systems while cutting end-to-end latency by more than two orders of magnitude.","feed_headline":"Video content distilled into weights for 100x faster queries","feed_subtitle":"A compact set of optimized weights mounts on a frozen VLM to answer questions without reloading or re-encoding the original video.","key_machinery":"Neural Knowledge Representation (NKR): a small, optimizable portion of network weights attached to the VLM backbone that encapsulates the video's semantic content.","core_discovery":"By distilling a video's content into a compact Neural Knowledge Representation via Agentic Knowledge Distillation and then mounting those weights on a frozen VLM, the approach supports accurate, query-driven understanding of arbitrarily long videos with end-to-end latency reduced by over two orders of magnitude and performance comparable to existing methods on LVBench.","pith_inferences":["The same distillation pattern could be applied to long audio or document collections by swapping the underlying modality encoder.","Pre-computed NKRs could be shared or cached like model checkpoints, enabling collaborative analysis without raw video transfer.","Performance on videos with high event density or rare objects would provide a direct test of encapsulation limits."],"forward_implications":["End-to-end latency drops by more than two orders of magnitude on LVBench.","Accuracy remains comparable to state-of-the-art long-video methods.","Video length no longer determines inference cost.","The NKR becomes a portable, reusable asset for repeated or multi-turn queries.","Amortized cost for interactive video understanding becomes much lower."],"fun_headline_variants":["Video content encoded as NKR weights attached to VLM","NKR created via AKD for portable video knowledge","Query videos by mounting NKR weights without re-encoding","NKR decouples long video length from VLM inference cost"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The distillation process can pack the video's complete semantic information into a small fixed set of weights without loss that would degrade accuracy on diverse queries.","fun_headline_variants_meta":{"raw":{"variants":["Video content encoded as NKR weights attached to VLM","NKR created via AKD for portable video knowledge","Query videos by mounting NKR weights without re-encoding","NKR decouples long video length from VLM inference cost"]},"model":"grok-4.3","cost_usd":0.010936,"raw_usage":{"total_tokens":4803,"prompt_tokens":642,"num_sources_used":0,"completion_tokens":57,"cost_in_usd_ticks":109362000,"prompt_tokens_details":{"text_tokens":642,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":4104,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":642,"tokens_out":57,"duration_ms":40135,"temperature":1.0,"reasoning_tokens":4104,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-27T10:07:54.825542+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A controlled experiment on videos containing subtle temporal or visual details where accuracy on targeted questions drops sharply for the NKR method compared with full-video baselines.","supporting_citations":[],"review_version":1}