Pith. sign in

REVIEW 4 major objections 6 minor 53 references

This paper claims that aggressive frame compression applied before the vision encoder—guided by inter-frame similarity and enriched with decay-weighted temporal blending—cuts video-LLM computation by 40–50% while holding or slightly improvi

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 11:45 UTC pith:GEYPGAGU

load-bearing objection Simple, honest pre-encoding frame reduction for video LLMs that delivers the FLOPs savings it claims, but the accuracy-maintenance story is weaker than the abstract implies and the hyperparameters look tuned on the test benchmarks. the 4 major comments →

arxiv 2607.22726 v1 pith:GEYPGAGU submitted 2026-07-22 cs.CV

PCA: Persistence-Aware Compression and Aggregation for Fast Video Large Language Models

classification cs.CV
keywords video large language modelsframe compressiontraining-free accelerationtemporal aggregationdynamic downsamplingvisual persistenceinference efficiencyMVBench
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that the dominant inefficiency in video large language models is the encoding and reasoning over redundant frames, and that compression should happen on raw frames before the vision encoder, not on tokens after it. It introduces a training-free, plug-and-play pipeline: Dynamic Downsampling keeps a uniform sample of frames plus the frames most dissimilar to their predecessors, and Persistence-Aware Motion Enhancement replaces each kept frame with a decay-weighted blend of itself and recent neighbors, mimicking human visual persistence. On a 0.5B and 7B video-LLM backbone, this reduces total FLOPs by 40–50% while maintaining or slightly improving accuracy on MVBench, PerceptionTest, YouCook2, and VideoMME, yielding 1.8–2.5× end-to-end speedups. The method generalizes to other VLLM architectures and degrades gracefully under sparse frame conditions.

Core claim

The central claim is that frame-level redundancy in video inputs can be removed before visual encoding without harming—and sometimes improving—video understanding, provided the surviving frames are made persistence-aware by aggregating short-range temporal context. The paper shows that a similarity-based selector (cosine similarity between flattened consecutive frames) combined with a decay-weighted temporal blending of neighbors preserves the cues VLLMs need for temporal reasoning, while cutting computation in both the encoder and the language model. In experiments, PCA reduces total FLOPs to roughly half of baseline on a 7B video-LLM (41–54% of original FLOPs depending on settings) and rai

What carries the argument

The argument rests on two mechanisms. Dynamic Downsampling (DD) selects frames by taking every k-th frame and then restoring the p fraction of frames whose cosine similarity to the previous frame is lowest, so scene transitions and motion onsets survive uniform subsampling. Persistence-Aware Motion Enhancement (PAME) then replaces each selected frame with a weighted average of itself and up to J preceding frames, where weights decay exponentially with temporal distance (α^(t_k−i)); this encodes short-term motion cues into a single still image. PAME is the carrier of the paper's core idea: a cheap, training-free pixel-level aggregation that compensates for dropped frames.

Load-bearing premise

The frame selector assumes that high cosine similarity between consecutive raw frames means redundancy, and low similarity means important new content; if that mapping fails—during pans, zooms, or lighting changes—the wrong frames get dropped or kept.

What would settle it

Take a video with constant camera pan across a static scene: adjacent frames have low pixel similarity yet no new semantic content, so PCA would restore many 'dissimilar' frames and waste its frame budget, while a pure uniform sampler at the same rate would do as well. An ablation on such a benchmark comparing PCA's selection to random selection at matched frame counts would settle whether similarity-based selection earns its complexity.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Compressing before the vision encoder removes the encoder itself as a bottleneck, which token-pruning methods leave untouched; PCA's speedup grows as video length or resolution rises because its own overhead is linear while encoder/LLM cost is superlinear.
  • Because both modules are training-free and operate on raw frames, PCA can be dropped onto any VLLM; the paper demonstrates gains on three distinct VLLM architectures.
  • PCA degrades gracefully under sparse input: at 25–50% frame rates it keeps higher accuracy than token-pruning baselines, suggesting the aggregation step reconstructs missing temporal context.
  • On the longest videos in MVBench, PCA stays within 1.6 points of the full baseline, indicating the local aggregation window prevents error accumulation over long horizons.
  • The decay coefficient matters: uniform averaging (α=1) collapses accuracy, while moderate exponential decay (α around 0.1) is stable, indicating that recent frames should dominate the blend.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The similarity proxy in DD is a risk: cosine similarity between raw pixels will flag camera pans and global illumination shifts as 'novel' even when no semantic event occurs, so the restored-frame budget may be spent on uninformative transitions; a semantic or object-level similarity could be a direct improvement.
  • PAME is a form of temporal low-pass filtering on pixels; one could test whether applying it before uniform downsampling rather than after, or integrating it with learned temporal modules, would yield further gains.
  • The method suggests a broader design principle: in video LLMs, input-side compression that preserves temporal context can outperform output-side token pruning; this encourages revisiting where in the pipeline redundancy reduction should occur.
  • A natural extension is streaming video, where PAME's causal weighted window fits online inference; the paper's current experiments are offline, but the module's local nature makes it compatible with frame-by-frame processing.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper proposes PCA, a training-free, pre-encoding frame-compression method for Video LLMs. It consists of Dynamic Downsampling (DD), which uniformly samples frames and then restores a fraction of the frames with lowest raw-pixel cosine similarity to their predecessors, and Persistence-Aware Motion Enhancement (PAME), which replaces each selected frame with a weighted temporal average of its preceding frames. Experiments on MVBench, PerceptionTest, YouCook2, and VideoMME with LLaVA-OV-0.5B/7B, plus MVBench results on VideoChatGPT and InternVL2, report a 40–50% total FLOPs reduction and a 1.8–2.5× speedup with accuracy roughly maintained relative to the untouched baseline and improved relative to token-pruning baselines.

Significance. If the results hold, this is a useful contribution: it targets the vision-encoder cost that post-encoding token-pruning methods cannot reduce, it is training-free and plug-and-play in design, and it is accompanied by open-source code. The ablations isolate the two modules, and the random-pruning comparison in Table 2 gives some evidence that the similarity heuristic carries signal rather than acting merely as a frame-count reducer. The derivation is simple and the computational reasoning is sound. The main open question is whether the reported accuracy retention is an artifact of hyperparameter selection on the evaluation benchmarks, and whether the generality claims in the abstract and Section 4.6 are supported by the data.

major comments (4)
  1. [§4.4, Tables 1–3] The accuracy-retention claim rests on configurations whose hyperparameters appear to be selected on the reported benchmarks. Table 1 reports only two PCA settings (K=2,J=3,P=0.2 and K=3,J=3,P=0.4), while Table 2 sweeps K∈{2,3,4}, J∈{2,4}, P∈{0.4,0.6} on the same MVBench/PercepTest/YouCook2/VideoMME numbers, and Table 3 sweeps α on MVBench/VideoMME. No development/held-out split, repeated runs, or error bars are given. Several decisive margins are small relative to this selection effect: e.g., Table 1 7B Avg 36.96 vs baseline 37.37 and MVBench 56.69 vs 57.25; Table 3 α=0.2 gives 56.55 vs α=0.1 56.69. Selection on the test set can therefore fully account for the 'maintaining accuracy' conclusion. Please provide an unbiased protocol (e.g., tune on a subset, use cross-validation, or report mean±std across seeds) for the reported configurations.
  2. [Abstract and §4.6, Table 4] The abstract claims PCA 'enhances the performance of the baseline model' and 'consistently outperforms existing state-of-the-art approaches in both efficiency and accuracy.' Table 4 contradicts the first part: on MVBench, InternVL2 origin 40.27 → PCA 39.07 and VideoChatGPT origin 58.60 → PCA 57.22. Table 1 also shows several drops relative to origin (e.g., 0.5B YouCook2 METEOR 8.31 vs 9.03; 7B VideoMME wo-subs 55.70 vs 57.89). PCA is better than the token-pruning baselines at the reported compression, but the stronger claim that baseline performance is enhanced is not supported. Please revise the abstract and conclusion to state the trade-offs precisely, or provide additional evidence for the stronger claim.
  3. [§4.3–§4.4, Table 1] The FLOPs comparison is asymmetric by construction: token-pruning baselines reduce only LLM-side tokens, leaving encoder FLOPs unchanged, while PCA reduces frames before encoding. This explains much of the efficiency margin and is not itself a flaw. However, the accuracy comparison is only interpretable if compute budgets are actually matched. The paper says pruning ratios are 'determined based on the total FLOPs' (§4.3) but does not state the procedure or the per-method settings; Table 1 reports one operating point per baseline. Please report accuracy-vs-FLOPs curves, or at least matched-FLOPs operating points, and state how each baseline's pruning strength was chosen, including whether it was tuned on the same benchmarks.
  4. [§3.3, Eq. (6)] The DD module assumes that raw-pixel cosine similarity between consecutive frames identifies semantically important transitions. The random-pruning ablation (Table 2) suggests this heuristic carries signal on the tested benchmarks, so I do not view this as an internal inconsistency. Nevertheless, raw-pixel similarity is sensitive to camera motion, illumination changes, and compression artifacts, which can produce large pixel differences without new semantic content. A concrete robustness test would help: evaluate on videos with controlled camera pan/zoom or photometric perturbation and compare DD against random selection at matched frame counts. If the gap closes, the paper should state the scope of validity. This concern is secondary to the validation-protocol issue but matters for the generality claim.
minor comments (6)
  1. [Throughout] There are several typos and spacing artifacts: 'Persistence-Aware Motion Enhencement' in Section 1, 'Frame Rate Radio' in Figure 6, 'LLA V A-OV-7B' in Figure 2, and 'na¨ıve' spacing in Section 4.5. Please proofread.
  2. [Eqs. (10)–(11) and Algorithm 1] The weighting in Eq. (11) uses the condition t_k-i < j, while Algorithm 1 sets start_idx = max(0, h.idx - window) and then includes the frame at distance exactly window. These definitions are inconsistent for the boundary frame. Also, α=0 in Table 3 involves 0^0 for the current frame; please clarify the intended convention.
  3. [References] References [5] and [6] are identical (both are 'Expanding performance boundaries of open-source multimodal models...'), but §4.2 cites InternVL2 as [6]. The cited reference does not appear to be the actual InternVL2 paper. Please correct the citation and remove the duplicate.
  4. [Eq. (7)] The notation in Eq. (7), i / ||F∖F′|| < p, is not self-explanatory because i is used as an index into the sorted sequence P. Please clarify that this selects the lowest-similarity p-fraction of the pairs in P.
  5. [Table 2] The Random Pruning baseline is not compute-matched to the PCA rows (e.g., 0.5B: random 8.3 TFLOPs vs PCA 5.7–9.3; 7B: random 50.1 vs PCA 35.3–56.9). Consider reporting a random-pruning curve across several frame counts so the comparison is not confounded with FLOPs.
  6. [§3.1, Eqs. (1)–(2)] The derivation of ΔV_rel is correct but compressed. Defining V explicitly as responses per unit time and stating that the frame ratio is α would improve readability.

Circularity Check

0 steps flagged

No significant circularity: PCA is a constructive heuristic with empirical evaluation; the claimed predictions are not forced by construction.

full rationale

I walked the paper's derivation chain. Equations (1)–(2) are plain algebra relating frame-rate reduction to relative speed improvement; they are not predictions and do not feed back into the accuracy claims. The Dynamic Downsampling module (Eqs. 3–8) defines a frame-selection rule based on cosine similarity between flattened pixel vectors (Eq. 6). This is a heuristic selection procedure, not a derivation of accuracy: the claim that low similarity corresponds to informative frames is an empirical assumption evaluated in Tables 1–4 and Figure 8, not an equation whose output equals its input. The Persistence-Aware Motion Enhancement module (Eqs. 9–12) is likewise an explicit weighted-averaging definition; it enriches frames but does not mathematically imply benchmark performance. No fitted parameter is renamed as a prediction: K, J, P, and alpha are discrete hyperparameters, and the reported configurations are empirical choices. The concern that these hyperparameters may have been selected using the reported test benchmarks is a real evaluation-validity risk, but it is not a circularity in the sense of an equation reducing to its own inputs, and the paper does not describe a fitting procedure that would make the benchmark numbers forced. The paper's self-citations (e.g., Phys-LLM, CAT+, ROD-MLLM, PHASE-Net, Multimodal Deception Detection) are listed as related prior work; none is used as a load-bearing uniqueness theorem or as the sole justification for the central efficiency/accuracy claim. The central claim rests on the experimental tables, which are external empirical evidence rather than a self-referential derivation. Therefore no specific circular step can be quoted, and the appropriate score is 0.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 0 invented entities

The central claim rests on a small set of hyperparameters that are tuned on the reported benchmarks, and on two domain assumptions: pixel-level similarity as an informativeness signal and pixel-level temporal averaging as an information-preserving enhancement. No new physical or mathematical entities are introduced.

free parameters (4)
  • K (uniform downsampling step) = 2 or 3
    Chosen based on ablations (Table 2) to balance FLOPs and accuracy; no separate validation split.
  • J (temporal window size) = 3
    Ablated in Table 2 with values 2, 3, 4; J=3 used in main results.
  • P (restore rate) = 0.2 or 0.4
    Ablated in Table 2 with 0.4, 0.6; main results use 0.2 and 0.4.
  • α (decay coefficient) = 0.1
    Sensitivity analysis in Table 3 shows α=0.1 gives best MVBench; α=1 is much worse.
axioms (5)
  • domain assumption Cosine similarity between flattened pixel vectors (Eq. 6) is a valid proxy for semantic informativeness of frames.
    The Dynamic Downsampling module selects keyframes based on this similarity; if it fails (e.g., camera pan), the selected frames may miss informative content.
  • domain assumption Exponentially weighted pixel averaging over a temporal window (Eqs. 10–11) preserves or enhances the temporal cues needed by the VLLM.
    PAME assumes that motion information survives pixel-level averaging; this is not guaranteed for fast motion or small objects.
  • domain assumption The reported benchmarks and 32-frame setting are representative of VLLM use cases.
    All experiments use 32 frames and N_v=196; conclusions may not transfer to longer videos or different token densities.
  • domain assumption Total FLOPs is an adequate proxy for runtime speedup.
    The paper reports FLOPs and some latency (Fig. 5), but the central claim is expressed in FLOPs reduction; memory and I/O effects are ignored.
  • standard math Transformer attention scales quadratically with token count.
    Used to motivate frame reduction as a way to reduce both encoder and LLM computational cost.

pith-pipeline@v1.3.0-alltime-deepseek · 14678 in / 13207 out tokens · 118415 ms · 2026-08-01T11:45:16.459570+00:00 · methodology

0 comments
read the original abstract

Despite advances in Video Large Language Models (VLLMs) that have displayed promising outcomes in video understanding, the redundancy in the long-duration frames remains a hindrance to efficient reasoning. This paper introduces a training-free $\mathbf{P}$ersistence-Aware $\mathbf{C}$ompression and $\mathbf{A}$ggregation (PCA) method designed to preserve high-fidelity raw visual information before the encoding stage. PCA can be built on arbitrary VLLMs and consists of two modules: 1) A Dynamic Downsampling (DD) module that adaptively removes redundant frames by analyzing frame-wise similarity. 2) A Persistence-Aware Motion Enhancement (PAME) module that enriches each selected keyframe by aggregating the temporal context of its neighbors, ensuring that essential information is preserved even after aggressive frame reduction. Our approach substantially reduces the computation of long-context modeling, while enhancing the performance of the baseline model. Extensive experiments demonstrate that PCA consistently outperforms existing state-of-the-art approaches in both efficiency and accuracy, achieving a speedup of 1.8$\times$ to 2.5$\times$ compared to the baseline VLLM. The code is open-sourced at https://github.com/Heisenberg10110/PCA.

Figures

Figures reproduced from arXiv: 2607.22726 by Bo Zhao, Jiayu Zhang, Ruixin Zhang, Shouhong Ding, Shuo Ye, Zihan Song, Zitong Yu.

Figure 1
Figure 1. Figure 1: Pairwise similarity heatmap between adja [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Comparison of models in terms of mean accuracy and total FLOPs on the MVBench dataset. Our proposed PCA substantially reduces compu￾tational cost while maintaining high model perfor￾mance, outperforming DyCoke and FastV with fewer FLOPs. “Origin” denotes the LLaVA-OV baseline without pruning or compression. that is particularly pronounced when encoding every frame in full resolution, resulting in excessive… view at source ↗
Figure 3
Figure 3. Figure 3: Overview of the proposed PCA framework, which consists of two main modules: Dynamic Down [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Ablation study on the effect of the PAME. [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Comparison of component-wise time con￾sumption across methods. The chart illustrates the time distribution for encoder, LLM, and other com￾ponents in different visual language models. Our method achieves a significant reduction in both en￾coder and LLM computation, highlighting the advan￾tage of pre-encoding pruning. 50% 33% 25% Frame Rate Radio 50 52 54 56 58 60 Accuracy Origin PCA DyCoke FastV VisionZip … view at source ↗
Figure 6
Figure 6. Figure 6: Comparison of Methods on Sparse Frame Conditions. Results are evaluated on MVBench, with LLaVA-OV-7B as the baseline model. The frame rate ratio indicates the proportion of frames retained as input to the LLM, reflecting model ro￾bustness under limited frame availability. challenging case for temporal modeling. As illustrated in Fig￾ure 6, PCA maintains high accuracy even when the input be￾comes sparse, wh… view at source ↗
Figure 7
Figure 7. Figure 7: Performance on the top 20% longest videos. [PITH_FULL_IMAGE:figures/full_fig_p008_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Visualization of Similarity-Based Down￾sampling. Keyframes are preserved by restoring frames with low similarity after uniform downsam￾pling. The similarity score for each frame is com￾puted to its preceding frame. persistence-aware and temporally enriched for robust down￾stream reasoning. Extensive experiments show that PCA consistently outperforms state-of-the-art baselines in both accuracy and efficienc… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

53 extracted references · 3 canonical work pages

  1. [1]

    Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai, Peizhao Zhang, Christoph Feichtenhofer, and Judy Hoffman. 2023. Token Merg- ing: Your ViT But Faster. InICLR. OpenReview.net

  2. [2]

    Gongwei Chen, Leyang Shen, Rui Shao, Xiang Deng, and Liqiang Nie. 2024. Lion: Empowering multimodal large language model with dual-level visual knowledge. InCVPR. 26540–26550

  3. [3]

    Liang Chen, Haozhe Zhao, Tianyu Liu, Shuai Bai, Junyang Lin, Chang Zhou, and Baobao Chang. 2024. An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models. InECCV (Lecture Notes in Computer Science, Vol. 15139), Ales Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten Sattler, and G¨ ul Va...

  4. [4]

    Shuimu Chen, Yuteng Chen, Yuanshen Guan, Zebang Cheng, Zeyu Zhang, Shengqian Qin, Bin Xia, Jiaran Li, Wenming Yang, and Fei Ma. 2026. Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding. InEuropean Con- ference on Computer Vision

  5. [6]

    Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Erfei Cui, Jinguo Zhu, Shenglong Ye, Hao Tian, Zhaoyang Liu, et al. 2024. Expanding performance boundaries of open- source multimodal models with model, data, and test-time scal- ing.arXiv preprint arXiv:2412.05271(2024)

  6. [7]

    Zebang Cheng, Shuimu Chen, Boxue Yang, Yuanshen Guan, Jingyi Chen, Zheng Lian, Xiaojiang Peng, Fei Ma, Laizhong Cui, and Qi Tian. 2026. OmniOPSD: Rationale-Privileged On- Policy Self-Distillation for Affective Computing.arXiv preprint arXiv:2606.15920(2026)

  7. [8]

    Zesen Cheng, Sicong Leng, Hang Zhang, Yifei Xin, Xin Li, Guanzheng Chen, Yongxin Zhu, Wenqi Zhang, Ziyang Luo, Deli Zhao, et al. 2024. Videollama 2: Advancing spatial-temporal modeling and audio understanding in video-llms.arXiv preprint arXiv:2406.07476(2024)

  8. [9]

    Xiangxiang Chu, Limeng Qiao, Xinyang Lin, Shuang Xu, Yang Yang, Yiming Hu, Fei Wei, Xinyu Zhang, Bo Zhang, Xiaolin Wei, et al. 2023. Mobilevlm: A fast, reproducible and strong vision language assistant for mobile devices.arXiv preprint arXiv:2312.168862, 6 (2023), 7

  9. [10]

    Xiangxiang Chu, Limeng Qiao, Xinyu Zhang, Shuang Xu, Fei Wei, Yang Yang, Xiaofei Sun, Yiming Hu, Xinyang Lin, Bo Zhang, et al. 2024. Mobilevlm v2: Faster and stronger baseline for vision language model.arXiv preprint arXiv:2402.03766(2024)

  10. [11]

    Max Coltheart. 1980. Iconic memory and visible persistence.Per- ception & psychophysics27, 3 (1980), 183–228

  11. [12]

    Ziyang Fan, Keyu Chen, Ruilong Xing, Yulin Li, Li Jiang, and Zhuotao Tian. 2026. FlashVID: Efficient Video Large Language Models via Training-free Tree-based Spatiotemporal Token Merg- ing.ICLR(2026)

  12. [13]

    Luciano Floridi and Massimo Chiriatti. 2020. GPT-3: Its na- ture, scope, limits, and consequences.Minds and machines30, 4 (2020), 681–694

  13. [14]

    Chaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Meng- dan Zhang, et al. 2025. Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis. In CVPR. 24108–24118

  14. [15]

    Libo Huang, Xiangqi Li, Jiarui Zhao, Zhulin An, Chuanguang Yang, Boyu Diao, Fei Wang, Yan Zeng, Zhifeng Hao, and Yongjun Xu. 2026. PrePrompt: Predictive Prompting for Class- Incremental Learning. InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD)

  15. [16]

    Libo Huang, Yan Zeng, Chuanguang Yang, Zhulin An, Boyu Diao, and Yongjun Xu. 2024. eTag: Class-Incremental Learning via Em- bedding Distillation and Task-Oriented Generation. InProceed- ings of the AAAI Conference on Artificial Intelligence (AAAI), Vol. 38. 12591–12599

  16. [17]

    Jitesh Jain, Jianwei Yang, and Humphrey Shi. 2024. Vcoder: Versatile vision encoders for multimodal large language models. InCVPR. 27992–28002

  17. [18]

    Bohao Li, Yuying Ge, Yixiao Ge, Guangzhi Wang, Rui Wang, Ruimao Zhang, and Ying Shan. 2024. Seed-bench: Benchmarking multimodal large language models. InCVPR. 13299–13308

  18. [19]

    Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Peiyuan Zhang, Yanwei Li, Ziwei Liu, et al. 2024. Llava-onevision: Easy visual task transfer.arXiv preprint arXiv:2408.03326(2024)

  19. [20]

    KunChang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao. 2023. Videochat: Chat-centric video understanding.arXiv preprint arXiv:2305.06355(2023)

  20. [21]

    Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Luo, et al. 2024. Mvbench: A comprehensive multi-modal video understanding benchmark. InCVPR. 22195–22206

  21. [22]

    Fei Ma, Yucheng Yuan, Yifan Xie, Hongwei Ren, Ivan Liu, Ying He, Fuji Ren, Fei Richard Yu, and Shiguang Ni. 2025. Gener- ative Technology for Human Emotion Recognition: A Scoping Review.Information Fusion115 (2025), 102753. doi:10.1016/j. inffus.2024.102753

  22. [23]

    Muhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, and Fahad Khan. 2024. Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models. InACL, Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). Associ- ation for Computational Linguistics, 12585–12602. doi:10.18653/ V1/2024.ACL-LONG.679

  23. [24]

    Xin Man, Jie Shao, Feiyu Chen, Mingxing Zhang, and Heng Tao Shen. 2023. TEVL: Trilinear Encoder for Video-language Rep- resentation Learning.ACM Trans. Multim. Comput. Commun. Song et al. Appl.19, 5s (2023), 168:1–168:20. doi:10.1145/3585388

  24. [25]

    Richard H Masland. 2012. The neuronal organization of the retina.Neuron76, 2 (2012), 266–280

  25. [26]

    Viorica Patraucean, Lucas Smaira, Ankush Gupta, Adria Re- casens, Larisa Markeeva, Dylan Banarse, Skanda Koppula, Ma- teusz Malinowski, Yi Yang, Carl Doersch, et al. 2023. Perception test: A diagnostic benchmark for multimodal video models.NIPS 36 (2023), 42748–42761

  26. [27]

    Rui Qian, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Shuangrui Ding, Dahua Lin, and Jiaqi Wang. 2024. Streaming long video understanding with large language models.NIPS37 (2024), 119336–119360

  27. [28]

    Shuhuai Ren, Sishuo Chen, Shicheng Li, Xu Sun, and Lu Hou

  28. [29]

    Shuhuai Ren, Linli Yao, Shicheng Li, Xu Sun, and Lu Hou. 2024. TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding. InCVPR. 14313–14323. doi:10. 1109/CVPR52733.2024.01357

  29. [30]

    Steven H Schwartz. 2004. Visual perception: A clinical orienta- tion. (2004)

  30. [31]

    Yuzhang Shang, Mu Cai, Bingxin Xu, Yong Jae Lee, and Yan Yan. 2024. Llava-prumerge: Adaptive token reduction for effi- cient large multimodal models.arXiv preprint arXiv:2403.15388 (2024)

  31. [32]

    Leqi Shen, Tianxiang Hao, Tao He, Sicheng Zhao, Yifeng Zhang, Pengzhang Liu, Yongjun Bao, and Guiguang Ding. 2025. TempMe: Video Temporal Token Merging for Efficient Text- Video Retrieval. InICLR. OpenReview.net

  32. [33]

    Keda Tao, Can Qin, Haoxuan You, Yang Sui, and Huan Wang

  33. [34]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth´ ee Lacroix, Baptiste Rozi` ere, Na- man Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models.arXiv preprint arXiv:2302.13971(2023)

  34. [35]

    Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al

  35. [36]

    Xiaohan Wang, Yuhui Zhang, Orr Zohar, and Serena Yeung-Levy

  36. [37]

    Yi Wang, Kunchang Li, Yizhuo Li, Yinan He, Bingkun Huang, Zhiyu Zhao, Hongjie Zhang, Jilan Xu, Yi Liu, Zun Wang, et al. 2022. Internvideo: General video foundation models via generative and discriminative learning.arXiv preprint arXiv:2212.03191(2022)

  37. [38]

    Yiping Xie, Bo Zhao, Mingtong Dai, Jian-Ping Zhou, Yue Sun, Tao Tan, Weicheng Xie, Linlin Shen, and Zitong Yu. 2026. Phys- LLM: Harnessing Large Language Models for Cross-Modal Re- mote Physiological Sensing. InInternational Conference on Learning Representations

  38. [39]

    Haiwei Xue, Xiangyang Luo, Zhanghao Hu, Xin Zhang, Xunzhi Xiang, Yuqin Dai, Jianzhuang Liu, Zhensong Zhang, Minglei Li, Jian Yang, Fei Ma, Zhiyong Wu, Changpeng Yang, Zonghong Dai, and Fei Richard Yu. 2025. Human Motion Video Genera- tion: A Survey.IEEE Transactions on Pattern Analysis and Machine Intelligence47, 11 (2025), 10709–10730. doi:10.1109/ TPAMI...

  39. [40]

    InEuropean Conference on Computer Vision

    Videoagent: Long-form video understanding with large lan- guage model as agent. InEuropean Conference on Computer Vision. Springer, 58–76

  40. [41]

    Qilang Ye, Zitong Yu, Rui Shao, Yawen Cui, Xiangui Kang, Xin Liu, Philip H. S. Torr, and Xiaochun Cao. 2025. CAT+: Investi- gating and Enhancing Audio-Visual Understanding in Large Lan- guage Models.IEEE Transactions on Pattern Analysis and Machine Intelligence(2025)

  41. [42]

    Heng Yin, Yuqiang Ren, Ke Yan, Shouhong Ding, and Yongtao Hao. 2025. ROD-MLLM: Towards More Reliable Object Detec- tion in Multimodal Large Language Models. InCVPR. 14358– 14368

  42. [43]

    Hang Zhang, Xin Li, and Lidong Bing. 2023. Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Un- derstanding. InProceedings of the 2023 Conference on Empir- ical Methods in Natural Language Processing, EMNLP 2023 - System Demonstrations, Singapore, December 6-10, 2023, Yan- song Feng and Els Lefever (Eds.). Association for Computatio...

  43. [44]

    Senqiao Yang, Yukang Chen, Zhuotao Tian, Chengyao Wang, Jingyao Li, Bei Yu, and Jiaya Jia. 2024. VisionZip: Longer is Better but Not Necessary in Vision Language Models.arXiv preprint arXiv:2412.04467(2024)

  44. [45]

    Bo Zhao, Dan Guo, Junzhe Cao, Yong Xu, Tao Tan, Yue Sun, Bochao Zou, Jie Zhang, and Zitong Yu. 2026. PHASE-Net: Physics-Grounded Harmonic Attention System for Efficient Re- mote Photoplethysmography Measurement. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

  45. [46]

    Shiyu Zhao, Zhenting Wang, Felix Juefei-Xu, Xide Xia, Miao Liu, Xiaofang Wang, Mingfu Liang, Ning Zhang, Dimitris N Metaxas, and Licheng Yu. 2025. Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction. InCVPR. 29869–29879

  46. [47]

    Yue Zhao, Ishan Misra, Philipp Kr¨ ahenb¨ uhl, and Rohit Girdhar

  47. [48]

    Jiayu Zhang, Xun Lin, Jiajian Huang, Shuo Ye, Xiaobao Guo, Dongliang Zhu, Ruimin Hu, Dan Guo, Yanyan Liang, Zitong Yu, and Xiaochun Cao. 2026. Multimodal Deception Detection: A Survey.Machine Intelligence Research23, 2 (2026), 284–307

  48. [49]

    Orr Zohar, Xiaohan Wang, Yann Dubois, Nikhil Mehta, Tong Xiao, Philippe Hansen-Estruch, Licheng Yu, Xiaofang Wang, Fe- lix Juefei-Xu, Ning Zhang, et al. 2025. Apollo: An exploration of video understanding in large multimodal models. InCVPR. 18891–18901

  49. [52]

    Learning video representations from large language models. InCVPR. 6586–6597

  50. [53]

    Luowei Zhou, Chenliang Xu, and Jason Corso. 2018. Towards automatic learning of procedures from web instructional videos. InAAAI, Vol. 32

  51. [2023]

    InFindings of the Associ- ation for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.)

    TESTA: Temporal-Spatial Token Aggregation for Long- form Video-Language Understanding. InFindings of the Associ- ation for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, 932–947. doi:10.18653/V1/2023.FINDINGS-EMNLP.66

  52. [2024]

    Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution.arXiv preprint arXiv:2409.12191 (2024)

  53. [2025]

    DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models. InCVPR. 18992–19001