REVIEW 4 major objections 6 minor 53 references
This paper claims that aggressive frame compression applied before the vision encoder—guided by inter-frame similarity and enriched with decay-weighted temporal blending—cuts video-LLM computation by 40–50% while holding or slightly improvi
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 11:45 UTC pith:GEYPGAGU
load-bearing objection Simple, honest pre-encoding frame reduction for video LLMs that delivers the FLOPs savings it claims, but the accuracy-maintenance story is weaker than the abstract implies and the hyperparameters look tuned on the test benchmarks. the 4 major comments →
PCA: Persistence-Aware Compression and Aggregation for Fast Video Large Language Models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that frame-level redundancy in video inputs can be removed before visual encoding without harming—and sometimes improving—video understanding, provided the surviving frames are made persistence-aware by aggregating short-range temporal context. The paper shows that a similarity-based selector (cosine similarity between flattened consecutive frames) combined with a decay-weighted temporal blending of neighbors preserves the cues VLLMs need for temporal reasoning, while cutting computation in both the encoder and the language model. In experiments, PCA reduces total FLOPs to roughly half of baseline on a 7B video-LLM (41–54% of original FLOPs depending on settings) and rai
What carries the argument
The argument rests on two mechanisms. Dynamic Downsampling (DD) selects frames by taking every k-th frame and then restoring the p fraction of frames whose cosine similarity to the previous frame is lowest, so scene transitions and motion onsets survive uniform subsampling. Persistence-Aware Motion Enhancement (PAME) then replaces each selected frame with a weighted average of itself and up to J preceding frames, where weights decay exponentially with temporal distance (α^(t_k−i)); this encodes short-term motion cues into a single still image. PAME is the carrier of the paper's core idea: a cheap, training-free pixel-level aggregation that compensates for dropped frames.
Load-bearing premise
The frame selector assumes that high cosine similarity between consecutive raw frames means redundancy, and low similarity means important new content; if that mapping fails—during pans, zooms, or lighting changes—the wrong frames get dropped or kept.
What would settle it
Take a video with constant camera pan across a static scene: adjacent frames have low pixel similarity yet no new semantic content, so PCA would restore many 'dissimilar' frames and waste its frame budget, while a pure uniform sampler at the same rate would do as well. An ablation on such a benchmark comparing PCA's selection to random selection at matched frame counts would settle whether similarity-based selection earns its complexity.
If this is right
- Compressing before the vision encoder removes the encoder itself as a bottleneck, which token-pruning methods leave untouched; PCA's speedup grows as video length or resolution rises because its own overhead is linear while encoder/LLM cost is superlinear.
- Because both modules are training-free and operate on raw frames, PCA can be dropped onto any VLLM; the paper demonstrates gains on three distinct VLLM architectures.
- PCA degrades gracefully under sparse input: at 25–50% frame rates it keeps higher accuracy than token-pruning baselines, suggesting the aggregation step reconstructs missing temporal context.
- On the longest videos in MVBench, PCA stays within 1.6 points of the full baseline, indicating the local aggregation window prevents error accumulation over long horizons.
- The decay coefficient matters: uniform averaging (α=1) collapses accuracy, while moderate exponential decay (α around 0.1) is stable, indicating that recent frames should dominate the blend.
Where Pith is reading between the lines
- The similarity proxy in DD is a risk: cosine similarity between raw pixels will flag camera pans and global illumination shifts as 'novel' even when no semantic event occurs, so the restored-frame budget may be spent on uninformative transitions; a semantic or object-level similarity could be a direct improvement.
- PAME is a form of temporal low-pass filtering on pixels; one could test whether applying it before uniform downsampling rather than after, or integrating it with learned temporal modules, would yield further gains.
- The method suggests a broader design principle: in video LLMs, input-side compression that preserves temporal context can outperform output-side token pruning; this encourages revisiting where in the pipeline redundancy reduction should occur.
- A natural extension is streaming video, where PAME's causal weighted window fits online inference; the paper's current experiments are offline, but the module's local nature makes it compatible with frame-by-frame processing.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes PCA, a training-free, pre-encoding frame-compression method for Video LLMs. It consists of Dynamic Downsampling (DD), which uniformly samples frames and then restores a fraction of the frames with lowest raw-pixel cosine similarity to their predecessors, and Persistence-Aware Motion Enhancement (PAME), which replaces each selected frame with a weighted temporal average of its preceding frames. Experiments on MVBench, PerceptionTest, YouCook2, and VideoMME with LLaVA-OV-0.5B/7B, plus MVBench results on VideoChatGPT and InternVL2, report a 40–50% total FLOPs reduction and a 1.8–2.5× speedup with accuracy roughly maintained relative to the untouched baseline and improved relative to token-pruning baselines.
Significance. If the results hold, this is a useful contribution: it targets the vision-encoder cost that post-encoding token-pruning methods cannot reduce, it is training-free and plug-and-play in design, and it is accompanied by open-source code. The ablations isolate the two modules, and the random-pruning comparison in Table 2 gives some evidence that the similarity heuristic carries signal rather than acting merely as a frame-count reducer. The derivation is simple and the computational reasoning is sound. The main open question is whether the reported accuracy retention is an artifact of hyperparameter selection on the evaluation benchmarks, and whether the generality claims in the abstract and Section 4.6 are supported by the data.
major comments (4)
- [§4.4, Tables 1–3] The accuracy-retention claim rests on configurations whose hyperparameters appear to be selected on the reported benchmarks. Table 1 reports only two PCA settings (K=2,J=3,P=0.2 and K=3,J=3,P=0.4), while Table 2 sweeps K∈{2,3,4}, J∈{2,4}, P∈{0.4,0.6} on the same MVBench/PercepTest/YouCook2/VideoMME numbers, and Table 3 sweeps α on MVBench/VideoMME. No development/held-out split, repeated runs, or error bars are given. Several decisive margins are small relative to this selection effect: e.g., Table 1 7B Avg 36.96 vs baseline 37.37 and MVBench 56.69 vs 57.25; Table 3 α=0.2 gives 56.55 vs α=0.1 56.69. Selection on the test set can therefore fully account for the 'maintaining accuracy' conclusion. Please provide an unbiased protocol (e.g., tune on a subset, use cross-validation, or report mean±std across seeds) for the reported configurations.
- [Abstract and §4.6, Table 4] The abstract claims PCA 'enhances the performance of the baseline model' and 'consistently outperforms existing state-of-the-art approaches in both efficiency and accuracy.' Table 4 contradicts the first part: on MVBench, InternVL2 origin 40.27 → PCA 39.07 and VideoChatGPT origin 58.60 → PCA 57.22. Table 1 also shows several drops relative to origin (e.g., 0.5B YouCook2 METEOR 8.31 vs 9.03; 7B VideoMME wo-subs 55.70 vs 57.89). PCA is better than the token-pruning baselines at the reported compression, but the stronger claim that baseline performance is enhanced is not supported. Please revise the abstract and conclusion to state the trade-offs precisely, or provide additional evidence for the stronger claim.
- [§4.3–§4.4, Table 1] The FLOPs comparison is asymmetric by construction: token-pruning baselines reduce only LLM-side tokens, leaving encoder FLOPs unchanged, while PCA reduces frames before encoding. This explains much of the efficiency margin and is not itself a flaw. However, the accuracy comparison is only interpretable if compute budgets are actually matched. The paper says pruning ratios are 'determined based on the total FLOPs' (§4.3) but does not state the procedure or the per-method settings; Table 1 reports one operating point per baseline. Please report accuracy-vs-FLOPs curves, or at least matched-FLOPs operating points, and state how each baseline's pruning strength was chosen, including whether it was tuned on the same benchmarks.
- [§3.3, Eq. (6)] The DD module assumes that raw-pixel cosine similarity between consecutive frames identifies semantically important transitions. The random-pruning ablation (Table 2) suggests this heuristic carries signal on the tested benchmarks, so I do not view this as an internal inconsistency. Nevertheless, raw-pixel similarity is sensitive to camera motion, illumination changes, and compression artifacts, which can produce large pixel differences without new semantic content. A concrete robustness test would help: evaluate on videos with controlled camera pan/zoom or photometric perturbation and compare DD against random selection at matched frame counts. If the gap closes, the paper should state the scope of validity. This concern is secondary to the validation-protocol issue but matters for the generality claim.
minor comments (6)
- [Throughout] There are several typos and spacing artifacts: 'Persistence-Aware Motion Enhencement' in Section 1, 'Frame Rate Radio' in Figure 6, 'LLA V A-OV-7B' in Figure 2, and 'na¨ıve' spacing in Section 4.5. Please proofread.
- [Eqs. (10)–(11) and Algorithm 1] The weighting in Eq. (11) uses the condition t_k-i < j, while Algorithm 1 sets start_idx = max(0, h.idx - window) and then includes the frame at distance exactly window. These definitions are inconsistent for the boundary frame. Also, α=0 in Table 3 involves 0^0 for the current frame; please clarify the intended convention.
- [References] References [5] and [6] are identical (both are 'Expanding performance boundaries of open-source multimodal models...'), but §4.2 cites InternVL2 as [6]. The cited reference does not appear to be the actual InternVL2 paper. Please correct the citation and remove the duplicate.
- [Eq. (7)] The notation in Eq. (7), i / ||F∖F′|| < p, is not self-explanatory because i is used as an index into the sorted sequence P. Please clarify that this selects the lowest-similarity p-fraction of the pairs in P.
- [Table 2] The Random Pruning baseline is not compute-matched to the PCA rows (e.g., 0.5B: random 8.3 TFLOPs vs PCA 5.7–9.3; 7B: random 50.1 vs PCA 35.3–56.9). Consider reporting a random-pruning curve across several frame counts so the comparison is not confounded with FLOPs.
- [§3.1, Eqs. (1)–(2)] The derivation of ΔV_rel is correct but compressed. Defining V explicitly as responses per unit time and stating that the frame ratio is α would improve readability.
Circularity Check
No significant circularity: PCA is a constructive heuristic with empirical evaluation; the claimed predictions are not forced by construction.
full rationale
I walked the paper's derivation chain. Equations (1)–(2) are plain algebra relating frame-rate reduction to relative speed improvement; they are not predictions and do not feed back into the accuracy claims. The Dynamic Downsampling module (Eqs. 3–8) defines a frame-selection rule based on cosine similarity between flattened pixel vectors (Eq. 6). This is a heuristic selection procedure, not a derivation of accuracy: the claim that low similarity corresponds to informative frames is an empirical assumption evaluated in Tables 1–4 and Figure 8, not an equation whose output equals its input. The Persistence-Aware Motion Enhancement module (Eqs. 9–12) is likewise an explicit weighted-averaging definition; it enriches frames but does not mathematically imply benchmark performance. No fitted parameter is renamed as a prediction: K, J, P, and alpha are discrete hyperparameters, and the reported configurations are empirical choices. The concern that these hyperparameters may have been selected using the reported test benchmarks is a real evaluation-validity risk, but it is not a circularity in the sense of an equation reducing to its own inputs, and the paper does not describe a fitting procedure that would make the benchmark numbers forced. The paper's self-citations (e.g., Phys-LLM, CAT+, ROD-MLLM, PHASE-Net, Multimodal Deception Detection) are listed as related prior work; none is used as a load-bearing uniqueness theorem or as the sole justification for the central efficiency/accuracy claim. The central claim rests on the experimental tables, which are external empirical evidence rather than a self-referential derivation. Therefore no specific circular step can be quoted, and the appropriate score is 0.
Axiom & Free-Parameter Ledger
free parameters (4)
- K (uniform downsampling step) =
2 or 3
- J (temporal window size) =
3
- P (restore rate) =
0.2 or 0.4
- α (decay coefficient) =
0.1
axioms (5)
- domain assumption Cosine similarity between flattened pixel vectors (Eq. 6) is a valid proxy for semantic informativeness of frames.
- domain assumption Exponentially weighted pixel averaging over a temporal window (Eqs. 10–11) preserves or enhances the temporal cues needed by the VLLM.
- domain assumption The reported benchmarks and 32-frame setting are representative of VLLM use cases.
- domain assumption Total FLOPs is an adequate proxy for runtime speedup.
- standard math Transformer attention scales quadratically with token count.
read the original abstract
Despite advances in Video Large Language Models (VLLMs) that have displayed promising outcomes in video understanding, the redundancy in the long-duration frames remains a hindrance to efficient reasoning. This paper introduces a training-free $\mathbf{P}$ersistence-Aware $\mathbf{C}$ompression and $\mathbf{A}$ggregation (PCA) method designed to preserve high-fidelity raw visual information before the encoding stage. PCA can be built on arbitrary VLLMs and consists of two modules: 1) A Dynamic Downsampling (DD) module that adaptively removes redundant frames by analyzing frame-wise similarity. 2) A Persistence-Aware Motion Enhancement (PAME) module that enriches each selected keyframe by aggregating the temporal context of its neighbors, ensuring that essential information is preserved even after aggressive frame reduction. Our approach substantially reduces the computation of long-context modeling, while enhancing the performance of the baseline model. Extensive experiments demonstrate that PCA consistently outperforms existing state-of-the-art approaches in both efficiency and accuracy, achieving a speedup of 1.8$\times$ to 2.5$\times$ compared to the baseline VLLM. The code is open-sourced at https://github.com/Heisenberg10110/PCA.
Figures
Reference graph
Works this paper leans on
-
[1]
Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai, Peizhao Zhang, Christoph Feichtenhofer, and Judy Hoffman. 2023. Token Merg- ing: Your ViT But Faster. InICLR. OpenReview.net
2023
-
[2]
Gongwei Chen, Leyang Shen, Rui Shao, Xiang Deng, and Liqiang Nie. 2024. Lion: Empowering multimodal large language model with dual-level visual knowledge. InCVPR. 26540–26550
2024
-
[3]
Liang Chen, Haozhe Zhao, Tianyu Liu, Shuai Bai, Junyang Lin, Chang Zhou, and Baobao Chang. 2024. An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models. InECCV (Lecture Notes in Computer Science, Vol. 15139), Ales Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten Sattler, and G¨ ul Va...
-
[4]
Shuimu Chen, Yuteng Chen, Yuanshen Guan, Zebang Cheng, Zeyu Zhang, Shengqian Qin, Bin Xia, Jiaran Li, Wenming Yang, and Fei Ma. 2026. Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding. InEuropean Con- ference on Computer Vision
2026
-
[6]
Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Erfei Cui, Jinguo Zhu, Shenglong Ye, Hao Tian, Zhaoyang Liu, et al. 2024. Expanding performance boundaries of open- source multimodal models with model, data, and test-time scal- ing.arXiv preprint arXiv:2412.05271(2024)
Pith/arXiv arXiv 2024
-
[7]
Zebang Cheng, Shuimu Chen, Boxue Yang, Yuanshen Guan, Jingyi Chen, Zheng Lian, Xiaojiang Peng, Fei Ma, Laizhong Cui, and Qi Tian. 2026. OmniOPSD: Rationale-Privileged On- Policy Self-Distillation for Affective Computing.arXiv preprint arXiv:2606.15920(2026)
Pith/arXiv arXiv 2026
-
[8]
Zesen Cheng, Sicong Leng, Hang Zhang, Yifei Xin, Xin Li, Guanzheng Chen, Yongxin Zhu, Wenqi Zhang, Ziyang Luo, Deli Zhao, et al. 2024. Videollama 2: Advancing spatial-temporal modeling and audio understanding in video-llms.arXiv preprint arXiv:2406.07476(2024)
Pith/arXiv arXiv 2024
-
[9]
Xiangxiang Chu, Limeng Qiao, Xinyang Lin, Shuang Xu, Yang Yang, Yiming Hu, Fei Wei, Xinyu Zhang, Bo Zhang, Xiaolin Wei, et al. 2023. Mobilevlm: A fast, reproducible and strong vision language assistant for mobile devices.arXiv preprint arXiv:2312.168862, 6 (2023), 7
Pith/arXiv arXiv 2023
-
[10]
Xiangxiang Chu, Limeng Qiao, Xinyu Zhang, Shuang Xu, Fei Wei, Yang Yang, Xiaofei Sun, Yiming Hu, Xinyang Lin, Bo Zhang, et al. 2024. Mobilevlm v2: Faster and stronger baseline for vision language model.arXiv preprint arXiv:2402.03766(2024)
Pith/arXiv arXiv 2024
-
[11]
Max Coltheart. 1980. Iconic memory and visible persistence.Per- ception & psychophysics27, 3 (1980), 183–228
1980
-
[12]
Ziyang Fan, Keyu Chen, Ruilong Xing, Yulin Li, Li Jiang, and Zhuotao Tian. 2026. FlashVID: Efficient Video Large Language Models via Training-free Tree-based Spatiotemporal Token Merg- ing.ICLR(2026)
2026
-
[13]
Luciano Floridi and Massimo Chiriatti. 2020. GPT-3: Its na- ture, scope, limits, and consequences.Minds and machines30, 4 (2020), 681–694
2020
-
[14]
Chaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Meng- dan Zhang, et al. 2025. Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis. In CVPR. 24108–24118
2025
-
[15]
Libo Huang, Xiangqi Li, Jiarui Zhao, Zhulin An, Chuanguang Yang, Boyu Diao, Fei Wang, Yan Zeng, Zhifeng Hao, and Yongjun Xu. 2026. PrePrompt: Predictive Prompting for Class- Incremental Learning. InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD)
2026
-
[16]
Libo Huang, Yan Zeng, Chuanguang Yang, Zhulin An, Boyu Diao, and Yongjun Xu. 2024. eTag: Class-Incremental Learning via Em- bedding Distillation and Task-Oriented Generation. InProceed- ings of the AAAI Conference on Artificial Intelligence (AAAI), Vol. 38. 12591–12599
2024
-
[17]
Jitesh Jain, Jianwei Yang, and Humphrey Shi. 2024. Vcoder: Versatile vision encoders for multimodal large language models. InCVPR. 27992–28002
2024
-
[18]
Bohao Li, Yuying Ge, Yixiao Ge, Guangzhi Wang, Rui Wang, Ruimao Zhang, and Ying Shan. 2024. Seed-bench: Benchmarking multimodal large language models. InCVPR. 13299–13308
2024
-
[19]
Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Peiyuan Zhang, Yanwei Li, Ziwei Liu, et al. 2024. Llava-onevision: Easy visual task transfer.arXiv preprint arXiv:2408.03326(2024)
Pith/arXiv arXiv 2024
-
[20]
KunChang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao. 2023. Videochat: Chat-centric video understanding.arXiv preprint arXiv:2305.06355(2023)
Pith/arXiv arXiv 2023
-
[21]
Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Luo, et al. 2024. Mvbench: A comprehensive multi-modal video understanding benchmark. InCVPR. 22195–22206
2024
-
[22]
Fei Ma, Yucheng Yuan, Yifan Xie, Hongwei Ren, Ivan Liu, Ying He, Fuji Ren, Fei Richard Yu, and Shiguang Ni. 2025. Gener- ative Technology for Human Emotion Recognition: A Scoping Review.Information Fusion115 (2025), 102753. doi:10.1016/j. inffus.2024.102753
arXiv 2025
-
[23]
Muhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, and Fahad Khan. 2024. Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models. InACL, Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). Associ- ation for Computational Linguistics, 12585–12602. doi:10.18653/ V1/2024.ACL-LONG.679
2024
-
[24]
Xin Man, Jie Shao, Feiyu Chen, Mingxing Zhang, and Heng Tao Shen. 2023. TEVL: Trilinear Encoder for Video-language Rep- resentation Learning.ACM Trans. Multim. Comput. Commun. Song et al. Appl.19, 5s (2023), 168:1–168:20. doi:10.1145/3585388
-
[25]
Richard H Masland. 2012. The neuronal organization of the retina.Neuron76, 2 (2012), 266–280
2012
-
[26]
Viorica Patraucean, Lucas Smaira, Ankush Gupta, Adria Re- casens, Larisa Markeeva, Dylan Banarse, Skanda Koppula, Ma- teusz Malinowski, Yi Yang, Carl Doersch, et al. 2023. Perception test: A diagnostic benchmark for multimodal video models.NIPS 36 (2023), 42748–42761
2023
-
[27]
Rui Qian, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Shuangrui Ding, Dahua Lin, and Jiaqi Wang. 2024. Streaming long video understanding with large language models.NIPS37 (2024), 119336–119360
2024
-
[28]
Shuhuai Ren, Sishuo Chen, Shicheng Li, Xu Sun, and Lu Hou
-
[29]
Shuhuai Ren, Linli Yao, Shicheng Li, Xu Sun, and Lu Hou. 2024. TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding. InCVPR. 14313–14323. doi:10. 1109/CVPR52733.2024.01357
arXiv 2024
-
[30]
Steven H Schwartz. 2004. Visual perception: A clinical orienta- tion. (2004)
2004
-
[31]
Yuzhang Shang, Mu Cai, Bingxin Xu, Yong Jae Lee, and Yan Yan. 2024. Llava-prumerge: Adaptive token reduction for effi- cient large multimodal models.arXiv preprint arXiv:2403.15388 (2024)
arXiv 2024
-
[32]
Leqi Shen, Tianxiang Hao, Tao He, Sicheng Zhao, Yifeng Zhang, Pengzhang Liu, Yongjun Bao, and Guiguang Ding. 2025. TempMe: Video Temporal Token Merging for Efficient Text- Video Retrieval. InICLR. OpenReview.net
2025
-
[33]
Keda Tao, Can Qin, Haoxuan You, Yang Sui, and Huan Wang
-
[34]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth´ ee Lacroix, Baptiste Rozi` ere, Na- man Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models.arXiv preprint arXiv:2302.13971(2023)
Pith/arXiv arXiv 2023
-
[35]
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al
-
[36]
Xiaohan Wang, Yuhui Zhang, Orr Zohar, and Serena Yeung-Levy
-
[37]
Yi Wang, Kunchang Li, Yizhuo Li, Yinan He, Bingkun Huang, Zhiyu Zhao, Hongjie Zhang, Jilan Xu, Yi Liu, Zun Wang, et al. 2022. Internvideo: General video foundation models via generative and discriminative learning.arXiv preprint arXiv:2212.03191(2022)
Pith/arXiv arXiv 2022
-
[38]
Yiping Xie, Bo Zhao, Mingtong Dai, Jian-Ping Zhou, Yue Sun, Tao Tan, Weicheng Xie, Linlin Shen, and Zitong Yu. 2026. Phys- LLM: Harnessing Large Language Models for Cross-Modal Re- mote Physiological Sensing. InInternational Conference on Learning Representations
2026
-
[39]
Haiwei Xue, Xiangyang Luo, Zhanghao Hu, Xin Zhang, Xunzhi Xiang, Yuqin Dai, Jianzhuang Liu, Zhensong Zhang, Minglei Li, Jian Yang, Fei Ma, Zhiyong Wu, Changpeng Yang, Zonghong Dai, and Fei Richard Yu. 2025. Human Motion Video Genera- tion: A Survey.IEEE Transactions on Pattern Analysis and Machine Intelligence47, 11 (2025), 10709–10730. doi:10.1109/ TPAMI...
arXiv 2025
-
[40]
InEuropean Conference on Computer Vision
Videoagent: Long-form video understanding with large lan- guage model as agent. InEuropean Conference on Computer Vision. Springer, 58–76
-
[41]
Qilang Ye, Zitong Yu, Rui Shao, Yawen Cui, Xiangui Kang, Xin Liu, Philip H. S. Torr, and Xiaochun Cao. 2025. CAT+: Investi- gating and Enhancing Audio-Visual Understanding in Large Lan- guage Models.IEEE Transactions on Pattern Analysis and Machine Intelligence(2025)
2025
-
[42]
Heng Yin, Yuqiang Ren, Ke Yan, Shouhong Ding, and Yongtao Hao. 2025. ROD-MLLM: Towards More Reliable Object Detec- tion in Multimodal Large Language Models. InCVPR. 14358– 14368
2025
-
[43]
Hang Zhang, Xin Li, and Lidong Bing. 2023. Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Un- derstanding. InProceedings of the 2023 Conference on Empir- ical Methods in Natural Language Processing, EMNLP 2023 - System Demonstrations, Singapore, December 6-10, 2023, Yan- song Feng and Els Lefever (Eds.). Association for Computatio...
-
[44]
Senqiao Yang, Yukang Chen, Zhuotao Tian, Chengyao Wang, Jingyao Li, Bei Yu, and Jiaya Jia. 2024. VisionZip: Longer is Better but Not Necessary in Vision Language Models.arXiv preprint arXiv:2412.04467(2024)
arXiv 2024
-
[45]
Bo Zhao, Dan Guo, Junzhe Cao, Yong Xu, Tao Tan, Yue Sun, Bochao Zou, Jie Zhang, and Zitong Yu. 2026. PHASE-Net: Physics-Grounded Harmonic Attention System for Efficient Re- mote Photoplethysmography Measurement. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2026
-
[46]
Shiyu Zhao, Zhenting Wang, Felix Juefei-Xu, Xide Xia, Miao Liu, Xiaofang Wang, Mingfu Liang, Ning Zhang, Dimitris N Metaxas, and Licheng Yu. 2025. Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction. InCVPR. 29869–29879
2025
-
[47]
Yue Zhao, Ishan Misra, Philipp Kr¨ ahenb¨ uhl, and Rohit Girdhar
-
[48]
Jiayu Zhang, Xun Lin, Jiajian Huang, Shuo Ye, Xiaobao Guo, Dongliang Zhu, Ruimin Hu, Dan Guo, Yanyan Liang, Zitong Yu, and Xiaochun Cao. 2026. Multimodal Deception Detection: A Survey.Machine Intelligence Research23, 2 (2026), 284–307
2026
-
[49]
Orr Zohar, Xiaohan Wang, Yann Dubois, Nikhil Mehta, Tong Xiao, Philippe Hansen-Estruch, Licheng Yu, Xiaofang Wang, Fe- lix Juefei-Xu, Ning Zhang, et al. 2025. Apollo: An exploration of video understanding in large multimodal models. InCVPR. 18891–18901
2025
-
[52]
Learning video representations from large language models. InCVPR. 6586–6597
-
[53]
Luowei Zhou, Chenliang Xu, and Jason Corso. 2018. Towards automatic learning of procedures from web instructional videos. InAAAI, Vol. 32
2018
-
[2023]
TESTA: Temporal-Spatial Token Aggregation for Long- form Video-Language Understanding. InFindings of the Associ- ation for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, 932–947. doi:10.18653/V1/2023.FINDINGS-EMNLP.66
-
[2024]
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution.arXiv preprint arXiv:2409.12191 (2024)
Pith/arXiv arXiv 2024
-
[2025]
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models. InCVPR. 18992–19001
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.