REVIEW 3 major objections 4 minor 71 references
GMoT shows that compact gated motion tokens let pretrained video language models recognize micro-gestures from kinematics rather than static posture, improving accuracy on two benchmarks and grounding their rationales in body regions.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 04:35 UTC pith:GO5D34DH
load-bearing objection The motion-token module is plausible but the headline gain overstates its contribution — the matched gain is ~+4.36 and the module is only ablated before RL. the 3 major comments →
GMoT: Gated Motion-Aware Tokenization for Fine-Grained Micro-Gesture Video Reasoning with Multimodal LLMs
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
GMoT's central claim is that a lightweight gated motion-aware tokenization module—not heavy 3D convolutions or dense space-time attention—can supply the kinematic evidence that pretrained MLLMs lack. The module computes a learned spatial saliency map per frame, weights patch tokens into a single frame feature, takes explicit adjacent-frame differences, decouples them into mean magnitude (motion energy) and signed direction, and projects the concatenation into a compact motion token. That token is fused into the aligned visual stream through a residual gate initialized near zero, so adaptation starts from the pretrained representation. The paper further claims that a staged recipe—GMoT warm-u
What carries the argument
The load-bearing object is the GMoT module, a three-part motion tokenizer: a spatial saliency scorer that softmax-weights patch tokens into a frame-level feature, adjacent-frame temporal differencing decoupled into mean-magnitude (motion energy) and signed-difference (direction) streams, and a near-closed gated residual fusion that injects the resulting compact motion token into the aligned visual stream. The gate is initialized so that the fused features start nearly identical to the original visual features, preserving the pretrained visual-language interface while letting the model learn how much motion evidence to trust. Around this module, the paper builds a four-stage training schedule
Load-bearing premise
The load-bearing premise is that the semi-automatically generated chain-of-thought annotations are reliable enough to supervise evidence-grounded reasoning; the paper's own audit found a 24.86% severe-hallucination rate among accepted SMG descriptions, so if the model inherits those hallucination patterns the reasoning-grounding claims weaken.
What would settle it
A decisive test: build a micro-gesture set of static-pose-matched pairs (same posture, different subtle motion) and compare GMoT against the vanilla backbone; if accuracy does not improve when motion is the only discriminative cue, the motion token is not supplying the kinematic evidence claimed. A cheaper check is a length-matched BRG comparison—if the grounding gap collapses when rationales are equalized for length, the grounding gain is a verbosity artifact.
If this is right
- The motion-token branch is the main source of the accuracy gain: removing the spatial scorer, temporal differencing, or the gate each costs 4.25–5.54 points in the ablation.
- The staged training schedule is necessary: label-only SFT slightly hurts accuracy (59.85%), chain-of-thought SFT recovers it (61.19%), and policy refinement delivers the largest gain (67.32%).
- The augmented model retains accuracy gains over its backbone under the tested label-preserving corruptions and improves Accuracy and Macro-F1 in the restricted overlapping-label transfer diagnostic in both directions.
- BRG Recall is high on correct predictions (97.70% on iMiGUE after refinement), and the paper explicitly reports rationale length as a confound rather than as evidence of better reasoning.
Where Pith is reading between the lines
- Editorial inference: the gated residual design suggests a general adapter recipe—any auxiliary signal (optical flow, audio onsets, joint positions) could be injected as a compact token with a near-closed gate, letting the pretrained model decide how much of the signal to trust.
- Editorial inference: because the chain-of-thought supervision carries a 24.86% severe-hallucination rate, the reasoning-grounding results may overstate faithfulness; a length-matched grounding metric (for example, truncating rationales to equal word counts before measuring body-region overlap) would separate genuine grounding from verbosity.
- Editorial inference: the cross-domain diagnostic rests on 22 target samples, so the 4.54-point accuracy gap is a single sample; the transfer benefit is not yet established. A larger overlapping-label set with matched per-class support would be the decisive test.
- Editorial inference: adjacent-frame differencing captures only short-range motion; micro-gestures that unfold over multiple seconds or involve cumulative drift would likely need a longer-range temporal model, a direction the paper itself flags.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes GMoT, a gated motion-aware tokenization module for MLLMs that computes spatial saliency, adjacent-frame temporal differencing, and gated fusion to inject a compact motion token into the visual stream. The method is trained with a four-stage recipe (GMoT warm-up, video-label pairing, CoT SFT, and reward-guided policy refinement), using semi-automatically generated CoT annotations. Evaluations on iMiGUE and SMG report Top-1 accuracy improvements over Qwen3-VL-8B (67.32% vs. 60.52% on iMiGUE; 73.11% vs. 70.00% on SMG), along with a new BRG Recall metric and an overlapping-label cross-domain transfer diagnostic. The paper claims that GMoT improves in-domain accuracy, retains gains under label-preserving corruptions, and maintains high anatomical grounding in generated rationales, while explicitly acknowledging several limitations.
Significance. If the central attribution claim holds, GMoT is a useful, lightweight plug-in for injecting motion evidence into video MLLMs, and the four-stage training recipe with reward-guided refinement is a reusable recipe for subtle-motion reasoning tasks. The paper has several strengths: the module is simple and backbone-agnostic (demonstrated on Qwen2.5-VL and Qwen3-VL at two scales); the temporal-differencing and conservative gate design is principled; and the authors go beyond Top-1 accuracy by introducing BRG Recall and a transfer diagnostic, with explicit caveats about small splits and lexical proxy limitations. The human audit of annotation quality is transparent and is reported in the main text. However, the headline attribution of the accuracy gain to the GMoT module itself is not fully supported by the ablations as presented, and the reasoning-grounding evidence is weakened by a known length confound and noisy supervision. The contribution is potentially sound, but the evidence needs tightening before the claims are acceptable.
major comments (3)
- [Abstract / §4.4.1, Table 6] The headline gain "+6.80" (Abstract and Table 1) is computed against Qwen3-VL-8B at 60.52%, but Table 6 shows S1 = S0 + Policy Refinement reaches 62.96%. Since S1 is the strongest no-GMoT pipeline with the same final RL stage, the matched GMoT contribution is S4 − S1 = +4.36 points, not +6.80. This gap is material because the largest absolute jump in Table 6 comes from Stage 3 RL (+6.13 on the GMoT-initialized model and +2.44 on the vanilla model). Moreover, Table 7 ablates the spatial scorer, temporal differencing, and semantic gate only at Stage 2 (61.19%), i.e., before policy refinement; there is no final-stage ablation. Without a full-pipeline no-GMoT control (or component ablations at the final stage), the claim that the motion token module drives the improvement is not established. Please report S4 without each GMoT component, or at least S4 without the entire module, with repeated
- [§3.2 / §4.2, Table 3 and human audit] The reasoning-grounding claim is only weakly supported. The CoT annotations were generated by LMMs under prompts that explicitly emphasize action-relevant body regions and suppress background/identity cues, and BRG Recall (§4.2, Eq. 11) checks for lexical overlap with class-specific keyword sets. Unsurprisingly, models trained on these annotations score high on BRG. Table 3 shows Avg. Len. increases from 45.08 to 64.84 words, a direct confound for BRG; the authors acknowledge this but provide no matched-length comparison. The human audit of 181 accepted SMG descriptions reports a 24.86% severe-hallucination rate (45/181 under the conservative rule), which further undermines the claim of "high anatomical grounding" as more than a lexical artifact. I recommend either (a) a human-annotated grounding evaluation on final model outputs, (b) a BRG comparison controlled for rationale length, or
- [Tables 1, 2, 6] No error bars or statistical significance tests are reported for any main accuracy comparison. Given the small datasets (iMiGUE and SMG test splits are not large, and the iMiGUE→SMG transfer set has only 22 samples), differences of 1–4 points may not be reliable. For example, Table 2 reports Weighted-F1 decreases for the GMoT model in both transfer directions while Accuracy/Macro-F1 increase; the authors correctly downplay this, but similar caution should apply to the in-domain gains. Please report mean±std over at least 3 seeds, or otherwise justify that the differences exceed run-to-run variability.
minor comments (4)
- [§3.4 / Table 6] The notation for stages is confusing: §3.4 defines Stage 0–3, while Table 6 calls the ablation rows S0–S4. In particular, "Stage 2" in Table 7 appears to mean the pre-RL state (S3 in Table 6), not Stage 2 as defined in §3.4. Please unify the stage labels (e.g., use S0–S4 consistently in both text and tables).
- [Table 3] The row "Qwen3-VL-8B [3]" in Table 3 reports 62.96%, which is S1 from Table 6, not the vanilla 60.52% from Table 1. This should be stated explicitly in the caption to avoid confusion with the backbone baseline.
- [§4.1] The video sampling protocol (4 FPS, at most 12 frames, 896 visual tokens) is a critical hyperparameter that strongly affects motion-sensitive processing. It is mentioned only in Implementation Details; a brief discussion of the sensitivity to sampling rate would strengthen the paper.
- [§5 / §4.2] The Limitations section is honest but should be cross-referenced in the Abstract and Conclusion. Currently the abstract states "improves accuracy-oriented cross-domain transfer under explicit small-split caveats while maintaining high anatomical grounding," which reads as a stronger claim than the limitations that follow.
Circularity Check
Central accuracy claim is externally validated and not circular; BRG reasoning metric is partially self-aligned with the annotation prompt.
specific steps
-
self definitional
[§3.2 Semi-automatic Reasoning Annotation; §4.2 Quantitative Results, Eq. (11)]
"the models are required to focus on action-relevant regions, with primary attention to the hands and arms, and to describe only observable motion patterns that support the MG label. ... We therefore define Body-Region Grounding (BRG) Recall by mapping each class c to a region-specific keyword set Kc and checking whether the generated reasoning trace Ti overlaps with Kc"
The CoT supervision used in Stage 2 is generated under prompts that explicitly require descriptions to focus on action-relevant body regions. BRG then scores generated rationales by lexical overlap with body-region keyword sets, so it checks for exactly the property that the annotation prompt enforced and that the SFT recipe trains the model to reproduce. High BRG is therefore partly a byproduct of the annotation/training construction rather than independent evidence of visual grounding. The paper discloses the lexical/verbosity confound and explicitly says BRG does not establish causal faithfulness, which limits severity, but the reported BRG comparison is not an independent test of grounding.
full rationale
The main Top-1 accuracy results are measured against external ground-truth labels, so the central claim that GMoT improves iMiGUE/SMG accuracy is not circular. The paper contains no load-bearing self-citation chain: citations to Qwen3-VL, DeepSeek-Math/GRPO, and the datasets are independent or non-architectural. The headline +6.80 gain is computed against the un-refined SFT baseline (60.52); Table 6 S1 shows the same backbone with policy refinement reaches 62.96, implying a more matched GMoT increment of about +4.36. That is an experimental-attribution and reporting issue, not a circularity. The only partial circularity is in the reasoning evaluation: BRG Recall rewards body-region keyword overlap, and the CoT annotations were generated under prompts demanding exactly such body-region focus, so high BRG is partly by construction. The paper candidly labels BRG a lexical proxy and discloses the length confound and the 24.86% severe-hallucination audit, which confines the issue to a secondary, well-flagged metric rather than the accuracy claim.
Axiom & Free-Parameter Ledger
free parameters (7)
- Near-closed gate initialization =
sigma(g) ≈ 0 at start
- Label reward coefficients =
+3.0, -1.0, -2.0
- Lazy-prediction penalty coefficient and high-frequency set H =
-0.5; predefined class set
- Format reward coefficients =
+0.5/-0.5, -1.5
- Observation penalty coefficient and keyword set K_bg =
-1.0; predefined keyword set
- Video sampling protocol =
4 FPS, ≤12 frames, 786,432 px, 896 tokens
- BRG keyword sets K_c =
per-class keyword lists
axioms (5)
- domain assumption Adjacent-frame temporal differencing captures discriminative micro-gesture motion
- domain assumption Sampling at 4 FPS with at most 12 frames preserves subtle motion evidence
- ad hoc to paper Two-LMM candidate generation with Qwen3-VL-MoE selection yields reliable CoT annotations
- domain assumption Class-level reference descriptions from Deepseek are accurate semantic anchors
- ad hoc to paper BRG keyword overlap measures anatomical grounding
read the original abstract
Micro-gesture recognition demands the detection of fleeting, spatially localized movements that are frequently overwhelmed by dominant static appearances and background noise. While Multimodal Large Language Models (MLLMs) excel at general video understanding, they inherently struggle with subtle kinematics and often rely on static posture priors. To this end, we propose GMoT, a Gated Motion-Aware Tokenization module that explicitly distills sparse kinematic evidence into a compact sequence prior to temporal modeling. GMoT dynamically spotlights action-relevant regions via spatially weighted pooling, extracts adjacent-frame temporal differencing to capture precise motion energy, and adaptively fuses these cues into the visual stream using a conservatively initialized semantic gate. To transition from simple classification to evidence-grounded reasoning, we further introduce a progressive reward-guided policy refinement paradigm, supported by a semi-supervised annotation pipeline that generates anatomically focused captions. Beyond achieving the best Top-1 accuracy among the compared methods on iMiGUE (67.32\%) and SMG (73.11\%), improving the Qwen3-VL-8B baseline by +6.80 and +3.11 points, our framework introduces Body-Region Grounding (BRG) Recall as an anatomical-grounding proxy conditioned on correct predictions, together with an overlapping-label cross-domain transfer protocol between iMiGUE and SMG. Extensive evaluations demonstrate that our GMoT-augmented model improves in-domain accuracy, retains clear gains under label-preserving corruptions, and improves accuracy-oriented cross-domain transfer under explicit small-split caveats while maintaining high anatomical grounding in its generated rationales.
Figures
Reference graph
Works this paper leans on
-
[1]
Axtell and Mike Fornwald
Roger E. Axtell and Mike Fornwald. 1998.Gestures: The Do’s and Taboos of Body Language Around the World. John Wiley & Sons, New York
1998
-
[2]
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Jun- yang Lin, Chang Zhou, and Jingren Zhou. 2023. Qwen-VL: A Versatile Vision- Language Model for Understanding, Localization, Text Reading, and Beyond. arXiv:2308.12966 [cs.CV]
Pith/arXiv arXiv 2023
-
[3]
Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, Wenbin Ge, Zhifang Guo, Qidong Huang, Jie Huang, Fei Huang, Binyuan Hui, Shutong Jiang, Zhaohai Li, Mingsheng Li, Mei Li, Kaixin Li, Zicheng Lin, Junyang Lin, Xuejing Liu, Jiawei Liu, Chenglong Liu, Yang Liu, Dayiheng Liu, Shixuan ...
Pith/arXiv arXiv 2025
-
[4]
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. 2025. Qwen2.5-VL Technical Rep...
Pith/arXiv arXiv 2025
-
[5]
Gedas Bertasius, Heng Wang, and Lorenzo Torresani. 2021. Is Space-Time Attention All You Need for Video Understanding? arXiv:2102.05095 [cs.CV] https://arxiv.org/abs/2102.05095
Pith/arXiv arXiv 2021
-
[6]
Joao Carreira and Andrew Zisserman. 2017. Quo vadis, action recognition? a new model and the kinetics dataset. Inproceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 6299–6308
2017
-
[7]
Boyu Chen, Zhengrong Yue, Siran Chen, Zikang Wang, Yang Liu, Peng Li, and Yali Wang. 2025. Lvagent: Long video understanding by multi-round dynamical collaboration of mllm agents.arXiv preprint arXiv:2503.10200(2025)
arXiv 2025
-
[8]
Guoliang Chen, Fei Wang, Kun Li, Zhiliang Wu, Hehe Fan, Yi Yang, Meng Wang, and Dan Guo. 2024. Prototype learning for micro-gesture classification. In Proceedings of the IJCAI 2024 Workshop and Challenge on Micro-gesture Analysis for Hidden Emotion Understanding (CEUR Workshop Proceedings, Vol. 3848). https: //ceur-ws.org/Vol-3848/paper_3.pdf
2024
-
[9]
Guo Chen, Yin-Dong Zheng, Jiahao Wang, Jilan Xu, Yifei Huang, Junting Pan, Yi Wang, Yali Wang, Yu Qiao, Tong Lu, and Limin Wang. 2023. VideoLLM: Modeling Video Sequence with Large Language Models. arXiv:2305.13292 [cs.CV] https://arxiv.org/abs/2305.13292
Pith/arXiv arXiv 2023
-
[10]
Haoyu Chen, Xin Liu, Xiaobai Li, Henglin Shi, and Guoying Zhao. 2019. Analyze Spontaneous Gestures for Emotional Stress State Recognition: A Micro-gesture Dataset and Analysis with Deep Learning. In2019 14th IEEE International Con- ference on Automatic Face and Gesture Recognition. IEEE, 1–8
2019
-
[11]
Haoyu Chen, Henglin Shi, Xin Liu, Xiaobai Li, and Guoying Zhao. 2023. Smg: A micro-gesture dataset towards spontaneous body gestures for emotional stress state analysis.International Journal of Computer Vision131, 6 (2023), 1346–1366
2023
-
[12]
Shuimu Chen, Yuteng Chen, Yuanshen Guan, Zebang Cheng, Zeyu Zhang, Shengqian Qin, Bin Xia, Jiaran Li, Wenming Yang, and Fei Ma. 2026. Reflect- R1: Evidence-Driven Reflection for Self-Correction in Long Video Understand- ing. InEuropean Conference on Computer Vision. arXiv:2606.27922 [cs.CV] doi:10.48550/arXiv.2606.27922
-
[13]
Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Erfei Cui, Jinguo Zhu, Shenglong Ye, Hao Tian, Zhaoyang Liu, et al . 2024. Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling. arXiv:2412.05271 [cs.CV] doi:10.48550/arXiv.2412.05271
-
[14]
Zebang Cheng, Shuimu Chen, Boxue Yang, Yuanshen Guan, Jingyi Chen, Zheng Lian, Xiaojiang Peng, Fei Ma, Laizhong Cui, and Qi Tian. 2026. OmniOPSD: Rationale-Privileged On-Policy Self-Distillation for Affective Computing.arXiv preprint arXiv:2606.15920(2026). arXiv:2606.15920 [cs.CV] doi:10.48550/arXiv. 2606.15920
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2606.15920 2026
-
[15]
Gheorghe Comanici, Eric Bieber, Mike Schaekermann, et al. 2025. Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities. arXiv:2507.06261 [cs.CL] doi:10.48550/ arXiv.2507.06261
-
[16]
Kaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo, Yibing Wang, Tianshuo Peng, Junfei Wu, Xiaoying Zhang, Benyou Wang, and Xiangyu Yue. 2025. Video- R1: Reinforcing Video Reasoning in MLLMs. arXiv:2503.21776 [cs.CV] https: //arxiv.org/abs/2503.21776
Pith/arXiv arXiv 2025
-
[17]
Jihao Gu, Fei Wang, Kun Li, Yanyan Wei, Zhiliang Wu, and Dan Guo. 2025. MM- Gesture: Towards Precise Micro-Gesture Recognition through Multimodal Fusion. InProceedings of the IJCAI-2025 Workshop and Challenge on Human Behavior Analysis for Emotion Understanding (CEUR Workshop Proceedings, Vol. 4168). https://ceur-ws.org/Vol-4168/paper_2.pdf
2025
-
[18]
Dan Guo, Kun Li, Bin Hu, Yan Zhang, and Meng Wang. 2024. Benchmarking micro-action recognition: Dataset, methods, and applications.IEEE Transactions on Circuits and Systems for Video Technology34, 7 (2024), 6238–6252
2024
-
[19]
Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. LoRA: Low-Rank Adaptation of Large Language Models. arXiv:2106.09685 [cs.CL] https://arxiv.org/abs/2106.09685
Pith/arXiv arXiv 2021
-
[20]
Hexiang Huang, Yuhan Wang, Kerui Linghu, and Zhaoqiang Xia. 2024. Multi- modal micro-gesture classification via multiscale heterogeneous ensemble net- work. InProceedings of the IJCAI 2024 Workshop and Challenge on Micro- gesture Analysis for Hidden Emotion Understanding (CEUR Workshop Proceedings, Vol. 3848). https://ceur-ws.org/Vol-3848/paper_2.pdf
2024
-
[21]
Libo Huang, Xiangqi Li, Jiarui Zhao, Zhulin An, Chuanguang Yang, Boyu Diao, Fei Wang, Yan Zeng, Zhifeng Hao, and Yongjun Xu. 2026. PrePrompt: Predictive Prompting for Class-Incremental Learning. InProceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining
2026
-
[22]
Libo Huang, Yan Zeng, Chuanguang Yang, Zhulin An, Boyu Diao, and Yongjun Xu. 2024. eTag: Class-Incremental Learning via Embedding Distillation and Task- Oriented Generation.Proceedings of the AAAI Conference on Artificial Intelligence 38, 11 (2024), 12591–12599. doi:10.1609/aaai.v38i11.29153
-
[23]
Deng Li, Xin Liu, Bohao Xing, Baiqiang Xia, Yuan Zong, Bihan Wen, and Heikki Kälviäinen. 2024. EALD-MLLM: Emotion Analysis in Long- sequential and De-identity videos with Multi-modal Large Language Model. arXiv:2405.00574 [cs.CV] doi:10.48550/arXiv.2405.00574
-
[24]
Deng Li, Jun Shao, Bohao Xing, Rong Gao, Bihan Wen, Heikki Kälviäinen, and Xin Liu. 2026. MSF-Mamba: Motion-Aware State Fusion Mamba for Efficient Micro-Gesture Recognition.IEEE Transactions on Multimedia(2026), 1–12. doi:10. 1109/TMM.2026.3668511
arXiv 2026
-
[25]
Deng Li, Bohao Xing, Xin Liu, Baiqiang Xia, Bihan Wen, and Heikki Kälviäinen
-
[26]
Kun Li, Dan Guo, Guoliang Chen, Xinge Peng, and Meng Wang. 2023. Joint Skeletal and Semantic Embedding Loss for Micro-gesture Classification. arXiv:2307.10624 [cs.CV] https://arxiv.org/abs/2307.10624
Pith/arXiv arXiv 2023
-
[27]
Kunchang Li, Xinhao Li, Yi Wang, Yinan He, Yali Wang, Limin Wang, and Yu Qiao. 2024. Videomamba: State space model for efficient video understanding. InEuropean conference on computer vision. Springer, 237–255
2024
-
[28]
Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Luo, et al. 2024. Mvbench: A comprehensive multi-modal video understanding benchmark. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 22195–22206
2024
-
[29]
Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Limin Wang, and Yu Qiao
-
[30]
Zheng Lian, Haoyu Chen, Lan Chen, Haiyang Sun, Licai Sun, Yong Ren, Zebang Cheng, Bin Liu, Rui Liu, Xiaojiang Peng, Jiangyan Yi, and Jianhua Tao. 2025. AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language Models. InProceedings of the 42nd Interna- tional Conference on Machine Learning (Proceedings of Mach...
2025
-
[31]
Ji Lin, Chuang Gan, and Song Han. 2019. Tsm: Temporal shift module for efficient video understanding. InProceedings of the IEEE/CVF international conference on computer vision. 7083–7093
2019
-
[32]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual In- struction Tuning. arXiv:2304.08485 [cs.CV] https://arxiv.org/abs/2304.08485
Pith/arXiv arXiv 2023
-
[33]
Xin Liu, Henglin Shi, Haoyu Chen, Zitong Yu, Xiaobai Li, and Guoying Zhao
-
[34]
Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu
-
[35]
Zhishu Liu, Kaishen Yuan, Bo Zhao, Hui Ma, and Zitong Yu. 2026. AULLM++: Structured-Token-Conditioned Large Language Models for Micro-Expression Action Unit Detection.arXiv preprint arXiv:2603.08387(2026)
Pith/arXiv arXiv 2026
-
[36]
Fei Ma, Yucheng Yuan, Yifan Xie, Hongwei Ren, Ivan Liu, Ying He, Fuji Ren, Fei Richard Yu, and Shiguang Ni. 2025. Generative Technology for Human Emotion Recognition: A Scoping Review.Information Fusion115 (2025), 102753. doi:10.1016/j.inffus.2024.102753
arXiv 2025
-
[37]
OpenAI. 2024. GPT-4o System Card. arXiv:2410.21276 [cs.CL] doi:10.48550/ arXiv.2410.21276
-
[38]
arXiv:2106.13230 [cs.CV] https://arxiv.org/abs/ 2106.13230
Video Swin Transformer. arXiv:2106.13230 [cs.CV] https://arxiv.org/abs/ 2106.13230
-
[39]
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020. ZeRO: Memory Optimizations Toward Training Trillion Parameter Models. arXiv:1910.02054 [cs.LG] https://arxiv.org/abs/1910.02054
Pith/arXiv arXiv 2020
-
[40]
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo. 2024. DeepSeek- Math: Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv:2402.03300 [cs.CL] https://arxiv.org/abs/2402.03300
Pith/arXiv arXiv 2024
-
[41]
Kun Su, Xiulong Liu, and Eli Shlizerman. 2020. Predict & cluster: Unsupervised skeleton based action recognition. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 9631–9640
2020
-
[42]
Santosh Patapati, Trisanth Srinivasan, and Amith Adiraju. 2025. CLIP-MG: Guiding Semantic Attention with Skeletal Pose Features and RGB Data for Micro- Gesture Recognition on the iMiGUE Dataset. InProceedings of the IJCAI-2025 Workshop and Challenge on Human Behavior Analysis for Emotion Understanding (CEUR Workshop Proceedings, Vol. 4168). https://ceur-w...
2025
-
[43]
Michael Tschannen, Alexey Gritsenko, Xiao Wang, Muhammad Ferjad Naeem, Ibrahim Alabdulmohsin, Nikhil Parthasarathy, Talfan Evans, Lucas Beyer, Ye Xia, Basil Mustafa, Olivier Hénaff, Jeremiah Harmsen, Andreas Steiner, and Xiaohua Zhai. 2025. SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Featur...
Pith/arXiv arXiv 2025
-
[44]
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool. 2018. Temporal segment networks for action recognition in videos.IEEE transactions on pattern analysis and machine intelligence41, 11 (2018), 2740–2755
2018
-
[45]
Tao Wang, Xue Lin, Yixing Xu, Qian Ye, Dan Guo, Sergio Escalera, George Khoriba, and Zhiyong Yu. 2026. Micro-gesture recognition: A comprehensive survey of datasets, methods, and challenges.Machine Intelligence Research23, 2 (2026), 308–331. doi:10.1007/s11633-025-1629-x
-
[46]
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri
-
[47]
Yelin Wang, Zijia Song, Shuo Ye, Chuanguang Yang, Miaoyu Wang, Yong Xu, Zhulin An, Yongjun Xu, and Zitong Yu. 2026. RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning.arXiv preprint arXiv:2606.28266(2026)
Pith/arXiv arXiv 2026
-
[48]
Zeheng Wang, Zitong Yu, Yijie Zhu, Bo Zhao, Haochen Liang, Taorui Wang, Wei Xia, Jiayu Zhang, Zhishu Liu, Hui Ma, Fei Ma, and Qi Tian. 2026. AffectA- gent: Collaborative Multi-Agent Reasoning for Retrieval-Augmented Multimodal Emotion Recognition. arXiv:2604.12735 [cs.CV] https://arxiv.org/abs/2604.12735
Pith/arXiv arXiv 2026
-
[49]
Zeheng Wang, Bo Zhao, Yijie Zhu, Zhishu Liu, Hui Ma, Ruixin Zhang, Shouhong Ding, Qianyu Xie, and Zitong Yu. 2026. Navigating the Emo- tion Tree: Hierarchical Hyperbolic RAG for Multimodal Emotion Recognition. arXiv:2605.18884 [cs.LG] https://arxiv.org/abs/2605.18884
Pith/arXiv arXiv 2026
-
[50]
Yiping Xie, Bo Zhao, Mingtong Dai, Jian-Ping Zhou, Yue Sun, Tao Tan, Weicheng Xie, Linlin Shen, and Zitong Yu. 2026. PhysLLM: Harnessing Large Language Models for Cross-Modal Remote Physiological Sensing. InThe Fourteenth Inter- national Conference on Learning Representations. https://openreview.net/forum? id=aR43t8OEeW
2026
-
[51]
Yelin Wang, Zijia Song, Chuanguang Yang, Miaoyu Wang, Zhulin An, Libo Huang, and Yongjun Xu. 2026. DFM: Difference Feature Modeling with Text- Guided Gated Contrastive Loss for Remote Sensing Image Change Captioning. arXiv preprint arXiv:2606.27410(2026)
Pith/arXiv arXiv 2026
-
[52]
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jing Zhou, Jingren Zhou, Junyang Lin, Kai Dang, Keqin Bao, Kexin Yang, ...
Pith/arXiv arXiv 2025
-
[53]
Qilang Ye, Wei Zeng, Meng Liu, Jie Zhang, Yupeng Hu, Zitong Yu, and Yu Zhou. 2026. When Eyes and Ears Disagree: Can MLLMs Discern Audio-Visual Confusion?. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 40. 11955–11963
2026
-
[54]
Xiangyu Zeng, Kunchang Li, Chenting Wang, Xinhao Li, Tianxiang Jiang, Ziang Yan, Songze Li, Yansong Shi, Zhengrong Yue, Yi Wang, et al. 2025. TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning. In International Conference on Learning Representations. https://openreview.net/ forum?id=nAVejJURqZ
2025
-
[55]
Hang Zhang, Xin Li, and Lidong Bing. 2023. Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. Association for Computational Linguistics, Singapore, 543–553. doi:10.18653/v1/2023.emnlp-demo.49
-
[56]
Haiwei Xue, Xiangyang Luo, Zhanghao Hu, Xin Zhang, Xunzhi Xiang, Yuqin Dai, Jianzhuang Liu, Zhensong Zhang, Minglei Li, Jian Yang, Fei Ma, Zhiyong Wu, Changpeng Yang, Zonghong Dai, and Fei Richard Yu. 2025. Human Motion Video Generation: A Survey.IEEE Transactions on Pattern Analysis and Machine Intelligence47, 11 (2025), 10709–10730. doi:10.1109/TPAMI.20...
arXiv 2025
-
[57]
Jie Zhang, Qilang Ye, Hao Zhou, Haochen Liang, and Fei Luo. 2026. MAVIS: Multi-Agent Video Retrieval via Structured Video Understanding. InFindings of the Association for Computational Linguistics: ACL 2026. 21751–21764
2026
-
[58]
Jiayu Zhang, Shuo Ye, Jiajian Huang, Yawen Cui, Taorui Wang, Wei Xia, Zeheng Wang, Haowen Tang, Hui Ma, and Zitong Yu. 2026. DeceptionX: Explainable Deception Detection with Multimodal Large Language Models.arXiv preprint arXiv:2606.11385(2026). arXiv:2606.11385 [cs.CV] doi:10.48550/arXiv.2606.11385
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2606.11385 2026
-
[59]
Xingjian Zhang, Xi Weng, Yihao Yue, Zhaoxin Fan, Wenjun Wu, and Lei Huang
-
[60]
Yuze Zhao, Jintao Huang, Jinghan Hu, Xingjun Wang, Yunlin Mao, Daoze Zhang, Hong Zhang, Zeyinzi Jiang, Zhikai Wu, Baole Ai, Ang Wang, Wenmeng Zhou, and Yingda Chen. 2025. SWIFT:A Scalable lightWeight Infrastructure for Fine- Tuning. arXiv:2408.05517 [cs.CL] https://arxiv.org/abs/2408.05517
Pith/arXiv arXiv 2025
-
[61]
Jiayu Zhang, Xun Lin, Jiajian Huang, Shuo Ye, Xiaobao Guo, Dongliang Zhu, Ruimin Hu, Dan Guo, Yanyan Liang, Zitong Yu, and Xiaochun Cao. 2026. Multi- modal Deception Detection: A Survey.Machine Intelligence Research23, 2 (2026), 284–307. doi:10.1007/s11633-025-1625-x
-
[62]
Yijie Zhu, Yibo Lyu, Zitong Yu, Rui Shao, Kaiyang Zhou, and Liqiang Nie. 2025. EmoSym: A Symbiotic Framework for Unified Emotional Understanding and Generation via Latent Reasoning. InProceedings of the 33nd ACM International Conference on Multimedia
2025
-
[63]
Yijie Zhu, Rui Shao, Ziyang Liu, Jie He, Jizhihui Liu, Jiuru Wang, and Zitong Yu. 2026. H-GAR: A Hierarchical Interaction Framework via Goal-Driven Observation-Action Refinement for Robotic Manipulation. InProceedings of the AAAI Conference on Artificial Intelligence
2026
-
[64]
Yijie Zhu, Lingsen Zhang, Zitong Yu, Rui Shao, Tao Tan, and Liqiang Nie. 2025. UniEmo: Unifying Emotional Understanding and Generation with Learnable Expert Queries.arXiv preprint arXiv:2507.23372(2025)
Pith/arXiv arXiv 2025
-
[65]
arXiv:2501.15513 [cs.CV] https://arxiv.org/abs/2501.15513
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler. arXiv:2501.15513 [cs.CV] https://arxiv.org/abs/2501.15513
-
[67]
Yijie Zhu, Jie He, Rui Shao, Kaishen Yuan, Tao Tan, Xiaochen Yuan, and Zitong Yu
-
[2015]
In Proceedings of the IEEE international conference on computer vision
Learning spatiotemporal features with 3d convolutional networks. In Proceedings of the IEEE international conference on computer vision. 4489–4497
-
[2021]
InProceedings of the IEEE/CVF conference on computer vision and pattern recognition
imigue: An identity-free video dataset for micro-gesture understanding and emotion analysis. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 10631–10642
-
[2023]
InProceedings of the IEEE/CVF International Conference on Computer Vision
Uniformerv2: Unlocking the potential of image vits for video understanding. InProceedings of the IEEE/CVF International Conference on Computer Vision. 1632– 1643
-
[2025]
In Proceedings of the 33rd ACM International Conference on Multimedia
DEEMO: De-identity Multimodal Emotion Recognition and Reasoning. In Proceedings of the 33rd ACM International Conference on Multimedia. Association for Computing Machinery, 5707–5716. doi:10.1145/3746027.3755411
-
[2026]
ΔVLA: Prior-Guided Vision-Language-Action Models via World Knowledge Variation.arXiv preprint arXiv:2603.08361(2026)
arXiv 2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.