Pith. sign in

REVIEW 3 major objections 4 minor 71 references

GMoT shows that compact gated motion tokens let pretrained video language models recognize micro-gestures from kinematics rather than static posture, improving accuracy on two benchmarks and grounding their rationales in body regions.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 04:35 UTC pith:GO5D34DH

load-bearing objection The motion-token module is plausible but the headline gain overstates its contribution — the matched gain is ~+4.36 and the module is only ablated before RL. the 3 major comments →

arxiv 2607.16322 v1 pith:GO5D34DH submitted 2026-07-15 cs.CV cs.AI

GMoT: Gated Motion-Aware Tokenization for Fine-Grained Micro-Gesture Video Reasoning with Multimodal LLMs

classification cs.CV cs.AI
keywords micro-gesture recognitionmotion-aware tokenizationmultimodal large language modelstemporal differencingvideo reasoningchain-of-thought supervisionreward-guided policy refinementbody-region grounding
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Micro-gestures—brief finger taps, shoulder shrugs, lip presses—are usually drowned out by static appearance, so video language models tend to guess from posture. GMoT tries to fix this by distilling motion into a compact token before the language model reasons: a learned spatial scorer highlights action-relevant regions, adjacent-frame differencing captures motion energy and direction, and a conservatively initialized gate injects the result into the visual stream. On top of this, the paper proposes a four-stage training schedule that moves from label supervision to chain-of-thought supervision and finally reward-guided policy refinement, plus a lexical metric (Body-Region Grounding Recall) for anatomical grounding and an overlapping-label cross-domain transfer protocol. If the claims hold, the result is a reusable recipe for making MLLMs reason from kinematics rather than static priors, with top-1 accuracy gains on both iMiGUE (67.32%) and SMG (73.11%) over the pretrained 8B backbone.

Core claim

GMoT's central claim is that a lightweight gated motion-aware tokenization module—not heavy 3D convolutions or dense space-time attention—can supply the kinematic evidence that pretrained MLLMs lack. The module computes a learned spatial saliency map per frame, weights patch tokens into a single frame feature, takes explicit adjacent-frame differences, decouples them into mean magnitude (motion energy) and signed direction, and projects the concatenation into a compact motion token. That token is fused into the aligned visual stream through a residual gate initialized near zero, so adaptation starts from the pretrained representation. The paper further claims that a staged recipe—GMoT warm-u

What carries the argument

The load-bearing object is the GMoT module, a three-part motion tokenizer: a spatial saliency scorer that softmax-weights patch tokens into a frame-level feature, adjacent-frame temporal differencing decoupled into mean-magnitude (motion energy) and signed-difference (direction) streams, and a near-closed gated residual fusion that injects the resulting compact motion token into the aligned visual stream. The gate is initialized so that the fused features start nearly identical to the original visual features, preserving the pretrained visual-language interface while letting the model learn how much motion evidence to trust. Around this module, the paper builds a four-stage training schedule

Load-bearing premise

The load-bearing premise is that the semi-automatically generated chain-of-thought annotations are reliable enough to supervise evidence-grounded reasoning; the paper's own audit found a 24.86% severe-hallucination rate among accepted SMG descriptions, so if the model inherits those hallucination patterns the reasoning-grounding claims weaken.

What would settle it

A decisive test: build a micro-gesture set of static-pose-matched pairs (same posture, different subtle motion) and compare GMoT against the vanilla backbone; if accuracy does not improve when motion is the only discriminative cue, the motion token is not supplying the kinematic evidence claimed. A cheaper check is a length-matched BRG comparison—if the grounding gap collapses when rationales are equalized for length, the grounding gain is a verbosity artifact.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • The motion-token branch is the main source of the accuracy gain: removing the spatial scorer, temporal differencing, or the gate each costs 4.25–5.54 points in the ablation.
  • The staged training schedule is necessary: label-only SFT slightly hurts accuracy (59.85%), chain-of-thought SFT recovers it (61.19%), and policy refinement delivers the largest gain (67.32%).
  • The augmented model retains accuracy gains over its backbone under the tested label-preserving corruptions and improves Accuracy and Macro-F1 in the restricted overlapping-label transfer diagnostic in both directions.
  • BRG Recall is high on correct predictions (97.70% on iMiGUE after refinement), and the paper explicitly reports rationale length as a confound rather than as evidence of better reasoning.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the gated residual design suggests a general adapter recipe—any auxiliary signal (optical flow, audio onsets, joint positions) could be injected as a compact token with a near-closed gate, letting the pretrained model decide how much of the signal to trust.
  • Editorial inference: because the chain-of-thought supervision carries a 24.86% severe-hallucination rate, the reasoning-grounding results may overstate faithfulness; a length-matched grounding metric (for example, truncating rationales to equal word counts before measuring body-region overlap) would separate genuine grounding from verbosity.
  • Editorial inference: the cross-domain diagnostic rests on 22 target samples, so the 4.54-point accuracy gap is a single sample; the transfer benefit is not yet established. A larger overlapping-label set with matched per-class support would be the decisive test.
  • Editorial inference: adjacent-frame differencing captures only short-range motion; micro-gestures that unfold over multiple seconds or involve cumulative drift would likely need a longer-range temporal model, a direction the paper itself flags.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper proposes GMoT, a gated motion-aware tokenization module for MLLMs that computes spatial saliency, adjacent-frame temporal differencing, and gated fusion to inject a compact motion token into the visual stream. The method is trained with a four-stage recipe (GMoT warm-up, video-label pairing, CoT SFT, and reward-guided policy refinement), using semi-automatically generated CoT annotations. Evaluations on iMiGUE and SMG report Top-1 accuracy improvements over Qwen3-VL-8B (67.32% vs. 60.52% on iMiGUE; 73.11% vs. 70.00% on SMG), along with a new BRG Recall metric and an overlapping-label cross-domain transfer diagnostic. The paper claims that GMoT improves in-domain accuracy, retains gains under label-preserving corruptions, and maintains high anatomical grounding in generated rationales, while explicitly acknowledging several limitations.

Significance. If the central attribution claim holds, GMoT is a useful, lightweight plug-in for injecting motion evidence into video MLLMs, and the four-stage training recipe with reward-guided refinement is a reusable recipe for subtle-motion reasoning tasks. The paper has several strengths: the module is simple and backbone-agnostic (demonstrated on Qwen2.5-VL and Qwen3-VL at two scales); the temporal-differencing and conservative gate design is principled; and the authors go beyond Top-1 accuracy by introducing BRG Recall and a transfer diagnostic, with explicit caveats about small splits and lexical proxy limitations. The human audit of annotation quality is transparent and is reported in the main text. However, the headline attribution of the accuracy gain to the GMoT module itself is not fully supported by the ablations as presented, and the reasoning-grounding evidence is weakened by a known length confound and noisy supervision. The contribution is potentially sound, but the evidence needs tightening before the claims are acceptable.

major comments (3)
  1. [Abstract / §4.4.1, Table 6] The headline gain "+6.80" (Abstract and Table 1) is computed against Qwen3-VL-8B at 60.52%, but Table 6 shows S1 = S0 + Policy Refinement reaches 62.96%. Since S1 is the strongest no-GMoT pipeline with the same final RL stage, the matched GMoT contribution is S4 − S1 = +4.36 points, not +6.80. This gap is material because the largest absolute jump in Table 6 comes from Stage 3 RL (+6.13 on the GMoT-initialized model and +2.44 on the vanilla model). Moreover, Table 7 ablates the spatial scorer, temporal differencing, and semantic gate only at Stage 2 (61.19%), i.e., before policy refinement; there is no final-stage ablation. Without a full-pipeline no-GMoT control (or component ablations at the final stage), the claim that the motion token module drives the improvement is not established. Please report S4 without each GMoT component, or at least S4 without the entire module, with repeated
  2. [§3.2 / §4.2, Table 3 and human audit] The reasoning-grounding claim is only weakly supported. The CoT annotations were generated by LMMs under prompts that explicitly emphasize action-relevant body regions and suppress background/identity cues, and BRG Recall (§4.2, Eq. 11) checks for lexical overlap with class-specific keyword sets. Unsurprisingly, models trained on these annotations score high on BRG. Table 3 shows Avg. Len. increases from 45.08 to 64.84 words, a direct confound for BRG; the authors acknowledge this but provide no matched-length comparison. The human audit of 181 accepted SMG descriptions reports a 24.86% severe-hallucination rate (45/181 under the conservative rule), which further undermines the claim of "high anatomical grounding" as more than a lexical artifact. I recommend either (a) a human-annotated grounding evaluation on final model outputs, (b) a BRG comparison controlled for rationale length, or
  3. [Tables 1, 2, 6] No error bars or statistical significance tests are reported for any main accuracy comparison. Given the small datasets (iMiGUE and SMG test splits are not large, and the iMiGUE→SMG transfer set has only 22 samples), differences of 1–4 points may not be reliable. For example, Table 2 reports Weighted-F1 decreases for the GMoT model in both transfer directions while Accuracy/Macro-F1 increase; the authors correctly downplay this, but similar caution should apply to the in-domain gains. Please report mean±std over at least 3 seeds, or otherwise justify that the differences exceed run-to-run variability.
minor comments (4)
  1. [§3.4 / Table 6] The notation for stages is confusing: §3.4 defines Stage 0–3, while Table 6 calls the ablation rows S0–S4. In particular, "Stage 2" in Table 7 appears to mean the pre-RL state (S3 in Table 6), not Stage 2 as defined in §3.4. Please unify the stage labels (e.g., use S0–S4 consistently in both text and tables).
  2. [Table 3] The row "Qwen3-VL-8B [3]" in Table 3 reports 62.96%, which is S1 from Table 6, not the vanilla 60.52% from Table 1. This should be stated explicitly in the caption to avoid confusion with the backbone baseline.
  3. [§4.1] The video sampling protocol (4 FPS, at most 12 frames, 896 visual tokens) is a critical hyperparameter that strongly affects motion-sensitive processing. It is mentioned only in Implementation Details; a brief discussion of the sensitivity to sampling rate would strengthen the paper.
  4. [§5 / §4.2] The Limitations section is honest but should be cross-referenced in the Abstract and Conclusion. Currently the abstract states "improves accuracy-oriented cross-domain transfer under explicit small-split caveats while maintaining high anatomical grounding," which reads as a stronger claim than the limitations that follow.

Circularity Check

1 steps flagged

Central accuracy claim is externally validated and not circular; BRG reasoning metric is partially self-aligned with the annotation prompt.

specific steps
  1. self definitional [§3.2 Semi-automatic Reasoning Annotation; §4.2 Quantitative Results, Eq. (11)]
    "the models are required to focus on action-relevant regions, with primary attention to the hands and arms, and to describe only observable motion patterns that support the MG label. ... We therefore define Body-Region Grounding (BRG) Recall by mapping each class c to a region-specific keyword set Kc and checking whether the generated reasoning trace Ti overlaps with Kc"

    The CoT supervision used in Stage 2 is generated under prompts that explicitly require descriptions to focus on action-relevant body regions. BRG then scores generated rationales by lexical overlap with body-region keyword sets, so it checks for exactly the property that the annotation prompt enforced and that the SFT recipe trains the model to reproduce. High BRG is therefore partly a byproduct of the annotation/training construction rather than independent evidence of visual grounding. The paper discloses the lexical/verbosity confound and explicitly says BRG does not establish causal faithfulness, which limits severity, but the reported BRG comparison is not an independent test of grounding.

full rationale

The main Top-1 accuracy results are measured against external ground-truth labels, so the central claim that GMoT improves iMiGUE/SMG accuracy is not circular. The paper contains no load-bearing self-citation chain: citations to Qwen3-VL, DeepSeek-Math/GRPO, and the datasets are independent or non-architectural. The headline +6.80 gain is computed against the un-refined SFT baseline (60.52); Table 6 S1 shows the same backbone with policy refinement reaches 62.96, implying a more matched GMoT increment of about +4.36. That is an experimental-attribution and reporting issue, not a circularity. The only partial circularity is in the reasoning evaluation: BRG Recall rewards body-region keyword overlap, and the CoT annotations were generated under prompts demanding exactly such body-region focus, so high BRG is partly by construction. The paper candidly labels BRG a lexical proxy and discloses the length confound and the 24.86% severe-hallucination audit, which confines the issue to a secondary, well-flagged metric rather than the accuracy claim.

Axiom & Free-Parameter Ledger

7 free parameters · 5 axioms · 0 invented entities

No physical constants are derived; this is an empirical ML paper. The central claim depends on hand-set reward coefficients, sampling choices, keyword sets, and the reliability of semi-automatic annotations. Standard learned network weights and standard LLM pretraining are omitted from the ledger.

free parameters (7)
  • Near-closed gate initialization = sigma(g) ≈ 0 at start
    Eq. (2); chosen to avoid disrupting the pretrained representation; if the gate stays closed, GMoT has no effect.
  • Label reward coefficients = +3.0, -1.0, -2.0
    Eq. (5); hand-set to strongly reward correct labels and penalize out-of-domain hallucinations.
  • Lazy-prediction penalty coefficient and high-frequency set H = -0.5; predefined class set
    Eq. (6); H is hand-curated from observed long-tail guessing behavior.
  • Format reward coefficients = +0.5/-0.5, -1.5
    Eq. (7); hand-set for <think> and <answer> block structure.
  • Observation penalty coefficient and keyword set K_bg = -1.0; predefined keyword set
    Eq. (8); hand-curated background keywords; no sensitivity analysis in main text.
  • Video sampling protocol = 4 FPS, ≤12 frames, 786,432 px, 896 tokens
    §4.1; governs temporal resolution for detecting micro-motion; no ablation on sampling rate.
  • BRG keyword sets K_c = per-class keyword lists
    Eq. (11); authors define region-specific keywords, so BRG numbers depend on these choices.
axioms (5)
  • domain assumption Adjacent-frame temporal differencing captures discriminative micro-gesture motion
    §3.3.2 and Table 7: removing temporal differencing causes the largest component drop (-5.54), so the method relies on this inductive bias.
  • domain assumption Sampling at 4 FPS with at most 12 frames preserves subtle motion evidence
    §4.1; low frame rate may alias or miss sub-frame gestures; no sampling-rate ablation is reported.
  • ad hoc to paper Two-LMM candidate generation with Qwen3-VL-MoE selection yields reliable CoT annotations
    §3.2; human audit reports 24.86% severe-hallucination rate on accepted SMG descriptions, partially contradicting this assumption.
  • domain assumption Class-level reference descriptions from Deepseek are accurate semantic anchors
    §3.2; used to condition candidate selection; no independent audit of these anchors is provided.
  • ad hoc to paper BRG keyword overlap measures anatomical grounding
    Eq. (11) and §5; authors explicitly state it is a lexical proxy, confounded by rationale length, and not causal faithfulness.

pith-pipeline@v1.3.0-alltime-deepseek · 16915 in / 15631 out tokens · 158039 ms · 2026-08-02T04:35:09.767310+00:00 · methodology

0 comments
read the original abstract

Micro-gesture recognition demands the detection of fleeting, spatially localized movements that are frequently overwhelmed by dominant static appearances and background noise. While Multimodal Large Language Models (MLLMs) excel at general video understanding, they inherently struggle with subtle kinematics and often rely on static posture priors. To this end, we propose GMoT, a Gated Motion-Aware Tokenization module that explicitly distills sparse kinematic evidence into a compact sequence prior to temporal modeling. GMoT dynamically spotlights action-relevant regions via spatially weighted pooling, extracts adjacent-frame temporal differencing to capture precise motion energy, and adaptively fuses these cues into the visual stream using a conservatively initialized semantic gate. To transition from simple classification to evidence-grounded reasoning, we further introduce a progressive reward-guided policy refinement paradigm, supported by a semi-supervised annotation pipeline that generates anatomically focused captions. Beyond achieving the best Top-1 accuracy among the compared methods on iMiGUE (67.32\%) and SMG (73.11\%), improving the Qwen3-VL-8B baseline by +6.80 and +3.11 points, our framework introduces Body-Region Grounding (BRG) Recall as an anatomical-grounding proxy conditioned on correct predictions, together with an overlapping-label cross-domain transfer protocol between iMiGUE and SMG. Extensive evaluations demonstrate that our GMoT-augmented model improves in-domain accuracy, retains clear gains under label-preserving corruptions, and improves accuracy-oriented cross-domain transfer under explicit small-split caveats while maintaining high anatomical grounding in its generated rationales.

Figures

Figures reproduced from arXiv: 2607.16322 by Hui Ma, Jiayu Zhang, Taorui Wang, Wei Xia, Yong Xu, Zeheng Wang, Zijia Song, Zitong Yu.

Figure 1
Figure 1. Figure 1: Problem setting and motivation. Generic video MLLMs often over-emphasize static appearance, whereas GMoT [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Overview of GMoT. GMoT extracts motion-sensitive evidence from short video clips and injects the resulting motion [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Overview of the four-stage training schedule. The training process progressively moves through GMoT warm-up, [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Qualitative visualization on a Moving torso sam￾ple. Compared with w/o GMoT, the model with GMoT shows more concentrated visual responses on the head-torso region. GMoT Focus line denotes the response difference map be￾tween models. 5 Limitations The iMiGUE→SMG diagnostic has only 22 target samples and a highly skewed class distribution, so its result should not be general￾ized to broad cross-domain robust… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

71 extracted references · 4 canonical work pages · 2 internal anchors

  1. [1]

    Axtell and Mike Fornwald

    Roger E. Axtell and Mike Fornwald. 1998.Gestures: The Do’s and Taboos of Body Language Around the World. John Wiley & Sons, New York

  2. [2]

    Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Jun- yang Lin, Chang Zhou, and Jingren Zhou. 2023. Qwen-VL: A Versatile Vision- Language Model for Understanding, Localization, Text Reading, and Beyond. arXiv:2308.12966 [cs.CV]

  3. [3]

    Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, Wenbin Ge, Zhifang Guo, Qidong Huang, Jie Huang, Fei Huang, Binyuan Hui, Shutong Jiang, Zhaohai Li, Mingsheng Li, Mei Li, Kaixin Li, Zicheng Lin, Junyang Lin, Xuejing Liu, Jiawei Liu, Chenglong Liu, Yang Liu, Dayiheng Liu, Shixuan ...

  4. [4]

    Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. 2025. Qwen2.5-VL Technical Rep...

  5. [5]

    Gedas Bertasius, Heng Wang, and Lorenzo Torresani. 2021. Is Space-Time Attention All You Need for Video Understanding? arXiv:2102.05095 [cs.CV] https://arxiv.org/abs/2102.05095

  6. [6]

    Joao Carreira and Andrew Zisserman. 2017. Quo vadis, action recognition? a new model and the kinetics dataset. Inproceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 6299–6308

  7. [7]

    Boyu Chen, Zhengrong Yue, Siran Chen, Zikang Wang, Yang Liu, Peng Li, and Yali Wang. 2025. Lvagent: Long video understanding by multi-round dynamical collaboration of mllm agents.arXiv preprint arXiv:2503.10200(2025)

  8. [8]

    Guoliang Chen, Fei Wang, Kun Li, Zhiliang Wu, Hehe Fan, Yi Yang, Meng Wang, and Dan Guo. 2024. Prototype learning for micro-gesture classification. In Proceedings of the IJCAI 2024 Workshop and Challenge on Micro-gesture Analysis for Hidden Emotion Understanding (CEUR Workshop Proceedings, Vol. 3848). https: //ceur-ws.org/Vol-3848/paper_3.pdf

  9. [9]

    Guo Chen, Yin-Dong Zheng, Jiahao Wang, Jilan Xu, Yifei Huang, Junting Pan, Yi Wang, Yali Wang, Yu Qiao, Tong Lu, and Limin Wang. 2023. VideoLLM: Modeling Video Sequence with Large Language Models. arXiv:2305.13292 [cs.CV] https://arxiv.org/abs/2305.13292

  10. [10]

    Haoyu Chen, Xin Liu, Xiaobai Li, Henglin Shi, and Guoying Zhao. 2019. Analyze Spontaneous Gestures for Emotional Stress State Recognition: A Micro-gesture Dataset and Analysis with Deep Learning. In2019 14th IEEE International Con- ference on Automatic Face and Gesture Recognition. IEEE, 1–8

  11. [11]

    Haoyu Chen, Henglin Shi, Xin Liu, Xiaobai Li, and Guoying Zhao. 2023. Smg: A micro-gesture dataset towards spontaneous body gestures for emotional stress state analysis.International Journal of Computer Vision131, 6 (2023), 1346–1366

  12. [12]

    Shuimu Chen, Yuteng Chen, Yuanshen Guan, Zebang Cheng, Zeyu Zhang, Shengqian Qin, Bin Xia, Jiaran Li, Wenming Yang, and Fei Ma. 2026. Reflect- R1: Evidence-Driven Reflection for Self-Correction in Long Video Understand- ing. InEuropean Conference on Computer Vision. arXiv:2606.27922 [cs.CV] doi:10.48550/arXiv.2606.27922

  13. [13]

    Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Erfei Cui, Jinguo Zhu, Shenglong Ye, Hao Tian, Zhaoyang Liu, et al . 2024. Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling. arXiv:2412.05271 [cs.CV] doi:10.48550/arXiv.2412.05271

  14. [14]

    Zebang Cheng, Shuimu Chen, Boxue Yang, Yuanshen Guan, Jingyi Chen, Zheng Lian, Xiaojiang Peng, Fei Ma, Laizhong Cui, and Qi Tian. 2026. OmniOPSD: Rationale-Privileged On-Policy Self-Distillation for Affective Computing.arXiv preprint arXiv:2606.15920(2026). arXiv:2606.15920 [cs.CV] doi:10.48550/arXiv. 2606.15920

  15. [15]

    Gheorghe Comanici, Eric Bieber, Mike Schaekermann, et al. 2025. Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities. arXiv:2507.06261 [cs.CL] doi:10.48550/ arXiv.2507.06261

  16. [16]

    Kaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo, Yibing Wang, Tianshuo Peng, Junfei Wu, Xiaoying Zhang, Benyou Wang, and Xiangyu Yue. 2025. Video- R1: Reinforcing Video Reasoning in MLLMs. arXiv:2503.21776 [cs.CV] https: //arxiv.org/abs/2503.21776

  17. [17]

    Jihao Gu, Fei Wang, Kun Li, Yanyan Wei, Zhiliang Wu, and Dan Guo. 2025. MM- Gesture: Towards Precise Micro-Gesture Recognition through Multimodal Fusion. InProceedings of the IJCAI-2025 Workshop and Challenge on Human Behavior Analysis for Emotion Understanding (CEUR Workshop Proceedings, Vol. 4168). https://ceur-ws.org/Vol-4168/paper_2.pdf

  18. [18]

    Dan Guo, Kun Li, Bin Hu, Yan Zhang, and Meng Wang. 2024. Benchmarking micro-action recognition: Dataset, methods, and applications.IEEE Transactions on Circuits and Systems for Video Technology34, 7 (2024), 6238–6252

  19. [19]

    Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

    Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. LoRA: Low-Rank Adaptation of Large Language Models. arXiv:2106.09685 [cs.CL] https://arxiv.org/abs/2106.09685

  20. [20]

    Hexiang Huang, Yuhan Wang, Kerui Linghu, and Zhaoqiang Xia. 2024. Multi- modal micro-gesture classification via multiscale heterogeneous ensemble net- work. InProceedings of the IJCAI 2024 Workshop and Challenge on Micro- gesture Analysis for Hidden Emotion Understanding (CEUR Workshop Proceedings, Vol. 3848). https://ceur-ws.org/Vol-3848/paper_2.pdf

  21. [21]

    Libo Huang, Xiangqi Li, Jiarui Zhao, Zhulin An, Chuanguang Yang, Boyu Diao, Fei Wang, Yan Zeng, Zhifeng Hao, and Yongjun Xu. 2026. PrePrompt: Predictive Prompting for Class-Incremental Learning. InProceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining

  22. [22]

    Libo Huang, Yan Zeng, Chuanguang Yang, Zhulin An, Boyu Diao, and Yongjun Xu. 2024. eTag: Class-Incremental Learning via Embedding Distillation and Task- Oriented Generation.Proceedings of the AAAI Conference on Artificial Intelligence 38, 11 (2024), 12591–12599. doi:10.1609/aaai.v38i11.29153

  23. [23]

    Deng Li, Xin Liu, Bohao Xing, Baiqiang Xia, Yuan Zong, Bihan Wen, and Heikki Kälviäinen. 2024. EALD-MLLM: Emotion Analysis in Long- sequential and De-identity videos with Multi-modal Large Language Model. arXiv:2405.00574 [cs.CV] doi:10.48550/arXiv.2405.00574

  24. [24]

    Deng Li, Jun Shao, Bohao Xing, Rong Gao, Bihan Wen, Heikki Kälviäinen, and Xin Liu. 2026. MSF-Mamba: Motion-Aware State Fusion Mamba for Efficient Micro-Gesture Recognition.IEEE Transactions on Multimedia(2026), 1–12. doi:10. 1109/TMM.2026.3668511

  25. [25]

    Deng Li, Bohao Xing, Xin Liu, Baiqiang Xia, Bihan Wen, and Heikki Kälviäinen

  26. [26]

    Kun Li, Dan Guo, Guoliang Chen, Xinge Peng, and Meng Wang. 2023. Joint Skeletal and Semantic Embedding Loss for Micro-gesture Classification. arXiv:2307.10624 [cs.CV] https://arxiv.org/abs/2307.10624

  27. [27]

    Kunchang Li, Xinhao Li, Yi Wang, Yinan He, Yali Wang, Limin Wang, and Yu Qiao. 2024. Videomamba: State space model for efficient video understanding. InEuropean conference on computer vision. Springer, 237–255

  28. [28]

    Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Luo, et al. 2024. Mvbench: A comprehensive multi-modal video understanding benchmark. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 22195–22206

  29. [29]

    Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Limin Wang, and Yu Qiao

  30. [30]

    Zheng Lian, Haoyu Chen, Lan Chen, Haiyang Sun, Licai Sun, Yong Ren, Zebang Cheng, Bin Liu, Rui Liu, Xiaojiang Peng, Jiangyan Yi, and Jianhua Tao. 2025. AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language Models. InProceedings of the 42nd Interna- tional Conference on Machine Learning (Proceedings of Mach...

  31. [31]

    Ji Lin, Chuang Gan, and Song Han. 2019. Tsm: Temporal shift module for efficient video understanding. InProceedings of the IEEE/CVF international conference on computer vision. 7083–7093

  32. [32]

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual In- struction Tuning. arXiv:2304.08485 [cs.CV] https://arxiv.org/abs/2304.08485

  33. [33]

    Xin Liu, Henglin Shi, Haoyu Chen, Zitong Yu, Xiaobai Li, and Guoying Zhao

  34. [34]

    Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu

  35. [35]

    Zhishu Liu, Kaishen Yuan, Bo Zhao, Hui Ma, and Zitong Yu. 2026. AULLM++: Structured-Token-Conditioned Large Language Models for Micro-Expression Action Unit Detection.arXiv preprint arXiv:2603.08387(2026)

  36. [36]

    Fei Ma, Yucheng Yuan, Yifan Xie, Hongwei Ren, Ivan Liu, Ying He, Fuji Ren, Fei Richard Yu, and Shiguang Ni. 2025. Generative Technology for Human Emotion Recognition: A Scoping Review.Information Fusion115 (2025), 102753. doi:10.1016/j.inffus.2024.102753

  37. [37]

    OpenAI. 2024. GPT-4o System Card. arXiv:2410.21276 [cs.CL] doi:10.48550/ arXiv.2410.21276

  38. [38]

    arXiv:2106.13230 [cs.CV] https://arxiv.org/abs/ 2106.13230

    Video Swin Transformer. arXiv:2106.13230 [cs.CV] https://arxiv.org/abs/ 2106.13230

  39. [39]

    Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020. ZeRO: Memory Optimizations Toward Training Trillion Parameter Models. arXiv:1910.02054 [cs.LG] https://arxiv.org/abs/1910.02054

  40. [40]

    Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo. 2024. DeepSeek- Math: Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv:2402.03300 [cs.CL] https://arxiv.org/abs/2402.03300

  41. [41]

    Kun Su, Xiulong Liu, and Eli Shlizerman. 2020. Predict & cluster: Unsupervised skeleton based action recognition. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 9631–9640

  42. [42]

    Santosh Patapati, Trisanth Srinivasan, and Amith Adiraju. 2025. CLIP-MG: Guiding Semantic Attention with Skeletal Pose Features and RGB Data for Micro- Gesture Recognition on the iMiGUE Dataset. InProceedings of the IJCAI-2025 Workshop and Challenge on Human Behavior Analysis for Emotion Understanding (CEUR Workshop Proceedings, Vol. 4168). https://ceur-w...

  43. [43]

    Michael Tschannen, Alexey Gritsenko, Xiao Wang, Muhammad Ferjad Naeem, Ibrahim Alabdulmohsin, Nikhil Parthasarathy, Talfan Evans, Lucas Beyer, Ye Xia, Basil Mustafa, Olivier Hénaff, Jeremiah Harmsen, Andreas Steiner, and Xiaohua Zhai. 2025. SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Featur...

  44. [44]

    Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool. 2018. Temporal segment networks for action recognition in videos.IEEE transactions on pattern analysis and machine intelligence41, 11 (2018), 2740–2755

  45. [45]

    Tao Wang, Xue Lin, Yixing Xu, Qian Ye, Dan Guo, Sergio Escalera, George Khoriba, and Zhiyong Yu. 2026. Micro-gesture recognition: A comprehensive survey of datasets, methods, and challenges.Machine Intelligence Research23, 2 (2026), 308–331. doi:10.1007/s11633-025-1629-x

  46. [46]

    Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri

  47. [47]

    Yelin Wang, Zijia Song, Shuo Ye, Chuanguang Yang, Miaoyu Wang, Yong Xu, Zhulin An, Yongjun Xu, and Zitong Yu. 2026. RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning.arXiv preprint arXiv:2606.28266(2026)

  48. [48]

    Zeheng Wang, Zitong Yu, Yijie Zhu, Bo Zhao, Haochen Liang, Taorui Wang, Wei Xia, Jiayu Zhang, Zhishu Liu, Hui Ma, Fei Ma, and Qi Tian. 2026. AffectA- gent: Collaborative Multi-Agent Reasoning for Retrieval-Augmented Multimodal Emotion Recognition. arXiv:2604.12735 [cs.CV] https://arxiv.org/abs/2604.12735

  49. [49]

    Zeheng Wang, Bo Zhao, Yijie Zhu, Zhishu Liu, Hui Ma, Ruixin Zhang, Shouhong Ding, Qianyu Xie, and Zitong Yu. 2026. Navigating the Emo- tion Tree: Hierarchical Hyperbolic RAG for Multimodal Emotion Recognition. arXiv:2605.18884 [cs.LG] https://arxiv.org/abs/2605.18884

  50. [50]

    Yiping Xie, Bo Zhao, Mingtong Dai, Jian-Ping Zhou, Yue Sun, Tao Tan, Weicheng Xie, Linlin Shen, and Zitong Yu. 2026. PhysLLM: Harnessing Large Language Models for Cross-Modal Remote Physiological Sensing. InThe Fourteenth Inter- national Conference on Learning Representations. https://openreview.net/forum? id=aR43t8OEeW

  51. [51]

    Yelin Wang, Zijia Song, Chuanguang Yang, Miaoyu Wang, Zhulin An, Libo Huang, and Yongjun Xu. 2026. DFM: Difference Feature Modeling with Text- Guided Gated Contrastive Loss for Remote Sensing Image Change Captioning. arXiv preprint arXiv:2606.27410(2026)

  52. [52]

    An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jing Zhou, Jingren Zhou, Junyang Lin, Kai Dang, Keqin Bao, Kexin Yang, ...

  53. [53]

    Qilang Ye, Wei Zeng, Meng Liu, Jie Zhang, Yupeng Hu, Zitong Yu, and Yu Zhou. 2026. When Eyes and Ears Disagree: Can MLLMs Discern Audio-Visual Confusion?. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 40. 11955–11963

  54. [54]

    Xiangyu Zeng, Kunchang Li, Chenting Wang, Xinhao Li, Tianxiang Jiang, Ziang Yan, Songze Li, Yansong Shi, Zhengrong Yue, Yi Wang, et al. 2025. TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning. In International Conference on Learning Representations. https://openreview.net/ forum?id=nAVejJURqZ

  55. [55]

    Hang Zhang, Xin Li, and Lidong Bing. 2023. Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. Association for Computational Linguistics, Singapore, 543–553. doi:10.18653/v1/2023.emnlp-demo.49

  56. [56]

    Haiwei Xue, Xiangyang Luo, Zhanghao Hu, Xin Zhang, Xunzhi Xiang, Yuqin Dai, Jianzhuang Liu, Zhensong Zhang, Minglei Li, Jian Yang, Fei Ma, Zhiyong Wu, Changpeng Yang, Zonghong Dai, and Fei Richard Yu. 2025. Human Motion Video Generation: A Survey.IEEE Transactions on Pattern Analysis and Machine Intelligence47, 11 (2025), 10709–10730. doi:10.1109/TPAMI.20...

  57. [57]

    Jie Zhang, Qilang Ye, Hao Zhou, Haochen Liang, and Fei Luo. 2026. MAVIS: Multi-Agent Video Retrieval via Structured Video Understanding. InFindings of the Association for Computational Linguistics: ACL 2026. 21751–21764

  58. [58]

    Jiayu Zhang, Shuo Ye, Jiajian Huang, Yawen Cui, Taorui Wang, Wei Xia, Zeheng Wang, Haowen Tang, Hui Ma, and Zitong Yu. 2026. DeceptionX: Explainable Deception Detection with Multimodal Large Language Models.arXiv preprint arXiv:2606.11385(2026). arXiv:2606.11385 [cs.CV] doi:10.48550/arXiv.2606.11385

  59. [59]

    Xingjian Zhang, Xi Weng, Yihao Yue, Zhaoxin Fan, Wenjun Wu, and Lei Huang

  60. [60]

    Yuze Zhao, Jintao Huang, Jinghan Hu, Xingjun Wang, Yunlin Mao, Daoze Zhang, Hong Zhang, Zeyinzi Jiang, Zhikai Wu, Baole Ai, Ang Wang, Wenmeng Zhou, and Yingda Chen. 2025. SWIFT:A Scalable lightWeight Infrastructure for Fine- Tuning. arXiv:2408.05517 [cs.CL] https://arxiv.org/abs/2408.05517

  61. [61]

    Jiayu Zhang, Xun Lin, Jiajian Huang, Shuo Ye, Xiaobao Guo, Dongliang Zhu, Ruimin Hu, Dan Guo, Yanyan Liang, Zitong Yu, and Xiaochun Cao. 2026. Multi- modal Deception Detection: A Survey.Machine Intelligence Research23, 2 (2026), 284–307. doi:10.1007/s11633-025-1625-x

  62. [62]

    Yijie Zhu, Yibo Lyu, Zitong Yu, Rui Shao, Kaiyang Zhou, and Liqiang Nie. 2025. EmoSym: A Symbiotic Framework for Unified Emotional Understanding and Generation via Latent Reasoning. InProceedings of the 33nd ACM International Conference on Multimedia

  63. [63]

    Yijie Zhu, Rui Shao, Ziyang Liu, Jie He, Jizhihui Liu, Jiuru Wang, and Zitong Yu. 2026. H-GAR: A Hierarchical Interaction Framework via Goal-Driven Observation-Action Refinement for Robotic Manipulation. InProceedings of the AAAI Conference on Artificial Intelligence

  64. [64]

    Yijie Zhu, Lingsen Zhang, Zitong Yu, Rui Shao, Tao Tan, and Liqiang Nie. 2025. UniEmo: Unifying Emotional Understanding and Generation with Learnable Expert Queries.arXiv preprint arXiv:2507.23372(2025)

  65. [65]

    arXiv:2501.15513 [cs.CV] https://arxiv.org/abs/2501.15513

    TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler. arXiv:2501.15513 [cs.CV] https://arxiv.org/abs/2501.15513

  66. [67]

    Yijie Zhu, Jie He, Rui Shao, Kaishen Yuan, Tao Tan, Xiaochen Yuan, and Zitong Yu

  67. [2015]

    In Proceedings of the IEEE international conference on computer vision

    Learning spatiotemporal features with 3d convolutional networks. In Proceedings of the IEEE international conference on computer vision. 4489–4497

  68. [2021]

    InProceedings of the IEEE/CVF conference on computer vision and pattern recognition

    imigue: An identity-free video dataset for micro-gesture understanding and emotion analysis. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 10631–10642

  69. [2023]

    InProceedings of the IEEE/CVF International Conference on Computer Vision

    Uniformerv2: Unlocking the potential of image vits for video understanding. InProceedings of the IEEE/CVF International Conference on Computer Vision. 1632– 1643

  70. [2025]

    In Proceedings of the 33rd ACM International Conference on Multimedia

    DEEMO: De-identity Multimodal Emotion Recognition and Reasoning. In Proceedings of the 33rd ACM International Conference on Multimedia. Association for Computing Machinery, 5707–5716. doi:10.1145/3746027.3755411

  71. [2026]

    ΔVLA: Prior-Guided Vision-Language-Action Models via World Knowledge Variation.arXiv preprint arXiv:2603.08361(2026)