REVIEW 3 major objections 6 minor 180 references
LightAIR: Lightweight Action Inversion and Riemannian Rectification for Text-based Person Anomaly Search
T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read LightAIR claims strict action-appearance decoupling via text anchors, null-space projection, and Riemannian gradient rectification, setting state-of-the-art scores on PAB and four TIPR benchmarks.
desk verdict Solid empirical gains and a sensible new combination for TPAS, but the 'strict forward decoupling' claim is overstated: the projection is rank-1, the paper's own appendix shows residual leakage, and the missing multi-basis ablation leaves the core assumption untested. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the orthogonal projection operator $\Pi_{\text{act}} = z_{\text{act}} z_{\text{act}}^{\top}/(\lVert z_{\text{act}} \rVert^2+\epsilon)$, built from the single reconstructed action vector. It does double duty: in the forward pass it defines the null space $(I-\Pi_{\text{act}})v$ that yields the appearance feature, and in the backward pass it defines the tangent space onto which the Euclidean gradient is projected to obtain the Riemannian gradient. The action vector itself is produced by the Action Inversion Operator, which estimates sparse Top-K coefficients over a frozen codebook of action-word embeddings and reconstructs $z_{\text{act}} = D^{\top} \boldsymbol{\alpha} / \lVert D^{\top} \boldsymbol{\alpha} \rVert_2$. An entropy-guided weight $\omega = \exp(-H/\tau)$ interpolates between the Euclidean and Riemannian gradients during training, so early optimization explores freely and later updates respect the decoupling constraint.
What would settle it
Compare LightAIR against a variant whose action subspace contains several orthogonal action directions rather than one; if the single-basis model already decouples perfectly, the multi-basis variant should perform identically, while a gain would indicate the core assumption is too strong. A direct measurement is to compute, for same-clothing/different-action pairs, the cosine similarity between $z_{\text{app}}$ and the text feature of the action description, which should be near zero for both actions under strict decoupling.
Extended reading notes
Core claim
The paper's central claim is that action and appearance can be strictly decoupled in a shared image-text space using text priors alone. An action inversion operator maps the global visual feature $v$ to sparse coefficients over a frozen action-word codebook and reconstructs $z_{\text{act}}$, then defines the appearance feature as $z_{\text{app}} = (I-\Pi_{\text{act}})v$, an orthogonal projection onto the null space of $z_{\text{act}}$. The paper shows that the Euclidean gradient of the contrastive loss with respect to $v$ splits into a tangent component and a harmful normal component, and that replacing it with the Riemannian gradient $g_{\text{Riem}} = (I-\Pi_{\text{act}})g_{\text{Euc}}$, blended by an entropy weight $\omega = \exp(-H/\tau)$, blocks the normal component's shortcut. Empirically, this yields 84.73% R@1 and 91.93% mAP on PAB at 0.1M data, and 85.49% R@1 and 92.20% mAP at 1M, surpassing the CMP baseline, while the same pipeline improves all four TIPR benchmarks.
Load-bearing premise
The decoupling argument rests on the assumption that the action content of an image is fully captured by the single reconstructed action vector $z_{\text{act}}$, so the orthogonal complement $(I-\Pi_{\text{act}})v$ contains no action information and all appearance information.
Editorial extensions
If this is right
- External pose estimators become unnecessary for TPAS: action information is supplied by text anchors, so retrieval should survive occlusion, unusual poses, and low-resolution surveillance frames.
- Hard-negative 'same appearance, different action' pairs no longer force appearance shortcuts, since the normal gradient component is blocked and updates stay on the tangent space.
- The same decoupling machinery transfers to conventional TIPR, producing best reported averages on CUHK-PEDES, ICFG-PEDES, RSTPReid, and UFineBench, with an 8.69-point average Recall gain on UFine3C.
- Out-of-distribution robustness improves: with only 0.1M training pairs LightAIR beats CMP trained on the full 1M data on the UCC set (62.36 vs 55.23 R@1; 51.53 vs 44.35 mAP).
Reading between the lines
- Beyond the paper: the whole decoupling scheme hinges on the action subspace being one-dimensional. If a pedestrian performs two simultaneous or sequential actions, a single $z_{\text{act}}$ cannot span the action content, and the null-space projection would either leak action into appearance or discard part of the action; a multi-basis version of the codebook projection is the natural stress test.
- Beyond the paper: the entropy-weighted Riemannian gradient rectification is a general anti-shortcut mechanism for any contrastive retrieval task with a weak semantic signal and a dominant distractor modality; the paper demonstrates it only for person search, but the gradient decomposition argument is task-agnostic.
- Beyond the paper: because the codebook is frozen and built from training vocabulary, deployment to genuinely novel action words may require codebook extension; the UCC experiment shifts scene distribution, not action vocabulary, so it does not test vocabulary generalization.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper targets Text-based Person Anomaly Search (TPAS), where a query describes both macro-level appearance and micro-level abnormal actions. It proposes LightAIR, combining three modules: an Action Inversion Operator (AIO) that reconstructs an action feature z_act from a frozen text-derived semantic codebook, an Orthogonal Null-Space Projection (ONSP) that obtains an appearance feature z_app by projecting the visual feature onto the null space of z_act, and a Gradient Rectification (GR) module that projects the Euclidean gradient onto a tangent space with an entropy-based adaptive weight. The method is evaluated on the PAB TPAS benchmark, its Multi-Weather and UCC variants, and four TIPR benchmarks, reporting state-of-the-art results, e.g., PAB 0.1M R@1 of 84.73% and mAP of 91.93%, and UCC R@1 of 62.36% with only 0.1M training data. Ablations are provided for each module across the TPAS settings.
Significance. If the reported results are reproducible, LightAIR would be a strong new state of the art for TPAS and also improves conventional TIPR benchmarks. The design avoids external pose estimators and the code is released, which are practical strengths. The ablations are reasonably complete at the module level and the reported gains over CMP are consistent across PAB, Multi-Weather, and UCC. However, the central mathematical claim of 'strict forward decoupling' is only established for the single direction defined by z_act; the paper's own failure cases in Appendix E acknowledge residual action–appearance interference, and the current experiments do not test a multi-dimensional action-subspace projector. The core idea is promising, but the decoupling guarantee and its empirical support need to be tightened before the headline claims are acceptable.
major comments (3)
- [§3.3, Eqs. (6)–(7); Abstract; Contributions] The claim that ONSP 'mathematically eliminates' action information from z_app is not established by the construction. The projector Pi_act in Eq. (6) is rank-1, built from the single vector z_act of Eq. (4), so Eq. (7) removes only the component of v parallel to z_act. If TPAS action semantics occupy more than one direction in the embedding space, the residual action information in z_app remains. This is not merely hypothetical: Appendix E, Figure 11(c) states that 'weak local action features are easily overshadowed by explicit large-area appearance attributes' and that 'the current forward decoupling mechanism still faces the risk of interference from dominant appearance information.' Please either restrict the claim to removal of the component along the estimated action direction, or provide evidence that a single direction suffices. A concrete test would be to train a linear probe on z_app to measure remaining action information, and to ablate a multi-dimensional action-basis projector built from the top-K codebook directions.
- [§3.2, Eq. (5); Algorithm 1, lines 6 and 10] The action discriminative loss L_cls aligns z_act with t_y = Phi_T(x_T), which is the feature of the full text query, not an action-only feature. Because the query text also describes appearance (e.g., clothing, background), this supervision can pull z_act toward appearance directions, weakening the claim that AIO extracts 'pure action features' and potentially reintroducing appearance into the action branch before ONSP operates. Please clarify whether t_y is restricted to action words (for example, by projecting the text feature onto the action codebook D) or provide an ablation that replaces the full-text feature with an action-only feature in L_cls.
- [§4.3.2, Figure 4 and accompanying text] The text states that 'using multiple action bases' lowers Multi-Weather R@1 by 3.02 points, but the ablation D#6 is labeled 'w/o num_K' (removing Top-K sparsification). Removing Top-K changes z_act to a dense combination of codebook atoms, yet the projection operator in Eq. (6) remains rank-1. Thus the experiment does not test a multi-dimensional action-basis projector and cannot support the conclusion that a single action direction is sufficient. Please correct the description and add an explicit experiment with a projector onto a multi-dimensional action subspace, or discuss why the rank-1 projector is adequate despite the Appendix E failure cases.
minor comments (6)
- [§3.4, Eq. (8)] The two terms in the gradient decomposition are called 'orthogonal components,' but no orthogonality is shown; the second term, which involves the derivative of Pi_act, is not generally contained in the range of Pi_act. Please clarify the claim or soften the wording. Also, the symbol ×1 is used without definition.
- [§3.4, Eq. (9)] The term 'Riemannian gradient' is used for an orthogonal projection onto a linear subspace T_vM, but the manifold M is never formally defined. If the tangent space is just the null space of Pi_act, the analysis is a linear projection; please either define the manifold or use a more neutral term such as 'projected gradient.'
- [§4.3, Figures 3–6] The captions of the ablation figures do not state whether the reported numbers correspond to the 0.1M or 1M training setting; the main text refers to both. Please state the training data scale in each caption or in the surrounding text.
- [Table 2] The rows for CLIP and X-VLM without a #Data value are not directly comparable to the 0.1M and 1M rows. Please specify the training data used for those baseline rows or remove them from the comparison table.
- [References] Reference [150] is listed as 'MRA (arXiv'25)' in Table 2, but the reference entry is identical to [59], the CMP ICCV'25 paper. Please correct the duplicate or replace it with the intended MRA reference.
- [§4.1.2, Implementation Details] The statement 'We equivalently implement Riemannian gradient rectification by applying a gradient clipping operation to Pi_act' is vague. Please specify how the clipping is applied to the projection operator and how it relates to Eq. (11).
Circularity Check
No circularity: the action feature is text-supervised, the projections are explicit geometric constructions, and all reported metrics are held-out test results.
full rationale
LightAIR's derivation chain is not circular. The action feature z_act (Eq. 4) is trained against the text-derived action feature t_y via the discriminative loss L_cls (Eq. 5); the projector Pi_act (Eq. 6) and the null-space appearance feature z_app (Eq. 7) are then built directly from z_act, and retrieval is evaluated on held-out PAB, UCC, and TIPR test splits. No fitted constant is later relabeled as a prediction: every R@1/mAP value is measured on test data, and the ablations compare functionally distinct variants (w/o L_cls, w/o Ortho, w/o Rectification, etc.). The only self-referential element is that z_app is orthogonal to z_act by construction, but removing the component along the reconstructed action direction is the intended geometric operation rather than a hidden reuse of the target metric. The rank-1 subspace assumption behind Eqs. 6-7 is a genuine correctness risk: Appendix E, Figure 11(c), admits that 'weak local action features are easily overshadowed by explicit large-area appearance attributes' and that 'the current forward decoupling mechanism still faces the risk of interference from dominant appearance information,' and no ablation tests a multi-dimensional action-basis projector. However, an overclaim or an untested assumption is a robustness limitation, not circularity. Although the paper cites many works by its own authors in the Related Work sections, none of those citations is load-bearing: the null-space projection is attributed to Ravfogel et al. [139], the Riemannian gradient to Absil et al. [140], and shortcut learning to Geirhos et al. [73]. There is no imported uniqueness theorem, no ansatz smuggled in via self-citation, and no fitted input presented as a prediction. The empirical claims are self-contained and evaluated against external benchmarks, so the appropriate circularity score is 0.
Assumptions & free parameters
free parameters (5)
- Top-K sparsity K =
3
- image-text matching loss weight gamma_1 =
4
- action discriminative loss weight gamma_2 =
1
- entropy temperature tau =
unspecified
- semantic codebook size L =
unspecified
assumptions (4)
- ad hoc to paper The action subspace is fully spanned by the single normalized feature z_act from Eq. 4.
- domain assumption The frozen text semantic codebook D provides reliable and complete anchors for all relevant actions, including OOD actions.
- standard math Riemannian gradient on the action semantic manifold equals orthogonal projection onto the null space of Pi_act.
- ad hoc to paper The entropy-guided weight omega=exp(-H/tau) is a valid measure of action mapping confidence.
Cite this review
Pith. "Pith review of LightAIR: Lightweight Action Inversion and Riemannian Rectification for Text-based Person Anomaly Search." pith.science (2026). https://pith.science/paper/Z6UMUAEB
@misc{pith2026260809152,
author = {Pith},
title = {Pith review of: LightAIR: Lightweight Action Inversion and Riemannian Rectification for Text-based Person Anomaly Search},
year = {2026},
howpublished = {\url{https://pith.science/paper/Z6UMUAEB}},
note = {Machine review of arXiv:2608.09152}
}
read the original abstract
Traditional Text-based Person Search (TPS) is typically limited to matching static appearance attributes, severely neglecting dynamic action information. The Text-based Person Anomaly Search (TPAS) task bridges this gap, requiring models to locate micro-level specific abnormal behaviors while matching macro-level appearance of pedestrians. However, current TPAS methods face fundamental limitations: external explicit pose estimators are fragile in unconstrained surveillance scenarios, and implicit learning encounters visual decoupling failure under pixel-level entanglement, causing dominant appearance information to easily swallow and contaminate subtle action features. Furthermore, performing contrastive optimization on hard negative samples (``same appearance, different actions'') in conventional Euclidean spaces induces severe shortcut learning. To address these, we propose the Lightweight Action Inversion and Riemannian rectification network (LightAIR). First, it introduces textual semantic priors as anchors via a lightweight action inversion operator to extract pure action features, thereby overcoming visual-inherent coupling. Subsequently, it employs orthogonal null-space projection to constrain appearance features within the orthogonal complement space of action features, guaranteeing strict forward decoupling. Finally, we designed a gradient rectification module that computes the Riemannian gradient to constrain the backpropagation trajectory, forcing the gradient flow to update strictly along the tangent space that preserves decoupling properties, thereby cutting off harmful shortcuts. Extensive experiments on the widely used TPAS and TIPR datasets demonstrate that LightAIR significantly outperforms existing state-of-the-art methods. Codes are available at https://github.com/rainy-london/LightAIR
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
-
[1]
Min Cao, Xinyu Zhou, Ding Jiang, Bo Du, Mang Ye, and Min Zhang. 2025. Mul- tilingual Text-to-Image Person Retrieval via Bidirectional Relation Reasoning and Aligning.IEEE Transactions on Pattern Analysis and Machine Intelligence (2025)
2025
-
[3]
Ding Jiang and Mang Ye. 2023. Cross-modal implicit relation reasoning and aligning for text-to-image person retrieval. InProceedings of the IEEE/CVF con- ference on computer vision and pattern recognition. 2787–2797
2023
-
[4]
Zhengxian Wu, Chuanrui Zhang, Shen’Ao Jiang, Hangrui Xu, Zirui Liao, Luyuan Zhang, Li Huaqiu, Peng Jiao, and Haoqian Wang. 2026. Language-guided and motion-aware gait representation for generalizable recognition. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 40. 10871–10878
2026
-
[5]
Hao Li, Yuhao Wang, Wenning Hao, Pingping Zhang, Dong Wang, and Huchuan Lu. 2026. RAGTrack: Language-aware RGBT Tracking with Retrieval- Augmented Generation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 28179–28189
2026
-
[6]
Jiale Huang, Zixu Li, Zhiheng Fu, Zhiwei Chen, Qinlei Huang, and Yupeng Hu. 2026. RankVR: Low-Rank Structure Perception and Value Recalibration for Robust Composed Image Retrieval. InProceedings of the 2026 International Conference on Multimedia Retrieval. 269–278
2026
-
[7]
Zixu Li, Yupeng Hu, Zhiheng Fu, Zhiwei Chen, Weili Guan, and Liqiang Nie. 2026. R3: Composed Video Retrieval via Reasoning-Guided Recalling and Re-ranking. arXiv preprint arXiv:2606.01113(2026)
arXiv 2026
-
[8]
Mingyu Zhang, Zixu Li, Zhiwei Chen, Zhiheng Fu, Xiaowei Zhu, Jiajia Nie, Yinwei Wei, and Yupeng Hu. 2026. Hint: Composed image retrieval with dual- path compositional contextualized network. InICASSP. IEEE, 13002–13006
2026
-
[9]
Jinhe Bi, Aniri, Minglai Yang, Xingcheng Zhou, Wenke Huang, Sikuan Yan, Yujun Wang, Zixuan Cao, Michael Färber, Xun Xiao, Volker Tresp, and Yunpu Ma. 2026. EchoRL: Reinforcement Learning via Rollout Echoing. InForty-third International Conference on Machine Learning. https://openreview.net/forum? id=A6az59SGtF
2026
Show all 180 references
-
[10]
Guozhi Qiu, Zhiwei Chen, Zixu Li, Qinlei Huang, Zhiheng Fu, Xuemeng Song, and Yupeng Hu. 2026. Melt: Improve composed image retrieval via the modifi- cation frequentation-rarity balance network. InICASSP. IEEE, 13007–13011
2026
-
[11]
Xi Xiao, Xingjian Li, Yunbei Zhang, Cheng Han, Tianming Liu, Tianyang Wang, Runmin Jiang, Jihun Hamm, Xiao Wang, and Min Xu. 2026. Layer-Specific Prompt Fusion Discovery via Differentiable Search in Vision Foundation Models. arXiv preprint arXiv:2606.26379(2026)
2026 arXiv
-
[12]
Xi Xiao, Chenrui Ma, Yunbei Zhang, Chen Liu, Zhuxuanzi Wang, Yanshu Li, Lin Zhao, Guosheng Hu, Tianyang Wang, and Hao Xu. 2026. Not all directions matter: Towards structured and task-aware low-rank model adaptation. In Proceedings of the 64th Annual Meeting of the Association ...
2026
-
[13]
Lin Zhao, Xinru Jiang, Xi Xiao, Qihui Fan, Lei Lu, Yanzhi Wang, Xue Lin, Octavia Camps, Pu Zhao, and Jianyang Gu. 2026. Hieramp: Coarse-to-fine autoregressive amplification for generative dataset distillation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pat...
2026
-
[14]
Hao Li, Yuhao Wang, Xiantao Hu, Wenning Hao, Pingping Zhang, Dong Wang, and Huchuan Lu. 2026. Cadtrack: Learning contextual aggregation with de- formable alignment for robust rgbt tracking. InProceedings of the AAAI Confer- ence on Artificial Intelligence, Vol. 40. 6109–6117
2026
-
[15]
Weilin Wu, Shifan Yang, Qizhao Lin, XingHong Chen, Kunping Yang, Jing Wang, and Guannan Chen. 2025. A Novel Perspective on Low-Light Image Enhancement: Leveraging Artifact Regularization and Walsh-Hadamard Trans- form. InProceedings of the 33rd ACM International Conference on ...
2025
-
[16]
Yuan Sun, Xu Wang, Dezhong Peng, Zhenwen Ren, and Xiaobo Shen. 2023. Hierarchical hashing learning for image set classification.IEEE TIP32 (2023), 1732–1744
2023
-
[17]
Qianyun Yang, Zhiwei Chen, Yupeng Hu, Zixu Li, Zhiheng Fu, and Liqiang Nie. 2026. STABLE: Efficient Hybrid Nearest Neighbor Search via Magnitude- Uniformity and Cardinality-Robustness.IEEE TKDE(2026)
2026
-
[18]
Yang Qin, Yuan Sun, Dezhong Peng, Joey Tianyi Zhou, Xi Peng, and Peng Hu
-
[19]
Jiale Huang, Zixu Li, Zhiwei Chen, Zhiheng Fu, Chunxiao Wang, and Yupeng Hu. 2026. IMAGINE: Adaptive Schema-Imagery Enhanced Composition for Composed Video Retrieval. InProceedings of the 2026 International Conference on Multimedia Retrieval. 288–297
2026
-
[20]
Jinhe Bi, Danqi Yan, Yifan Wang, Wenke Huang, Haokun Chen, Guancheng Wan, Mang Ye, Xun Xiao, Hinrich Schuetze, Volker Tresp, and Yunpu Ma. 2026. The Geometry of Reasoning: Self-Evaluation via Layerwise Trajectory Evolution. In Forty-third International Conference on Machine Le...
2026
-
[21]
Honglin Yuan, Yuan Sun, Fei Zhou, Jing Wen, Shihua Yuan, Xiaojian You, and Zhenwen Ren. 2025. Prototype matching learning for incomplete multi-view clustering.IEEE TIP34 (2025), 828–841
2025
-
[22]
Zixu Li, Yupeng Hu, Zhiwei Chen, Qinlei Huang, Guozhi Qiu, Zhiheng Fu, and Meng Liu. 2026. ReTrack: Evidence-Driven Dual-Stream Directional Anchor Calibration Network for Composed Video Retrieval. InAAAI, Vol. 40. 23373– 23381
2026
-
[23]
Yupeng Hu, Zixu Li, Zhiwei Chen, Qinlei Huang, Zhiheng Fu, Mingzhu Xu, and Liqiang Nie. 2026. REFINE: Composed Video Retrieval via Shared and Differential Semantics Enhancement.ACM ToMM(2026)
2026
-
[24]
Yupeng Hu, Liqiang Nie, Meng Liu, Kun Wang, Yinglong Wang, and Xian- Sheng Hua. 2021. Coarse-to-fine semantic alignment for cross-modal moment localization.IEEE Transactions on Image Processing30 (2021), 5933–5943
2021
-
[25]
Yupeng Hu, Kun Wang, Meng Liu, Haoyu Tang, and Liqiang Nie. 2023. Semantic collaborative learning for cross-modal moment localization.ACM Transactions on Information Systems42, 2 (2023), 1–26
2023
-
[26]
Yupeng Hu, Meng Liu, Xiaobin Su, Zan Gao, and Liqiang Nie. 2021. Video moment localization via deep cross-modal hashing.IEEE Transactions on Image Processing30 (2021), 4667–4677
2021
-
[27]
Lixian Chen, Jingchao Wang, Zhaorong Dai, Hanqian Liu, Danxiang Ai, and Yang Shi. 2026. When task performance deceives: Task-geometry decoupling in learnable-curvature hyperbolic GNNs.Neural Networks(2026), 109172
2026
-
[28]
Wei Zhang, Yihang Wu, Shengkai Yu, Songhua Li, Qiang Li, and Qi Wang
-
[29]
Xi Xiao, Yunbei Zhang, Lin Zhao, Yiyang Liu, Xiaoying Liao, Zheda Mai, Xingjian Li, Xiao Wang, Hao Xu, Jihun Hamm, Xue Lin, Min Xu, Qifan Wang, Tianyang Wang, and Cheng Han. 2026. Prompt-based Adaptation in Large-scale Vision Models: A Survey.Transactions on Machine Learning R...
2026
-
[30]
Xingfeng Li, Yinghui Sun, Quansen Sun, Zhenwen Ren, and Yuan Sun. 2023. Cross-view graph matching guided anchor alignment for incomplete multi-view clustering.Information Fusion100 (2023), 101941
2023
-
[31]
Qianyun Yang, Peizhuo Lv, Yingjiu Li, Shengzhi Zhang, Yuxuan Chen, Zhiwei Chen, Zixu Li, and Yupeng Hu. 2026. ERASE: Bypassing Collaborative Detection of AI Counterfeit Via Comprehensive Artifacts Elimination.IEEE TDSC(March 2026), 1–18. doi:10.1109/TDSC.2026.3677794
2026
-
[32]
Jinhe Bi, Yujun Wang, Haokun Chen, Xun Xiao, Artur Hecker, Volker Tresp, and Yunpu Ma. 2025. LLaVA steering: Visual instruction tuning with 500x fewer parameters through modality linear representation-steering. InProceedings of the 63rd Annual Meeting of the Association for Co...
2025
-
[33]
Zixu Li, Zhiwei Chen, Zhiheng Fu, Wenbo Wang, Yupeng Hu, Weili Guan, and Liqiang Nie. 2026. OmniEgo-R 2: A Routed Reasoning Framework for the 1st Cross-Domain EgoCross Challenge at CVPR 2026.arXiv preprint arXiv:2605.24481(2026)
2026 arXiv
-
[34]
Xingfeng Li, Yuangang Pan, Yuan Sun, Quansen Sun, Yinghui Sun, Ivor W Tsang, and Zhenwen Ren. 2024. Incomplete multi-view clustering with paired and balanced dynamic anchor learning.IEEE TMM27 (2024), 1486–1497
2024
-
[35]
Xi Xiao, Yunbei Zhang, Xingjian Li, Tianyang Wang, Xiao Wang, Yuxiang Wei, Jihun Hamm, and Min Xu. 2025. Visual instance-aware prompt tuning. In Proceedings of the 33rd ACM International Conference on Multimedia. 2880–2889
2025
-
[36]
Liu Yu, Fenghui Tian, Ping Kuang, ZhiKun Feng, and Fan Zhou. 2025. Knowl- edge Graphs Acquisition via Forward-Reverse Relation Enhanced Contrastive Pretraining from Large-scale Models. In2025 IEEE International Conference on Multimedia and Expo (ICME). IEEE, 1–6
2025
-
[37]
Tao Huang, Rui Wang, Xiaofei Liu, Yi Qin, Li Duan, and Liping Jing. 2026. Detect- ing Misbehaviors of Large Vision-Language Models by Evidential Uncertainty Quantification.arXiv preprint arXiv:2602.05535(2026)
2026
-
[38]
Wei Zhang, Qiang Li, Yuan Yuan, and Qi Wang. 2024. Visual Consistency Enhancement for Multiview Stereo Reconstruction in Remote Sensing.IEEE Transactions on Geoscience and Remote Sensing62 (2024), 1–11. doi:10.1109/ TGRS.2024.3482697
2024
-
[39]
Yan Zhong, Chenxi Yang, Suyuan Zhao, and Tingting Jiang. 2025. Semi- supervised blind quality assessment with confidence-quantifiable pseudo-label learning for authentic images. InForty-second International Conference on Ma- chine Learning
2025
-
[40]
Guangtao Lyu, Qi Liu, Chenghao Xu, Jiexi Yan, Muli Yang, Xueting Li, Fen Fang, and Cheng Deng. 2026. Revealing and Enhancing Core Visual Regions: Harnessing Internal Attention Dynamics for Hallucination Mitigation in LVLMs. ACL Findings(2026)
2026
-
[41]
Zixu Li, Zhiwei Chen, Haokun Wen, Zhiheng Fu, Yupeng Hu, and Weili Guan
-
[42]
Jinhe Bi, Danqi Yan, Yifan Wang, Wenke Huang, Haokun Chen, Guancheng Wan, Mang Ye, Xun Xiao, Hin rich Schuetze, Volker Tresp, and Yunpu Ma. 2025. CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process.ArXiv abs/2505.13408 (2025). https://api.semanticscholar.org/C...
2025 arXiv
-
[43]
Yujun Wang, Jinhe Bi, Soren Pirk, Yunpu Ma, et al . 2026. Ascd: Attention- steerable contrastive decoding for reducing hallucination in mllm. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 40. 10306–10314
2026
-
[44]
Jinhe Bi, Yifan Wang, Danqi Yan, Xun Xiao, Artur Hecker, Volker Tresp, and Yunpu Ma. 2025. PRISM: Self-Pruning Intrinsic Selection Method for Training- Free Multimodal Data Selection.ArXivabs/2502.12119 (2025). https://api. semanticscholar.org/CorpusID:276421326
2025 arXiv
-
[45]
Zixu Li, Zhiheng Fu, Yupeng Hu, Zhiwei Chen, Haokun Wen, and Liqiang Nie
-
[46]
Canran Xiao, Tianxiang Xu, Siyuan Ma, Yiyang Jiang, Haoyu Gao, and Yuhan Wu. 2026. Reversible primitive–composition alignment for continual vision– language learning. InThe Fourteenth International Conference on Learning Rep- resentations
2026
-
[47]
Zhiheng Fu, Zixu Li, Zhiwei Chen, Chunxiao Wang, Xuemeng Song, Yupeng Hu, and Liqiang Nie. 2025. PAIR: Complementarity-guided Disentanglement for Composed Image Retrieval. InProceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE, 1–5
2025
-
[48]
Qinlei Huang, Zhiwei Chen, Zixu Li, Chunxiao Wang, Xuemeng Song, Yu- peng Hu, and Liqiang Nie. 2025. MEDIAN: Adaptive Intermediate-grained Aggregation Network for Composed Image Retrieval. InProceedings of the IEEE International Conference on Acoustics, Speech and Signal Proce...
2025
-
[49]
FineCIR: Explicit Parsing of Fine-Grained Modification Semantics for Composed Image Retrieval.https://arxiv.org/abs/2503.21309(2025)
2025 arXiv
-
[50]
Zhiwei Chen, Yupeng Hu, Zixu Li, Zhiheng Fu, Xuemeng Song, and Liqiang Nie. 2025. OFFSET: Segmentation-based Focus Shift Revision for Composed Image Retrieval. InACM MM. 6113–6122
2025
-
[51]
Guangtao Lyu, Xinyi Cheng, Chenghao Xu, Qi Liu, Muli Yang, Fen Fang, Huilin Chen, Jiexi Yan, Xu Yang, and Cheng Deng. 2025. Revealing Perception and Generation Dynamics in LVLMs: Mitigating Hallucinations via Validated Domi- nance Correction.arXiv preprint arXiv:2512.18813(2025)
2025
-
[52]
Liangsi Lu, Jingchao Wang, Zhaorong Dai, Hanqian Liu, and Yang Shi. 2026. Riemannian liquid spatio-temporal graph network. InProceedings of the ACM Web Conference 2026. 463–474
2026
-
[53]
Yuan Sun, Yang Qin, Yongxiang Li, Dezhong Peng, Xi Peng, and Peng Hu. 2024. Robust multi-view clustering with noisy correspondence.IEEE TKDE36, 12 (2024), 9150–9162
2024
-
[54]
Zhiheng Fu, Zixu Li, Zhiwei Chen, Fangxu Liu, Yupeng Hu, Weili Guan, and Liqiang Nie. 2026. EgoAction: Egocentric Action Composition with Reliability- Aware Temporal Fusion for the EPIC-KITCHENS Action Detection Challenge at CVPR 2026.arXiv preprint arXiv:2605.24496(2026)
2026 arXiv
-
[55]
Zhiwei Chen, Yupeng Hu, Zixu Li, Zhiheng Fu, Guozhi Qiu, Weili Guan, and Liqiang Nie. 2026. EgoAdapt: A Multi-Scene Egocentric Adaptation Method for CVPR 2026 HD-EPIC VQA Challenge.arXiv preprint arXiv:2605.24500(2026)
2026 arXiv
-
[56]
Zhiming Lin, Canran Xiao, and Kai Zhao. 2026. Beyond More Context: Retrieval Diversity Boosts Multi-Turn Intent Understanding. InProceedings of the ACM Web Conference 2026. 2320–2329
2026
-
[57]
Min Cao, Yang Bai, Ziyin Zeng, Mang Ye, and Min Zhang. 2024. An empirical study of clip for text-based person search. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 465–473
2024
-
[58]
Qicheng Zhao, Qi Sun, and Zheyu Yan. 2026. Seeing the End at Step Zero: Accelerating Diffusion MLLMs via MLP Sparsity-Aware Truncation.arXiv preprint arXiv:2607.14557(2026)
2026 arXiv
-
[60]
Guangtao Lyu, Chenghao Xu, Jiexi Yan, Muli Yang, and Cheng Deng. 2025. To- wards Unified Human Motion-Language Understanding via Sparse Interpretable Characterization. InThe Thirteenth International Conference on Learning Repre- sentations(ICLR)
2025
-
[61]
Xi Xiao, Chen Liu, Chih-Ting Liao, Yunbei Zhang, Qizhen Lan, Yuxiang Wei, Lin Zhao, Janet Wang, Jianyang Gu, Muchao Ye, et al. 2026. Staying VIGILant: Mitigating Visual Laziness via Counterfactual Visual Alignment in MLLMs. arXiv preprint arXiv:2606.26387(2026)
2026 arXiv
-
[62]
Kaifang Long, Lianbo Ma, Jiaqi Liu, Liming Liu, and Guoyang Xie. 2026. To- wards an Incremental Unified Multimodal Anomaly Detection: Augmenting Multimodal Denoising From an Information Bottleneck Perspective. InProceed- ings of the IEEE/CVF Conference on Computer Vision and P...
2026
-
[63]
Haitian Li, Yanghao Zhou, Heyan Huang, Liangji Chen, YiMing Cheng, Xu Liu, Dian Jin, Jiajun Xu, Jingyun Liao, Tian Lan, et al . 2026. MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation.arXiv preprint arXiv:2605.28035(2026)
2026 arXiv
-
[64]
Guangtao Lyu, Chenghao Xu, Qi Liu, Jiexi Yan, Muli Yang, Fen Fang, and Cheng Deng. 2025. Tempo as the Stable Cue: Hierarchical Mixture of Tempo and Beat Experts for Music to 3D Dance Generation.arXiv preprint arXiv:2512.18804 (2025)
2025
-
[65]
Zhengxian Wu, Chuanrui Zhang, Hangrui Xu, Peng Jiao, and Haoqian Wang
-
[66]
In2025 IEEE International Conference on Multimedia and Expo (ICME)
DAGait: Generalized skeleton-guided data alignment for gait recognition. In2025 IEEE International Conference on Multimedia and Expo (ICME). IEEE, 1–6
-
[67]
Kaifang Long, Guoyang Xie, Lianbo Ma, Qing Li, Min Huang, Jianhui Lv, and Zhichao Lu. 2025. Enhancing Multimodal Learning via Hierarchical Fusion Architecture Search With Inconsistency Mitigation.IEEE Transactions on Image Processing(2025)
2025
-
[68]
Kaifang Long, Guoyang Xie, Lianbo Ma, Jiaqi Liu, and Zhichao Lu. 2025. Re- visiting multimodal fusion for 3D anomaly detection from an architectural perspective. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 12273–12281
2025
-
[69]
Guangtao Lyu, Xinyi Cheng, Qi Liu, Chenghao Xu, Jiexi Yan, Muli Yang, Fen Fang, and Cheng Deng. 2026. Towards Interpretable Hallucination Analysis and Mitigation in LVLMs via Contrastive Neuron Steering.arXiv preprint arXiv:2602.00621(2026)
2026
-
[70]
Guangtao Lyu, Xinyi Cheng, Qi Liu, Chenghao Xu, Jiexi Yan, Muli Yang, Fen Fang, and Cheng Deng. [n. d.]. COME: Advancing Representation Learning and Generative Modeling for High-Quality Text-to-Motion Generation. ([n. d.])
-
[71]
Yang-Hao Zhou, Haitian Li, Rexar Lin, Heyan Huang, Jinxing Zhou, Changsen Yuan, Tian Lan, Ziqin Zhou, Yudong Li, Jiajun Xu, et al. 2026. MTAVG-Bench: A Comprehensive Benchmark for Evaluating Multi-Talker Dialogue-Centric Audio-Video Generation.arXiv preprint arXiv:2602.00607(2026)
2026 arXiv
-
[72]
Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. 2021. Align before fuse: Vision and lan- guage representation learning with momentum distillation.Advances in neural information processing systems34 (2021), 9694–9705
2021
-
[73]
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann. 2020. Shortcut learning in deep neural networks.Nature Machine Intelligence2, 11 (2020), 665–673
2020
-
[74]
Yuhao Chen, Guoqing Zhang, Yujiang Lu, Zhenxing Wang, and Yuhui Zheng
-
[75]
Andra Acsintoae, Andrei Florescu, Mariana-Iuliana Georgescu, Tudor Mare, Paul Sumedrea, Radu Tudor Ionescu, Fahad Shahbaz Khan, and Mubarak Shah. 2022. Ubnormal: New benchmark for supervised open-set video anomaly detection. In Proceedings of the IEEE/CVF conference on compute...
2022
-
[76]
Yutian Lin, Liang Zheng, Zhedong Zheng, Yu Wu, Zhilan Hu, Chenggang Yan, and Yi Yang. 2019. Improving person re-identification by attribute and identity learning.Pattern recognition95 (2019), 151–161
2019
-
[77]
Haigang Deng, Qingyang Yang, Chengwei Li, Hanzhong Liang, and Chuanxu Wang. 2025. Video anomaly detection via pseudo-anomaly generation and multi- grained feature learning.Journal of Electronic Imaging34, 1 (2025), 013044– 013044
2025
-
[78]
Yan Zhong, Xinping Zhao, Guangzhi Zhao, Bohua Chen, Fei Hao, Ruoyu Zhao, Jiaqi He, Lei Shi, and Li Zhang. 2025. Ctd-inpainting: Towards the coherence of text-driven inpainting with blended diffusion.Information Fusion122 (2025), 103163
2025
-
[79]
Yan Zhong, Ruoyu Zhao, Chao Wang, Jiaqi He, Qinghai Guo, Jianguo Zhang, Zhichao Lu, and Luziwei Leng. 2026. Dyn-SSM: Towards the Efficient Long Se- quence Learning via Bio-interpretable Dynamics in Spiking State Space Models. IEEE Transactions on Cognitive and Developmental Sy...
2026
-
[80]
ZhengXian Wu, Hangrui Xu, Kai Shi, Zhuohong Chen, Yunyao Yu, Chuanrui Zhang, Zirui Liao, Jun Yang, Zhenyu Yang, Haonan Lu, et al. 2026. ProMSA: Progressive Multimodal Search Agents for Knowledge-Based Visual Question Answering.arXiv preprint arXiv:2606.27974(2026)
2026 arXiv
-
[81]
Xueming Qian, Dan Lu, Yaxiong Wang, Li Zhu, Yuan Yan Tang, and Meng Wang
-
[82]
Zhiwei Chen, Yupeng Hu, Zixu Li, Zhiheng Fu, Haokun Wen, and Weili Guan
-
[83]
Hari Lee. 2025. Knowledge-Guided Textual Reasoning for Explainable Video Anomaly Detection via LLMs.arXiv preprint arXiv:2511.07429(2025)
2025
-
[84]
Chunlei Meng, Guanhong Huang, Rong Fu, Runmin Jian, Zhongxue Gan, and Chun Ouyang. 2026. Clcr: Cross-level semantic collaborative representation for multimodal learning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 1606–1615. LightAIR: L...
2026
-
[85]
Rong Fu, Zijian Zhang, Haiyun Wei, Jiekai Wu, Kun Liu, Xianda Li, Haoyu Zhao, Yang Li, Yongtai Liu, Ziming Wang, et al . 2026. LiveGraph: Active- Structure Neural Re-ranking for Exercise Recommendation.arXiv preprint arXiv:2602.17036(2026)
2026 arXiv
-
[86]
Qicheng Zhao, Yu Li, Qi Sun, and Zheyu Yan. 2026. ResilPhase: Plug-and- Play Phase Mapping and Noise-Resilient Macro-Trajectory Extrapolation for Diffusion Acceleration.arXiv preprint arXiv:2606.26769(2026)
2026 arXiv
-
[87]
Yunyao Zhang, Yihao Ai, Zuocheng Ying, Qirui Mi, Junqing Yu, Wei Yang, and Zikai Song. 2026. Coupling Macro Dynamics and Micro States for Long-Horizon Social Simulation.arXiv preprint arXiv:2604.05516(2026)
2026 arXiv
-
[88]
Zixu Li, Yupeng Hu, Zhiwei Chen, Shiqi Zhang, Qinlei Huang, Zhiheng Fu, and Yinwei Wei. 2026. HABIT: Chrono-Synergia Robust Progressive Learning Framework for Composed Image Retrieval. InAAAI, Vol. 40. 6762–6770
2026
-
[89]
Zixu Li, Yupeng Hu, Zhiheng Fu, Zhiwei Chen, Yongqi Li, and Liqiang Nie
-
[90]
InACM MM
HUD: Hierarchical Uncertainty-Aware Disambiguation Network for Composed Video Retrieval. InACM MM. 6143–6152
-
[91]
Zixu Li, Yupeng Hu, Zhiwei Chen, Haokun Wen, Xuemeng Song, and Liqiang Nie. 2026. COMBINER: Composed Image Retrieval Guided by Attribute-based Neighbor Relations.IEEE TIP(2026)
2026
-
[92]
Zijian Zhang, Rong Fu, Yangfan He, Xinze Shen, Yanlong Wang, Xiaojing Du, Haochen You, Keyan Jin, Jiazhao Shi, and Simon Fong. 2026. FinSentLLM: Multi-LLM and structured semantic signals for enhanced financial sentiment forecasting. InICASSP 2026-2026 IEEE International Confer...
2026
-
[93]
Wenjie Zhu, Yabin Zhang, Xin Jin, Wenjun Zeng, and Lei Zhang. 2026. Ants: Adaptive negative textual space shaping for ood detection via test-time mllm understanding and reasoning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 20–30
2026
-
[94]
Xinglang Zhang, Yunyao Zhang, ZeLiang Chen, Junqing Yu, Wei Yang, and Zikai Song. 2026. Logical Phase Transitions: Understanding Collapse in LLM Logical Reasoning.arXiv preprint arXiv:2601.02902(2026)
2026 arXiv
-
[95]
Yan Zhong, Xingyu Wu, Xinping Zhao, Li Zhang, Xinyuan Song, Lei Shi, and Bingbing Jiang. 2026. Semi-supervised multi-label feature selection with consis- tent sparse graph learning.Neural Networks(2026), 109265
2026
-
[96]
Hangrui Xu, Zhengxian Wu, Chuanrui Zhang, Zhuohong Chen, Zhifang Liu, Peng Jiao, and Haoqian Wang. 2026. Psgait: Gait recognition using parsing skeleton. InICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 10427–10431
2026
-
[97]
Yan Zhong, Xingyu Wu, Li Zhang, Chenxi Yang, and Tingting Jiang. 2024. Causal-IQA: Towards the Generalization of Image Quality Assessment Based on Causal Inference.. InICML
2024
-
[98]
InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Tema: Anchor the image, follow the text for multi-modification composed image retrieval. InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 24421–24442
-
[99]
Rong Fu, Yemin Wang, Tianxiang Xu, Yongtai Liu, Weizhi Tang, Wangyu Wu, Xiaowen Ma, and Simon Fong. 2026. S-Path-RAG: Semantic-Aware Shortest-Path Retrieval Augmented Generation for Multi-Hop Knowledge Graph Question Answering. InProceedings of the ACM Web Conference 2026. 4057–4068
2026
-
[100]
Tailong Luo, Hao Li, Rong Fu, Xinyue Jiang, Huaxuan Ding, Yiduo Zhang, Zilin Zhao, Simon Fong, Guangyin Jin, and Jianyuan Ni. 2026. Multipress: A multi-agent framework for interpretable multimodal news classification.arXiv preprint arXiv:2604.03586(2026)
2026 arXiv
-
[101]
Zixu Li, Yupeng Hu, Zhiwei Chen, Mingyu Zhang, Zhiheng Fu, and Liqiang Nie
-
[102]
Zikai Song, Ying Tang, Run Luo, Lintao Ma, Junqing Yu, Yi-Ping Phoebe Chen, and Wei Yang. 2024. Autogenic language embedding for coherent point tracking. InProceedings of the 32nd ACM International Conference on Multimedia. 2021– 2030
2024
-
[103]
Zhiheng Fu, Yupeng Hu, Qianyun Yang, Shiqi Zhang, Zhiwei Chen, and Zixu Li. 2026. Air-know: Arbiter-calibrated knowledge-internalizing robust network for composed image retrieval. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2658–2670
2026
-
[104]
Wenbing Li, Hang Zhou, Junqing Yu, Zikai Song, and Wei Yang. 2024. Coupled mamba: Enhanced multimodal fusion with coupled state space model.Advances in Neural Information Processing Systems37 (2024), 59808–59832
2024
-
[105]
Hao Ju, Hu Zhang, and Zhedong Zheng. 2025. AnomalyLMM: Bridging Genera- tive Knowledge and Discriminative Retrieval for Text-Based Person Anomaly Search.arXiv preprint arXiv:2509.04376(2025)
2025 arXiv
-
[106]
Yang Shi, Jingchao Wang, Liangsi Lu, Haiying Huang, Sutong Xiao, Zhaorong Dai, Ying Wang, and Boyan Xu. 2026. Enhancing Robustness of Constant Curvature Graph Convolutional Network with Lipschitz Regularization.ACM Transactions on Knowledge Discovery from Data(2026)
2026
-
[107]
Zikai Song, Junqing Yu, Yi-Ping Phoebe Chen, and Wei Yang. 2022. Transformer tracking with cyclic shifting window attention. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 8791–8800
2022
-
[108]
Zikai Song, Run Luo, Junqing Yu, Yi-Ping Phoebe Chen, and Wei Yang. 2023. Compact transformer tracker with correlative masked modeling. InProceedings of the AAAI conference on artificial intelligence, Vol. 37. 2321–2329
2023
-
[109]
Yan Zhong, Xinping Zhao, Li Zhang, Xinyuan Song, and Tingting Jiang. 2025. Adaptive Prompt Learning for Blind Image Quality Assessment with Multi- modal Mixed-datasets Training. InProceedings of the 33rd ACM International Conference on Multimedia. 7453–7462
2025
-
[110]
Yangliu Hu, Zikai Song, Na Feng, Yawei Luo, Junqing Yu, Yi-Ping Phoebe Chen, and Wei Yang. 2025. Sf2t: Self-supervised fragment finetuning of video-llms for fine-grained understanding. InProceedings of the Computer Vision and Pattern Recognition Conference. 29108–29117
2025
-
[111]
InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Conesep: Cone-based robust noise-unlearning compositional network for composed image retrieval. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 16897–16909
-
[112]
Damien Teney, Ehsan Abbasnejad, Simon Lucey, and Anton Van den Hengel
-
[113]
Liu Yu, Ludie Guo, Ping Kuang, and Fan Zhou. 2025. Bridging the fairness gap: Enhancing pre-trained models with llm-generated sentences. InICASSP 2025- 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1–5
2025
-
[114]
Wenjie Zhu, Yabin Zhang, Liang Xu, Xin Jin, Wenjun Zeng, and Lei Zhang. 2026. Dual Distribution Estimation for Zero-shot Noisy Test-Time Adaptation with VLMs.arXiv preprint arXiv:2606.25758(2026)
2026 arXiv
-
[115]
Liu Yu, Fenghui Tian, Ping Kuang, and Fan Zhou. 2025. Amplifying common- sense knowledge via bi-directional relation integrated graph-based contrastive pre-training from large language models.Information Processing & Management 62, 3 (2025), 104068
2025
-
[116]
Zikai Song, Run Luo, Lintao Ma, Ying Tang, Yi-Ping Phoebe Chen, Junqing Yu, and Wei Yang. 2025. Temporal Coherent Object Flow for Multi-Object Tracking. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 6978–6986
2025
-
[117]
Liu Yu, Can Chen, Ping Kuang, Zhikun Feng, Fan Zhou, and Gillian Dobbie
-
[118]
Dismantling Pathological Shortcuts: A Causal Framework for Faithful LVLM Decoding.arXiv preprint arXiv:2606.27596(2026)
2026 arXiv
-
[119]
Zixu Li, Yupeng Hu, Zhiwei Chen, Zhiheng Fu, Xiaowei Zhu, Weili Guan, and Liqiang Nie. 2026. TempRet: Temporal Enhancement and Two-Stage Reranking for CVPR 2026 EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge.arXiv preprint arXiv:2605.24470(2026)
2026 arXiv
-
[120]
Zhiwei Chen, Yupeng Hu, Zhiheng Fu, Zixu Li, Jiale Huang, Qinlei Huang, and Yinwei Wei. 2026. INTENT: Invariance and Discrimination-aware Noise Mitigation for Robust Composed Image Retrieval. InAAAI, Vol. 40. 20463–20471
2026
-
[121]
Siming Fu, Sijun Dong, and Xiaoliang Meng. 2025. Disentangling Content from Style to Overcome Shortcut Learning: A Hybrid Generative-Discriminative Learning Framework.arXiv preprint arXiv:2509.11598(2025)
2025
-
[122]
Yu Yang, Eric Gan, Gintare Karolina Dziugaite, and Baharan Mirzasoleiman
-
[123]
Ziheng Chi, Yifan Hou, Chenxi Pang, Shaobo Cui, Mubashara Akhtar, and Mrinmaya Sachan. 2025. Chimera: Diagnosing Shortcut Learning in Visual- Language Understanding.arXiv preprint arXiv:2509.22437(2025)
2025
-
[124]
Yunyao Zhang, Zikai Song, Hang Zhou, Wenfeng Ren, Yi-Ping Phoebe Chen, Junqing Yu, and Wei Yang. 2025. 𝐺𝐴−𝑆 3: Comprehensive Social Network Simulation with Group Agents. InFindings of the Association for Computational Linguistics: ACL 2025. 8950–8970
2025
-
[125]
InProceedings of the IEEE/CVF conference on computer vision and pattern recognition
Evading the simplicity bias: Training a diverse set of models discov- ers solutions with superior ood generalization. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 16761–16772
-
[126]
Liu Yu, Yuzhou Mao, Jin Wu, and Fan Zhou. 2023. Mixup-based unified frame- work to overcome gender bias resurgence. InProceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval. 1755–1759. MM ’26, November 10–14, 2026, Rio d...
2023
-
[127]
Shunxin Wang, Raymond Veldhuis, and Nicola Strisciuglio. 2025. Do ImageNet- trained models learn shortcuts? The impact of frequency shortcuts on general- ization. InProceedings of the Computer Vision and Pattern Recognition Conference. 25198–25207
2025
-
[128]
Yihe Deng, Yu Yang, Baharan Mirzasoleiman, and Quanquan Gu. 2023. Robust learning with progressive data expansion against spurious correlation.Advances in neural information processing systems36 (2023), 1390–1402
2023
-
[129]
Yangyi Fang, Jiaye Lin, Xiaoliang Fu, Cong Qin, Haolin Shi, and Chang Liu. 2026. Proximity-based multi-turn optimization: Practical credit assignment for llm agent training. InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026). 285–307
2026
-
[130]
Xiaoliang Fu, Jiaye Lin, Yangyi Fang, Binbin Zheng, Chaowen Hu, Zekai Shao, Cong Qin, Lu Pan, Ke Zeng, and Xunliang Cai. 2026. Maspo: Unifying gradient utilization, probability mass, and signal reliability for robust and sample-efficient llm reasoning. InProceedings of the 64t...
2026
-
[131]
Liu Yu, Zhonghao Chen, Ping Kuang, Zhikun Feng, Fan Zhou, Lan Wang, and Gillian Dobbie. 2026. Causally-Grounded Dual-Path Attention Intervention for Object Hallucination Mitigation in LVLMs. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 40. 36021–36029
2026
-
[132]
Liu Yu, Jiajun Sun, Ping Kuang, Rui Zhou, Fan Zhou, and Zhikun Feng. 2025. Bimodal Debiasing for Text-to-Image Diffusion: Adaptive Guidance in Textual and Visual Spaces. InProceedings of the 33rd ACM International Conference on Multimedia. 11249–11258
2025
-
[133]
Yunyao Zhang, Xinglang Zhang, Junxi Sheng, Wenbing Li, Junqing Yu, Yi- Ping Phoebe Chen, Wei Yang, and Zikai Song. 2026. Semantic-Aware Logical Reasoning via a Semiotic Framework. arXiv:2509.24765 [cs.AI]
2026 arXiv
-
[134]
Jicheol Park, Dongwon Kim, Boseung Jeong, and Suha Kwak. 2024. Plot: Text- based person search with part slot attention for corresponding part discovery. InEuropean Conference on Computer Vision. Springer, 474–490
2024
-
[135]
Fan Zhou, Yuzhou Mao, Liu Yu, Yi Yang, and Ting Zhong. 2023. Causal-debias: Unifying debiasing in pretrained language models and fine-tuning via causal invariant learning. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long...
2023
-
[136]
Steven Bird. 2006. NLTK: the natural language toolkit. InProceedings of the COLING/ACL 2006 interactive presentation sessions. 69–72
2006
-
[137]
Matthew Honnibal. 2017. spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing.(No Title) (2017)
2017
-
[138]
Jiaye Lin, Yifu Guo, Yuzhen Han, Sen Hu, Ziyi Ni, Licheng Wang, Ming- guang Chen, Hongzhang Liu, Ronghao Chen, Yangfan He, Daxin Jiang, Binx- ing Jiao, Chen Hu, and Huacan Wang. 2025. SE-Agent: Self-Evolution Trajectory Optimization in Multi-Step Reasoning with LLM-Based Agent...
2025
-
[139]
Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, and Yoav Goldberg
-
[140]
2008.Optimization algo- rithms on matrix manifolds
P-A Absil, Robert Mahony, and Rodolphe Sepulchre. 2008.Optimization algo- rithms on matrix manifolds. Princeton University Press
2008
-
[141]
Jiaye Lin, Mengdi Li, Xufeng Zhao, Wenhao Lu, Peilin Zhao, Stefan Wermter, and Di Wang. 2026. Curriculum-RLAIF: Curriculum Alignment with Reinforcement Learning from AI Feedback. arXiv:2505.20075 [cs.AI] https://arxiv.org/abs/2505. 20075
2026 arXiv
-
[142]
Shuang Li, Tong Xiao, Hongsheng Li, Bolei Zhou, Dayu Yue, and Xiaogang Wang
-
[143]
Rong Fu and Simon Fong. 2025. Adaptive Multi-Backbone Fusion for UAV- Centric Cross-View Geo-Localization with Partial Street–Satellite Matching. In Proceedings of the 3rd International Workshop on UA Vs in Multimedia: Capturing the World from a New Perspective. 31–36
2025
-
[144]
Rong Fu, Yibo Meng, Jia Yee Tan, Jiaxuan Lu, Rui Lu, Jiekai Wu, Zhaolu Kang, and Simon Fong. 2026. CityGuard: Graph-Aware Private Descriptors for Bias- Resilient Identity Search Across Urban Cameras.arXiv preprint arXiv:2602.18047 (2026)
2026 arXiv
-
[145]
Jialong Zuo, Hanyu Zhou, Ying Nie, Feng Zhang, Tianyu Guo, Nong Sang, Yunhe Wang, and Changxin Gao. 2024. Ufinebench: Towards text-based person retrieval with ultra-fine granularity. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 22010–22019
2024
-
[146]
Research Team. 2024. TIPS: A Text-Image Pairs Synthesis Framework for Robust Text-based Person Retrieval.OpenReview(2024)
2024
-
[147]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sand- hini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al
-
[148]
Hang Yu, Jiahao Wen, and Zhedong Zheng. 2025. CAMeL: Cross-modality Adaptive Meta-Learning for Text-based Person Retrieval.IEEE Transactions on Information Forensics and Security(2025)
2025
-
[149]
Shuyu Yang, Yinan Zhou, Zhedong Zheng, Yaxiong Wang, Li Zhu, and Yujiao Wu
-
[150]
Shuyu Yang, Yaxiong Wang, Li Zhu, and Zhedong Zheng. 2025. Beyond walking: A large-scale image-text benchmark for text-based person anomaly search. In ICCV. 11720–11730
2025
-
[151]
Francesco Locatello, Stefan Bauer, Mario Lucic, Gunnar Raetsch, Sylvain Gelly, Bernhard Schölkopf, and Olivier Bachem. 2019. Challenging common assump- tions in the unsupervised learning of disentangled representations. Ininterna- tional conference on machine learning. PMLR, 4114–4124
2019
-
[152]
Chenyang Gao, Guanyu Cai, Xinyang Jiang, Feng Zheng, Jun Zhang, Yifei Gong, Pai Peng, Xiaowei Guo, and Xing Sun. 2021. Contextual non-local alignment over full-scale representation for text-based person search.arXiv preprint arXiv:2101.03036(2021)
2021 arXiv
-
[153]
Yucheng Chen, Rui Huang, Hong Chang, Chuanqi Tan, Tao Xue, and Bingpeng Ma. 2021. Cross-modal knowledge adaptation for language-based person search. IEEE Transactions on Image Processing30 (2021), 4057–4069
2021
-
[154]
Yushuang Wu, Zizheng Yan, Xiaoguang Han, Guanbin Li, Changqing Zou, and Shuguang Cui. 2021. Lapscore: language-guided person search via color reasoning. InProceedings of the IEEE/CVF International Conference on Computer Vision. 1624–1633
2021
-
[156]
Zhiyin Shao, Xinyu Zhang, Meng Fang, Zhifeng Lin, Jian Wang, and Changxing Ding. 2022. Learning granularity-unified representations for text-to-image person re-identification. InProceedings of the 30th acm international conference on multimedia. 5566–5574
2022
-
[157]
Person search with natural language description. InCVPR. 1970–1979
1970
-
[159]
Aichun Zhu, Zijie Wang, Yifeng Li, Xili Wan, Jing Jin, Tian Wang, Fangqiang Hu, and Gang Hua. 2021. Dssl: Deep surroundings-person separation learning for text-based person retrieval. InProceedings of the 29th ACM international conference on multimedia. 209–217
2021
-
[161]
Yan Zeng, Xinsong Zhang, Hang Li, Jiawei Wang, Jipeng Zhang, and Wangchun- shu Zhou. 2023. X 2-VLM: All-in-One Pre-Trained Model for Vision-Language Tasks.IEEE transactions on pattern analysis and machine intelligence46, 5 (2023), 3156–3168
2023
-
[162]
Dixuan Lin, Yi-Xing Peng, Jingke Meng, and Wei-Shi Zheng. 2024. Cross-modal adaptive dual association for text-to-image person retrieval.IEEE Transactions on Multimedia26 (2024), 6609–6620
2024
-
[163]
Zhiwei Zhao, Bin Liu, Yan Lu, Qi Chu, and Nenghai Yu. 2024. Unifying multi- modal uncertainty modeling and semantic alignment for text-to-image person re-identification. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 7534–7542
2024
-
[164]
Yang Bai, Min Cao, Daming Gao, Ziqiang Cao, Chen Chen, Zhenfeng Fan, Liqiang Nie, and Min Zhang. 2023. Rasa: Relation and sensitivity aware representation learning for text-based person search.arXiv preprint arXiv:2305.13653(2023)
2023 arXiv
-
[165]
Shuanglin Yan, Jun Liu, Neng Dong, Liyan Zhang, and Jinhui Tang. 2024. Proto- typical prompting for text-to-image person re-identification. InProceedings of the 32nd ACM International Conference on Multimedia. 2331–2340
2024
-
[166]
InProceedings of the 31st ACM international conference on multimedia
Towards unified text-based person retrieval: A large-scale multi-attribute and language search benchmark. InProceedings of the 31st ACM international conference on multimedia. 4492–4501
-
[168]
Jintao Sun, Hao Fei, Gangyi Ding, and Zhedong Zheng. 2025. From data deluge to data curation: A filtering-wora paradigm for efficient text-based person search. InProceedings of the ACM on Web Conference 2025. 2341–2351
2025
-
[172]
Zefeng Ding, Changxing Ding, Zhiyin Shao, and Dacheng Tao. 2021. Semanti- cally self-aligned network for text-to-image part-aware person re-identification. arXiv preprint arXiv:2107.12666(2021)
2021 arXiv
-
[174]
Ammarah Farooq, Muhammad Awais, Josef Kittler, and Syed Safwan Khalid
-
[175]
InProceedings of the AAAI conference on artificial intelligence, Vol
Axm-net: Implicit cross-modal feature alignment for person re- identification. InProceedings of the AAAI conference on artificial intelligence, Vol. 36. 4477–4485
-
[176]
Shiping Li, Min Cao, and Min Zhang. 2022. Learning semantic-aligned feature representation for text-based person search. InICASSP 2022-2022 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2724–2728
2022
-
[177]
Shuanglin Yan, Hao Tang, Liyan Zhang, and Jinhui Tang. 2023. Image-specific information suppression and implicit local alignment for text-based person search.IEEE transactions on neural networks and learning systems(2023)
2023
-
[178]
Shuanglin Yan, Neng Dong, Liyan Zhang, and Jinhui Tang. 2023. Clip-driven fine-grained text-image person re-identification.IEEE Transactions on Image Processing32 (2023), 6032–6046
2023
-
[179]
Takuro Fujii and Shuhei Tarashima. 2023. Bilma: Bidirectional local-matching for text-based person re-identification. InProceedings of the IEEE/CVF International Conference on Computer Vision. 2786–2790
2023
-
[182]
Di Wang, Feng Yan, Yifeng Wang, Lin Zhao, Xiao Liang, Haodi Zhong, and Ronghua Zhang. 2024. Fine-grained semantics-aware representation learning for text-based person retrieval. InProceedings of the 2024 International Conference on Multimedia Retrieval. 92–100
2024
-
[184]
Yang Qin, Yingke Chen, Dezhong Peng, Xi Peng, Joey Tianyi Zhou, and Peng Hu
-
[185]
LightAIR: Lightweight Action Inversion and Riemannian Rectification for Text-based Person Anomaly Search
Noisy-correspondence learning for text-to-image person re-identification. InCVPR. 27197–27206. LightAIR: Lightweight Action Inversion and Riemannian Rectification for Text-based Person Anomaly Search MM ’26, November 10–14, 2026, Rio de Janeiro, Brazil Supplementary Material f...
2026
-
[2017]
Image re-ranking based on topic diversity.IEEE Transactions on Image Processing26, 8 (2017), 3734–3747
2017
-
[2020]
InProceedings of the 58th annual meeting of the association for computational linguistics
Null it out: Guarding protected attributes by iterative nullspace projection. InProceedings of the 58th annual meeting of the association for computational linguistics. 7237–7256
-
[2021]
InInternational conference on machine learning
Learning transferable visual models from natural language supervision. InInternational conference on machine learning. PMLR, 8748–8763
-
[2022]
TIPCB: A simple but effective part-based convolutional baseline for text-based person search.Neurocomputing494 (2022), 171–181
2022
-
[2023]
Cross-modal active complementary learning with self-refining correspon- dence.NeurIPS36 (2023), 24829–24840
2023
-
[2024]
InInternational conference on artificial intelligence and statistics
Identifying spurious biases early in training through the lens of simplicity bias. InInternational conference on artificial intelligence and statistics. PMLR, 2953–2961
-
[2025]
InProceedings of the AAAI Conference on Artificial Intel- ligence
ENCODER: Entity Mining and Modification Relation Binding for Com- posed Image Retrieval. InProceedings of the AAAI Conference on Artificial Intel- ligence. MM ’26, November 10–14, 2026, Rio de Janeiro, Brazil Yulun Zhang, Zixu Li, Zhiwei Chen, Zhiheng Fu, Wenbo Wang, Zihang Qi...
2026
-
[2026]
doi:10.1109/TGRS.2026.3710063
GPR-MVS: Global Propagation Regularization for Large Scale Multi- view Stereo.IEEE Transactions on Geoscience and Remote Sensing(2026), 1–1. doi:10.1109/TGRS.2026.3710063
2026
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.