Pith. sign in

REVIEW 3 major objections 6 minor 180 references

LightAIR: Lightweight Action Inversion and Riemannian Rectification for Text-based Person Anomaly Search

T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read LightAIR claims strict action-appearance decoupling via text anchors, null-space projection, and Riemannian gradient rectification, setting state-of-the-art scores on PAB and four TIPR benchmarks.

desk verdict Solid empirical gains and a sensible new combination for TPAS, but the 'strict forward decoupling' claim is overstated: the projection is rank-1, the paper's own appendix shows residual leakage, and the missing multi-basis ablation leaves the core assumption untested. read the letter →

arxiv 2608.09152 v1 pith:Z6UMUAEB submitted 2026-08-10 cs.CV

classification cs.CV
keywords Text-basedPersonAnomalySearchText-to-ImageRetrievalaction-appearancedecouplingnull-spaceprojectionRiemanniangradientshortcutlearningsemanticcodebookhardnegativesamples
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Text-based Person Anomaly Search (TPAS) asks a model to retrieve a pedestrian by matching both static appearance and a fine-grained abnormal action. The paper tries to establish that this can be done without fragile external pose estimators: a lightweight action inversion operator uses a frozen codebook of action words to reconstruct a pure action vector from the visual feature, a null-space projection forces the appearance vector onto the orthogonal complement of that action direction, and a Riemannian gradient rectification confines backpropagation to the tangent space of the action semantic manifold so hard negatives cannot exploit appearance shortcuts. If the claim holds, the resulting LightAIR network is the top scorer on the PAB anomaly benchmark (84.73% R@1 and 91.93% mAP at 0.1M training pairs, improving to 85.49% and 92.20% at 1M) and also improves four standard text-to-image person retrieval benchmarks, including a large gain on the ultra-fine-grained UFine3C setting.

What carries the argument

The load-bearing object is the orthogonal projection operator $\Pi_{\text{act}} = z_{\text{act}} z_{\text{act}}^{\top}/(\lVert z_{\text{act}} \rVert^2+\epsilon)$, built from the single reconstructed action vector. It does double duty: in the forward pass it defines the null space $(I-\Pi_{\text{act}})v$ that yields the appearance feature, and in the backward pass it defines the tangent space onto which the Euclidean gradient is projected to obtain the Riemannian gradient. The action vector itself is produced by the Action Inversion Operator, which estimates sparse Top-K coefficients over a frozen codebook of action-word embeddings and reconstructs $z_{\text{act}} = D^{\top} \boldsymbol{\alpha} / \lVert D^{\top} \boldsymbol{\alpha} \rVert_2$. An entropy-guided weight $\omega = \exp(-H/\tau)$ interpolates between the Euclidean and Riemannian gradients during training, so early optimization explores freely and later updates respect the decoupling constraint.

What would settle it

Compare LightAIR against a variant whose action subspace contains several orthogonal action directions rather than one; if the single-basis model already decouples perfectly, the multi-basis variant should perform identically, while a gain would indicate the core assumption is too strong. A direct measurement is to compute, for same-clothing/different-action pairs, the cosine similarity between $z_{\text{app}}$ and the text feature of the action description, which should be near zero for both actions under strict decoupling.

Watch

Extended reading notes

Core claim

The paper's central claim is that action and appearance can be strictly decoupled in a shared image-text space using text priors alone. An action inversion operator maps the global visual feature $v$ to sparse coefficients over a frozen action-word codebook and reconstructs $z_{\text{act}}$, then defines the appearance feature as $z_{\text{app}} = (I-\Pi_{\text{act}})v$, an orthogonal projection onto the null space of $z_{\text{act}}$. The paper shows that the Euclidean gradient of the contrastive loss with respect to $v$ splits into a tangent component and a harmful normal component, and that replacing it with the Riemannian gradient $g_{\text{Riem}} = (I-\Pi_{\text{act}})g_{\text{Euc}}$, blended by an entropy weight $\omega = \exp(-H/\tau)$, blocks the normal component's shortcut. Empirically, this yields 84.73% R@1 and 91.93% mAP on PAB at 0.1M data, and 85.49% R@1 and 92.20% mAP at 1M, surpassing the CMP baseline, while the same pipeline improves all four TIPR benchmarks.

Load-bearing premise

The decoupling argument rests on the assumption that the action content of an image is fully captured by the single reconstructed action vector $z_{\text{act}}$, so the orthogonal complement $(I-\Pi_{\text{act}})v$ contains no action information and all appearance information.

Editorial extensions

If this is right

  • External pose estimators become unnecessary for TPAS: action information is supplied by text anchors, so retrieval should survive occlusion, unusual poses, and low-resolution surveillance frames.
  • Hard-negative 'same appearance, different action' pairs no longer force appearance shortcuts, since the normal gradient component is blocked and updates stay on the tangent space.
  • The same decoupling machinery transfers to conventional TIPR, producing best reported averages on CUHK-PEDES, ICFG-PEDES, RSTPReid, and UFineBench, with an 8.69-point average Recall gain on UFine3C.
  • Out-of-distribution robustness improves: with only 0.1M training pairs LightAIR beats CMP trained on the full 1M data on the UCC set (62.36 vs 55.23 R@1; 51.53 vs 44.35 mAP).

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: the whole decoupling scheme hinges on the action subspace being one-dimensional. If a pedestrian performs two simultaneous or sequential actions, a single $z_{\text{act}}$ cannot span the action content, and the null-space projection would either leak action into appearance or discard part of the action; a multi-basis version of the codebook projection is the natural stress test.
  • Beyond the paper: the entropy-weighted Riemannian gradient rectification is a general anti-shortcut mechanism for any contrastive retrieval task with a weak semantic signal and a dominant distractor modality; the paper demonstrates it only for person search, but the gradient decomposition argument is task-agnostic.
  • Beyond the paper: because the codebook is frozen and built from training vocabulary, deployment to genuinely novel action words may require codebook extension; the UCC experiment shifts scene distribution, not action vocabulary, so it does not test vocabulary generalization.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper targets Text-based Person Anomaly Search (TPAS), where a query describes both macro-level appearance and micro-level abnormal actions. It proposes LightAIR, combining three modules: an Action Inversion Operator (AIO) that reconstructs an action feature z_act from a frozen text-derived semantic codebook, an Orthogonal Null-Space Projection (ONSP) that obtains an appearance feature z_app by projecting the visual feature onto the null space of z_act, and a Gradient Rectification (GR) module that projects the Euclidean gradient onto a tangent space with an entropy-based adaptive weight. The method is evaluated on the PAB TPAS benchmark, its Multi-Weather and UCC variants, and four TIPR benchmarks, reporting state-of-the-art results, e.g., PAB 0.1M R@1 of 84.73% and mAP of 91.93%, and UCC R@1 of 62.36% with only 0.1M training data. Ablations are provided for each module across the TPAS settings.

Significance. If the reported results are reproducible, LightAIR would be a strong new state of the art for TPAS and also improves conventional TIPR benchmarks. The design avoids external pose estimators and the code is released, which are practical strengths. The ablations are reasonably complete at the module level and the reported gains over CMP are consistent across PAB, Multi-Weather, and UCC. However, the central mathematical claim of 'strict forward decoupling' is only established for the single direction defined by z_act; the paper's own failure cases in Appendix E acknowledge residual action–appearance interference, and the current experiments do not test a multi-dimensional action-subspace projector. The core idea is promising, but the decoupling guarantee and its empirical support need to be tightened before the headline claims are acceptable.

major comments (3)
  1. [§3.3, Eqs. (6)–(7); Abstract; Contributions] The claim that ONSP 'mathematically eliminates' action information from z_app is not established by the construction. The projector Pi_act in Eq. (6) is rank-1, built from the single vector z_act of Eq. (4), so Eq. (7) removes only the component of v parallel to z_act. If TPAS action semantics occupy more than one direction in the embedding space, the residual action information in z_app remains. This is not merely hypothetical: Appendix E, Figure 11(c) states that 'weak local action features are easily overshadowed by explicit large-area appearance attributes' and that 'the current forward decoupling mechanism still faces the risk of interference from dominant appearance information.' Please either restrict the claim to removal of the component along the estimated action direction, or provide evidence that a single direction suffices. A concrete test would be to train a linear probe on z_app to measure remaining action information, and to ablate a multi-dimensional action-basis projector built from the top-K codebook directions.
  2. [§3.2, Eq. (5); Algorithm 1, lines 6 and 10] The action discriminative loss L_cls aligns z_act with t_y = Phi_T(x_T), which is the feature of the full text query, not an action-only feature. Because the query text also describes appearance (e.g., clothing, background), this supervision can pull z_act toward appearance directions, weakening the claim that AIO extracts 'pure action features' and potentially reintroducing appearance into the action branch before ONSP operates. Please clarify whether t_y is restricted to action words (for example, by projecting the text feature onto the action codebook D) or provide an ablation that replaces the full-text feature with an action-only feature in L_cls.
  3. [§4.3.2, Figure 4 and accompanying text] The text states that 'using multiple action bases' lowers Multi-Weather R@1 by 3.02 points, but the ablation D#6 is labeled 'w/o num_K' (removing Top-K sparsification). Removing Top-K changes z_act to a dense combination of codebook atoms, yet the projection operator in Eq. (6) remains rank-1. Thus the experiment does not test a multi-dimensional action-basis projector and cannot support the conclusion that a single action direction is sufficient. Please correct the description and add an explicit experiment with a projector onto a multi-dimensional action subspace, or discuss why the rank-1 projector is adequate despite the Appendix E failure cases.
minor comments (6)
  1. [§3.4, Eq. (8)] The two terms in the gradient decomposition are called 'orthogonal components,' but no orthogonality is shown; the second term, which involves the derivative of Pi_act, is not generally contained in the range of Pi_act. Please clarify the claim or soften the wording. Also, the symbol ×1 is used without definition.
  2. [§3.4, Eq. (9)] The term 'Riemannian gradient' is used for an orthogonal projection onto a linear subspace T_vM, but the manifold M is never formally defined. If the tangent space is just the null space of Pi_act, the analysis is a linear projection; please either define the manifold or use a more neutral term such as 'projected gradient.'
  3. [§4.3, Figures 3–6] The captions of the ablation figures do not state whether the reported numbers correspond to the 0.1M or 1M training setting; the main text refers to both. Please state the training data scale in each caption or in the surrounding text.
  4. [Table 2] The rows for CLIP and X-VLM without a #Data value are not directly comparable to the 0.1M and 1M rows. Please specify the training data used for those baseline rows or remove them from the comparison table.
  5. [References] Reference [150] is listed as 'MRA (arXiv'25)' in Table 2, but the reference entry is identical to [59], the CMP ICCV'25 paper. Please correct the duplicate or replace it with the intended MRA reference.
  6. [§4.1.2, Implementation Details] The statement 'We equivalently implement Riemannian gradient rectification by applying a gradient clipping operation to Pi_act' is vague. Please specify how the clipping is applied to the projection operator and how it relates to Eq. (11).

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the action feature is text-supervised, the projections are explicit geometric constructions, and all reported metrics are held-out test results.

full rationale

LightAIR's derivation chain is not circular. The action feature z_act (Eq. 4) is trained against the text-derived action feature t_y via the discriminative loss L_cls (Eq. 5); the projector Pi_act (Eq. 6) and the null-space appearance feature z_app (Eq. 7) are then built directly from z_act, and retrieval is evaluated on held-out PAB, UCC, and TIPR test splits. No fitted constant is later relabeled as a prediction: every R@1/mAP value is measured on test data, and the ablations compare functionally distinct variants (w/o L_cls, w/o Ortho, w/o Rectification, etc.). The only self-referential element is that z_app is orthogonal to z_act by construction, but removing the component along the reconstructed action direction is the intended geometric operation rather than a hidden reuse of the target metric. The rank-1 subspace assumption behind Eqs. 6-7 is a genuine correctness risk: Appendix E, Figure 11(c), admits that 'weak local action features are easily overshadowed by explicit large-area appearance attributes' and that 'the current forward decoupling mechanism still faces the risk of interference from dominant appearance information,' and no ablation tests a multi-dimensional action-basis projector. However, an overclaim or an untested assumption is a robustness limitation, not circularity. Although the paper cites many works by its own authors in the Related Work sections, none of those citations is load-bearing: the null-space projection is attributed to Ravfogel et al. [139], the Riemannian gradient to Absil et al. [140], and shortcut learning to Geirhos et al. [73]. There is no imported uniqueness theorem, no ansatz smuggled in via self-citation, and no fitted input presented as a prediction. The empirical claims are self-contained and evaluated against external benchmarks, so the appropriate circularity score is 0.

Assumptions & free parameters 5 free parameters · 4 assumptions · 0 invented entities

The central method depends on several validation-tuned hyperparameters (K, gamma_1, gamma_2, tau) and on the assumption that a single action direction spans all action semantics. The codebook construction and entropy schedule are design choices without independent evidence.

free parameters (5)
  • Top-K sparsity K = 3
    K is the number of activated action codebook entries; sensitivity analysis (Appendix B.3) shows optimal at K=3 for PAB.
  • image-text matching loss weight gamma_1 = 4
    Tuned on PAB; Sec 4.4 shows optimum at 4.0.
  • action discriminative loss weight gamma_2 = 1
    Tuned on PAB; Sec 4.4 shows optimum at 1.0.
  • entropy temperature tau = unspecified
    Eq. 10 uses tau in omega=exp(-H/tau), but no value or sensitivity analysis is reported.
  • semantic codebook size L = unspecified
    Number of action words extracted from the training captions is not reported; it determines the codebook dimension.
assumptions (4)
  • ad hoc to paper The action subspace is fully spanned by the single normalized feature z_act from Eq. 4.
    ONSP projects out only the direction of z_act (Eq. 7); if actions span multiple directions or are entangled with appearance, the claimed strict decoupling does not follow.
  • domain assumption The frozen text semantic codebook D provides reliable and complete anchors for all relevant actions, including OOD actions.
    The codebook is built from training captions via NLTK/SpaCy; OOD test sets (UCC) may contain action words not in the codebook (Sec 3.2, App A.1).
  • standard math Riemannian gradient on the action semantic manifold equals orthogonal projection onto the null space of Pi_act.
    Footnote 1 and Eq. 9 define the Riemannian gradient this way, following Absil et al.; for a Euclidean embedding with a linear constraint, this is the standard projection.
  • ad hoc to paper The entropy-guided weight omega=exp(-H/tau) is a valid measure of action mapping confidence.
    Eq. 10 introduces this schedule without a derivation or comparison to other confidence measures beyond the ablations in Fig. 6.

how reviews work

0 comments
Cite this review

Pith. "Pith review of LightAIR: Lightweight Action Inversion and Riemannian Rectification for Text-based Person Anomaly Search." pith.science (2026). https://pith.science/paper/Z6UMUAEB

@misc{pith2026260809152,
  author       = {Pith},
  title        = {Pith review of: LightAIR: Lightweight Action Inversion and Riemannian Rectification for Text-based Person Anomaly Search},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Z6UMUAEB}},
  note         = {Machine review of arXiv:2608.09152}
}
read the original abstract

Traditional Text-based Person Search (TPS) is typically limited to matching static appearance attributes, severely neglecting dynamic action information. The Text-based Person Anomaly Search (TPAS) task bridges this gap, requiring models to locate micro-level specific abnormal behaviors while matching macro-level appearance of pedestrians. However, current TPAS methods face fundamental limitations: external explicit pose estimators are fragile in unconstrained surveillance scenarios, and implicit learning encounters visual decoupling failure under pixel-level entanglement, causing dominant appearance information to easily swallow and contaminate subtle action features. Furthermore, performing contrastive optimization on hard negative samples (``same appearance, different actions'') in conventional Euclidean spaces induces severe shortcut learning. To address these, we propose the Lightweight Action Inversion and Riemannian rectification network (LightAIR). First, it introduces textual semantic priors as anchors via a lightweight action inversion operator to extract pure action features, thereby overcoming visual-inherent coupling. Subsequently, it employs orthogonal null-space projection to constrain appearance features within the orthogonal complement space of action features, guaranteeing strict forward decoupling. Finally, we designed a gradient rectification module that computes the Riemannian gradient to constrain the backpropagation trajectory, forcing the gradient flow to update strictly along the tangent space that preserves decoupling properties, thereby cutting off harmful shortcuts. Extensive experiments on the widely used TPAS and TIPR datasets demonstrate that LightAIR significantly outperforms existing state-of-the-art methods. Codes are available at https://github.com/rainy-london/LightAIR

Figures

Figures reproduced from arXiv: 2608.09152 by the authors.

Figure 1
Figure 1. (a) Examples of TIPR and TPAS tasks. (b) Pixel-level [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. LightAIR consists of (a) Action Inversion Operator, which utilizes textual semantic priors to extract pure action features; (b) Orthogonal Null-Space Projection, which ensures strict forward decoupling by projecting raw visual features; and (c) Gradient Rectification, which computes riemannian gradients to effectively cut off harmful shortcuts. preliminary explorations of this highly challenging emerging task with t… view at source ↗
Figure 3
Figure 3. Ablation study of AIO module on PAB, Multi-Weather, and UCC datasets. Figure best viewed in color. [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: Ablation study of ONSP module on PAB, Multi [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Ablation study of GR module on PAB, Multi [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: Ablation study of Entropy-Guided Retraction pro [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 8
Figure 8. Figure 8: Case study on LightAIR and CMP. Matched images [PITH_FULL_IMAGE:figures/full_fig_p008_8.png]
Figure 7
Figure 7. Figure 7: Hyper-parameters sensitivity analysis of loss [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]
Figure 9
Figure 9. Figure 9: Sensitivity analysis of the sparsity parameter [PITH_FULL_IMAGE:figures/full_fig_p015_9.png]
Figure 10
Figure 10. Figure 10: Qualitative comparison of successful retrieval cases between LightAIR and the baseline CMP. Matched images are [PITH_FULL_IMAGE:figures/full_fig_p019_10.png]
Figure 11
Figure 11. Figure 11: Qualitative analysis of failure cases for both LightAIR and CMP. Matched images are marked by green boxes, and [PITH_FULL_IMAGE:figures/full_fig_p020_11.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

180 extracted references · 51 canonical work pages

  1. [1]

    Min Cao, Xinyu Zhou, Ding Jiang, Bo Du, Mang Ye, and Min Zhang. 2025. Mul- tilingual Text-to-Image Person Retrieval via Bidirectional Relation Reasoning and Aligning.IEEE Transactions on Pattern Analysis and Machine Intelligence (2025)

  2. [3]

    Ding Jiang and Mang Ye. 2023. Cross-modal implicit relation reasoning and aligning for text-to-image person retrieval. InProceedings of the IEEE/CVF con- ference on computer vision and pattern recognition. 2787–2797

  3. [4]

    Zhengxian Wu, Chuanrui Zhang, Shen’Ao Jiang, Hangrui Xu, Zirui Liao, Luyuan Zhang, Li Huaqiu, Peng Jiao, and Haoqian Wang. 2026. Language-guided and motion-aware gait representation for generalizable recognition. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 40. 10871–10878

  4. [5]

    Hao Li, Yuhao Wang, Wenning Hao, Pingping Zhang, Dong Wang, and Huchuan Lu. 2026. RAGTrack: Language-aware RGBT Tracking with Retrieval- Augmented Generation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 28179–28189

  5. [6]

    Jiale Huang, Zixu Li, Zhiheng Fu, Zhiwei Chen, Qinlei Huang, and Yupeng Hu. 2026. RankVR: Low-Rank Structure Perception and Value Recalibration for Robust Composed Image Retrieval. InProceedings of the 2026 International Conference on Multimedia Retrieval. 269–278

  6. [7]

    Zixu Li, Yupeng Hu, Zhiheng Fu, Zhiwei Chen, Weili Guan, and Liqiang Nie. 2026. R3: Composed Video Retrieval via Reasoning-Guided Recalling and Re-ranking. arXiv preprint arXiv:2606.01113(2026)

  7. [8]

    Mingyu Zhang, Zixu Li, Zhiwei Chen, Zhiheng Fu, Xiaowei Zhu, Jiajia Nie, Yinwei Wei, and Yupeng Hu. 2026. Hint: Composed image retrieval with dual- path compositional contextualized network. InICASSP. IEEE, 13002–13006

  8. [9]

    Jinhe Bi, Aniri, Minglai Yang, Xingcheng Zhou, Wenke Huang, Sikuan Yan, Yujun Wang, Zixuan Cao, Michael Färber, Xun Xiao, Volker Tresp, and Yunpu Ma. 2026. EchoRL: Reinforcement Learning via Rollout Echoing. InForty-third International Conference on Machine Learning. https://openreview.net/forum? id=A6az59SGtF

Show all 180 references
  1. [10]

    Guozhi Qiu, Zhiwei Chen, Zixu Li, Qinlei Huang, Zhiheng Fu, Xuemeng Song, and Yupeng Hu. 2026. Melt: Improve composed image retrieval via the modifi- cation frequentation-rarity balance network. InICASSP. IEEE, 13007–13011

  2. [11]

    Xi Xiao, Xingjian Li, Yunbei Zhang, Cheng Han, Tianming Liu, Tianyang Wang, Runmin Jiang, Jihun Hamm, Xiao Wang, and Min Xu. 2026. Layer-Specific Prompt Fusion Discovery via Differentiable Search in Vision Foundation Models. arXiv preprint arXiv:2606.26379(2026)

  3. [12]

    Xi Xiao, Chenrui Ma, Yunbei Zhang, Chen Liu, Zhuxuanzi Wang, Yanshu Li, Lin Zhao, Guosheng Hu, Tianyang Wang, and Hao Xu. 2026. Not all directions matter: Towards structured and task-aware low-rank model adaptation. In Proceedings of the 64th Annual Meeting of the Association ...

  4. [13]

    Lin Zhao, Xinru Jiang, Xi Xiao, Qihui Fan, Lei Lu, Yanzhi Wang, Xue Lin, Octavia Camps, Pu Zhao, and Jianyang Gu. 2026. Hieramp: Coarse-to-fine autoregressive amplification for generative dataset distillation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pat...

  5. [14]

    Hao Li, Yuhao Wang, Xiantao Hu, Wenning Hao, Pingping Zhang, Dong Wang, and Huchuan Lu. 2026. Cadtrack: Learning contextual aggregation with de- formable alignment for robust rgbt tracking. InProceedings of the AAAI Confer- ence on Artificial Intelligence, Vol. 40. 6109–6117

  6. [15]

    Weilin Wu, Shifan Yang, Qizhao Lin, XingHong Chen, Kunping Yang, Jing Wang, and Guannan Chen. 2025. A Novel Perspective on Low-Light Image Enhancement: Leveraging Artifact Regularization and Walsh-Hadamard Trans- form. InProceedings of the 33rd ACM International Conference on ...

  7. [16]

    Yuan Sun, Xu Wang, Dezhong Peng, Zhenwen Ren, and Xiaobo Shen. 2023. Hierarchical hashing learning for image set classification.IEEE TIP32 (2023), 1732–1744

  8. [17]

    Qianyun Yang, Zhiwei Chen, Yupeng Hu, Zixu Li, Zhiheng Fu, and Liqiang Nie. 2026. STABLE: Efficient Hybrid Nearest Neighbor Search via Magnitude- Uniformity and Cardinality-Robustness.IEEE TKDE(2026)

  9. [18]

    Yang Qin, Yuan Sun, Dezhong Peng, Joey Tianyi Zhou, Xi Peng, and Peng Hu

  10. [19]

    Jiale Huang, Zixu Li, Zhiwei Chen, Zhiheng Fu, Chunxiao Wang, and Yupeng Hu. 2026. IMAGINE: Adaptive Schema-Imagery Enhanced Composition for Composed Video Retrieval. InProceedings of the 2026 International Conference on Multimedia Retrieval. 288–297

  11. [20]

    Jinhe Bi, Danqi Yan, Yifan Wang, Wenke Huang, Haokun Chen, Guancheng Wan, Mang Ye, Xun Xiao, Hinrich Schuetze, Volker Tresp, and Yunpu Ma. 2026. The Geometry of Reasoning: Self-Evaluation via Layerwise Trajectory Evolution. In Forty-third International Conference on Machine Le...

  12. [21]

    Honglin Yuan, Yuan Sun, Fei Zhou, Jing Wen, Shihua Yuan, Xiaojian You, and Zhenwen Ren. 2025. Prototype matching learning for incomplete multi-view clustering.IEEE TIP34 (2025), 828–841

  13. [22]

    Zixu Li, Yupeng Hu, Zhiwei Chen, Qinlei Huang, Guozhi Qiu, Zhiheng Fu, and Meng Liu. 2026. ReTrack: Evidence-Driven Dual-Stream Directional Anchor Calibration Network for Composed Video Retrieval. InAAAI, Vol. 40. 23373– 23381

  14. [23]

    Yupeng Hu, Zixu Li, Zhiwei Chen, Qinlei Huang, Zhiheng Fu, Mingzhu Xu, and Liqiang Nie. 2026. REFINE: Composed Video Retrieval via Shared and Differential Semantics Enhancement.ACM ToMM(2026)

  15. [24]

    Yupeng Hu, Liqiang Nie, Meng Liu, Kun Wang, Yinglong Wang, and Xian- Sheng Hua. 2021. Coarse-to-fine semantic alignment for cross-modal moment localization.IEEE Transactions on Image Processing30 (2021), 5933–5943

  16. [25]

    Yupeng Hu, Kun Wang, Meng Liu, Haoyu Tang, and Liqiang Nie. 2023. Semantic collaborative learning for cross-modal moment localization.ACM Transactions on Information Systems42, 2 (2023), 1–26

  17. [26]

    Yupeng Hu, Meng Liu, Xiaobin Su, Zan Gao, and Liqiang Nie. 2021. Video moment localization via deep cross-modal hashing.IEEE Transactions on Image Processing30 (2021), 4667–4677

  18. [27]

    Lixian Chen, Jingchao Wang, Zhaorong Dai, Hanqian Liu, Danxiang Ai, and Yang Shi. 2026. When task performance deceives: Task-geometry decoupling in learnable-curvature hyperbolic GNNs.Neural Networks(2026), 109172

  19. [28]

    Wei Zhang, Yihang Wu, Shengkai Yu, Songhua Li, Qiang Li, and Qi Wang

  20. [29]

    Xi Xiao, Yunbei Zhang, Lin Zhao, Yiyang Liu, Xiaoying Liao, Zheda Mai, Xingjian Li, Xiao Wang, Hao Xu, Jihun Hamm, Xue Lin, Min Xu, Qifan Wang, Tianyang Wang, and Cheng Han. 2026. Prompt-based Adaptation in Large-scale Vision Models: A Survey.Transactions on Machine Learning R...

  21. [30]

    Xingfeng Li, Yinghui Sun, Quansen Sun, Zhenwen Ren, and Yuan Sun. 2023. Cross-view graph matching guided anchor alignment for incomplete multi-view clustering.Information Fusion100 (2023), 101941

  22. [31]

    Qianyun Yang, Peizhuo Lv, Yingjiu Li, Shengzhi Zhang, Yuxuan Chen, Zhiwei Chen, Zixu Li, and Yupeng Hu. 2026. ERASE: Bypassing Collaborative Detection of AI Counterfeit Via Comprehensive Artifacts Elimination.IEEE TDSC(March 2026), 1–18. doi:10.1109/TDSC.2026.3677794

  23. [32]

    Jinhe Bi, Yujun Wang, Haokun Chen, Xun Xiao, Artur Hecker, Volker Tresp, and Yunpu Ma. 2025. LLaVA steering: Visual instruction tuning with 500x fewer parameters through modality linear representation-steering. InProceedings of the 63rd Annual Meeting of the Association for Co...

  24. [33]

    Zixu Li, Zhiwei Chen, Zhiheng Fu, Wenbo Wang, Yupeng Hu, Weili Guan, and Liqiang Nie. 2026. OmniEgo-R 2: A Routed Reasoning Framework for the 1st Cross-Domain EgoCross Challenge at CVPR 2026.arXiv preprint arXiv:2605.24481(2026)

  25. [34]

    Xingfeng Li, Yuangang Pan, Yuan Sun, Quansen Sun, Yinghui Sun, Ivor W Tsang, and Zhenwen Ren. 2024. Incomplete multi-view clustering with paired and balanced dynamic anchor learning.IEEE TMM27 (2024), 1486–1497

  26. [35]

    Xi Xiao, Yunbei Zhang, Xingjian Li, Tianyang Wang, Xiao Wang, Yuxiang Wei, Jihun Hamm, and Min Xu. 2025. Visual instance-aware prompt tuning. In Proceedings of the 33rd ACM International Conference on Multimedia. 2880–2889

  27. [36]

    Liu Yu, Fenghui Tian, Ping Kuang, ZhiKun Feng, and Fan Zhou. 2025. Knowl- edge Graphs Acquisition via Forward-Reverse Relation Enhanced Contrastive Pretraining from Large-scale Models. In2025 IEEE International Conference on Multimedia and Expo (ICME). IEEE, 1–6

  28. [37]

    Tao Huang, Rui Wang, Xiaofei Liu, Yi Qin, Li Duan, and Liping Jing. 2026. Detect- ing Misbehaviors of Large Vision-Language Models by Evidential Uncertainty Quantification.arXiv preprint arXiv:2602.05535(2026)

  29. [38]

    Wei Zhang, Qiang Li, Yuan Yuan, and Qi Wang. 2024. Visual Consistency Enhancement for Multiview Stereo Reconstruction in Remote Sensing.IEEE Transactions on Geoscience and Remote Sensing62 (2024), 1–11. doi:10.1109/ TGRS.2024.3482697

  30. [39]

    Yan Zhong, Chenxi Yang, Suyuan Zhao, and Tingting Jiang. 2025. Semi- supervised blind quality assessment with confidence-quantifiable pseudo-label learning for authentic images. InForty-second International Conference on Ma- chine Learning

  31. [40]

    Guangtao Lyu, Qi Liu, Chenghao Xu, Jiexi Yan, Muli Yang, Xueting Li, Fen Fang, and Cheng Deng. 2026. Revealing and Enhancing Core Visual Regions: Harnessing Internal Attention Dynamics for Hallucination Mitigation in LVLMs. ACL Findings(2026)

  32. [41]

    Zixu Li, Zhiwei Chen, Haokun Wen, Zhiheng Fu, Yupeng Hu, and Weili Guan

  33. [42]

    Jinhe Bi, Danqi Yan, Yifan Wang, Wenke Huang, Haokun Chen, Guancheng Wan, Mang Ye, Xun Xiao, Hin rich Schuetze, Volker Tresp, and Yunpu Ma. 2025. CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process.ArXiv abs/2505.13408 (2025). https://api.semanticscholar.org/C...

  34. [43]

    Yujun Wang, Jinhe Bi, Soren Pirk, Yunpu Ma, et al . 2026. Ascd: Attention- steerable contrastive decoding for reducing hallucination in mllm. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 40. 10306–10314

  35. [44]

    Jinhe Bi, Yifan Wang, Danqi Yan, Xun Xiao, Artur Hecker, Volker Tresp, and Yunpu Ma. 2025. PRISM: Self-Pruning Intrinsic Selection Method for Training- Free Multimodal Data Selection.ArXivabs/2502.12119 (2025). https://api. semanticscholar.org/CorpusID:276421326

  36. [45]

    Zixu Li, Zhiheng Fu, Yupeng Hu, Zhiwei Chen, Haokun Wen, and Liqiang Nie

  37. [46]

    Canran Xiao, Tianxiang Xu, Siyuan Ma, Yiyang Jiang, Haoyu Gao, and Yuhan Wu. 2026. Reversible primitive–composition alignment for continual vision– language learning. InThe Fourteenth International Conference on Learning Rep- resentations

  38. [47]

    Zhiheng Fu, Zixu Li, Zhiwei Chen, Chunxiao Wang, Xuemeng Song, Yupeng Hu, and Liqiang Nie. 2025. PAIR: Complementarity-guided Disentanglement for Composed Image Retrieval. InProceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE, 1–5

  39. [48]

    Qinlei Huang, Zhiwei Chen, Zixu Li, Chunxiao Wang, Xuemeng Song, Yu- peng Hu, and Liqiang Nie. 2025. MEDIAN: Adaptive Intermediate-grained Aggregation Network for Composed Image Retrieval. InProceedings of the IEEE International Conference on Acoustics, Speech and Signal Proce...

  40. [49]

    FineCIR: Explicit Parsing of Fine-Grained Modification Semantics for Composed Image Retrieval.https://arxiv.org/abs/2503.21309(2025)

  41. [50]

    Zhiwei Chen, Yupeng Hu, Zixu Li, Zhiheng Fu, Xuemeng Song, and Liqiang Nie. 2025. OFFSET: Segmentation-based Focus Shift Revision for Composed Image Retrieval. InACM MM. 6113–6122

  42. [51]

    Guangtao Lyu, Xinyi Cheng, Chenghao Xu, Qi Liu, Muli Yang, Fen Fang, Huilin Chen, Jiexi Yan, Xu Yang, and Cheng Deng. 2025. Revealing Perception and Generation Dynamics in LVLMs: Mitigating Hallucinations via Validated Domi- nance Correction.arXiv preprint arXiv:2512.18813(2025)

  43. [52]

    Liangsi Lu, Jingchao Wang, Zhaorong Dai, Hanqian Liu, and Yang Shi. 2026. Riemannian liquid spatio-temporal graph network. InProceedings of the ACM Web Conference 2026. 463–474

  44. [53]

    Yuan Sun, Yang Qin, Yongxiang Li, Dezhong Peng, Xi Peng, and Peng Hu. 2024. Robust multi-view clustering with noisy correspondence.IEEE TKDE36, 12 (2024), 9150–9162

  45. [54]

    Zhiheng Fu, Zixu Li, Zhiwei Chen, Fangxu Liu, Yupeng Hu, Weili Guan, and Liqiang Nie. 2026. EgoAction: Egocentric Action Composition with Reliability- Aware Temporal Fusion for the EPIC-KITCHENS Action Detection Challenge at CVPR 2026.arXiv preprint arXiv:2605.24496(2026)

  46. [55]

    Zhiwei Chen, Yupeng Hu, Zixu Li, Zhiheng Fu, Guozhi Qiu, Weili Guan, and Liqiang Nie. 2026. EgoAdapt: A Multi-Scene Egocentric Adaptation Method for CVPR 2026 HD-EPIC VQA Challenge.arXiv preprint arXiv:2605.24500(2026)

  47. [56]

    Zhiming Lin, Canran Xiao, and Kai Zhao. 2026. Beyond More Context: Retrieval Diversity Boosts Multi-Turn Intent Understanding. InProceedings of the ACM Web Conference 2026. 2320–2329

  48. [57]

    Min Cao, Yang Bai, Ziyin Zeng, Mang Ye, and Min Zhang. 2024. An empirical study of clip for text-based person search. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 465–473

  49. [58]

    Qicheng Zhao, Qi Sun, and Zheyu Yan. 2026. Seeing the End at Step Zero: Accelerating Diffusion MLLMs via MLP Sparsity-Aware Truncation.arXiv preprint arXiv:2607.14557(2026)

  50. [60]

    Guangtao Lyu, Chenghao Xu, Jiexi Yan, Muli Yang, and Cheng Deng. 2025. To- wards Unified Human Motion-Language Understanding via Sparse Interpretable Characterization. InThe Thirteenth International Conference on Learning Repre- sentations(ICLR)

  51. [61]

    Xi Xiao, Chen Liu, Chih-Ting Liao, Yunbei Zhang, Qizhen Lan, Yuxiang Wei, Lin Zhao, Janet Wang, Jianyang Gu, Muchao Ye, et al. 2026. Staying VIGILant: Mitigating Visual Laziness via Counterfactual Visual Alignment in MLLMs. arXiv preprint arXiv:2606.26387(2026)

  52. [62]

    Kaifang Long, Lianbo Ma, Jiaqi Liu, Liming Liu, and Guoyang Xie. 2026. To- wards an Incremental Unified Multimodal Anomaly Detection: Augmenting Multimodal Denoising From an Information Bottleneck Perspective. InProceed- ings of the IEEE/CVF Conference on Computer Vision and P...

  53. [63]

    Haitian Li, Yanghao Zhou, Heyan Huang, Liangji Chen, YiMing Cheng, Xu Liu, Dian Jin, Jiajun Xu, Jingyun Liao, Tian Lan, et al . 2026. MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation.arXiv preprint arXiv:2605.28035(2026)

  54. [64]

    Guangtao Lyu, Chenghao Xu, Qi Liu, Jiexi Yan, Muli Yang, Fen Fang, and Cheng Deng. 2025. Tempo as the Stable Cue: Hierarchical Mixture of Tempo and Beat Experts for Music to 3D Dance Generation.arXiv preprint arXiv:2512.18804 (2025)

  55. [65]

    Zhengxian Wu, Chuanrui Zhang, Hangrui Xu, Peng Jiao, and Haoqian Wang

  56. [66]

    In2025 IEEE International Conference on Multimedia and Expo (ICME)

    DAGait: Generalized skeleton-guided data alignment for gait recognition. In2025 IEEE International Conference on Multimedia and Expo (ICME). IEEE, 1–6

  57. [67]

    Kaifang Long, Guoyang Xie, Lianbo Ma, Qing Li, Min Huang, Jianhui Lv, and Zhichao Lu. 2025. Enhancing Multimodal Learning via Hierarchical Fusion Architecture Search With Inconsistency Mitigation.IEEE Transactions on Image Processing(2025)

  58. [68]

    Kaifang Long, Guoyang Xie, Lianbo Ma, Jiaqi Liu, and Zhichao Lu. 2025. Re- visiting multimodal fusion for 3D anomaly detection from an architectural perspective. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 12273–12281

  59. [69]

    Guangtao Lyu, Xinyi Cheng, Qi Liu, Chenghao Xu, Jiexi Yan, Muli Yang, Fen Fang, and Cheng Deng. 2026. Towards Interpretable Hallucination Analysis and Mitigation in LVLMs via Contrastive Neuron Steering.arXiv preprint arXiv:2602.00621(2026)

  60. [70]

    Guangtao Lyu, Xinyi Cheng, Qi Liu, Chenghao Xu, Jiexi Yan, Muli Yang, Fen Fang, and Cheng Deng. [n. d.]. COME: Advancing Representation Learning and Generative Modeling for High-Quality Text-to-Motion Generation. ([n. d.])

  61. [71]

    Yang-Hao Zhou, Haitian Li, Rexar Lin, Heyan Huang, Jinxing Zhou, Changsen Yuan, Tian Lan, Ziqin Zhou, Yudong Li, Jiajun Xu, et al. 2026. MTAVG-Bench: A Comprehensive Benchmark for Evaluating Multi-Talker Dialogue-Centric Audio-Video Generation.arXiv preprint arXiv:2602.00607(2026)

  62. [72]

    Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. 2021. Align before fuse: Vision and lan- guage representation learning with momentum distillation.Advances in neural information processing systems34 (2021), 9694–9705

  63. [73]

    Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann. 2020. Shortcut learning in deep neural networks.Nature Machine Intelligence2, 11 (2020), 665–673

  64. [74]

    Yuhao Chen, Guoqing Zhang, Yujiang Lu, Zhenxing Wang, and Yuhui Zheng

  65. [75]

    Andra Acsintoae, Andrei Florescu, Mariana-Iuliana Georgescu, Tudor Mare, Paul Sumedrea, Radu Tudor Ionescu, Fahad Shahbaz Khan, and Mubarak Shah. 2022. Ubnormal: New benchmark for supervised open-set video anomaly detection. In Proceedings of the IEEE/CVF conference on compute...

  66. [76]

    Yutian Lin, Liang Zheng, Zhedong Zheng, Yu Wu, Zhilan Hu, Chenggang Yan, and Yi Yang. 2019. Improving person re-identification by attribute and identity learning.Pattern recognition95 (2019), 151–161

  67. [77]

    Haigang Deng, Qingyang Yang, Chengwei Li, Hanzhong Liang, and Chuanxu Wang. 2025. Video anomaly detection via pseudo-anomaly generation and multi- grained feature learning.Journal of Electronic Imaging34, 1 (2025), 013044– 013044

  68. [78]

    Yan Zhong, Xinping Zhao, Guangzhi Zhao, Bohua Chen, Fei Hao, Ruoyu Zhao, Jiaqi He, Lei Shi, and Li Zhang. 2025. Ctd-inpainting: Towards the coherence of text-driven inpainting with blended diffusion.Information Fusion122 (2025), 103163

  69. [79]

    Yan Zhong, Ruoyu Zhao, Chao Wang, Jiaqi He, Qinghai Guo, Jianguo Zhang, Zhichao Lu, and Luziwei Leng. 2026. Dyn-SSM: Towards the Efficient Long Se- quence Learning via Bio-interpretable Dynamics in Spiking State Space Models. IEEE Transactions on Cognitive and Developmental Sy...

  70. [80]

    ZhengXian Wu, Hangrui Xu, Kai Shi, Zhuohong Chen, Yunyao Yu, Chuanrui Zhang, Zirui Liao, Jun Yang, Zhenyu Yang, Haonan Lu, et al. 2026. ProMSA: Progressive Multimodal Search Agents for Knowledge-Based Visual Question Answering.arXiv preprint arXiv:2606.27974(2026)

  71. [81]

    Xueming Qian, Dan Lu, Yaxiong Wang, Li Zhu, Yuan Yan Tang, and Meng Wang

  72. [82]

    Zhiwei Chen, Yupeng Hu, Zixu Li, Zhiheng Fu, Haokun Wen, and Weili Guan

  73. [83]

    Hari Lee. 2025. Knowledge-Guided Textual Reasoning for Explainable Video Anomaly Detection via LLMs.arXiv preprint arXiv:2511.07429(2025)

  74. [84]

    Chunlei Meng, Guanhong Huang, Rong Fu, Runmin Jian, Zhongxue Gan, and Chun Ouyang. 2026. Clcr: Cross-level semantic collaborative representation for multimodal learning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 1606–1615. LightAIR: L...

  75. [85]

    Rong Fu, Zijian Zhang, Haiyun Wei, Jiekai Wu, Kun Liu, Xianda Li, Haoyu Zhao, Yang Li, Yongtai Liu, Ziming Wang, et al . 2026. LiveGraph: Active- Structure Neural Re-ranking for Exercise Recommendation.arXiv preprint arXiv:2602.17036(2026)

  76. [86]

    Qicheng Zhao, Yu Li, Qi Sun, and Zheyu Yan. 2026. ResilPhase: Plug-and- Play Phase Mapping and Noise-Resilient Macro-Trajectory Extrapolation for Diffusion Acceleration.arXiv preprint arXiv:2606.26769(2026)

  77. [87]

    Yunyao Zhang, Yihao Ai, Zuocheng Ying, Qirui Mi, Junqing Yu, Wei Yang, and Zikai Song. 2026. Coupling Macro Dynamics and Micro States for Long-Horizon Social Simulation.arXiv preprint arXiv:2604.05516(2026)

  78. [88]

    Zixu Li, Yupeng Hu, Zhiwei Chen, Shiqi Zhang, Qinlei Huang, Zhiheng Fu, and Yinwei Wei. 2026. HABIT: Chrono-Synergia Robust Progressive Learning Framework for Composed Image Retrieval. InAAAI, Vol. 40. 6762–6770

  79. [89]

    Zixu Li, Yupeng Hu, Zhiheng Fu, Zhiwei Chen, Yongqi Li, and Liqiang Nie

  80. [90]

    InACM MM

    HUD: Hierarchical Uncertainty-Aware Disambiguation Network for Composed Video Retrieval. InACM MM. 6143–6152

  81. [91]

    Zixu Li, Yupeng Hu, Zhiwei Chen, Haokun Wen, Xuemeng Song, and Liqiang Nie. 2026. COMBINER: Composed Image Retrieval Guided by Attribute-based Neighbor Relations.IEEE TIP(2026)

  82. [92]

    Zijian Zhang, Rong Fu, Yangfan He, Xinze Shen, Yanlong Wang, Xiaojing Du, Haochen You, Keyan Jin, Jiazhao Shi, and Simon Fong. 2026. FinSentLLM: Multi-LLM and structured semantic signals for enhanced financial sentiment forecasting. InICASSP 2026-2026 IEEE International Confer...

  83. [93]

    Wenjie Zhu, Yabin Zhang, Xin Jin, Wenjun Zeng, and Lei Zhang. 2026. Ants: Adaptive negative textual space shaping for ood detection via test-time mllm understanding and reasoning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 20–30

  84. [94]

    Xinglang Zhang, Yunyao Zhang, ZeLiang Chen, Junqing Yu, Wei Yang, and Zikai Song. 2026. Logical Phase Transitions: Understanding Collapse in LLM Logical Reasoning.arXiv preprint arXiv:2601.02902(2026)

  85. [95]

    Yan Zhong, Xingyu Wu, Xinping Zhao, Li Zhang, Xinyuan Song, Lei Shi, and Bingbing Jiang. 2026. Semi-supervised multi-label feature selection with consis- tent sparse graph learning.Neural Networks(2026), 109265

  86. [96]

    Hangrui Xu, Zhengxian Wu, Chuanrui Zhang, Zhuohong Chen, Zhifang Liu, Peng Jiao, and Haoqian Wang. 2026. Psgait: Gait recognition using parsing skeleton. InICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 10427–10431

  87. [97]

    Yan Zhong, Xingyu Wu, Li Zhang, Chenxi Yang, and Tingting Jiang. 2024. Causal-IQA: Towards the Generalization of Image Quality Assessment Based on Causal Inference.. InICML

  88. [98]

    InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

    Tema: Anchor the image, follow the text for multi-modification composed image retrieval. InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 24421–24442

  89. [99]

    Rong Fu, Yemin Wang, Tianxiang Xu, Yongtai Liu, Weizhi Tang, Wangyu Wu, Xiaowen Ma, and Simon Fong. 2026. S-Path-RAG: Semantic-Aware Shortest-Path Retrieval Augmented Generation for Multi-Hop Knowledge Graph Question Answering. InProceedings of the ACM Web Conference 2026. 4057–4068

  90. [100]

    Tailong Luo, Hao Li, Rong Fu, Xinyue Jiang, Huaxuan Ding, Yiduo Zhang, Zilin Zhao, Simon Fong, Guangyin Jin, and Jianyuan Ni. 2026. Multipress: A multi-agent framework for interpretable multimodal news classification.arXiv preprint arXiv:2604.03586(2026)

  91. [101]

    Zixu Li, Yupeng Hu, Zhiwei Chen, Mingyu Zhang, Zhiheng Fu, and Liqiang Nie

  92. [102]

    Zikai Song, Ying Tang, Run Luo, Lintao Ma, Junqing Yu, Yi-Ping Phoebe Chen, and Wei Yang. 2024. Autogenic language embedding for coherent point tracking. InProceedings of the 32nd ACM International Conference on Multimedia. 2021– 2030

  93. [103]

    Zhiheng Fu, Yupeng Hu, Qianyun Yang, Shiqi Zhang, Zhiwei Chen, and Zixu Li. 2026. Air-know: Arbiter-calibrated knowledge-internalizing robust network for composed image retrieval. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2658–2670

  94. [104]

    Wenbing Li, Hang Zhou, Junqing Yu, Zikai Song, and Wei Yang. 2024. Coupled mamba: Enhanced multimodal fusion with coupled state space model.Advances in Neural Information Processing Systems37 (2024), 59808–59832

  95. [105]

    Hao Ju, Hu Zhang, and Zhedong Zheng. 2025. AnomalyLMM: Bridging Genera- tive Knowledge and Discriminative Retrieval for Text-Based Person Anomaly Search.arXiv preprint arXiv:2509.04376(2025)

  96. [106]

    Yang Shi, Jingchao Wang, Liangsi Lu, Haiying Huang, Sutong Xiao, Zhaorong Dai, Ying Wang, and Boyan Xu. 2026. Enhancing Robustness of Constant Curvature Graph Convolutional Network with Lipschitz Regularization.ACM Transactions on Knowledge Discovery from Data(2026)

  97. [107]

    Zikai Song, Junqing Yu, Yi-Ping Phoebe Chen, and Wei Yang. 2022. Transformer tracking with cyclic shifting window attention. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 8791–8800

  98. [108]

    Zikai Song, Run Luo, Junqing Yu, Yi-Ping Phoebe Chen, and Wei Yang. 2023. Compact transformer tracker with correlative masked modeling. InProceedings of the AAAI conference on artificial intelligence, Vol. 37. 2321–2329

  99. [109]

    Yan Zhong, Xinping Zhao, Li Zhang, Xinyuan Song, and Tingting Jiang. 2025. Adaptive Prompt Learning for Blind Image Quality Assessment with Multi- modal Mixed-datasets Training. InProceedings of the 33rd ACM International Conference on Multimedia. 7453–7462

  100. [110]

    Yangliu Hu, Zikai Song, Na Feng, Yawei Luo, Junqing Yu, Yi-Ping Phoebe Chen, and Wei Yang. 2025. Sf2t: Self-supervised fragment finetuning of video-llms for fine-grained understanding. InProceedings of the Computer Vision and Pattern Recognition Conference. 29108–29117

  101. [111]

    InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Conesep: Cone-based robust noise-unlearning compositional network for composed image retrieval. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 16897–16909

  102. [112]

    Damien Teney, Ehsan Abbasnejad, Simon Lucey, and Anton Van den Hengel

  103. [113]

    Liu Yu, Ludie Guo, Ping Kuang, and Fan Zhou. 2025. Bridging the fairness gap: Enhancing pre-trained models with llm-generated sentences. InICASSP 2025- 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1–5

  104. [114]

    Wenjie Zhu, Yabin Zhang, Liang Xu, Xin Jin, Wenjun Zeng, and Lei Zhang. 2026. Dual Distribution Estimation for Zero-shot Noisy Test-Time Adaptation with VLMs.arXiv preprint arXiv:2606.25758(2026)

  105. [115]

    Liu Yu, Fenghui Tian, Ping Kuang, and Fan Zhou. 2025. Amplifying common- sense knowledge via bi-directional relation integrated graph-based contrastive pre-training from large language models.Information Processing & Management 62, 3 (2025), 104068

  106. [116]

    Zikai Song, Run Luo, Lintao Ma, Ying Tang, Yi-Ping Phoebe Chen, Junqing Yu, and Wei Yang. 2025. Temporal Coherent Object Flow for Multi-Object Tracking. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 6978–6986

  107. [117]

    Liu Yu, Can Chen, Ping Kuang, Zhikun Feng, Fan Zhou, and Gillian Dobbie

  108. [118]

    Dismantling Pathological Shortcuts: A Causal Framework for Faithful LVLM Decoding.arXiv preprint arXiv:2606.27596(2026)

  109. [119]

    Zixu Li, Yupeng Hu, Zhiwei Chen, Zhiheng Fu, Xiaowei Zhu, Weili Guan, and Liqiang Nie. 2026. TempRet: Temporal Enhancement and Two-Stage Reranking for CVPR 2026 EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge.arXiv preprint arXiv:2605.24470(2026)

  110. [120]

    Zhiwei Chen, Yupeng Hu, Zhiheng Fu, Zixu Li, Jiale Huang, Qinlei Huang, and Yinwei Wei. 2026. INTENT: Invariance and Discrimination-aware Noise Mitigation for Robust Composed Image Retrieval. InAAAI, Vol. 40. 20463–20471

  111. [121]

    Siming Fu, Sijun Dong, and Xiaoliang Meng. 2025. Disentangling Content from Style to Overcome Shortcut Learning: A Hybrid Generative-Discriminative Learning Framework.arXiv preprint arXiv:2509.11598(2025)

  112. [122]

    Yu Yang, Eric Gan, Gintare Karolina Dziugaite, and Baharan Mirzasoleiman

  113. [123]

    Ziheng Chi, Yifan Hou, Chenxi Pang, Shaobo Cui, Mubashara Akhtar, and Mrinmaya Sachan. 2025. Chimera: Diagnosing Shortcut Learning in Visual- Language Understanding.arXiv preprint arXiv:2509.22437(2025)

  114. [124]

    Yunyao Zhang, Zikai Song, Hang Zhou, Wenfeng Ren, Yi-Ping Phoebe Chen, Junqing Yu, and Wei Yang. 2025. 𝐺𝐴−𝑆 3: Comprehensive Social Network Simulation with Group Agents. InFindings of the Association for Computational Linguistics: ACL 2025. 8950–8970

  115. [125]

    InProceedings of the IEEE/CVF conference on computer vision and pattern recognition

    Evading the simplicity bias: Training a diverse set of models discov- ers solutions with superior ood generalization. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 16761–16772

  116. [126]

    Liu Yu, Yuzhou Mao, Jin Wu, and Fan Zhou. 2023. Mixup-based unified frame- work to overcome gender bias resurgence. InProceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval. 1755–1759. MM ’26, November 10–14, 2026, Rio d...

  117. [127]

    Shunxin Wang, Raymond Veldhuis, and Nicola Strisciuglio. 2025. Do ImageNet- trained models learn shortcuts? The impact of frequency shortcuts on general- ization. InProceedings of the Computer Vision and Pattern Recognition Conference. 25198–25207

  118. [128]

    Yihe Deng, Yu Yang, Baharan Mirzasoleiman, and Quanquan Gu. 2023. Robust learning with progressive data expansion against spurious correlation.Advances in neural information processing systems36 (2023), 1390–1402

  119. [129]

    Yangyi Fang, Jiaye Lin, Xiaoliang Fu, Cong Qin, Haolin Shi, and Chang Liu. 2026. Proximity-based multi-turn optimization: Practical credit assignment for llm agent training. InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026). 285–307

  120. [130]

    Xiaoliang Fu, Jiaye Lin, Yangyi Fang, Binbin Zheng, Chaowen Hu, Zekai Shao, Cong Qin, Lu Pan, Ke Zeng, and Xunliang Cai. 2026. Maspo: Unifying gradient utilization, probability mass, and signal reliability for robust and sample-efficient llm reasoning. InProceedings of the 64t...

  121. [131]

    Liu Yu, Zhonghao Chen, Ping Kuang, Zhikun Feng, Fan Zhou, Lan Wang, and Gillian Dobbie. 2026. Causally-Grounded Dual-Path Attention Intervention for Object Hallucination Mitigation in LVLMs. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 40. 36021–36029

  122. [132]

    Liu Yu, Jiajun Sun, Ping Kuang, Rui Zhou, Fan Zhou, and Zhikun Feng. 2025. Bimodal Debiasing for Text-to-Image Diffusion: Adaptive Guidance in Textual and Visual Spaces. InProceedings of the 33rd ACM International Conference on Multimedia. 11249–11258

  123. [133]

    Yunyao Zhang, Xinglang Zhang, Junxi Sheng, Wenbing Li, Junqing Yu, Yi- Ping Phoebe Chen, Wei Yang, and Zikai Song. 2026. Semantic-Aware Logical Reasoning via a Semiotic Framework. arXiv:2509.24765 [cs.AI]

  124. [134]

    Jicheol Park, Dongwon Kim, Boseung Jeong, and Suha Kwak. 2024. Plot: Text- based person search with part slot attention for corresponding part discovery. InEuropean Conference on Computer Vision. Springer, 474–490

  125. [135]

    Fan Zhou, Yuzhou Mao, Liu Yu, Yi Yang, and Ting Zhong. 2023. Causal-debias: Unifying debiasing in pretrained language models and fine-tuning via causal invariant learning. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long...

  126. [136]

    Steven Bird. 2006. NLTK: the natural language toolkit. InProceedings of the COLING/ACL 2006 interactive presentation sessions. 69–72

  127. [137]

    Matthew Honnibal. 2017. spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing.(No Title) (2017)

  128. [138]

    Jiaye Lin, Yifu Guo, Yuzhen Han, Sen Hu, Ziyi Ni, Licheng Wang, Ming- guang Chen, Hongzhang Liu, Ronghao Chen, Yangfan He, Daxin Jiang, Binx- ing Jiao, Chen Hu, and Huacan Wang. 2025. SE-Agent: Self-Evolution Trajectory Optimization in Multi-Step Reasoning with LLM-Based Agent...

  129. [139]

    Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, and Yoav Goldberg

  130. [140]

    2008.Optimization algo- rithms on matrix manifolds

    P-A Absil, Robert Mahony, and Rodolphe Sepulchre. 2008.Optimization algo- rithms on matrix manifolds. Princeton University Press

  131. [141]

    Jiaye Lin, Mengdi Li, Xufeng Zhao, Wenhao Lu, Peilin Zhao, Stefan Wermter, and Di Wang. 2026. Curriculum-RLAIF: Curriculum Alignment with Reinforcement Learning from AI Feedback. arXiv:2505.20075 [cs.AI] https://arxiv.org/abs/2505. 20075

  132. [142]

    Shuang Li, Tong Xiao, Hongsheng Li, Bolei Zhou, Dayu Yue, and Xiaogang Wang

  133. [143]

    Rong Fu and Simon Fong. 2025. Adaptive Multi-Backbone Fusion for UAV- Centric Cross-View Geo-Localization with Partial Street–Satellite Matching. In Proceedings of the 3rd International Workshop on UA Vs in Multimedia: Capturing the World from a New Perspective. 31–36

  134. [144]

    Rong Fu, Yibo Meng, Jia Yee Tan, Jiaxuan Lu, Rui Lu, Jiekai Wu, Zhaolu Kang, and Simon Fong. 2026. CityGuard: Graph-Aware Private Descriptors for Bias- Resilient Identity Search Across Urban Cameras.arXiv preprint arXiv:2602.18047 (2026)

  135. [145]

    Jialong Zuo, Hanyu Zhou, Ying Nie, Feng Zhang, Tianyu Guo, Nong Sang, Yunhe Wang, and Changxin Gao. 2024. Ufinebench: Towards text-based person retrieval with ultra-fine granularity. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 22010–22019

  136. [146]

    Research Team. 2024. TIPS: A Text-Image Pairs Synthesis Framework for Robust Text-based Person Retrieval.OpenReview(2024)

  137. [147]

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sand- hini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al

  138. [148]

    Hang Yu, Jiahao Wen, and Zhedong Zheng. 2025. CAMeL: Cross-modality Adaptive Meta-Learning for Text-based Person Retrieval.IEEE Transactions on Information Forensics and Security(2025)

  139. [149]

    Shuyu Yang, Yinan Zhou, Zhedong Zheng, Yaxiong Wang, Li Zhu, and Yujiao Wu

  140. [150]

    Shuyu Yang, Yaxiong Wang, Li Zhu, and Zhedong Zheng. 2025. Beyond walking: A large-scale image-text benchmark for text-based person anomaly search. In ICCV. 11720–11730

  141. [151]

    Francesco Locatello, Stefan Bauer, Mario Lucic, Gunnar Raetsch, Sylvain Gelly, Bernhard Schölkopf, and Olivier Bachem. 2019. Challenging common assump- tions in the unsupervised learning of disentangled representations. Ininterna- tional conference on machine learning. PMLR, 4114–4124

  142. [152]

    Chenyang Gao, Guanyu Cai, Xinyang Jiang, Feng Zheng, Jun Zhang, Yifei Gong, Pai Peng, Xiaowei Guo, and Xing Sun. 2021. Contextual non-local alignment over full-scale representation for text-based person search.arXiv preprint arXiv:2101.03036(2021)

  143. [153]

    Yucheng Chen, Rui Huang, Hong Chang, Chuanqi Tan, Tao Xue, and Bingpeng Ma. 2021. Cross-modal knowledge adaptation for language-based person search. IEEE Transactions on Image Processing30 (2021), 4057–4069

  144. [154]

    Yushuang Wu, Zizheng Yan, Xiaoguang Han, Guanbin Li, Changqing Zou, and Shuguang Cui. 2021. Lapscore: language-guided person search via color reasoning. InProceedings of the IEEE/CVF International Conference on Computer Vision. 1624–1633

  145. [156]

    Zhiyin Shao, Xinyu Zhang, Meng Fang, Zhifeng Lin, Jian Wang, and Changxing Ding. 2022. Learning granularity-unified representations for text-to-image person re-identification. InProceedings of the 30th acm international conference on multimedia. 5566–5574

  146. [157]

    Person search with natural language description. InCVPR. 1970–1979

  147. [159]

    Aichun Zhu, Zijie Wang, Yifeng Li, Xili Wan, Jing Jin, Tian Wang, Fangqiang Hu, and Gang Hua. 2021. Dssl: Deep surroundings-person separation learning for text-based person retrieval. InProceedings of the 29th ACM international conference on multimedia. 209–217

  148. [161]

    Yan Zeng, Xinsong Zhang, Hang Li, Jiawei Wang, Jipeng Zhang, and Wangchun- shu Zhou. 2023. X 2-VLM: All-in-One Pre-Trained Model for Vision-Language Tasks.IEEE transactions on pattern analysis and machine intelligence46, 5 (2023), 3156–3168

  149. [162]

    Dixuan Lin, Yi-Xing Peng, Jingke Meng, and Wei-Shi Zheng. 2024. Cross-modal adaptive dual association for text-to-image person retrieval.IEEE Transactions on Multimedia26 (2024), 6609–6620

  150. [163]

    Zhiwei Zhao, Bin Liu, Yan Lu, Qi Chu, and Nenghai Yu. 2024. Unifying multi- modal uncertainty modeling and semantic alignment for text-to-image person re-identification. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 7534–7542

  151. [164]

    Yang Bai, Min Cao, Daming Gao, Ziqiang Cao, Chen Chen, Zhenfeng Fan, Liqiang Nie, and Min Zhang. 2023. Rasa: Relation and sensitivity aware representation learning for text-based person search.arXiv preprint arXiv:2305.13653(2023)

  152. [165]

    Shuanglin Yan, Jun Liu, Neng Dong, Liyan Zhang, and Jinhui Tang. 2024. Proto- typical prompting for text-to-image person re-identification. InProceedings of the 32nd ACM International Conference on Multimedia. 2331–2340

  153. [166]

    InProceedings of the 31st ACM international conference on multimedia

    Towards unified text-based person retrieval: A large-scale multi-attribute and language search benchmark. InProceedings of the 31st ACM international conference on multimedia. 4492–4501

  154. [168]

    Jintao Sun, Hao Fei, Gangyi Ding, and Zhedong Zheng. 2025. From data deluge to data curation: A filtering-wora paradigm for efficient text-based person search. InProceedings of the ACM on Web Conference 2025. 2341–2351

  155. [172]

    Zefeng Ding, Changxing Ding, Zhiyin Shao, and Dacheng Tao. 2021. Semanti- cally self-aligned network for text-to-image part-aware person re-identification. arXiv preprint arXiv:2107.12666(2021)

  156. [174]

    Ammarah Farooq, Muhammad Awais, Josef Kittler, and Syed Safwan Khalid

  157. [175]

    InProceedings of the AAAI conference on artificial intelligence, Vol

    Axm-net: Implicit cross-modal feature alignment for person re- identification. InProceedings of the AAAI conference on artificial intelligence, Vol. 36. 4477–4485

  158. [176]

    Shiping Li, Min Cao, and Min Zhang. 2022. Learning semantic-aligned feature representation for text-based person search. InICASSP 2022-2022 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2724–2728

  159. [177]

    Shuanglin Yan, Hao Tang, Liyan Zhang, and Jinhui Tang. 2023. Image-specific information suppression and implicit local alignment for text-based person search.IEEE transactions on neural networks and learning systems(2023)

  160. [178]

    Shuanglin Yan, Neng Dong, Liyan Zhang, and Jinhui Tang. 2023. Clip-driven fine-grained text-image person re-identification.IEEE Transactions on Image Processing32 (2023), 6032–6046

  161. [179]

    Takuro Fujii and Shuhei Tarashima. 2023. Bilma: Bidirectional local-matching for text-based person re-identification. InProceedings of the IEEE/CVF International Conference on Computer Vision. 2786–2790

  162. [182]

    Di Wang, Feng Yan, Yifeng Wang, Lin Zhao, Xiao Liang, Haodi Zhong, and Ronghua Zhang. 2024. Fine-grained semantics-aware representation learning for text-based person retrieval. InProceedings of the 2024 International Conference on Multimedia Retrieval. 92–100

  163. [184]

    Yang Qin, Yingke Chen, Dezhong Peng, Xi Peng, Joey Tianyi Zhou, and Peng Hu

  164. [185]

    LightAIR: Lightweight Action Inversion and Riemannian Rectification for Text-based Person Anomaly Search

    Noisy-correspondence learning for text-to-image person re-identification. InCVPR. 27197–27206. LightAIR: Lightweight Action Inversion and Riemannian Rectification for Text-based Person Anomaly Search MM ’26, November 10–14, 2026, Rio de Janeiro, Brazil Supplementary Material f...

  165. [2017]

    Image re-ranking based on topic diversity.IEEE Transactions on Image Processing26, 8 (2017), 3734–3747

  166. [2020]

    InProceedings of the 58th annual meeting of the association for computational linguistics

    Null it out: Guarding protected attributes by iterative nullspace projection. InProceedings of the 58th annual meeting of the association for computational linguistics. 7237–7256

  167. [2021]

    InInternational conference on machine learning

    Learning transferable visual models from natural language supervision. InInternational conference on machine learning. PMLR, 8748–8763

  168. [2022]

    TIPCB: A simple but effective part-based convolutional baseline for text-based person search.Neurocomputing494 (2022), 171–181

  169. [2023]

    Cross-modal active complementary learning with self-refining correspon- dence.NeurIPS36 (2023), 24829–24840

  170. [2024]

    InInternational conference on artificial intelligence and statistics

    Identifying spurious biases early in training through the lens of simplicity bias. InInternational conference on artificial intelligence and statistics. PMLR, 2953–2961

  171. [2025]

    InProceedings of the AAAI Conference on Artificial Intel- ligence

    ENCODER: Entity Mining and Modification Relation Binding for Com- posed Image Retrieval. InProceedings of the AAAI Conference on Artificial Intel- ligence. MM ’26, November 10–14, 2026, Rio de Janeiro, Brazil Yulun Zhang, Zixu Li, Zhiwei Chen, Zhiheng Fu, Wenbo Wang, Zihang Qi...

  172. [2026]

    doi:10.1109/TGRS.2026.3710063

    GPR-MVS: Global Propagation Regularization for Large Scale Multi- view Stereo.IEEE Transactions on Geoscience and Remote Sensing(2026), 1–1. doi:10.1109/TGRS.2026.3710063

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.