REVIEW 6 major objections 5 minor 37 references
HCMRM: A High-Consistency Multimodal Relevance Model for Search Ads
T0 review · 6 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read Pre-training a search-ad relevance model on pseudo-queries cut from a video's own keywords, then fine-tuning with a hierarchical softmax loss, reduced irrelevant ads by 6.1% and raised ad revenue by 1.4% in the deployed system.
desk verdict A useful industrial contribution with a clever pseudo-query pre-training task and a sensible ordinal loss, backed by a year of deployment; the main weakness is that the offline gains are small and the synthetic-to-real query gap is only partially addressed. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Two mechanisms carry the argument. PQVM: during pre-training, the video keyword sequence is split at a random index into a pseudo-query and the remaining video text; positive and hard-negative triplet pairs are trained with a binary cross-entropy classifier on the fusion encoder's [CLS] embedding, making pre-training structurally identical to the downstream query-video relevance task. Symmetric hierarchical softmax: the four relevance labels form a two-level binary tree, so the label distribution factors into three sigmoid probabilities ($p_{pos}$, $p_{le}$, $p_{ex}$); the expected label value (Equation 9) is the relevance score used for ranking, and the loss (Equation 10) combines the hierarchical softmax likelihood with an MSE regression term. The architecture itself is reused from ALBEF with minimal modification: a text encoder, a vision encoder, and a fusion encoder with cross-attention.
What would settle it
A controlled experiment where PQVM is replaced with real user queries from the same platform during pre-training, holding the model and all other objectives fixed: if the real-query variant matches or exceeds HCMRM's offline AUC and Spearman on the 230K-sample evaluation set, the paper's claim that pseudo-queries are sufficient—and better than click-derived queries—would be called into question. Alternatively, removing PQVM from HCMRM while keeping the hierarchical softmax loss and all other settings identical should measurably lower Spearman on the same evaluation set if the mechanism carries the reported ranking gain.
Extended reading notes
Core claim
HCMRM is an ALBEF-style dual-stream vision-language model—a text encoder, an image encoder, and a fusion encoder—pre-trained with four objectives. Three are standard (image-text contrastive learning, image-text matching, masked language modeling); the fourth, PQVM, is the paper's contribution. For each video, its keyword sequence is split at a random length in [a,b]: the left part becomes the pseudo-query and the right part remains the video text, with the full keyword sequence excluded to avoid information leakage. The model must classify whether a pseudo-query matches a video, using hard negatives sampled from the in-batch similarity distribution. For fine-tuning, the four relevance labels are organized as a balanced binary tree ('Is Positive?', then 'Is Less?' among negatives and 'Is Excellent?' among positives), producing three binary classifiers whose probability products give the label distribution and an expected relevance score in [0,3]. The paper reports that this architecture outperforms unimodal and multimodal baselines and both Qwen-VL and MiniCPM-V-2.6 on offline AUC, Spearman, and Pearson metrics, and beats the deployed BERT-VL online.
Load-bearing premise
The approach assumes that splitting a video's own keyword sequence at a random position produces a pseudo-query that behaves like a real user query for learning relevance; if real queries systematically target content outside the top keywords, the PQVM pre-training could teach the model the wrong kind of query-awareness.
Editorial extensions
If this is right
- Relevance models for search ads can be made query-aware without any query logs, using only video self-text.
- The PQVM pre-training task transfers to other base architectures: BERT, BERT-VL, and ViLT all improve when PQVM is added (Table 1).
- The hierarchical softmax loss improves Spearman and Pearson rank correlation substantially over binary cross-entropy, aligning relevance scores better with the downstream ad-ranking objective.
- Multimodal large language models fine-tuned on the same data do not beat the domain-pre-trained HCMRM, suggesting that domain-specific pre-training still matters.
- The deployed impact is a 6.1% relative reduction in irrelevant ads and a 1.4% ad revenue increase.
Reading between the lines
- Inference: The pseudo-query synthesis assumes the head of the keyword list is a faithful proxy for user intent; this is plausible for title-like texts but may fail for videos where the query targets an aspect not captured by top keywords, and varying the split range or weighting keywords by field type could be a testable refinement.
- Inference: The same synthetic-query pre-training idea could transfer to other retrieval domains with rich item-side text, such as product search or news video search, wherever a query is a short text about the item.
- Inference: The symmetric hierarchical softmax loss is a general ordinal-classification technique that could improve other graded-relevance ranking systems beyond ads, such as answer ranking or recommendation explanations.
- Inference: The paper's comparison against click-based pre-training (Table 5) suggests that synthetic queries from item text may be a cheaper and less noisy alternative to behavior logs; verifying this on a public dataset would be a direct next step.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes HCMRM, an ALBEF-based multimodal relevance model for query-to-video matching in Kuaishou's search advertising system. The two main contributions are (i) a pre-training task called Pseudo-Query-Video Matching (PQVM), which synthesizes pseudo-queries by splitting a video's importance-ranked keyword sequence into a query prefix and a remaining text suffix, and (ii) a symmetric hierarchical softmax fine-tuning loss, defined in Equation 10, that combines three binary classifiers with an MSE term. Offline experiments on 2.5M labeled query-video pairs report improvements in AUC, Spearman, and Pearson over BERT, BERT-VL, ViLT, Query-LIFE, ALBEF, Qwen-VL, and MiniCPM-V-2.6. Online A/B tests report reductions in the irrelevant-ad ratio and increases in conversions and ad revenue, and the paper states the method has been deployed for over a year.
Significance. If the reported gains are reliable, this is a valuable industrial case study: it shows that a lightweight modification to ALBEF can improve query-video relevance without relying on noisy click logs, and it demonstrates a practical ordinal relevance loss for ranking. The paper's strengths include the scale of the pre-training data (200M videos), the 2.5M labeled query-video fine-tuning set, the breadth of baselines including multimodal LLMs, and the evidence from a real deployed advertising system. However, the incremental offline gains over ALBEF are small (0.003 AUC, 0.005 Spearman), no confidence intervals or significance tests are reported, the PQVM pre-training and downstream QVM input distributions differ in a way that weakens the 'high consistency' claim, and the abstract's aggregate online numbers appear to be obtained by summing two separate A/B tests. These issues currently leave the attribution of the reported gains to the two proposed components less secure than the presentation suggests.
major comments (6)
- [Section 3.4, Eq. (2); Section 3.5] The PQVM pre-training task feeds [CLS] pseudoQ_i [SEP] rem.T_i to the text encoder, while the downstream QVM task uses the full video keyword sequence T. Thus the model is never pre-trained on the exact downstream input configuration (real query plus full video text). Moreover, because pseudoQ_i and rem.T_i are contiguous pieces of the same original keyword sequence, positive PQVM pairs can be solved by lexical/structural continuity between the prefix and suffix rather than by a general query-video relevance function. The comparison in Section 4.7.1 against click-derived queries does not resolve this, since it does not test with clean real-user queries or quantify lexical overlap. Please report how often pseudo-query terms appear in rem.T, compare with a variant that uses the full T on the video side during PQVM, and consider a pseudo-query generation mechanism that is not a prefix of the same video's own text.
- [Section 3.5, Eq. (10); Table 2] The proposed loss is the sum of a hierarchical softmax negative log-likelihood and an MSE term (r - l)^2. The paper never ablates these two terms. Table 2 compares 'hierarchical softmax' (i.e., the combined loss) against BCE, MSE, and ordinal regression, so the reported gains could be driven partly or wholly by the MSE regression component rather than by the hierarchical structure. Please report hierarchical softmax NLL alone, the MSE term alone, and the combination, and justify why both are needed.
- [Section 4.5, Tables 1 and 5] No confidence intervals, standard deviations, or significance tests are reported for any offline metric. The headline improvement over ALBEF is 0.003 in AUC and 0.005 in Spearman, which is the same order as typical run-to-run variation for large-scale fine-tuning. Please provide results over multiple random seeds (or a paired significance test on the evaluation set) for at least the ALBEF vs HCMRM comparison and the loss-function comparison, so the reader can judge whether the proposed components improve over the baselines beyond noise.
- [Abstract; Section 4.6, Tables 3 and 4] The abstract and conclusion state a 6.1% reduction in the irrelevant-ad ratio and a 1.4% increase in ad revenue, but the online tables report separate A/B experiments: -4.0%/+1.0% for the model change versus BERT-VL and -2.1%/+0.4% for the loss change versus BCE. Simply adding percentages from independent tests is not a valid estimate of the combined effect. Please report a single combined A/B test, or clearly state how the aggregate figures were computed and whether the two effects were actually measured jointly.
- [Section 4.6.1; Appendix A] Appendix A says the model deployed online is a lightweight BERT-VL-based student distilled from the offline HCMRM teacher, yet Section 4.6.1 states the A/B test compares HCMRM with BERT-VL. Please clarify what exactly differs between the online experimental buckets: is the online comparison between students distilled from HCMRM and BERT-VL teachers, or between the full teachers? If the former, the online results validate the distillation pipeline and the teacher's scores, not HCMRM itself, and this should be stated explicitly.
- [Section 4.7.1, Table 5] The comparison intended to show that pseudo-queries 'surpass real queries' confounds data source with data size: row 2 uses 60M click-derived query-video pairs, while row 4 uses 200M video materials. The two-stage row 3 partially addresses this, but the total amount of pre-training data still differs. Please add a matched comparison, e.g., ALBEF pre-trained on a 60M subset of the material data with a QVM task on the click data, or HCMRM with pseudo-queries on the 60M click-derived video subset, so the source of the gain is identifiable.
minor comments (5)
- [Section 3.4, Eq. (2)] The range [a, b] for the pseudo-query length is never specified in the paper; please provide the value used in the experiments and, ideally, a small sensitivity analysis.
- [Section 3.2, Algorithm 1] The field importance weights alpha_i are defined but their values are not reported; please state the weights used for title, description, OCR, ASR, and the other fields.
- [Section 4.2] The AUC metric is computed on binary labels derived by collapsing 0/1 and 2/3, while the model outputs a continuous score in [0, 3]; please clarify exactly how AUC is computed over the evaluation set.
- [Section 3.5, Figure 3] The description 'balanced binary tree' is informal; Equation 8 does follow from the tree, but the sentence should say 'a two-level hierarchical decomposition' rather than implying a standard binary search tree over labels.
- [Section 4.7.2, Table 6] The input formatting for Qwen-VL and MiniCPM-V-2.6 is not described; since these models have different tokenizers and image processors, please specify how the query, video text, and video frames were packaged for fine-tuning.
Circularity Check
No significant circularity: PQVM is an auxiliary pre-training task whose benefits are measured on held-out human labels and online A/B tests, not recovered by construction.
full rationale
The central claim is that adding a pseudo-query-video matching (PQVM) pre-training task and a hierarchical softmax fine-tuning loss improves query-video relevance. The pseudo-query is constructed by splitting a video's own keyword sequence (Eq. 2), but the downstream QVM task uses real user queries and full video text with human-annotated four-level labels (Eq. 10). There is no fitted parameter or label information from the evaluation set that is fed back into the pre-training loss; offline metrics are computed on a separate 230K held-out sample, and the online gains are measured in A/B tests against a deployed baseline. The hierarchical softmax loss is a deterministic reparameterization of the four label probabilities using three sigmoid outputs (Eqs. 7-9), introducing no constants fitted to the downstream data. The assumption that a pseudo-query behaves like a real user query is an explicit modeling assumption, and the paper provides a direct empirical comparison against click-derived real queries in Section 4.7.1. References to 'our previous relevance model' are descriptive of the industrial baseline rather than load-bearing citations. No equation in the paper reduces by construction to its own inputs, and no claimed prediction is equivalent to a fitted input. The paper is therefore self-contained against external benchmarks and should not be scored as circular.
Assumptions & free parameters
free parameters (3)
- Field importance weights alpha_i =
not specified
- Pseudo-query length range [a, b] =
not specified
- MSE term weight in L_QVM =
1 (implicit)
assumptions (3)
- domain assumption A query should represent the key content of a short video, so the highest-importance keywords of the video text are a valid stand-in for real user queries.
- domain assumption The four relevance levels form a balanced binary tree with conditionally independent binary questions (Is Positive, Is Less, Is Excellent), so the probability factorization in Equation 8 is appropriate.
- domain assumption Knowledge distillation from the teacher HCMRM to the lightweight BERT-VL student preserves the relative relevance ordering well enough that online metrics reflect the teacher's quality.
Cite this review
Pith. "Pith review of HCMRM: A High-Consistency Multimodal Relevance Model for Search Ads." pith.science (2026). https://pith.science/paper/DQ67D6MX
@misc{pith2026250205822,
author = {Pith},
title = {Pith review of: HCMRM: A High-Consistency Multimodal Relevance Model for Search Ads},
year = {2026},
howpublished = {\url{https://pith.science/paper/DQ67D6MX}},
note = {Machine review of arXiv:2502.05822}
}
read the original abstract
Search advertising is essential for merchants to reach the target users on short video platforms. Short video ads aligned with user search intents are displayed through relevance matching and bid ranking mechanisms. This paper focuses on improving query-to-video relevance matching to enhance the effectiveness of ranking in ad systems. Recent vision-language pre-training models have demonstrated promise in various multimodal tasks. However, their contribution to downstream query-video relevance tasks is limited, as the alignment between the pair of visual signals and text differs from the modeling of the triplet of the query, visual signals, and video text. In addition, our previous relevance model provides limited ranking capabilities, largely due to the discrepancy between the binary cross-entropy fine-tuning objective and the ranking objective. To address these limitations, we design a high-consistency multimodal relevance model (HCMRM). It utilizes a simple yet effective method to enhance the consistency between pre-training and relevance tasks. Specifically, during the pre-training phase, along with aligning visual signals and video text, several keywords are extracted from the video text as pseudo-queries to perform the triplet relevance modeling. For the fine-tuning phase, we introduce a hierarchical softmax loss, which enables the model to learn the order within labels while maximizing the distinction between positive and negative samples. This promotes the fusion ranking of relevance and bidding in the subsequent ranking stage. The proposed method has been deployed in the Kuaishou search advertising system for over a year, contributing to a 6.1% reduction in the proportion of irrelevant ads and a 1.4% increase in ad revenue.
Figures
Reference graph
Works this paper leans on
-
[1]
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. 2023. Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond. ArXiv (2023). https://arxiv.org/abs/2308.12966
arXiv 2023
- [2]
-
[3]
Wei-Cheng Chang, Daniel Jiang, Hsiang-Fu Yu, Choon Hui Teo, Jiong Zhang, Kai Zhong, Kedarnath Kolluri, Qie Hu, Nikhil Shandilya, Vyacheslav Ievgrafov, Japinder Singh, and Inderjit S. Dhillon. 2021. Extreme Multi-label Learning for Semantic Matching in Product Search. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining (KD...
-
[4]
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Zhong Muyan, Qinglong Zhang, Xizhou Zhu, Lewei Lu, Bin Li, Ping Luo, Tong Lu, Yu Qiao, and Jifeng Dai. 2023. Intern VL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 24185...
work page 2023
-
[5]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). Association for Comput...
-
[6]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xi- aohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In Interna- tional Conference on Learning Representations . http...
2021
-
[7]
Yuxin Fang, Wen Wang, Binhui Xie, Quan-Sen Sun, Ledell Yu Wu, Xinggang Wang, Tiejun Huang, Xinlong Wang, and Yue Cao. 2023. EVA: Exploring the Limits of Masked Visual Representation Learning at Scale. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE Computer Society, Los Alamitos, CA, USA, 19358–19369. doi:10.1109/CVPR5...
-
[8]
Jingjia Huang, Yinan Li, Jiashi Feng, Xiaoshuai Sun, and Rongrong Ji. 2022. Clover: Towards A Unified Video-Language Alignment and Fusion Model. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2022), 14856–14866. https://cvpr.thecvf.com/virtual/2023/poster/22766
work page 2022
Show all 37 references
-
[9]
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. 2021. Scaling Up Visual and Vision- Language Representation Learning With Noisy Text Supervision. In Proceedings of the 38th International Conference on Mac...
2021
-
[10]
Wonjae Kim, Bokyung Son, and Ildoo Kim. 2021. ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision. In Proceedings of the 38th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 139). PMLR, 5583–5594. https:/...
2021
-
[11]
Chunyuan Li, Zhe Gan, Zhengyuan Yang, Jianwei Yang, Linjie Li, Lijuan Wang, and Jianfeng Gao. 2024. Multimodal Foundation Models: From Specialists to General-Purpose Assistants. Found. Trends Comput. Graph. Vis. 16, 1–2 (2024). https://doi.org/10.1561/0600000110
2024 doi
-
[12]
Feng Li, Renrui Zhang, Hao Zhang, Yuanhan Zhang, Bo Li, Wei Li, Zejun Ma, and Chunyuan Li. 2025. LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models. In The Thirteenth International Conference on Learning Representations. https://openreview.ne...
2025
-
[13]
Junnan Li, Dongxu Li, Silvio Savarese, and Steven C. H. Hoi. 2023. BLIP- 2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models. In International Conference on Machine Learn- ing (Proceedings of Machine Learning Research, Vol. 202) . ...
2023
-
[14]
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation. In Proceedings of the 39th International Conference on Machine Learn- ing (Proceedings of Machine Learning Resea...
2022
-
[15]
Selvaraju, Akhilesh Deepak Gotmare, Shafiq R
Junnan Li, Ramprasaath R. Selvaraju, Akhilesh Deepak Gotmare, Shafiq R. Joty, Caiming Xiong, and Steven C. H. Hoi. 2021. Align before Fuse: Vision and Language Representation Learning with Momentum Distillation. In Advances in Neural Information Processing Systems . https://op...
2021
-
[16]
Yangguang Li, Feng Liang, Lichen Zhao, Yufeng Cui, Wanli Ouyang, Jing Shao, Fengwei Yu, and Junjie Yan. 2022. Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm. In International Conference on Learning Representations. https://open...
2022
-
[17]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual In- struction Tuning. In Thirty-seventh Conference on Neural Information Processing Systems, Vol. 36. Curran Associates, Inc., 34892–34916. https://openreview.net/ forum?id=w0H2xGHlkw
2023
-
[19]
Huaishao Luo, Lei Ji, Ming Zhong, Yang Chen, Wen Lei, Nan Duan, and Tianrui Li. 2022. CLIP4Clip: An empirical study of CLIP for end to end video clip retrieval and captioning. Neurocomput. 508, C (Oct. 2022), 293–304. doi:10.1016/j.neucom. 2022.07.028
2022 doi
-
[20]
Yiwei Ma, Guohai Xu, Xiaoshuai Sun, Ming Yan, Ji Zhang, and Rongrong Ji. 2022. X-CLIP: End-to-End Multi-grained Contrastive Learning for Video-Text Retrieval. In Proceedings of the 30th ACM International Conference on Multimedia (MM ’22) . Association for Computing Machinery, ...
2022
-
[21]
Wagner, and Saining Xie
Norman Mu, Alexander Kirillov, David A. Wagner, and Saining Xie. 2022. SLIP: Self-supervision Meets Language-Image Pre-training. In Computer Vision - ECCV 2022 - 17th European Conference, Tel A viv, Israel, October 23-27, 2022, Proceedings, Part XXVI (Lecture Notes in Computer...
2022 doi
-
[22]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From Natural Language Supervision. In Proceedings...
2021
-
[23]
Howard, Menglong Zhu, Andrey Zhmoginov, and Liang- Chieh Chen
Mark Sandler, Andrew G. Howard, Menglong Zhu, Andrey Zhmoginov, and Liang- Chieh Chen. 2018. MobileNetV2: Inverted Residuals and Linear Bottlenecks. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition . IEEE Computer Society, 4510–4520. doi:10.1109/CVPR.2018.00474
2018
-
[24]
Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, Wo- jciech Galuba, Marcus Rohrbach, and Douwe Kiela. 2021. FLAVA: A Foun- dational Language And Vision Alignment Model. 2022 IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR) (2021), 15617...
2021
-
[25]
Wesley Tansey, Karl Pichotta, and James G. Scott. 2018. Leaf-Smoothed Hierarchi- cal Softmax for Ordinal Prediction. In AAAI Conference on Artificial Intelligence , Vol. 32. doi:10.1609/aaai.v32i1.11754
2018 doi
-
[26]
Zirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai, Yulia Tsvetkov, and Yuan Cao. 2022. SimVLM: Simple Visual Language Model Pretraining with Weak Supervision. In International Conference on Learning Representations . https: //openreview.net/forum?id=GUrhfTuf_3
2022
-
[27]
Zhoufutu Wen, Xinyu Zhao, Zhipeng Jin, Yi Yang, Wei Jia, Xiaodong Chen, Shuanglong Li, and Lin Liu. 2023. Enhancing Dynamic Image Advertising with Vision-Language Pre-training. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Informa...
2023
-
[28]
Lin Xiao, Xiaofeng Li, and Yucheng Zhang. 2023. Exploring the factors influencing consumer engagement behavior regarding short-form video advertising: A big data perspective. Journal of Retailing and Consumer Services 70 (2023), 103170. doi:10.1016/j.jretconser.2022.103170
2023
-
[29]
Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer. 2021. VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding. In Conference on Empirical Methods in Natural Language Processing...
2021 doi
-
[30]
Haiyang Xu, Qinghao Ye, Mingshi Yan, Yaya Shi, Jiabo Ye, Yuanhong Xu, Chen- liang Li, Bin Bi, Qiuchen Qian, Wei Wang, Guohai Xu, Ji Zhang, Songfang Huang, Feiran Huang, and Jingren Zhou. 2023. mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video. In...
2023
-
[31]
Shaowei Yao, Jiwei Tan, Xi Chen, Keping Yang, Rong Xiao, Hongbo Deng, and Xiaojun Wan. 2021. Learning a Product Relevance Model from Click-Through Data in E-Commerce. In Proceedings of the Web Conference 2021 (WWW ’21) . Association for Computing Machinery, 2890–2899. doi:10.1...
2021
-
[32]
Yuan Yao, Tianyu Yu, Ao Zhang, Chongyi Wang, Junbo Cui, Hongji Zhu, Tianchi Cai, Haoyu Li, Weilin Zhao, Zhihui He, Qi-An Chen, Huarong Zhou, Zhensheng Zou, Haoye Zhang, Shengding Hu, Zhi Zheng, Jie Zhou, Jie Cai, Xu Han, Guoyang Zeng, Dahai Li, Zhiyuan Liu, and Maosong Sun. 20...
2024 arXiv
-
[33]
Chengcan Ye, Ting Peng, Tim Chang, Zhiyi Zhou, and Feng Wang. 2023. Query- aware Multi-modal based Ranking Relevance in Video Search. In Conference on WWW Companion ’25, April 28-May 2, 2025, Sydney, NSW, Australia Guobing Gan, Kaiming Gao, Li Wang, Shen Jiang, and Peng Jiang ...
2023 doi
-
[34]
Shukang Yin, Chaoyou Fu, Sirui Zhao, Ke Li, Xing Sun, Tong Xu, and En- hong Chen. 2023. A Survey on Multimodal Large Language Models. ArXiv abs/2306.13549 (2023). https://arxiv.org/pdf/2306.13549
2023 arXiv
-
[35]
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu. 2022. CoCa: Contrastive Captioners are Image-Text Foundation Models. Transactions on Machine Learning Research (2022). https://openreview. net/forum?id=Ee277P3AYC
2022
-
[36]
Duzhen Zhang, Yahan Yu, Jiahua Dong, Chenxing Li, Dan Su, Chenhui Chu, and Dong Yu. 2024. MM-LLMs: Recent Advances in MultiModal Large Language Models. In Findings of the Association for Computational Linguistics: ACL 2024 . Association for Computational Linguistics, Bangkok, ...
2024 doi
-
[37]
Hai Zhu, Yuankai Guo, Ronggang Dou, and Kai Liu. 2025. Query-LIFE: Query- aware Language Image Fusion Embedding for E-Commerce Relevance. In Inter- national Conference on Computational Linguistics . Association for Computational Linguistics. https://aclanthology.org/2025.colin...
2025
-
[38]
Lixin Zou, Shengqiang Zhang, Hengyi Cai, Dehong Ma, Suqi Cheng, Shuaiqiang Wang, Daiting Shi, Zhicong Cheng, and Dawei Yin. 2021. Pre-trained Language Model based Ranking in Baidu Search. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining (KD...
2021
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.