REVIEW 3 major objections 4 minor 81 references
Generative Universal Multimodal Retrieval with Dual-role Identifiers
T0 review · 3 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Reusing one identifier as both a sequence and a set lifts generative retrieval accuracy.
desk verdict DrIG's dual-role identifier idea is new and the empirical gains over GENIUS are consistent and substantial, but the set-based prior in Equation 15 is computationally underspecified and the efficiency claim needs an approximation or an algorithm before it stands. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the dual-role identifier: a length-$L$ residual-quantized code $m = [m_1,\dots,m_L]$ whose first codebook has size 3 for modality and whose later codebooks encode increasingly fine semantics. In its sequential role the identifier is decoded under a Trie constraint with a score that sums token-embedding dot products; in its set-based role the same tokens are mapped through a fixed global codebook table to token-level query scores, and a candidate-level order-invariant score is computed as a sum over the identifier's tokens. The mechanism that carries the argument is the combined expansion score $f(t_{\le i}; z_q) = \delta(t_{\le i}) + \eta(t_{<i}) + E^{dec}_i[t_i]\cdot h_i + \lambda \phi(t_{\le i})$, where $\phi(t_{\le i})$ is the maximum order-invariant score over all candidates sharing the current prefix. This term injects global, prefix-independent relevance into every beam expansion, so a branch that the local decoder underrates can survive if its reachable candidates match the query as a set.
What would settle it
Train DrIG with the set-based prior computed from the additive token score replaced by a mean or a learned attention pooling, and measure Recall@K on M-BEIR; if performance does not drop, then the additive token-level prior is not the active ingredient behind the reported gains.
Extended reading notes
Core claim
The central claim is that a single residual-quantized identifier can simultaneously serve as an ordered generative target and as an unordered set-based relevance representation, and that this dual role is what lets constrained beam search avoid prefix-pruning errors. Concretely, DrIG encodes each candidate with a first token that names the modality (image, text, or image-text) followed by residual-quantized semantic tokens; during decoding the same codebook tokens are scored independently against the query via an order-invariant score, and the best score among all candidates sharing a prefix is added, weighted by a hyperparameter, to the autoregressive score. The authors report that this consistently improves over the GENIUS baseline across all M-BEIR task types, with particularly large gains on knowledge-intensive InfoSeek and compositional CIRR, and that the improvement is complementary to dense reranking. When the top-k generated candidates are reranked by continuous cosine similarity, the method moves from generative-only averages of 38.0 and 36.4 to 50.4 and 48.9, surpassing CLIP-SF and BLIP-FF in the global-pool setting while remaining below stronger LMM-based rerankers.
Load-bearing premise
The load-bearing premise is that a candidate's relevance can be estimated by adding up independent per-token scores, and that the best-scoring candidate under a prefix reveals the value of keeping that prefix.
Editorial extensions
If this is right
- On M-BEIR, DrIG outperforms the previous generative universal multimodal retriever on every task type, raising average Recall from 29.5 to 38.0 in local-pool retrieval and from 28.6 to 36.4 in global-pool retrieval.
- Reranking the top-k generated candidates with dense cosine similarity brings the global-pool average to 47.1 and 48.9 for the CLIP-based and LMM-based reranking variants, surpassing the CLIP-SF and BLIP-FF dense baselines while remaining below stronger LMM-based rerankers.
- The set-based prior is what carries the prefix-independence claim: removing it costs one to two Recall points across tasks, while removing the Trie constraint costs far more, showing that valid-prefix control and global relevance guidance are complementary.
- Throughput stays nearly flat as the candidate pool grows from 5K to 300K, because the online cost is dominated by fixed-length identifier decoding and a small top-k reranking step rather than by scoring the full corpus.
- Identifier expressiveness, not decoder size, is the main lever on knowledge-intensive tasks: deeper quantization and larger codebooks improve text-centric Recall@1 substantially, while scaling the decoder from T5-small to T5-large gives only mixed gains and hurts visually fine-grained tasks.
Reading between the lines
- If the dual-role reuse is the active ingredient, the same one-identifier-two-views design should transfer to other structured decoding tasks where the output vocabulary is built from quantized embeddings, such as hierarchical product or entity search; the paper leaves this implicit.
- The additive token score is the paper's simplest aggregation choice; replacing it with a learned or attention-weighted pooling over the set tokens would be a direct test of whether token independence is a limitation or a feature, and could improve hard cases like NIGHTS.
- The flat throughput curve suggests that generative retrieval with dual-role identifiers could support dynamic corpora more gracefully than dense indexes, since inserting a candidate only requires adding its code tokens to the Trie and the global codebook table; the paper lists dynamic corpora as future work, not as a demonstrated capability.
- Reading the ablations as a whole, the remaining gap on text-to-text and knowledge tasks points to quantization depth as the likely bottleneck rather than reranker strength; this is an interpretation, not a claim the paper makes.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes DrIG, a generative framework for universal multimodal retrieval. Each candidate is assigned a single residual-quantized identifier whose first token encodes modality and whose remaining tokens encode progressively finer semantics. The same identifier is used in two roles: a sequential role for Trie-constrained autoregressive decoding, and a set-based role in which the tokens are treated as an unordered set to produce a prefix-independent relevance prior. During inference, the two scores are combined in Equation (16), and an optional dense reranker refines the top-k generated candidates. Experiments on M-BEIR, Flickr30K, and MSCOCO compare DrIG against dense baselines, GENIUS, and text-to-image generative baselines, and include ablations over the set-based role, Trie constraint, query augmentation, ranking loss, contrastive loss, modality codebook, codebook configuration, beam size, reranking depth, and decoder backbone.
Significance. If the reported results hold, DrIG is a meaningful advance for generative universal multimodal retrieval. The paper evaluates on a broad and challenging benchmark (M-BEIR), shows consistent improvements over the closest generative baseline GENIUS, and provides unusually thorough ablations and hyperparameter analyses that give practical design guidance. The hybrid generative-plus-dense reranking strategy is a useful contribution, and the paper is honest about the remaining gap to strong dense retrievers. However, the central inference mechanism in Equation (15) is not specified at an implementable level, and the empirical comparison relies on single-seed results without released code or significance tests. These issues are load-bearing for the paper's main claims and need to be addressed before the manuscript is ready for publication.
major comments (3)
- [Section 4.3.1, Equation (15)] The exact definition of the global relevance prior phi(t<=i; z_q) = max_{m in C_{t<=i}} s_oi(z_q, m) is not computationally feasible as written. For early prefixes, C_{t<=i} contains a substantial fraction of the candidate corpus, which has up to 5.6M candidates on M-BEIR. Computing an exact maximum over this set for every expanded prefix, for every query, gives an inference cost that scales with |C|, contradicting the corpus-size-independent decoding claim in Section 4.4 and the flat QPS curve for DrIG in Figure 3a. If the implementation instead uses an approximation, such as per-level maxima over codebook tokens, the paper must say so explicitly, because such a sum can select a token combination that is not an actual candidate identifier and therefore is not a valid relevance score for any real candidate. Please provide the exact algorithm or data structure used to compute Equation (15), state whether it is exact or approximate, and analyze the validity and cost of that approximation.
- [Section 5.2, Table 3] The main comparative claim that DrIG 'consistently outperforms' GENIUS is supported only by single-seed runs with no significance tests and no released code. Some of the task-level differences are small; for example, in local-pool retrieval on MSCOCO qi->ct, DrIG-LT reaches R@5 of 90.0 versus GENIUS-C at 89.9, and several DrIG-C versus GENIUS-C differences are around two points. Given the stochasticity of LMM fine-tuning, identifier construction, and decoder training, the absence of multiple seeds or significance testing makes it difficult to know which of the task-level gains are reliable. Please report multiple seeds with means and variances, or provide significance tests for the main comparisons, and make the code available so the baseline numbers can be reproduced.
- [Section 4.2.2, Equations (8) and (15); Table 5] The set-based relevance prior assumes that token-level evidence is additive and order-free, and that the maximum score over all candidates sharing a prefix is a reliable global relevance signal. The paper does not test this assumption directly. The ablations show consistent but modest gains from the set-based role, on the order of 0.6 to 1.7 R@5 points across tasks in Table 5. To strengthen the central claim, please add ablations with alternative aggregation functions, such as mean pooling, max pooling, or a learned attention over identifier tokens, and report the relative contribution of the modality token versus the semantic tokens to s_oi. This would clarify whether the order-free additive model is essential or whether a simpler aggregation would perform equally well.
minor comments (4)
- [Section 2.2.1] The sentence 'Inspired these two studies, DrIG designs identifiers with two complementary roles' is missing the word 'by' and should read 'Inspired by these two studies'.
- [References] Reference [50] contains the typo 'PmLR'; the correct publisher abbreviation is 'PMLR'.
- [Table 4] The table would be clearer if the zero-shot rows were explicitly marked with the dagger symbol in the method name itself, since the current formatting places dagger symbols inconsistently and the caption requires careful reading to distinguish M-BEIR-trained from in-domain-trained models.
- [Section 5.4.2] The t-SNE visualization in Figure 2 is qualitative; please state the number of points visualized and whether the same query or candidate embeddings were used for both panels, so that the visual comparison is meaningful.
Circularity Check
No significant circularity: every retrieval outcome is produced by trained components evaluated on held-out splits; the hybrid reranker and the ComGTIR self-citation are disclosed and not used to derive the central claim.
full rationale
DrIG's derivation chain is self-contained in the sense required by the circularity test. The identifier is built by residual quantization with an explicit RQ loss, a contrastive loss, and an MSE loss (Eqs. 3-6); the set-based scoring function s_oi is trained with its own contrastive and margin-ranking losses (Eqs. 8-10); and the decoder is trained with a token-level generative loss plus a pairwise ranking loss (Eqs. 18-21). None of these components is fitted to the test labels or to the metrics that are later reported. The set-based prior phi in Eq. (15) is a scoring heuristic composed of already-trained quantities; whether its exact max over C_{t<=i} is computationally feasible at corpus scale is a legitimate correctness/efficiency concern, but it is not a circularity because the prior is not defined in terms of the final retrieval output. The main empirical comparisons are against external baselines (GENIUS, GRACE, IRGen, AVG, ComGTIR-D, CLIP/BLIP variants, LamRA) on held-out M-BEIR, Flickr30K, and MSCOCO splits, so the reported recalls are predictions rather than re-statements of training objectives. Reuse of LamRA's fine-tuning recipe and of LamRA embeddings in DrIG-LT is explicitly disclosed in Sections 4.1.1-4.1.2 and 5.1.4; it is a hybrid composition, not a fitted parameter renamed as a prediction. The only self-citation, ComGTIR [30], is presented in related work as a prior text-to-image method with separate dual identifiers, and the paper explicitly differentiates DrIG by using a single identifier in both roles; no load-bearing argument is reduced to that citation. Therefore no circular step can be quoted from the paper.
Assumptions & free parameters
free parameters (9)
- beta_1, beta_2 =
100.0, 100.0
- contrastive temperature tau =
0.01
- margin rho =
0.1
- query interpolation alpha =
2.0
- adaptive margin scale gamma =
not reported
- set-based prior weight lambda =
1.0
- beam_size =
50
- rerank_depth_k =
50
- RQ levels and vocabulary (L, K) =
L=8, K=4096
assumptions (6)
- domain assumption Qwen2-VL embeddings, after LamRA-style two-stage contrastive fine-tuning, form an instruction-aware shared embedding space where cosine similarity reflects retrieval relevance.
- domain assumption Residual quantization codebooks assign identifiers whose first token encodes modality and later tokens encode progressively finer semantics.
- domain assumption Candidate relevance can be estimated by summing independent token-level scores in the global codebook vocabulary, as in Eq. 8.
- domain assumption The maximum set-based score over candidates sharing a prefix in Eq. 15 is a valid global relevance prior for beam search.
- standard math InfoNCE contrastive loss and EMA codebook updates produce stable, retrieval-oriented discrete codes.
- domain assumption M-BEIR relevance labels and instruction templates are accurate supervision for universal multimodal retrieval.
Cite this review
Pith. "Pith review of Generative Universal Multimodal Retrieval with Dual-role Identifiers." pith.science (2026). https://pith.science/paper/LQDHVD3V
@misc{pith2026260812987,
author = {Pith},
title = {Pith review of: Generative Universal Multimodal Retrieval with Dual-role Identifiers},
year = {2026},
howpublished = {\url{https://pith.science/paper/LQDHVD3V}},
note = {Machine review of arXiv:2608.12987}
}
read the original abstract
Generative information retrieval (GIR) has emerged as a compelling alternative to the conventional index-retrieve-then-rank retrieval pipeline by training a generator to produce the identifiers of relevant items directly. Despite its promise, a number of open challenges still remain. First, constrained left-to-right decoding is vulnerable to prefix-level errors and local optima. Second, most prior GIR research remains largely unimodal, leaving instruction-aware retrieval across text, image, and mixed image-text items underexplored. Third, although discrete identifier-based GIR offers higher efficiency, its retrieval accuracy still lags behind that of the cutting-edge dense-vector-based retrieval methods. Motivated by these challenges, we propose DrIG, a novel Generative framework for universal multimodal retrieval featuring Dual-role Identifiers, which supports diverse retrieval tasks across multiple modalities and domains. Each candidate is assigned a single residual-quantized identifier that serves two complementary roles. In its sequential role, the identifier is decoded autoregressively, where the first token explicitly models modality and the remaining tokens capture progressively finer semantics. In its set-based role, the same tokens are reinterpreted as an unordered set to provide a prefix-independent relevance prior, which guides constrained beam search and alleviates local-optimum errors. Extensive experiments on the M-BEIR benchmark and the text-to-image evaluation datasets show that:(1)DrIG consistently outperforms state-of-the-art generative multimodal baselines across diverse tasks, while hybrid reranking achieves a favorable efficiency-effectiveness trade-off against strong dense retrievers. (2)Ablation and scaling analyses reveal how the base LMM, beam size, reranking depth, and fusion strategy affect retrieval performance, providing practical guidance for system design.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Michele Bevilacqua, Giuseppe Ottaviano, Patrick Lewis, Scott Yih, Sebastian Riedel, and Fabio Petroni. 2022. Autore- gressive search engines: Generating substrings as document identifiers.Advances in Neural Information Processing Systems35 (2022), 31668–31683
2022
-
[2]
Christopher Burges, Robert Ragno, and Quoc Le. 2006. Learning to rank with nonsmooth cost functions.Advances in neural information processing systems19 (2006)
work page 2006
-
[3]
Hanqun Cao, Cheng Tan, Zhangyang Gao, Yilun Xu, Guangyong Chen, Pheng-Ann Heng, and Stan Z Li. 2024. A survey on generative diffusion models.IEEE transactions on knowledge and data engineering36, 7 (2024), 2814–2830
work page 2024
-
[4]
Min Cao, Shiping Li, Juntao Li, Liqiang Nie, and Min Zhang. 2022. Image-text Retrieval: A Survey on Recent Research and Development. InProceedings of the Thirty-First International Joint Conference on Artificial Intelligence, IJCAI-22. International Joint Conferences on Artificial Intelligence Organization, 5410–5417. Survey Track
work page 2022
-
[5]
Yingshan Chang, Mridu Narang, Hisami Suzuki, Guihong Cao, Jianfeng Gao, and Yonatan Bisk. 2022. Webqa: Multihop and multimodal qa. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 16495–16504
2022
-
[6]
Jiacheng Chen, Hexiang Hu, Hao Wu, Yuning Jiang, and Changhu Wang. 2021. Learning the best pooling strategy for visual semantic embedding. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 15789–15798
work page 2021
-
[7]
Jiangui Chen, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke, Wei Chen, Yixing Fan, and Xueqi Cheng. 2023. Continual learning for generative retrieval over dynamic corpora. InProceedings of the 32nd ACM international conference on information and knowledge management. 306–315
work page 2023
-
[8]
Xiaoyang Chen, Yanjiang Liu, Ben He, Le Sun, and Yingfei Sun. 2023. Understanding Differential Search Index for Text Retrieval. InFindings of the Association for Computational Linguistics: ACL 2023. Association for Computational Linguistics, Toronto, Canada, 10701–10717
work page 2023
Show all 81 references
-
[9]
Yang Chen, Hexiang Hu, Yi Luan, Haitian Sun, Soravit Changpinyo, Alan Ritter, and Ming-Wei Chang. 2023. Can pre-trained vision and language models answer visual information-seeking questions?. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Proces...
2023
-
[10]
Minghui Fang, Shengpeng Ji, Jialong Zuo, Hai Huang, Yan Xia, Jieming Zhu, Xize Cheng, Xiaoda Yang, Wenrui Liu, Gang Wang, Zhenhua Dong, and Zhou Zhao. 2025. CART: A Generative Cross-Modal Retrieval Framework With Coarse-To-Fine Semantic Modeling. InProceedings of the 63rd Annu...
2025 doi
-
[11]
Stephanie Fu, Netanel Tamir, Shobhita Sundaram, Lucy Chai, Richard Zhang, Tali Dekel, and Phillip Isola. 2023. DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data. arXiv:2306.09344 [cs.CV] https://arxiv.org/abs/2306.09344
2023 arXiv
-
[12]
Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021. SimCSE: Simple Contrastive Learning of Sentence Embeddings. InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (Eds...
2021 doi
-
[13]
Yunchao Gong, Liwei Wang, Micah Hodosh, Julia Hockenmaier, and Svetlana Lazebnik. 2014. Improving image- sentence embeddings using large weakly annotated photo collections. InComputer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Procee...
2014
-
[14]
Xintong Han, Zuxuan Wu, Phoenix X Huang, Xiao Zhang, Menglong Zhu, Yuan Li, Yang Zhao, and Larry S Davis
-
[15]
Mariya Hendriksen, Shuo Zhang, Ridho Reinanda, Mohamed Yahya, Edgar Meij, and Maarten de Rijke. 2025. Benchmark Granularity and Model Robustness for Image-Text Retrieval: A Reproducibility Study. InProceedings of the 48th International ACM SIGIR Conference on Research and Deve...
2025
-
[16]
Hexiang Hu, Yi Luan, Yang Chen, Urvashi Khandelwal, Mandar Joshi, Kenton Lee, Kristina Toutanova, and Ming-Wei Chang. 2023. Open-domain visual entity recognition: Towards recognizing millions of wikipedia entities. InProceedings of the IEEE/CVF International Conference on Comp...
2023
-
[17]
Jiho Jang, Chaerin Kong, Donghyeon Jeon, Seonhoon Kim, and Nojun Kwak. 2023. Unifying vision-language repre- sentation space with single-tower transformer. InProceedings of the AAAI conference on artificial intelligence, Vol. 37. 30 Li et al. 980–988
2023
-
[18]
Ting Jiang, Shaohan Huang, Zhongzhi Luan, Deqing Wang, and Fuzhen Zhuang. 2024. Scaling sentence embeddings with large language models. InFindings of the association for computational linguistics: EMNLP 2024. 3182–3196
2024
-
[19]
Ting Jiang, Minghui Song, Zihan Zhang, Haizhen Huang, Weiwei Deng, Feng Sun, Qi Zhang, Deqing Wang, and Fuzhen Zhuang. 2024. E5-v: Universal embeddings with multimodal large language models.arXiv preprint arXiv:2407.12580 (2024)
2024 arXiv
-
[20]
Ziyan Jiang, Rui Meng, Xinyi Yang, Semih Yavuz, Yingbo Zhou, and Wenhu Chen. 2025. VLM2Vec: Training Vision- Language Models for Massive Multimodal Embedding Tasks. InThe Thirteenth International Conference on Learning Representations
2025
-
[21]
Omar Khattab and Matei Zaharia. 2020. Colbert: Efficient and effective passage search via contextualized late interaction over bert. InProceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval. 39–48
2020
-
[22]
Chaeeun Kim, Soyoung Yoon, Hyunji Lee, Joel Jang, Sohee Yang, and Minjoon Seo. 2024. Exploring the Practicality of Generative Retrieval on Dynamic Corpora. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational L...
2024
-
[23]
Sungyeon Kim, Xinliang Zhu, Xiaofan Lin, Muhammet Bastan, Douglas Gray, and Suha Kwak. 2025. GENIUS: A generative framework for universal multimodal search. InProceedings of the Computer Vision and Pattern Recognition Conference. 19659–19669
2025
-
[24]
Chankyu Lee, Rajarshi Roy, Mengyao Xu, Jonathan Raiman, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping. 2025. NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models. InInternational Conference on Learning Representations, Y. Yue, A. Garg, N. Peng, F. S...
2025
-
[25]
Doyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho, and Wook-Shin Han. 2022. Autoregressive image generation using residual quantization. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 11523–11532
2022
-
[26]
Kuang-Huei Lee, Xi Chen, Gang Hua, Houdong Hu, and Xiaodong He. 2018. Stacked cross attention for image-text matching. InProceedings of the European conference on computer vision (ECCV). 201–216
2018
-
[27]
Haoxuan Li, Yi Bin, Yunshan Ma, Guoqing Wang, Yang Yang, See-Kiong Ng, and Tat-Seng Chua. 2025. SemCORE: A Semantic-Enhanced Generative Cross-Modal Retrieval Framework with MLLMs.arXiv preprint arXiv:2504.13172 (2025)
2025 arXiv
-
[28]
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. InInternational conference on machine learning. PMLR, 12888–12900
2022
-
[29]
Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. 2021. Align before fuse: Vision and language representation learning with momentum distillation.Advances in neural information processing systems34 (2021), 9694–9705
2021
-
[30]
Kaipeng Li, Haitao Yu, Yubo Fang, and Chao Lei. 2025. A Combination-based Framework for Generative Text-image Retrieval: Dual Identifiers and Hybrid Retrieval Strategies. InProceedings of the 2025 Annual International ACM SIGIR Conference on Research and Development in Informa...
2025
-
[31]
Xiaoxi Li, Jiajie Jin, Yujia Zhou, Yuyao Zhang, Peitian Zhang, Yutao Zhu, and Zhicheng Dou. 2025. From matching to generation: A survey on generative information retrieval.ACM Transactions on Information Systems43, 3 (2025), 1–62
2025
-
[32]
Yongqi Li, Hongru Cai, Wenjie Wang, Leigang Qu, Yinwei Wei, Wenjie Li, Liqiang Nie, and Tat-Seng Chua. 2025. Revolutionizing text-to-image retrieval as autoregressive token-to-voken generation. InProceedings of the 48th International ACM SIGIR Conference on Research and Develo...
2025
-
[33]
Yongqi Li, Wenjie Wang, Leigang Qu, Liqiang Nie, Wenjie Li, and Tat-Seng Chua. 2024. Generative cross-modal retrieval: Memorizing images in multimodal language models for retrieval and beyond. InProceedings of the 62nd Annual Meeting of the Association for Computational Lingui...
2024
-
[34]
Yongqi Li, Nan Yang, Liang Wang, Furu Wei, and Wenjie Li. 2023. Multiview identifiers enhanced generative retrieval. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 6636–6648
2023
-
[35]
Yongqi Li, Nan Yang, Liang Wang, Furu Wei, and Wenjie Li. 2024. Learning to rank in generative retrieval. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 8716–8723
2024
-
[36]
Sheng-Chieh Lin, Chankyu Lee, Mohammad Shoeybi, Jimmy Lin, Bryan Catanzaro, and Wei Ping. 2025. Mm- embed: Universal multimodal retrieval with multimodal llms.The Thirteenth International Conference on Learning Representations(2025). Generative Universal Multimodal Retrieval w...
2025
-
[37]
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014. Microsoft coco: Common objects in context. InEuropean conference on computer vision. Springer, 740–755
2014
-
[38]
Fuxiao Liu, Yinghan Wang, Tianlu Wang, and Vicente Ordonez. 2021. Visual news: Benchmark and challenges in news image captioning. InProceedings of the 2021 conference on empirical methods in natural language processing. 6761–6771
2021
-
[39]
Siqi Liu, Weixi Feng, Tsu-Jui Fu, Wenhu Chen, and William Wang. 2023. Edis: Entity-driven image search over multimodal web content. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 4877–4894
2023
-
[40]
Yikun Liu, Yajie Zhang, Jiayin Cai, Xiaolong Jiang, Yao Hu, Jiangchao Yao, Yanfeng Wang, and Weidi Xie. 2025. Lamra: Large multimodal model as your advanced retrieval assistant. InProceedings of the Computer Vision and Pattern Recognition Conference. 4015–4025
2025
-
[41]
Zheyuan Liu, Cristian Rodriguez-Opazo, Damien Teney, and Stephen Gould. 2021. Image retrieval on real-life images with pre-trained vision-and-language models. InProceedings of the IEEE/CVF international conference on computer vision. 2125–2134
2021
-
[42]
Sanket Vaibhav Mehta, Jai Gupta, Yi Tay, Mostafa Dehghani, Vinh Q Tran, Jinfeng Rao, Marc Najork, Emma Strubell, and Donald Metzler. 2023. Dsi++: Updating transformer memory with new documents. InProceedings of the 2023 conference on empirical methods in natural language proce...
2023
-
[43]
Kidist Amde Mekonnen, Yubao Tang, and Maarten de Rijke. 2025. Lightweight and Direct Document Relevance Optimization for Generative Information Retrieval. InProceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval. 1327–1338
2025
-
[44]
Donald Metzler, Yi Tay, Dara Bahri, and Marc Najork. 2021. Rethinking search: making domain experts out of dilettantes. InAcm sigir forum, Vol. 55. ACM New York, NY, USA, 1–27
2021
-
[45]
Hyeonseob Nam, Jung-Woo Ha, and Jeonghee Kim. 2017. Dual attention networks for multimodal reasoning and matching. InProceedings of the IEEE conference on computer vision and pattern recognition. 299–307
2017
-
[46]
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2018. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748(2018)
2018 arXiv
-
[47]
Ronak Pradeep, Kai Hui, Jai Gupta, Adam Lelkes, Honglei Zhuang, Jimmy Lin, Donald Metzler, and Vinh Tran. 2023. How Does Generative Retrieval Scale to Millions of Passages?. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association f...
2023
-
[48]
Shanbao Qiao, Xuebing Liu, and Seung-Hoon Na. 2023. DiffusionRet: Diffusion-Enhanced Generative Retriever using Constrained Decoding. InFindings of the Association for Computational Linguistics: EMNLP 2023. 9515–9529
2023
-
[49]
Leigang Qu, Haochuan Li, Tan Wang, Wenjie Wang, Yongqi Li, Liqiang Nie, and Tat-Seng Chua. 2025. Tiger: Unifying text-to-image generation and retrieval with large multimodal models. InThe Thirteenth International Conference on Learning Representations
2025
-
[50]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. InInternational conference on machine learnin...
2021
-
[51]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research21, 140 (2020), 1–67
2020
-
[52]
Shashank Rajput, Nikhil Mehta, Anima Singh, Raghunandan Hulikal Keshavan, Trung Vu, Lukasz Heldt, Lichan Hong, Yi Tay, Vinh Tran, Jonah Samost, et al. 2023. Recommender systems with generative retrieval.Advances in Neural Information Processing Systems36 (2023), 10299–10315
2023
-
[53]
Ali Razavi, Aaron Van den Oord, and Oriol Vinyals. 2019. Generating diverse high-fidelity images with vq-vae-2. Advances in neural information processing systems32 (2019)
2019
-
[54]
Stephen Edward Robertson, Steve Walker, Susan Jones, Micheline M Hancock-Beaulieu, Mike Gatford, et al. 1994. Okapi at TREC. (1994)
1994
-
[55]
Zihua Si, Zhongxiang Sun, Jiale Chen, Guozhang Chen, Xiaoxue Zang, Kai Zheng, Yang Song, Xiao Zhang, Jun Xu, and Kun Gai. 2024. Generative Retrieval with Semantic Tree-Structured Identifiers and Contrastive Learning. InProceedings of the 2024 Annual International ACM SIGIR Con...
2024
-
[56]
Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, Wojciech Galuba, Marcus Rohrbach, and Douwe Kiela. 2022. Flava: A foundational language and vision alignment model. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 15638–15650
2022
-
[57]
Karen Sparck Jones. 1972. A statistical interpretation of term specificity and its application in retrieval.Journal of documentation28, 1 (1972), 11–21. 32 Li et al
1972
-
[58]
Weiwei Sun, Lingyong Yan, Zheng Chen, Shuaiqiang Wang, Haichao Zhu, Pengjie Ren, Zhumin Chen, Dawei Yin, Maarten Rijke, and Zhaochun Ren. 2023. Learning to tokenize for generative retrieval.Advances in Neural Information Processing Systems36 (2023), 46345–46361
2023
-
[59]
Yubao Tang, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke, Shihao Liu, Shuaiqiang Wang, Dawei Yin, and Xueqi Cheng. 2025. Generative Retrieval for Book Search. InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1 (KDD ’25). Association for C...
2025
-
[60]
Yubao Tang, Ruqing Zhang, Jiafeng Guo, Maarten Rijke, Wei Chen, and Xueqi Cheng. 2024. Generative retrieval meets multi-graded relevance.Advances in Neural Information Processing Systems37 (2024), 72790–72817
2024
-
[61]
Yi Tay, Vinh Tran, Mostafa Dehghani, Jianmo Ni, Dara Bahri, Harsh Mehta, Zhen Qin, Kai Hui, Zhe Zhao, Jai Gupta, et al. 2022. Transformer memory as a differentiable search index.Advances in Neural Information Processing Systems 35 (2022), 21831–21843
2022
-
[62]
Michael Tschannen, Basil Mustafa, and Neil Houlsby. 2023. Clippo: Image-and-language understanding from pixels only. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 11006–11017
2023
-
[63]
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al. 2024. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution.arXiv preprint arXiv:2409.12191(2024)
2024 arXiv
-
[64]
Yujing Wang, Yingyan Hou, Haonan Wang, Ziming Miao, Shibin Wu, Qi Chen, Yuqing Xia, Chengmin Chi, Guoshuai Zhao, Zheng Liu, et al. 2022. A neural corpus indexer for document retrieval.Advances in Neural Information Processing Systems35 (2022), 25600–25614
2022
-
[65]
Zihan Wang, Yujia Zhou, Yiteng Tu, and Zhicheng Dou. 2023. Novo: Learnable and interpretable document identifiers for model-based ir. InProceedings of the 32nd ACM International Conference on Information and Knowledge Management. 2656–2665
2023
-
[66]
Cong Wei, Yang Chen, Haonan Chen, Hexiang Hu, Ge Zhang, Jie Fu, Alan Ritter, and Wenhu Chen. 2024. Uniir: Training and benchmarking universal multimodal information retrievers. InEuropean Conference on Computer Vision. Springer, 387–404
2024
-
[67]
Hui Wu, Yupeng Gao, Xiaoxiao Guo, Ziad Al-Halah, Steven Rennie, Kristen Grauman, and Rogerio Feris. 2021. Fashion iq: A new dataset towards retrieving images by natural language feedback. InProceedings of the IEEE/CVF Conference on computer vision and pattern recognition. 11307–11317
2021
-
[68]
Shiguang Wu, Zhaochun Ren, Xin Xin, Jiyuan Yang, Mengqi Zhang, Zhumin Chen, Maarten de Rijke, and Pengjie Ren
-
[69]
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. 2014. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions.Transactions of the association for computational linguistics2 (2014), 67–78
2014
-
[70]
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu. 2022. CoCa: Contrastive Captioners are Image-Text Foundation Models.Transactions on Machine Learning Research(2022). https://openreview. net/forum?id=Ee277P3AYC
2022
-
[71]
Hansi Zeng, Chen Luo, Bowen Jin, Sheikh Muhammad Sarwar, Tianxin Wei, and Hamed Zamani. 2024. Scalable and effective generative information retrieval. InProceedings of the ACM Web Conference 2024. 1441–1452
2024
-
[72]
Hansi Zeng, Chen Luo, and Hamed Zamani. 2024. Planning ahead in generative retrieval: Guiding autoregressive generation through simultaneous decoding. InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. 469–480
2024
-
[73]
Fuwei Zhang, Xiaoyu Liu, Xinyu Jia, Yingfei Zhang, Shuai Zhang, Xiang Li, Fuzhen Zhuang, Wei Lin, and Zhao Zhang
-
[74]
Peitian Zhang, Zheng Liu, Yujia Zhou, Zhicheng Dou, Fangchao Liu, and Zhao Cao. 2024. Generative Retrieval via Term Set Generation. InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. doi:10.1145/3626772.3657797
2024
-
[75]
Yidan Zhang, Ting Zhang, Dong Chen, Yujing Wang, Qi Chen, Xing Xie, Hao Sun, Weiwei Deng, Qi Zhang, Fan Yang, et al. 2024. Irgen: Generative modeling for image retrieval. InEuropean Conference on Computer Vision. Springer, 21–41
2024
-
[76]
InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics
Multi-level Relevance Document Identifier Learning for Generative Retrieval. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics. 10066–10080. doi:10.18653/v1/2025.acl-long.497
2025 doi
-
[77]
Zhedong Zheng, Liang Zheng, Michael Garrett, Yi Yang, Mingliang Xu, and Yi-Dong Shen. 2020. Dual-path convolu- tional image-text embeddings with instance loss.ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM)16, 2 (2020), 1–23. Generative Univer...
2020
-
[78]
Shengyao Zhuang, Houxing Ren, Linjun Shou, Jian Pei, Ming Gong, Guido Zuccon, and Daxin Jiang. 2023. Bridging the gap between indexing and retrieval for differentiable search index with query generation.Proceedings of the First Workshop on Generative Information Retrieval (Gen...
2023
-
[79]
Zhen Zhang, Xinyu Ma, Weiwei Sun, Pengjie Ren, Zhumin Chen, Shuaiqiang Wang, Dawei Yin, Maarten de Rijke, and Zhaochun Ren. 2025. Replication and exploration of generative retrieval over dynamic corpora. InProceedings of the 48th International ACM SIGIR Conference on Research ...
2025
-
[2017]
InProceedings of the IEEE international conference on computer vision
Automatic spatially-aware fashion concept discovery. InProceedings of the IEEE international conference on computer vision. 1463–1471
-
[2025]
InProceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval
Constrained Auto-Regressive Decoding Constrains Generative Retrieval. InProceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval. 2429–2440
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.