REVIEW 3 major objections 5 minor 51 references
Cross-encoder reranking runs four times faster with only the essential token interactions.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 22:34 UTC pith:AEJSV3FM
load-bearing objection MICE's masking analysis is a solid contribution, but the abstract overclaims OOD gains (contradicted by Ettin) and the 4x speedup is precomputation-dependent. the 3 major comments →
MICE: Minimal Interaction Cross-Encoders for efficient Re-ranking
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central discovery is that an information-retrieval cross-encoder can be reduced, by masking blocks of the self-attention matrix, to a minimal interaction pattern without losing effectiveness — and that training the model with those masks applied improves out-of-domain robustness. The authors establish this through a masking study: blocking attention to special tokens, blocking the classification token's direct view of the document, and blocking query-to-document flow all preserve or improve ranking scores, whereas blocking document-to-query flow hurts. Feeding this into a new architecture, query and document are encoded independently in the early layers, then a small number of interactio
What carries the argument
The load-bearing mechanism is the interaction-masking analysis translated into an architecture: a self-attention mask that leaves only (1) independent query and document contextualization in the first few layers, (2) one-way cross-attention from a frozen document representation to the query, (3) no direct document-to-CLS flow, and (4) pruning of top backbone layers. The mask is not a post-hoc patch; it is applied during fine-tuning, so the model learns to rank under the minimal interaction pattern. This both cuts compute and acts as a regularizer.
Load-bearing premise
The 4x speedup assumes document token representations can be precomputed offline and reused at query time; if that is not possible, the gain over a standard cross-encoder drops to about two times.
What would settle it
Run an end-to-end retrieve-and-rerank system with MICE, measuring total latency including corpus encoding, indexing, and the cost of storing and fetching document vectors. If the realised per-query speedup over a standard cross-encoder is about 2x rather than 4x, the central efficiency claim fails.
If this is right
- Document vectors can be encoded once offline, so the per-query work shrinks to a short query encode plus a few cross-attention layers; at that point MICE matches late-interaction models in latency while keeping cross-encoder-like ranking quality.
- Training with the masks applied does not hurt in-domain ranking, and on out-of-domain benchmarks it improves nDCG@10 by several points, suggesting the masked-out attention paths carry corpus-specific shortcuts.
- Only three interaction layers are needed; the top layers of the pretrained backbone can be dropped without hurting effectiveness, so the deployed model is smaller and faster than the original cross-encoder.
- On collections where a lexical baseline is already strong, MICE gains more than a late-interaction model of the same backbone size, which the authors attribute to cross-attention capturing exact-match signals better than a max-similarity operator.
Where Pith is reading between the lines
- Inference: the headline fourfold speedup depends on precomputing document token representations offline; if a deployment must encode documents live, the paper's own benchmark shows the gain over a standard cross-encoder is about two times, not four.
- Inference: because MICE freezes document representations after the initial layers, it should combine naturally with compressed or cached document stores, potentially reducing the memory footprint that the paper identifies as its main practical cost.
- Inference: the masking analysis transfers to a new cross-encoder backbone only when the model's attention sinks and information flow match the assumptions; the paper's results on one of its backbones show the starting checkpoint matters, so new backbones would need their own masking study before adopting the architecture.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. MICE derives a late-interaction-like re-ranking architecture from a standard cross-encoder by masking attention flows between query, document, [CLS], and [SEP] parts, then introducing mid-fusion, a frozen-document light cross-attention, and dropping late backbone layers. The method is evaluated on MiniLM-v2 and two Ettin backbones over MS MARCO, TREC DL19/DL20, and 13 BEIR datasets, with five seeds, and compared against cross-encoder baselines, Sparse CE, PreTTR, and ColBERTv2. The headline claims are a 4x inference speedup over standard cross-encoders, ColBERT-matching latency under precomputed document representations, retained in-domain effectiveness, and superior out-of-domain generalization.
Significance. The paper addresses a real bottleneck — cross-encoder re-ranking cost — and proposes a principled, interpretability-driven way to derive a late-interaction-style ranker from a cross-encoder. If the claims were fully supported, the architecture would be a useful contribution to efficient IR: it is reproducible, evaluated extensively across three backbones, multiple seeds, and many datasets, and the code is public. The masking study itself yields interesting findings, e.g., that blocking [CLS] attention to the document can improve OOD robustness, and that blocking query-to-document flow is less harmful than document-to-query flow. However, two central claims are currently overbroad: the OOD superiority claim does not hold for the Ettin backbones, and the 4x latency claim only holds with precomputed document representations.
major comments (3)
- [Abstract / Conclusion and §4.3 / Table 4] The abstract and conclusion claim that MICE shows "superior generalization abilities in OOD" without qualification. This is contradicted by Table 4: on average OOD nDCG@10, Ettin-32M MICE-ℓ6+3 obtains 46.4 vs. its unmasked baseline's 48.2, and Ettin-17M MICE-ℓ3+3 obtains 43.9 vs. 44.0. Section 4.3 itself states: "Ettin-based models show a different trend: MICE underperforms or matches the average OOD performance of the unmasked baselines." Since OOD robustness is a central value proposition of MICE, this internal inconsistency must be fixed: the claim should be restricted to the MiniLM backbone or the paper should provide an explanation and additional evidence for why the unqualified claim is justified.
- [Abstract / §4.4 / Table 5] The abstract's "decreases fourfold the inference latency" is unqualified. Table 5 reports a 241 ms vs. 470 ms latency for MICE-ℓ4+3 without precomputed document representations — a 2x speedup — and 113 ms vs. 470 ms only when document vectors are precomputed. Section 5 explicitly states "we do not implement the full indexing and first-stage retrieval pipeline." The 4x figure is thus conditional on an offline precomputation setup whose full end-to-end cost is not measured. The abstract and conclusion should state the precomputation condition, and ideally quantify the indexing/storage overhead.
- [§4.1 / §4.4 / Figure 4] The choice of first interaction layer ℓ* and the number of interaction layers is validated on in-domain data only (Figures 3 and 4). The OOD comparisons, especially the strong MiniLM OOD gains, are therefore selected rather than independently predicted. This is not a fatal flaw, but the paper should acknowledge that the OOD robustness claim is not a parameter-free prediction; it is an observed property of models selected on ID performance.
minor comments (5)
- [§4 title] Typo: "Minimal Interation Cross Encoder" should be "Minimal Interaction Cross-Encoder".
- [§3.4.2] Typo: "do not exacly" should be "do not exactly"; also "We can draw to conclusions" should be "We can draw two conclusions".
- [§3.2] Typo: "as long a the information" should be "as long as the information".
- [§2 / §3.3] The acronym "Sparse CE" is used both for a baseline model [32] and for the general notion of a sparsified cross-encoder. Consider clarifying the distinction in the text.
- [Table 4] The arrow notation is dense and the reader must infer which baseline the significance test compares against. It would help to spell out in the caption that all arrows are relative to each backbone's own unmasked cross-encoder reproduction, and to note the number of seeds used for significance testing.
Circularity Check
No significant circularity: MICE's architecture and claims are empirically evaluated against external baselines; the efficiency claim is conditional but not circular.
full rationale
The paper's derivation chain is empirical rather than definitional. Section 3 tests concrete masking patterns (Mask 0-3) on off-the-shelf and fine-tuned cross-encoders, using external interpretability findings [23,49] and external baselines (Sparse CE [32], PreTTR [25], ColBERTv2 [31], BEIR [38]) as reference points. The MICE architecture in Section 4 is then assembled from the measured effects of these ablations: mid-fusion uses the ℓ* found by Mask 3, light cross-attention is motivated by Mask 2 results, and layer dropping is validated in Figure 4. The ℓ* value and number of interaction layers are selected on ID validation performance, not fitted to the headline outcome, so the reported effectiveness numbers are measurements rather than identities. No equation defines a predicted quantity in terms of itself, and no fitted parameter is renamed as a prediction. The only author-overlapping citation is SPLADE [14] in related-work background, which is not load-bearing. The 4× latency claim is explicitly conditional on precomputed document representations (Table 5, Section 4.4), and the body admits that the full indexing/first-stage pipeline is not implemented (Section 5), making the claim conditional but not circular. The abstract's unqualified 'superior OOD' statement conflicts with Table 4 and Section 4.3's own admission that Ettin-based MICE 'underperforms or matches' the unmasked baselines on average OOD; this is an internal-consistency or overclaiming issue, not a circularity issue under the specified criteria. Overall, the central results are self-contained, externally benchmarked, and not forced by construction.
Axiom & Free-Parameter Ledger
free parameters (2)
- first interaction layer ℓ* =
MiniLM: 4; Ettin-17M: 3; Ettin-32M: 6
- number of interaction layers kept =
3
axioms (3)
- domain assumption Interpretability findings from Lu et al. on MiniLM-L12 cross-encoder transfer to other backbones.
- domain assumption Top layers of an MLM-pretrained transformer are non-essential for reranking and can be pruned without retraining the remaining layers.
- domain assumption Freezing document representations during interaction layers does not reduce effectiveness.
read the original abstract
Cross-encoders deliver state-of-the-art ranking effectiveness in information retrieval, but have a high inference cost. This prevents them from being used as first-stage rankers, but also incurs a cost when re-ranking documents. Prior work has addressed this bottleneck from two largely separate directions: accelerating cross-encoder inference by sparsifying the attention process or improving first-stage retrieval effectiveness using more complex models, e.g. late-interaction ones. In this work, we propose to bridge these two approaches, based on an in-depth understanding of the internal mechanisms of cross-encoders. Starting from cross-encoders, we show that it is possible to derive a new late-interaction-like architecture by carefully removing detrimental or unnecessary interactions. We name this architecture MICE (Minimal Interaction Cross-Encoders). We extensively evaluate MICE across both in-domain (ID) and out-of-domain (OOD) datasets. MICE decreases fourfold the inference latency compared to standard cross-encoders, matching late-interaction models like ColBERT while retaining most of cross-encoder ID effectiveness and demonstrating superior generalization abilities in OOD.
Figures
Reference graph
Works this paper leans on
-
[1]
Gianni Amati and Cornelis Joost Van Rijsbergen. 2002. Probabilistic models of information retrieval based on measuring the divergence from randomness. ACM Trans. Inf. Syst.20, 4 (Oct. 2002), 357–389. doi:10.1145/582415.582416
arXiv 2002
-
[2]
Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, et al
-
[3]
Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020. Longformer: The Long-Document Transformer.ArXivabs/2004.05150 (2020). https://api. semanticscholar.org/CorpusID:215737171
Pith/arXiv arXiv 2020
-
[4]
Cesare Campagnano, Antonio Mallia, Jack Pertschuk, and Fabrizio Silvestri. 2025. E2Rank: Efficient and Effective Layer-Wise Reranking. InEuropean Conference on Information Retrieval. Springer, 417–426
2025
-
[5]
Daniel Campos, Alexandre Marques, Tuan Nguyen, Mark Kurtz, and ChengXiang Zhai. 2023. Sparse*BERT: Sparse Models Generalize To New tasks and Domains. arXiv:2205.12452 [cs.CL] https://arxiv.org/abs/2205.12452
Pith/arXiv arXiv 2023
-
[6]
Qingqing Cao, Harsh Trivedi, Aruna Balasubramanian, and Niranjan Balasub- ramanian. 2020. DeFormer: Decomposing Pre-trained Transformers for Faster Question Answering. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Association for Computation...
-
[7]
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019. What Does BERT Look at? An Analysis of BERT‘s Attention. InProceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, Tal Linzen, Grzegorz Chrupała, Yonatan Belinkov, and Dieuwke Hupkes (Eds.). Association for Computational Linguistics,...
2019
-
[8]
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, and Daniel Campos. 2021. Overview of the TREC 2020 deep learning track. InText REtrieval Conference (TREC). TREC. https://www.microsoft.com/en-us/research/publication/overview-of- the-trec-2020-deep-learning-track/
2021
-
[9]
Voorhees
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, and Ellen M. Voorhees. 2020. Overview of the TREC 2019 deep learning track. InText RE- trieval Conference (TREC). TREC. https://www.microsoft.com/en-us/research/ publication/overview-of-the-trec-2019-deep-learning-track/
2020
-
[10]
Hervé Déjean and Stéphane Clinchant. 2025. Reranking with Compressed Docu- ment Representation. arXiv:2505.15394 [cs] doi:10.48550/arXiv.2505.15394
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2505.15394 2025
-
[11]
Hervé Déjean, Stéphane Clinchant, and Thibault Formal. 2024. A Thorough Comparison of Cross-Encoders and LLMs for Reranking SPLADE.ArXiv abs/2403.10407 (2024). https://api.semanticscholar.org/CorpusID:268510535
Pith/arXiv arXiv 2024
-
[12]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), Jill Burstein, Christy...
2019
-
[13]
Costa-jussà
Javier Ferrando, Gabriele Sarti, Arianna Bisazza, and Marta R. Costa-jussà
-
[14]
Thibault Formal, Benjamin Piwowarski, and Stéphane Clinchant. 2021. SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking. InProceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, 2288–2292. doi:10.1145/3404835.3463098
arXiv 2021
-
[15]
Jonathan Frankle and Michael Carbin. 2019. The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks. In7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. Open- Review.net. https://openreview.net/forum?id=rJl-b3RcF7
2019
-
[16]
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015. Distilling the Knowledge in a Neural Network. arXiv:1503.02531 [stat.ML] https://arxiv.org/abs/1503.02531
Pith/arXiv arXiv 2015
-
[17]
Schröder, Mete Sertkan, and A
Sebastian Hofstätter, Sophia Althammer, M. Schröder, Mete Sertkan, and A. Han- bury. 2020. Improving Efficient Neural Ranking Models with Cross-Architecture Knowledge Distillation.ArXiv(Oct. 2020)
2020
-
[18]
Samuel Humeau, Kurt Shuster, Marie-Anne Lachaux, and Jason Weston. 2020. Poly-encoders: Architectures and Pre-training Strategies for Fast and Accurate Multi-sentence Scoring. InInternational Conference on Learning Representations. https://openreview.net/forum?id=SkxgnnNFvH
2020
-
[19]
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense Passage Retrieval for Open- Domain Question Answering. InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguistics, Online, 6769–6781. doi:10.18653/v1/...
-
[21]
Yibin Lei, Shwai He, Ang Li, and Andrew Yates. 2025. Making Large Language Models Efficient Dense Retrievers. arXiv:2512.20612 [cs.IR] https://arxiv.org/ abs/2512.20612
arXiv 2025
-
[22]
Zhouhan Lin, Minwei Feng, Cicero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou, and Yoshua Bengio. 2017. A STRUCTURED SELF-ATTENTIVE SENTENCE EMBEDDING. InInternational Conference on Learning Representa- tions. https://openreview.net/forum?id=BJC_jUqxe
2017
-
[23]
Meng Lu, Catherine Chen, and Carsten Eickhoff. 2025. Pathway to Relevance: How Cross-Encoders Implement a Semantic Variant of BM25. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng (Eds.). Association for Computational Linguistics, Suzho...
-
[24]
Xueguang Ma, Liang Wang, Nan Yang, Furu Wei, and Jimmy Lin. 2024. Fine- Tuning LLaMA for Multi-Stage Text Retrieval. InProceedings of the 47th Inter- national ACM SIGIR Conference on Research and Development in Information Retrieval(Washington DC, USA)(SIGIR ’24). Association for Computing Machin- ery, New York, NY, USA, 2421–2425. doi:10.1145/3626772.3657951
arXiv 2024
-
[25]
Sean MacAvaney, Franco Maria Nardini, Raffaele Perego, Nicola Tonellotto, Nazli Goharian, and Ophir Frieder. 2020. Efficient Document Re-Ranking for Transformers by Precomputing Term Representations. InProceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’20). Association for Computing Mach...
arXiv 2020
-
[26]
Chuan Meng, Negar Arabzadeh, Arian Askari, Mohammad Aliannejadi, and Maarten de Rijke. 2024. Ranked List Truncation for Large Language Model- based Re-Ranking. InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval(Washington DC, USA) (SIGIR ’24). Association for Computing Machinery, New York, NY,...
arXiv 2024
-
[27]
Rodrigo Nogueira and Kyunghyun Cho. 2019. Passage Re-ranking with BERT. arXiv preprint arXiv:1901.04085(2019)
Pith/arXiv arXiv 2019
-
[28]
Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2020. A Primer in BERTology: What We Know About How BERT Works.Transactions of the Association for Computational Linguistics8 (2020), 842–866. doi:10.1162/tacl_a_00349
-
[29]
Guilherme Rosa, Luiz Bonifacio, Vitor Jeronymo, Hugo Abonizio, Marzieh Fadaee, Roberto Lotufo, and Rodrigo Nogueira. 2022. In Defense of Cross-Encoders for Zero-Shot Retrieval. arXiv:2212.06121 [cs.IR] https://arxiv.org/abs/2212.06121
Pith/arXiv arXiv 2022
-
[30]
Keshav Santhanam, Omar Khattab, Christopher Potts, and Matei Zaharia. 2022. PLAID: an efficient engine for late interaction retrieval. InProceedings of the 31st ACM International Conference on Information & Knowledge Management. 1747–1756. MICE : Minimal Interaction Cross-Encoders for efficient Re-ranking
2022
-
[31]
Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, and Matei Zaharia. 2022. ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction. InProceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Marine Carpuat, Marie-Catherine de Marn...
2022
-
[32]
Ferdinand Schlatt, Maik Fröbe, and Matthias Hagen. 2024. Investigating the Effects of Sparse Attention on Cross-Encoders. Vol. 14608. 173–190. arXiv:2312.17649 [cs] doi:10.1007/978-3-031-56027-9_11
Pith/arXiv arXiv 2024
-
[33]
Ferdinand Schlatt, Maik Fröbe, Harrisen Scells, Shengyao Zhuang, Bevan Koop- man, Guido Zuccon, Benno Stein, Martin Potthast, and Matthias Hagen. 2025. Rank-DistiLLM: Closing the Effectiveness Gap Between Cross-Encoders and LLMs for Passage Re-ranking. Springer Nature Switzerland, 323–334. doi:10.1007/978- 3-031-88714-7_31
doi:10.1007/978- 2025
-
[34]
Ivan Sekulić, Amir Soleimani, Mohammad Aliannejadi, and Fabio Crestani. 2020. Longformer for MS MARCO Document Re-ranking Task. arXiv:2009.09392 [cs.IR] https://arxiv.org/abs/2009.09392
Pith/arXiv arXiv 2020
-
[35]
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. 2024. RoFormer: Enhanced transformer with Rotary Position Embedding. Neurocomputing568 (2024), 127063. doi:10.1016/j.neucom.2023.127063
arXiv 2024
-
[36]
Hao Tan and Mohit Bansal. 2019. LXMERT: Learning Cross-Modality Encoder Representations from Transformers. InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan (Eds.). Associati...
-
[37]
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. 2022. Efficient Transformers: A Survey.ACM Comput. Surv.55, 6, Article 109 (Dec. 2022), 28 pages. doi:10.1145/3530811
doi:10.1145/3530811 2022
-
[38]
Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021. BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models. InThirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2). https://openreview. net/forum?id=wCu6T5xFjeJ
2021
-
[39]
Gomez, Lukasz Kaiser, and Illia Poloshukin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoeit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Poloshukin. 2017. Attention Is All You Need. In31st Conference on Neural Information Processing Systems (NIPS 2017). doi:10.48550/arXiv.1706.03762
-
[40]
Sinong Wang, Belinda Z Li, Madian Khabsa, Han Fang, and Hao Ma. 2020. Lin- former: Self-attention with linear complexity.arXiv preprint arXiv:2006.04768 (2020)
Pith/arXiv arXiv 2020
-
[41]
Wenhui Wang, Hangbo Bao, Shaohan Huang, Li Dong, and Furu Wei. 2021. MiniLMv2: Multi-Head Self-Attention Relation Distillation for Compressing Pre- trained Transformers. InFindings of the Association for Computational Linguistics: ACL-IJCNLP 2021. Association for Computational Linguistics, Online, 2140–2151. doi:10.18653/v1/2021.findings-acl.188
-
[42]
Benjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller, Oskar Hall- ström, Said Taghadouini, Alexis Gallagher, Raja Biswas, Faisal Ladhak, Tom Aarsen, Nathan Cooper, Griffin Adams, Jeremy Howard, and Iacopo Poli. 2024. Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Infer...
Pith/arXiv arXiv 2024
-
[43]
Orion Weller, Kathryn Ricci, Marc Marone, Antoine Chaffin, Dawn Lawrie, and Benjamin Van Durme. 2025. Seq vs Seq: An Open Suite of Paired Encoders and Decoders. (2025). arXiv:2507.11412 [cs.CL] https://arxiv.org/abs/2507.11412
arXiv 2025
-
[44]
Chuhan Wu, Fangzhao Wu, Tao Qi, and Yongfeng Huang. 2021. Fastformer: Additive Attention Can Be All You Need.ArXivabs/2108.09084 (2021). https: //api.semanticscholar.org/CorpusID:237266377
Pith/arXiv arXiv 2021
-
[45]
Zhichao Xu, Zhiqi Huang, Shengyao Zhuang, and Vivek Srikumar. 2025. Distillation versus Contrastive Learning: How to Train Your Rerankers. arXiv:2507.08336 [cs.CL] https://arxiv.org/abs/2507.08336
arXiv 2025
-
[46]
Andrew Yates, Rodrigo Nogueira, and Jimmy Lin. 2021. Pretrained transformers for text ranking: BERT and beyond. InProceedings of the 14th ACM International Conference on web search and data mining. 1154–1156
2021
-
[47]
Andrew Yates, Rodrigo Nogueira, and Jimmy Lin. 2021. Pretrained Transformers for Text Ranking: BERT and Beyond. InProceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval(Virtual Event, Canada)(SIGIR ’21). Association for Computing Machinery, New York, NY, USA, 2666–2668. doi:10.1145/3404835.3462812
arXiv 2021
-
[48]
Manzil Zaheer, Guru Guruganesh, Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, and Amr Ahmed. 2020. Big bird: transformers for longer sequences. InProceedings of the 34th International Conference on Neural Information Processing Systems (Vancouver, BC, Canada)(NIPS ’20). Curran Associates ...
2020
-
[49]
Jingtao Zhan, Jiaxin Mao, Yiqun Liu, Min Zhang, and Shaoping Ma. 2020. An Analysis of BERT in Document Ranking. InProceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval (Virtual Event, China)(SIGIR ’20). Association for Computing Machinery, New York, NY, USA, 1941–1944. doi:10.1145/3397271.3401325
arXiv 2020
-
[2016]
arXiv preprint arXiv:1611.09268(2016)
Ms marco: A human generated machine reading comprehension dataset. arXiv preprint arXiv:1611.09268(2016)
Pith/arXiv arXiv 2016
-
[2024]
arXiv:2405.00208 [cs.CL] https://arxiv.org/abs/2405.00208
A Primer on the Inner Workings of Transformer-based Language Models. arXiv:2405.00208 [cs.CL] https://arxiv.org/abs/2405.00208
-
[3734]
doi:10.18653/v1/2022.naacl-main.272
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.