Pith. sign in

REVIEW 3 major objections 5 minor 90 references

CoRank: LLM-Based Compact Reranking with Document Features for Scientific Retrieval

T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read CoRank claims that LLM reranking should first score 200 papers from compact semantic summaries, then re-score the top 20 on full text, and reports consistent top-10 ranking gains on five scientific retrieval benchmarks.

desk verdict CoRank is a practical, well-evaluated two-stage reranking method, but the reported gains confound semantic features with simply seeing more candidates. read the letter →

arxiv 2505.13757 v2 pith:XZS2UQRI submitted 2025-05-19 cs.IR cs.CL

classification cs.IRcs.CL
keywords InformationRetrievalScientificDocumentSearchLargeLanguageModelsLLM-BasedRerankinglistwisezero-shotcompactrepresentationscandidatecoverage
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that standard LLM listwise reranking is self-limiting for scientific search: because full text eats the context window, only about 20 candidates get reranked, and when the first-stage retriever is weak, the truly relevant papers are often outside that window. CoRank instead reranks in two passes. In the first pass, each paper is reduced offline to a compact semantic representation—hierarchical category, one query-relevant section, and a few top keywords—so 200 candidates fit in the same prompt; the top 20 from that coarse ranking are then reranked on full text. The authors report that this training-free, model-agnostic recipe lifts average nDCG@10 from 50.6 to 55.5 across five scientific retrieval datasets and several LLM backbones, at a fraction of the token cost of sliding-window expansion.

What carries the argument

The load-bearing object is the compact semantic document representation, specifically the paper's Form 4: a hierarchical category string, one query-selected section, and five query-selected keywords, produced offline by an LLM and cached. This cuts per-document tokens from roughly 200 for full text to tens, which is what lets the same context window hold 200 candidates instead of 20. The pipeline pairs this with a coarse-to-fine listwise reranking design: the LLM ranks the 200 summaries, the top 20 move to a second prompt with full text, and that second ranking is the final output. Adaptive selection uses embedding cosine similarity to keep only the section and keywords most relevant to the query.

What would settle it

Run CoRank on a benchmark where the first-stage retriever already has perfect recall within its top 20: if CoRank still beats vanilla full-text reranking over the same 20 documents, the gain is not coming from candidate coverage; equivalently, compare the coarse stage's recall at 20 against the retriever's recall at 20 and check whether the extra pool is actually adding relevant documents.

Watch

Extended reading notes

Core claim

The central claim is that a coarse reranking pass over compact LLM-extracted features can expand the candidate pool by an order of magnitude, and a second full-text pass then restores the precision lost to compression. Concretely, CoRank takes 200 first-stage candidates, replaces each document with category, section, and keywords, uses an LLM to listwise-rank those summaries, keeps the top 20, and then listwise-ranks those 20 on full text. The paper reports that this beats vanilla full-text reranking over 20 documents on every tested dataset and backbone, and that both the adaptive selection of query-relevant sections and keywords and the fine-grained full-text stage contribute. It frames the result as evidence that information extraction and retrieval are synergistic: structured features improve coverage without sacrificing final ranking accuracy.

Load-bearing premise

The load-bearing premise is that the LLM-extracted category, section, and keywords keep enough query-relevant signal that ranking 200 summaries finds relevant papers a 20-document full-text ranking would miss; if the extraction discards the decisive cues, the wider pool cannot compensate, and the paper's own Section 3.3.3 reports that summaries underperform full text when the document count is held fixed.

Editorial extensions

If this is right

  • The same LLM can consider ten times more candidates without extra model capacity or training, because the context budget is spent on coverage instead of per-document detail.
  • CoRank's gains are not tied to one retriever or one model family: the paper reports improvements across four different first-stage retrievers and with both open and proprietary LLM backbones.
  • Sliding-window expansion and CoRank stack, so the two breadth-increasing strategies can be combined rather than treated as alternatives.
  • Because feature extraction happens offline and is cached, the extra coverage comes with limited per-query latency and roughly 40% of the token cost of sliding windows.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable transfer: the same breadth-then-depth recipe should help in any long-document search domain with imperfect first-stage retrieval, such as patents, legal opinions, or clinical notes, since the bottleneck is context length rather than scientific text specifically.
  • If coverage is the active mechanism, CoRank's edge should shrink as first-stage retrieval improves; an adaptive variant could size the coarse pool from the retriever's estimated recall.
  • The paper's human evaluation suggests extraction quality is high on a small sample, which leaves open whether stronger extraction would further improve reranking or whether the coarse stage already saturates the available signal.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes CoRank, a training-free, model-agnostic three-stage pipeline for LLM listwise reranking in scientific retrieval. In stage one, document-level semantic features (category, sections, keywords) are extracted offline by an LLM. In stage two, a coarse listwise reranking is performed over up to 200 compact feature-based document representations, producing a shortlist. In stage three, the top 20 shortlisted documents are reranked again using their full text. The authors motivate the design by the observation that first-stage retrieval in the scientific domain is weak and that full-text listwise reranking is limited to roughly 20 candidates per prompt. They evaluate on five datasets (LitSearch, CSFCube, NFCorpus, SciFact, Trec-Covid) with three LLM backbones, reporting consistent improvements over vanilla listwise reranking and over a sliding-window baseline, with a headline average gain of +4.9 nDCG@10. The paper also includes ablations, hyperparameter studies, retriever robustness tests, small-backbone experiments, and a human evaluation of the extracted features.

Significance. If the claimed mechanism is correct, the paper offers a practical and broadly applicable recipe for improving zero-shot listwise reranking under tight context budgets: use compact, structured document representations to widen the candidate pool, then refine with full text. The empirical scope is substantial (five datasets, three reranking backbones, four first-stage retrievers, two small backbones), and the ablation and human-evaluation studies are thoughtful. The idea is simple and orthogonal to existing reranking techniques, which makes it potentially useful to practitioners. However, the central attribution of the gains to the semantic-feature representation is not established, because the main comparisons confound representation type with candidate pool size, and the paper itself concedes that at a fixed document count compact features underperform full text.

major comments (3)
  1. [§5.1, Table 1; §3.3.3] The headline comparison varies candidate pool size and representation type simultaneously. CoRank reranks 200 compact documents before refining over 20 full-text documents, whereas the vanilla baseline reranks 20 full-text documents and the sliding-window baseline reranks 100 full-text documents. The paper itself states in §3.3.3 that when the number of documents is held constant, feature-based representations underperform full text. Therefore the consistent gains in Table 1 and the advertised +4.9 nDCG@10 are equally compatible with the hypothesis that the improvement comes from considering 200 candidates instead of 20, and they do not establish that LLM-extracted semantic features are the reason. To separate the two factors, please add controls such as (i) CoRank with 20 compact documents versus 20 full-text documents; (ii) a token-equivalent non-semantic compact representation (e.g., the first N tokens or a truncated abstract) reranked over 200 candidates; and (iii) a full-text-based expanded-pool baseline of the same candidate count (e.g., via multiple windows or a long-context model). Without these, the claimed advantage of the semantic-feature representation over mere pool expansion is not demonstrated.
  2. [§5.3.2, Table 4] The efficiency comparison is subject to the same confound. The claim that CoRank needs only 40% of the sliding-window token budget attributes the savings to the semantic features, but any token-efficient representation, including simple truncation, would exhibit a similarly low token profile. Please report a cost-performance comparison at equal performance or equal candidate coverage, including a non-IE compact baseline, and clarify whether the reported token usage includes the offline feature-extraction cost. This would make the efficiency claim interpretable.
  3. [§5.3.1, Table 2] The default configuration uses 5 keywords, yet Table 2 shows that 15 keywords give the best coarse-stage nDCG@10 (42.6 vs. 41.1) at roughly 1.6x the token cost. Please either justify the choice of 5 as the default or present the main results with the better-performing configuration and discuss the sensitivity of the headline gain to this choice.
minor comments (5)
  1. [§5.3.5] The text says 'As exhibited in Table 5.3.5' but the table is numbered Table 5; please correct the reference.
  2. [Abstract; §5.2.1] The abstract reports an average gain 'from 50.6 to 55.5' while §5.2.1 reports 'from 47.2 to 54.9' without sliding windows. The two numbers use different aggregations; please state the aggregation rule explicitly so readers can reconcile them.
  3. [Table 1] The label 'CoRank Full' is ambiguous because CoRank is a two-stage method (compact features followed by full text). A label such as 'CoRank (compact→full)' would be clearer.
  4. [§5.3.2, Table 4] The table caption should state explicitly that the score is an average over LitSearch and CSFCube and over the three tested models; the current text leaves this implicit.
  5. [§3.3.2, Figure 3] The average full-text token length of roughly 200 tokens with a maximum of 5,774 is surprising for scientific papers; please clarify whether the 'full text' input is the complete paper or a truncated/abstract-only version, since this affects the interpretation of the token-efficiency analysis.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: CoRank is an empirical pipeline evaluated against external ground-truth labels, not a derivation that reduces to its own inputs.

full rationale

CoRank is an empirical reranking paper; there is no derivation chain in which an output is defined as an input. The headline result, an average absolute improvement of +4.9 nDCG@10, is measured against ground-truth relevance labels on five external benchmarks, so it cannot reduce to the method's own construction. The preliminary analysis in Section 3.3.2 motivates Form 4 on LitSearch, and hyperparameters such as the number of keywords and the fine-grained pool size are studied on LitSearch and then reused in the main LitSearch results; this is tuning on an evaluation set, which is a legitimate experimental concern, but it is not circularity under the definitions used here because the reported gains are not algebraically forced by that tuning. Several cited works [20, 21, 22, 72, 83, 87] share authors with the present paper, but they support peripheral claims such as suboptimal first-stage scientific retrieval, pseudo-query generation, positional bias, and zero-shot information extraction; those claims are also supported by external citations or by the paper's own first-stage baselines in Table 1, and none is invoked as a uniqueness theorem or as the sole justification of the central result. Section 3.3.3 explicitly concedes that feature-based representations underperform full text when the document count is held fixed, which correctly identifies the skeptic's confound between pool expansion and feature quality; however, Section 5.3.2 compares CoRank against sliding-window pool expansion and reports a +1.4 nDCG@10 gain at 40% of the token budget, so the central comparison is not merely a restatement of the input representations. No specific equation, definition, or fitted parameter can be exhibited that makes a claimed prediction equivalent to an input by construction.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

The method has no mathematical derivation; it relies on empirical assumptions about retriever quality, LLM extraction fidelity, and hyperparameter settings tuned on one dataset. No new entities are introduced.

free parameters (4)
  • top-k selected keywords = 5 of 30
    Adaptive selection retains the top 5 keywords per document; Table 2 shows nDCG varies with keyword count, so this is a tuned value rather than a derived one.
  • top-k selected sections = 1 of 3
    The method keeps only the most query-relevant section per document; this is a design choice not derived from theory.
  • coarse candidate count = 200
    The coarse stage reranks 200 compact representations per query; this number is selected to fit the context window and is a hyperparameter.
  • fine-grained pool size = 20
    The fine stage reranks the top 20 full-text documents; Table 3 shows a plateau around 20, so this value is empirically chosen.
assumptions (5)
  • domain assumption First-stage retrievers are often suboptimal in the scientific domain.
    Section 3.2 asserts that sparse and dense retrievers struggle with long-tail scientific concepts; this motivates the need for wider candidate coverage, but it is an empirical premise.
  • domain assumption LLM-extracted semantic features preserve enough relevance signal for coarse reranking.
    Section 3.3.3 concedes that feature representations underperform full text when the number of documents is fixed, so the success of the coarse stage depends on this premise.
  • domain assumption Zero-shot LLM information extraction produces accurate categories, sections, and keywords at scale.
    Section 5.3.5 reports human evaluation on 29 documents with high accuracy, but this is a small sample and the method assumes it generalizes.
  • domain assumption Fine-grained reranking on full text of the top candidates yields a more accurate final ordering than the coarse ranking.
    This is the design rationale for the second stage; ablation in Figure 4 supports it, but it remains an assumption about LLM behavior.
  • domain assumption The five benchmark datasets provide reliable ground-truth relevance judgments.
    Standard IR evaluation assumption, but dataset construction and judgment quality are not examined in this paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of CoRank: LLM-Based Compact Reranking with Document Features for Scientific Retrieval." pith.science (2026). https://pith.science/paper/XZS2UQRI

@misc{pith2026250513757,
  author       = {Pith},
  title        = {Pith review of: CoRank: LLM-Based Compact Reranking with Document Features for Scientific Retrieval},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XZS2UQRI}},
  note         = {Machine review of arXiv:2505.13757}
}
read the original abstract

Scientific retrieval is essential for advancing scientific knowledge discovery. Within this process, document reranking plays a critical role in refining first-stage retrieval results. However, standard LLM listwise reranking faces challenges in the scientific domain. First-stage retrieval is often suboptimal in the scientific domain, so relevant documents are ranked lower. Meanwhile, conventional listwise reranking places the full text of candidates into the context window, limiting the number of candidates that can be considered. As a result, many relevant documents are excluded before reranking, constraining overall retrieval performance. To address these challenges, we explore semantic-feature-based compact document representations (e.g., categories, sections, and keywords) and propose CoRank, a training-free, model-agnostic reranking framework for scientific retrieval. It presents a three-stage solution: (i) offline extraction of document features, (ii) coarse-grained reranking using these compact representations, and (iii) fine-grained reranking on full texts of the top candidates from (ii). This integrated process addresses suboptimal first-stage retrieval: Compact representations allow more documents to fit within the context window, improving candidate set coverage, while the final fine-grained ranking ensures a more accurate ordering. Experiments on 5 academic retrieval datasets show that CoRank significantly improves reranking performance across different LLM backbones (average nDCG@10 from 50.6 to 55.5). Overall, these results underscore the synergistic interaction between information extraction and information retrieval, demonstrating how structured semantic features can enhance reranking in the scientific domain.

Figures

Figures reproduced from arXiv: 2505.13757 by the authors.

Figure 1
Figure 1. We extract semantic features, rerank a larger candi [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Overview of our feature extraction pipeline: from unstructured documents, we apply zero-shot LLM information [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Token efficiency and performance comparison across different document representations. (a) Per-document token [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Ablation study results. Foreground bars report per [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Comparison of reranking methods (None, Vanilla, [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

90 extracted references · 23 canonical work pages

  1. [1]

    Anirudh Ajith, Mengzhou Xia, Alexis Chevalier, Tanya Goyal, Danqi Chen, and Tianyu Gao. 2024. LitSearch: A Retrieval Benchmark for Scientific Literature Search. https://arxiv.org/abs/2407.18940

  2. [2]

    Ricardo Baeza-Yates, Berthier Ribeiro-Neto, et al . 1999. Modern information retrieval. Vol. 463. ACM press, New York, NY, USA

  3. [3]

    Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, et al. 2016. MS MARCO: A Human Generated MAchine Reading COmprehension Dataset. https://arxiv.org/abs/1611.09268

  4. [4]

    Luiz Henrique Bonifacio, Hugo Queiroz Abonizio, Marzieh Fadaee, and Ro- drigo Frassetto Nogueira. 2022. InPars: Unsupervised Dataset Generation for Information Retrieval. In SIGIR ’22: The 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, Madrid, Spain, July 11 - 15, 2022, Enrique Amigó, Pablo Castells, Julio Go...

  5. [5]

    Vera Boteva, Demian Gholipour, Artem Sokolov, and Stefan Riezler. 2016. A Full- text Learning to Rank Dataset for Medical Information Retrieval. In Proceedings of the European Conference on Information Retrieval (ECIR) (Lecture Notes in Computer Science, Vol. 9626). Springer, Cham, 716–722

  6. [6]

    Zhe Cao, Tao Qin, Tie-Yan Liu, Ming-Feng Tsai, and Hang Li. 2007. Learning to rank: from pairwise approach to listwise approach. In Machine Learning, Proceedings of the Twenty-Fourth International Conference (ICML 2007), Corvallis, Oregon, USA, June 20-24, 2007 (ACM International Conference Proceeding Series, Vol. 227), Zoubin Ghahramani (Ed.). ACM, New Y...

  7. [7]

    Carbonell and Jade Goldstein-Stewart

    J. Carbonell and Jade Goldstein-Stewart. 1998. The Use of MMR, Diversity-Based Reranking for Reordering Documents and Producing Summaries. ACM SIGIR Forum 51 (1998), 209–210. http://dl.acm.org/citation.cfm?id=3130369

  8. [8]

    Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov. 2019. Transformer-XL: Attentive Language Models beyond a Fixed-Length Context. InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Anna Korhonen, David Traum, and Lluís Màrquez (Eds.). Association for Computational Linguistics, ...

Show all 90 references
  1. [9]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human...

  2. [10]

    Revanth Gangi Reddy, JaeHyeok Doo, Yifei Xu, Md Arafat Sultan, Deevya Swain, Avirup Sil, and Heng Ji. 2024. FIRST: Faster Improved Listwise Reranking with Sin- gle Token Decoding. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , Yaser...

  3. [11]

    Luyu Gao, Zhuyun Dai, and Jamie Callan. 2021. COIL: Revisit Exact Lexical Match in Information Retrieval with Contextualized Inverted List. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Com- putational Linguistics: Human Language Te...

  4. [12]

    Luyu Gao, Zhuyun Dai, and Jamie Callan. 2021. Rethink training of BERT rerankers in multi-stage retrieval pipeline. InEuropean Conference on Information Retrieval. Springer, Cham, 280–286

  5. [13]

    Google DeepMind. 2025. Gemini 2.5: Our most intelligent AI model. https://blog.google/technology/google-deepmind/gemini-model- thinking-updates-march-2025/. https://blog.google/technology/google- deepmind/gemini-model-thinking-updates-march-2025/ Google Blog

  6. [14]

    Aaron Grattafiori et al . 2024. The Llama 3 Herd of Models. arXiv:2407.21783 [cs.AI] https://arxiv.org/abs/2407.21783

  7. [15]

    Guo, Yixing Fan, Liang Pang, Liu Yang, Qingyao Ai, Hamed Zamani, Chen Wu, W

    J. Guo, Yixing Fan, Liang Pang, Liu Yang, Qingyao Ai, Hamed Zamani, Chen Wu, W. Bruce Croft, and Xueqi Cheng. 2019. A Deep Look into Neural Ranking Models for Information Retrieval. Inf. Process. Manag. 57 (2019), 102067. https: //api.semanticscholar.org/CorpusID:81977235

  8. [16]

    Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, and Boris Ginsburg. 2024. RULER: What’s the Real Context Size of Your Long-Context Language Models? arXiv:2404.06654 [cs.CL] https://arxiv.org/abs/2404.06654

  9. [17]

    Yunpeng Huang, Jingwei Xu, Junyu Lai, Zixu Jiang, Taolue Chen, Zenan Li, Yuan Yao, Xiaoxing Ma, Lijuan Yang, Hao Chen, et al. 2023. Advancing transformer architecture in long-context large language models: A comprehensive survey

  10. [18]

    Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bo- janowski, Armand Joulin, and Edouard Grave. 2021. Unsupervised Dense Infor- mation Retrieval with Contrastive Learning. https://arxiv.org/abs/2112.09118

  11. [19]

    Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, De- vendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thoma...

  12. [20]

    SeongKu Kang, Shivam Agarwal, Bowen Jin, Dongha Lee, Hwanjo Yu, and Jiawei Han. 2024. Improving Retrieval in Theme-specific Applications using a Corpus Topical Taxonomy. InProceedings of the ACM on Web Conference 2024 (WWW ’24), Singapore, May 13–17, 2024 . ACM, New York, NY, ...

  13. [21]

    SeongKu Kang, Bowen Jin, Wonbin Kweon, Yu Zhang, Dongha Lee, Jiawei Han, and Hwanjo Yu. 2025. Improving Scientific Document Retrieval with Concept Coverage-based Query Set Generation. https://arxiv.org/abs/2502.11181

  14. [22]

    SeongKu Kang, Yunyi Zhang, Pengcheng Jiang, Dongha Lee, Jiawei Han, and Hwanjo Yu. 2024. Taxonomy-guided Semantic Indexing for Academic Paper Search. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Yaser Al-Onaizan, Mohit Bansal, and ...

  15. [23]

    Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense Passage Retrieval for Open- Domain Question Answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP...

  16. [24]

    Jihyuk Kim, Minsoo Kim, Joonsuk Park, and Seung-won Hwang. 2023. Relevance- assisted Generation for Robust Zero-shot Retrieval. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: Industry Track , CoRank: LLM-Based Compact Reranking with ...

  17. [25]

    Oren Kurland and Lillian Lee. 2005. PageRank without hyperlinks: structural re-ranking using links induced by language models. In Proceedings of the 28th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (Salvador, Brazil) (SIGIR ’0...

  18. [26]

    Lee Giles

    Steve Lawrence, Kurt Bollacker, and C. Lee Giles. 1999. Indexing and retrieval of scientific literature. In Proceedings of the Eighth International Conference on Information and Knowledge Management (Kansas City, Missouri, USA) (CIKM ’99). Association for Computing Machinery, ...

  19. [27]

    Wanhae Lee, Minki Chun, Hyeonhak Jeong, and Hyunggu Jung. 2023. Toward Keyword Generation through Large Language Models. InCompanion Proceedings of the 28th International Conference on Intelligent User Interfaces (Sydney, NSW, Australia) (IUI ’23 Companion). Association for Co...

  20. [28]

    Haitao Li, Qingyao Ai, Jia Chen, Qian Dong, Yueyue Wu, Yiqun Liu, Chong Chen, and Qi Tian. 2023. SAILER: Structure-aware Pre-trained Language Model for Legal Case Retrieval. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Informatio...

  21. [29]

    Junyi Li, Tianyi Tang, Wayne Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen. 2024. Pre-trained language models for text generation: A survey. Comput. Surveys 56, 9 (2024), 1–39

  22. [30]

    Gaussier, Juntao Li, and Guodong Zhou

    Minghan Li, É. Gaussier, Juntao Li, and Guodong Zhou. 2024. KeyB2: Selecting Key Blocks is Also Important for Long Document Ranking with Large Language Models. https://arxiv.org/abs/2411.06254

  23. [31]

    Robert Litschko, Ivan Vulić, and Goran Glavaš. 2022. Parameter-Efficient Neu- ral Reranking for Cross-Lingual and Multilingual Retrieval. In Proceedings of the 29th International Conference on Computational Linguistics , Nicoletta Calzo- lari, Chu-Ren Huang, Hansaem Kim, James...

  24. [32]

    Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang

    Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024. Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics 12 (2024), 157–173. doi:10.1162/tacl_a_00638

  25. [33]

    Qi Liu, Bo Wang, Nan Wang, and Jiaxin Mao. 2025. Leveraging Passage Em- beddings for Efficient Listwise Reranking with Large Language Models. In Pro- ceedings of the ACM on Web Conference 2025 (Sydney NSW, Australia) (WWW ’25). Association for Computing Machinery, New York, NY...

  26. [34]

    Tie-Yan Liu et al. 2009. Learning to rank for information retrieval. Foundations and Trends® in Information Retrieval 3, 3 (2009), 225–331

  27. [35]

    Wenhan Liu, Xinyu Ma, Yutao Zhu, Ziliang Zhao, Shuaiqiang Wang, Dawei Yin, and Zhicheng Dou. 2024. Sliding Windows Are Not the End: Exploring Full Ranking with Long-Context Large Language Models. https://arxiv.org/abs/2412. 14574

  28. [36]

    Ye Liu, Kazuma Hashimoto, Yingbo Zhou, Semih Yavuz, Caiming Xiong, and Philip Yu. 2021. Dense Hierarchical Retrieval for Open-domain Question An- swering. In Findings of the Association for Computational Linguistics: EMNLP 2021, Marie-Francine Moens, Xuanjing Huang, Lucia Spec...

  29. [37]

    Yijun Liu, Jinzheng Yu, Yang Xu, Zhongyang Li, and Qingfu Zhu. 2025. A survey on transformer context extension: Approaches and evaluation. arXiv preprint arXiv:2503.13299 (2025)

  30. [38]

    Xueguang Ma, Xinyu Zhang, Ronak Pradeep, and Jimmy Lin. 2023. Zero-Shot Listwise Document Reranking with a Large Language Model. https://arxiv.org/ abs/2305.02156

  31. [39]

    Bonan Min, Hayley Ross, Elior Sulem, Amir Pouran Ben Veyseh, Thien Huu Nguyen, Oscar Sainz, Eneko Agirre, Ilana Heintz, and Dan Roth. 2023. Recent advances in natural language processing via large pre-trained language models: A survey. Comput. Surveys 56, 2 (2023), 1–40

  32. [40]

    Sheshera Mysore, Tim O’Gorman, Andrew McCallum, and Hamed Zamani. 2021. CSFCube – A Test Collection of Computer Science Research Articles for Faceted Query by Example. https://arxiv.org/abs/2103.12906

  33. [41]

    Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernandez Abrego, Ji Ma, Vincent Zhao, Yi Luan, Keith Hall, Ming-Wei Chang, and Yinfei Yang. 2022. Large Dual Encoders Are Generalizable Retrievers. InProceedings of the 2022 Conference on Empirical Methods in Natural Language P...

  34. [42]

    Christina Niklaus, Matthias Cetto, André Freitas, and Siegfried Handschuh. 2018. A Survey on Open Information Extraction. InProceedings of the 27th International Conference on Computational Linguistics, Emily M. Bender, Leon Derczynski, and Pierre Isabelle (Eds.). Association ...

  35. [43]

    Rodrigo Nogueira, Zhiying Jiang, Ronak Pradeep, and Jimmy Lin. 2020. Docu- ment Ranking with a Pretrained Sequence-to-Sequence Model. In Findings of the Association for Computational Linguistics: EMNLP 2020 , Trevor Cohn, Yulan He, and Yang Liu (Eds.). Association for Computat...

  36. [44]

    Rodrigo Nogueira and Jimmy Lin. 2019. From doc2query to docTTTTT- query. https://cs.uwaterloo.ca/~jimmylin/publications/Nogueira_Lin_2019_ docTTTTTquery-v2.pdf. Micropublication; version 2, Dec 7, 2019

  37. [45]

    Rodrigo Nogueira, Wei Yang, Kyunghyun Cho, and Jimmy Lin. 2019. Multi-Stage Document Ranking with BERT. https://arxiv.org/abs/1910.14424

  38. [46]

    Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, et al

    OpenAI, :, Aaron Hurst, Adam Lerer, Adam P. Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, et al. 2024. GPT-4o System Card. arXiv:2410.21276 [cs.CL] https://arxiv.org/abs/2410.21276

  39. [47]

    OpenAI. 2024. Introducing GPT-4.1 in the API. https://openai.com/index/gpt-4-1/. Accessed: 2025-05-19

  40. [48]

    OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, et al. 2023. GPT-4 Technical Report. https://arxiv.org/abs/2303.08774

  41. [49]

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback.Advances in neural information processing systems 35 ...

  42. [50]

    Saurav Pawar, SM Tonmoy, SM Zaman, Vinija Jain, Aman Chadha, and Amitava Das. 2024. The What, Why, and How of Context Length Extension Techniques in Large Language Models–A Detailed Survey

  43. [51]

    Ronak Pradeep, Sahel Sharifymoghaddam, and Jimmy Lin. 2023. RankVicuna: Zero-Shot Listwise Document Reranking with Open-Source Large Language Models. https://arxiv.org/abs/2309.15088

  44. [52]

    Ronak Pradeep, Sahel Sharifymoghaddam, and Jimmy Lin. 2023. RankZephyr: Effective and Robust Zero-Shot Listwise Reranking is a Breeze! https://arxiv. org/abs/2312.02724

  45. [53]

    Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, et al. 2024. Qwen2.5 Technical Report. https://arxiv.org/abs/2412.15115

  46. [54]

    Revanth Gangi Reddy, Pradeep Dasigi, Md Arafat Sultan, Arman Cohan, Avirup Sil, Heng Ji, and Hannaneh Hajishirzi. 2023. ReFIT: Relevance Feedback from a Reranker during Inference. https://arxiv.org/abs/2305.11744

  47. [55]

    Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Confer- ence on Natural Language Processing (EMNLP-IJ...

  48. [56]

    Ruiyang Ren, Shangwen Lv, Yingqi Qu, Jing Liu, Wayne Xin Zhao, QiaoQiao She, Hua Wu, Haifeng Wang, and Ji-Rong Wen. 2021. PAIR: Leveraging Passage- Centric Similarity Relation for Improving Dense Passage Retrieval. InFindings of the Association for Computational Linguistics: A...

  49. [57]

    Ruiyang Ren, Yuhao Wang, Kun Zhou, Wayne Xin Zhao, Wenjie Wang, Jing Liu, Ji-Rong Wen, and Tat-Seng Chua. 2024. Self-Calibrated Listwise Reranking with Large Language Models. https://arxiv.org/abs/2411.04602

  50. [58]

    Ruiyang Ren, Wayne Xin Zhao, Jing Liu, Hua Wu, Ji-Rong Wen, and Haifeng Wang. 2023. TOME: A Two-stage Approach for Model-based Retrieval. InProceed- ings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Anna Rogers, Jordan Bo...

  51. [59]

    Stephen Robertson, Hugo Zaragoza, et al . 2009. The probabilistic relevance framework: BM25 and beyond. Foundations and Trends® in Information Retrieval 3, 4 (2009), 333–389

  52. [60]

    Robertson and Karen Spärck Jones

    Stephen E. Robertson and Karen Spärck Jones. 1976. Relevance weighting of search terms. J. Am. Soc. Inf. Sci. 27 (1976), 129–146. https://api.semanticscholar. org/CorpusID:45186038

  53. [61]

    Stuart Rose, Dave Engel, Nick Cramer, and Wendy Cowley. 2010. Automatic Keyword Extraction from Individual Documents. In Text Mining: Applications and Theory, Michael W. Berry and Jacob Kogan (Eds.). John Wiley & Sons, Ltd, Chichester, UK, 1–20. doi:10.1002/9780470689646.ch1

  54. [62]

    Devendra Sachan, Mike Lewis, Mandar Joshi, Armen Aghajanyan, Wen-tau Yih, Joelle Pineau, and Luke Zettlemoyer. 2022. Improving Passage Retrieval with Zero-Shot Question Generation. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , Yoav ...

  55. [63]

    Gerard Salton, Anita Wong, and Chung-Shu Yang. 1975. A vector space model for automatic indexing. Commun. ACM 18 (1975), 613–620. https: //api.semanticscholar.org/CorpusID:6473756

  56. [64]

    Christopher Sciavolino, Zexuan Zhong, Jinhyuk Lee, and Danqi Chen. 2021. Simple Entity-Centric Questions Challenge Dense Retrievers. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , Marie- Francine Moens, Xuanjing Huang, Lucia Specia,...

  57. [65]

    Amanpreet Singh, Mike D’Arcy, Arman Cohan, Doug Downey, and Sergey Feldman. 2023. SciRepEval: A Multi-Format Benchmark for Scientific Document Representations. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Houda Bouamor, Juan Pino, ...

  58. [66]

    Amit Singhal et al. 2001. Modern information retrieval: A brief overview. IEEE Data Eng. Bull. 24, 4 (2001), 35–43

  59. [67]

    Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. 2024. Scaling llm test- time compute optimally can be more effective than scaling model parameters

  60. [68]

    Weiwei Sun, Lingyong Yan, Xinyu Ma, Shuaiqiang Wang, Pengjie Ren, Zhumin Chen, Dawei Yin, and Zhaochun Ren. 2023. Is ChatGPT Good at Search? Inves- tigating Large Language Models as Re-Ranking Agents. In Proceedings of the 2023 Conference on Empirical Methods in Natural Langua...

  61. [69]

    Xiaofei Sun, Xiaoya Li, Jiwei Li, Fei Wu, Shangwei Guo, Tianwei Zhang, and Guoyin Wang. 2023. Text Classification via Large Language Models. In Findings of the Association for Computational Linguistics: EMNLP 2023 , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association...

  62. [70]

    Raphael Tang, Crystina Zhang, Xueguang Ma, Jimmy Lin, and Ferhan Ture

  63. [71]

    Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021. BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models. https://arxiv.org/abs/2104.08663

  64. [72]

    Runchu Tian, Yanghao Li, Yuepeng Fu, Siyang Deng, Qinyu Luo, Cheng Qian, Shuo Wang, Xin Cong, Zhong Zhang, Yesai Wu, et al. 2025. Distance between Relevant Information Pieces Causes Bias in Long-Context LLMs. In Findings of the Association for Computational Linguistics: ACL 20...

  65. [73]

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open Foundation and Fine-Tuned Chat Models. https://arxiv.org/abs/2307.09288

  66. [74]

    Davison, and Jeff Heflin

    Mohamed Ali Trabelsi, Zhiyu Chen, Brian D. Davison, and Jeff Heflin. 2021. Neural ranking models for document retrieval. Information Retrieval Journal 24 (2021), 400–444. https://api.semanticscholar.org/CorpusID:232035744

  67. [75]

    Gomez, Lukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention Is All You Need. In Advances in Neural Information Processing Systems 30 (NIPS 2017) . Cur- ran Associates, Inc., Red Hook, NY, USA, 59...

  68. [76]

    Hersh, Kyle Lo, Kirk Roberts, Ian Soboroff, and Lucy Lu Wang

    Ellen Voorhees, Tasmeer Alam, Steven Bedrick, Dina Demner-Fushman, William R. Hersh, Kyle Lo, Kirk Roberts, Ian Soboroff, and Lucy Lu Wang. 2021. TREC-COVID: constructing a pandemic information retrieval test collection. SIGIR Forum 54, 1, Article 1 (Feb. 2021), 12 pages. doi:...

  69. [77]

    David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi. 2020. Fact or Fiction: Verifying Scientific Claims. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , Bonnie Webber...

  70. [78]

    Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. 2022. Text Embeddings by Weakly-Supervised Contrastive Pre-training. https://arxiv.org/abs/2212.03533

  71. [79]

    Howard D White, H Cooper, LV Hedges, et al. 2009. Scientific communication and literature retrieval. The handbook of research synthesis and meta-analysis 2 (2009), 51–71

  72. [80]

    Daya C Wimalasuriya and Dejing Dou. 2010. Ontology-based information extraction: An introduction and a survey of current approaches. 306–323 pages

  73. [81]

    Shijie Xia, Yiwei Qin, Xuefeng Li, Yan Ma, Run-Ze Fan, Steffi Chern, Haoyang Zou, Fan Zhou, Xiangkun Hu, Jiahe Jin, et al. 2025. Generative AI Act II: Test Time Scaling Drives Cognition Engineering. https://arxiv.org/abs/2504.13828

  74. [82]

    Wenhan Xiong, Jingyu Liu, Igor Molybog, Hejia Zhang, Prajjwal Bhargava, Rui Hou, Louis Martin, Rashi Rungta, Karthik Abinav Sankararaman, Barlas Oguz, et al . 2024. Effective Long-Context Scaling of Foundation Models. In Proceedings of the 2024 Conference of the North American...

  75. [83]

    Xueqiang Xu, Jinfeng Xiao, James Barry, Mohab Elkaref, Jiaru Zou, Pengcheng Jiang, Yunyi Zhang, Max Giammona, Geeth de Mel, and Jiawei Han. 2025. Zero- Shot Open-Schema Entity Structure Discovery. arXiv:2506.04458 [cs.CL] https: //arxiv.org/abs/2506.04458

  76. [84]

    Andrew Yates, Rodrigo Frassetto Nogueira, and Jimmy Lin. 2021. Pretrained Transformers for Text Ranking: BERT and Beyond. In SIGIR ’21: The 44th In- ternational ACM SIGIR Conference on Research and Development in Information Retrieval, Virtual Event, Canada, July 11-15, 2021 ,...

  77. [85]

    Chenhan Yuan, Qianqian Xie, and Sophia Ananiadou. 2023. Zero-shot Tempo- ral Relation Extraction with ChatGPT. In Proceedings of the 22nd Workshop on Biomedical Natural Language Processing and BioNLP Shared Tasks , Dina Demner- fushman, Sophia Ananiadou, and Kevin Cohen (Eds.)...

  78. [86]

    Kai Zhang, Bernal Jimenez Gutierrez, and Yu Su. 2023. Aligning Instruction Tasks Unlocks Large Language Models as Zero-Shot Relation Extractors. In Findings of the Association for Computational Linguistics: ACL 2023 , Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Eds.)....

  79. [87]

    Yunyi Zhang, Ruozhen Yang, Xueqiang Xu, Rui Li, Jinfeng Xiao, Jiaming Shen, and Jiawei Han. 2025. TELEClass: Taxonomy Enrichment and LLM-Enhanced Hierarchical Text Classification with Minimal Supervision. In Proceedings of the ACM on Web Conference 2025 (Sydney NSW, Australia)...

  80. [88]

    Xuanhe Zhou, Guoliang Li, and Zhiyuan Liu. 2023. LLM As DBA. arXiv:2308.05481 [cs.DB] https://arxiv.org/abs/2308.05481

  81. [89]

    Honglei Zhuang, Zhen Qin, Rolf Jagerman, Kai Hui, Ji Ma, Jing Lu, Jianmo Ni, Xuanhui Wang, and Michael Bendersky. 2023. RankT5: Fine-Tuning T5 for Text Ranking with Ranking Losses. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Inf...

  82. [2024]

    Found in the Middle: Permutation Self-Consistency Improves Listwise Ranking in Large Language Models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , Kev...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.