Pith. sign in

REVIEW 2 major objections 1 minor 154 references

Artificial Intelligence and Misinformation in Art: Can Vision Language Models Judge the Hand or the Machine Behind the Canvas?

T0 review · 2 major / 1 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The paper claims that vision language models are limited at attributing paintings to their artists and at detecting AI-generated images, warning that this could spread misinformation in art.

desk verdict The submission is a title and abstract about VLM art attribution wrapped around a completely unrelated hybrid-search paper; there is no experiment to review. read the letter →

arxiv 2508.01408 v1 pith:6F3WZECK submitted 2025-08-02 cs.CY

classification cs.CY
keywords visionlanguagemodelspaintingattributionAI-generatedimagesartmisinformationartistidentificationcanvasimagegenerationmultimodalAIevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that vision-language models—AI systems that answer questions about images—are limited at two tasks: attributing a painting to its artist and spotting AI-generated images. The claimed basis is an experimental study of close to 40,000 paintings from 128 artists, using state-of-the-art models both to create imitation artworks and to analyze real and synthetic images. If the claim is right, users who ask AI models about art would frequently receive wrong attributions, and AI-generated forgeries could circulate with false labels. The delivered full text, however, is a different manuscript on trade-offs in hybrid database search, and it does not describe the painting-attribution experiment or present any of its results. A sympathetic reader should therefore treat the abstract's assertion as the paper's intended contribution, while recognizing that its supporting evidence is absent from the provided text.

What carries the argument

The central object is the vision language model (VLM) under test, together with the evaluation apparatus described in the abstract: a dataset of close to 40,000 paintings from 128 artists, plus state-of-the-art AI image-generation models that produce style-mimicking fakes. The machinery works by presenting real and generated images to VLMs and measuring whether the models can name the artist and say whether the image is AI-made. The abstract's conclusion of "limited capabilities" rests entirely on this experimental setup, which the delivered full text does not actually describe.

What would settle it

Run the described study: take a balanced set of real paintings by the named artists, generate style-imitation images with current AI models, and ask a panel of vision language models to attribute each image and classify it as real or AI-generated. If the models achieve high accuracy on both tasks, the abstract's claim of limited capability would be refuted. As an immediate check, the delivered full text contains no methods, dataset description, or results for such an experiment.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that current vision language models are not reliable judges of visual art provenance or of AI authorship. The author would describe an experiment in which state-of-the-art AI image generators are used to create paintings that mimic the style of 128 known artists, and then both real and generated paintings are shown to VLMs. The outcome, according to the abstract, is that the models have limited capability at canvas attribution and at identifying AI-generated images, meaning they frequently misattribute works and cannot reliably flag synthetic ones. Because users increasingly obtain information from AI queries, the author takes this result as evidence that VLMs need improved artist-attribution and AI-detection capabilities to prevent the spread of incorrect information about art.

Load-bearing premise

The conclusion presumes that the experiment named in the abstract—40,000 paintings, 128 artists, and the chosen generation and analysis models—was actually run and reported in a valid way; the delivered full text does not include or describe that experiment.

Editorial extensions

If this is right

  • Users who ask a vision language model whether a painting is genuine or who made it can receive incorrect answers, and may spread those answers as facts.
  • AI-generated images that mimic a painter's style are unlikely to be reliably flagged as synthetic by current VLMs, so imitations could be accepted as authentic works.
  • The claimed results argue for developing specialized VLM training or evaluation for artist attribution and AI-image detection before these tools are used in art-historical or authenticity contexts.
  • If the finding is correct, the risk of AI-assisted misinformation in art is concrete rather than hypothetical, matching the paper's stated motivation.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The delivered full text is a manuscript on hybrid search in databases, not the painting-attribution study; any reader relying on the full text for methods, data, or numbers would find no support for the abstract's claim.
  • Absent the described methods, the abstract's quantitative details (40,000 paintings, 128 artists) cannot be verified, so the strength of the "limited capabilities" claim is currently indeterminate.
  • If the abstract's qualitative conclusion is nevertheless taken as a hypothesis, a natural testable extension would be a public benchmark where multiple VLMs are scored on a fixed set of real and AI-generated artworks, to see whether the claimed limitation holds across models and prompt styles.
  • The paper's framing suggests a broader concern: as AI-generated images improve, the gap between generation and detection may widen, making provenance tools that rely on off-the-shelf VLMs increasingly unreliable over time.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 1 minor

Summary. The manuscript as submitted consists of an abstract that claims an experimental study of vision language models (VLMs) for two tasks: attributing paintings to artists and identifying AI-generated images, using a dataset of nearly 40,000 paintings from 128 artists. The abstract reports that VLMs have limited capabilities in both tasks. However, the delivered full text is an unrelated paper titled 'Balancing the Blend: An Experimental Analysis of Trade-offs in Hybrid Search' by Mengzhao Wang et al., which discusses retrieval paradigms, fusion methods, and re-ranking strategies for database systems. The full text contains no dataset of paintings, no VLMs, no prompts, no evaluation metrics, and no results tables relevant to the abstract's claims. The submission therefore provides no verifiable evidence for its central empirical conclusion.

Significance. If the claimed result were true, it would be relevant to ongoing discussions about AI-generated misinformation in art and the reliability of AI tools for art attribution. However, the manuscript's central claim is entirely unsupported by its content. The only substantive text is a hybrid-search paper with its own topic, methodology, and results, which does not address the abstract at all. The submission as it stands cannot be evaluated as a research contribution on VLMs and art, and it provides no basis for the stated conclusion. There are no machine-checked proofs, reproducible code, or parameter-free derivations for the claimed experiment; the only reproducible artifact is the hybrid-search framework, which is outside the scope of this submission.

major comments (2)
  1. [Abstract vs. Full Text] The abstract's central claim that 'the results show that vision language models have limited capabilities' for canvas attribution and AI-image detection is contradicted by the delivered full text, which is a paper about hybrid search in databases (arXiv:2508.01405v2). The full text contains no mention of paintings, artists, vision language models, or any experimental protocol for the claimed study. The submission therefore does not contain the study it reports; the empirical conclusion in the abstract is an unsupported assertion rather than a result.
  2. [Abstract only] Even considering the abstract in isolation, it provides insufficient detail for evaluation: it does not name the specific VLMs tested, the source or composition of the 40,000-painting dataset, the prompt templates used, or the evaluation metrics. No methods section, dataset description, or results table exists anywhere in the submitted manuscript. This omission is load-bearing because it prevents any check of the validity, fairness, or reproducibility of the claimed experiment.
minor comments (1)
  1. [Abstract] The phrase 'have limited capabilities to: 1) perform canvas attribution and 2) to identify AI generated images' contains an inconsistent infinitive marker; consider writing 'to: 1) perform canvas attribution and 2) identify AI-generated images'.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation exists to audit; the abstract's art-attribution experiment is absent from the delivered full text, which is an unrelated hybrid-search paper, so the mismatch is a verifiability problem rather than a circularity step.

full rationale

I find no circular step because the submission contains no derivation that could reduce to its own inputs. The abstract for 'Artificial Intelligence and Misinformation in Art' claims that 'both problems are experimentally studied using state-of-the-art AI models for image generation and analysis on a large dataset with close to 40,000 paintings from 128 artists' and that 'the results show that vision language models have limited capabilities.' The delivered full text, however, is 'Balancing the Blend: An Experimental Analysis of Trade-offs in Hybrid Search' by Mengzhao Wang and colleagues, self-labeled arXiv:2508.01405v2, with a different author list, a different subject, and no mention of paintings, artists, vision language models, or AI-generated images. There is consequently no method section, dataset description, prompt protocol, metric, or results table against which to check for equivalence-by-construction, fitted-parameter-as-prediction, or self-citation chains. The mismatch is a serious verifiability and integrity problem for the abstract's conclusion, and I flag it explicitly under the review rule, but it is not one of the circularity patterns this pass is charged to detect: no equation in the manuscript is defined in terms of the claimed answer, no fitted parameter is renamed as a prediction, and no load-bearing self-citation exists. Under the hard rule that circularity claims require a quotable reduction, the appropriate circularity score is 0.

Assumptions & free parameters 2 free parameters · 3 assumptions · 0 invented entities

The abstract's claims rest entirely on an experiment that is absent from the delivered manuscript. The free parameters listed are the experimental conditions the missing study would have had to fix: model and prompt choice, and dataset composition. The axioms are the validity conditions the measurement would need: correct labels, representative synthetic images, and a protocol that measures attribution rather than memorization. No new entities are introduced. The dominant feature of this ledger is the absence of the methods section that would let any of these entries be checked.

free parameters (2)
  • Choice of VLM models and prompt templates = not stated in abstract
    The abstract reports that 'vision language models' have limited capabilities, but names no models and no prompts; both choices determine the measured accuracy and would be fixed by the authors rather than derived.
  • Composition of the 40,000-painting dataset = not stated in abstract
    Which 128 artists and which paintings were included, and in what proportions, determines the difficulty of both the attribution and the AI-detection tasks; these choices are unstated and would act as hand-set experimental conditions.
assumptions (3)
  • domain assumption The ground-truth artist labels on the 40,000-painting dataset are correct.
    Attribution accuracy is meaningless without reliable labels; the abstract gives no provenance or verification for the labels.
  • domain assumption The AI-generated images used in the detection test resemble the images users actually encounter.
    If the synthetic images are either too easy (obvious artifacts) or too hard (perfectly matching a single generator) to detect, the reported limitation does not generalize; the abstract says nothing about the generators or prompting used.
  • domain assumption The evaluation protocol measures attribution skill rather than memorization of famous paintings.
    VLMs may recognize well-known works by memorization; the abstract gives no split between famous and obscure works, so the 'limited capability' conclusion presumes the protocol rules out memorization.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Artificial Intelligence and Misinformation in Art: Can Vision Language Models Judge the Hand or the Machine Behind the Canvas?." pith.science (2026). https://pith.science/paper/6F3WZECK

@misc{pith2026250801408,
  author       = {Pith},
  title        = {Pith review of: Artificial Intelligence and Misinformation in Art: Can Vision Language Models Judge the Hand or the Machine Behind the Canvas?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6F3WZECK}},
  note         = {Machine review of arXiv:2508.01408}
}
read the original abstract

The attribution of artworks in general and of paintings in particular has always been an issue in art. The advent of powerful artificial intelligence models that can generate and analyze images creates new challenges for painting attribution. On the one hand, AI models can create images that mimic the style of a painter, which can be incorrectly attributed, for example, by other AI models. On the other hand, AI models may not be able to correctly identify the artist for real paintings, inducing users to incorrectly attribute paintings. In this paper, both problems are experimentally studied using state-of-the-art AI models for image generation and analysis on a large dataset with close to 40,000 paintings from 128 artists. The results show that vision language models have limited capabilities to: 1) perform canvas attribution and 2) to identify AI generated images. As users increasingly rely on queries to AI models to get information, these results show the need to improve the capabilities of VLMs to reliably perform artist attribution and detection of AI generated images to prevent the spread of incorrect information.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

154 extracted references · 66 canonical work pages

  1. [1]

    Apress, Berkeley, CA, 217–237

    2008.Full-Text Search. Apress, Berkeley, CA, 217–237

  2. [2]

    Small but Mighty: Introducing answerai-colbert-small

    2014. Small but Mighty: Introducing answerai-colbert-small. https://www. answer.ai/posts/2024-08-13-small-but-mighty-colbert.html. [Online; accessed 05-November-2024]

  3. [3]

    CQADupStack

    2015. CQADupStack. http://nlp.cis.unimelb.edu.au/resources/cqadupstack/. [Online; accessed 02-November-2024]

  4. [4]

    Financial Opinion Mining and Question Answering

    2018. Financial Opinion Mining and Question Answering. https://sites.google. com/view/fiqa/. [Online; accessed 02-November-2024]

  5. [5]

    Getting Started with Hybrid Search

    2023. Getting Started with Hybrid Search. https://www.pinecone.io/learn/ hybrid-search-intro/. [Online; accessed 06-February-2025]

  6. [6]

    Elasticsearch: The heart of the Elastic Stack

    2024. Elasticsearch: The heart of the Elastic Stack. https://www.elastic.co/ elasticsearch. [Online; accessed 20-October-2024]

  7. [7]

    Full-text search vs vector search

    2024. Full-text search vs vector search. https://www.meilisearch.com/blog/full- text-search-vs-vector-search. [Online; accessed 09-October-2024]

  8. [8]

    IK Analysis for Elasticsearch and OpenSearch

    2024. IK Analysis for Elasticsearch and OpenSearch. https://github.com/ infinilabs/analysis-ik. [Online; accessed 20-October-2024]

Show all 154 references
  1. [9]

    Standard tokenizer

    2024. Standard tokenizer. https://www.elastic.co/guide/en/elasticsearch/ reference/current/analysis-standard-tokenizer.html. [Online; accessed 10- December-2024]

  2. [10]

    We Make AI Work

    2024. We Make AI Work. https://vespa.ai/. [Online; accessed 21-September- 2025]

  3. [11]

    Chroma is the open-source search and retrieval database for AI applica- tions

    2025. Chroma is the open-source search and retrieval database for AI applica- tions. https://www.trychroma.com/. [Online; accessed 21-September-2025]

  4. [12]

    Find the meaning in your data

    2025. Find the meaning in your data. https://opensearch.org/. [Online; accessed 21-September-2024]

  5. [13]

    High-Performance Vector Search at Scale

    2025. High-Performance Vector Search at Scale. https://qdrant.tech/. [Online; accessed 21-September-2025]

  6. [14]

    Hybrid Search

    2025. Hybrid Search. https://turbopuffer.com/docs/hybrid. [Online; accessed 20-September-2025]

  7. [15]

    Hybrid Search Explained

    2025. Hybrid Search Explained. https://weaviate.io/blog/hybrid-search- explained. [Online; accessed 06-February-2025]

  8. [16]

    Hybrid Search with Milvus

    2025. Hybrid Search with Milvus. https://milvus.io/docs/hybrid_search_with_ milvus.md. [Online; accessed 06-February-2025]

  9. [17]

    Relevance scoring in hybrid search using Reciprocal Rank Fusion (RRF)

    2025. Relevance scoring in hybrid search using Reciprocal Rank Fusion (RRF). https://learn.microsoft.com/en-us/azure/search/hybrid-search-ranking. [On- line; accessed 20-March-2025]

  10. [18]

    Reranking

    2025. Reranking. https://milvus.io/docs/reranking.md. [Online; accessed 20-March-2025]

  11. [19]

    A Review of Hybrid Search in Milvus

    2025. A Review of Hybrid Search in Milvus. https://zilliz.com/blog/a-review- of-hybrid-search-in-milvus. [Online; accessed 20-September-2025]

  12. [20]

    Supplementary Materials: Detailed Experimental Results

    2025. Supplementary Materials: Detailed Experimental Results. https://github. com/whenever5225/infinity/tree/main/exps/results. [Online; accessed 28-July- 2025]

  13. [21]

    The vector database for scale in production

    2025. The vector database for scale in production. https://www.pinecone.io/. [Online; accessed 21-September-2025]

  14. [22]

    Cecilia Aguerrebere, Ishwar Singh Bhati, Mark Hildebrand, Mariano Tepper, and Theodore L. Willke. 2023. Similarity search in the blink of an eye with compressed indices.Proc. VLDB Endow.16, 11 (2023), 3433–3446

  15. [23]

    Gianni Amati and C. J. van Rijsbergen. 2002. Probabilistic models of information retrieval based on measuring the divergence from randomness.ACM Trans. Inf. Syst.20, 4 (2002), 357–389

  16. [24]

    Fabien André, Anne-Marie Kermarrec, and Nicolas Le Scouarnec. 2015. Cache locality is not enough: High-Performance Nearest Neighbor Search with Product Quantization Fast Scan.Proc. VLDB Endow.9, 4 (2015), 288–299

  17. [25]

    Ilias Azizi, Karima Echihabi, and Themis Palpanas. 2025. Graph-Based Vector Search: An Experimental Evaluation of the State-of-the-Art.Proc. ACM Manag. Data3, 1 (2025), 43:1–43:31

  18. [26]

    Alexander Bondarenko, Maik Fröbe, Meriem Beloucif, Lukas Gienapp, Ya- men Ajjour, Alexander Panchenko, Chris Biemann, Benno Stein, Henning Wachsmuth, Martin Potthast, and Matthias Hagen. 2020. Overview of Touché 2020: Argument Retrieval. InWorking Notes of CLEF 2020 - Conferen...

  19. [27]

    Broder, David Carmel, Michael Herscovici, Aya Soffer, and Jason Y

    Andrei Z. Broder, David Carmel, Michael Herscovici, Aya Soffer, and Jason Y. Zien. 2003. Efficient query evaluation using a two-level retrieval process. In Proceedings of the 2003 ACM CIKM International Conference on Information and Knowledge Management, New Orleans, Louisiana...

  20. [28]

    Sebastian Bruch, Siyu Gai, and Amir Ingber. 2024. An Analysis of Fusion Functions for Hybrid Retrieval.ACM Trans. Inf. Syst.42, 1 (2024), 20:1–20:35

  21. [29]

    Sebastian Bruch, Franco Maria Nardini, Amir Ingber, and Edo Liberty. 2024. Bridging Dense and Sparse Maximum Inner Product Search.ACM Trans. Inf. Syst.42, 6 (2024), 151:1–151:38

  22. [31]

    Sebastian Bruch, Franco Maria Nardini, Cosimo Rulli, and Rossano Venturini

  23. [32]

    Sebastian Bruch, Franco Maria Nardini, Cosimo Rulli, Rossano Venturini, and Leonardo Venuta. 2025. Investigating the Scalability of Approximate Sparse Retrieval Algorithms to Massive Datasets.CoRRabs/2501.11628 (2025)

  24. [33]

    InProceedings of the 33rd ACM International Conference on Information and Knowledge Management, CIKM 2024, Boise, ID, USA, October 21-25, 2024

    Pairing Clustered Inverted Indexes with �-NN Graphs for Fast Approx- imate Retrieval over Learned Sparse Representations. InProceedings of the 33rd ACM International Conference on Information and Knowledge Management, CIKM 2024, Boise, ID, USA, October 21-25, 2024. 3642–3646

  25. [34]

    Kaushik Chakrabarti, Surajit Chaudhuri, and Venkatesh Ganti. 2011. Interval- based pruning for top-k processing over compressed lists. InProceedings of the 27th International Conference on Data Engineering, ICDE 2011, April 11-16, 2011, Hannover, Germany. 709–720

  26. [35]

    Yuzheng Cai, Jiayang Shi, Yizhuo Chen, and Weiguo Zheng. 2024. Navigat- ing Labels and Vectors: A Unified Approach to Filtered Approximate Nearest Neighbor Search.Proc. ACM Manag. Data2, 6 (2024), 246:1–246:27

  27. [36]

    Haonan Chen, Carlos Lassance, and Jimmy Lin. 2023. End-to-End Retrieval with Learned Dense and Sparse Representations Using Lucene.CoRRabs/2311.18503 (2023)

  28. [37]

    Cheng Chen, Chenzhe Jin, Yunan Zhang, Sasha Podolsky, Chun Wu, Szu- Po Wang, Eric Hanson, Zhou Sun, Robert Walzer, and Jianguo Wang. 2024. SingleStore-V: An Integrated Vector Database System in SingleStore.Proc. VLDB Endow.17, 12 (2024), 3772–3785

  29. [38]

    Jianlv Chen, Shitao Xiao, Peitian Zhang, Kun Luo, Defu Lian, and Zheng Liu. 2024. BGE M3-Embedding: Multi-Lingual, Multi-Functionality, Multi- Granularity Text Embeddings Through Self-Knowledge Distillation.CoRR abs/2402.03216 (2024)

  30. [39]

    Jianlyu Chen, Nan Wang, Chaofan Li, Bo Wang, Shitao Xiao, Han Xiao, Hao Liao, Defu Lian, and Zheng Liu. 2024. AIR-Bench: Automated Heterogeneous Information Retrieval Benchmark.CoRRabs/2412.13102 (2024)

  31. [40]

    Tao Chen, Mingyang Zhang, Jing Lu, Michael Bendersky, and Marc Najork

  32. [41]

    Sibei Chen, Ju Fan, Bin Wu, Nan Tang, Chao Deng, Pengyi Wang, Ye Li, Jian Tan, Feifei Li, Jingren Zhou, and Xiaoyong Du. 2025. Automatic Database Configuration Debugging using Retrieval-Augmented Language Models.Proc. ACM Manag. Data3, 1 (2025), 13:1–13:27

  33. [42]

    Benjamin Clavié. 2024. JaColBERTv2.5: Optimising Multi-Vector Retrievers to Create State-of-the-Art Japanese Retrievers with Constrained Resources.CoRR abs/2407.20750 (2024)

  34. [43]

    Arman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey, and Daniel S. Weld

  35. [44]

    Yaoqi Chen, Ruicheng Zheng, Qi Chen, Shuotao Xu, Qianxi Zhang, Xue Wu, Weihao Han, Hua Yuan, Mingqin Li, Yujing Wang, Jason Li, Fan Yang, Hao Sun, Weiwei Deng, Feng Sun, Qi Zhang, and Mao Yang. 2024. OneSparse: A Unified System for Multi-index Vector Search. InCompanion Procee...

  36. [45]

    Cormack, Charles L

    Gordon V. Cormack, Charles L. A. Clarke, and Stefan Büttcher. 2009. Reciprocal rank fusion outperforms condorcet and individual rank learning methods. In Proceedings of the 32nd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, SIG...

  37. [46]

    Shane Culpepper, Jimmy Lin, Joel M

    Matt Crane, J. Shane Culpepper, Jimmy Lin, Joel M. Mackenzie, and Andrew Trotman. 2017. A Comparison of Document-at-a-Time and Score-at-a-Time Query Evaluation. InProceedings of the Tenth ACM International Conference on Web Search and Data Mining, WSDM 2017, Cambridge, United ...

  38. [47]

    Constantinos Dimopoulos, Sergey Nepomnyachiy, and Torsten Suel. 2013. A candidate filtering mechanism for fast top-k query processing on modern cpus. InThe 36th International ACM SIGIR conference on research and development in Information Retrieval, SIGIR ’13, Dublin, Ireland ...

  39. [48]

    Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guil- laume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised Cross-lingual Representation Learn- ing at Scale. InProceedings of the 58th Annual Mee...

  40. [49]

    Xin Luna Dong. 2024. The Journey to a Knowledgeable Assistant with Retrieval- Augmented Generation (RAG). InCompanion of the 2024 International Conference on Management of Data, SIGMOD/PODS 2024, Santiago AA, Chile, June 9-15,

  41. [50]

    Manuel Faysse, Hugues Sibille, Tony Wu, Bilel Omrani, Gautier Viaud, Céline Hudelot, and Pierre Colombo. 2025. ColPali: Efficient Document Retrieval with Vision Language Models. InThe Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April ...

  42. [51]

    Thibault Formal, Carlos Lassance, Benjamin Piwowarski, and Stéphane Clin- chant. 2021. SPLADE v2: Sparse Lexical and Expansion Model for Information Retrieval.CoRRabs/2109.10086 (2021)

  43. [52]

    Shuai Ding and Torsten Suel. 2011. Faster top-k document retrieval using block- max indexes. InProceeding of the 34th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2011, Beijing, China, July 25-29, 2011. 993–1002

  44. [53]

    Cong Fu, Chao Xiang, Changxu Wang, and Deng Cai. 2019. Fast Approximate Nearest Neighbor Search With The Navigating Spreading-out Graph.Proc. VLDB Endow.12, 5 (2019), 461–474

  45. [54]

    Debasis Ganguly, Dwaipayan Roy, Mandar Mitra, and Gareth J. F. Jones. 2015. Word Embedding based Generalized Language Model for Information Retrieval. InProceedings of the 38th International ACM SIGIR Conference on Research and Development in Information Retrieval, Santiago, C...

  46. [55]

    Jinyang Gao, H. V. Jagadish, Wei Lu, and Beng Chin Ooi. 2014. DSH: data sensitive hashing for high-dimensional k-nnsearch. InInternational Conference on Management of Data, SIGMOD 2014, Snowbird, UT, USA, June 22-27, 2014. 1127–1138

  47. [56]

    Thibault Formal, Benjamin Piwowarski, and Stéphane Clinchant. 2021. SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking. InSIGIR ’21: The 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, Virtual Event, Canada, July 11...

  48. [57]

    Jianyang Gao and Cheng Long. 2024. RaBitQ: Quantizing High-Dimensional Vectors with a Theoretical Error Bound for Approximate Nearest Neighbor Search.Proc. ACM Manag. Data2, 3 (2024), 167

  49. [58]

    Luyu Gao, Zhuyun Dai, and Jamie Callan. 2021. COIL: Revisit Exact Lexical Match in Information Retrieval with Contextualized Inverted List. InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Techn...

  50. [59]

    Aristides Gionis, Piotr Indyk, and Rajeev Motwani. 1999. Similarity Search in High Dimensions via Hashing. InVLDB’99, Proceedings of 25th International Conference on Very Large Data Bases, September 7-10, 1999, Edinburgh, Scotland, UK. 518–529

  51. [60]

    Jianyang Gao and Cheng Long. 2023. High-Dimensional Approximate Nearest Neighbor Search: with Reliable and Efficient Distance Comparison Operations. Proc. ACM Manag. Data1, 2 (2023), 137:1–137:27

  52. [61]

    Christophe Van Gysel, Maarten de Rijke, and Evangelos Kanoulas. 2018. Neural Vector Spaces for Unsupervised Information Retrieval.ACM Trans. Inf. Syst.36, 4 (2018), 38:1–38:25

  53. [62]

    Marios Hadjieleftheriou, Nick Koudas, and Divesh Srivastava. 2009. Incremental maintenance of length normalized indexes for approximate string matching. InProceedings of the ACM SIGMOD International Conference on Management of Data, SIGMOD 2009, Providence, Rhode Island, USA, ...

  54. [63]

    Faegheh Hasibi, Fedor Nikolaev, Chenyan Xiong, Krisztian Balog, Svein Erik Bratsberg, Alexander Kotov, and Jamie Callan. 2017. DBpedia-Entity v2: A Test Collection for Entity Search. InProceedings of the 40th International ACM SIGIR Conference on Research and Development in In...

  55. [64]

    Adrien Grand, Robert Muir, Jim Ferenczi, and Jimmy Lin. 2020. From MAXS- CORE to Block-Max Wand: The Story of How Lucene Significantly Improved Query Evaluation Performance. InAdvances in Information Retrieval - 42nd European Conference on IR Research, ECIR 2020, Lisbon, Portu...

  56. [65]

    Qiang Huang, Jianlin Feng, Yikai Zhang, Qiong Fang, and Wilfred Ng. 2015. Query-Aware Locality-Sensitive Hashing for Approximate Nearest Neighbor Search.Proc. VLDB Endow.9, 1 (2015), 1–12

  57. [66]

    Hervé Jégou, Matthijs Douze, and Cordelia Schmid. 2011. Product Quantization for Nearest Neighbor Search.IEEE Trans. Pattern Anal. Mach. Intell.33, 1 (2011), 117–128

  58. [67]

    Rohan Jha, Bo Wang, Michael Günther, Saba Sturua, Mohammad Kalim Akram, and Han Xiao. 2024. Jina-ColBERT-v2: A General-Purpose Multilingual Late Interaction Retriever.CoRRabs/2408.16672 (2024)

  59. [68]

    Ipeirotis

    Vagelis Hristidis, Yuheng Hu, and Panagiotis G. Ipeirotis. 2010. Ranked queries over sources with Boolean query interfaces without ranking support. InPro- ceedings of the 26th International Conference on Data Engineering, ICDE 2010, March 1-6, 2010, Long Beach, California, USA...

  60. [69]

    Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick S. H. Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense Passage Retrieval for Open-Domain Question Answering. InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, E...

  61. [70]

    Omar Khattab and Matei Zaharia. 2020. ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT. InProceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval, SIGIR 2020, Virtual Event, Ch...

  62. [71]

    Carlos Lassance, Hervé Déjean, Thibault Formal, and Stéphane Clinchant. 2024. SPLADE-v3: New baselines for SPLADE.CoRRabs/2403.06789 (2024)

  63. [72]

    Wenqi Jiang, Marco Zeller, Roger Waleffe, Torsten Hoefler, and Gustavo Alonso

  64. [73]

    VLDB Endow.18, 1 (2024), 42–52

    Chameleon: a Heterogeneous and Disaggregated Accelerator System for Retrieval-Augmented Language Models.Proc. VLDB Endow.18, 1 (2024), 42–52

  65. [74]

    Minghan Li, Sheng-Chieh Lin, Barlas Oguz, Asish Ghoshal, Jimmy Lin, Yashar Mehdad, Wen-tau Yih, and Xilun Chen. 2023. CITADEL: Conditional Token Interaction via Dynamic Lexical Routing for Efficient and Effective Multi-Vector Retrieval. InProceedings of the 61st Annual Meeting...

  66. [75]

    Wen Li, Ying Zhang, Yifang Sun, Wei Wang, Mingjie Li, Wenjie Zhang, and Xuemin Lin. 2020. Approximate Nearest Neighbor Search on High Dimensional Data - Experiments, Analyses, and Improvement.IEEE Trans. Knowl. Data Eng. 32, 8 (2020), 1475–1488

  67. [76]

    Xiangyang Li, Kuicai Dong, Yi Quan Lee, Wei Xia, Yichun Yin, Hao Zhang, Yong Liu, Yasheng Wang, and Ruiming Tang. 2024. CoIR: A Comprehensive Benchmark for Code Information Retrieval Models.CoRRabs/2407.02883 (2024)

  68. [77]

    Chankyu Lee, Rajarshi Roy, Mengyao Xu, Jonathan Raiman, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping. 2025. NV-Embed: Improved Tech- niques for Training LLMs as Generalist Embedding Models. InThe Thirteenth International Conference on Learning Representations, ICLR 2025, Si...

  69. [78]

    Andersen, and Yuxiong He

    Conglong Li, Minjia Zhang, David G. Andersen, and Yuxiong He. 2020. Improv- ing Approximate Nearest Neighbor Search through Learned Adaptive Early Termination. InProceedings of the 2020 International Conference on Management of Data, SIGMOD Conference 2020, online conference [...

  70. [79]

    Antoine Louis, Vageesh Kumar Saxena, Gijs van Dijck, and Gerasimos Spanakis

  71. [80]

    Kejing Lu, Mineichi Kudo, Chuan Xiao, and Yoshiharu Ishikawa. 2021. HVS: Hierarchical Graph Structure Based on Voronoi Diagrams for Solving Approxi- mate Nearest Neighbor Search.Proc. VLDB Endow.15, 2 (2021), 246–258

  72. [81]

    Yi Luan, Jacob Eisenstein, Kristina Toutanova, and Michael Collins. 2021. Sparse, Dense, and Attentional Representations for Text Retrieval.Trans. Assoc. Comput. Linguistics9 (2021), 329–345

  73. [82]

    Jimmy Lin and Xueguang Ma. 2021. A Few Brief Notes on DeepImpact, COIL, and a Conceptual Framework for Information Retrieval Techniques.CoRR abs/2106.14807 (2021)

  74. [83]

    Yingfan Liu, Jiangtao Cui, Zi Huang, Hui Li, and Heng Tao Shen. 2014. SK-LSH: An Efficient Index Structure for Approximate Nearest Neighbor Search.Proc. VLDB Endow.7, 9 (2014), 745–756

  75. [84]

    Joel Mackenzie, Antonio Mallia, Alistair Moffat, and Matthias Petri. 2022. Ac- celerating Learned Sparse Indexes Via Term Impact Decomposition. InFindings of the Association for Computational Linguistics: EMNLP 2022, Abu Dhabi, United Arab Emirates, December 7-11, 2022. Associ...

  76. [85]

    Joel Mackenzie, Matthias Petri, and Alistair Moffat. 2022. Anytime Ranking on Document-Ordered Indexes.ACM Trans. Inf. Syst.40, 1 (2022), 13:1–13:32

  77. [86]

    Malkov and Dmitry A

    Yury A. Malkov and Dmitry A. Yashunin. 2020. Efficient and Robust Approx- imate Nearest Neighbor Search Using Hierarchical Navigable Small World Graphs.IEEE Trans. Pattern Anal. Mach. Intell.42, 4 (2020), 824–836

  78. [87]

    Antonio Mallia, Giuseppe Ottaviano, Elia Porciani, Nicola Tonellotto, and Rossano Venturini. 2017. Faster BlockMax WAND with Variable-sized Blocks. InProceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval, Shinjuku, Toky...

  79. [88]

    Hall, and Ryan T

    Ji Ma, Ivan Korotkov, Keith B. Hall, and Ryan T. McDonald. 2020. Hybrid First- stage Retrieval Models for Biomedical Literature. InWorking Notes of CLEF 2020 - Conference and Labs of the Evaluation Forum, Thessaloniki, Greece, September 22-25, 2020 (CEUR Workshop Proceedings),...

  80. [89]

    Hall, and Ryan T

    Ji Ma, Ivan Korotkov, Yinfei Yang, Keith B. Hall, and Ryan T. McDonald. 2021. Zero-shot Neural Passage Retrieval via Domain-targeted Synthetic Question Generation. InProceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Ma...

  81. [90]

    Antonio Mallia, Torsten Suel, and Nicola Tonellotto. 2024. Faster Learned Sparse Retrieval with Block-Max Pruning. InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2024, Washington DC, USA, July 14-18, 202...

  82. [91]

    Niklas Muennighoff, Nouamane Tazi, Loïc Magne, and Nils Reimers. 2023. MTEB: Massive Text Embedding Benchmark. InProceedings of the 17th Confer- ence of the European Chapter of the Association for Computational Linguistics, EACL 2023, Dubrovnik, Croatia, May 2-6, 2023. 2006–2029

  83. [92]

    Franco Maria Nardini, Cosimo Rulli, and Rossano Venturini. 2024. Efficient Multi-vector Dense Retrieval with Bit Vectors. InAdvances in Information Retrieval - 46th European Conference on Information Retrieval, ECIR 2024, Glasgow, UK, March 24-28, 2024, Proceedings, Part II, V...

  84. [93]

    Thong Nguyen, Sean MacAvaney, and Andrew Yates. 2023. A Unified Frame- work for Learned Sparse Retrieval. InAdvances in Information Retrieval - 45th European Conference on Information Retrieval, ECIR 2023, Dublin, Ireland, April 2-6, 2023, Proceedings, Part III (Lecture Notes ...

  85. [94]

    Antonio Mallia, Michal Siedlaczek, and Torsten Suel. 2019. An Experimental Study of Index Compression and DAAT Query Processing Methods. InAd- vances in Information Retrieval - 41st European Conference on IR Research, ECIR 2019, Cologne, Germany, April 14-18, 2019, Proceedings...

  86. [95]

    Antonio Mallia, Michal Siedlaczek, and Torsten Suel. 2021. Fast Disjunctive Candidate Generation Using Live Block Filtering. InWSDM ’21, The Fourteenth ACM International Conference on Web Search and Data Mining, Virtual Event, Israel, March 8-12, 2021. 671–679

  87. [96]

    James Jie Pan, Jianguo Wang, and Guoliang Li. 2024. Survey of vector database management systems.VLDB J.33, 5 (2024), 1591–1615

  88. [97]

    Cheoneum Park, Seohyeong Jeong, Minsang Kim, KyungTae Lim, and Yong- Hun Lee. 2025. SCV: Light and Effective Multi-Vector Retrieval with Sequence Compressive Vectors. InProceedings of the 31st International Conference on Computational Linguistics: Industry Track. 760–770

  89. [98]

    Derek Paulsen, Yash Govind, and AnHai Doan. 2023. Sparkly: A Simple yet Surprisingly Strong TF/IDF Blocker for Entity Matching.Proc. VLDB Endow.16, 6 (2023), 1507–1519

  90. [99]

    Yun Peng, Byron Choi, Tsz Nam Chan, Jianye Yang, and Jianliang Xu. 2023. Efficient Approximate Nearest Neighbor Search in Multi-dimensional Databases. Proc. ACM Manag. Data1, 1 (2023), 54:1–54:27

  91. [100]

    Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016. MS MARCO: A Human Generated MAchine Reading COmprehension Dataset. InProceedings of the Workshop on Cognitive Computation: Integrating neural and symbolic approaches 2016 co-...

  92. [101]

    Jiaul H. Paik. 2013. A novel TF-IDF weighting scheme for effective ranking. InThe 36th International ACM SIGIR conference on research and development in Information Retrieval, SIGIR ’13, Dublin, Ireland - July 28 - August 01, 2013. 343–352

  93. [102]

    Giovanni Puccetti, Alessio Miaschi, and Felice Dell’Orletta. 2021. How Do BERT Embeddings Organize Linguistic Knowledge?. InProceedings of Deep Learning Inside Out: The 2nd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures, DeeLIO@NAACL-HLT 2021,...

  94. [103]

    Yifan Qiao, Yingrui Yang, Haixin Lin, and Tao Yang. 2023. Optimizing Guided Traversal for Fast Learned Sparse Retrieval. InProceedings of the ACM Web Conference 2023, WWW 2023, Austin, TX, USA, 30 April 2023 - 4 May 2023. 3375–3385

  95. [104]

    Tadeusz Radecki. 1988. Trends in research on information retrieval – The potential for improvements in conventional Boolean retrieval systems.Inf. Process. Manag.24, 3 (1988), 219–227

  96. [105]

    Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. InProceedings of the 2019 Conference on Em- pirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJ...

  97. [106]

    Giulio Ermanno Pibiri and Rossano Venturini. 2021. Techniques for Inverted Index Compression.ACM Comput. Surv.53, 6 (2021), 125:1–125:36

  98. [107]

    Stefan Pohl, Alistair Moffat, and Justin Zobel. 2012. Efficient Extended Boolean Retrieval.IEEE Trans. Knowl. Data Eng.24, 6 (2012), 1014–1024

  99. [108]

    Keshav Santhanam, Omar Khattab, Christopher Potts, and Matei Zaharia. 2022. PLAID: An Efficient Engine for Late Interaction Retrieval. InProceedings of the 31st ACM International Conference on Information & Knowledge Management, Atlanta, GA, USA, October 17-21, 2022. 1747–1756

  100. [109]

    Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, and Matei Zaharia. 2022. ColBERTv2: Effective and Efficient Retrieval via Light- weight Late Interaction. InProceedings of the 2022 Conference of the North American Chapter of the Association for Computational...

  101. [110]

    Kunal Sawarkar, Abhilasha Mangal, and Shivam Raj Solanki. 2024. Blended RAG: Improving RAG (Retriever-Augmented Generation) Accuracy with Semantic Search and Hybrid Query-Based Retrievers. In7th IEEE International Conference on Multimedia Information Processing and Retrieval, ...

  102. [111]

    Chanop Silpa-Anan and Richard I. Hartley. 2008. Optimised KD-trees for fast image descriptor matching. In2008 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR 2008), 24-26 June 2008, Anchorage, Alaska, USA

  103. [112]

    Antoinette Renouf, Andrew Kehoe, and Jayeeta Banerjee. 2007. WebCorp: an integrated system for web text search. InCorpus linguistics and the web. 47–67

  104. [113]

    Robertson, Steve Walker, Susan Jones, Micheline Hancock-Beaulieu, and Mike Gatford

    Stephen E. Robertson, Steve Walker, Susan Jones, Micheline Hancock-Beaulieu, and Mike Gatford. 1994. Okapi at TREC-3. InProceedings of The Third Text REtrieval Conference, TREC 1994, Gaithersburg, Maryland, USA, November 2-4, 1994 (NIST Special Publication), Vol. 500-225. Nati...

  105. [114]

    Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. [n.d.]. BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models. InThirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Tr...

  106. [115]

    Nicola Tonellotto, Craig Macdonald, and Iadh Ounis. 2018. Efficient Query Processing for Scalable Web Search.Found. Trends Inf. Retr.12, 4-5 (2018), 319–500

  107. [116]

    Andrew Trotman and David Keeler. 2011. Ad hoc IR: not much room for improvement. InProceeding of the 34th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2011, Beijing, China, July 25-29, 2011. 1095–1096

  108. [117]

    Andrew Trotman, Antti Puurula, and Blake Burgess. 2014. Improvements to BM25 and Language Models Examined. InProceedings of the 2014 Australasian Document Computing Symposium, ADCS 2014, Melbourne, VIC, Australia, No- vember 27-28, 2014. 58

  109. [118]

    R. Smith. 2007. An Overview of the Tesseract OCR Engine. In9th International Conference on Document Analysis and Recognition (ICDAR 2007), 23-26 September, Curitiba, Paraná, Brazil. IEEE Computer Society, 629–633

  110. [119]

    Suhas Jayaram Subramanya, Devvrit, Harsha Vardhan Simhadri, Ravishankar Krishnaswamy, and Rohan Kadekodi. 2019. DiskANN: Fast Accurate Billion- point Nearest Neighbor Search on a Single Node. InAdvances in Neural Informa- tion Processing Systems 32: Annual Conference on Neural...

  111. [120]

    David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi. 2020. Fact or Fiction: Verifying Sci- entific Claims. InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, N...

  112. [121]

    Jianguo Wang, Chunbin Lin, Ruining He, Moojin Chae, Yannis Papakonstanti- nou, and Steven Swanson. 2017. MILC: Inverted List Compression in Memory. Proc. VLDB Endow.10, 8 (2017), 853–864

  113. [122]

    Jianguo Wang, Chunbin Lin, Yannis Papakonstantinou, and Steven Swanson

  114. [123]

    Jianguo Wang, Xiaomeng Yi, Rentong Guo, Hai Jin, Peng Xu, Shengjun Li, Xiangyu Wang, Xiangzhou Guo, Chengming Li, Xiaohai Xu, Kun Yu, Yuxing Yuan, Yinghao Zou, Jiquan Long, Yudong Cai, Zhenxiang Li, Zhifeng Zhang, Yihua Mo, Jun Gu, Ruiyi Jiang, Yi Wei, and Charles Xie. 2021. M...

  115. [124]

    Turtle and James Flood

    Howard R. Turtle and James Flood. 1995. Query Evaluation: Strategies and Optimizations.Inf. Process. Manag.31, 6 (1995), 831–850

  116. [125]

    Voorhees, Tasmeer Alam, Steven Bedrick, Dina Demner-Fushman, William R

    Ellen M. Voorhees, Tasmeer Alam, Steven Bedrick, Dina Demner-Fushman, William R. Hersh, Kyle Lo, Kirk Roberts, Ian Soboroff, and Lucy Lu Wang. 2020. TREC-COVID: constructing a pandemic information retrieval test collection. SIGIR Forum54, 1 (2020), 1:1–1:12

  117. [126]

    Mengzhao Wang, Xiaoliang Xu, Qiang Yue, and Yuxiang Wang. 2021. A Com- prehensive Survey and Experimental Comparison of Graph-Based Approximate Nearest Neighbor Search.Proc. VLDB Endow.14, 11 (2021), 1964–1978

  118. [127]

    Shuai Wang, Shengyao Zhuang, and Guido Zuccon. 2021. BERT-based Dense Retrievers Require Interpolation with BM25 for Effective Passage Retrieval. InICTIR ’21: The 2021 ACM SIGIR International Conference on the Theory of Information Retrieval, Virtual Event, Canada, July 11, 20...

  119. [128]

    Xiaoyue Wang, Jianyou Wang, Weili Cao, Kaicheng Wang, Ramamohan Paturi, and Leon Bergen. 2024. BIRCO: A Benchmark of Information Retrieval Tasks with Complex Objectives.CoRRabs/2402.14151 (2024)

  120. [129]

    Yifan Wang, Haodi Ma, and Daisy Zhe Wang. 2022. LIDER: An Efficient High- dimensional Learned Index for Large-scale Dense Passage Retrieval.Proc. VLDB Endow.16, 2 (2022), 154–166

  121. [130]

    Yining Wang, Liwei Wang, Yuanzhi Li, Di He, and Tie-Yan Liu. 2013. A The- oretical Analysis of NDCG Type Ranking Measures. InCOLT 2013 - The 26th Annual Conference on Learning Theory, June 12-14, 2013, Princeton University, NJ, USA (JMLR Workshop and Conference Proceedings), V...

  122. [131]

    Mengzhao Wang, Haotian Wu, Xiangyu Ke, Yunjun Gao, Xiaoliang Xu, and Lu Chen. 2024. An Interactive Multi-modal Query Answering System with Retrieval-Augmented Large Language Models.Proc. VLDB Endow.17, 12 (2024), 4333–4336. 15

  123. [132]

    Mengzhao Wang, Weizhi Xu, Xiaomeng Yi, Songlin Wu, Zhangyang Peng, Xiangyu Ke, Yunjun Gao, Xiaoliang Xu, Rentong Guo, and Charles Xie. 2024. Starling: An I/O-Efficient Disk-Resident Graph Index Framework for High- Dimensional Vector Similarity Search on Data Segment.Proc. ACM ...

  124. [133]

    Shiguang Wu, Wenda Wei, Mengqi Zhang, Zhumin Chen, Jun Ma, Zhaochun Ren, Maarten de Rijke, and Pengjie Ren. 2024. Generative Retrieval as Multi- Vector Dense Retrieval. InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retr...

  125. [134]

    Xiang Wu, Ruiqi Guo, David Simcha, Dave Dopson, and Sanjiv Kumar. 2019. Efficient Inner Product Approximation in Hybrid Spaces.CoRRabs/1903.08690 (2019)

  126. [135]

    Jasper Xian, Tommaso Teofili, Ronak Pradeep, and Jimmy Lin. 2024. Vector Search with OpenAI Embeddings: Lucene Is All You Need. InProceedings of the 17th ACM International Conference on Web Search and Data Mining, WSDM 2024, Merida, Mexico, March 4-8, 2024. ACM, 1090–1093

  127. [136]

    Zhichao Xu, Fengran Mo, Zhiqi Huang, Crystina Zhang, Puxuan Yu, Bei Wang, Jimmy Lin, and Vivek Srikumar. 2025. A Survey of Model Architectures in Information Retrieval.CoRRabs/2502.14822 (2025)

  128. [137]

    Peilin Yang, Hui Fang, and Jimmy Lin. 2017. Anserini: Enabling the Use of Lucene for Information Retrieval Research. InProceedings of the 40th Inter- national ACM SIGIR Conference on Research and Development in Information Retrieval, Shinjuku, Tokyo, Japan, August 7-11, 2017. ...

  129. [138]

    Zeyu Wang, Qitong Wang, Xiaoxing Cheng, Peng Wang, Themis Palpanas, and Wei Wang. 2024. Steiner-Hardness: A Query Hardness Measure for Graph-Based ANN Indexes.Proc. VLDB Endow.17, 13 (2024), 4668–4682

  130. [139]

    Manuel Widmoser, Daniel Kocher, and Nikolaus Augsten. 2024. Scalable Dis- tributed Inverted List Indexes in Disaggregated Memory.Proc. ACM Manag. Data2, 3 (2024), 171

  131. [140]

    Yianilos

    Peter N. Yianilos. 1993. Data Structures and Algorithms for Nearest Neigh- bor Search in General Metric Spaces. InProceedings of the Fourth Annual ACM/SIGACT-SIAM Symposium on Discrete Algorithms, 25-27 January 1993, Austin, Texas, USA. 311–321

  132. [141]

    Runjie Yu, Weizhou Huang, Shuhan Bai, Jian Zhou, and Fei Wu. 2025. AquaPipe: A Quality-Aware Pipeline for Knowledge Retrieval and Large Language Models. Proc. ACM Manag. Data3, 1 (2025), 11:1–11:26

  133. [142]

    Jingtao Zhan, Jiaxin Mao, Yiqun Liu, Jiafeng Guo, Min Zhang, and Shaoping Ma

  134. [143]

    Haoyu Zhang, Jun Liu, Zhenhua Zhu, Shulin Zeng, Maojia Sheng, Tao Yang, Guo- hao Dai, and Yu Wang. 2024. Efficient and Effective Retrieval of Dense-Sparse Hybrid Vectors using Graph-based Approximate Nearest Neighbor Search.CoRR abs/2410.20381 (2024)

  135. [144]

    Xin Zhang, Yanzhao Zhang, Dingkun Long, Wen Xie, Ziqi Dai, Jialong Tang, Huan Lin, Baosong Yang, Pengjun Xie, Fei Huang, Meishan Zhang, Wenjie Li, and Min Zhang. 2024. mGTE: Generalized Long-Context Text Representation and Reranking Models for Multilingual Text Retrieval. InPr...

  136. [145]

    Wen Yang, Tao Li, Gai Fang, and Hong Wei. 2020. PASE: PostgreSQL Ultra-High- Dimensional Approximate Nearest Neighbor Search Extension. InProceedings of the 2020 International Conference on Management of Data, SIGMOD Conference 2020, online conference [Portland, OR, USA], June...

  137. [146]

    Yingrui Yang, Parker Carlson, Shanxiu He, Yifan Qiao, and Tao Yang. 2024. Cluster-based Partial Dense Retrieval Fused with Sparse Text Retrieval. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2024, Was...

  138. [147]

    Chaoji Zuo, Miao Qiao, Wenchao Zhou, Feifei Li, and Dong Deng. 2024. SeRF: Segment Graph for Range-Filtering Approximate Nearest Neighbor Search. Proc. ACM Manag. Data2, 1 (2024), 69:1–69:26. 16

  139. [153]

    Xi Zhao, Zhonghan Chen, Kai Huang, Ruiyuan Zhang, Bolong Zheng, and Xiaofang Zhou. 2024. Efficient Approximate Maximum Inner Product Search Over Sparse Vectors. In40th IEEE International Conference on Data Engineering, ICDE 2024, Utrecht, The Netherlands, May 13-16, 2024. 3961–3974

  140. [154]

    Xinyang Zhao, Xuanhe Zhou, and Guoliang Li. 2024. Chat2Data: An Interactive Data Analysis System with RAG, Vector Databases and LLMs.Proc. VLDB Endow.17, 12 (2024), 4481–4484

  141. [2017]

    Inverted List Compres- sion

    An Experimental Study of Bitmap Compression vs. Inverted List Compres- sion. InProceedings of the 2017 ACM International Conference on Management of Data, SIGMOD Conference 2017, Chicago, IL, USA, May 14-19, 2017. 993–1008

  142. [2020]

    InProceedings of the 58th Annual Meeting of the Associa- tion for Computational Linguistics, ACL 2020, Online, July 5-10, 2020

    SPECTER: Document-level Representation Learning using Citation- informed Transformers. InProceedings of the 58th Annual Meeting of the Associa- tion for Computational Linguistics, ACL 2020, Online, July 5-10, 2020. 2270–2282

  143. [2021]

    InCIKM ’21: The 30th ACM International Conference on Information and Knowledge Management, Virtual Event, Queensland, Australia, November 1 - 5, 2021

    Jointly Optimizing Query Encoder and Product Quantization to Improve Retrieval Performance. InCIKM ’21: The 30th ACM International Conference on Information and Knowledge Management, Virtual Event, Queensland, Australia, November 1 - 5, 2021. 2487–2496

  144. [2022]

    Out-of-Domain Semantics to the Rescue! Zero-Shot Hybrid Retrieval Models. InAdvances in Information Retrieval - 44th European Conference on IR Research, ECIR 2022, Stavanger, Norway, April 10-14, 2022, Proceedings, Part I (Lecture Notes in Computer Science), Vol. 13185. 95–110

  145. [2024]

    InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2024, Washington DC, USA, July 14-18, 2024

    Efficient Inverted Indexes for Approximate Retrieval over Learned Sparse Representations. InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2024, Washington DC, USA, July 14-18, 2024. 152–162

  146. [2025]

    InProceedings of the 31st International Conference on Computational Linguistics, COLING 2025, Abu Dhabi, UAE, Janu- ary 19-24, 2025

    ColBERT-XM: A Modular Multi-Vector Representation Model for Zero- Shot Multilingual Information Retrieval. InProceedings of the 31st International Conference on Computational Linguistics, COLING 2025, Abu Dhabi, UAE, Janu- ary 19-24, 2025. 4370–4383

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.