{"work":{"id":"f641f433-ce28-4c6b-bce6-4f77e0d21705","openalex_id":"https://openalex.org/W4366559971","doi":"10.48550/arxiv.2304.09542","arxiv_id":"2304.09542","raw_key":null,"title":"Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking Agents","authors":null,"authors_text":"W","year":2023,"venue":"cs.CL","abstract":"Large Language Models (LLMs) have demonstrated remarkable zero-shot generalization across various language-related tasks, including search engines. However, existing work utilizes the generative ability of LLMs for Information Retrieval (IR) rather than direct passage ranking. The discrepancy between the pre-training objectives of LLMs and the ranking objective poses another challenge. In this paper, we first investigate generative LLMs such as ChatGPT and GPT-4 for relevance ranking in IR. Surprisingly, our experiments reveal that properly instructed LLMs can deliver competitive, even superior results to state-of-the-art supervised methods on popular IR benchmarks. Furthermore, to address concerns about data contamination of LLMs, we collect a new test set called NovelEval, based on the latest knowledge and aiming to verify the model's ability to rank unknown knowledge. Finally, to improve efficiency in real-world applications, we delve into the potential for distilling the ranking capabilities of ChatGPT into small specialized models using a permutation distillation scheme. Our evaluation results turn out that a distilled 440M model outperforms a 3B supervised model on the BEIR benchmark. The code to reproduce our results is available at www.github.com/sunnweiwei/RankGPT.","external_url":"https://arxiv.org/abs/2304.09542","cited_by_count":23,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2304.09542","created_at":"2026-05-10T08:27:51.859672+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"[Online; accessed 2025-07-26]","render_title":"[Online; accessed 2025-07-26]"},"hub":{"state":{"work_id":"f641f433-ce28-4c6b-bce6-4f77e0d21705","tier":"hub","tier_reason":"10+ Pith inbound or 1,000+ external citations","pith_inbound_count":28,"external_cited_by_count":23,"distinct_field_count":6,"first_pith_cited_at":"2023-12-05T12:39:00+00:00","last_pith_cited_at":"2026-07-07T08:04:05+00:00","author_build_status":"not_needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"not_needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-23T10:09:37.500128+00:00","tier_text":"hub"},"tier":"hub","role_counts":[{"context_role":"background","n":7},{"context_role":"baseline","n":1},{"context_role":"dataset","n":1}],"polarity_counts":[{"context_polarity":"background","n":7},{"context_polarity":"baseline","n":1},{"context_polarity":"use_dataset","n":1}],"runs":{},"summary":{},"graph":{},"authors":[]}}