{"id":"f73e3879-0c1a-4324-ba7f-2adb451df1ae","arxiv_id":"2507.08890","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A TREC benchmark track found that LLM-prompted ranking systems beat the previous four-time winner, and that synthetic queries from T5 and GPT-4 rank systems nearly as consistently as human queries (Kendall tau = 0.8487).","lead":"This paper reports outcomes of the fifth and final TREC Deep Learning track, where runs using large language models with prompting outperformed fine-tuned neural ranking models for the first time. It also tests whether synthetic queries generated by T5 and GPT-4 can replace human queries in building evaluation benchmarks, finding similar system rankings.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Synthetic-query reliability (tau=0.8487) rests on only 31 filtered queries; the human selection step may inflate agreement, so the abstract's substitution claim is overbroad.","rationale":"The paper's two headline claims are (1) prompt-based LLM runs outperform nnlm and (2) synthetic queries give similar system ordering to human queries (tau=0.8487). Claim (1) is well supported by Table 2: the best prompt run (naverloo-rgpt4, NDCG@10 0.6994) exceeds the best nnlm run (naverloo_fs_RR, 0.5972), and the comparison is descriptive. Claim (2) is the more novel and more load-bearing contribution, since it supports using synthetic queries for test collection construction. The reader's weakest assumption—that the surviving 31 synthetic queries are representative—is correct and is the main soft spot. The paper explicitly acknowledges the filtering step and calls the result 'initial,' but the abstract states the agreement without this qualifier. Because the filtering is designed to remove queries that are noisy or uninformative, the estimate is likely optimistic. This does not invalidate the paper; it means the claim should be presented as contingent on human curation, which is exactly the CONDITIONAL verdict. No other concern is as load-bearing: the prompt-vs-nnlm comparison is clearly labeled as self-reported run types and is not over-interpreted. Therefore we agree with the reader and recommend no change to the verdict.","tokens_in":14698,"tokens_out":4467,"duration_ms":47208,"concrete_test":"Recompute the Kendall tau and top-system rankings using all 97 synthetic queries that were submitted to the NIST assessors (48 T5 + 49 GPT-4), including those that were rejected, provided their relevance judgments were recorded; if the rejected queries were not judged, generate a fresh unfiltered set of synthetic queries from the same pipeline and have them judged at the same depth. If tau drops materially (e.g., below 0.8) or the top-5 system order changes, the 0.8487 estimate is an artifact of the curator. Additionally, bootstrap the 31 accepted queries (10,000 resamples) and report a 95% CI for tau; a lower bound below 0.7 would show the reliability claim is not statistically supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 5 reports that NIST assessors selected only 13 of 48 T5-generated and 18 of 49 GPT-4-generated queries for the test collection, alongside 51 human queries. The abstract's headline agreement tau=0.8487 (Figure 5, right) is computed on these 31 surviving synthetic queries. The selection is explicitly not random: assessors removed queries that 'do not look reasonable' or have 'too few or too many relevant documents' because they are 'noisy or not very informative for evaluation purposes.' If this human filter preferentially removes ambiguous, difficult, or otherwise atypical synthetic queries—precisely the cases where synthetic and human judgments diverge—the observed Kendall tau overstates the agreement that an unfiltered synthetic-query test collection would achieve. The paper's conclusion that 'test collections consisting of synthetically generated queries could be reliably used' therefore depends on an unstated and untested premise: that the filtering step is either unnecessary or does not change the system ranking. No confidence interval or significance test is reported for tau, and with n=31 queries the estimate is fragile. The abstract phrase 'human effort was needed to select a subset' undercuts the substitution claim, since a fully synthetic pipeline would need to automate or eliminate that curation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper is the overview of the TREC 2023 Deep Learning Track. It describes the passage and document ranking tasks, the MS MARCO v2 datasets, the submitted runs, and the main evaluation results. The authors report two headline findings: (1) runs that use large language model (LLM) prompting in some part of the pipeline outperformed runs using the previous best \"nnlm\" approach, and (2) evaluation using synthetically generated queries (from T5 and GPT-4) produced system orderings similar to those from human queries, with a Kendall tau of 0.8487. The paper also analyzes potential bias from using T5 or GPT-4 generated queries toward systems based on the same model family, finding no clear bias. The track is the final year of the Deep Learning Track, and the paper emphasizes the construction of a reusable test collection.","tokens_in":14961,"tokens_out":4831,"duration_ms":48653,"significance":"If the findings hold, they are of considerable significance to the IR community. The prompt-vs-nnlm result suggests that LLM prompting is now a state-of-the-art technique on this benchmark, while the synthetic-query analysis speaks directly to the cost and feasibility of building test collections without human queries. The paper is transparent about its query generation pipeline, including the prompts used, the filtering rates, and the final query counts, and it makes the official NIST judgments the basis of all reported metrics. These are genuine strengths. However, the synthetic-query claim is based on a small, human-filtered subset of the generated queries, and the prompt-vs-nnlm comparison relies on self-classified run categories without significance testing. Both limitations constrain the strength of the conclusions as stated in the abstract, although the paper itself uses cautious language in places (e.g., \"initial results suggest\" and \"more analysis is needed\").","major_comments":[{"comment":"The headline agreement tau = 0.8487 is computed on only 31 synthetic queries (13 T5-generated and 18 GPT-4-generated) that survived a human filtering step. The paper states that assessors removed queries that \"do not look reasonable\" or that contain \"too few or too many relevant documents\" because they are \"noisy or not very informative for evaluation purposes.\" This filtering is not random, and it plausibly removes exactly the difficult, ambiguous, or atypical synthetic queries where system agreement with human-query evaluation could be lowest. The reported Kendall tau therefore may substantially overstate the agreement that an unfiltered synthetic-query test collection would achieve. The conclusion in Section 5 that \"test collections consisting of synthetically generated queries could be reliably used\" is not supported for unfiltered synthetic generation. The authors should either report the agreement on the full set of generated queries (if judgments exist), provide a confidence interval or significance test for the tau value, or explicitly qualify the claim to apply only to synthetic queries that pass a human quality filter. The abstract's substitution claim currently goes beyond what the data show.","section":"Section 4, Table 2, Figure 2"},{"comment":"The headline claim about prompt runs outperforming nnlm runs is not supported by significance testing, and the self-classified run categories plus the concentration of top runs in two groups make the comparison less controlled than the abstract implies.","section":"Section 4, Table 2, Figure 2"}],"minor_comments":[{"comment":"There is a typo in \"one or mpre phases\" (should be \"more\"), and the sentence \"The best 'prompt' run outperforms the best 'nnlm' on on the majority of queries\" has a duplicated \"on.\"","section":"Section 4"},{"comment":"The caption says \"As in the previous two years, 'nnlm' runs continue to outperform over 'trad' runs for both tasks,\" but the figure includes the 'prompt' category; the caption should describe all three run types shown.","section":"Figure 2 caption"},{"comment":"The caption refers to \"mean performance between 'prompt' and 'llm' runs,\" but the correct category name is \"nnlm\".","section":"Figure 4 caption"},{"comment":"Table 6 reports the final number of queries per type (82 total) but not the numbers of queries initially generated and provided to assessors (200 human, 250 T5, 250 GPT-4, of which 147/48/49 were provided to assessors). Adding the initial counts and the selection rates would make the filtering step clearer.","section":"Section 5, Table 6"},{"comment":"The sentence \"For all query types, depth-10 pooling was used to select the documents to be judged by the NIST assessors\" is potentially confusing because the track judged passages, not documents, and document labels were inferred from passage labels; please clarify that the depth-10 pooling applies to passage pooling, with document labels propagated afterward.","section":"Section 5"},{"comment":"The Conclusion contains a typo: \"we repeated the updats that were first introduced last year\" should be \"updates.\"","section":"Conclusion"}],"recommendation":"major_revision","confidential_remarks":"The paper is a TREC overview, and the community may accept observational results without significance tests as a matter of convention. However, the synthetic-query claim is central to the abstract and is based on only 31 human-filtered queries; this needs either additional analysis or a substantially more qualified statement. The prompt-vs-nnlm claim, while supported by the official run tables in the narrow sense, would also benefit from a significance test or a clear caveat that the comparison is not controlled. The track being final makes it unlikely that additional data will be collected, so the revision should focus on appropriately scoping the claims rather than expecting new experiments."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the synthetic-query headline (tau = 0.8487) is the weakest part of an otherwise solid track overview. The stress-test note is right: that number comes from 13 T5 + 18 GPT-4 queries that survived human filtering, and the filtering is not random—assessors removed 'noisy' or 'unreasonable' queries. So the agreement with human queries is likely inflated. The paper itself is more cautious than the abstract ('initial results suggest', 'more analysis needed'), but the abstract oversells it.\n\nWhat's genuinely useful: the track delivers a new, reusable test collection built on held-out MS MARCO v2 queries, with deduplication and passage-to-document labels. The prompt-vs-nnlm result is descriptively interesting—top runs this year used LLM prompting—but it's not a controlled comparison: runs are self-classified, there are no significance tests, and the five passage-ranking groups include two that supplied most of the prompt runs. Still, as a data point it's valuable, and participant papers give details.\n\nThe synthetic-query bias analysis is the most novel piece: they checked whether T5- or GPT-4-generated queries favor systems using the same model, and found no clear evidence. That's a useful negative result, but again, on a handful of queries.\n\nWhere it's soft: besides the n=31 issue, there's no uncertainty quantification on Kendall tau, and the human filter is a confound the paper acknowledges but doesn't address. Also, the paper mostly repeats last year's design, so novelty is incremental. The references are appropriate; self-citations are for the track's own prior overviews, which is fine.\n\nBottom line: this is a resource paper for IR evaluation folks, not a methodological breakthrough. It deserves a serious referee—the data and the bias analysis are worth publishing—but the authors should soften the substitution claim and add a confidence interval or at least a caveat about the filter. I'd accept it with minor revisions.","headline":"Solid track overview; the synthetic-query reliability claim in the abstract is overstated given it rests on only 31 human-filtered queries.","tokens_in":15503,"tokens_out":2360,"would_cite":true,"duration_ms":25265,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LLM-prompted rerankers outscored the four-year-best nnlm approach at TREC 2023, and human-filtered synthetic queries reproduced human-query system rankings with Kendall tau 0.8487.","keywords":["TREC Deep Learning track","passage ranking","document ranking","MS MARCO v2","synthetic queries","LLM prompting","test collection reusability","Kendall tau"],"falsifier":"Take all 97 T5- and GPT-4-generated queries that went to the assessors, including the 66 rejected ones, judge them under the same relevance-assessment protocol, and recompute Kendall's tau between system NDCG@10 on those synthetic queries and on the 51 human queries; if the tau drops materially below 0.8487, the reported agreement is an artifact of the human filter rather than a property of synthetic queries in general.","tokens_in":14531,"feed_emoji":"📊","tokens_out":10887,"duration_ms":112600,"temperature":0.7,"pith_summary":"On its fifth and final run, the TREC Deep Learning track set out to build another reusable test collection on the MS MARCO v2 passage and document corpora and to ask whether synthetic queries could stand in for real user queries during evaluation. The track reports two headline results. First, runs that use an LLM through prompting outperformed runs built on the nnlm recipe, a fine-tuned neural language model stack that was the best approach for the previous four years; the best prompt run scores NDCG@10 0.6994 on passage ranking. Second, after human assessors filtered unusable synthetic queries, evaluating systems on the remaining 31 T5- and GPT-4-generated queries agreed closely with evaluation on 51 human queries, with a Kendall tau of 0.8487, and the track found no clear sign that a query generator favored systems built with the same model. The paper presents this as evidence that synthetic queries can be a reliable component of test collection construction, while cautioning that human filtering was required.","feed_headline":"LLM-prompted rerankers beat four-year best at TREC 2023","feed_subtitle":"Filtered synthetic queries match human-query system rankings at tau 0.8487, a step toward cheaper evaluation.","key_machinery":"The load-bearing mechanism is the reusable test collection: 82 judged queries spanning three query types, with four-point relevance judgments and expanded qrels that copy a judged passage's label to every near-duplicate passage in its cluster. The comparison engine is NDCG@10, with Kendall's tau measuring whether systems keep their order when human queries are swapped for synthetic ones; tau = 0.8487 is the number that carries the synthetic-query conclusion. The synthetic pipeline has four stages: a GPT-4 prompt scores sampled passages for self-containedness, T5 generates many candidate queries stratified to match human query length and lexical overlap, GPT-4 generates one zero-shot query per seed passage, and human assessors reject most candidates, 66 of 97, before judging. The prompt-run result rests on comparing prompt-classified runs against nnlm runs at the system level and on per-query comparisons.","core_discovery":"The paper's central claim is that prompt-based LLM ranking overtook fine-tuned nnlm stacks as the best-performing approach on the TREC 2023 Deep Learning track, and that synthetic queries, once filtered by human assessors, rank systems almost as well as real user queries do. The track judged 82 passage queries on the MS MARCO v2 collections: 51 from held-out human queries, 13 produced by a fine-tuned T5 query generator, and 18 by a GPT-4 prompt; labels were then propagated from passages to source documents for the document task, and expanded qrels spread each judgment through its cluster of near-duplicate passages. The best prompt run reaches NDCG@10 of 0.6994 on passage ranking versus 0.5972 for the best nnlm run, and the prompt methods win on most individual queries. Across submitted systems, Kendall's tau between system orderings on human queries and on synthetic queries is 0.8487, and 0.9395 when the collection is evaluated on all real plus synthetic track queries. The authors conclude that synthetic queries can be reliably used in test collection construction, but only after human selection, and they find no clear evidence that GPT-4-generated queries inflate GPT-based systems or that T5-generated queries inflate T5-based systems.","pith_inferences":["Editorial inference: the 0.8487 tau is measured on only 31 accepted synthetic queries, and 66 of 97 generated queries were discarded by human assessors, so the practical cost saving depends on automating or outsourcing that filter, which this paper does not attempt.","Editorial inference: GPT-4 queries are nearly twice as long as human queries, and both synthetic types return fewer relevant documents per query, so a synthetic-query collection is a harder and somewhat different evaluation surface, not a free replacement for the historical human-query series.","Editorial inference: the no-bias conclusion is drawn from very few systems in each family, for example one document-ranking prompt run, so the absence of clear bias is a weak bound; a larger paired sampling of generators and system families could detect smaller or conditional biases."],"forward_implications":["Prompt-based LLM ranking now becomes the reference point for retrieval track benchmarks, replacing the fine-tuned nnlm stacks that had dominated the previous four years.","Test collection builders can generate candidate queries with T5 or GPT-4 and still obtain system orderings close to those from human queries, provided a human quality filter is applied first.","Passage-level judgments remain sufficient for document ranking: propagating passage labels to source documents yields usable document qrels without separate document judging.","With two consecutive years of harder held-out query sets, year-over-year comparisons on this collection are more discriminative than the earlier MS MARCO test queries.","Even the best prompt runs mostly self-reported using MS MARCO training data somewhere in the stack, so prompting does not remove the need for fine-tuned retrieval components."],"supporting_citations":[{"why":"Introduces the MS MARCO dataset whose human-annotated training labels and corpora the entire track benchmark builds on.","marker":"[Bajaj et al., 2016]"},{"why":"The 2022 track overview that introduced the harder held-out query sampling, near-duplicate expansion, and passage-only judging design this year repeats.","marker":"[Craswell et al., 2023]"},{"why":"Documents the too-many-relevants judgment budget problem and the 2021 collection issues that motivate the harder, more reusable test design.","marker":"[V oorhees et al., 2022, Craswell et al., 2022]"},{"why":"Validates passage-to-document label propagation on the 2021 test collection, the method used to form document qrels here.","marker":"[Craswell et al., 2022]"},{"why":"Supplies the shared top-100 candidate sets for both reranking subtasks, making reranking runs directly comparable.","marker":"[Lin et al., 2021b]"},{"why":"Provides the evidence that passage-level relevance judgments can be transferred to document-level relevance.","marker":"[Wu et al., 2019]"},{"why":"Defines NDCG, the metric used for all run comparisons and for the Kendall tau agreement between query types.","marker":"[Järvelin and Kekäläinen, 2002]"}],"fun_headline_variants":["LLM prompts dethrone nnlm at TREC Deep Learning 2023","Synthetic queries rival human ones for ranking eval: TREC 2023","TREC 2023: Prompted LLMs top four-year champion","LLM rerankers surpass nnlm; synthetic queries rank systems alike"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 31 synthetic queries that survived human filtering are representative of synthetic queries in general; if the filter removes exactly the difficult, ambiguous, or biased queries, the observed agreement of tau = 0.8487 will not carry over to unfiltered synthetic queries.","fun_headline_variants_meta":{"raw":{"variants":["LLM prompts dethrone nnlm at TREC Deep Learning 2023","Synthetic queries rival human ones for ranking eval: TREC 2023","TREC 2023: Prompted LLMs top four-year champion","LLM rerankers surpass nnlm; synthetic queries rank systems alike"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000216,"raw_usage":{"total_tokens":1532,"prompt_tokens":1142,"completion_tokens":390,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":758,"completion_tokens_details":{"reasoning_tokens":308}},"tokens_in":758,"tokens_out":390,"duration_ms":3791,"temperature":1.0,"reasoning_tokens":308,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T18:25:37.917055+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take all 97 T5- and GPT-4-generated queries that went to the assessors, including the 66 rejected ones, judge them under the same relevance-assessment protocol, and recompute Kendall's tau between system NDCG@10 on those synthetic queries and on the 51 human queries; if the tau drops materially below 0.8487, the reported agreement is an artifact of the human filter rather than a property of synthetic queries in general.","supporting_citations":[],"review_version":1}