Pith. sign in

REVIEW 1 cited by

UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of Rerankers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.00807 v3 pith:2AI3B6MI submitted 2023-03-01 cs.IR cs.CL

classification cs.IRcs.CL
keywords largedomainqueriessyntheticdatasetsexpensivemethodmodels
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Many information retrieval tasks require large labeled datasets for fine-tuning. However, such datasets are often unavailable, and their utility for real-world applications can diminish quickly due to domain shifts. To address this challenge, we develop and motivate a method for using large language models (LLMs) to generate large numbers of synthetic queries cheaply. The method begins by generating a small number of synthetic queries using an expensive LLM. After that, a much less expensive one is used to create large numbers of synthetic queries, which are used to fine-tune a family of reranker models. These rerankers are then distilled into a single efficient retriever for use in the target domain. We show that this technique boosts zero-shot accuracy in long-tail domains and achieves substantially lower latency than standard reranking methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluating the Robustness of Dense Retrievers in Interdisciplinary Domains

    cs.IR 2025-06 reject novelty 4.0 of 10

    Domain adaptation gains for ColBERTv2 appear 3.6 times larger on a benchmark with overlapping topics, but the benchmark also differs in corpus size, question-per-context ratio, and baseline headroom, so the causal att...

Pith tools