Pith. sign in

REVIEW 1 cited by

AugTriever: Unsupervised Dense Retrieval and Domain Adaptation by Scalable Data Augmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.08841 v4 pith:BEEEVGNL submitted 2022-12-17 cs.CL cs.IR

classification cs.CLcs.IR
keywords densequeryretrievalunsupervisedgenerationmodelspseudoadaptation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Dense retrievers have made significant strides in text retrieval and open-domain question answering. However, most of these achievements have relied heavily on extensive human-annotated supervision. In this study, we aim to develop unsupervised methods for improving dense retrieval models. We propose two approaches that enable annotation-free and scalable training by creating pseudo querydocument pairs: query extraction and transferred query generation. The query extraction method involves selecting salient spans from the original document to generate pseudo queries. On the other hand, the transferred query generation method utilizes generation models trained for other NLP tasks, such as summarization, to produce pseudo queries. Through extensive experimentation, we demonstrate that models trained using these augmentation methods can achieve comparable, if not better, performance than multiple strong dense baselines. Moreover, combining these strategies leads to further improvements, resulting in superior performance of unsupervised dense retrieval, unsupervised domain adaptation and supervised finetuning, benchmarked on both BEIR and ODQA datasets. Code and datasets are publicly available at https://github.com/salesforce/AugTriever.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. More Than Efficiency: Embedding Compression Improves Domain Adaptation in Dense Retrieval

    cs.IR 2026-01 conditional novelty 6.0 of 10

    Query-only PCA compression improves retrieval NDCG@10 in most tested model-dataset pairs at 90% retention, with gains largest in structured domains like SpartQA.

Pith tools