Pith. sign in

REVIEW 3 cited by

Training Data is More Valuable than You Think: A Simple and Effective Method by Retrieving from Training Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.08773 v1 pith:7SM7UZFR submitted 2022-03-16 cs.CL cs.AIcs.IR

classification cs.CLcs.AIcs.IR
keywords taskstrainingdatamethodretrievingeffectiveinputreina
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Retrieval-based methods have been shown to be effective in NLP tasks via introducing external knowledge. However, the indexing and retrieving of large-scale corpora bring considerable computational cost. Surprisingly, we found that REtrieving from the traINing datA (REINA) only can lead to significant gains on multiple NLG and NLU tasks. We retrieve the labeled training instances most similar to the input text and then concatenate them with the input to feed into the model to generate the output. Experimental results show that this simple method can achieve significantly better performance on a variety of NLU and NLG tasks, including summarization, machine translation, language modeling, and question answering tasks. For instance, our proposed method achieved state-of-the-art results on XSum, BigPatent, and CommonsenseQA. Our code is released, https://github.com/microsoft/REINA .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TabPack: Efficient Hyperparameter Ensembles for Tabular Deep Learning

    cs.LG 2026-07 accept novelty 6.0 of 10

    TabPack packs MLPs with diverse sampled hyperparameters into one vectorized model, selects ensemble members online during training, and matches tuned baselines at a fraction of the compute cost.

  2. RELexED: Retrieval-Enhanced Legal Summarization with Exemplar Diversity

    cs.CL 2025-01 conditional novelty 5.0 of 10

    Diverse, influence-function-scored exemplar summaries retrieved with a DPP improve legal summarization over no-exemplar and similarity-only baselines on SuperSCOTUS and CivilSum, with modest and statistically partial gains.

  3. A Survey on Retrieval And Structuring Augmented Generation with Large Language Models

    cs.CL 2025-09 conditional novelty 2.0 of 10

    The paper presents a comprehensive survey and taxonomy of RAS methods, covering retrieval, text structuring, and LLM integration.

Pith tools