Pith. sign in

REVIEW 10 cited by

Reliable, Adaptable, and Attributable Language Models with Retrieval

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.03187 v1 pith:NZLTJ745 submitted 2024-03-05 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords retrieval-augmentedadaptableattributabledatadatastoresinferenceinfrastructureinteraction
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Parametric language models (LMs), which are trained on vast amounts of web data, exhibit remarkable flexibility and capability. However, they still face practical challenges such as hallucinations, difficulty in adapting to new data distributions, and a lack of verifiability. In this position paper, we advocate for retrieval-augmented LMs to replace parametric LMs as the next generation of LMs. By incorporating large-scale datastores during inference, retrieval-augmented LMs can be more reliable, adaptable, and attributable. Despite their potential, retrieval-augmented LMs have yet to be widely adopted due to several obstacles: specifically, current retrieval-augmented LMs struggle to leverage helpful text beyond knowledge-intensive tasks such as question answering, have limited interaction between retrieval and LM components, and lack the infrastructure for scaling. To address these, we propose a roadmap for developing general-purpose retrieval-augmented LMs. This involves a reconsideration of datastores and retrievers, the exploration of pipelines with improved retriever-LM interaction, and significant investment in infrastructure for efficient training and inference.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RH-RAG: Trustworthy Long-Form Generation for Privacy-Constrained Settings

    cs.CL 2026-08 reject novelty 6.0 of 10

    A multi-agent RAG framework that adds planning, bounded memory, and NLI-based revision to local 7-8B models, reported to improve faithfulness and coherence in long-form generation.

  2. Med-R$^3$: Enhancing Medical Retrieval-Augmented Reasoning of LLMs via Progressive Reinforcement Learning

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A progressive three-stage RL framework with medical-specific rewards improves retrieval-augmented reasoning on medical QA benchmarks, reportedly surpassing GPT-4o-mini with an 8B model.

  3. SUCEA: Reasoning-Intensive Retrieval for Adversarial Fact-checking through Claim Decomposition and Editing

    cs.CL 2025-06 conditional novelty 6.0 of 10

    SUCEA improves adversarial fact-checking by decomposing claims into atomic sub-claims, editing each sub-claim toward retrieved evidence, and re-retrieving before predicting the final label.

  4. RaDeR: Reasoning-aware Dense Retrieval Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A math-trained dense retriever and reranker, built from MCTS reasoning trajectories and self-reflection, outperforms strong baselines on reasoning-intensive retrieval benchmarks and beats BM25 on chain-of-thought queries.

  5. Safety Degradation in AI Agents

    cs.CY 2025-05 conditional novelty 6.0 of 10

    Adding retrieval to aligned LLMs degrades safety: refusal rates fall, bias and harmfulness rise, and prompt-based mitigation only partially restores alignment.

  6. Adaptable Non-parametric Approach for Speech-based Symptom Assessment: Isolating Private Medical Data in a Retrieval Datastore

    eess.AS 2025-06 conditional novelty 5.0 of 10

    NoNPSA shows that retrieval from a datastore of speech embeddings can assess respiratory symptoms as accurately as fine-tuned self-supervised models, while making data updates and removal easier.

  7. Evaluating the Retrieval Robustness of Large Language Models

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A realistic benchmark with three metrics shows modern LLMs are generally robust to imperfect retrieval, though not perfectly so.

  8. Small Encoders Can Rival Large Decoders in Detecting Groundedness

    cs.CL 2025-06 conditional novelty 4.0 of 10

    Task-specific encoders (e.g., RoBERTa-large) rival large decoders such as Llama-3-8B and GPT-4o on binary groundedness detection, within 5 to 10 accuracy points while requiring one to three orders of magnitude fewer FLOPs.

  9. CCRS: A Zero-Shot LLM-as-a-Judge Framework for Comprehensive RAG Evaluation

    cs.CL 2025-06 conditional novelty 4.0 of 10

    CCRS is a zero-shot LLM-as-a-judge framework whose five metrics discriminate between RAG systems on BioASQ with comparable or better power than RAGChecker at lower compute.

  10. AI-Driven Climate Policy Scenario Generation for Sub-Saharan Africa

    cs.AI 2025-05 conditional novelty 4.0 of 10

    A RAG pipeline using llama3.2-3B and UN COP documents generated 34 policy scenarios for Sub-Saharan Africa, 30 passed author validation, but automated evaluation showed mixed agreement with human judgment.

Pith tools