Pith. sign in

REVIEW 3 cited by

WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.14727 v1 pith:JCLFS6KU submitted 2025-02-20 cs.SD cs.AIeess.AS

classification cs.SDcs.AIeess.AS
keywords retrievalwavragaudioaugmentedgenerationknowledgemodelsdialogue
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Retrieval Augmented Generation (RAG) has gained widespread adoption owing to its capacity to empower large language models (LLMs) to integrate external knowledge. However, existing RAG frameworks are primarily designed for text-based LLMs and rely on Automatic Speech Recognition to process speech input, which discards crucial audio information, risks transcription errors, and increases computational overhead. Therefore, we introduce WavRAG, the first retrieval augmented generation framework with native, end-to-end audio support. WavRAG offers two key features: 1) Bypassing ASR, WavRAG directly processes raw audio for both embedding and retrieval. 2) WavRAG integrates audio and text into a unified knowledge representation. Specifically, we propose the WavRetriever to facilitate the retrieval from a text-audio hybrid knowledge base, and further enhance the in-context capabilities of spoken dialogue models through the integration of chain-of-thought reasoning. In comparison to state-of-the-art ASR-Text RAG pipelines, WavRAG achieves comparable retrieval performance while delivering a 10x acceleration. Furthermore, WavRAG's unique text-audio hybrid retrieval capability extends the boundaries of RAG to the audio modality.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style

    eess.AS 2025-08 conditional novelty 6.0 of 10

    A contrastively trained speech-text model retrieves speech segments from free-form text descriptions of speaking style across 22 categories.

  2. Event-Grounded Question Answering over Long Audio via Structured Retrieval

    eess.AS 2026-02 reject novelty 5.0 of 10

    LA-RAG stores timestamped audio events from a grounding model in SQL and uses intent-aware retrieval plus an LLM to answer long-audio questions, claiming 76.88% accuracy on synthetic home audio.

  3. SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding

    cs.CV 2025-06 conditional novelty 5.0 of 10

    SceneRAG uses LLM-driven scene segmentation and a scene-level knowledge graph to retrieve and answer questions about long videos, reporting higher LLM-judged win-rates than chunk-based RAG baselines on the LongerVideo...

Pith tools