Pith. sign in

REVIEW 1 cited by

SpeechDPR: End-to-End Spoken Passage Retrieval for Open-Domain Spoken Question Answering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.13463 v3 pith:AWSSPVM3 submitted 2024-01-24 cs.CL cs.IRcs.SDeess.AS

SpeechDPR: End-to-End Spoken Passage Retrieval for Open-Domain Spoken Question Answering

classification cs.CL cs.IRcs.SDeess.AS
keywords spokenpassagequestionspeechspeechdpruasransweranswering
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Spoken Question Answering (SQA) is essential for machines to reply to user's question by finding the answer span within a given spoken passage. SQA has been previously achieved without ASR to avoid recognition errors and Out-of-Vocabulary (OOV) problems. However, the real-world problem of Open-domain SQA (openSQA), in which the machine needs to first retrieve passages that possibly contain the answer from a spoken archive in addition, was never considered. This paper proposes the first known end-to-end framework, Speech Dense Passage Retriever (SpeechDPR), for the retrieval component of the openSQA problem. SpeechDPR learns a sentence-level semantic representation by distilling knowledge from the cascading model of unsupervised ASR (UASR) and text dense retriever (TDR). No manually transcribed speech data is needed. Initial experiments showed performance comparable to the cascading model of UASR and TDR, and significantly better when UASR was poor, verifying this approach is more robust to speech recognition errors.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. LuxSQA: Ask Me in Luxembourgish with TTS-Augmented Spoken Question Answering

    cs.CL 2026-07 conditional novelty 5.0

    Multi-source and voice-design TTS training data yield the strongest Luxembourgish spoken QA on real speakers, while no-reference MOS scores fail to rank systems by QA utility.