Pith. sign in

REVIEW 1 cited by

DUAL: Discrete Spoken Unit Adaptive Learning for Textless Spoken Question Answering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.04911 v3 pith:VL3US4GC submitted 2022-03-09 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords spokendatadualquestionwordsadaptiveansweringanswers
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Spoken Question Answering (SQA) is to find the answer from a spoken document given a question, which is crucial for personal assistants when replying to the queries from the users. Existing SQA methods all rely on Automatic Speech Recognition (ASR) transcripts. Not only does ASR need to be trained with massive annotated data that are time and cost-prohibitive to collect for low-resourced languages, but more importantly, very often the answers to the questions include name entities or out-of-vocabulary words that cannot be recognized correctly. Also, ASR aims to minimize recognition errors equally over all words, including many function words irrelevant to the SQA task. Therefore, SQA without ASR transcripts (textless) is always highly desired, although known to be very difficult. This work proposes Discrete Spoken Unit Adaptive Learning (DUAL), leveraging unlabeled data for pre-training and fine-tuned by the SQA downstream task. The time intervals of spoken answers can be directly predicted from spoken documents. We also release a new SQA benchmark corpus, NMSQA, for data with more realistic scenarios. We empirically showed that DUAL yields results comparable to those obtained by cascading ASR and text QA model and robust to real-world data. Our code and model will be open-sourced.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VoxRAG: A Step Toward Transcription-Free RAG Systems in Spoken Question Answering

    cs.IR 2025-05 conditional novelty 5.0 of 10

    VoxRAG shows that a spoken query can retrieve topically relevant podcast segments via CLAP audio embeddings and FAISS search, with Recall@10 of 0.60 for somewhat relevant segments, though precise answers remain rare.

Pith tools