Pith. sign in

REVIEW 2 cited by

Prompting the Hidden Talent of Web-Scale Speech Models for Zero-Shot Task Generalization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.11095 v3 pith:F4NTWXVS submitted 2023-05-18 eess.AS cs.AIcs.CLcs.LGcs.SD

classification eess.AScs.AIcs.CLcs.LGcs.SD
keywords promptsspeechtasksdefaultexperimentsmodelmodelsproposed
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We investigate the emergent abilities of the recently proposed web-scale speech model Whisper, by adapting it to unseen tasks with prompt engineering. We selected three tasks: audio-visual speech recognition (AVSR), code-switched speech recognition (CS-ASR), and speech translation (ST) on unseen language pairs. We design task-specific prompts, by either leveraging another large-scale model, or simply manipulating the special tokens in the default prompts. Experiments show that compared to the default prompts, our proposed prompts improve performance by 10% to 45% on the three zero-shot tasks, and even outperform SotA supervised models on some datasets. In addition, our experiments reveal many interesting properties of Whisper, including its robustness to prompts, bias on accents, and the multilingual understanding in its latent space. Code is available at https://github.com/jasonppy/PromptingWhisper

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Refining Transcripts With TV Subtitles by Prompt-Based Weakly Supervised Training of ASR

    cs.CL 2025-09 conditional novelty 6.0 of 10

    Using subtitles as prompts and pseudo transcripts as targets, with a Gini-based attention weighting, refines Whisper's transcripts on low-resource Flemish TV speech without verbatim labels.

  2. Optimizing ASR for Catalan-Spanish Code-Switching: A Comparative Analysis of Methodologies

    cs.CL 2025-07 conditional novelty 4.0 of 10

    For Catalan-Spanish code-switching ASR, fine-tuning Whisper on 17 hours of synthetic TTS data and decoding with the Catalan token outperforms audio concatenation and in-domain fine-tuning on held-out parliamentary test sets.

Pith tools