REVIEW 4 cited by
Can ChatGPT Detect Intent? Evaluating Large Language Models for Spoken Language Understanding
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recently, large pretrained language models have demonstrated strong language understanding capabilities. This is particularly reflected in their zero-shot and in-context learning abilities on downstream tasks through prompting. To assess their impact on spoken language understanding (SLU), we evaluate several such models like ChatGPT and OPT of different sizes on multiple benchmarks. We verify the emergent ability unique to the largest models as they can reach intent classification accuracy close to that of supervised models with zero or few shots on various languages given oracle transcripts. By contrast, the results for smaller models fitting a single GPU fall far behind. We note that the error cases often arise from the annotation scheme of the dataset; responses from ChatGPT are still reasonable. We show, however, that the model is worse at slot filling, and its performance is sensitive to ASR errors, suggesting serious challenges for the application of those textual models on SLU.
Forward citations
Cited by 4 Pith papers
-
Retrieval-Augmented Generation in LLMs for Mental Health: Quantifying the Incremental Contribution of Retrieval Within a Layered Safety Architecture
Retrieval augmentation improves mental-health chatbot intent classification for 4 of 6 tested LLMs, mainly by catching more high-risk cases, at the cost of more false alarms.
-
Leveraging Information Retrieval to Enhance Spoken Language Understanding Prompts in Few-Shot Learning
Using BM25 lexical retrieval to select prompt examples improves few-shot slot-filling F1 scores on ATIS, SNIPS, SLURP, and MEDIA compared to random or intent-only selection.
-
Do Large Language Models Need Intent? Revisiting Response Generation Strategies for Service Assistant
Adding a fine-tuned intent recognition step before response generation improves automatic quality scores on two customer-service datasets, but the results are reported without statistical or human validation.
-
A Computational Approach to Modeling Conversational Systems: Analyzing Large-Scale Quasi-Patterned Dialogue Flows
A dialogue-flow pipeline that embeds, clusters, labels, and prunes conversations into trees, evaluated with a semantic metric that largely reflects its own clustering choices.
Discussion (0). Sign in to comment.