Pith. sign in

REVIEW 4 cited by

Can ChatGPT Detect Intent? Evaluating Large Language Models for Spoken Language Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.13512 v2 pith:PCJAQTJV submitted 2023-05-22 cs.CL cs.AIcs.SDeess.AS

classification cs.CLcs.AIcs.SDeess.AS
keywords modelslanguagechatgptunderstandingintentlargespokenabilities
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, large pretrained language models have demonstrated strong language understanding capabilities. This is particularly reflected in their zero-shot and in-context learning abilities on downstream tasks through prompting. To assess their impact on spoken language understanding (SLU), we evaluate several such models like ChatGPT and OPT of different sizes on multiple benchmarks. We verify the emergent ability unique to the largest models as they can reach intent classification accuracy close to that of supervised models with zero or few shots on various languages given oracle transcripts. By contrast, the results for smaller models fitting a single GPU fall far behind. We note that the error cases often arise from the annotation scheme of the dataset; responses from ChatGPT are still reasonable. We show, however, that the model is worse at slot filling, and its performance is sensitive to ASR errors, suggesting serious challenges for the application of those textual models on SLU.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Retrieval-Augmented Generation in LLMs for Mental Health: Quantifying the Incremental Contribution of Retrieval Within a Layered Safety Architecture

    cs.IR 2026-07 conditional novelty 6.0 of 10

    Retrieval augmentation improves mental-health chatbot intent classification for 4 of 6 tested LLMs, mainly by catching more high-risk cases, at the cost of more false alarms.

  2. Leveraging Information Retrieval to Enhance Spoken Language Understanding Prompts in Few-Shot Learning

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Using BM25 lexical retrieval to select prompt examples improves few-shot slot-filling F1 scores on ATIS, SNIPS, SLURP, and MEDIA compared to random or intent-only selection.

  3. Do Large Language Models Need Intent? Revisiting Response Generation Strategies for Service Assistant

    cs.CL 2025-09 reject novelty 4.0 of 10

    Adding a fine-tuned intent recognition step before response generation improves automatic quality scores on two customer-service datasets, but the results are reported without statistical or human validation.

  4. A Computational Approach to Modeling Conversational Systems: Analyzing Large-Scale Quasi-Patterned Dialogue Flows

    cs.CL 2025-07 reject novelty 4.0 of 10

    A dialogue-flow pipeline that embeds, clusters, labels, and prunes conversations into trees, evaluated with a semantic metric that largely reflects its own clustering choices.

Pith tools