Pith. sign in

REVIEW 4 major objections 5 minor 48 references

An Interdisciplinary Review of Commonsense Reasoning and Intent Detection

T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read This review of 28 papers from ACL, EMNLP, and CHI (2020–2025) argues that commonsense reasoning and intent detection are converging on adaptive, zero-shot, and human-centered methods, with unresolved gaps in grounding, generalization, and…

desk verdict Readable survey with a genuinely useful HCI tilt, but the stated corpus doesn't match Table 1, the COMET attribution is wrong, and the abstract/table count doesn't add up. read the letter →

arxiv 2506.14040 v1 pith:S6BDUW6Y submitted 2025-06-16 cs.CL cs.HC

classification cs.CLcs.HC
keywords commonsensereasoningintentdetectionnaturallanguageunderstandingzero-shotlearningculturaladaptationmultilingualevaluationbenchmarkdesignhuman-computerinteraction
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This review synthesizes 28 papers from NLP and HCI venues published between 2020 and 2025 to show how commonsense reasoning and intent detection are changing. Its central claim is that both fields are moving away from fully supervised, fixed-label systems toward zero-shot, generative, contrastive, and human-centered approaches that aim to handle unseen intents and culturally varied situations. The paper organizes the literature into two themes—commonsense reasoning and intent detection—each with four methodological subthemes, and uses this structure to highlight a shared set of open problems: knowledge that is not grounded in specific contexts, English-centric biases in multilingual resources, and benchmarks that can be gamed by shallow heuristics. A sympathetic reader would take the main contribution to be the cross-venue map of trends and gaps, offered so that future systems can balance robustness, cultural sensitivity, and task specificity.

What carries the argument

The machinery is the review's two-theme, four-subtheme taxonomy. Commonsense reasoning is split into (i) self-supervised and zero-shot learning, (ii) multilingual and cultural adaptation, (iii) structured reasoning and evaluation analysis, and (iv) interactive, dialog-based, and applied commonsense; intent detection is split into (i) open-set and zero-shot detection, (ii) multi-intent modeling and generative formulation, (iii) contrastive learning and clustering, and (iv) human-centered and HCI applications. Each reviewed paper is assigned to a cell based on methodology (graph-based, generative, prompting, or hybrid) and reasoning type (causal, dialogic, social). The taxonomy does the work of turning 28 individual findings into trend claims—a shift away from supervised learning—and gap claims about grounding, generalization, and benchmark reliability. Named systems such as COMET, DrFact, CICERO, ExplaGraphs, AGIF, LABAN, and Gen-PINT serve as concrete anchors in the cells.

What would settle it

Re-run the review with a strict filter requiring every paper to come from ACL, EMNLP, or CHI 2020–2025 and check whether the four commonsense subthemes and four intent-detection subthemes still capture the dominant trends; if the corpus shrinks or the trends disappear, the central synthesis fails. A more direct test: on the benchmarks the review discusses, measure whether zero-shot, generative, and contrastive methods actually outperform the supervised baselines they are said to replace.

Watch

Extended reading notes

Core claim

The paper claims that, between 2020 and 2025, commonsense reasoning and intent detection have undergone a methodological shift: self-supervised and zero-shot methods (self-talk, DrFact, perturbation-refined Winograd models) reduce reliance on labeled data; multilingual resources such as X-CSQA and the Mickey Corpus extend coverage but still carry English-centric assumptions; structured reasoning benchmarks (ExplaGraphs, ATOMIC 2020) and studies of shortcut learning show that reported performance can reflect spurious patterns rather than genuine reasoning; and intent detection is being reformulated as open-set detection, generative label production (Gen-PINT), contrastive clustering, and human-centered applications, including self-harm query detection, voice interfaces for older adults, and gaze-based intent estimation. The synthesis concludes that the two fields are converging on a shared set of design challenges—grounding, generalization across languages and cultures, and reliable evaluation—rather than remaining separate classification and inference problems.

Load-bearing premise

The load-bearing premise is that the 28 papers the review claims to cover are all accurately summarized and actually come from ACL, EMNLP, or CHI 2020–2025; the review's own table includes AAAI, LREC-COLING, and ACM THRI entries and one COMET misattribution, so if the corpus is not reliable, the trend and gap conclusions do not follow.

Editorial extensions

If this is right

  • If the shift is real, future dialogue agents will increasingly couple commonsense knowledge with open-set intent detection, so they can flag utterances that do not fit any known intent and reason about user meaning in context.
  • Benchmark builders will need to treat cultural and linguistic diversity as a first-class evaluation axis, because translated datasets alone preserve English-centric logic.
  • Reported accuracy on commonsense benchmarks should be treated as provisional, since shortcut-learning results imply performance gains must be verified against artifact-free evaluation.
  • Generative intent-labeling methods offer flexibility in low-resource settings but introduce consistency and evaluation challenges that the field will have to standardize.
  • HCI cases such as mental-health query classification, older-adult voice assistance, and gaze-based intent will continue to pull intent detection toward design concerns like interpretability, fairness, and contextual awareness.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension implied by the review: benchmark designers could merge the two literatures by creating an intent-detection benchmark whose utterances require social or physical commonsense to disambiguate, then measure whether open-set and generative intent models improve when paired with a commonsense knowledge source.
  • The review's own corpus constraints (venue mismatches and a misattributed COMET reference) suggest that the trend claims would be more robust if re-run on a strictly filtered corpus; this is my inference, not the paper's.
  • One could quantify the grounding gap the paper identifies by evaluating COMET-style generated knowledge graphs against a contextual-anchoring metric, such as whether generated facts change when dialogue context changes, which the review describes qualitatively but does not measure.
  • The human-centered observations imply that intent detection may eventually be evaluated less by classification accuracy and more by downstream outcomes such as successful intervention or task completion; this is an editorial projection.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This manuscript presents itself as an interdisciplinary literature review of commonsense reasoning and intent detection, claiming to analyze 28 papers from ACL, EMNLP, and CHI (2020–2025) and organize them by methodology and application. The review is structured into two main themes with sub-themes, covering zero-shot and self-supervised commonsense reasoning, multilingual and cultural adaptation, structured reasoning evaluation, interactive commonsense, and intent detection approaches ranging from open-set and generative models to contrastive clustering and human-centered HCI applications. A discussion section derives aggregate trends and research gaps from the reviewed corpus, and an appendix lists search keywords. The paper's core contributions are the synthesis itself and the identified gaps in grounding, generalization, and benchmark design.

Significance. If the corpus were accurately defined and faithfully summarized, an updated interdisciplinary review spanning NLP and HCI would be a useful resource, and the explicit attempt to connect commonsense reasoning with intent detection across these communities is a genuine strength. The paper names concrete organizing dimensions, including zero-shot learning, cultural adaptation, structured evaluation, interactive contexts, open-set and generative intent detection, clustering, and human-centered applications, and it makes a plausible case for a methodological shift toward adaptive, context-aware models. However, the significance is conditional: the paper's conclusions are aggregate claims over a corpus that is not described reproducibly and that contains clear attribution and summary errors, so the value of the synthesis cannot be assessed as written.

major comments (4)
  1. [Section 3 and Table 1] The stated inclusion criteria in Section 3 ('peer-reviewed papers from top conferences between 2020 and 2025 (ACL, EMNLP, CHI)' and exclusion of preprints) are contradicted by multiple entries in Table 1. For example, Hwang et al. (2020) is an AAAI paper, Murata and Kawahara (2024) is LREC-COLING, Belardinelli (2024) is ACM THRI, Lin et al. (2021b) and Kumar et al. (2022) are NAACL, and Sencan (2024) has no listed venue. If 'ACL' is intended to include all ACL-affiliated venues, that convention needs to be stated explicitly; as written, the corpus definition and the actual table do not match, and the aggregate trends in Section 5 are computed over an ill-defined set.
  2. [Section 4.1.1] The text states that 'Another work introduces COMET, a model that uses transformers to generate commonsense knowledge graphs, building upon existing resources like ConceptNet (Murata and Kawahara, 2024).' COMET was introduced by Bosselut et al. (2019), which is in the reference list; Murata and Kawahara (2024) is Time-aware COMET, an extension. Because Section 5 later uses COMET as a key example of the grounding gap, this misattribution changes the evidentiary basis of the review's main discussion.
  3. [Section 4.1.2] The text names the multilingual dataset 'X-CSQA' and cites Sakai et al. (2024), but the reference list entry is titled 'mCSQA: Multilingual commonsense reasoning dataset with unified creation strategy by language models and humans.' Either the dataset name is wrong or the cited paper is the wrong one; in both cases the summary does not match the source.
  4. [Section 3 and Appendix A] The methodology does not provide a reproducible search protocol. It lists keywords but gives no databases, no search dates, no full Boolean query strings, and no screening or eligibility criteria beyond the venue restriction. Appendix A also uses inconsistent separators between keywords. Therefore the 28-paper corpus cannot be independently reconstructed, which is a load-bearing limitation for a review whose conclusions are aggregate over that corpus.
minor comments (5)
  1. [Table 1 and Section 4.1.3] The citation 'Hwangy et al.' is a typo for Hwang et al., and the reference entry 'D Jena Hwangy and 1 others. Atomic2020... AAAI2020' is malformed; the author list and venue information should be completed correctly.
  2. [References] The Bosselut et al. (2019) reference is listed as an arXiv preprint even though COMET appeared at ACL 2019; the published venue citation should be used.
  3. [Section 2] There is a stray period in the sentence that reads 'provide a comprehensive overview of natural language reasoning in NLP. and the integration of commonsense knowledge into NLP tasks.'
  4. [Section 4] The sentence 'will discuss about each subthemes' is grammatically awkward; the entire manuscript would benefit from a light copyedit.
  5. [Appendix A] The keyword list uses inconsistent separators, mixing spaces, commas, and periods; a single consistent separator should be used.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; this is a narrative literature review with no fitted inputs, self-citations, or derivation that reduces to its own assumptions.

full rationale

This paper is a narrative literature review rather than a derivation or empirical study. It reports no equations, fitted parameters, or predictive claims that could reduce to its inputs by construction. The synthesis is explicitly grounded in 28 external papers, and the author does not cite any of their own prior work, so the self-citation patterns that drive circularity findings are absent. The themes and gap analysis in Sections 4 and 5 are descriptive summaries of the reviewed corpus, which is the normal and non-circular operation of a survey. The manuscript's corpus-provenance and attribution problems, such as Table 1 including entries outside the stated ACL/EMNLP/CHI 2020-2025 scope and Section 4.1.1 attributing COMET to Murata and Kawahara (2024) rather than Bosselut et al. (2019), are factual-accuracy and reproducibility concerns about the review's inputs, not circularity: the review's claims are not equivalent to its own assumptions by definition. Similarly, Appendix A's keyword-only search documentation weakens reproducibility but does not make any conclusion self-justifying. Accordingly, the appropriate circularity score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

This paper is a qualitative literature review, so it introduces no free parameters or invented entities. Its central arguments rest on domain assumptions about the representativeness of the selected papers, the faithfulness of each summary, and the usefulness of the sub-theme taxonomy. These assumptions are challenged by internal inconsistencies: the paper claims 28 papers but presents 27, includes papers from venues outside its stated scope, and contains at least one clear misattribution of a foundational model.

assumptions (3)
  • domain assumption The selected papers are representative of recent advances in commonsense reasoning and intent detection.
    Section 3 asserts a keyword search and screening, but the criteria are not operationalized and the table lists 27 papers rather than 28, so representativeness cannot be verified.
  • domain assumption Each reviewed paper's contribution is faithfully summarized.
    Section 4.1.1 misattributes COMET to Murata and Kawahara (2024), so the summaries are not all reliable.
  • domain assumption The sub-theme categorization is a meaningful organization of the field.
    The taxonomy in Table 1 is presented without justification or comparison to alternative taxonomies in prior surveys.

how reviews work

0 comments
Cite this review

Pith. "Pith review of An Interdisciplinary Review of Commonsense Reasoning and Intent Detection." pith.science (2026). https://pith.science/paper/S6BDUW6Y

@misc{pith2026250614040,
  author       = {Pith},
  title        = {Pith review of: An Interdisciplinary Review of Commonsense Reasoning and Intent Detection},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/S6BDUW6Y}},
  note         = {Machine review of arXiv:2506.14040}
}
read the original abstract

This review explores recent advances in commonsense reasoning and intent detection, two key challenges in natural language understanding. We analyze 28 papers from ACL, EMNLP, and CHI (2020-2025), organizing them by methodology and application. Commonsense reasoning is reviewed across zero-shot learning, cultural adaptation, structured evaluation, and interactive contexts. Intent detection is examined through open-set models, generative formulations, clustering, and human-centered systems. By bridging insights from NLP and HCI, we highlight emerging trends toward more adaptive, multilingual, and context-aware models, and identify key gaps in grounding, generalization, and benchmark design.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

48 extracted references · 26 canonical work pages

  1. [1]

    Shourya Aggarwal, Divyanshu Mandowara, Vishwajeet Agrawal, Dinesh Khandelwal, Parag Singla, and Dinesh Garg. 2021. Explanations for commonsenseqa: New dataset and models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Pape...

  2. [2]

    Anna Belardinelli. 2024. Gaze-based intention estimation: principles, methodologies, and applications in hri. ACM Transactions on Human-Robot Interaction, 13(3):1--30

  3. [3]

    Antoine Bosselut, Hannah Rashkin, Maarten Sap, Chaitanya Malaviya, Asli Celikyilmaz, and Yejin Choi. 2019. Comet: Commonsense transformers for automatic knowledge graph construction. arXiv preprint arXiv:1906.05317

  4. [4]

    Ruben Branco, Ant \'o nio Branco, Jo \ a o Ant \'o nio Rodrigues, and Jo \ a o Ricardo Silva. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.113 Shortcutted commonsense: Data spuriousness in deep learning of commonsense reasoning . In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 1504--1521, Online and Pu...

  5. [5]

    Jiangjie Chen, Wei Shi, Ziquan Fu, Sijie Cheng, Lei Li, and Yanghua Xiao. 2023. https://doi.org/10.18653/v1/2023.acl-long.550 Say what you mean! large language models speak too positively about negative commonsense knowledge . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 9890--99...

  6. [6]

    Haoang Chi, He Li, Wenjing Yang, Feng Liu, Long Lan, Xiaoguang Ren, Tongliang Liu, and Bo Han. 2024. Unveiling causal reasoning in large language models: Reality or mirage? Advances in Neural Information Processing Systems, 37:96640--96670

  7. [7]

    DongHyun Choi, Myeong Cheol Shin, EungGyun Kim, and Dong Ryeol Shin. 2021. https://doi.org/10.18653/v1/2021.findings-acl.45 O ut F lip: Generating examples for unknown intent detection with natural language attack . In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 504--512, Online. Association for Computational Linguistics

  8. [8]

    Ernest Davis and Gary Marcus. 2015. Commonsense reasoning and commonsense knowledge in artificial intelligence. Communications of the ACM, 58(9):92--103

Show all 48 references
  1. [9]

    Lam, and Xiao-Ming Wu

    Lu Fan, Guangfeng Yan, Qimai Li, Han Liu, Xiaotong Zhang, Albert Y.S. Lam, and Xiao-Ming Wu. 2020. https://doi.org/10.18653/v1/2020.acl-main.99 Unknown intent detection using G aussian mixture model with an application to zero-shot intent classification . In Proceedings of the...

  2. [10]

    Deepanway Ghosal, Siqi Shen, Navonil Majumder, Rada Mihalcea, and Soujanya Poria. 2022. https://doi.org/10.18653/v1/2022.acl-long.344 CICERO : A dataset for contextualized commonsense inference in dialogues . In Proceedings of the 60th Annual Meeting of the Association for Com...

  3. [11]

    Jie He, Tao Wang, Deyi Xiong, and Qun Liu. 2020. https://doi.org/10.18653/v1/2020.findings-emnlp.327 The box is in the pen: Evaluating commonsense reasoning in neural machine translation . In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3662--36...

  4. [12]

    Atomic2020: On symbolic and neural commonsense knowledge graphs

    D Jena Hwangy and 1 others. Atomic2020: On symbolic and neural commonsense knowledge graphs. AAAI2020

  5. [13]

    Mourad Jbene, Abdellah Chehri, Rachid Saadane, Smail Tigani, and Gwanggil Jeon. 2025. Intent detection for task-oriented conversational agents: A comparative study of recurrent neural networks and transformer models. Expert Systems, 42(2):e13712

  6. [14]

    Tassilo Klein and Moin Nabi. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.688 Towards zero-shot commonsense reasoning with self-supervised refinement of language models . In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 8737...

  7. [15]

    Rajat Kumar, Mayur Patidar, Vaibhav Varshney, Lovekesh Vig, and Gautam Shroff. 2022. https://doi.org/10.18653/v1/2022.naacl-main.134 Intent detection and discovery from user logs via deep semi-supervised contrastive clustering . In Proceedings of the 2022 Conference of the Nor...

  8. [16]

    Bill Yuchen Lin, Xinyue Chen, Jamin Chen, and Xiang Ren. 2019. Kagnet: Knowledge-aware graph networks for commonsense reasoning. arXiv preprint arXiv:1909.02151

  9. [17]

    Bill Yuchen Lin, Seyeon Lee, Xiaoyang Qiao, and Xiang Ren. 2021 a . https://doi.org/10.18653/v1/2021.acl-long.102 Common sense beyond E nglish: Evaluating and improving multilingual language models for commonsense reasoning . In Proceedings of the 59th Annual Meeting of the As...

  10. [18]

    Bill Yuchen Lin, Haitian Sun, Bhuwan Dhingra, Manzil Zaheer, Xiang Ren, and William Cohen. 2021 b . https://doi.org/10.18653/v1/2021.naacl-main.366 Differentiable open-ended commonsense reasoning . In Proceedings of the 2021 Conference of the North American Chapter of the Asso...

  11. [19]

    Jiacheng Liu, Alisa Liu, Ximing Lu, Sean Welleck, Peter West, Ronan Le Bras, Yejin Choi, and Hannaneh Hajishirzi. 2021. Generated knowledge prompting for commonsense reasoning. arXiv preprint arXiv:2110.08387

  12. [20]

    Jiao Liu, Yanling Li, and Min Lin. 2019. Review of intent detection methods in the human-machine dialogue system. In Journal of physics: conference series, volume 1267, page 012059. IOP Publishing

  13. [21]

    Dylan P Losey, Craig G McDonald, Edoardo Battaglia, and Marcia K O'Malley. 2018. A review of intent detection, arbitration, and communication aspects of shared control for physical human--robot interaction. Applied Mechanics Reviews, 70(1):010804

  14. [22]

    i'd rather just go to bed

    Annie Louis, Dan Roth, and Filip Radlinski. 2020. " i'd rather just go to bed": Understanding indirect answers. arXiv preprint arXiv:2010.03450

  15. [23]

    Martino Mensio, Giuseppe Rizzo, and Maurizio Morisio. 2018. Multi-turn qa: A rnn contextual approach to intent classification for goal-oriented systems. In Companion Proceedings of the The Web Conference 2018, pages 1075--1080

  16. [24]

    Eiki Murata and Daisuke Kawahara. 2024. Time-aware comet: a commonsense knowledge model with temporal knowledge. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pages 16162--16174

  17. [25]

    Yawen Ouyang, Jiasheng Ye, Yu Chen, Xinyu Dai, Shujian Huang, and Jiajun Chen. 2021. https://doi.org/10.18653/v1/2021.findings-acl.252 Energy-based unknown intent detection with data manipulation . In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, ...

  18. [26]

    Silvia Pareti, Tim O’keefe, Ioannis Konstas, James R Curran, and Irena Koprinska. 2013. Automatically detecting and attributing indirect quotations. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, pages 989--999

  19. [27]

    Libo Qin, Xiao Xu, Wanxiang Che, and Ting Liu. 2020. https://doi.org/10.18653/v1/2020.findings-emnlp.163 AGIF : An adaptive graph-interactive framework for joint multiple intent detection and slot filling . In Findings of the Association for Computational Linguistics: EMNLP 20...

  20. [28]

    Yincen Qu, Ningyu Zhang, Hui Chen, Zelin Dai, Chengming Wang, Xiaoyu Wang, Qiang Chen, and Huajun Chen. 2022. https://doi.org/10.18653/v1/2022.findings-emnlp.2 Commonsense knowledge salience evaluation with a benchmark dataset in E -commerce . In Findings of the Association fo...

  21. [29]

    Hannah Rashkin, Eric Michael Smith, Margaret Li, and Y-Lan Boureau. 2018. Towards empathetic open-domain conversation models: A new benchmark and dataset. arXiv preprint arXiv:1811.00207

  22. [30]

    May Lynn Reese and Anastasia Smirnova. 2024. Comparing chatgpt and humans on world knowledge and common-sense reasoning tasks: A case study of the japanese winograd schema challenge. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, pages 1--9

  23. [31]

    Christopher Richardson and Larry Heck. 2023. Commonsense reasoning for conversational ai: A survey of the state of the art. arXiv preprint arXiv:2302.07926

  24. [32]

    Julien Romero and Simon Razniewski. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.752 Do children texts hold the key to commonsense knowledge? In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 10954--10959, Abu Dhabi, United A...

  25. [33]

    Swarnadeep Saha, Prateek Yadav, Lisa Bauer, and Mohit Bansal. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.609 E xpla G raphs: An explanation graph generation task for structured commonsense reasoning . In Proceedings of the 2021 Conference on Empirical Methods in Natural...

  26. [34]

    Yusuke Sakai, Hidetaka Kamigaito, and Taro Watanabe. 2024. https://doi.org/10.18653/v1/2024.findings-acl.844 m CSQA : Multilingual commonsense reasoning dataset with unified creation strategy by language models and humans . In Findings of the Association for Computational Ling...

  27. [35]

    Maarten Sap, Hannah Rashkin, Derek Chen, Ronan LeBras, and Yejin Choi. 2019. Socialiqa: Commonsense reasoning about social interactions. arXiv preprint arXiv:1904.09728

  28. [36]

    Maarten Sap, Vered Shwartz, Antoine Bosselut, Yejin Choi, and Dan Roth. 2020. Commonsense reasoning for natural language processing. In Proceedings of the 58th annual meeting of the association for computational linguistics: Tutorial abstracts, pages 27--33

  29. [37]

    Cevdet S encan. 2024. Intention mining: surfacing and reshaping deep intentions by proactive human computer interaction

  30. [38]

    Vered Shwartz, Peter West, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.373 Unsupervised commonsense question answering with self-talk . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Proc...

  31. [39]

    Oyvind Tafjord, Matt Gardner, Kevin Lin, and Peter Clark. 2019. Quartz: An open-domain dataset of qualitative relationship questions. arXiv preprint arXiv:1909.03553

  32. [40]

    Henry Weld, Xiaoqi Huang, Siqu Long, Josiah Poon, and Soyeon Caren Han. 2022. A survey of joint intent detection and slot filling models in natural language understanding. ACM Computing Surveys, 55(8):1--38

  33. [41]

    Sixing Wu, Ying Li, Dawei Zhang, Yang Zhou, and Zhonghai Wu. 2020. Diverse and informative dialogue generation with context-specific commonsense knowledge awareness. In Proceedings of the 58th annual meeting of the association for computational linguistics, pages 5811--5820

  34. [42]

    Ting-Wei Wu, Ruolin Su, and Biing Juang. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.399 A label-aware BERT attention network for zero-shot multi-intent detection in spoken language understanding . In Proceedings of the 2021 Conference on Empirical Methods in Natural Lan...

  35. [43]

    Erqian Xu, Hecong Wang, and Zhen Bai. 2023. Engage ai and child in explanatory dialogue on commonsense reasoning. In Extended Abstracts of the 2023 CHI Conference on Human Factors in Computing Systems, pages 1--8

  36. [44]

    Da Yin, Liunian Harold Li, Ziniu Hu, Nanyun Peng, and Kai-Wei Chang. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.162 Broaden the vision: Geo-diverse visual commonsense reasoning . In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, ...

  37. [45]

    Fei Yu, Hongbo Zhang, Prayag Tiwari, and Benyou Wang. 2024. Natural language reasoning, a survey. ACM Computing Surveys, 56(12):1--39

  38. [46]

    where is history

    Ja Eun Yu, Natalie Parde, and Debaleena Chattopadhyay. 2023. “where is history”: Toward designing a voice assistant to help older adults locate interface features quickly. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pages 1--19

  39. [47]

    Feng Zhang, Wei Chen, Fei Ding, Meng Gao, Tengjiao Wang, Jiahui Yao, and Jiabin Zheng. 2024. https://doi.org/10.18653/v1/2024.findings-acl.605 From discrimination to generation: Low-resource intent detection with language model instruction tuning . In Findings of the Associati...

  40. [48]

    Yuwei Zhang, Haode Zhang, Li-Ming Zhan, Xiao-Ming Wu, and Albert Lam. 2022. https://doi.org/10.18653/v1/2022.acl-long.21 New intent discovery with pre-training and contrastive learning . In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.