{"id":"927d0b76-7671-46d3-9913-64c6ce3437d8","arxiv_id":"2506.14040","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey that organizes 28 (actually 27) papers on commonsense reasoning and intent detection into eight themes, but with factual misattributions and inconsistent venue selection.","lead":"This paper reviews 28 papers on commonsense reasoning and intent detection, grouping them by method and application. It aims to bridge natural language processing and human-computer interaction, though its paper count and venue selection do not match its claims.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Corpus provenance and summary fidelity are the load-bearing foundation of this review; Table 1 and §4.1.1 violate the stated ACL/EMNLP/CHI 2020–2025 corpus, so the synthesis is not trustworthy as written.","rationale":"The reader's weakest_assumption identifies the same load-bearing issue: the review's value depends on an accurate, consistently sourced corpus and faithful summaries. The manuscript provides direct counterexamples: §4.1.1 misattributes COMET to Murata & Kawahara (2024) despite Bosselut et al. (2019) being in the references, and Table 1 includes entries from AAAI, LREC-COLING, ACM THRI, and NAACL that violate §3's stated ACL/EMNLP/CHI scope. The Discussion builds on COMET as a grounding limitation, so the error affects the paper's substantive conclusions, not just the reference formatting. I did not find a separate technical flaw in the review's qualitative framing; the load-bearing problem is the integrity of the corpus and its summaries. The concrete audit above would settle whether this is a handful of citation slips or a systemic provenance failure; either way, the paper as written cannot support its central claim of synthesizing 28 ACL/EMNLP/CHI 2020–2025 papers. One detail in the reader's rationale, the claim that Table 1 lists 27 papers, does not match my count of 28 entries; the venue and COMET problems stand independently. This supports the reader's REJECT verdict; a corrected version with a reproducible search protocol and fixed citations might be worth reconsidering.","tokens_in":8459,"tokens_out":10408,"duration_ms":88180,"concrete_test":"Run a corpus audit: extract all 28 citations from Table 1, align each with the reference list, and query DBLP/ACL Anthology/ACM DL for official venue and year; flag every entry outside ACL/EMNLP/CHI 2020–2025. Separately, read Murata & Kawahara (2024) and Bosselut et al. (2019) to determine which paper introduces COMET. If any Table 1 venue falls outside the stated set, or if COMET is not introduced by Murata & Kawahara (2024), the review's central corpus claim fails and the affected trend statements must be revised.","verdict_should_be":"REJECT","load_bearing_attack":"The central claim is an accurate, reproducible synthesis of 28 papers from ACL, EMNLP, and CHI (2020–2025). That claim requires each reviewed paper to be correctly attributed and correctly summarized. The manuscript contradicts this on both counts. §3 restricts inclusion to peer-reviewed papers from ACL, EMNLP, or CHI 2020–2025 and excludes preprints, yet Table 1 includes AAAI (Hwang et al., ATOMIC2020), LREC-COLING (Murata & Kawahara, 2024), ACM THRI (Belardinelli, 2024), and NAACL (Lin et al., 2021b; Kumar et al., 2022) entries. §4.1.1 states that 'Another work introduces COMET ... building upon existing resources like ConceptNet (Murata and Kawahara, 2024)'; COMET was introduced by Bosselut et al. (2019), which appears in the reference list as an arXiv preprint, while Murata & Kawahara (2024) is Time-aware COMET, an extension. The Discussion later uses COMET as evidence of a grounding gap, so the misattribution is not cosmetic. Because the review's trends and gap analysis are aggregate conclusions over this corpus, these provenance and summary errors make the synthesis unreliable as stated. Appendix A lists only keywords; there is no search log, database list, search dates, or screening protocol, so the 28-paper corpus cannot be independently reconstructed.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript presents itself as an interdisciplinary literature review of commonsense reasoning and intent detection, claiming to analyze 28 papers from ACL, EMNLP, and CHI (2020–2025) and organize them by methodology and application. The review is structured into two main themes with sub-themes, covering zero-shot and self-supervised commonsense reasoning, multilingual and cultural adaptation, structured reasoning evaluation, interactive commonsense, and intent detection approaches ranging from open-set and generative models to contrastive clustering and human-centered HCI applications. A discussion section derives aggregate trends and research gaps from the reviewed corpus, and an appendix lists search keywords. The paper's core contributions are the synthesis itself and the identified gaps in grounding, generalization, and benchmark design.","tokens_in":8664,"tokens_out":8435,"duration_ms":74152,"significance":"If the corpus were accurately defined and faithfully summarized, an updated interdisciplinary review spanning NLP and HCI would be a useful resource, and the explicit attempt to connect commonsense reasoning with intent detection across these communities is a genuine strength. The paper names concrete organizing dimensions, including zero-shot learning, cultural adaptation, structured evaluation, interactive contexts, open-set and generative intent detection, clustering, and human-centered applications, and it makes a plausible case for a methodological shift toward adaptive, context-aware models. However, the significance is conditional: the paper's conclusions are aggregate claims over a corpus that is not described reproducibly and that contains clear attribution and summary errors, so the value of the synthesis cannot be assessed as written.","major_comments":[{"comment":"The stated inclusion criteria in Section 3 ('peer-reviewed papers from top conferences between 2020 and 2025 (ACL, EMNLP, CHI)' and exclusion of preprints) are contradicted by multiple entries in Table 1. For example, Hwang et al. (2020) is an AAAI paper, Murata and Kawahara (2024) is LREC-COLING, Belardinelli (2024) is ACM THRI, Lin et al. (2021b) and Kumar et al. (2022) are NAACL, and Sencan (2024) has no listed venue. If 'ACL' is intended to include all ACL-affiliated venues, that convention needs to be stated explicitly; as written, the corpus definition and the actual table do not match, and the aggregate trends in Section 5 are computed over an ill-defined set.","section":"Section 3 and Table 1"},{"comment":"The text states that 'Another work introduces COMET, a model that uses transformers to generate commonsense knowledge graphs, building upon existing resources like ConceptNet (Murata and Kawahara, 2024).' COMET was introduced by Bosselut et al. (2019), which is in the reference list; Murata and Kawahara (2024) is Time-aware COMET, an extension. Because Section 5 later uses COMET as a key example of the grounding gap, this misattribution changes the evidentiary basis of the review's main discussion.","section":"Section 4.1.1"},{"comment":"The text names the multilingual dataset 'X-CSQA' and cites Sakai et al. (2024), but the reference list entry is titled 'mCSQA: Multilingual commonsense reasoning dataset with unified creation strategy by language models and humans.' Either the dataset name is wrong or the cited paper is the wrong one; in both cases the summary does not match the source.","section":"Section 4.1.2"},{"comment":"The methodology does not provide a reproducible search protocol. It lists keywords but gives no databases, no search dates, no full Boolean query strings, and no screening or eligibility criteria beyond the venue restriction. Appendix A also uses inconsistent separators between keywords. Therefore the 28-paper corpus cannot be independently reconstructed, which is a load-bearing limitation for a review whose conclusions are aggregate over that corpus.","section":"Section 3 and Appendix A"}],"minor_comments":[{"comment":"The citation 'Hwangy et al.' is a typo for Hwang et al., and the reference entry 'D Jena Hwangy and 1 others. Atomic2020... AAAI2020' is malformed; the author list and venue information should be completed correctly.","section":"Table 1 and Section 4.1.3"},{"comment":"The Bosselut et al. (2019) reference is listed as an arXiv preprint even though COMET appeared at ACL 2019; the published venue citation should be used.","section":"References"},{"comment":"There is a stray period in the sentence that reads 'provide a comprehensive overview of natural language reasoning in NLP. and the integration of commonsense knowledge into NLP tasks.'","section":"Section 2"},{"comment":"The sentence 'will discuss about each subthemes' is grammatically awkward; the entire manuscript would benefit from a light copyedit.","section":"Section 4"},{"comment":"The keyword list uses inconsistent separators, mixing spaces, commas, and periods; a single consistent separator should be used.","section":"Appendix A"}],"recommendation":"major_revision","confidential_remarks":"The identified errors may be symptoms of a broader reliability problem in the summaries. I would ask the author to re-verify every entry in Table 1 against the full text of each cited paper, to correct the COMET attribution and the dataset name, and to add a full search and screening protocol. If the revision does not address these issues systematically, I would not support publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: this is a useful survey to have on your radar, but it's not trustworthy as written. The good news first. The eight-subtheme organization is sensible, and the decision to pull in CHI papers alongside ACL and EMNLP work is a real expansion over prior surveys like Richardson and Heck (2023) and Weld et al. (2022). The summaries in Section 4 are mostly readable and accurate, particularly on the HCI side; seeing self-harm detection, voice assistance for older adults, and gaze-based intent estimation next to standard NLP intent detection is genuinely informative. The gap discussion—English-centric bias, benchmark artifacts, grounding—hits the right notes.\n\nThe soft spots are the kind that sink a review. Section 3 says the corpus is peer-reviewed ACL, EMNLP, and CHI papers from 2020–2025, excluding preprints. Table 1 contains at least four entries from other venues: AAAI (Hwang et al., ATOMIC2020), LREC-COLING (Murata & Kawahara 2024), ACM THRI (Belardinelli 2024), and NAACL (Lin et al. 2021b; Kumar et al. 2022). The abstract says 28 papers; the table has 27. And Section 4.1.1 says 'Another work introduces COMET ... building upon existing resources like ConceptNet (Murata and Kawahara, 2024).' That's backwards: COMET is Bosselut et al. (2019), which the reference list includes as an arXiv preprint, and Murata & Kawahara is Time-Aware COMET, an extension. Because the Discussion uses COMET to illustrate a grounding gap, the misattribution is load-bearing rather than cosmetic. There's also no search log or screening protocol—Appendix A is just a keyword list—so the corpus can't be reconstructed independently.\n\nThese are fixable problems, but they're the core of what makes a review reliable. The paper doesn't need new experiments; it needs a corrected table, a stated venue policy (or a justified expansion), and a fact-check pass on attributions. The author's organizational thinking is sound, and the interdisciplinary angle is worth preserving.\n\nWho's this for? Someone new to the area who wants a quick map, or an NLP researcher checking what HCI has to offer on these tasks. It deserves a serious referee with a request for major revision rather than a desk reject, because the flaws are specific and correctable and the useful core is visible.","headline":"Readable survey with a genuinely useful HCI tilt, but the stated corpus doesn't match Table 1, the COMET attribution is wrong, and the abstract/table count doesn't add up.","tokens_in":9194,"tokens_out":3681,"would_cite":false,"duration_ms":33452,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This review of 28 papers from ACL, EMNLP, and CHI (2020–2025) argues that commonsense reasoning and intent detection are converging on adaptive, zero-shot, and human-centered methods, with unresolved gaps in grounding, generalization, and…","keywords":["commonsense reasoning","intent detection","natural language understanding","zero-shot learning","cultural adaptation","multilingual evaluation","benchmark design","human-computer interaction"],"falsifier":"Re-run the review with a strict filter requiring every paper to come from ACL, EMNLP, or CHI 2020–2025 and check whether the four commonsense subthemes and four intent-detection subthemes still capture the dominant trends; if the corpus shrinks or the trends disappear, the central synthesis fails. A more direct test: on the benchmarks the review discusses, measure whether zero-shot, generative, and contrastive methods actually outperform the supervised baselines they are said to replace.","tokens_in":8195,"feed_emoji":"🧠","tokens_out":9773,"duration_ms":87496,"temperature":0.7,"pith_summary":"This review synthesizes 28 papers from NLP and HCI venues published between 2020 and 2025 to show how commonsense reasoning and intent detection are changing. Its central claim is that both fields are moving away from fully supervised, fixed-label systems toward zero-shot, generative, contrastive, and human-centered approaches that aim to handle unseen intents and culturally varied situations. The paper organizes the literature into two themes—commonsense reasoning and intent detection—each with four methodological subthemes, and uses this structure to highlight a shared set of open problems: knowledge that is not grounded in specific contexts, English-centric biases in multilingual resources, and benchmarks that can be gamed by shallow heuristics. A sympathetic reader would take the main contribution to be the cross-venue map of trends and gaps, offered so that future systems can balance robustness, cultural sensitivity, and task specificity.","feed_headline":"28 papers show commonsense and intent detection converging","feed_subtitle":"A 28-paper NLP and HCI review tracks the shift to zero-shot, culturally aware, human-centered models.","key_machinery":"The machinery is the review's two-theme, four-subtheme taxonomy. Commonsense reasoning is split into (i) self-supervised and zero-shot learning, (ii) multilingual and cultural adaptation, (iii) structured reasoning and evaluation analysis, and (iv) interactive, dialog-based, and applied commonsense; intent detection is split into (i) open-set and zero-shot detection, (ii) multi-intent modeling and generative formulation, (iii) contrastive learning and clustering, and (iv) human-centered and HCI applications. Each reviewed paper is assigned to a cell based on methodology (graph-based, generative, prompting, or hybrid) and reasoning type (causal, dialogic, social). The taxonomy does the work of turning 28 individual findings into trend claims—a shift away from supervised learning—and gap claims about grounding, generalization, and benchmark reliability. Named systems such as COMET, DrFact, CICERO, ExplaGraphs, AGIF, LABAN, and Gen-PINT serve as concrete anchors in the cells.","core_discovery":"The paper claims that, between 2020 and 2025, commonsense reasoning and intent detection have undergone a methodological shift: self-supervised and zero-shot methods (self-talk, DrFact, perturbation-refined Winograd models) reduce reliance on labeled data; multilingual resources such as X-CSQA and the Mickey Corpus extend coverage but still carry English-centric assumptions; structured reasoning benchmarks (ExplaGraphs, ATOMIC 2020) and studies of shortcut learning show that reported performance can reflect spurious patterns rather than genuine reasoning; and intent detection is being reformulated as open-set detection, generative label production (Gen-PINT), contrastive clustering, and human-centered applications, including self-harm query detection, voice interfaces for older adults, and gaze-based intent estimation. The synthesis concludes that the two fields are converging on a shared set of design challenges—grounding, generalization across languages and cultures, and reliable evaluation—rather than remaining separate classification and inference problems.","pith_inferences":["A testable extension implied by the review: benchmark designers could merge the two literatures by creating an intent-detection benchmark whose utterances require social or physical commonsense to disambiguate, then measure whether open-set and generative intent models improve when paired with a commonsense knowledge source.","The review's own corpus constraints (venue mismatches and a misattributed COMET reference) suggest that the trend claims would be more robust if re-run on a strictly filtered corpus; this is my inference, not the paper's.","One could quantify the grounding gap the paper identifies by evaluating COMET-style generated knowledge graphs against a contextual-anchoring metric, such as whether generated facts change when dialogue context changes, which the review describes qualitatively but does not measure.","The human-centered observations imply that intent detection may eventually be evaluated less by classification accuracy and more by downstream outcomes such as successful intervention or task completion; this is an editorial projection."],"forward_implications":["If the shift is real, future dialogue agents will increasingly couple commonsense knowledge with open-set intent detection, so they can flag utterances that do not fit any known intent and reason about user meaning in context.","Benchmark builders will need to treat cultural and linguistic diversity as a first-class evaluation axis, because translated datasets alone preserve English-centric logic.","Reported accuracy on commonsense benchmarks should be treated as provisional, since shortcut-learning results imply performance gains must be verified against artifact-free evaluation.","Generative intent-labeling methods offer flexibility in low-resource settings but introduce consistency and evaluation challenges that the field will have to standardize.","HCI cases such as mental-health query classification, older-adult voice assistance, and gaze-based intent will continue to pull intent detection toward design concerns like interpretability, fairness, and contextual awareness."],"supporting_citations":[{"why":"supplies the social-commonsense dataset that grounds the review's framing of intent as needing world knowledge.","marker":"Sap et al., 2019"},{"why":"provides the zero-shot self-talk method that anchors the claimed shift away from supervised training.","marker":"Shwartz et al., 2020"},{"why":"supports the claim that open-ended commonsense retrieval (DrFact) works but depends on high-quality corpora.","marker":"Lin et al., 2021b"},{"why":"anchors the benchmark-reliability critique by showing models exploit spurious patterns in commonsense reasoning.","marker":"Branco et al., 2021"},{"why":"supplies the X-CSQA multilingual benchmark used to ground the cultural-bias and translation-artifact discussion.","marker":"Sakai et al., 2024"},{"why":"supports the multi-intent modeling trend with the AGIF graph-interactive framework for joint intent detection and slot filling.","marker":"Qin et al., 2020"},{"why":"anchors the contrastive-clustering trend in intent discovery from partially labeled user logs.","marker":"Kumar et al., 2022"},{"why":"supports the cross-linguistic limitation claim connecting commonsense reasoning and HCI through the Japanese Winograd Schema Challenge.","marker":"Reese and Smirnova, 2024"}],"fun_headline_variants":["Commonsense and intent: a 28-paper convergence review","How a 28-paper review ties commonsense to intent detection","28-paper review: commonsense and intent detection converge","NLU review: 28 papers unite commonsense and intent detection"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 28 papers the review claims to cover are all accurately summarized and actually come from ACL, EMNLP, or CHI 2020–2025; the review's own table includes AAAI, LREC-COLING, and ACM THRI entries and one COMET misattribution, so if the corpus is not reliable, the trend and gap conclusions do not follow.","fun_headline_variants_meta":{"raw":{"variants":["Commonsense and intent: a 28-paper convergence review","How a 28-paper review ties commonsense to intent detection","28-paper review: commonsense and intent detection converge","NLU review: 28 papers unite commonsense and intent detection"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000336,"raw_usage":{"total_tokens":1801,"prompt_tokens":829,"completion_tokens":972,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":445,"completion_tokens_details":{"reasoning_tokens":899}},"tokens_in":445,"tokens_out":972,"duration_ms":8236,"temperature":1.0,"reasoning_tokens":899,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:55:40.192300+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the review with a strict filter requiring every paper to come from ACL, EMNLP, or CHI 2020–2025 and check whether the four commonsense subthemes and four intent-detection subthemes still capture the dominant trends; if the corpus shrinks or the trends disappear, the central synthesis fails. A more direct test: on the benchmarks the review discusses, measure whether zero-shot, generative, and contrastive methods actually outperform the supervised baselines they are said to replace.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"anchors the contrastive-clustering trend in intent discovery from partially labeled user logs."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supports the cross-linguistic limitation claim connecting commonsense reasoning and HCI through the Japanese Winograd Schema Challenge."}],"review_version":1}