{"id":"b18e52ea-1073-4c4a-8335-5301e63796d5","arxiv_id":"2507.22934","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of deep learning methods for intent recognition, tracing the field from unimodal text, audio, vision, and EEG approaches to multimodal fusion, alignment, knowledge-augmented, and multi-task models.","lead":"This paper surveys how intent recognition moved from single-modality text, audio, vision, and EEG systems to multimodal deep learning models that fuse such signals. It organizes datasets, methods, evaluation metrics, applications, and open challenges, and is meant as a reference map for researchers entering this area.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'first systematic review' claim is not backed by any disclosed search protocol or inclusion criteria, so the curated dataset/method sample and four-paradigm taxonomy may be unrepresentative and the synthesis is not yet checkable.","rationale":"Reader's verdict is fair; I agree with the weakest assumption. The central contribution is a synthesis, not a derivation, so the test of correctness is whether the synthesis is reproducible and representative. The paper is readable, and individual method/dataset write-ups are mostly faithful to cited sources, which earns credit. But no disclosed methodology means the selection could be cherry-picked, and the strong 'first systematic' claim cannot be verified from the manuscript alone. Duplicate references and a numeric inconsistency in the MultiWOZ entry are warning signs that the curation layer was not carefully checked; they do not affect any particular model result, but they do affect trust in the survey-level claims. A scoping search plus a recall check on the citation pool would settle whether the concern actually lands. Since the reader already conditioned the verdict on exactly this, no verdict change is needed.","tokens_in":32010,"tokens_out":9146,"duration_ms":97914,"concrete_test":"Conduct a pre-registered scoping search to test the two load-bearing premises: (1) Query DBLP, Scopus, and the ACL Anthology with variants of TITLE-ABS-KEY('intent recognition' OR 'intent detection') AND ('multimodal' OR 'survey' OR 'review') with no start-date restriction, and list all surveys or monographs published before 24 Jul 2025 that cover the unimodal-to-multimodal evolution; if any same-scope prior survey exists, the 'first' claim fails. (2) Build a citation-based pool from the top 50 most-cited papers citing MIntRec, MIntRec2.0, or IntentQA up to the submission date, and compute the recall of this survey's Table 4 within that pool. If recall is below 80%, or if the search surfaces a same-scope earlier survey, the 'first systematic review' framing is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 1's first contribution is the claim of 'the first systematic review that traces the development of intent recognition from early unimodal approaches to modern multimodal techniques.' For that claim to license the paper's comparative conclusions, the sample in Table 2 and Table 4 must be representative, and the four-paradigm taxonomy in Section 4 must partition MIR work in a way that other researchers can reproduce. The paper gives no PRISMA-style search strategy, no query strings, no inclusion/exclusion criteria, no screening procedure, and no comparison with prior surveys of similar scope (e.g., the 2023 Springer monograph cited as [143], or earlier intent-detection reviews cited as [81]). Without these, the 'systematic' qualifier is an assertion, not an established property, and every section-level generalization in Section 4.5 — that fusion methods 'effectively combine text, audio, and visual signals,' that knowledge-augmented methods are limited in scalability, etc. — rests on an unverified sample. Curation errors compound the problem: MultiWOZ is described as '10,438 samples (8,438 conversations)' when 10,438 is the dialogue count and 8,438 is the training split; references [76] and [77] are duplicate entries for the same MBCFNet paper, as are [134] and [135] for the same MGC paper; MDID is used in Section 3.2 but absent from Table 2; and methods discussed in the text (MIntOOD, KDSL, MMSAIR) are omitted from Table 4. These issues are individually minor, but together they show the curation is not reproducible, which is the load-bearing property behind the 'systematic review' claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey reviews deep learning methods for intent recognition, tracing the field from unimodal text/vision/audio/EEG approaches to multimodal intent recognition (MIR). It proposes a three-stage processing pipeline (feature extraction, multimodal representation learning, intent classification) and organizes MIR methods into four paradigms: fusion, alignment and disentanglement, knowledge-augmented, and multi-task coordination. The paper compiles benchmark datasets, representative methods, evaluation metrics, applications, and challenges, and it claims to be the first systematic review of this trajectory. The survey is curated rather than derived: there are no fitted parameters or formal derivations, and the central contributions are the taxonomy, the resource inventory, and the field-level synthesis.","tokens_in":32309,"tokens_out":6156,"duration_ms":66420,"significance":"If the inventory and taxonomy are reliable, the paper is a useful reference for researchers entering multimodal intent recognition, especially because it covers many 2024–2025 works and organizes them into an explicit taxonomy. The evaluation-metrics section, with equations for Macro F1, Weighted F1, F1-IS/F1-OOS, and EER, is a practical contribution. The application and challenge sections are broad and well organized, and the paper explicitly names open problems such as modal asynchrony, long-tail distributions, and cross-lingual generalization. The main caveat is that the 'systematic review' claim is not currently supported by a disclosed methodology, and the dataset/method inventory contains internal inconsistencies. These issues are fixable, but they are load-bearing for the paper's comparative and synthetic conclusions.","major_comments":[{"comment":"The first contribution bullet claims 'the first systematic review that traces the development of intent recognition from early unimodal approaches to modern multimodal techniques,' but the manuscript provides no search protocol, query strings, inclusion/exclusion criteria, screening procedure, or comparison with existing surveys of similar scope such as [81], [143], and [3]. Without this methodology, the dataset and method sample in Tables 2 and 4 cannot be checked for representativeness, and the field-level generalizations in §4.5 — e.g., that fusion methods 'effectively combine text, audio, and visual signals' or that knowledge-augmented methods are 'limited in scalability' — rest on an unverified curated sample. Please either add a transparent systematic methodology or revise the claim to 'structured survey' and temper the corresponding generalizations.","section":"§1 (first contribution bullet), §2.1, §4.5"},{"comment":"Several internal consistency errors affect the resource inventory. MultiWOZ is described as containing '10,438 samples (8,438 conversations)'; in the original corpus 10,438 is the dialogue count and 8,438 is the training split, so the sample/dialogue counts are inverted. MDID is used as a benchmark in §3.2 (Tang et al. [124]) but is absent from Table 2. Methods discussed in the text — MIntOOD [164], KDSL [10], and MMSAIR [116] — are missing from Table 4. Please correct the MultiWOZ statistics and align the tables with the narrative so that the 'standardized foundation' claim is credible.","section":"§2.1, Table 2, §3.2, Table 4"},{"comment":"The four methodological paradigms are presented as a partition of MIR research, but membership rules are not given, and several methods fall into multiple categories: CaVIR appears under fusion, alignment, and knowledge-augmented methods; MGC appears under fusion and alignment; A-MESS appears under alignment and knowledge-augmented methods. It is not clear whether these are mutually exclusive classes or complementary aspects of a design space. Please state the classification criterion explicitly, and either assign each method to one primary paradigm or state that the categories are overlapping facets.","section":"§4.1–§4.4"}],"minor_comments":[{"comment":"References [76] and [77] are duplicate entries for the same MBCFNet paper, and references [134] and [135] are duplicate entries for the same MGC paper; please merge the duplicates and update the in-text citations accordingly.","section":"References"},{"comment":"The metrics list introduces Weighted Precision (WP) and Pearson Correlation (Corr), but neither is defined or used later; please provide formulas or remove them from the list.","section":"§5"},{"comment":"The Dataset column uses the abbreviation 'Comm.' in several rows without explanation; please define this abbreviation or replace it with the actual dataset names.","section":"Table 3"},{"comment":"In the discussion of MuProCL, the text says 'SLURP and MintRec'; the dataset name should be 'MIntRec'.","section":"§3.3"},{"comment":"Item (1) covers five text datasets while items (2)–(10) each cover a single dataset; renumbering or restructuring would make the dataset list easier to navigate.","section":"§2.1"}],"recommendation":"major_revision","confidential_remarks":"This is a survey with no derivations or fitted parameters, so circularity is not an issue. The main risk is the unsupported 'first systematic review' claim and the internal curation errors in the dataset/table inventory. In my view these are fixable within the manuscript's scope: adding a search protocol (or softening the claim), correcting the MultiWOZ count, reconciling Tables 2 and 4 with the text, and clarifying the taxonomy's membership rules would make the paper suitable for publication. I would not reject on novelty grounds, but the 'systematic' qualifier is likely to draw scrutiny from editors and readers."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth knowing: this is a genuinely useful map of multimodal intent recognition, but it overclaims the word 'systematic.' The three-stage pipeline and four-paradigm taxonomy are reasonable and will help newcomers find their way. The coverage is broader than most prior reviews—text, vision, audio, EEG, then multimodal fusion, alignment/disentanglement, knowledge-augmented, and multi-task methods—and the dataset table, metric summary, and application list give a decent starting point. I checked several method descriptions against the cited literature and they are mostly faithful.\n\nThe soft spots are real but not fatal. The paper calls itself 'the first systematic review' in Section 1, yet gives no search protocol, no inclusion/exclusion criteria, and no comparison with earlier surveys such as the 2019 review [81] or the 2023 Springer monograph [143]. Without that, the 'systematic' qualifier is an assertion, and the representativeness of the curated sample is unverified. Curation errors compound the issue: MultiWOZ is described as 10,438 samples with 8,438 conversations, but those are dialogue counts and a training split; references [76]/[77] and [134]/[135] are duplicate entries for the same papers; MDID is discussed in Section 3.2 but absent from Table 2; and MIntOOD, KDSL, and MMSAIR appear in the text but not in Table 4. Individually these are minor, but together they undercut the reproducibility the paper claims.\n\nThe central taxonomy itself holds up. It is inspired by MIntRec rather than derived from it, which is fine for a survey. The method summaries in Section 4 are coherent, and the discussion of challenges in Section 7 is level-headed. The 'systematic' framing and the curation problems are what need work, not the core synthesis.\n\nThis is a paper I would send to peer review rather than desk-reject. A serious referee could ask for a stated search strategy, fixed tables, deduplicated references, and a tempered contribution claim. After those revisions it would be a citable resource for the subfield. As it stands, I would not cite it yet, and I am not sure I would put it on the reading group schedule until the numbers and tables are cleaned up.","headline":"A useful organizing survey whose 'first systematic review' claim is not backed by any disclosed methodology, and whose curation errors make the synthesis hard to check.","tokens_in":32865,"tokens_out":1474,"would_cite":false,"duration_ms":21241,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey traces intent recognition from text-only models to multimodal deep learning, organizing the field around ten benchmark datasets and a four-paradigm taxonomy.","keywords":["multimodal intent recognition","intent recognition","multimodal learning","deep learning","survey","benchmark datasets","intent taxonomy","large language models"],"falsifier":"A comprehensive, reproducible literature search over the same period that finds either an earlier survey covering the same unimodal-to-multimodal trajectory, or a substantial cluster of multimodal intent recognition methods that cannot be placed into any of the four paradigms (fusion, alignment and disentanglement, knowledge-augmented, multi-task coordination), would refute the paper's central claims; a concrete version would count what fraction of a random sample of 2019–2025 multimodal intent recognition papers falls outside the four paradigms.","tokens_in":31798,"feed_emoji":"🗺️","tokens_out":6176,"duration_ms":62033,"temperature":0.7,"pith_summary":"This paper sets out to establish a structured map of deep learning for intent recognition, tracing the field from text-only models through vision, audio, and EEG to multimodal systems that fuse several signals at once. Its central claim is that multimodal intent recognition is the natural next stage because single modalities fail under noise, ambiguity, and missing context. To make the field tractable, the paper organizes methods around a three-stage pipeline and a four-paradigm taxonomy, and it compiles the widely used datasets and evaluation metrics into a single reference. A reader would care because the survey gives researchers a common vocabulary, a benchmark baseline, and an explicit list of open problems—ambiguity, multi-intent utterances, evolving dialogue intent, modality asynchrony, out-of-domain inputs, long-tail labels, cross-lingual gaps, and continuous reasoning in dynamic environments.","feed_headline":"Survey maps intent recognition's shift from text to multimodality","feed_subtitle":"Ten benchmark datasets, four modeling paradigms, and eight open challenges in one structured overview.","key_machinery":"The load-bearing organizing device is the three-stage multimodal intent recognition pipeline—Feature Extraction, Multimodal Representation Learning, and Intent Classification—with representation learning treated as the core stage. Within that stage, the paper's central taxonomy divides methods into four paradigms: Fusion Methods (feature-, decision-, and hybrid-level), Alignment and Disentanglement Methods (contrastive learning, cross-modal attention, disentangled encoders), Knowledge-Augmented Methods (LLM-based and retrieval-based), and Multi-Task Coordination Methods (joint optimization of intent with related tasks such as emotion recognition). This machinery is what lets the survey place dozens of separate model papers into a single comparative structure.","core_discovery":"In the paper's own framing, the core discovery is that intent recognition has undergone a coherent evolution that can be systematically reviewed: initial rule- and feature-based text methods gave way to deep learning, and increasingly to Transformer- and LLM-based models, while the field simultaneously expanded from text to vision, audio, EEG, and multimodal combinations. The authors claim to present the first systematic review covering this whole trajectory, and they support the claim with a catalog of ten datasets, a coarse-grained intent taxonomy (Emotion and Attitude, Goal Achievement, Information and Declaration), and a classification of multimodal methods into fusion, alignment and disentanglement, knowledge-augmented, and multi-task coordination paradigms. They further assemble evaluation metrics, including Accuracy, Precision, Recall, F1, and specialized out-of-scope measures, and identify application areas from human-computer interaction to automotive systems and sports.","pith_inferences":["The survey's implicit emphasis on text as the anchor modality suggests that robustness to missing modalities—models that must still work when video or audio is absent—could become the next standard evaluation criterion; the paper lists this issue but does not develop a benchmark protocol for it.","A testable consequence of the four-paradigm taxonomy is coverage: applying the taxonomy to a broader, systematically sampled set of 2019–2025 papers would reveal whether any substantial family of multimodal intent methods falls outside the four paradigms.","The inclusion of EEG and eye-tracking datasets points to cognitive-signal fusion as an underexplored frontier, where intent is inferred from neural and gaze data rather than spoken or typed language; this direction is visible in the survey's dataset table but not developed as a research program.","The recurring use of large language models for label description, knowledge extraction, and reasoning suggests that the field may converge on LLM-generated pseudo-labels or knowledge as a standard data-augmentation step, an implication the survey leaves implicit."],"forward_implications":["New multimodal intent recognition systems can be described and compared against a common pipeline, so performance numbers from different papers become easier to relate to one another.","The dataset catalog of ten benchmarks gives researchers a ready-made evaluation foundation for unimodal and multimodal settings, including out-of-scope detection benchmarks.","The four-paradigm taxonomy suggests that future work will increasingly combine paradigms—for example, knowledge augmentation used inside a contrastive alignment framework—rather than staying within one of them.","The eight named challenges act as a research agenda: work on ambiguity, multi-intent structure, dialogue-level intent evolution, modality asynchrony, OOD detection, long-tail labels, cross-lingual generalization, and continuous reasoning are all identified as open problems."],"supporting_citations":[{"why":"Supplies the MIntRec benchmark that motivates the coarse-grained intent taxonomy and anchors multimodal intent recognition experiments.","marker":"[163]"},{"why":"Provides the large-scale MIntRec2.0 benchmark for in-scope classification and out-of-scope detection in multi-party dialogues.","marker":"[162]"},{"why":"Contributes the MINE dataset, the first real-world benchmark with modality incompleteness and joint emotion-intention labels.","marker":"[149]"},{"why":"Offers the MC-EIU multilingual dataset for coupled emotion and intent understanding across text, audio, and video.","marker":"[84]"},{"why":"Supplies the SLURP spoken language understanding benchmark with hierarchical scenario, action, and entity annotations.","marker":"[5]"},{"why":"Defines the Intentonomy visual intent dataset and its social-psychology-based taxonomy of 28 intent categories.","marker":"[53]"},{"why":"Provides the IntentQA video-question-answering dataset for context-aware intent reasoning and the CaVIR baseline.","marker":"[66]"},{"why":"Introduces BERT, the pre-trained transformer model that the survey identifies as the key breakthrough for modern text intent recognition.","marker":"[23]"},{"why":"Introduces the transformer architecture underpinning the pre-trained and LLM-based methods surveyed across modalities.","marker":"[130]"},{"why":"Contributes the MindGaze dataset combining EEG and eye tracking for navigation and information intent prediction.","marker":"[111]"}],"fun_headline_variants":["Survey: deep learning drives intent recognition beyond text","Intent recognition: multimodal deep learning, surveyed","Deep learning's path from text to EEG in intent recognition","Ten datasets, four paradigms: the intent recognition landscape","How deep learning expanded intent recognition to vision and audio"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's map of the field relies on a hand-picked collection of datasets and methods chosen without a stated systematic search or inclusion criteria, so the proposed pipeline and taxonomy hold only if that selection is representative.","fun_headline_variants_meta":{"raw":{"variants":["Survey: deep learning drives intent recognition beyond text","Intent recognition: multimodal deep learning, surveyed","Deep learning's path from text to EEG in intent recognition","Ten datasets, four paradigms: the intent recognition landscape","How deep learning expanded intent recognition to vision and audio"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000586,"raw_usage":{"total_tokens":2682,"prompt_tokens":804,"completion_tokens":1878,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":420,"completion_tokens_details":{"reasoning_tokens":1804}},"tokens_in":420,"tokens_out":1878,"duration_ms":13208,"temperature":1.0,"reasoning_tokens":1804,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T14:31:43.808709+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A comprehensive, reproducible literature search over the same period that finds either an earlier survey covering the same unimodal-to-multimodal trajectory, or a substantial cluster of multimodal intent recognition methods that cannot be placed into any of the four paradigms (fusion, alignment and disentanglement, knowledge-augmented, multi-task coordination), would refute the paper's central claims; a concrete version would count what fraction of a random sample of 2019–2025 multimodal intent recognition papers falls outside the four paradigms.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the MIntRec benchmark that motivates the coarse-grained intent taxonomy and anchors multimodal intent recognition experiments."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the large-scale MIntRec2.0 benchmark for in-scope classification and out-of-scope detection in multi-party dialogues."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Contributes the MINE dataset, the first real-world benchmark with modality incompleteness and joint emotion-intention labels."}],"review_version":1}