{"id":"2775064f-b8da-4a66-972d-a293eec206fe","arxiv_id":"2412.17616","paper_version":3,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey that unifies macro-expression and micro-expression analysis under one learning-paradigm taxonomy and maps them onto Internet of Things applications.","lead":"This paper surveys how facial expression recognition, covering both macro-expressions and micro-expressions, is being combined with Internet of Things systems for healthcare, security, and emotion-aware services. It organizes a large body of recent methods by learning paradigm and application, which makes it a useful entry point for engineers and researchers choosing where facial analysis fits in edge computing.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim of a comprehensive, gap-bridging survey rests on an informal selection of papers; without a documented search strategy, its completeness and its asserted novelty over prior surveys cannot be checked, and omitted integrated surveys could falsify the gap.","rationale":"The reader identified the same weakest assumption, so I agree. My read adds a concrete falsification route: the novelty claim 'existing surveys treat MaE and MiE in isolation' is checkable by structured search. The paper's internal organization does not formally define the 'holistic framework' — Sections 4 and 5 are parallel taxonomies and Section 6 lists applications separately — but a survey can still be valuable as a map; the main unverified premise is coverage, not mathematical correctness. Since the reader already marked the paper UNVERDICTED on exactly this basis, my concern does not move the verdict; it reinforces it. I would keep UNVERDICTED as the appropriate scholarly status until the authors document their selection methodology or the coverage check is run. No ad hominem; the issue is methodological transparency.","tokens_in":51188,"tokens_out":3132,"duration_ms":33088,"concrete_test":"Run a systematic coverage check: search DBLP/Scopus/arXiv (2015-2024) with a Boolean query combining 'facial expression', 'micro-expression', 'macro-expression', 'survey'/'review', and 'IoT'/'Internet of Things'/'edge computing'; screen against explicit inclusion criteria; then compare the retrieved set with the references cited in Sections 2 and 6. If the search returns one or more prior surveys that already treat MaE and MiE jointly or that survey facial-expression IoT applications, and those surveys are absent from the paper, the claimed gap and comprehensiveness are falsified; if the retrieved set is covered, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's contribution, stated in Section 2.2, is that existing surveys treat MaE and MiE in isolation and that this work 'bridges this gap' with a holistic framework and a learning-paradigm taxonomy. The load-bearing premise is that the selected literature is representative enough to support that novelty claim. Yet Section 2 gives no reproducible selection protocol: no databases, query strings, time window, inclusion/exclusion criteria, or screening steps. Without such a protocol, comprehensiveness is an assertion, not a demonstrated property. The risk is concrete: if a structured search finds prior surveys (or application-focused reviews) that already integrate MaE and MiE, or that cover facial-expression IoT applications with comparable breadth, then the claimed gap over prior work and the framing as a 'contemporary survey' both weaken. This is a completeness premise rather than an internal inconsistency, but it is load-bearing because the map is the contribution; a map with undocumented coverage cannot be assessed. The misplaced appendices and typos are secondary correctness risks, not the decisive issue.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of facial expression analysis, covering both macro-expression (MaE) and micro-expression (MiE) recognition, and discusses their potential applications in Internet-of-Things (IoT) systems. The authors organize recent work by learning paradigm (e.g., ensemble learning, transfer learning, multi-task learning, attention-based learning, self-supervised learning) for MaE analysis, and by spotting versus recognition for MiE analysis. They also catalog datasets, summarize representative methods with performance numbers, and dedicate sections to IoT applications in emotion recognition, healthcare, security monitoring, negotiation, and mental disorder detection. The paper claims three unique contributions: bridging the gap between MaE and MiE surveys, proposing a learning-paradigm-based framework, and emphasizing practical IoT deployment.","tokens_in":51356,"tokens_out":3391,"duration_ms":34895,"significance":"If the survey's coverage is complete and its comparisons are accurate, it would be a useful structured reference for researchers working at the intersection of facial expression analysis and IoT systems. The paper contains a large number of references, detailed tables of datasets and methods with performance figures that are generally traceable to the cited literature, and a taxonomy that could help readers navigate the field. The appendices provide useful paradigm-level comparisons. However, the main contribution is the map itself, so the credibility of the survey hinges on the completeness and representativeness of the selected literature. The absence of a documented, reproducible survey methodology is a significant weakness: it prevents a reader from verifying that the claimed gap over prior surveys is real and that no major integrated surveys were overlooked. The internal inconsistencies in dataset statistics also need correction before the paper can serve as a reliable reference.","major_comments":[{"comment":"The paper's central novelty claim is that 'existing surveys treat MaE and MiE in isolation' and that this work 'bridges this gap' with a holistic framework. This is a load-bearing claim, yet the survey provides no reproducible selection protocol: no databases searched, no query strings, no time window, no inclusion/exclusion criteria, and no screening steps are reported. Without such a protocol, 'comprehensive overview' is an assertion rather than a demonstrated property. The risk is concrete: if prior surveys or application-focused reviews that already integrate MaE and MiE were missed, the claimed gap and the framing as a contemporary survey would be weakened. This issue must be addressed, at minimum by adding a methodology subsection that documents the search and screening process, and by re-evaluating the novelty claim in light of the full set of prior surveys found by that search.","section":"§2.2"},{"comment":"There is an internal inconsistency in the EmotioNet dataset count: the text states 'a total of 9500,000 images were annotated with AUs, AU intensities, and emotion categories,' while Table 1 reports '950,000 images.' This is not a trivial typo because the dataset scale is a key piece of information in a survey's dataset overview. The authors should correct the number and verify that all dataset statistics in Tables 1 and 2 are consistent with the primary sources and with the accompanying text.","section":"§3.1 and Table 1"},{"comment":"The text repeatedly refers to 'Appendix 3.1,' 'Appendix 3.2,' 'Appendix 4.1,' 'Appendix 4.2,' and 'Appendix 5,' but these appendices appear after the reference list as Sections 12, 13, and 14, and they are numbered inconsistently. For example, the appendix for static MaE comparisons is labeled '12.1' in the appendix itself but is referred to as 'Appendix 3.1' in the main text. This cross-reference mismatch makes it unnecessarily difficult for readers to locate the comparative analyses that are supposed to support the survey's insights. The authors should renumber the appendices or update the in-text references so that they match.","section":"§4.2.6, §5.1.3, and §13"}],"minor_comments":[{"comment":"There are several typographical errors in Table 1 and the text, including 'imgaes' for 'images,' '35762 imgaes' for '35,762 images,' and inconsistent use of commas in numbers. These should be corrected throughout.","section":"§3.1, Table 1"},{"comment":"The phrase 'Since the seminar work [165]' should be 'seminal work.' The term 'seminar' changes the meaning and is clearly a typo.","section":"§5.2"},{"comment":"In the sentence about traditional 3D CNNs, 'can increased latency' should be 'can increase latency.' Also in the same section, 'hinders the the training' contains a duplicated 'the.'","section":"§4.2.6"},{"comment":"The description of MaE and MiE characteristics in Section 1 is repeated almost verbatim in Section 10. Since Section 10 is an appendix-like illustration section, the duplication should be removed or one of the passages should be shortened to avoid redundancy.","section":"§1 and §10"},{"comment":"The term 'holographic MiE analysis' is used without definition or motivation. If 'holographic' is intended to mean 'holistic' or 'complete,' the authors should either define the term or replace it with clearer wording.","section":"§5 title and throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper is broadly within the scope of the journal and addresses a timely topic. The main risk is the undocumented survey methodology: as a survey paper, its contribution is the map of the literature, and without a reproducible selection process the completeness claim cannot be assessed. The editorial decision should weigh whether the authors can supply a methodology section and revalidate the novelty claim. The dataset inconsistencies and appendix cross-reference errors are correctable but should be fixed in revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a competent survey of facial expression analysis that separates macro- and micro-expressions, organizes methods by learning paradigm, and ties both to IoT applications. The dataset tables are genuinely useful, especially the MiE table with annotation details, and the method summaries mostly track the cited literature. The IoT applications section is the most distinctive part; it is thinner than the methods sections but it does connect the two fields in a way I have not seen neatly packaged before. Credit is due for that.\n\nThe soft spots are real but not fatal. The load-bearing claim is that existing surveys treat MaE and MiE in isolation and this one bridges the gap. That claim is checked only against the handful of surveys the authors chose to discuss. There is no search protocol, no inclusion criteria, no time window, so 'comprehensive' is asserted rather than demonstrated. The stress-test note is right: if a structured search finds an integrated survey, the novelty framing weakens. That is a legitimate completeness concern, not a manufactured one.\n\nThere are also several typos that should have been caught: EmotioNet is described as '9500,000 images' in one place and 950,000 in the table; the appendices appear after the references and duplicate text from the main body; the 'Structure of Our Work' section is stranded in an appendix; and a few citations look mismatched (Chung et al.'s GRU paper appears in the security monitoring paragraph without a clear tie to the claim). None of this undermines the core summaries, but it signals a final proofread was skipped.\n\nFor a survey, the absence of a new method or result is not a flaw. The value here is organizational: a newcomer to facial expression analysis with an IoT interest could use this to get oriented quickly. I would not cite it as the definitive authority on coverage, but I would point students to the dataset tables and the IoT discussion.\n\nMy recommendation: send it to peer review. A serious referee can ask for a documented selection strategy, a toned-down novelty claim, and cleanup of the typos and formatting. The paper has enough substance to warrant that investment.","headline":"A solid but not exceptional survey; the IoT lens is the real differentiator, but the 'gap-bridging' claim rests on an undocumented literature selection and needs cleanup before it earns that framing.","tokens_in":51925,"tokens_out":1639,"would_cite":true,"duration_ms":19707,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey claims that macro-expressions and micro-expressions are two halves of one facial-emotion analysis problem, and proposes a learning-paradigm-based framework that connects both to edge-driven IoT systems for healthcare and…","keywords":["facial expression analysis","macro-expression recognition","micro-expression analysis","Internet of Things","edge computing","emotion recognition","deep learning","survey"],"falsifier":"Find a peer-reviewed survey, published before this one, that already reviews both macro-expression and micro-expression analysis together with IoT applications under a single framework; if such a survey exists and is not referenced, the paper's central claim of a missing holistic integration collapses.","tokens_in":50989,"feed_emoji":"😊","tokens_out":4539,"duration_ms":40320,"temperature":0.7,"pith_summary":"This survey argues that facial expression analysis should be understood as two complementary problems—macro-expressions (MaEs), which are voluntary and last 0.5–4 seconds, and micro-expressions (MiEs), which are involuntary, last under 0.5 seconds, and reveal concealed emotion—and that both belong together in a single framework for Internet-of-Things (IoT) systems. It claims that prior surveys treated MaE and MiE in isolation, leaving a gap: no existing review offered a holistic, learning-paradigm-based map of methods, datasets, and IoT applications for both. The paper organizes MaE methods into ensemble, transfer, multi-task, attention-based, and self-supervised paradigms, and separates MiE analysis into spotting (finding the fleeting expression) and recognition (classifying it). It then shows how each can be deployed on edge-driven IoT architectures for healthcare, security, negotiation, and mental-health monitoring. If the framework is right, a reader gets a structured route from raw facial video to emotion-aware IoT services.","feed_headline":"One map links macro- and micro-expressions to IoT systems","feed_subtitle":"A single taxonomy of methods and datasets shows how emotion-sensing can run on edge devices for healthcare and security.","key_machinery":"The load-bearing structure is a dual distinction plus a taxonomy. The first distinction is temporal and intentional: MaEs last roughly 0.5–4 seconds and are voluntary; MiEs last under 0.5 seconds, are involuntary, and carry about 55% of emotional messages. The second is the learning-paradigm taxonomy: for MaE, the paper partitions methods into ensemble learning, transfer learning, multi-task learning, attention-based learning, and self-supervised learning, for both static images and dynamic videos; for MiE, it partitions the pipeline into spotting (locating onset-apex-offset intervals) and recognition, and reviews descriptors (LBP variants, HOG, optical flow), deep transfer, multi-task, self-supervised, lightweight, meta-learning, and GAN-based generation. This taxonomy is what allows the paper to claim a holistic framework that connects fundamental research to IoT applications, with edge devices collecting and preprocessing before offloading inference.","core_discovery":"The paper's central claim is that macro-expression recognition and micro-expression analysis, usually surveyed as separate fields, are two halves of one problem that a learning-paradigm-based taxonomy can unify, and that this unification is the missing bridge to practical IoT deployment. It asserts that MaE analysis has matured (over 97% accuracy in controlled labs) while MiE analysis remains hard (about 47% accuracy even after training), so a holistic framework must treat them with different pipelines—static and dynamic deep recognition for MaEs, spotting-then-recognition for MiEs—under one structure. It further claims that edge computing is the natural hosting layer because it provides low latency and privacy for biometric facial data. On the application side, it argues that MaE-driven IoT handles real-time emotion monitoring, while MiE-driven IoT handles concealed-emotion tasks such as lie detection, security surveillance, and depression screening.","pith_inferences":["The framework implies a natural product architecture: a single edge device could run a MaE classifier continuously and trigger a slower MiE spotter-and-recognizer only when an expression is too brief or too suppressed to be classified as MaE.","Because the taxonomy is paradigm-based rather than architecture-based, it could be lifted to other subtle behavior tasks (e.g., gesture or gaze micro-movements) where spotting precedes recognition.","The paper underplays the possibility that MiE recognition categories are not standardized across datasets; cross-dataset few-shot and AU-grounded methods point to where the field would need to standardize labels.","If the IoT integration is taken seriously, the next bottleneck is not accuracy but privacy and real-time constraints; the survey's own challenges section lists these but leaves concrete protocol design open."],"forward_implications":["A researcher entering facial expression analysis can use the taxonomy to locate which learning paradigm (e.g., transfer vs. self-supervised) fits their data size and deployment target.","IoT systems for smart healthcare can adopt the edge-driven architecture described here, running MaE models for real-time emotion monitoring and MiE models for detecting concealed distress.","Security applications can combine MaE and MiE pipelines: MaE for visible state, MiE for deception or threat cues in high-stakes settings.","The survey's comparison tables give concrete accuracy expectations across datasets and paradigms, helping practitioners set realistic baselines.","Future methods can position themselves against this framework, filling gaps such as MiE spotting in long, unconstrained videos."],"supporting_citations":[{"why":"Prior survey of deep macro-expression recognition; the MaE baseline this paper builds on.","marker":"[108]"},{"why":"Prior comprehensive survey of micro-expression analysis; one of the papers that treated MiE in isolation.","marker":"[10]"},{"why":"Prior survey of deep learning for micro-expression recognition; defines the taxonomy this paper extends.","marker":"[120]"},{"why":"Recent overview of micro-expression research; shows the MiE literature's scope and gaps.","marker":"[271]"},{"why":"Systematic review of MaE recognition methods; establishes the MaE survey landscape.","marker":"[13]"},{"why":"Reports about 47% trained-human accuracy on micro-expressions; grounds the difficulty claim that motivates MiE research.","marker":"[50]"},{"why":"Provides duration ranges for MaE (0.5-4 s) and MiE (<0.5 s); the core temporal distinction used throughout the paper.","marker":"[185]"},{"why":"The Extended Cohn-Kanade dataset; a load-bearing MaE benchmark used across reviewed methods.","marker":"[135]"},{"why":"SMIC spontaneous micro-expression corpus; foundational MiE dataset used for spotting and recognition.","marker":"[110]"}],"fun_headline_variants":["Survey unifies macro and micro expressions for IoT","IoT meets facial expressions: macro and micro in one view","One taxonomy for emotion-driven IoT systems","Macro and micro expressions mapped to IoT applications"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's usefulness as a map depends on its informal, non-reproducible selection of papers; if prior work already integrated MaE and MiE with IoT, or if key integrated surveys were missed, the claimed gap and the map itself would be incomplete.","fun_headline_variants_meta":{"raw":{"variants":["Survey unifies macro and micro expressions for IoT","IoT meets facial expressions: macro and micro in one view","One taxonomy for emotion-driven IoT systems","Macro and micro expressions mapped to IoT applications"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000224,"raw_usage":{"total_tokens":1452,"prompt_tokens":928,"completion_tokens":524,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":544,"completion_tokens_details":{"reasoning_tokens":465}},"tokens_in":544,"tokens_out":524,"duration_ms":5084,"temperature":1.0,"reasoning_tokens":465,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T05:20:38.017538+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Find a peer-reviewed survey, published before this one, that already reviews both macro-expression and micro-expression analysis together with IoT applications under a single framework; if such a survey exists and is not referenced, the paper's central claim of a missing holistic integration collapses.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Recent overview of micro-expression research; shows the MiE literature's scope and gaps."}],"review_version":1}