{"id":"10ccaaac-0bf7-4edc-9196-1cef3c6f8ff9","arxiv_id":"2507.08344","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Combining joint, limb, RGB, Taylor-video, optical-flow, and depth streams with two video backbones and a validation-tuned weighted ensemble reaches 73.213% top-1 accuracy on iMiGUE, the best MiGA challenge result to date.","lead":"A team that won the 2025 micro-gesture classification challenge describes a six-modality fusion system that recognizes subtle gestures by combining skeletal, RGB, motion, and depth cues, reaching 73.213% top-1 accuracy on the iMiGUE benchmark. The result tops all previous MiGA challenge entries and shows incremental gains from adding each extra modality.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The SOTA claim is not established: only MiGA challenge leaderboards are compared, not all published iMiGUE methods.","rationale":"The reader's verdict is CONDITIONAL, and the reader's rationale already notes 'no comparison to published non-challenge SOTA methods' as a deficiency. I agree with that concern and consider it the most load-bearing: the central claim of 'state-of-the-art' or 'current best-performing method on iMiGUE' is directly falsifiable by a single counterexample from the published literature, and the paper provides no evidence against that possibility. I disagree with the reader's choice of weakest_assumption, which is the generalization of validation-tuned ensemble weights. That assumption does not threaten the reported test accuracy: the test set is a fixed, independent evaluation, and the 73.213% figure is an unbiased point estimate for the selected configuration. The weights being unreported affects reproducibility but not the truth of the competition result. By contrast, the missing non-challenge comparison affects the truth of the broader SOTA claim. The concrete test I propose—searching for published iMiGUE results and comparing to 73.213%—would settle whether the concern lands. If a higher published number exists, the paper must be revised to either add the comparison or soften the claim to 'first in the MiGA 2025 challenge.' If no higher number exists, the concern is resolved and the paper's SOTA claim is acceptable. Either way, the reader's CONDITIONAL verdict is appropriate: the paper needs a comparative table of published iMiGUE results and explicit numerical ensemble weights. Therefore, I recommend UNCHANGED.","tokens_in":9051,"tokens_out":16784,"duration_ms":189716,"concrete_test":"Perform a systematic literature search (arXiv, OpenReview, CVF/IEEE/ACM) for all papers reporting top-1 accuracy on the iMiGUE test set, including methods not participating in MiGA challenges. Compare the maximum published accuracy to 73.213%. If any non-challenge published method exceeds 73.213%, the 'state-of-the-art' claim is false; if none does, the claim stands. As a secondary check, run the released code on the iMiGUE test set using the described weighted ensemble and verify the reproduced accuracy is within the standard error (~1%) of 73.213%.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim, as framed in the abstract and conclusion, is that MM-Gesture achieves 'state-of-the-art performance' and, per the reader's formulation, is 'the current best-performing method on the iMiGUE benchmark.' The only comparative evidence is Table 1, which lists results from MiGA challenge leaderboards (2023-2025). These leaderboards include only participating teams, not the full set of published methods that report iMiGUE top-1 accuracy. The paper does not compare against published non-challenge methods, such as prototype-learning or cross-attention approaches from the same research community (e.g., refs. [10], [19], [24], or any later work citing iMiGUE [2]). The precise claim 'the highest reported accuracy in previous MiGA challenges' is narrower than 'state-of-the-art,' yet the abstract and conclusion use the broader term. If any published method reports >73.213% on iMiGUE, the broad SOTA claim is false, even though the competition result itself is a fact. This is a correctness risk, not a matter of consensus: the paper's evidence is systematically incomplete. The validation-set-tuned ensemble weights (Section 3.5, Eq. 7) are a secondary reproducibility concern—the weights are not reported numerically—but they do not invalidate the reported test accuracy, which is a single evaluation on an independent test set. The load-bearing gap is the missing systematic comparison to published iMiGUE results.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents MM-Gesture, a multimodal ensemble system for micro-gesture classification on the iMiGUE dataset. The system integrates six modalities (joint, limb, RGB, Taylor-series video, optical flow, depth) using PoseConv3D and Video Swin Transformer backbones, plus MA-52 pretraining for the RGB stream, and combines the resulting probability streams through a weighted late-fusion ensemble. The authors report a top-1 accuracy of 73.213% on the iMiGUE test set, which ranked first in the 3rd MiGA Challenge at IJCAI 2025, and they claim state-of-the-art performance on this benchmark.","tokens_in":9321,"tokens_out":4889,"duration_ms":52313,"significance":"The external competition result is a strong form of validation: the 73.213% figure comes from an official leaderboard, and the paper releases code, which aids reproducibility. The system appears to be a well-engineered combination of existing components, and the empirical gains over prior MiGA challenge entries (roughly +3% over the 2024 winner) are substantial for this benchmark. The main scientific value is the demonstration that a carefully tuned multimodal ensemble, together with MA-52 transfer learning, achieves the current best reported result in the MiGA challenge series. The methodological novelty is modest, however, as the components are off-the-shelf and the fusion is standard late fusion; the claimed state-of-the-art status also requires a broader comparison than the challenge leaderboards alone.","major_comments":[{"comment":"The claim that MM-Gesture achieves 'state-of-the-art performance' and 'superior performance compared to previous state-of-the-art methods' (abstract and §5) is not supported by the evidence presented. Table 1 compares only the top-3 entries from MiGA challenge leaderboards for 2023–2025. Several published methods that report iMiGUE top-1 accuracy are cited in the related work (e.g., refs. [10], [19], [24]) but are absent from the comparison. If any published method exceeds 73.213% on iMiGUE, the state-of-the-art claim is false, even though the competition ranking itself is an externally verified fact. The authors should add a systematic comparison with all publicly reported iMiGUE results, or explicitly restrict their claim to 'best among MiGA challenge entries.'","section":"§4.2, Table 1"},{"comment":"The ensemble weights w_i are stated to be 'empirically determined weights obtained via validation-set performance,' but the paper does not report the weight values, the search procedure, or the validation accuracy used to select them. This makes the final 73.213% result, and specifically the 0.569% improvement from the 'optimized multimodal fusion weighting strategy' (Table 2, 72.644% to 73.213%), impossible to reproduce from the paper alone. The authors should report the selected weights, describe the validation-based selection method (e.g., grid search or a learning algorithm), and ideally include an ablation comparing equal weighting against the tuned weights.","section":"§3.5, Eq. (7)"},{"comment":"The checkmark layout in Table 2 is ambiguous for the rows with a single backbone name, because the checkmarks are not visually aligned with the column headers in the text, making the modality-to-accuracy mapping unclear (for instance, which single modality yields 65.256%?). Additionally, the final row 'MM-Gesture (Ours)' shows seven checkmarks, while Eq. (7) defines an ensemble of six probability streams; the relationship between the seven table columns (Joint, Limb, RGB, RGB*, Taylor, Flow, Depth) and the six ensemble terms (R+J, R+L, R, T, F, D) needs clarification. The text also reports '72.096%' for the Taylor-inclusive row while Table 2 lists 72.095%; this inconsistency should be fixed.","section":"Table 2"}],"minor_comments":[{"comment":"The Taylor-series expansion parameters are not reported: the maximum order K and the temporal window τ used in ℱ_taylor are never specified, although they are free hyperparameters that may affect the Taylor modality's contribution. Please provide the values used in the experiments.","section":"§3.1, Eq. (1)"},{"comment":"The selection of the 36 skeleton keypoints from the original 137 is described only as focusing on upper body, hands, and facial joints. Please specify the exact keypoint indices or the selection criterion, since this choice influences the joint and limb modalities.","section":"§3.1"},{"comment":"The symbol R is used for both the RGB input to PoseConv3D (Eq. (5)) and the RGB model in the VideoSwinT stream (Eq. (6)), where the latter is pretrained on MA-52. This notation conflict makes it unclear which RGB probability is used in the ensemble Eq. (7). Please disambiguate (e.g., R_P and R_S or R and R* consistently).","section":"§3.4–§3.5"},{"comment":"The paper reports a single evaluation without variance, confidence intervals, or significance tests. The competition leaderboard provides a fixed external result, but for the ablation findings in Table 2, the lack of repeated runs makes it hard to judge whether differences such as 72.227% vs. 72.644% are meaningful. A brief statement acknowledging this limitation would be appropriate.","section":"§4.1"}],"recommendation":"major_revision","confidential_remarks":"This is essentially a competition system report. The externally verified leaderboard result is a real strength, but the missing comparison to published (non-challenge) iMiGUE methods is a substantive gap that must be addressed before the state-of-the-art claim can stand. If the journal is willing to accept benchmark-focused contributions, the paper is potentially suitable after revision; otherwise, its methodological novelty may be too thin for the venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The one thing to know: this is a genuine competition win, externally confirmed on the MiGA 2025 leaderboard, and the 73.213% top-1 accuracy on iMiGUE is probably the best number reported so far. The paper does not open new ground—every component (PoseConv3D, VideoSwinT, Taylor videos, MemFlow, monocular depth, late fusion) is off-the-shelf—but the specific six-modality combination plus MA-52 transfer learning on the RGB stream is a legitimate extension of prior MiGA work, and the authors are honest that it is a systems/ensemble solution. Code is released, which is more than many competition write-ups do.\n\nThe main soft spot is the comparison set. Table 1 lists only MiGA challenge leaderboards from 2023 to 2025. That supports the narrow claim \"highest reported accuracy in previous MiGA challenges,\" which is likely true. But the abstract and conclusion say \"state-of-the-art,\" and that is not established unless the paper compares against published iMiGUE methods outside the challenges—for example, the prototype-learning and cross-attention work the authors cite. If any of those report >73.213%, the broad claim is wrong even though the competition result is a fact. This is a correctness risk, not a style preference. Please fix the wording or add the missing comparisons.\n\nThe other concerns are minor. The ensemble weights in Eq. (7) are tuned on the validation split and not reported numerically, which hurts reproducibility but does not invalidate the independent test evaluation. There are no variance or significance measures—common in competition papers, but worth asking for. The ablation table's checkmark layout for the first six PoseConv3D rows is ambiguous; a reader cannot tell which single modality each row refers to. That should be cleaned up.\n\nOverall: the paper is a solid competition report with an externally verified result and a plausible engineering recipe. My own take is closer to the reader's conditional verdict, but I would not call the validation-set tuning a load-bearing flaw. The missing systematic comparison to published non-challenge methods is the real issue.\n\nFor peer review: yes, send it to reviewers, especially for a workshop venue, with a request to fix the SOTA comparison and report the weights and error bars. It deserves referee time because the empirical result is externally confirmed and the method is clearly described; the claims just need to match the evidence.","headline":"Credible competition result with an externally verified first place, but the SOTA claim needs a broader comparison; the ensemble/transfer-learning recipe is a solid engineering contribution, not a scientific breakthrough.","tokens_in":9896,"tokens_out":1752,"would_cite":false,"duration_ms":22124,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MM-Gesture claims 73.213% top-1 on iMiGUE by fusing six modalities.","keywords":["Micro-Gesture Recognition","Multimodal Fusion","Ensemble Learning","Transfer Learning","PoseConv3D","Video Swin Transformer","iMiGUE","Action Recognition"],"falsifier":"Retrain the six modality models, tune the ensemble weights on the validation set, and evaluate on the test set; if the RGB* model alone, with a weight of 1 for that modality and 0 for the others, matches or exceeds 73.213% on the test set, then the claimed benefit of the full multimodal ensemble is falsified.","tokens_in":8843,"feed_emoji":"🤏","tokens_out":4084,"duration_ms":42670,"temperature":0.7,"pith_summary":"This paper tries to show that micro-gesture recognition improves substantially when six complementary modalities are fused instead of relying on RGB or skeleton alone. The authors present MM-Gesture, which combines PoseConv3D on skeleton joint and limb heatmaps with Video Swin Transformer on RGB, Taylor-series, optical-flow, and depth videos, then blends the probability outputs with weights tuned on the validation split. On the iMiGUE benchmark, this reaches a top-1 accuracy of 73.213%, the highest reported in the MiGA challenges from 2023 to 2025 and about 3 percentage points above the previous best entry. The underlying message is that off-the-shelf modality extractors plus a simple weighted ensemble, aided by pretraining the RGB branch on the larger MA-52 dataset, are enough to set a new state of the art.","feed_headline":"Six modalities push micro-gesture accuracy to 73.2%","feed_subtitle":"MM-Gesture tops the 2025 MiGA challenge on iMiGUE by fusing skeleton, RGB, Taylor, flow, and depth cues.","key_machinery":"The central mechanism is a probability-level weighted ensemble (Eq. 7) that combines outputs from six modality-specific models. Skeleton joints and limbs are converted into Gaussian heatmap volumes and processed by PoseConv3D, with RGB fused through paired training losses. RGB, Taylor-series, optical-flow, and depth videos are each encoded by Video Swin Transformer, and the RGB branch is pretrained on MA-52 before fine-tuning on iMiGUE. The weights $w_i$ are set empirically from validation performance, which is what allows the final accuracy of 73.213%.","core_discovery":"The paper's central claim is that a six-modality ensemble, with each modality-specific network trained separately and combined at the probability level, outperforms every prior micro-gesture method on the iMiGUE dataset. The authors report that a top-1 accuracy of 73.213% is achieved, ranking first in the 3rd MiGA Challenge at IJCAI 2025. The ablation results show a consistent incremental gain from adding modalities: skeleton plus RGB reaches 71.416%, adding the Taylor modality gives 72.096%, optical flow gives 72.227%, and depth gives 72.644%, with the final optimized weighting reaching 73.213%. Transfer learning from the MA-52 dataset improves the RGB branch, and the final ensemble weights are chosen empirically by validation performance.","pith_inferences":["The ensemble weights are selected on only 777 validation samples, so the 73.213% test figure could be sensitive to that particular validation split; repeated evaluation across different validation splits would show how stable the result is.","Because the modality extractors are all off-the-shelf, the same weighted late-fusion recipe could transfer to other fine-grained behavior benchmarks, with the relative contribution of each modality possibly shifting.","Depth and optical flow together add less than one percentage point over skeleton plus RGB plus Taylor, so a cheaper two- or three-stream system with MA-52 pretraining might recover most of the benefit at lower computational cost."],"forward_implications":["If MM-Gesture is correct, fusing skeleton, RGB, Taylor, optical flow, and depth yields 73.213% top-1 accuracy on iMiGUE, about 3 points above the previous best challenge entry.","Each added modality contributes a small but consistent gain, which suggests that the six modalities carry complementary information rather than redundant cues.","Pretraining the RGB branch on the MA-52 dataset improves RGB-only accuracy, making transfer learning from a larger micro-action dataset a reusable ingredient.","The winning result comes from a simple late-fusion weighted ensemble with validation-tuned weights, not from end-to-end joint training of all modalities.","The ranking outcome implies that this combination of backbone architectures, modality extractors, and ensemble weighting is currently the best-performing method on the iMiGUE benchmark."],"supporting_citations":[{"why":"Supplies the iMiGUE dataset, its 32-class micro-gesture annotation, and the train/validation/test split used in all experiments.","marker":"[2]"},{"why":"Provides the Video Swin Transformer backbone used to encode RGB, Taylor, optical-flow, and depth modalities.","marker":"[9]"},{"why":"Provides the MA-52 dataset used to pretrain the RGB branch before fine-tuning on iMiGUE.","marker":"[11]"},{"why":"Provides the PoseConv3D backbone and the 3D heatmap representation used for joint, limb, and RGB-plus-skeleton fusion.","marker":"[21]"},{"why":"Provides the Taylor-series temporal expansion method used to generate the Taylor video modality.","marker":"[25]"},{"why":"Provides the MemFlow network used to estimate optical flow from consecutive frames.","marker":"[26]"},{"why":"Provides the monocular depth estimation method used to generate the depth video modality.","marker":"[27]"}],"fun_headline_variants":["73.2% accuracy: MM-Gesture wins micro-gesture challenge","Six modalities fuse to top micro-gesture leaderboard at 73.2%","MM-Gesture: multimodal ensemble nets 73.2% in MiGA challenge","73.2% top-1 accuracy via six-way fusion of gesture cues","Precise micro-gesture recognition: winner fuses six modalities at 73.2%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the ensemble weights $w_i$ tuned on the 777-sample validation split generalize to the 4,562-sample test split; if that premise fails, the reported 73.213% accuracy is not a reliable estimate of true performance.","fun_headline_variants_meta":{"raw":{"variants":["73.2% accuracy: MM-Gesture wins micro-gesture challenge","Six modalities fuse to top micro-gesture leaderboard at 73.2%","MM-Gesture: multimodal ensemble nets 73.2% in MiGA challenge","73.2% top-1 accuracy via six-way fusion of gesture cues","Precise micro-gesture recognition: winner fuses six modalities at 73.2%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000335,"raw_usage":{"total_tokens":1833,"prompt_tokens":898,"completion_tokens":935,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":514,"completion_tokens_details":{"reasoning_tokens":824}},"tokens_in":514,"tokens_out":935,"duration_ms":7743,"temperature":1.0,"reasoning_tokens":824,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T18:20:41.047695+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retrain the six modality models, tune the ensemble weights on the validation set, and evaluate on the test set; if the RGB* model alone, with a weight of 1 for that modality and 0 for the others, matches or exceeds 73.213% on the test set, then the claimed benefit of the full multimodal ensemble is falsified.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the PoseConv3D backbone and the 3D heatmap representation used for joint, limb, and RGB-plus-skeleton fusion."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Taylor-series temporal expansion method used to generate the Taylor video modality."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the MemFlow network used to estimate optical flow from consecutive frames."}],"review_version":1}