{"id":"5eeaf0fa-f4a0-4cb5-9de9-d20bea323d78","arxiv_id":"2606.00170","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"UF-AMA fuses EEG and eye-tracking via transformers and cross-attention, applies confidence-based screening with alignment and distillation, and uses multi-level domain adaptation to reach SOTA on cross-subject and cross-session tasks in SEED and SEED-IV datasets.","lead":"The paper proposes UF-AMA, a framework combining transformer-based cross-modal fusion, confidence-aware sample screening, and multi-level domain adaptation to handle distribution shifts in EEG and eye-tracking data for emotion recognition across subjects and sessions. A smart generalist might read it to see how multimodal physiological signals can be made more robust for real-world affective computing applications despite individual differences.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Generalization of confidence-aware screening + multi-level adaptation to truly unseen subjects/sessions without new bias remains the least-secured assumption","rationale":"Reader's weakest_assumption directly identifies the same load-bearing point. Full text does not add external validation or theoretical guarantees that would remove the concern, so the UNVERDICTED / LOW rating is unaffected.","tokens_in":1776,"tokens_out":324,"duration_ms":14181,"concrete_test":"Re-run the cross-subject leave-one-subject-out protocol on SEED while varying the screening threshold over [0.6, 0.7, 0.8, 0.9] and report mean accuracy plus standard deviation across the 15 folds; if any threshold produces >4% absolute drop relative to the reported best or if variance exceeds 3%, the screening step is not robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The SOTA claim on cross-subject and cross-session tasks rests on the premise that the confidence-aware screening (which partitions samples by predictive reliability) and the joint marginal/conditional alignment at local and global levels will reliably shrink distribution shift on target-domain samples without overfitting or introducing selection bias. This premise is tested only via internal ablations and results on the same two datasets (SEED, SEED-IV) whose subject/session statistics are already known; no external hold-out corpus, no sensitivity sweep on the screening threshold, and no reporting of per-fold variance or statistical significance tests are described that would confirm the mechanisms remain stable when the underlying subject distribution changes.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript proposes UF-AMA, a unified framework for cross-domain multimodal emotion recognition from EEG and eye-tracking signals. It consists of a cross-modal feature fusion network using Transformer encoders and multi-head cross-attention, a confidence-aware screening mechanism that partitions target-domain samples by predictive reliability into quality subsets for global consistency alignment and cross-modal distillation, and a multi-level domain adaptation module that jointly aligns marginal and conditional distributions at local modality-specific and global fusion levels. The central claim is that this approach achieves state-of-the-art performance on cross-subject and cross-session tasks on the SEED and SEED-IV datasets.","tokens_in":1907,"tokens_out":434,"duration_ms":21227,"significance":"If the performance claims are substantiated with full experimental details, the work could advance robust multimodal physiological signal processing for emotion recognition by addressing distribution shifts through adaptive screening and multi-granularity alignment. The public release of source code at the cited GitHub repository is a clear strength that supports reproducibility.","major_comments":[{"comment":"Abstract: The state-of-the-art performance claim is asserted without any quantitative results, baseline comparisons, ablation studies, error bars, or statistical significance tests, which is load-bearing for the central empirical claim and prevents verification of the magnitude of improvement or the contribution of the confidence-aware screening and multi-level adaptation components.","section":"Abstract"},{"comment":"Experimental validation (as summarized in the abstract): The generalization premise that the confidence-aware screening mechanism and multi-level marginal/conditional alignment reliably reduce distribution shifts without introducing selection bias or overfitting is tested only on the SEED and SEED-IV datasets; no external hold-out corpus, no sensitivity analysis on the screening threshold, and no per-fold variance reporting are described, leaving the robustness claim under-supported.","section":"Experimental validation"}],"minor_comments":[{"comment":"The abstract could be strengthened by briefly noting the key quantitative improvements over prior methods once the experimental section is expanded.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback. We address the major comments point by point below, indicating where revisions will be made to the manuscript.","responses":[{"response":"We agree that the abstract would benefit from quantitative support for the SOTA claim. In the revised manuscript, we will update the abstract to include specific performance metrics (e.g., accuracies on cross-subject and cross-session tasks for SEED and SEED-IV), baseline comparisons, and references to the ablation studies and statistical tests detailed in the experimental section.","revision_made":"yes","referee_comment":"[Abstract] Abstract: The state-of-the-art performance claim is asserted without any quantitative results, baseline comparisons, ablation studies, error bars, or statistical significance tests, which is load-bearing for the central empirical claim and prevents verification of the magnitude of improvement or the contribution of the confidence-aware screening and multi-level adaptation components."},{"response":"SEED and SEED-IV are the standard benchmarks for cross-subject and cross-session multimodal emotion recognition, and our protocols directly target the distribution shifts in these datasets. We will add sensitivity analysis on the screening threshold and per-fold variance reporting in the revision to strengthen the robustness evidence. An external hold-out corpus is not part of the current evaluation.","revision_made":"partial","referee_comment":"[Experimental validation] Experimental validation (as summarized in the abstract): The generalization premise that the confidence-aware screening mechanism and multi-level marginal/conditional alignment reliably reduce distribution shifts without introducing selection bias or overfitting is tested only on the SEED and SEED-IV datasets; no external hold-out corpus, no sensitivity analysis on the screening threshold, and no per-fold variance reporting are described, leaving the robustness claim under-supported."}],"tokens_in":1443,"tokens_out":398,"duration_ms":34804,"standing_objections":["Validation on an external hold-out corpus beyond the SEED and SEED-IV benchmarks"]},"desk_editor":{"model":"grok-4.3","letter":"The main point is that this paper gives a concrete engineering pipeline for handling distribution shifts in multimodal physiological emotion recognition. It fuses EEG and eye-tracking via transformer encoders plus cross-attention, adds a confidence-aware screen that splits target samples by reliability and routes them to alignment or distillation, and runs joint marginal-plus-conditional adaptation at both local modality and global fusion levels.\n\nThe combination is new as a single framework even if the pieces draw from prior domain-adaptation and multimodal work. Releasing the code helps, and the problem it targets—individual and session variability in affective computing—is real and worth addressing.\n\nThe soft spots are exactly where the stress-test note flags them. The abstract states SOTA results on SEED and SEED-IV for cross-subject and cross-session tasks but supplies no numbers, baselines, ablations, or variance measures, so the performance claim cannot be checked. The screening step risks introducing selection bias, and everything is tested only on the same two datasets with no external hold-out or threshold sensitivity runs. Without those details the generalization premise stays unproven.\n\nThis is for people who build practical systems in affective computing or physiological signal domain adaptation. A reader who wants a worked example of confidence-based routing plus multi-granularity alignment could extract useful pieces. It deserves a serious referee because the architecture is fully specified and the code is public; the experiments simply need to be examined in the full manuscript.","headline":"UF-AMA stitches together transformer fusion, confidence screening, and multi-level adaptation into a usable pipeline for cross-subject EEG emotion work, but the SOTA claim sits on thin evidence from the abstract alone.","tokens_in":2418,"tokens_out":372,"would_cite":false,"duration_ms":22633,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"An adaptive alignment framework for brain and eye signals enables emotion recognition that generalizes across people and sessions.","keywords":["emotion recognition","multimodal fusion","domain adaptation","EEG signals","eye-tracking","cross-subject","cross-session","physiological signals"],"falsifier":"Running the framework on a fresh collection of subjects or sessions outside the SEED datasets and finding that accuracy does not exceed prior methods, or that the confidence scores fail to predict which modality branches perform well, would undermine the central claim.","tokens_in":2653,"feed_emoji":"🧠","tokens_out":714,"duration_ms":30417,"temperature":0.7,"pith_summary":"The paper develops a unified framework that fuses EEG and eye-tracking data through transformer encoders and cross-attention modules to create integrated features. It adds a screening step that evaluates each modality's reliability on new samples and applies different alignment strategies accordingly, followed by multi-level domain adaptation that matches both marginal and conditional distributions at local and global scales. This combination targets the distribution shifts that arise from individual differences and session variations, which currently limit how well models transfer. A reader would care if the approach succeeds because it could make physiological emotion detection practical without extensive per-user data collection. The reported results on the SEED and SEED-IV datasets show improved accuracy in the targeted cross-domain settings.","feed_headline":"New method aligns brain and eye data to recognize emotions across people","feed_subtitle":"Fusing signals with transformers, screening reliable modalities by confidence, and adapting distributions at multiple levels reaches top acc","key_machinery":"The adaptive multimodal alignment process that combines confidence-aware screening of modality reliability with multi-level optimization of marginal and conditional distributions on fused features.","core_discovery":"The framework constructs a cross-modal feature fusion network with Transformer encoders and multi-head cross-attention modules for deep integration of EEG and eye-tracking signals, introduces a confidence-aware screening mechanism that partitions target samples by predictive reliability and applies global consistency alignment plus cross-modal distillation accordingly, and proposes a multi-level domain adaptation framework that jointly optimizes marginal and conditional distributions of both modality-specific and global fusion features, thereby reducing cross-domain shifts at multiple granularities and achieving state-of-the-art performance on SEED and SEED-IV datasets in cross-subject and c","pith_inferences":["The screening step could be examined on other signal types such as heart-rate variability to test whether modality reliability estimation transfers.","If the multi-level adaptation proves decisive, similar hierarchical matching might reduce calibration effort in related physiological classification problems.","Practical systems built on this pattern could lower the data requirements for deploying emotion-aware interfaces in everyday settings."],"forward_implications":["Deep fusion of EEG and eye-tracking produces richer representations than single-modality approaches for emotion tasks.","Dynamic screening by confidence allows the model to use global alignment only where sample quality supports it and distillation where one modality is weaker.","Joint marginal and conditional alignment at both local and global levels reduces shifts more thoroughly than single-level methods.","The resulting model supports generalization in both cross-subject and cross-session scenarios without separate retraining."],"fun_headline_variants":["UF-AMA aligns EEG and eye data across subjects and sessions","Adaptive multimodal alignment for cross-domain EEG emotion recognition","Transformer cross-attention integrates EEG and eye signals across domains","UF-AMA reduces cross-domain shifts using adaptive multimodal alignment"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The screening mechanism will correctly identify reliable modalities and the multi-level adaptation will reduce distribution shifts without creating new biases or overfitting when applied to subjects and sessions not seen during training.","fun_headline_variants_meta":{"raw":{"variants":["UF-AMA aligns EEG and eye data across subjects and sessions","Adaptive multimodal alignment for cross-domain EEG emotion recognition","Transformer cross-attention integrates EEG and eye signals across domains","UF-AMA reduces cross-domain shifts using adaptive multimodal alignment"]},"model":"grok-4.3","cost_usd":0.009676,"raw_usage":{"total_tokens":4349,"prompt_tokens":742,"num_sources_used":0,"completion_tokens":64,"cost_in_usd_ticks":96762000,"prompt_tokens_details":{"text_tokens":742,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3543,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":742,"tokens_out":64,"duration_ms":25490,"temperature":1.0,"reasoning_tokens":3543,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-28T21:00:24.537846+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Running the framework on a fresh collection of subjects or sessions outside the SEED datasets and finding that accuracy does not exceed prior methods, or that the confidence scores fail to predict which modality branches perform well, would undermine the central claim.","supporting_citations":[],"review_version":1}