{"id":"8dbb2b31-2a8b-45b6-af84-1439521d33b3","arxiv_id":"2507.12889","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"Ordered sequences of gaze fixations mapped to semantic objects, encoded by the new SIO representation and EmoGazeNet, are reported to recognize six emotions with accuracy close to EEG-based methods on self-collected data.","lead":"The paper presents a camera-only system that tracks eye movements and maps them to objects in the environment, then uses these gaze-object sequences to recognize six emotion states. It matters because, if the results hold, ordinary HD cameras could become low-cost, continuous, user-unaware emotion monitors in classrooms, vehicles, and public spaces.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline near-EEG accuracy (88.74%) comes from eye-tracker gaze inputs, not the proposed camera-based gaze pipeline—the camera system is only validated at ~80-84% accuracy on separate real-scene tests, so the central claim is unsupported.","rationale":"The reader's CONDITIONAL verdict focused on label validity, but the more immediate problem is that the single headline number—the basis for the 'close to EEG' claim—was produced with eye-tracker gaze, not the camera-based gaze the paper introduces. This is not a dispute about the value of gaze-semantic modeling; it is a mismatch between the claimed system and the evaluated system. The camera-based pipeline is validated only on smaller real-scene datasets with substantially lower accuracy. The paper's own component evaluation (Supplementary Figure 11) shows gaze-point prediction method matters, and the eye-tracker comparison (Figure 7-b) shows a 3.5-point accuracy gap between tracker and camera gaze. Since the central claim is 'camera-based ... close to EEG,' the absence of an end-to-end camera evaluation leaves the claim unsupported as stated. The label-validity concern remains relevant for interpreting any emotion-recognition result in this paradigm, but it does not need to be resolved to see that the camera system's performance was not measured on the headline benchmark. The concrete check would settle whether the camera pipeline can approach EEG-level accuracy; until that is shown, the paper should be accepted only conditionally, with the claim re-scoped to 'EmoGazeNet with eye-tracker gaze' or re-evaluated with camera gaze.","tokens_in":31679,"tokens_out":6378,"duration_ms":66645,"concrete_test":"Re-run the EmoGaze2D-50 and EmoGaze360-1K benchmark replacing the eye-tracker-derived scanpaths with scanpaths generated by the proposed eight-camera gaze estimation pipeline (Methods: Camera-based gaze tracking method; Algorithm 1-2), keeping the EmoGazeNet architecture and training procedure identical. If overall accuracy/cawF1 on these datasets drops from ~88.74% to the ~72-84% range observed on Real360/field experiments, then the headline claim is unsupported and must be re-scoped. Independent access to the camera-gaze data and code (or a public release) would make this check reproducible.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The headline result (88.74% overall, close to EEG's 88.94%) is computed on the EmoGaze2D-50 dataset in the section 'Performance comparison between proposed deep model EmoGazeNet with existing emotion recognition methods on 2D screen dataset.' The supplementary Methods state that EmoGaze2D-50 'shares the same collection setup as EmoGaze360-1K,' and EmoGaze360-1K gaze data were collected with 'a standard eye tracker set up to record the data' using the WinDB HMD-free approach. Therefore, the gaze scanpaths feeding EmoGazeNet in this benchmark come from a precise commercial eye tracker, not from the eight-HD-camera, user-unaware gaze estimation pipeline that is the paper's claimed contribution. The camera pipeline is validated separately only on Real360 and the two field experiments, where accuracy is 80.22% (Real360), 81.45% (campus), and 83.85% (driving simulator)—all well below the 88.74% headline. The paper never reports end-to-end performance of the camera-based gaze pipeline on EmoGaze2D-50 or EmoGaze360-1K. Consequently, the central claim that a camera-only system matches EEG-level emotion recognition is not actually tested; the 88.74% result reflects a system with laboratory-grade eye-tracking input. This is an internal mismatch between the claim and the measurement, not a matter of external consensus.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a camera-based, user-unaware emotion recognition pipeline that estimates gaze from multi-view HD camera images of the eye and head, maps gaze onto semantic objects in a reconstructed panoramic environment, and feeds the resulting scanpath, represented as Semantic Interactive Orders (SIO), into a GAN-based model called EmoGazeNet. The authors introduce three datasets (EmoGaze2D-50, EmoGaze360-1K, Real360), a new metric cawF1, an online calibration method, and a third-person panoramic modeling approach. They report that the context gaze-based method reaches 88.74% overall accuracy on the 2D screen dataset, close to an EEG-based method (88.94%), and that it outperforms gaze-only and facial-expression baselines in real-scene field experiments. The central claim is that ordinary HD cameras can achieve near-EEG-level emotion recognition without wearable sensors, in a user-unaware manner.","tokens_in":32027,"tokens_out":3728,"duration_ms":41361,"significance":"If substantiated, the claim that a camera-only gaze pipeline can match EEG-based emotion recognition would be an important practical result for affective computing, enabling unobtrusive, scalable monitoring in education, driving, and public safety. The paper contains several commendable elements: it attempts real-world deployment in campus and driving-simulator settings, includes longitudinal stability experiments over 17 days, provides component ablations for EmoGazeNet, and proposes a richer representation of gaze as a sequence of semantically meaningful fixations rather than raw coordinates. The scanpath visualizations by gender and emotion are also a useful descriptive contribution. However, the headline quantitative claim is not actually tested end-to-end, and the evaluation metrics and ground-truth labeling have self-referential elements, so the significance of the reported numbers is currently uncertain.","major_comments":[{"comment":"The headline result of 88.74% overall accuracy, which is the basis for the claim of near-EEG performance, is computed on EmoGaze2D-50 using gaze data collected with a standard eye tracker, not with the proposed eight-HD-camera gaze pipeline. The supplementary states that EmoGaze2D-50 shares the collection setup of EmoGaze360-1K, and that EmoGaze360-1K used the WinDB HMD-free approach with a standard eye tracker. The camera-based gaze acquisition method is validated separately on Real360 and the two field experiments, where accuracy ranges from 80.22% to 83.85%. The paper never reports end-to-end performance of the camera-based gaze pipeline on EmoGaze2D-50 or EmoGaze360-1K. Consequently, the central claim that a camera-only system matches EEG-level emotion recognition is not supported by the reported experiments; the 88.74% figure reflects a system with laboratory-grade eye-tracking input. This is a load-bearing mismatch between the claim and the measurement and must be addressed, either by reporting end-to-end camera results on the benchmark datasets or by substantially revising the claim.","section":"Results, 'Performance comparison between proposed deep model EmoGazeNet with existing emotion recognition methods on…"},{"comment":"The cawF1 metric is proposed in this paper and is used to rank all methods on EmoGaze360-1K and in the field experiments, but it is not computable in a meaningful way for baselines that do not produce gaze or fixation predictions. Facial-expression and EEG baselines have no fixation-context consistency term, so it is unclear whether FCC is set to a constant, omitted, or computed from some proxy; without this definition, the cawF1 comparisons in Figure 4-b and Figures 6-d/6-e are not interpretable. Additionally, the weights alpha and beta in Eq. (5) are never reported. The paper should state the values of alpha and beta, describe how cawF1 is applied to each baseline, and provide a version of the comparison using standard metrics only, or justify why cawF1 is appropriate for all methods.","section":"Methods, 'Proposed evaluation metric', Eqs. (4)-(5)"},{"comment":"The ground-truth emotion labels are based on instructed emotion induction or self-simulation, and the paper does not verify that participants actually experienced the target emotion. The 'Collection setting' section states that participants underwent emotion induction through video and image stimuli, while the campus field experiment explicitly asked participants to 'simulate six different emotional states'. No manipulation check, self-report rating, or physiological verification is reported. If participants only acted the emotion, the system may be learning to classify posed gaze behavior rather than internal emotional states, which would invalidate the 'mind reading' framing and the claim of recognizing 'real emotions' in the abstract. The authors should add a manipulation check or clearly restrict their claims to acted/induced emotional behavior.","section":"Methods, 'Collection setting'; 'Field experiment to evaluate the practical application of our eye gaze collection…"},{"comment":"The baselines ACTNN, Toisoul, and CCER are mentioned by name but no implementation details are given for how they were adapted to the new datasets, what input modalities they received, whether they were retrained, or how their cawF1 scores were obtained. The statistical tests are also under-specified: the two-sample t-test and one-way ANOVA p-values are reported without describing the number of subjects, the folds, whether the tests are paired, or the exact comparisons being made. Without this information, the claimed statistical superiority over the EEG-based method on deceptive emotions cannot be assessed. Please provide a clear evaluation protocol, including dataset splits, subject independence, and baseline configurations.","section":"Results, 'Performance comparison between proposed deep model EmoGazeNet with existing emotion recognition methods on…"}],"minor_comments":[{"comment":"There is a typo: 'Furthre' should be 'Further'. Similar typographical issues appear elsewhere, including 'ANOV A' instead of 'ANOVA' and 'fiaxtion' instead of 'fixation' in the supplementary architecture description.","section":"Methods, 'Camera-based gaze tracking method'"},{"comment":"The six indicators (emotion sensitivity, stability, etc.) are described as quantified on a 1-10 scale, but the method for assigning these ratings is not given; it is unclear whether they come from the model, from human annotators, or from the authors. Please specify the source of these ratings and any inter-rater reliability if humans were involved.","section":"Figure 5-b and Supplementary 'Detailed explanations of six indicators...'"},{"comment":"References 52 and 66 are duplicates of the same paper (ShanghaiTechGaze); one of them should be removed and the citation list renumbered.","section":"References"},{"comment":"The calibration method depends on several thresholds and windows (e.g., head movement yaw/pitch thresholds, the 200-300 ms online fine-tuning window) that are described qualitatively but never given concrete values. Please provide the actual thresholds used in the experiments.","section":"Methods, 'Online personalized calibration'"},{"comment":"The dataset construction section says EmoGaze360-1K contains 1,000 panoramic images but later mentions '500 emotion-inducing images and 50 emotion-inducing videos... resulting in a total of 2,500 images and 250 videos for each emotional state'; the relationship between these stimulus pools and the 1,000 annotated panoramas should be clarified.","section":"Supplementary, 'EmoGaze360-1K and EmoGaze2D-50 datasets construction'"},{"comment":"The long-term stability experiment is reported only as cawF1 fluctuations around 71-73%, but no statistical test or confidence interval is provided; adding a trend test or at least a variance estimate would strengthen the claim of stability.","section":"Results, 'Robustness evaluation'"}],"recommendation":"major_revision","confidential_remarks":"The core problem is that the paper's most impressive number (88.74%) comes from eye-tracker input, not from the camera pipeline that is the claimed contribution. This is fixable in principle by re-framing the claims or by adding end-to-end experiments, but as written it is a serious internal mismatch. The cawF1 metric and the simulated-emotion labels are additional load-bearing concerns. I would not reject on novelty because the SIO representation and the multi-camera panoramic gaze estimation are interesting ideas, but the evaluation needs substantial rework before the central claim can be accepted. I would also suggest that the authors consult the journal's policy on dataset availability and code release, since the datasets are author-built and no code is provided."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere is my honest read. The paper's useful core is real: SIO (ordering gaze fixations into object patches over time), EmoGazeNet, and three new datasets. The component ablations, scanpath prediction comparison, and field experiments are more than most submissions include. If the gaze-environment semantic pipeline worked as claimed end to end, it would be a meaningful advance for affective computing.\n\nThe problem is that the headline claim is not actually tested. The 88.74% accuracy on EmoGaze2D-50 — the number that is 'close to EEG' — was produced using eye-tracker-collected scanpaths from the WinDB setup, not from the eight-HD-camera user-unaware gaze estimation pipeline that the paper sells. The camera pipeline is validated separately on Real360 and the field tests, where accuracy lands at 80–84%. The paper never reports end-to-end camera-based gaze estimation on EmoGaze2D-50 or EmoGaze360-1K. So the sentence 'camera-based emotion recognition can match EEG' is currently unsupported. What the data support is gaze-based emotion recognition with laboratory-grade eye tracking, plus a separate camera-based gaze estimator that is less accurate and is never plugged into the full system.\n\nThere are secondary issues. cawF1 is proposed and used to rank all methods, but alpha and beta in Eq. 5 are never reported, and gaze-free baselines cannot be assigned a meaningful fixation-context consistency; the comparison under cawF1 is therefore ill-defined for those baselines. Emotion ground truth comes from instructed induction or self-simulation, with no check that participants actually experienced the labeled emotion; the model may be classifying posed gaze behavior. Data and code are not released, so the result is not independently reproducible. The authors do acknowledge limitations (lighting, occlusion, head movement, dataset diversity), but they do not acknowledge the eye-tracker/camera gap.\n\nAll of this is fixable. Re-run the headline experiment end to end through the camera pipeline, or reframe the claim. Report the metric weights, release the datasets and code, and validate on an external dataset. Then there is a credible paper here.\n\nRecommendation: yes, send this to peer review. It deserves serious referee time, but the verdict should hinge on whether the central claim is re-measured. As it stands, I would not cite the headline claim.","headline":"A substantial pipeline and new datasets are undermined by a central claim that is not actually tested: the near-EEG accuracy comes from eye-tracker input, not from the camera system.","tokens_in":32635,"tokens_out":2782,"would_cite":false,"duration_ms":28814,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that the order in which a person's gaze visits meaningful objects in a scene, captured with ordinary cameras, reveals emotional state at roughly EEG-level accuracy (88.74% vs 88.94%) and beats EEG on concealed emotions.","keywords":["gaze-based emotion recognition","environmental context","Semantic Interactive Orders","user-unaware monitoring","gaze scanpath","deceptive emotion detection","camera-based gaze tracking","emotion-attention interaction"],"falsifier":"Record gaze and scene data from people experiencing genuine, un-instructed emotions — for example during real events where the felt emotion is confirmed at the moment it happens by independent self-report or physiological recording — and run EmoGazeNet trained only on instructed emotions; if accuracy on genuinely felt emotions falls far below the reported 88.74% (or toward chance), the claim that the method reads internal states rather than acted behavior would be refuted.","tokens_in":31403,"feed_emoji":"👁️","tokens_out":16312,"duration_ms":154701,"temperature":0.7,"pith_summary":"This paper tries to establish that a person's emotional state can be read unobtrusively from the way their gaze moves through the meaningful contents of a scene. Instead of treating gaze as a static fixation point, it models the sequence of semantic objects a person looks at and in what order — the Semantic Interactive Orders — and claims this sequence reveals which of six basic emotions the person is in. On its own collected data the method reaches 88.74% overall accuracy, close to the 88.94% of an EEG-based system, and it does better than EEG (89.82% versus 87.13%) when people try to conceal their emotion. Because everything is captured by standard HD cameras with no wearable devices and no active user task, the authors argue, continuous real-time emotion monitoring becomes practical in classrooms, cockpits, and vehicles. The broader claim is that emotions are not purely internal physiological states but products of human-environment interaction, and that this interaction can substitute for physiological sensing.","feed_headline":"Hits 88.7% emotion accuracy from gaze plus scene","feed_subtitle":"A camera-only system matches EEG accuracy without wearables and beats it on concealed emotion.","key_machinery":"The load-bearing object is the Semantic Interactive Orders (SIO) representation: gaze fixation points predicted from eye appearance and head movement are mapped onto object-level regions of a panoramic scene, and the regions are arranged in viewing order, so each emotion is encoded as an ordered patch sequence over the scene's semantics. That sequence drives EmoGazeNet, a generative-adversarial classifier whose generator applies spatial-temporal positional encoding (object position plus viewing time) inside a Transformer encoder, while its discriminator enforces separation between ordinary appearance features and scanpath-derived temporal-spatial features through an adversarial reverse-suppression loss and an auxiliary scanpath-prediction task. Two supporting mechanisms carry the system: a third-person multi-camera pipeline (eight HD cameras, super-resolution, 3D reconstruction of eye appearances, and online personalized calibration fusing saliency-based 'objective' fixations with user-specific 'subjective' fixations) that produces gaze trajectories without wearables, and the cawF1 evaluation metric ($\\mathrm{cawF1}=\\sum_i \\mathrm{FCC}_i\\,\\mathrm{bF1}_i/\\sum_i \\mathrm{FCC}_i$, with FCC a fixation-context consistency score) that holds the model to predicting where people look, not only the emotion label.","core_discovery":"The paper's core claim is that emotion is legible in how visual attention travels through a scene's semantics over time, not just where the gaze rests. Concretely, predicted gaze fixation points are mapped onto object-level regions of a panoramic reconstruction of the environment, these regions are ordered by viewing sequence, and the ordered patch sequence is fed into EmoGazeNet, a GAN-based classifier with a Transformer encoder, scanpath-guided and auxiliary classification branches, and adversarial reverse suppression that keeps appearance features and scanpath features distinct. On the 2D screen dataset the method reports 89.82% accuracy on deceptive emotions and 87.65% on real emotions (88.74% overall), against 88.94% overall for the EEG baseline; on the 360-degree dataset its average cawF1 of 78.14% trails EEG by 1.26% while surpassing facial and gaze-only baselines, and in campus and driving-simulator field experiments it outperforms both facial-expression and physiological-signal methods (81.45% and 83.85% accuracy). The authors interpret these results as showing that gaze-environment interaction dynamics can replace physiological sensors, and that this 'user-unaware' monitoring — no wearables, no active participation — brings continuous emotion recognition into real-world settings.","pith_inferences":["Because all labels come from instructed induction or self-simulation, the reported accuracy may measure recognition of posed gaze behavior; a test with naturally occurring emotions and independently confirmed labels is the untaken step that would decide whether the 'mind reading' framing is justified.","The SIO hypothesis implies a falsifiable regularity: the same person in the same scene should produce statistically distinct object-visitation orders under different emotions, and if scanpath orders do not separate by emotion, the information channel the model relies on would be empty.","The current pipeline requires a pre-modeled panoramic reconstruction and fixed camera positions, so moving to unbounded or moving environments would need real-time semantic segmentation and view synthesis — a natural next step the paper only gestures at.","Because the system is designed to be 'user-unaware', deployment raises a consent tension: continuous emotion reading without the person's knowledge would require institutional oversight and opt-in policies before classroom or cockpit use becomes ethical."],"forward_implications":["If the accuracy numbers hold, continuous emotion monitoring becomes deployable with off-the-shelf HD cameras at EEG-level accuracy (88.74% versus 88.94%) in classrooms, cockpits, and driver monitoring, without wearable sensors or active user participation.","Concealed emotion, which defeats facial-expression systems (43.91% deceptive accuracy), is recognized at 89.82% accuracy — above the EEG baseline of 87.13% — so gaze-environment dynamics leak information even when facial behavior is controlled.","Field results in a campus square and a driving simulator (81.45% and 83.85% accuracy, with significant ANOVA contrasts against the facial and physiological baselines) support generalization beyond the laboratory, and 17-day monitoring stays stable near 71.5%–72% cawF1.","Adopting the proposed cawF1 metric raises the field's evaluation bar: a model must predict both the emotion and the attended regions of the scene, and under that stricter standard the complete system scores 72.22% on real-scene data."],"supporting_citations":[{"why":"Supplies the EEG and eye movement data template that the paper's 2D screen dataset is built to extend, establishing the multimodal baseline structure.","marker":"[47]"},{"why":"The EEG-based emotion recognition baseline whose 88.94% overall accuracy sets the bar the proposed method is compared against.","marker":"[48]"},{"why":"The facial-expression baseline that scores 94.36% on real emotions but collapses to 43.91% on deceptive ones, establishing the need for non-facial signals.","marker":"[49]"},{"why":"The gaze-only baseline at 74.85% overall accuracy that the proposed method claims to surpass by roughly 13%.","marker":"[50]"},{"why":"Grounds the six-category emotion label set (Angry, Disgust, Fear, Happy, Sad, Surprised) used across all datasets.","marker":"[51]"},{"why":"The gaze prediction model whose limited head-angle data is 3D-reconstructed into multi-angle eye appearance to retrain the camera-based gaze estimator.","marker":"[52]"},{"why":"Supports the claim that attention in the first 200–300 ms after a scene transition is driven by objective saliency, justifying the online calibration's fine-tuning window.","marker":"[53]"},{"why":"Supplies the head-mounted-display-free fixation collection methodology used to build the 360-degree dataset EmoGaze360-1K.","marker":"[56]"}],"fun_headline_variants":["Gaze paths in scenes reveal hidden emotions via camera","Camera reads emotions from how gaze moves through a scene","No wearables: emotion detection from gaze and environment","Gaze plus scene context decodes emotion with 89% accuracy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"All emotion labels come from instructed induction or self-simulation — participants watched emotion-triggering videos or were asked to simulate an emotional state — and the paper never verifies that anyone actually felt the labeled emotion, so the system may be learning posed gaze patterns rather than internal states.","fun_headline_variants_meta":{"raw":{"variants":["Gaze paths in scenes reveal hidden emotions via camera","Camera reads emotions from how gaze moves through a scene","No wearables: emotion detection from gaze and environment","Gaze plus scene context decodes emotion with 89% accuracy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000565,"raw_usage":{"total_tokens":2736,"prompt_tokens":1058,"completion_tokens":1678,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":674,"completion_tokens_details":{"reasoning_tokens":1612}},"tokens_in":674,"tokens_out":1678,"duration_ms":14082,"temperature":1.0,"reasoning_tokens":1612,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T16:35:45.634805+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Record gaze and scene data from people experiencing genuine, un-instructed emotions — for example during real events where the felt emotion is confirmed at the moment it happens by independent self-report or physiological recording — and run EmoGazeNet trained only on instructed emotions; if accuracy on genuinely felt emotions falls far below the reported 88.74% (or toward chance), the claim that the method reads internal states rather than acted behavior would be refuted.","supporting_citations":[{"cited_title":"& Lu, B.-L","cited_arxiv_id":null,"evidence_quote":"Supplies the EEG and eye movement data template that the paper's 2D screen dataset is built to extend, establishing the multimodal baseline structure."},{"cited_title":"& Chen, W","cited_arxiv_id":null,"evidence_quote":"The EEG-based emotion recognition baseline whose 88.94% overall accuracy sets the bar the proposed method is compared against."},{"cited_title":"& Pan- tic, M","cited_arxiv_id":null,"evidence_quote":"The facial-expression baseline that scores 94.36% on real emotions but collapses to 43.91% on deceptive ones, establishing the need for non-facial signals."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The gaze-only baseline at 74.85% overall accuracy that the proposed method claims to surpass by roughly 13%."},{"cited_title":"& Friesen, W","cited_arxiv_id":null,"evidence_quote":"Grounds the six-category emotion label set (Angry, Disgust, Fear, Happy, Sad, Surprised) used across all datasets."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The gaze prediction model whose limited head-angle data is 3D-reconstructed into multi-angle eye appearance to retrain the camera-based gaze estimator."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports the claim that attention in the first 200–300 ms after a scene transition is driven by objective saliency, justifying the online calibration's fine-tuning window."},{"cited_title":"& Fan, D.-P","cited_arxiv_id":null,"evidence_quote":"Supplies the head-mounted-display-free fixation collection methodology used to build the 360-degree dataset EmoGaze360-1K."}],"review_version":1}