{"id":"1e169b5e-8d52-4fca-aaf8-9fb52e07f4af","arxiv_id":"2606.31211","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"AA is a multi-view multimodal dataset for screen-based gaze estimation featuring eight screen-mounted and two side-view cameras with precise gaze targets and subject-independent splits.","lead":"The paper presents AA, a new dataset with synchronized multi-view facial images from eight screen-mounted cameras plus two side views, paired with precise gaze targets on screens under controlled conditions. Smart generalists might read it because better multi-view data could help build eye-tracking systems that work more reliably when faces move or are partially blocked.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.3","headline":"Controlled fixation may not reflect real screen-task variability","rationale":"The reader's weakest assumption directly identifies the same transferability gap. Because the paper is a dataset release without reported cross-domain experiments, the concern is internal to the claim rather than a disagreement with external consensus. No machine-checked proof or parameter-free derivation exists to override it.","tokens_in":1604,"tokens_out":286,"duration_ms":15988,"concrete_test":"Train the same gaze estimator on AA and on a single-view dataset of comparable size; evaluate both on an external real-world screen-gaze test set containing free head motion; if the multi-view model shows no statistically significant gain over the single-view baseline on the external set, the robustness claim does not hold.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on the assertion that eight screen-mounted + two side-view synchronized captures under controlled fixation produce observations and targets that support robust modeling under viewpoint variation and occlusion. This requires that the lab protocol (fixed head pose, instructed fixation points, uniform lighting) generates the same distribution of head pose, eye appearance, and partial occlusions encountered in unconstrained screen use. No evidence is supplied that the chosen camera placements or fixation protocol reproduce the statistics of natural head motion or spontaneous gaze shifts; if they do not, the multi-view advantage remains unproven for the stated use case.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript presents the AA dataset for screen-based gaze estimation. It captures synchronized facial observations from eight fixed screen-mounted cameras and two side-view cameras, paired with precise screen-space gaze targets under controlled fixation conditions. Each sample includes multi-view face observations and structured facial region crops for multimodal learning from global and local cues. The dataset provides subject-independent evaluation splits and a standardized processing pipeline, positioned as enabling more robust modeling under viewpoint variation and occlusion compared to existing single-view gaze datasets.","tokens_in":1685,"tokens_out":309,"duration_ms":22689,"significance":"If the capture protocol produces observations whose statistics of head pose, eye appearance, and partial occlusions match those of unconstrained screen use, the multi-view coverage from screen-mounted and side perspectives would constitute a useful resource for developing gaze estimators that are more robust to viewpoint changes and occlusions. The inclusion of subject-independent splits and a reproducible pipeline is a positive feature for community adoption.","major_comments":[{"comment":"Abstract (dataset capture description): the central claim that the dataset 'enables more robust modeling under viewpoint variation and occlusion' rests on the assumption that the controlled fixation conditions and fixed camera placements reproduce the distribution of natural head motion, spontaneous gaze shifts, and occlusions encountered in real screen-based tasks; no supporting statistics, comparisons to unconstrained data, or validation of this match are supplied.","section":"Abstract (dataset capture description)"}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback on the AA dataset manuscript. The recommendation for major revision is noted, and we address the single major comment below regarding the abstract's claims about robustness under viewpoint variation and occlusion.","responses":[{"response":"We agree that the manuscript provides no supporting statistics, comparisons to unconstrained screen-use data, or explicit validation that the controlled fixation protocol and fixed camera placements reproduce the distributions of natural head motion, spontaneous gaze shifts, or occlusions. The capture design prioritizes precise screen-space gaze targets and synchronized multi-view observations under controlled conditions to enable high-quality labeled data. The claim in the abstract is prospective, based on the availability of multi-view (screen-mounted and side-view) observations that can be used to train and evaluate models handling viewpoint changes and partial occlusions. We will revise the abstract to remove the implication of distributional match and instead state that the multi-view coverage supports development of models robust to such variations when applied to the provided data. No new empirical validation or external comparisons will be added, as they fall outside the scope of a dataset release paper.","revision_made":"yes","referee_comment":"the central claim that the dataset 'enables more robust modeling under viewpoint variation and occlusion' rests on the assumption that the controlled fixation conditions and fixed camera placements reproduce the distribution of natural head motion, spontaneous gaze shifts, and occlusions encountered in real screen-based tasks; no supporting statistics, comparisons to unconstrained data, or validation of this match are supplied."}],"tokens_in":1189,"tokens_out":327,"duration_ms":15984,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"This paper releases the AA dataset: synchronized captures from eight fixed screen-mounted cameras and two side-view ones, plus structured facial crops and precise screen-space gaze targets collected under instructed fixation. The main addition is the multi-view coverage from both screen and side angles, with subject-independent splits and a standardized pipeline. That combination is new relative to the single-view datasets referenced.\n\nThe work is straightforward about what it provides and includes the processing steps needed for others to use the data. For a dataset paper that is a reasonable baseline.\n\nThe soft spot is the reliance on controlled fixation conditions. The abstract gives no evidence that the head poses, eye appearances, or partial occlusions match the distribution seen in unconstrained screen use. If the lab protocol produces a narrower range of natural motion and spontaneous shifts, the multi-view advantage for robustness under variation stays untested. That is the central assumption and it is not backed by any reported statistics or comparisons.\n\nThis is for CV and HCI groups that need multi-view gaze data and are willing to work within the lab constraints. Readers who want datasets with documented ecological validity will find less here.\n\nIt is worth sending to peer review. Dataset releases can be useful even when the paper is mostly descriptive, provided the data itself is released and the collection details can be scrutinized.","headline":"New multi-view gaze dataset with eight screen cameras plus side views, but the controlled fixation protocol leaves the real-world robustness claim unproven.","tokens_in":2161,"tokens_out":336,"would_cite":false,"duration_ms":12442,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"The AA dataset supplies synchronized multi-view facial images from eight screen-mounted cameras plus two side views, paired with precise screen gaze targets under controlled fixations, to support models robust to viewpoint changes and occlu","keywords":["multi-view dataset","gaze estimation","screen-based interaction","multimodal learning","facial observations","viewpoint variation","occlusion robustness","subject-independent splits"],"falsifier":"A controlled comparison in which models trained only on AA show no accuracy gain over single-view models when both are evaluated on the same set of real-world screen recordings that contain natural head motion and partial occlusions.","tokens_in":2508,"feed_emoji":"👁","tokens_out":684,"duration_ms":24392,"temperature":0.7,"pith_summary":"The paper introduces the AA dataset to overcome the limitations of existing single-view collections for screen-based gaze estimation. It records simultaneous face observations across ten cameras while subjects fixate on known screen locations, then supplies both full-face frames and structured region crops for each sample. This multi-view coverage from screen and side angles is intended to let models learn features that remain stable when the viewpoint shifts or parts of the face are blocked. The release also includes subject-independent splits and a fixed processing pipeline so that different research groups can run comparable experiments.","feed_headline":"Ten-camera dataset pairs multi-view faces with exact screen gaze targets","feed_subtitle":"Eight screen-mounted views plus two side angles let models train on viewpoint changes and partial blocks that single-camera sets miss.","key_machinery":"The synchronized ten-camera array (eight screen-mounted, two side-view) that records facial observations together with exact screen-space gaze targets and structured region crops.","core_discovery":"The AA dataset captures synchronized facial observations from eight fixed screen-mounted cameras and two additional side-view cameras, paired with precise screen-space gaze targets collected under controlled fixation conditions. Each sample contains multi-view face observations together with structured facial region crops, enabling multimodal learning from both global and local visual cues. Unlike existing single-view gaze datasets, AA provides multi-view coverage from both screen-mounted and side-mounted perspectives, enabling more robust modeling under viewpoint variation and occlusion.","pith_inferences":["The same multi-view capture setup could be reused to collect data for dynamic tasks such as smooth pursuit or reading instead of static fixations.","Performance gains on AA might translate to laptop or tablet scenarios where the camera is not perfectly centered on the screen.","The dataset format supports future addition of depth or infrared channels from the same camera positions without changing the annotation protocol."],"forward_implications":["Models can be trained to combine information across multiple simultaneous viewpoints rather than relying on a single camera.","Structured facial crops make it possible to train on both global face appearance and local eye or mouth regions within the same sample.","Subject-independent splits allow direct comparison of methods without leakage from repeated identities.","The fixed processing pipeline removes one source of non-reproducibility when different groups benchmark new gaze estimators."],"fun_headline_variants":["Ten-camera dataset for multi-view screen gaze estimation","Synchronized ten views pair faces with screen gaze targets","Multi-view dataset from eight screen and two side cameras","Dataset enables robust gaze modeling under viewpoint changes"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The controlled fixation conditions and fixed camera placements generate gaze targets and multi-view observations that match the variability found in actual screen-based gaze estimation tasks.","fun_headline_variants_meta":{"raw":{"variants":["Ten-camera dataset for multi-view screen gaze estimation","Synchronized ten views pair faces with screen gaze targets","Multi-view dataset from eight screen and two side cameras","Dataset enables robust gaze modeling under viewpoint changes"]},"model":"grok-4.3","cost_usd":0.006346,"raw_usage":{"total_tokens":2928,"prompt_tokens":564,"num_sources_used":0,"completion_tokens":58,"cost_in_usd_ticks":63462000,"prompt_tokens_details":{"text_tokens":564,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2306,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":564,"tokens_out":58,"duration_ms":21084,"temperature":1.0,"reasoning_tokens":2306,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-01T06:08:36.594740+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A controlled comparison in which models trained only on AA show no accuracy gain over single-view models when both are evaluated on the same set of real-world screen recordings that contain natural head motion and partial occlusions.","supporting_citations":[],"review_version":1}