{"id":"095aa194-32b9-4025-9630-b13af8fb403d","arxiv_id":"2507.15443","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A head-tracking based method quantifies joint attention for four people across two rooms on wall-sized displays, with preliminary evidence that gaze correlates among collaborators.","lead":"This paper describes a method for measuring joint attention between people collaborating across two rooms on large wall-sized displays, using depth cameras that track head direction instead of eye-tracking glasses. The authors report early results from one session, suggesting that collaborators, especially those in the same room, tend to look at the same parts of the screen.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Head-gaze proxy is unvalidated and reported correlations could be inflated by interpolation and task structure; the paper's central empirical insight needs a ground-truth check.","rationale":"The reader's weakest_assumption correctly identifies the head-gaze proxy as the load-bearing premise. I agree that the paper's own Section 1 caveat about depth-camera head tracking accuracy applies directly to the joint-attention metric, and that no validation is provided. My concern also points to two additional, related threats that are not fully captured by the reader's phrasing: interpolation artifacts in Section 2 and task-induced correlation structure in Section 3. However, these do not overturn the paper's modest claim of proposing an approach and reporting preliminary insights; they do mean the empirical correlations should not be interpreted as established evidence about joint attention. The proposed ground-truth comparison would settle whether the proxy is adequate for this setting. Since the reader already arrived at CONDITIONAL with moderate confidence, my read does not move the verdict; the paper should remain conditional pending the validation, but the contribution is not rejected.","tokens_in":3926,"tokens_out":3615,"duration_ms":44589,"concrete_test":"Use one recorded session: for a 3-5 minute segment, obtain ground truth by having an annotator label, from video, the art piece each participant looks at every second (or run one participant with eye-tracking glasses as a reference). Compute the proposed head-gaze-based joint-attention indicator (Section 3) and compare it to the ground-truth joint-attention labels using agreement and Cohen's kappa or ROC. If agreement is below a pre-specified threshold (e.g., kappa < 0.6), the head-gaze proxy cannot support the paper's claims without correction. Additionally, recompute the correlations on the non-interpolated rows only; if Spearman rho drops materially, the interpolation is driving the result.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2 constructs the gaze time series by manually merging skeleton IDs, filtering out non-participants, and linearly interpolating lost tracking. Section 3 then computes Spearman correlations and joint-attention shares directly from these values. The load-bearing premise is that the resulting 'gaze target' from Azure Kinect head tracking reliably indicates where each participant looks on the wall display. The paper itself states in Section 1 that depth-camera head tracking is 'less accurate to analyse gaze' and cites [13] for the claim that it 'still provides a good idea' of attention, but no validation establishes this at room scale or for this task. If head direction deviates from true gaze (peripheral glances, gaze at hands or at the other participant), the metric mislabels joint attention. In addition, the positive same-room correlations in Figure 2 are compatible with a trivial explanation: participants in the same room are physically oriented toward the shared display because of the task, independent of shared attentional focus. Linear interpolation of missing segments can also create smooth shared trends across users, inflating Spearman rho. Thus the reported insight 'horizontal gaze values are indeed correlated' is not yet established as evidence about joint attention.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript proposes a pipeline for measuring joint attention in mixed-presence collaboration around wall-sized displays using head-gaze data from Azure Kinect depth cameras and the psi framework. Section 2 describes merging recordings from two rooms, applying a time offset, manually filtering and merging skeleton IDs, linearly interpolating missing tracking data, and normalizing gaze coordinates. Section 3 computes Spearman correlations between participants' horizontal gaze values and operationalizes joint attention as all-four or three-out-of-four participants having gaze targets within a distance threshold derived from maximum art-piece dimensions, reported as the share of time below threshold over 5-second windows. The empirical content is a single selected session, presented as preliminary.","tokens_in":4125,"tokens_out":2791,"duration_ms":33296,"significance":"The contribution is a practical, unobtrusive evaluation method rather than a fully validated finding. Strengths include a concrete and reusable pipeline, an explicit operational definition of joint attention, and an openly available dataset on Zenodo. The central empirical claim—that same-room participants' horizontal gaze values are correlated—is plausible but not yet established because the head-gaze proxy and the preprocessing choices are not validated. If confirmed, the method could enable room-scale collaboration studies without obtrusive eye trackers; at present, the paper should be read as a proof-of-concept demonstration of the pipeline.","major_comments":[{"comment":"The load-bearing premise that head-gaze direction from Azure Kinect is a valid proxy for actual gaze is not validated. Section 1 cites Stiefelhagen et al. [13] for the claim that head tracking 'still provides a good idea' of attention, but that reference comes from a different setting and does not establish accuracy for four users at room scale on wall-sized displays, where peripheral glances and gaze at hands or at the other participant can decouple head orientation from gaze. Without a ground-truth comparison (e.g., simultaneous eye tracking or manual coding of where people are looking), the joint-attention measure can be systematically distorted; please add such a validation, or present the results as an unvalidated proof of concept rather than as evidence about joint attention.","section":"1, 3"},{"comment":"Linear interpolation is used to fill missing tracking data before computing correlations and joint-attention shares, yet the amount and location of missing data are not reported. Interpolated segments can create smooth common trends across users, which can inflate Spearman's rho and artificially increase the share of time below the distance threshold. Please report per-participant missing-data rates and run a sensitivity analysis that recomputes the metrics on non-interpolated samples only or otherwise flags interpolated segments.","section":"2"},{"comment":"The correlations in Figure 2 are reported without confidence intervals, significance tests, or correction for temporal autocorrelation. With a single session and strongly autocorrelated time series, the nominal number of samples overstates precision, so the statement that 'horizontal gaze values are indeed correlated' is not statistically supported. Provide bootstrap or permutation-based intervals that account for the time-series structure, and show results for all sessions with a documented session-selection rule rather than an unspecified 'chosen session'.","section":"3, Figure 2"},{"comment":"The joint-attention measure depends on several unexamined parameter choices: the distance threshold derived from maximum art-piece dimensions, the 5-second time window, and the time offset used to synchronize the two rooms. Small changes in the threshold will directly change the reported share of joint attention. Please provide a sensitivity analysis for these parameters and justify the window length and the offset determination procedure.","section":"3"}],"minor_comments":[{"comment":"In the sentence 'we computed the the share of time spent below threshold', the definite article 'the' is duplicated.","section":"3"},{"comment":"Figure 1 has no axis labels or legend identifying the four participants, and Figure 2's right panel does not state whether the plotted share refers to all-four or three-out-of-four joint attention; please clarify both figures.","section":"Figures 1 and 2"},{"comment":"The cited validation for head tracking as a gaze proxy is from 1999; given advances in depth cameras and tracking algorithms, please justify the transfer to modern Azure Kinect data or cite a more recent validation.","section":"1, Reference [13]"},{"comment":"The abstract and Section 4 correctly frame the analysis as preliminary, but the Section 3 wording 'indeed correlated' is stronger than the evidence supports; aligning the wording with the exploratory scope would avoid overinterpretation.","section":"Abstract, 3"}],"recommendation":"major_revision","confidential_remarks":"This is a three-page companion/poster paper, so the empirical scope is necessarily limited. The main concern is the gap between the exploratory analysis of one session and the claim that correlations are 'indeed' present. The manuscript would be acceptable if the authors add a validation check for the head-gaze proxy, report missing-data and sensitivity analyses, and soften the claims to match the preliminary scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this is a three-page companion paper that does what it says. It proposes a pipeline for estimating joint attention from head gaze in a four-user, two-room wall-sized display study, and it backs the proposal with a published dataset. The novelty is incremental—pair-focused joint attention and depth-camera tracking are known—but the extension to four users across two rooms is a legitimate step for the mixed-presence CSCW subfield.\n\nWhat the paper does well: the method section is honest. It describes manual merging of skeleton IDs, filtering of non-participants, and linear interpolation of missing tracking, which is more transparency than most short papers offer. The threshold for joint attention is tied to a domain rationale (art piece dimensions) rather than fitted to the outcome. Spearman is a sensible choice. The authors explicitly frame the results as preliminary and note that more sessions are needed.\n\nThe soft spots are real but in proportion. Only one session is analyzed, and the selection is not justified. The correlations come without confidence intervals or significance tests, so the claim that same-room horizontal gaze values 'are indeed correlated' is descriptive at best. The deeper issue is the head-gaze proxy: the paper cites an old study to say head direction gives a good idea of attention, but no validation at room scale for this task. Same-room correlations could partly reflect physical orientation toward the shared display rather than shared attentional focus. Linear interpolation of lost tracking can also create smooth shared trends and inflate Spearman rho. The stress-test note is right about these threats; it is also right that they do not sink the paper, because the authors do not overclaim.\n\nBottom line: for researchers in mixed-presence collaboration who want an unobtrusive alternative to eye tracking, this is a usable starting point with a reproducible dataset. It is not a definitive empirical result. I would send it to peer review rather than desk reject it, and I would ask for ground-truth validation of the head-gaze proxy, multi-session analysis, and at least basic inferential statistics before the method is used to draw general conclusions. I would cite it if I were working in this area.","headline":"Honest short paper with a usable pipeline and a published dataset; the headline correlation result needs a ground-truth check before it supports general conclusions.","tokens_in":4627,"tokens_out":2342,"would_cite":true,"duration_ms":24216,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Head gaze from depth cameras can detect and quantify joint attention in mixed-presence collaboration on wall-sized displays.","keywords":["Wall-sized displays","Mixed-presence collaboration","Joint attention","Head gaze","Depth camera body tracking","Workspace awareness","Spearman correlation","Gaze normalization"],"falsifier":"Equip participants in one mixed-presence session with eye-tracking glasses while the depth cameras record head gaze, then compare the two gaze targets per moment; if the two systems disagree about whether all four users were focused on the same art piece in more than a small fraction of the session, the head-gaze proxy is not reliable for joint attention.","tokens_in":3699,"feed_emoji":"👀","tokens_out":5524,"duration_ms":57052,"temperature":0.7,"pith_summary":"This paper tries to establish that head gaze data captured by depth cameras, rather than wearable eye trackers, can serve as an unobtrusive and effective basis for measuring joint attention in mixed-presence collaboration on wall-sized displays. The authors describe a pipeline that records, normalizes, and merges head-gaze time series from two rooms, then quantify joint attention by checking whether all four participants' gaze targets stay within a distance threshold corresponding to the size of the art pieces they discuss. In a preliminary analysis of one session, they report that horizontal gaze values are correlated, particularly between participants in the same room, and also, more weakly, across the two sites. The reason this matters is that robust evaluation of collaboration quality around room-scale displays currently relies on eye-tracking hardware that is either too constrained or obtrusive, and the proposed method is meant to fill that gap.","feed_headline":"Head gaze from depth cameras tracks joint attention across rooms","feed_subtitle":"Four people in two rooms: head gaze shows collocated partners focus together, and a metric quantifies it.","key_machinery":"The central object is head gaze, i.e., the direction of a participant's head as estimated by depth-camera body tracking, used as a stand-in for eye gaze. The argument is carried by a processing pipeline: record body-tracking streams in both rooms, merge and normalize them into a single time series, interpolate over tracking losses, then apply Spearman correlation to examine coupling and a Euclidean-distance threshold over gaze targets to define joint attention; the threshold is anchored to the maximum dimensions of the art pieces in the shared web-based layout. A key implementation detail is that gaze values are normalized between 0 and 1 for both axes so that the two rooms' differently sized displays are comparable.","core_discovery":"On the paper's own terms, the central claim is that joint attention in a four-user, two-room wall-sized display setting can be detected and quantified from head gaze alone. The authors define joint attention operationally: at each timestamp, compute the maximum Euclidean distance between the normalized gaze targets of all four participants, and if that maximum falls below a threshold set by the maximum dimensions of an art piece, count the moment as joint attention; a second variant drops the worst participant to capture three-out-of-four joint attention. Analyzing one session, they find that horizontal gaze values are correlated across users, with the strongest correlations between collocated participants, and cross-site correlations present but weaker. These results are presented as promising early evidence that the pipeline yields meaningful insight, with the explicit next step being to aggregate across all sessions and conditions.","pith_inferences":["Editorial extension: if head gaze lags true gaze, the method likely underestimates joint attention during quick glances; a head-mounted eye-tracker comparison in the same sessions would show whether the systematic error is acceptable.","Editorial extension: the same pipeline applied to short windows around deictic references, as the authors suggest, could test whether joint attention spikes exactly when one participant points or names an art piece.","Editorial extension: comparing the screen-wide attention-cue condition against a no-cue condition in the full dataset would quantify how much the cue improves cross-site joint attention, which the single-session analysis cannot yet establish."],"forward_implications":["If head gaze suffices, mixed-presence collaboration can be evaluated without wearable or fixed eye trackers, removing a source of obtrusiveness that can alter natural behavior.","Joint attention becomes a time-varying quantity, so sessions can be compared on how much of the collaboration was spent with all four, or three of four, participants focused on the same display area.","Separating same-room from cross-site gaze correlations gives a direct measure of how strongly collocated partners coordinate their attention versus how much remote coupling the awareness cues produce.","The threshold-based definition ties joint attention to meaningful task objects (art pieces), making the metric interpretable rather than purely geometric."],"supporting_citations":[{"why":"Establishes the premise that head position and head gaze give a good indication of where a person's attention is, the basis for using head gaze instead of eye tracking.","marker":"[13]"},{"why":"Describes the depth-camera tracking and awareness pipeline used to record and extract the head gaze data in both rooms.","marker":"[3]"},{"why":"Provides the situated-intelligence platform underlying the recording infrastructure that produced the gaze time series.","marker":"[2]"},{"why":"Categorizes gaze as a component of workspace awareness in remote collaboration on wall-sized displays, motivating the measurement.","marker":"[4]"},{"why":"Defines workspace awareness, the broader construct that joint attention is meant to support.","marker":"[6]"},{"why":"Treats joint attention as a collaboration indicator and supplies the eye-tracking based approach that the paper adapts to head gaze.","marker":"[11]"},{"why":"Justifies the choice of Spearman's rho over Pearson for the correlation analysis because it is robust to outliers.","marker":"[5]"},{"why":"Makes the session dataset publicly available, allowing the reported analysis to be reproduced.","marker":"[1]"}],"fun_headline_variants":["Head gaze alone reveals joint attention across rooms","Gaze metric tracks shared focus on wall displays","Cross-room joint attention from head pose data","Quantifying joint attention with head gaze in mixed presence","Gaze tracks joint attention across two rooms"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The measure stands on head gaze from depth-camera body tracking being a trustworthy stand-in for where people are actually looking; if participants often look without turning their heads, joint attention is systematically undercounted.","fun_headline_variants_meta":{"raw":{"variants":["Head gaze alone reveals joint attention across rooms","Gaze metric tracks shared focus on wall displays","Cross-room joint attention from head pose data","Quantifying joint attention with head gaze in mixed presence","Gaze tracks joint attention across two rooms"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000185,"raw_usage":{"total_tokens":1237,"prompt_tokens":779,"completion_tokens":458,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":395,"completion_tokens_details":{"reasoning_tokens":389}},"tokens_in":395,"tokens_out":458,"duration_ms":5343,"temperature":1.0,"reasoning_tokens":389,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T15:31:05.401457+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Equip participants in one mixed-presence session with eye-tracking glasses while the depth cameras record head gaze, then compare the two gaze targets per moment; if the two systems disagree about whether all four users were focused on the same art piece in more than a small fraction of the session, the head-gaze proxy is not reliable for joint attention.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the premise that head position and head gaze give a good indication of where a person's attention is, the basis for using head gaze instead of eye tracking."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes the depth-camera tracking and awareness pipeline used to record and extract the head gaze data in both rooms."},{"cited_title":"Platform for Situated Intelligence","cited_arxiv_id":"2103.15975","evidence_quote":"Provides the situated-intelligence platform underlying the recording infrastructure that produced the gaze time series."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Categorizes gaze as a component of workspace awareness in remote collaboration on wall-sized displays, motivating the measurement."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Treats joint attention as a collaboration indicator and supplies the eye-tracking based approach that the paper adapts to head gaze."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Justifies the choice of Spearman's rho over Pearson for the correlation analysis because it is robust to outliers."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Makes the session dataset publicly available, allowing the reported analysis to be reproduced."}],"review_version":1}