{"id":"c172bc29-59ed-4724-88ff-2cfe47101205","arxiv_id":"2502.01597","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":1,"one_line_summary":"A pilot eye-tracking and frame-semantic method suggests that understanding far-future and pessimistic scenarios demands more cognitive effort.","lead":"This paper proposes FutureVision, a method that combines eye tracking with semantic annotation to measure how hard people's brains work when they interpret fictional ads about the future. A pilot with three speakers suggests that far-future and pessimistic scenarios trigger longer and more erratic eye movements, but the tiny sample makes the result preliminary.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The causal link from base-space fractures to cognitive load is untested because base-space violation is inferred from the same participant descriptions that gaze is meant to explain; the reported gaze differences cannot carry the abstract's causal conclusion.","rationale":"The reader's weakest assumption—that oculomotor metrics directly index cognitive effort—is related but not identical to my concern. Even if eye-tracking metrics are accepted as valid cognitive-load indicators, the paper still does not establish that the gaze differences are caused by base-space violations rather than by the circularly defined semantic-divergence measure. This is more load-bearing because it threatens the central interpretive claim even under the reader's optimistic view of eye tracking. The paper has genuine strengths: it is an explicitly framed pilot, the protocol received ethics approval, data and scripts are made available via OSF, and the multimodal frame-semantic annotation is a sophisticated contribution. However, the abstract and conclusion overstate the evidence by asserting a causal relationship that the experimental design cannot test. The verdict should remain CONDITIONAL, requiring an independent operationalization of base-space violation and a larger, pre-registered replication before the empirical claim can be accepted. I do not recommend REJECT because the methodology itself is plausibly sound and the authors acknowledge the pilot status; the fix is a design improvement rather than a wholesale abandonment of the approach.","tokens_in":9158,"tokens_out":4116,"duration_ms":40818,"concrete_test":"Have independent annotators, blind to the eye-tracking data and to participants' descriptions, rate each of the eight ads on the degree to which it violates or fractures the expected base frame (e.g., a Likert-scale judgment based on the ad alone). Then fit pre-registered mixed-effects models predicting fixation duration and saccadic variability from these independent violation ratings, with scenario valence and future distance as fixed effects and image complexity, text length, and AOI layout as covariates. If the independent violation ratings no longer predict gaze after controlling for those factors, the causal attribution to base-space fractures fails. As a secondary check, re-estimate the models with corrected statistics and at least 30 participants to ensure the reported effects replicate without the duplicated estimates.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing problem is not sample size alone, but that the paper's central causal construct—'base space violation/fracture'—is never independently operationalized. The statistical methods section lists fixed effects as scenario valence and future distance only; no model includes a base-space violation predictor. The only operational link is the cosine similarity between stimulus frame annotations and each participant's post-hoc verbal description (Table 1). Low similarity is then interpreted as evidence of base-space violation, as in the Privée Mystique discussion. But that similarity is a product of the same comprehension process whose effort the gaze metrics are meant to explain. A participant who finds an ad confusing, or who interprets it idiosyncratically, will produce both low semantic similarity and unstable gaze; the correlation cannot distinguish 'base-space violation caused cognitive load' from 'idiosyncratic interpretation caused both.' The abstract's causal claim therefore rests on a circular operationalization. Secondary statistical problems—duplicated β = −0.19, p = 0.041, R² = 0.27 across cluster and interaction models; Table 2 and Table 4 inconsistent with their surrounding text; only three participants—make the reported effect estimates unreliable, but the conceptual circularity would undermine the claim even with a larger sample.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes FutureVision, a methodology that combines multimodal frame-semantic annotation (FrameNet Brasil) with mobile eye-tracking to investigate cognitive effort in the comprehension of future scenarios. In an exploratory pilot, three pairs of participants viewed eight fictional advertisements from the FTI 2022 Tech Trends Report, differing in valence (optimistic/pessimistic) and counterfactuality (near/far future); one participant in each pair wore a Tobii Pro Glasses 3 eye tracker and described the scenario to a listener. Semantic similarity between the frame-semantic annotations of the stimuli and of the participants' descriptions is computed, and gaze metrics (fixation duration, saccadic variability, pupil dilation, AOI transitions) are compared across conditions. The paper reports that far-future and pessimistic scenarios are associated with longer fixations and more erratic saccades, and interprets this as support for the hypothesis that 'fractures in the base spaces' underlying future interpretations increase cognitive load. The abstract presents this as a central finding, while the Limitations section acknowledges the pilot's small sample.","tokens_in":9477,"tokens_out":5132,"duration_ms":46710,"significance":"If validated, the FutureVision methodology would be a useful interdisciplinary tool: it integrates well-established eye-tracking measures with rich multimodal frame-semantic annotation, and it targets a genuinely under-studied question about how future scenarios are understood. The OSF repository with data and scripts is a strength that supports reproducibility. However, the current pilot cannot carry the causal claims in the abstract. The operational circularity around 'base space violation' and the severe underpowering (N=3) mean that the present evidence is only feasibility-level; the methodological contribution is potentially sound, but the empirical conclusions are not yet supported.","major_comments":[{"comment":"The central construct, 'base space violation/fracture,' is never independently operationalized. The cosine similarity between the frame-semantic representations of the stimuli and of the participants' descriptions (Table 1) is the only quantitative link proposed, and the same similarity scores are used as the outcome in the mediation and regression analyses. The gaze metrics are then interpreted as a consequence of base-space violations, but no statistical model in the Statistical analysis section includes a base-space violation predictor. The Privée Mystique discussion in the Results section exemplifies the circularity: low cosine similarity is taken as evidence of a base-space violation, and the gaze data are then offered as independent support for that violation. A participant who finds an ad confusing or interprets it idiosyncratically will plausibly produce both low semantic similarity and unstable gaze, so the reported associations cannot distinguish 'base-space violation caused cognitive load' from 'idiosyncratic interpretation caused both.' Please provide an independent operationalization of base-space violation, such as a priori stimulus manipulation or expert annotation of frame incongruity, and model its effect on gaze separately from the interpretation outcome.","section":"Results, Table 1 and Integration of behavioral data"},{"comment":"The mediation analysis as reported is not a valid mediation model. Table 2 lists a multiple regression with Fut. Dist. and Pup. Dilat. as simultaneous predictors, while the text describes a causal chain in which future distance predicts pupil dilation and pupil dilation then predicts semantic similarity. There is no indirect effect estimate, no bootstrap confidence interval, and the claim that the effect of future distance 'decreased' from β = -0.21 to β = -0.14 is not, as stated, a comparison of a total effect with an indirect path; it is a comparison across two different models. The reported R² = 0.23 is absent from the table. Please re-run a proper mediation analysis reporting the a, b, c, and c' paths, bootstrapped indirect effects, and include the R² in the table.","section":"Eye-Tracking insights into cognitive load, Table 2"},{"comment":"The identical values β = -0.19, p = 0.041, R² = 0.27 appear for both Cluster 1 and Cluster 2 in the cluster analysis and for the 'Fix. Durat.' estimate in Table 4. This is statistically implausible and indicates a reporting error rather than a genuine finding. In addition, it is unclear how a K-means cluster assignment yields p-values for a cluster-level beta; p-values from models applied to cluster-derived labels are not standard. Please report the actual model estimates, standard errors, test statistics, and confidence intervals for each cluster and for the interaction model, and correct whatever duplication or misassignment produced these identical numbers.","section":"Eye-Tracking insights into cognitive load, Cluster analysis and Table 4"},{"comment":"The empirical results rest on N = 3 participants (the three speakers who wore the eye tracker). Mixed-effects models with random intercepts for participants cannot provide stable variance component estimates or trustworthy p-values at this sample size. The Limitations section appropriately acknowledges the small sample, but the Abstract and Conclusion nonetheless state that the results 'support' the base-space-fracture hypothesis and call the finding 'central.' Please either reframe all empirical claims as feasibility observations, or collect a larger sample and report a power analysis and effect-size justification. Without this, the reported p-values are not meaningful evidence for the proposed cognitive mechanism.","section":"Results, Participant selection and sample size"}],"minor_comments":[{"comment":"The header says 'Consine similarities' and should be corrected to 'Cosine similarities.'","section":"Table 1 header"},{"comment":"There are several typographical and consistency problems, including 'ellaborated' in the Introduction and inconsistent rendering of 'Privée Mystique' as 'Priv´ee Mystique'; a careful proofread would improve readability.","section":"Throughout"},{"comment":"The manuscript does not state whether the eight fictional ads were matched for text length, image complexity, or layout. Because those low-level visual properties can strongly affect eye movements, they are potential confounds for the gaze comparisons and should be either controlled or explicitly discussed.","section":"Methods, Data analysis"},{"comment":"The transition probabilities in Table 3 are reported without confidence intervals or statistical tests; as a result, it is unclear whether the differences between Cluster 1 and Cluster 2 are meaningful beyond describing the two clusters.","section":"Eye-Tracking insights into cognitive load, Table 3"},{"comment":"The Limitations section mentions unmeasured potential confounds such as working memory and cultural influences, but it does not mention the circularity issue or the duplicated statistics; both should be acknowledged.","section":"Limitations"},{"comment":"The citation to Torrent and Turner (2023) is used to ground the key theoretical claim of 'Persistence of the Base,' but the reference is only to a conference abstract; please provide a fuller account or cite a more detailed source.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper's interdisciplinary ambition is genuine, and the OSF repository is a welcome touch. However, the gap between the pilot data and the abstract's causal claim is large, and the circular operationalization of 'base space violation' is the kind of issue that would not be fixed simply by adding more participants. I would encourage the editor to require the authors to either (a) substantially reframe the contribution as a methodology proposal with no causal claims, or (b) redesign the analysis around an independent measure of base-space violation and report a properly powered study. The duplicated statistics in Section 4 also need a direct author explanation before evaluation can proceed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: the paper combines eye tracking with FrameNet-style multimodal annotation in a way I haven't seen before, and the pilot data are shared, which is good. But the central claim that 'base space fractures increase cognitive load' is not supported by the design. The gap isn't just N=3, which the authors concede. The deeper issue is operational circularity: base-space violation is inferred from the cosine similarity between the stimulus annotation and the participant's own verbal description. That description is a product of the same comprehension process whose effort the gaze metrics are meant to explain. A participant who is confused will produce both low similarity and erratic gaze; the correlation cannot tell you whether the violation caused the load or whether idiosyncratic interpretation caused both. No model actually includes a base-space violation predictor, only valence and future distance. So the causal language in the abstract is not backed by the analysis.\n\nWhat's genuinely new: the specific pipeline, the use of FTI ads, the multimodal annotation, the open OSF data, and the Privée Mystique case study, which is a nice qualitative illustration of reframing. The theoretical framing is clear and the authors are honest about the pilot status.\n\nSoft spots, in order: the statistical reporting is internally inconsistent—identical beta, p, and R² appear in two different analyses, and the mediation table's numbers don't match the text. That needs to be corrected or it undermines confidence. Then the circularity, which would survive a larger sample. Also, the gaze-as-cognitive-load assumption is asserted rather than validated, though that's a common assumption in eye-tracking work.\n\nWho's this for? Cognitive linguists interested in future-directed communication, and methodologically-minded foresight researchers. It would also be a good reading-group example of how a promising measurement idea can be undermined by a construct that is only measured through the same behavior it is meant to explain.\n\nMy recommendation: send it to peer review, yes, but expect heavy revision. The methodology deserves discussion; the strong claim should be stripped from the abstract, and the authors should be asked to pre-register a larger study and to operationalize base-space violation independently of the dependent measures.","headline":"A genuinely new pilot methodology for linking gaze to frame-semantic comprehension, but the central causal claim is undercut by circular operationalization and N=3.","tokens_in":9928,"tokens_out":2667,"would_cite":false,"duration_ms":24092,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that fractures in the base spaces underlying future-scenario interpretation are measurable through eye-tracking, and that a pilot shows far-future and pessimistic scenarios increase cognitive load.","keywords":["future cognition","eye tracking","cognitive load","frame semantics","multimodal annotation","counterfactual scenarios","base space","mental spaces"],"falsifier":"Take the same eight ads, match their layout, text length, and image complexity across conditions, and present them to new participants with only the described future distance and valence systematically swapped; the hypothesis predicts gaze patterns will track the described scenario rather than the ads' intrinsic features. If gaze patterns instead follow only the physical properties of the ads, or if adding an independent cognitive-load measure such as EEG fails to correlate with the gaze differences, the base-space account would be falsified.","tokens_in":149,"feed_emoji":"👀","tokens_out":5840,"duration_ms":96445,"temperature":0.7,"pith_summary":"This paper proposes a methodology that pairs frame-semantic annotation of multimodal advertisements with eye-tracking to measure how much cognitive effort it takes people to understand communicated future scenarios. To demonstrate it, the authors ran a pilot in which participants explained fictional futuristic ads to a partner while wearing a portable eye tracker. The pilot results indicate that far-future and pessimistic scenarios produce longer fixations and more erratic saccades than near-future and optimistic ones, and that these gaze patterns track with how close participants' semantic interpretations are to the ad's intended meaning. The paper argues that these effects reflect fractures in the base spaces that normally anchor interpretation of future scenarios, and that the proposed method can make such fractures measurable.","feed_headline":"Eye-tracking shows distant futures are harder to process","feed_subtitle":"Pilot data link far-future and pessimistic scenarios to longer fixations, erratic saccades, and cognitive load.","key_machinery":"The load-bearing object is the 'base space' from mental-spaces theory, the initial mental space from which a network of spaces and blends is built, together with the base frame that structures it. The method's machinery is a pipeline: multimodal frame-semantic annotation of both the ad stimulus and the participant's spoken description produces semantic representations; cosine similarity measures the interpretive gap; gaze metrics (fixation duration, saccadic variability, area-of-interest transition probabilities, pupil dilation) operationalize cognitive load; regression, clustering, and Markov-chain analyses link the two. The argument is that when the base space persists, interpretations align and gaze is stable, whereas when it fractures, gaze becomes erratic and interpretation diverges.","core_discovery":"The central discovery is that base-space violations in future-scenario communication are empirically trackable: when a scenario departs from the shared frames that structure everyday expectation, comprehenders show longer fixations, more regressions between text and image, and lower pupil-linked engagement, while their semantic representations diverge from the stimulus. The pilot's most striking case is an optimistic ad that all three participants read as pessimistic, invoking betrayal and theft, because they lacked the background story needed to keep the intended protecting frame intact; gaze heatmaps confirm they attended to the same areas as expected, so the interpretive divergence is attributed to base-space fracture rather than to low-level perceptual differences.","pith_inferences":["If the gaze–cognitive-load link holds, the same protocol could be applied to adaptive user interfaces, flagging the moment a user's interpretation of a scenario diverges and adjusting the message in real time.","The observation that a nominally optimistic ad was uniformly read as pessimistic suggests that base-space persistence depends on culturally shared background knowledge; future work could manipulate that knowledge directly and should show corresponding shifts in gaze and semantic similarity.","Because the paper's own mediation result is only partial, a stronger test would include an independent measure of cognitive load, such as EEG or dual-task performance, to validate the eye-tracking index.","The small pilot sample, acknowledged in the paper, leaves open how much of the gaze effect is driven by individual differences rather than by the scenario conditions; a larger replication could quantify that."],"forward_implications":["Far-future and pessimistic scenarios should be expected to cost more cognitive effort, so designers of future-oriented communications can anticipate comprehension breakdowns and add scaffolding.","The cosine similarity between intended and produced semantic representations can serve as a quantitative measure of communicative success in future-scenario messages.","Pupil dilation partially mediates the effect of future distance on interpretation, meaning a portion of the comprehension gap is explained by cognitive effort.","Cluster analysis suggests that individual differences in processing strategy can be recovered from gaze patterns alone, enabling person-specific predictions of cognitive load.","The methodology is portable: a lightweight eye tracker plus semantic annotation can be used outside a laboratory to evaluate real-world future-oriented materials."],"supporting_citations":[{"why":"Supplies the mental-spaces theory and the notion of base spaces that the hypothesis operationalizes.","marker":"Fauconnier (1997)"},{"why":"Provides frame semantics, the basis for defining base frames and for the annotation categories.","marker":"Fillmore (1982)"},{"why":"Establishes the persistence of base spaces in counterfactual thinking, the key theoretical claim being tested.","marker":"Torrent & Turner (2023)"},{"why":"Supplies the multimodal frame-semantic annotation method used on both stimuli and descriptions.","marker":"Belcavello et al. (2020)"},{"why":"Source of the eight fictional ad stimuli and their valence/counterfactuality grouping.","marker":"FTI (2022)"},{"why":"Provides the eye-tracking methodology and conventions for interpreting fixation and saccade metrics.","marker":"Conklin et al. (2018)"},{"why":"Supplies the cosine-similarity method for comparing frame-semantic representations of stimuli and descriptions.","marker":"Viridiano et al. (2022)"}],"fun_headline_variants":["When futures break frame, gaze goes erratic","Optimistic ads read as pessimistic when frames fail","Distant futures strain the eyes and the mind","Base-space fractures show up in eye movements"],"cache_read_input_tokens":12032,"weakest_assumption_plain":"The load-bearing premise is that eye-tracking metrics directly and monotonically index cognitive effort, and that the gaze differences between conditions are caused by conceptual base-space violations rather than by low-level differences in the ads' visual or textual complexity.","fun_headline_variants_meta":{"raw":{"variants":["When futures break frame, gaze goes erratic","Optimistic ads read as pessimistic when frames fail","Distant futures strain the eyes and the mind","Base-space fractures show up in eye movements"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000719,"raw_usage":{"total_tokens":3156,"prompt_tokens":801,"completion_tokens":2355,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":417,"completion_tokens_details":{"reasoning_tokens":2298}},"tokens_in":417,"tokens_out":2355,"duration_ms":14933,"temperature":1.0,"reasoning_tokens":2298,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T14:50:32.801069+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the same eight ads, match their layout, text length, and image complexity across conditions, and present them to new participants with only the described future distance and valence systematically swapped; the hypothesis predicts gaze patterns will track the described scenario rather than the ads' intrinsic features. If gaze patterns instead follow only the physical properties of the ads, or if adding an independent cognitive-load measure such as EEG fails to correlate with the gaze differences, the base-space account would be falsified.","supporting_citations":[{"cited_title":"APACrefauthors \\ 1997","cited_arxiv_id":null,"evidence_quote":"Supplies the mental-spaces theory and the notion of base spaces that the hypothesis operationalizes."},{"cited_title":"APACrefauthors \\ 1982","cited_arxiv_id":null,"evidence_quote":"Provides frame semantics, the basis for defining base frames and for the annotation categories."},{"cited_title":"\\ Turner, M","cited_arxiv_id":null,"evidence_quote":"Establishes the persistence of base spaces in counterfactual thinking, the key theoretical claim being tested."},{"cited_title":", Viridiano, M","cited_arxiv_id":null,"evidence_quote":"Supplies the multimodal frame-semantic annotation method used on both stimuli and descriptions."},{"cited_title":"APACrefauthors \\ 2022","cited_arxiv_id":null,"evidence_quote":"Source of the eight fictional ad stimuli and their valence/counterfactuality grouping."},{"cited_title":", Pellicer-Sánchez, A","cited_arxiv_id":null,"evidence_quote":"Provides the eye-tracking methodology and conventions for interpreting fixation and saccade metrics."},{"cited_title":", Torrent, T T","cited_arxiv_id":null,"evidence_quote":"Supplies the cosine-similarity method for comparing frame-semantic representations of stimuli and descriptions."}],"review_version":1}