{"id":"fb7ac25c-73b5-4093-ae75-b6afe9251ea6","arxiv_id":"1908.00611","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper proposes Vis4Vis, a subfield in which visualization is used to analyze and communicate the complex observational data collected during empirical visualization research.","lead":"This position paper argues that visualization researchers should use visualization tools to analyze the rich sensor and interaction data collected during their own studies, a new subfield it names Vis4Vis. Generalists might read it to see a proposed shift in how complex visualization systems are evaluated.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The key premise that established evaluation methods 'can no longer capture' complex visualizations is asserted without supporting evidence; the paper's examples mostly show visualization as a complement, so the subfield claim remains conditional.","rationale":"Good-faith reading: this is a position paper, not an empirical study. It asks the community to recognize a subfield and to adopt data-rich recording and visual analysis as standard evaluation practice. The strongest claim is the call to action, and the necessary condition for that call is the inadequacy of traditional methods. I looked for evidence of that inadequacy in the manuscript. The background sections cite BELIV work and eye-tracking research, but no review or case collection showing a systematic failure of established methods. The examples in Section 3 are most naturally read as cases where visualization adds insight—e.g., identifying reading strategies and building hypotheses for later statistical tests—rather than as cases where traditional evaluation broke down. This makes the paper's own evidence more consistent with a complementary role than with the strong necessity claim. Section 4.2 identifies reliability risks of interactive visual analysis, so the proposal cannot simply assume that data-rich visual analysis is more reliable than the methods it would supplement. These points do not amount to an internal inconsistency; the paper is clearly labeled as a subjective position statement. They mean, however, that the central premise is not yet supported. An empirical survey or a re-analysis of the cited case studies could settle whether the premise holds. Since the concern is the same one identified by the reader, and the verdict was already CONDITIONAL, no change is needed.","tokens_in":12601,"tokens_out":5500,"duration_ms":57471,"concrete_test":"Perform a structured review of user-study papers at IEEE VIS and EuroVis from 2016 to 2023: for each paper that evaluates a complex or immersive visualization system, classify whether the reported conclusions depended on data-rich visual analysis (e.g., gaze or interaction-log visualization) or would have been reachable with traditional measures such as completion time, accuracy, and questionnaires, and record whether the authors explicitly stated that traditional methods were insufficient. If most successful evaluations are fully handled by traditional measures and few papers report insufficiency, the premise is weakened; if a substantial set of papers demonstrates changed or otherwise unreachable conclusions through data-rich visual analysis, the premise is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central argument turns on the premise stated in the abstract and Section 1: 'many of the established methods of empirical studies can no longer capture the full complexity of the evaluation.' This premise is load-bearing because it converts a potentially useful technique into a needed subfield. The paper does not provide a systematic survey or documented cases where traditional methods demonstrably failed; instead, it offers a curated list of related work and illustrative examples drawn largely from the author's own eye-tracking studies. Those examples, as described in Section 3, tend to show visualization supporting hypothesis generation and the definition of new dependent measures that then feed statistical testing (e.g., Netzel et al. [64], [63]; Burch et al. [16]). This is evidence that visualization is a useful complement to traditional evaluation, not that traditional evaluation 'can no longer capture' the complexity. Moreover, Section 4.2 acknowledges that interactive visual analysis has its own reliability problems, so the assertion that data-rich visual analysis yields 'more reliable interpretations' is not established either. The paper is honest about being a subjective position statement (Section 5) and about not being the only answer (Section 6), but those caveats do not replace the missing empirical support. Consequently, the strong version of the proposal—establishing Vis4Vis as a distinct subfield with dedicated venues and keywords—rests on an unverified premise; a weaker version ('visualization is a valuable addition to the evaluation toolbox') would be much better supported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This position paper argues that empirical visualization research should establish a new subfield, 'Vis4Vis' (visualization for visualization), in which visualization methods are used to analyze and communicate the large, heterogeneous, time-dependent data acquired during empirical studies. The author claims that traditional evaluation methods can no longer capture the full complexity of increasingly sophisticated visualization systems, and that data-rich observations (eye tracking, physiological sensors, interaction logs) combined with interactive visual analysis offer a promising route. The paper grounds the argument in eye tracking research, describes a generalized problem characterization with data/visualization types and analysis/dissemination goals, and issues a call to action for integrating Vis4Vis into major venues and research practices. The manuscript is explicitly framed as a subjective position statement rather than a systematic review or empirical validation.","tokens_in":12868,"tokens_out":2156,"duration_ms":23194,"significance":"If the central claim is accepted, Vis4Vis could become a recognized application domain within visualization research, influencing how empirical studies are designed, analyzed, and reported, and potentially extending to HCI and other data-rich empirical disciplines. The paper's strengths include a concrete pipeline for eye tracking analysis (Figure 1), a clear enumeration of open challenges (e.g., data fusion, scalability, privacy), and an honest acknowledgment of its subjective nature and the reliability limits of interactive visual analysis. However, the significance rests on an unverified premise about the insufficiency of established evaluation methods; the presented examples show visualization as a valuable complement to statistical testing rather than as a necessary replacement, so the case for a distinct subfield remains conditional.","major_comments":[{"comment":"The load-bearing assertion that 'many of the established methods of empirical studies can no longer capture the full complexity of the evaluation' is stated as a fact without systematic supporting evidence. The examples in Section 3 (Netzel et al. [64], Netzel et al. [63], Burch et al. [16]) show visualization helping to formulate hypotheses and define new dependent measures that are then tested statistically, which demonstrates complementarity with traditional methods rather than their failure or insufficiency. To make the case for a new subfield, the paper should either provide documented cases where traditional methods demonstrably failed to capture relevant complexity, or clearly reframe this premise as a conjecture and argue for it on programmatic grounds.","section":"Abstract and Section 1"},{"comment":"The paper acknowledges that interactive visual analysis 'might lead to different findings, depending on the interaction steps taken by the analyst' and therefore is typically accompanied by statistical analysis to obtain controlled answers. This acknowledgment directly qualifies the abstract's claim that data-rich visual analysis provides 'more reliable interpretations of empirical research.' The manuscript should either reconcile this tension or soften the reliability claim, because the current wording overstates the epistemic status of visual analysis while its own limitation note undermines it.","section":"Section 4.2"},{"comment":"The call to action to integrate Vis4Vis into main conference keywords and venues presumes that existing venues (BELIV, ETVIS) are insufficient for the proposed agenda. The paper notes that BELIV 'implicitly supports' Vis4Vis and ETVIS does so for eye tracking, but it does not explain why these venues could not accommodate or be extended to the broader Vis4Vis scope. Without this argument, the institutional reform proposal is not fully justified, even if the technical direction is sound.","section":"Section 5"}],"minor_comments":[{"comment":"The caption lists labels A, B, C, D, and F but no E, and it references 'temporal navigation (F)' while the timeline is also described as 'F' earlier; please recheck the label assignments and correct the inconsistency.","section":"Section 3, Figure 2 caption"},{"comment":"The text says 'whether we have to analyze data for individual participants or groups of participates'; 'participates' should be 'participants'.","section":"Section 4.1"},{"comment":"Some reference formatting is inconsistent, e.g., reference [10] renders 'V A2' instead of 'VA2', and [46] abbreviates 'IEEE Transactions Visualization Computer Graphics' instead of the full journal name; please unify style.","section":"References"},{"comment":"The term 'Vis4Vis' is sometimes written with a space ('visualization for visualization ( Vis4Vis)') and sometimes without; please make the typography consistent.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper is a position statement and a preprint of a book chapter, so the novelty and scope expectations differ from a standard archival research article. The self-citation pattern is notable but not inappropriate for a position statement drawing on the author's own prior work; however, the editor may wish to ensure that the peer review emphasizes the distinction between a research agenda and an empirical claim. The central premise needs reframing or support before the manuscript can be accepted as a rigorous argument."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I’ll cut to the chase: this is a position paper, not a research paper, and it should be read as such. It names something that already exists—using visualization and visual analytics to make sense of eye-tracking and interaction data from visualization studies—and argues that this practice deserves recognition as its own subfield. That’s the actual contribution: a useful label and a call to action, supported by a clear synthesis of prior work.\n\nThe paper does several things well. The writing is direct and well organized. Section 3’s eye-tracking pipeline (foraging/sensemaking loops) is a nice frame, and the three examples from the author’s own work—scatterplots/parallel coordinates, metro maps, tree diagrams—concretely show how visual analysis generated hypotheses that later became testable metrics. That is real evidence, and it is the strongest part of the case. I also appreciate that the author explicitly acknowledges, in Section 4.2, that interactive analysis has reliability problems, and in Section 6 that Vis4Vis is not the only answer.\n\nThe soft spot is the load-bearing premise. The abstract says established methods ‘can no longer capture the full complexity of the evaluation,’ but the paper never demonstrates that any traditional method actually failed. The cited examples show visualization as a complement to statistical testing—visualization fed hypotheses, statistics confirmed them. That’s an argument for adding visualization to the toolbox, not for a new subfield built on the inadequacy of existing methods. If the author wants the strong version of the claim, we need documented cases where standard evaluation broke down and visual analysis succeeded. Without that, the proposal stands as a well-argued recommendation, not an established need.\n\nOn citations: the pattern is heavily self-referential, but the cited work is genuinely the relevant prior art, and I don’t see missing references. Self-citation here is not a flaw so much as a reminder that this is one group’s agenda.\n\nWho is this for? Researchers working on visualization evaluation, especially eye-tracking studies, and anyone in the BELIV community. They’ll get a clear framing and a useful term. It’s a good reading-group discussion piece.\n\nBottom line: yes, it deserves serious peer review as a position statement. The reviewer should ask for evidence for the strong premise, or a rewording that drops it. But the paper is coherent, honest, and citable.","headline":"A clear, honest position paper that names an existing practice; the case for a new subfield rests on an unproven premise that traditional evaluation methods are failing.","tokens_in":13367,"tokens_out":3102,"would_cite":true,"duration_ms":27091,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This position paper argues that the visualization community should establish Vis4Vis—visualization for visualization—as a subfield using visualization to analyze and communicate the rich data collected during empirical visualization…","keywords":["Vis4Vis","visualization for visualization","empirical visualization research","evaluation methodology","eye tracking","visual analytics","data-rich user studies","in-the-wild studies"],"falsifier":"A concrete test would be a systematic comparison of evaluation practice before and after adopting visual analytics of rich study data: if a set of complex visualization evaluations yields the same conclusions, reliability, and insight from conventional statistics alone as from adding interactive visual analysis of eye tracking, interaction logs, or physiological data, the claim that traditional methods are insufficient would be contradicted. A meta-analysis of in-the-wild visualization studies showing that standard methods already capture their findings would also settle the question.","tokens_in":12410,"feed_emoji":"👁️","tokens_out":6473,"duration_ms":59970,"temperature":0.7,"pith_summary":"This position paper argues that empirical evaluation in visualization research is outgrowing its inherited methods: sensors and logging now produce large, time-dependent, heterogeneous study data that traditional statistical user studies were not designed to handle. The author proposes a new subfield, Vis4Vis (visualization for visualization), defined as using visualization to analyze and communicate data from empirical studies in order to advance visualization research itself. If the proposal is right, visual analytics of data-rich study recordings becomes a standard part of evaluation, especially for in-the-wild, mobile, and collaborative settings where controlled laboratory conditions no longer apply. The paper makes the case through eye tracking as a worked example and then generalizes to other sensor, interaction, video, audio, and performance data.","feed_headline":"Make visualization research its own visualization application","feed_subtitle":"As sensors and logging make study data huge and messy, visualization should do the analysis.","key_machinery":"The central mechanism is the visual analytics pipeline adapted to study recordings, illustrated in the paper by an eye-tracking analysis pipeline: raw gaze and complementary data are processed and annotated, mapped to spatial, temporal, and relational visualizations, and explored through two linked loops—a foraging loop for investigating observables and a sensemaking loop for building, confirming, or rejecting hypotheses. The paper generalizes this pipeline to any timestamped observational data, since data from different sensors and logs can be registered along a common timeline even when sampled at different rates. The machinery does the work of turning unstructured sensor and log streams into interpretable evidence that informs and validates empirical visualization research.","core_discovery":"The central claim is that visualization researchers should make their own empirical studies an application domain for visualization. Established evaluation approaches, mostly adopted from other fields and earlier, data-poor eras, are increasingly unable to capture the complexity of sophisticated visualization systems; the proposed remedy is to collect rich observational data during studies and to analyze and report that data with visualization. Eye tracking serves as the worked example: visual exploration of scanpaths, areas of interest, and spatiotemporal gaze data has already helped form hypotheses and identify qualitative findings in the author's own studies. From that example the paper generalizes to a data model of timestamped, heterogeneous observational data and argues that the visual analytics pipeline—processing, mapping, interactive exploration, and sensemaking—should be applied to study data as a matter of course, and that Vis4Vis should be explicitly recognized in conference scopes and calls for papers.","pith_inferences":["A reader can infer that the proposal's success is measurable: if Vis4Vis is adopted, evaluation sections of visualization papers will begin to include interactive visual analysis systems or links to them, alongside traditional statistics.","The argument is strongest for studies where each participant sees different stimuli, such as mobile eye tracking in the wild; a natural extension is to test Vis4Vis methods first in those settings rather than in controlled lab trials.","The paper's emphasis on keeping raw data in the analysis pipeline implies a broader reproducibility agenda: published study results could come with the visual analysis setup itself, not just the raw data, making replication more concrete.","Because the paper frames evaluation as an application domain, it implicitly predicts that visualization of study data will become a source of new visualization techniques that later generalizes to fields outside visualization."],"forward_implications":["Data-rich recordings of gaze, interaction, physiology, video, and audio become a standard component of empirical visualization evaluation, especially for studies in the wild and mobile settings.","Visualization conferences and paper keywords explicitly include visualization research as an application domain, so Vis4Vis contributions are reviewed as first-class application papers.","Researchers develop new visualization-based methods for reporting complex study results, including storytelling, and for privacy-preserving release of open research data.","Visual analytics, statistical testing, and machine learning are combined into an analysis toolkit for study data, with the author expecting this to carry over to neighboring disciplines."],"supporting_citations":[{"why":"This work supplies the initial call for exploratory visual analysis and hypothesis building in eye-tracking evaluation, which the paper extends into Vis4Vis.","marker":"[49]"},{"why":"This work documents open issues such as scanpath comparison, data fusion, and missing tools, motivating the need for a dedicated subfield.","marker":"[50]"},{"why":"This work demonstrates a visual analytics approach for evaluating visual analytics systems with eye tracking, think-aloud, and interaction data, the direct precursor of Vis4Vis.","marker":"[10]"},{"why":"This work adds visual analysis and coding of user behavior to data-rich evaluation, supporting the paper's proposal to integrate qualitative and quantitative data.","marker":"[6]"},{"why":"This survey and taxonomy of eye-tracking visualization supplies the state of the art from which the paper selects appropriate techniques.","marker":"[12]"},{"why":"This work contributes space-time visual analytics for gaze data, an example of the visual analysis methods the paper argues should become standard.","marker":"[53]"},{"why":"This system combines spatiotemporal gaze analysis with area-of-interest analysis and serves as the paper's concrete example of integrated visual evaluation.","marker":"[51]"},{"why":"This work integrates eye tracking with interaction logs for evaluating interactive visualization systems, an instance of the multi-source data fusion Vis4Vis calls for.","marker":"[8]"}],"fun_headline_variants":["Visualize your visualization study data","Bring visual analytics to empirical visualization research","Vis4Vis: a call to visualize study logs","When sensors flood your study, visualize the flood","Make empirical study data the next visualization challenge"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that established empirical methods can no longer capture the full complexity of evaluating sophisticated visualization systems; if traditional controlled user studies and statistics still suffice for this purpose, the case for establishing Vis4Vis as a new subfield loses much of its force.","fun_headline_variants_meta":{"raw":{"variants":["Visualize your visualization study data","Bring visual analytics to empirical visualization research","Vis4Vis: a call to visualize study logs","When sensors flood your study, visualize the flood","Make empirical study data the next visualization challenge"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000276,"raw_usage":{"total_tokens":1657,"prompt_tokens":963,"completion_tokens":694,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":579,"completion_tokens_details":{"reasoning_tokens":628}},"tokens_in":579,"tokens_out":694,"duration_ms":7215,"temperature":1.0,"reasoning_tokens":628,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:42:57.323569+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete test would be a systematic comparison of evaluation practice before and after adopting visual analytics of rich study data: if a set of complex visualization evaluations yields the same conclusions, reliability, and insight from conventional statistics alone as from adding interactive visual analysis of eye tracking, interaction logs, or physiological data, the claim that traditional methods are insufficient would be contradicted. A meta-analysis of in-the-wild visualization studies showing that standard methods already capture their findings would also settle the question.","supporting_citations":[{"cited_title":"In: Proceedings of the Workshop on Beyond Time And Errors: Novel Evaluation Methods for Visualization (BELIV), pp","cited_arxiv_id":null,"evidence_quote":"This work supplies the initial call for exploratory visual analysis and hypothesis building in eye-tracking evaluation, which the paper extends into Vis4Vis."},{"cited_title":"Information Visualization 15(4), 340–358 (2016) Vis4Vis: Visualization for (Empirical) Visualization Research 15","cited_arxiv_id":null,"evidence_quote":"This work documents open issues such as scanpath comparison, data fusion, and missing tools, motivating the need for a dedicated subfield."},{"cited_title":"IEEE Transactions on Visualization and Computer Graphics 22, 61–70 (2016)","cited_arxiv_id":null,"evidence_quote":"This work demonstrates a visual analytics approach for evaluating visual analytics systems with eye tracking, think-aloud, and interaction data, the direct precursor of Vis4Vis."},{"cited_title":"In: Proceedings of the IEEE Conference on Visual Analytics Science and Technology, pp","cited_arxiv_id":null,"evidence_quote":"This work adds visual analysis and coding of user behavior to data-rich evaluation, supporting the paper's proposal to integrate qualitative and quantitative data."},{"cited_title":"Computer Graphics Forum36(8), 260–284 (2017)","cited_arxiv_id":null,"evidence_quote":"This survey and taxonomy of eye-tracking visualization supplies the state of the art from which the paper selects appropriate techniques."},{"cited_title":"IEEE Transactions on Visualization and Computer Graphics 19(12), 2129–2138 (2013)","cited_arxiv_id":null,"evidence_quote":"This work contributes space-time visual analytics for gaze data, an example of the visual analysis methods the paper argues should become standard."},{"cited_title":"In: Proceedings of the ACM Symposium on Eye Tracking Research & Applications, pp","cited_arxiv_id":null,"evidence_quote":"This system combines spatiotemporal gaze analysis with area-of-interest analysis and serves as the paper's concrete example of integrated visual evaluation."},{"cited_title":"In: Proceedings of the Workshop on Beyond Time And Errors: Novel Evalu- ation Methods for Visualization (BELIV), pp","cited_arxiv_id":null,"evidence_quote":"This work integrates eye tracking with interaction logs for evaluating interactive visualization systems, an instance of the multi-source data fusion Vis4Vis calls for."}],"review_version":1}