{"id":"161953d1-2629-4fd4-a4a7-0c688768d61c","arxiv_id":"2608.11364","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"QUARTZ provides a multimodal accessible system for four qualitative data visualization types and documents, through an 8-participant RITE study, how BLV users navigate them and which barriers remain.","lead":"Researchers built QUARTZ, a web tool that lets blind and low-vision (BLV) users explore qualitative data visualizations such as concept maps, network graphs, Sankey diagrams, and coding stripes through screen readers, sound, and keyboard navigation. A formative study with 8 BLV participants shows these interfaces are usable, but also reveals accessibility barriers that earlier quantitative-chart systems did not face.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'resolved barriers' claim rests on a single-facilitator RITE trajectory confounded by participant and AT differences; a fixed-system independent-observation check is needed before accepting the central claim.","rationale":"The reader's weakest_assumption correctly identifies the first author's dual role as facilitator and coder as a threat to internal validity. My review agrees that this matters, but the more structurally decisive problem is that RITE's design makes causal attribution of improvement impossible from the reported data: system version, participant, AT, and session order are perfectly confounded. The paper honestly flags this in Section 6.5.5, and the intervention graphs in Figures 7-8 show post-modification gains that are also explained by P7-P8's mouse/zoom usage, P5's untested ZDSR, and P4's skipped tasks. The first author's bias concern is real but secondary; even a perfectly unbiased facilitator could not disentangle system improvements from participant differences with this design. I do not think this requires rejecting the paper. The study is explicitly formative, the qualitative quotes are informative, many tasks were completed successfully, and the limitations section is candid. However, the abstract and conclusion use stronger language than the evidence supports: 'resolved' and 'our evidence indicates they can' should be tempered to 'progressively addressed' and 'suggest they can, pending fixed-system evaluation.' That is exactly the conditional status the reader assigned, so I recommend UNCHANGED rather than a move to accept or reject. The concrete step that would settle the concern is a fixed-system summative study with independent facilitation and coding, especially measuring intervention-free task completion and network-graph performance.","tokens_in":31057,"tokens_out":4167,"duration_ms":42786,"concrete_test":"Run a fixed-system summative evaluation of the final QUARTZ build with at least 8 keyboard-primary BLV participants, facilitated by someone outside the design team and coded by an independent rater using pre-registered task-success criteria and the paper's intervention taxonomy. Report network-graph task success and, crucially, the rate of tasks completed with zero facilitator interventions. If network-graph success remains below roughly 50% or intervention-free completion is rare, the 'resolved barriers' and 'evidence indicates they can independently navigate' claims are not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim has two parts: BLV users can navigate, comprehend, and derive meaning from qualitative visualizations, and iterative co-design resolved the barriers that initially blocked this. The second part is load-bearing for the paper's 'design knowledge' contribution, and it is the least secure. The evidence for resolution is a before/after comparison across sessions (Figures 7-8, Appendix B.4), but RITE deliberately changes the system between participants, so system version is confounded with participant identity, AT configuration, and session order. The paper concedes this in Section 6.5.5: 'direct comparison across sessions reflects both system improvement and individual differences... not controlled experimental evidence.' Additional confounds: P7 and P8 used zoom/mouse rather than keyboard navigation, so post-modification improvement is partly modality-driven; P5 used ZDSR, an untested screen reader, and required facilitator remote control; P4 skipped network-graph tasks, making his zero in Figure 8b structural rather than evidence of improvement; and P8, on the newest build, recorded the highest intervention count (16). Separately, the first author facilitated every session, categorized facilitator interventions, and performed the initial thematic coding (Section 3.1), so the pre/post outcome coding is not independent of the design investment. The task table (Figure 5) also shows many partial and failed outcomes, especially for network graphs, and it does not reveal whether 'success' outcomes were achieved without facilitator assistance. Therefore, the abstract's claim that barriers were 'resolved' and the conclusion that 'our evidence indicates they can' overstate what the data establish: the evidence supports a design trajectory, not a validated resolution.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents QUARTZ, a web-based system offering screen-reader-accessible, multimodal representations of four qualitative data visualization types (concept maps, network graphs, Sankey diagrams, and coding stripes). The authors report a RITE-based formative evaluation with 8 blind and low-vision participants completing 12 tasks, identifying accessibility barriers specific to qualitative visualizations and documenting iterative design modifications intended to resolve them. The paper contributes the system, empirical findings about non-linear navigation and semantic comprehension barriers, and six design guidelines for accessible qualitative visualization. The authors are notably transparent about confounds and limitations, including the single-facilitator design, the changing system across RITE sessions, and the mixed task outcomes for network graphs.","tokens_in":31204,"tokens_out":4757,"duration_ms":43876,"significance":"If the central claim were fully supported, this would be a valuable contribution to an underexplored area of accessible visualization: it would provide formative empirical evidence that BLV researchers can engage with qualitative visualizations when the tools are designed with them, and it would supply concrete design guidance where the literature has focused almost exclusively on quantitative charts. The paper has genuine strengths: the domain is well-motivated, the system covers four visualization types with coordinated multimodal access, the RITE process is documented in detail with a full change log (Appendix B.4), the rule-based parameters are disclosed completely (Appendix A), and the authors openly report partial and failed tasks, AT-specific conflicts, and their own positionality and bias risks. The design guidelines in Section 7.2 are concrete and grounded in observed participant behavior. However, the load-bearing claim that iterative co-design 'resolved' barriers is not established by the reported data, and the paper's own evidence suggests the conclusion should be scoped more carefully.","major_comments":[{"comment":"The abstract and conclusion claim that iterative co-design 'resolved' accessibility barriers, but the RITE trajectory confounds system version with participant identity, assistive-technology configuration, interaction modality, and session order. The paper acknowledges this in §6.5.5 ('not controlled experimental evidence'), yet the contribution language still asserts resolution. Specific confounds visible in the paper: P4's zero in Fig. 8b is structural because network-graph tasks were skipped; P5 used ZDSR, an untested screen reader, and required facilitator remote control; P7-P8 used zoom and mouse rather than keyboard navigation; and P8, on the newest build, recorded the highest intervention count (16). To support the 'resolved' claim, the authors should either add a fixed-system evaluation with an independent observer or reframe the contribution to 'surfaced and partially addressed barriers' across the RITE trajectory. This is load-bearing for the paper's design-knowledge contribution.","section":"§3.1, §5.6"},{"comment":"The first author built QUARTZ, facilitated every session, categorized facilitator interventions, and conducted the initial thematic coding. The paper acknowledges this bias risk in §3.1, and the mitigation (a priori task criteria and consensus with the second author) is reasonable but does not remove the structural dependence on a single invested evaluator for the pre/post comparisons that ground the 'resolution' claim. I recommend adding an independent coding of task outcomes and intervention categories, such as a second coder blind to session order with agreement statistics, or explicitly presenting the trajectory as exploratory and removing causal 'resolved' language from the contributions. Without this, the central empirical claim is not independently verifiable from the reported evidence.","section":"§6.3.1, Fig. 5, §6.4.2, Fig. 6"},{"comment":"The conclusion that BLV users 'can independently navigate, comprehend, and derive meaning' from qualitative data visualizations is stronger than the reported task data. Fig. 5 shows many partial and failed outcomes, especially for network graph tasks (T4-T6), and the same section notes that network graphs were consistently the most challenging visualization type. Additionally, Fig. 6 reports AUS scores aggregated across successive system versions (mean 70.6, s=21.6), so the 'above-average mean AUS' statement is a property of the design trajectory, not a benchmark of the final build, as the paper itself notes. Please report per-type success rates and explicitly scope the conclusion to the visualization types and task levels where success was actually observed.","section":"§5.2, Table 2, §6.2.3, §6.5.4"},{"comment":"There is an internal inconsistency about P4's assistive technology: Table 2 lists P4 as using NVDA on Windows, but §6.2.3 quotes P4 as making 'JAWS compatibility' the gating factor and §6.5.4 attributes 'JAWS virtual cursor' conflicts to P4. Please reconcile the participant table with the quoted AT attribution.","section":"§3.1, reference [46]"},{"comment":"In §3.1, the citation [46] appears to be the wrong reference: the claim about R2's daily screen-reader use surfacing interaction barriers is supported by [30] and likely should cite the VoxLens paper [47] rather than the language-preferences paper currently listed as [46].","section":"§4.2.3, Appendix A.2"},{"comment":"The Evaluation Panel is described in §4.2.3 as reporting a 1-5 readiness score with an export threshold of 3, while Appendix A.2 describes a 0-100 additive point allocation with a ready threshold of 60. Please clarify the relationship between these two scoring systems so readers can map the interface labels to the underlying calculation.","section":"Fig. 5 caption"},{"comment":"Figure 5 uses only symbols (✓, ∼, ✗) to present task outcomes; given the paper's focus on accessibility, the symbol-only legend may not be usable by screen reader users. Please add an accessible textual equivalent or a data table in the caption or body text.","section":"Fig. 7 caption"},{"comment":"Given the small sample and RITE-induced system changes, the AUS score distribution in Fig. 6 would benefit from a brief note that the scores are not directly comparable across sessions; the text makes this point in §6.4.2, but the figure caption does not.","section":"§6.4.2, Fig. 6"}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":"I see this as a genuine formative contribution with unusually candid reporting of limitations and mixed outcomes. The main gap is between the 'resolved barriers' framing and the confounded single-facilitator RITE evidence; a fixed-system follow-up with independent observation, or a carefully reframed contribution, would close it. The paper fits ASSETS well and should not be rejected, but the central claim needs to be made commensurate with the evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nYou should know this paper is a genuinely useful formative study of accessible qualitative visualization for BLV researchers, and the body is noticeably more careful than its abstract. The system contribution is real: QUARTZ provides screen-reader-accessible multimodal representations for concept maps, network graphs, Sankey diagrams, and coding stripes—four types that are not just quantitative charts relabeled. The empirical material, 12 tasks, 8 participants, a detailed change log, and a per-task outcome table that includes failures, gives the accessible-visualization community its first substantive look at how BLV users actually interact with these structures.\n\nThe paper is also honest about its own limits. Section 6.5.5 explicitly says comparisons across sessions reflect system improvement plus individual differences, not controlled experimental evidence. Section 7.4 says the same and calls for a fixed-system summative evaluation. The positionality statement in Section 3.1 flags the first author's dual role as developer/facilitator/coder and the risk of underweighting negative feedback. That is more candor than most HCI papers manage.\n\nThe soft spots are real but mostly acknowledged. The RITE design confounds system version with participant and AT configuration, as the stress test correctly notes: P5's ZDSR issues, P7's zoom and mouse modality, P4's skipped network tasks, P8's high intervention count on the newest build. The authors concede the trajectory is not a controlled result. The one place the paper transgresses is the abstract: \"document how iterative co-design with BLV users resolved them\" and the conclusion's \"our evidence indicates they can\" are stronger than the data warrant. The evidence supports a promising design trajectory, not a validated resolution. Also, the claim that qualitative visualizations \"received no attention\" is too strong given TADA on node-link diagrams, which the paper itself cites.\n\nThe lack of released code or data is a genuine reproducibility gap, especially since the design guidelines are the main contribution. The appendix documents the rule-based detection weights, which helps, but the system itself is not available.\n\nBottom line: this paper should go to peer review. A serious referee will want the abstract matched to the body, the dual-role bias mitigation made more transparent, and artifacts released. The work fills a real gap and will be cited.","headline":"A solid formative RITE study of accessible qualitative visualization, with an honest body and an abstract that oversells the evidence.","tokens_in":31897,"tokens_out":2966,"would_cite":true,"duration_ms":27671,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"QUARTZ reports formative evidence that blind and low-vision researchers can independently navigate, comprehend, and interpret qualitative data visualizations, provided the tools are designed together with them.","keywords":["accessibility","qualitative data visualization","blind and low-vision users","multimodal interaction","screen reader","sonification","iterative design","RITE method"],"falsifier":"Run the identical 12-task protocol with the final QUARTZ build, an independent facilitator who was not involved in the system's design, and a fresh sample of BLV users spanning JAWS, NVDA, and ZDSR. If task-completion rates fall substantially relative to the paper's later sessions, or if the barrier categories reported as resolved (workspace entry, relational meaning of edges, AT-intercepted shortcuts) reappear at similar frequencies, the claim that iterative co-design resolved those barriers would be refuted. A cheaper falsifier is to re-code the existing session recordings with participant and facilitator identities masked and compare the barrier and success codes; substantive disagreement on partial completions would undermine the central claim.","tokens_in":30718,"feed_emoji":"🎧","tokens_out":7647,"duration_ms":87533,"temperature":0.7,"pith_summary":"QUARTZ is a web-based system that gives blind and low-vision (BLV) researchers screen-reader-accessible, multimodal access to four qualitative visualization types: concept maps, network graphs, Sankey diagrams, and coding stripes. The paper's central claim is that BLV researchers can independently navigate, comprehend, and derive meaning from qualitative data visualizations, but only when the tools are designed with them rather than retrofitted to be accessible after the fact. To support that claim, the authors ran a RITE-based usability study with 8 BLV participants who completed 12 tasks across the four visualization types, documenting barriers unique to qualitative visualization: non-linear navigation breakdowns, semantic comprehension gaps, and screen-reader-specific interaction conflicts. The study reports that iterative co-design between sessions progressively resolved these barriers, and it derives design guidelines for a visualization domain that previously had none. If correct, the finding extends accessible-visualization infrastructure from quantitative charts to the analytical instruments of qualitative research.","feed_headline":"Blind and low-vision researchers can use qualitative charts","feed_subtitle":"Co-designed QUARTZ converts concept maps, networks, Sankey flows, and coding stripes into multimodal, screen-reader-accessible views.","key_machinery":"The load-bearing mechanism is the pairing of a fixed multimodal representation layer with the Rapid Iterative Testing and Evaluation (RITE) method. QUARTZ renders each visualization through four synchronized non-visual channels—semantic HTML with ARIA live regions, keyboard-first navigation patterned on each visualization's topology, interactive sonification that maps structural properties (hierarchy depth, degree centrality, flow magnitude, code identity) to pitch, volume, pan, and duration, and textual summary panels—so that a single navigation action updates all active channels. RITE then treats the evaluation itself as a design intervention: after each participant session, the team reviewed task outcomes, think-aloud data, and facilitator intervention logs, implemented targeted modifications, and recorded each change with its rationale as data. This loop converts observed barriers into system changes between sessions, and the reduction of specific intervention categories across P1 through P8 is the paper's evidence that the barriers were resolved rather than merely documented.","core_discovery":"Existing accessible-visualization systems cover quantitative charts, whose ordered axes, discrete enumerable data points, and deterministic data-to-visual mappings make sequential traversal and numeric sonification tractable. QUARTZ addresses qualitative visualizations, which encode semantic relationships, have non-linear topologies, and reflect interpretive judgment. The paper's discovery is that these differences are categorical, not a matter of degree: the interaction paradigms that make bar charts and scatter plots accessible do not transfer to concept maps, network graphs, Sankey diagrams, or coding stripes. Through 12 tasks with 8 BLV participants, the study identifies three barrier families—structural navigation failures in non-linear topologies, semantic comprehension gaps around relational meaning, and assistive-technology conflicts such as JAWS virtual-cursor interception—and shows that targeted RITE modifications (declared entry affordances, directional and weighted connection announcements, replacing global shortcuts with labeled buttons) reduced those specific barriers in later sessions. The authors present this as formative evidence that BLV researchers can independently navigate, comprehend, and derive meaning from qualitative visualizations, while explicitly noting that eight sessions did not saturate the barrier space and that summative confirmation is future work.","pith_inferences":["If the central claim generalizes, the same co-design-plus-topology-aware-modality recipe could extend to other qualitative displays the paper did not test, such as affinity diagrams, thematic maps, and journey maps; the paper itself flags these as open territory.","The epistemic-verification barrier—users cannot confirm that the non-visual representation matches what a sighted collaborator would see—suggests that fully independent use may require trust-building mechanisms such as inspectable rubrics, previewable exemplars, or collaborative verification, none of which the evaluation directly tests.","Because the system changed between sessions and the barrier space was not saturated after eight participants, a fixed-system summative study with a larger, stratified sample is the natural next test; the design-trajectory evidence cannot by itself establish stable usability.","The rule-based recommendation layer is deliberately free of generative models, so extending QUARTZ toward raw-text authoring would require a coding pass first; a testable extension would be an in-tool coding interface, which the paper lists as future work."],"forward_implications":["Accessibility solutions for quantitative charts cannot be assumed to carry over to qualitative visualizations; each qualitative type needs its own navigation model, such as tree traversal for concept maps, stage-grid movement for Sankey diagrams, connection-list browsing for network graphs, and segment-level focus for coding stripes.","Screen-reader support for graph structures must encode edge direction and relative importance, for example through outgoing and incoming connection lists ordered by weight, rather than rely on node-and-edge enumeration alone.","Sonification in this domain works best when it conveys shape and relational structure, such as flow width through pitch, rather than only numerical magnitude.","Replacing global keyboard shortcuts that collide with screen-reader virtual cursors with persistent labeled buttons removes a whole class of assistive-technology conflicts for the configurations tested.","Design guidelines for accessible qualitative visualization should include declared entry affordances, scaffolded vocabulary, explainable quality metrics, and a balance between recommendation guidance and researcher agency.","The three barrier families identified here are distinct from quantitative-chart accessibility barriers, so future qualitative-visualization accessibility work should expect to design for non-linear navigation and semantic comprehension rather than numeric lookup and axis traversal."],"supporting_citations":[{"why":"Documents that data analysis and visualization are the hardest research stage for BLV researchers and that many delegate visual evaluation; establishes the problem QUARTZ addresses.","marker":"[23]"},{"why":"Defines the Rapid Iterative Testing and Evaluation method used for the user study and for treating system changes between sessions as data.","marker":"[33]"},{"why":"Describes qualitative data displays and the inaccessible visual affordances of major QDA tools; grounds the coding-stripes and concept-map conventions.","marker":"[36]"},{"why":"Provides the main baseline accessible multimodal system for quantitative charts, including task taxonomy and assistive-technology configurations used for comparison.","marker":"[45]"},{"why":"Introduces stream-graph-style qualitative visualization used to ground the Sankey diagram representation.","marker":"[37]"},{"why":"Provides the network-analysis framework for representing thematic relationships as graphs, grounding QUARTZ's network graph view.","marker":"[42]"},{"why":"Supplies an extensible screen-reader visualization library that QUARTZ references for design patterns and as prior evidence that BLV users can interpret visualizations with appropriate representation.","marker":"[4]"},{"why":"Describes accessible structured editing of multimodal data representations, the approach QUARTZ extends toward qualitative visualization authoring.","marker":"[60]"}],"fun_headline_variants":["QUARTZ makes qualitative visualizations accessible to BLV researchers","Qualitative data viz gets screen-reader access via QUARTZ","For BLV researchers: QUARTZ makes concept maps and networks audible","QUARTZ co-design breaks BLV barriers in qualitative data visualization","Screen-reader access for qualitative charts: QUARTZ study shows how"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim depends on the assumption that the researcher who built QUARTZ, facilitated every session, logged every facilitator intervention, and performed the initial thematic coding could judge task success and barrier resolution without bias; if that judgment systematically favored success, the evidence that design iterations resolved the barriers would not hold.","fun_headline_variants_meta":{"raw":{"variants":["QUARTZ makes qualitative visualizations accessible to BLV researchers","Qualitative data viz gets screen-reader access via QUARTZ","For BLV researchers: QUARTZ makes concept maps and networks audible","QUARTZ co-design breaks BLV barriers in qualitative data visualization","Screen-reader access for qualitative charts: QUARTZ study shows how"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000627,"raw_usage":{"total_tokens":2906,"prompt_tokens":954,"completion_tokens":1952,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":570,"completion_tokens_details":{"reasoning_tokens":1860}},"tokens_in":570,"tokens_out":1952,"duration_ms":14293,"temperature":1.0,"reasoning_tokens":1860,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:12:26.336892+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the identical 12-task protocol with the final QUARTZ build, an independent facilitator who was not involved in the system's design, and a fresh sample of BLV users spanning JAWS, NVDA, and ZDSR. If task-completion rates fall substantially relative to the paper's later sessions, or if the barrier categories reported as resolved (workspace entry, relational meaning of edges, AT-intercepted shortcuts) reappear at similar frequencies, the claim that iterative co-design resolved those barriers would be refuted. A cheaper falsifier is to re-code the existing session recordings with participant and facilitator identities masked and compare the barrier and success codes; substantive disagreement on partial completions would undermine the central claim.","supporting_citations":[{"cited_title":"and Wixon, Dennis and McGee, Mick and Welsh, Dan","cited_arxiv_id":null,"evidence_quote":"Defines the Rapid Iterative Testing and Evaluation method used for the user study and for treating system changes between sessions as data."},{"cited_title":"Miles, A","cited_arxiv_id":null,"evidence_quote":"Describes qualitative data displays and the inaccessible visual affordances of major QDA tools; grounds the coding-stripes and concept-map conventions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces stream-graph-style qualitative visualization used to ground the Sankey diagram representation."},{"cited_title":"Pokorny, Alex Norman, Anthony P","cited_arxiv_id":null,"evidence_quote":"Provides the network-analysis framework for representing thematic relationships as graphs, grounding QUARTZ's network graph view."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies an extensible screen-reader visualization library that QUARTZ references for design patterns and as prior evidence that BLV users can interpret visualizations with appropriate representation."}],"review_version":1}