{"id":"684c81af-b36e-432c-b3fa-df9ea083bf1e","arxiv_id":"2607.14588","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A co-designed system called Graphy shows that blind users read charts primarily by touch and use a voice agent for calculations, verifying answers by tracing the tactile chart.","lead":"This paper co-designed Graphy with three blind co-designers: a tactile pin display plus a conversational AI agent that lets users explore charts by touch and ask questions by voice. It codifies design patterns—layered chart introduction, a feedback grammar, and a select-confirm-ask-verify loop—that could guide future accessible data tools.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Researcher-mediated filtering leaves a Wizard-of-Oz component in the 'fully instantiated' system, undercutting the central claim.","rationale":"Good-faith read: The paper is a valuable co-design study with detailed reporting. The authors explicitly acknowledge small sample size and limited generalizability in Section 6.6, which the Reader identified as the weakest assumption. My review, however, found a more concrete, internally verifiable problem: the 'fully instantiated' system is not fully instantiated because filtering is carried out by a researcher. This matters because the paper's framing repeatedly contrasts Graphy with the prior WOz study (Section 1, Section 6.3) and claims to 'reveal challenges WOz abstracted away.' Keeping a human in the loop for filtering means at least one interaction modality remains unvalidated as a system capability. The co-design finding that users filter by voice and then continue exploring is still valuable as a design target, but it is not evidence that the LLM-driven agent can perform the filtering reliably. The reader's generalizability concern is real but widely acknowledged; the filtering issue is not flagged as a limitation and is only detectable from a figure caption and a passing sentence. This is a missing support for the central claim. I recommend the paper remain conditional: authors should either automate filtering, re-run the affected activities, or explicitly scope the contribution to exclude filtering and revise the 'fully instantiated' language. This does not change the overall verdict (conditional) but adds a concrete condition to address. Hence verdict_should_be = UNCHANGED.","tokens_in":22096,"tokens_out":7452,"duration_ms":76237,"concrete_test":"Audit the open-source repository (github.com/accessible-data-vis/feelogue) for the filter command path: verify that a spoken 'Filter everything apart from X' triggers an automated change in the RTD renderer without human intervention. If the code lacks such a path, the 'fully instantiated' claim is false and the paper must be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central contribution rests on Graphy being a fully instantiated CTDI that moves beyond the authors' prior Wizard-of-Oz study (Section 1, Section 4). However, filtering - presented in Section 3 as a core system capability ('Filtering for focus. Users can isolate individual data series through voice commands to the agent') - was actually executed by a human. Figure 3's caption labels 'researcher-mediated filtering', and Section 5.6.1 states 'filtering commands were mediated by the researcher behind the scenes.' This is not acknowledged in Section 6.6's limitations. Because filtering is one of the five design areas and feeds the recommendation 'Provide complementary navigation mechanisms... Filtering can further aid exploration', the evidence for that recommendation, and for the division of labor between touch and agent, is partly derived from wizard behavior - exactly the kind of idealized help the paper claims to have eliminated (Section 1). The 'fully functioning implementation' claim is therefore overstated, and the observed filtering behavior cannot be attributed to the LLM agent. A reader cannot tell from the paper which other interactions, if any, were human-supported; the only explicit note is a figure caption and a single sentence. This is a missing limitation statement that should be surfaced prominently.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents Graphy, a conversational tactile data interface (CTDI) combining a refreshable tactile display (Dot Pad) with an LLM-powered conversational agent. Over four co-design workshops with three congenitally blind co-designers across eight months, the authors iteratively refined an initial prototype into a system supporting layered presentation of chart components, touch-driven point selection with audio/Braille/tactile feedback, deictic conversational queries, segmented agent responses with animated tactile highlights, and filtering of data series. The reported design knowledge is that touch is the primary sensemaking channel, the agent is reserved for computation and analysis that touch cannot resolve, and users verify agent responses by tracing the tactile chart. Additional contributions include a tactile feedback grammar distinguishing user- and agent-initiated highlights, and a select-confirm-ask-verify interaction pattern. The paper claims Graphy is the first CTDI and a 'fully instantiated' system that moves beyond the authors' prior Wizard-of-Oz study.","tokens_in":22363,"tokens_out":4809,"duration_ms":51541,"significance":"If the findings hold, this is a meaningful step for accessible data visualization: it demonstrates a concrete multimodal interaction paradigm and derives actionable design recommendations grounded in long-term co-design with BLV users. The longitudinal method, the inclusion of co-designers as co-authors, and the availability of open-source code are notable strengths. The authors are appropriately transparent that the study is formative rather than evaluative and that the small, congenitally blind, tactile-experienced sample limits generalizability. However, a load-bearing gap exists: filtering—presented in §3 as a core agent-driven capability and included in the design recommendations of §6.4—was actually executed by a human researcher behind the scenes (§5.6.1, Fig. 3). This partially undercuts the 'fully instantiated'/'beyond the wizard' claims and weakens the evidence for the filtering recommendation specifically.","major_comments":[{"comment":"Filtering is presented as a system capability ('Filtering for focus. Users can isolate individual data series through voice commands to the agent') and as a design recommendation ('Filtering can further aid exploration'). However, §5.6.1 states that 'filtering commands were mediated by the researcher behind the scenes,' and Fig. 3 marks filtering as 'researcher-mediated.' This is a Wizard-of-Oz component in an otherwise instantiated system, and it is not acknowledged in §6.6. It also conflicts with the claim in §5.6.2 that co-designers used 'the full range of interactions unassisted.' The evidence for the filtering recommendation is therefore partly based on researcher behavior rather than the agent's behavior. The authors should either implement filtering in the system or explicitly present it as a simulated capability, and revise the 'fully instantiated' claim, the WOz comparison in §6","section":"§3, §5.6.1, Fig. 3, §6.4, §6.6"}],"minor_comments":[{"comment":"The legend for the four state symbols (suggested, not yet implemented, implemented/refined, carried forward) is not visible in the provided version; please ensure the symbols render correctly.","section":"Table 2"},{"comment":"The caption 'T X Y D1 S…D2' is cryptic. Spell out the layer labels (Title, X-axis, Y-axis, Data series, Summary) for readability.","section":"Figure 3 caption"},{"comment":"The sentence 'By WS4, the co-designers used the full range of interactions unassisted' conflicts with the disclosure in §5.6.1 that filtering was researcher-mediated. Qualify the claim or exclude filtering from the 'full range' description.","section":"§5.6.2"}],"recommendation":"major_revision","confidential_remarks":"The central concern is the researcher-mediated filtering. The authors can resolve this by implementing filtering or by reframing it as a simulated capability and adjusting the 'fully instantiated' and 'unassisted' claims. The other findings—layered presentation, touch-first division of labor, and select-confirm-ask-verify—rest on interactions other than filtering and appear adequately supported by the qualitative data and interaction counts."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: this paper is a solid, honest co-design study that contributes working design patterns for combining an RTD with a conversational agent, but it overstates its 'fully instantiated' claim. The filtering capability, presented in Section 3 as a core voice-command feature, was actually executed by the researcher behind the scenes. The authors disclose this in a figure caption and one sentence in 5.6.1, but it doesn't make it into the limitations, and the abstract and intro say 'fully functioning implementation' without the caveat. That's a real discrepancy, not a nitpick: filtering is one of the five design areas and it feeds a design recommendation, so part of the observed behavior comes from a wizard, exactly what the paper says it moved beyond. The good news is that this doesn't undercut the whole study. The core patterns—touch as primary sensemaking, agent for calculations, verification by touch—are grounded in plenty of other interactions and in the co-designers' reflections. The layered presentation, the feedback grammar, and the select-confirm-ask-verify pattern are coherent and supported by quotes and interaction counts. The authors are also clear that this is formative, not a controlled evaluation, and they flag the small sample and possible lack of generalizability to acquired blindness or less tactile experience. Those are real limits but not flaws. The open-source code is a plus. The main revision point is to make the filtering limitation prominent and to qualify the 'fully instantiated' language in the abstract and Section 1. Also, it would help to say explicitly which other interaction components were fully system-driven; right now the reader has to hunt for the one disclosure. This paper is for researchers in accessible visualization and RTD design. It's a meaningful contribution to an emerging subfield, and the design patterns will be cited. Send it to peer review; it needs revisions, but the work deserves referee time.","headline":"Genuine co-design contribution with real design patterns, but the 'fully instantiated' claim needs a caveat: filtering was researcher-mediated.","tokens_in":22850,"tokens_out":2597,"would_cite":true,"duration_ms":24516,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Graphy, the first conversational tactile data interface, lets blind users explore charts by touch and ask an AI agent only for what touch cannot resolve.","keywords":["accessible data visualization","refreshable tactile display","conversational AI","co-design","blind and low vision","multimodal interaction","tactile feedback","data sensemaking"],"falsifier":"Run the WS4 free-form data-summary task with a group of blind or low-vision users who have acquired blindness and no prior RTD or tactile-graphics training, and count the order and proportion of touch explorations versus agent queries, plus whether they touch-verify agent answers. If these users default to agent-first queries and skip verification touches, the claimed touch-primary division of labor is specific to experienced tactile readers, not to blind users generally.","tokens_in":22025,"feed_emoji":"🖐️","tokens_out":5156,"duration_ms":56964,"temperature":0.7,"pith_summary":"This paper works out how to combine a refreshable tactile display — a pin-based screen that renders charts you can feel — with a conversational AI assistant, so that people who are blind or have low vision can explore data on their own. The authors co-designed a working system, Graphy, with three blind co-designers over four workshop rounds and eight months. Their central finding is a division of labor: users make sense of a chart's shape, trends, and relationships through touch first, ask the AI only for calculations and analysis that touch cannot provide, and then return to the chart to verify the AI's answer. The paper contributes three reusable design principles — layered presentation, a tactile feedback grammar, and a select-confirm-ask-verify interaction pattern — as validated starting points for this new class of interface.","feed_headline":"Blind chart readers lead with touch, then verify AI by feel","feed_subtitle":"In Graphy, users touch the chart to understand it, ask the AI for analysis, then touch again to verify.","key_machinery":"The central mechanism is the coupling of a refreshable tactile display (an RTD, a pin-based display that renders tactile graphics dynamically) with an LLM-backed conversational agent, bridged by deictic queries that refer to touch-selected data points and by tactile highlighting that lets the chart itself confirm or contradict the agent's speech. The RTD provides a persistent spatial representation of the data; the agent provides calculation and analytical depth; and a feedback grammar — static highlights for user selections, animated highlights for agent references, transitional animations for stepping — tells users who initiated the feedback so they can decide when to trust and when to ver","core_discovery":"Graphy is, to the authors' knowledge, the first conversational tactile data interface: a system combining a multi-line refreshable tactile display with an LLM-powered conversational agent, built through co-design with three blind co-designers. The central discovery is that the two modalities take distinct, complementary roles: touch is the primary sensemaking channel for spatial understanding of the data's shape, trends, and relationships; the conversational agent is reserved for what touch cannot resolve, such as calculation and analysis; and the chart on the tactile display is used to verify the agent's responses. The authors also claim three transferable design findings: a layered present","pith_inferences":["Beyond the paper: the touch-first, AI-second division of labor resembles a pattern seen in sighted users with high data literacy, who treat LLMs as specialists rather than crutches; a testable extension is to adapt CTDIs to the user's tactile and data experience rather than assuming one interaction style.","Beyond the paper: layered presentation could transfer to other tactile media — maps, diagrams, STEM graphics, and densely sampled charts — where isolating components may be the only way to keep them readable; the co-designers hinted at this, but the paper does not test it.","Beyond the paper: the verify step offers a natural, unobtrusive signal for measuring user trust in AI answers — tracing behavior after an agent response could serve as a proxy for whether the user found the answer credible.","Beyond the paper: if LLM accuracy were somehow guaranteed, users say they would still explore by touch, suggesting the RTD's value is spatial understanding, not just error checking; one could test whether touch-first interaction remains preferred even when the agent is provably correct."],"forward_implications":["If the pattern holds, a blind user can independently inspect a chart, select points, ask for calculations, and confirm the answer without sighted help.","Designers of accessible data tools should treat the tactile chart as the ground truth for verification, because co-designers consistently returned to touch to check agent responses.","Agent responses should be concise, answer-first, segmented into one-sentence chunks, and synchronized with animated tactile highlights so that speech does not compete with touch exploration.","The select-confirm-ask-verify pattern provides a concrete, reusable interaction sequence for future conversational tactile data interfaces.","Future refreshable displays should support multi-height pins, built-in multi-touch sensing, and reliable actuation under the finger to expand what these interfaces can encode."],"fun_headline_variants":["Touch leads, AI assists, tactile confirms","Blind users: touch to read, AI to analyze, touch to verify","Co-designing tactile data interfaces with AI for blind users","For blind users, touch first, AI second, verify by feel"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The observed touch-first pattern, and the design principles built on it, come from three congenitally blind co-designers who were comfortable with tactile graphics and RTDs; if those patterns do not hold for blind users with acquired blindness, less tactile experience, or residual vision, the design recommendations lose much of their force.","fun_headline_variants_meta":{"raw":{"variants":["Touch leads, AI assists, tactile confirms","Blind users: touch to read, AI to analyze, touch to verify","Co-designing tactile data interfaces with AI for blind users","For blind users, touch first, AI second, verify by feel"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000393,"raw_usage":{"total_tokens":1903,"prompt_tokens":747,"completion_tokens":1156,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":491,"completion_tokens_details":{"reasoning_tokens":1086}},"tokens_in":491,"tokens_out":1156,"duration_ms":12379,"temperature":1.0,"reasoning_tokens":1086,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T01:38:28.936039+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the WS4 free-form data-summary task with a group of blind or low-vision users who have acquired blindness and no prior RTD or tactile-graphics training, and count the order and proportion of touch explorations versus agent queries, plus whether they touch-verify agent answers. If these users default to agent-first queries and skip verification touches, the claimed touch-primary division of labor is specific to experienced tactile readers, not to blind users generally.","supporting_citations":[],"review_version":1}