Pith. sign in

REVIEW 4 major objections 4 minor 59 references

Querying Multimodal Scientific Papers with AI: Practices and Preferences Across Blind, Low-Vision, and Sighted Scientists

T0 review · 4 major / 4 minor · reviewed 2026-08-01 · deepseek-v4-flash

Pith's one-line read Blind and low-vision scientists rely on AI to read figures and tables in papers, but vague or wrong descriptions drive both them and sighted scientists away.

desk verdict A useful empirical study of real scientist–AI queries; the qualitative findings hold, but the headline percentages don't survive a re-count from the paper's own appendix. read the letter →

arxiv 2607.18514 v1 pith:SRBFBGRL submitted 2026-07-20 cs.HC cs.AI

classification cs.HCcs.AI
keywords blindandlow-visionaccessibilitymultimodalquestionansweringscientificpapersfigurestablesAItrustworkaroundsqueryclassificationuserstudy
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish how blind or low-vision (BLV) scientists, compared with sighted scientists, actually use AI question-answering to read multimodal scientific papers, and where current AI falls short. The authors report that BLV participants directed 49% of their queries at figures and tables, almost twice the 29% share from sighted participants, who instead spent 56% of queries on methods and findings. Both groups followed a common funnel strategy, starting with broad overviews before narrowing to specifics. The central claim is that AI-generated descriptions of figures are often too vague, incomplete, or wrong for scientific work, and that this failure is harsher for BLV scientists, who cannot validate answers against the original PDF, an inaccessible document. The paper also contributes a dataset of 115 real queries and responses to ground future systems and benchmarks.

What carries the argument

The study's load-bearing mechanism is a four-way query taxonomy — General/Overview, Figures/Tables, Methods/Findings, and Miscellaneous — applied to the 115 logged queries, which reveals sharply different distributions for BLV versus sighted users. The finding is explained by two roles: an 'access layer' for BLV scientists who cannot see the visual content, and an 'efficiency layer' for sighted scientists who want faster synthesis. The critical design insight is the observed 'validation asymmetry': sighted participants could return to the original paper to check an AI answer, while BLV participants had no accessible source to verify against, amplifying the cost of any hallucination or vague

What would settle it

Log real, unaided AI paper-QA sessions from BLV and sighted scientists — without any instruction to ask about figures or tables — and repeat the measurement after a model update; if the figure/table query split narrows to near equality or the 'vague descriptions' trust breakdown disappears, the paper's central claims would be features of the task design and model snapshot rather than of the underlying access gap.

Watch

Extended reading notes

Core claim

Based on semi-structured interviews with five BLV and five sighted scientists across STEM fields, the study watched participants query two commercial AI tools, ChatGPT and Gemini, against familiar and unfamiliar papers from their own domains, logging 115 queries. The discovery is that the systems play different roles for the two groups: for BLV scientists they are a primary access bridge to visual content, while for sighted scientists they are an efficiency layer for synthesis and verification. Across both groups, AI responses that described surface visual features (colors, line styles, chart types) without scientific meaning (axes, units, spatial locations, quantitative trends) led to loss

Load-bearing premise

The study assumes that the query behavior observed in the 90-minute sessions — including the instruction to make at least one query about a figure or table and the choice of two specific AI tools at one point in time — represents how scientists actually use AI paper-QA tools in their everyday work.

Editorial extensions

If this is right

  • QA systems for scientific papers should prioritize interpretive, domain-aware figure descriptions — axes, units, trends, spatial references — over literal visual semantics like colors and line styles.
  • AI tools should support verification without the original PDF, for example by citing exact figure panels or quoting source passages, because BLV scientists cannot independently open the paper.
  • Systems should scaffold the observed 'funnel' strategy, returning an overview first and then offering progressive detail, so users do not have to manage the structure themselves.
  • Transparency about whether a response is quoted, paraphrased, or inferred would help all scientists judge trustworthiness, a need that is acute for BLV users who cannot easily check hallucinations.
  • The released dataset of 115 scientist-authored queries and responses offers a grounding resource for benchmarking multimodal QA systems in real user practice rather than synthetic test sets.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • We note that the study explicitly asked participants to make at least one figure/table query, which likely inflates the figure/table share in both groups; a naturalistic log study without such a prompt might show a smaller gap, though the opening interviews suggest the direction is real.
  • The 'vague descriptions' failure is a snapshot of two specific AI assistants at one point in time; newer models may shift the failure mode from vagueness to overconfident but wrong interpretations, which would change the design priorities.
  • One testable extension beyond the paper: QA systems that return answers with inline source citations (e.g., figure coordinates or quoted sentences) should differentially improve BLV scientists' trust and comprehension, and that improvement could be measured directly.
  • The small sample (five BLV participants, all PhD students) limits the quantitative split; recruiting more senior BLV scientists could reveal different workarounds or an even stronger reliance on AI for visual access.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. This paper reports a qualitative study of ten scientists (five blind/low-vision, five sighted) using ChatGPT and Gemini to query multimodal scientific papers. It categorizes 115 queries and finds that BLV participants disproportionately ask about figures and tables (49%) while sighted participants focus on methods and findings (56%), with both groups following a shared "funnel" strategy. It also documents how vague or incorrect AI descriptions erode trust, and argues this is especially consequential for BLV scientists who cannot easily validate outputs against the source document. The paper contributes a dataset of 115 queries and responses, along with design implications for inclusive QA systems. The qualitative material is rich and transparently analyzed, but the quantitative headline contains internal inconsistencies and is vulnerable to task-design and sampling confounds.

Significance. If the findings stand, this is a valuable user-centered complement to benchmark-driven multimodal QA research. The dataset of 115 scientist-authored queries is a concrete resource for grounding future evaluation in real practices, and the thematic findings—domain-aware interpretation over surface description, progressive scaffolding, transparency about synthesis vs. quotation—are actionable. The study's methodological transparency (inductive/deductive coding, peer debriefing, affinity diagramming, candid limitations) is a strength. However, the central quantitative contrast in Fig. 1 currently rests on incorrect or unclearly denominated percentages and on an explicit study instruction that may inflate the very category being compared, so the headline claims need correction and scoping before they can be taken at face value.

major comments (4)
  1. [§4.2, Fig. 1, Appendix A] The reported sighted figure/table percentage (29%) and methods/findings percentage (56%) do not match Appendix A. Counting Table 2 gives sighted figure/table queries as 20/76 ≈ 26% and methods/findings as 46/76 ≈ 61% (the BLV 19/39 ≈ 49% is correct). If a different denominator was used (e.g., excluding General/Overview), the paper must state this; otherwise the percentages in the abstract and Figure 1 are inaccurate. The "almost double" contrast is qualitatively preserved under either count, but the specific numbers need correction.
  2. [§3.2, Fig. 1, §4.2] The instruction to "make at least one query about a figure or table" places a floor on the category being compared. The manuscript neither reports how many queries were attributable to this instruction nor gives per-participant query counts. Since BLV participants contributed 39 queries and sighted participants 76, even one mandated query per BLV participant accounts for ~13% of BLV queries versus ~6.6% of sighted queries, potentially inflating the 49% figure. Please report sensitivity analyses excluding the mandated queries, provide per-participant breakdowns, and soften the claim that QA systems function as "a primary access bridge" to visual content.
  3. [Table 1, §5.3] All five BLV participants are PhD students, whereas the sighted group includes industry researchers, a clinical researcher, and a program manager. The abstract and introduction generalize to "BLV scientists," but the observed query-type differences may partly reflect career stage or job tasks rather than vision status. §5.3 acknowledges the homogeneous BLV sample but does not connect it to the Fig. 1 contrast. Please scope the comparative claims to the studied population or provide evidence that role and experience do not confound the distribution.
  4. [§3.2, §4.3, §5.3] The trust/abandonment findings are observed with two specific model versions (GPT-5 and Gemini-2.5) at one point in time. The discussion and abstract sometimes refer to "AI systems" generically (e.g., §5.2). Please add an explicit limitation that the identified failure modes may be version-specific and that rapid model improvements could change the trust dynamics, rather than implying a timeless property of the paradigm.
minor comments (4)
  1. [Fig. 1 caption] State the within-group denominators (n=39 BLV, n=76 sighted) and clarify whether the percentages are out of all queries or a subset.
  2. [§4.2] Typo: "BLV1 noted that is responses become too cumbersome" should read "if responses become too cumbersome."
  3. [Appendix A] Add a column indicating which queries were explicitly mandated by the study instruction; this would directly address the sensitivity concern in Major Comment 2.
  4. [Abstract] The phrase "current AI systems" is time-sensitive; consider dating the observation or using "the systems studied here."

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: empirical interview study with self-contained findings; procedural constraints are validity concerns, not circular reductions.

full rationale

This paper is a qualitative interview study and does not claim a derivation, prediction, or first-principles result. The headline percentages (49% figures/tables for BLV participants, 56% methods/findings for sighted participants, Fig. 1, §4.2) are descriptive tabulations of the 115 logged queries, not parameters fitted to make a claim true. The procedure did instruct participants to "make at least one query about a figure or table" (§3.2), which could bias the marginal query-type distribution and weaken the 'primary access bridge' reading; however, that is a study-design validity threat, not a circular reduction—the counts are observed behavior, not defined by the claim. The paper's self-citations (GeoVisA11y [27], FigurA11y [45], SciA11y [51], and others) appear only as related work and as motivation for studying QA; none is invoked as an external theorem or as evidence for the new empirical findings. The paper also discloses relevant limitations in §5.3 (small sample, all BLV participants being PhD students, self-selected papers), which affects generalizability but does not create circularity. No step reduces to its own inputs by construction, and no fitted parameter is renamed as a prediction. Accordingly, the paper is self-contained against its own evidence base, and the circularity score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

This is an empirical qualitative study: there are no fitted mathematical parameters and no invented entities. The free-parameter count is zero in the model sense; the four-type query taxonomy is a hand-built coding scheme that determines the headline percentages, and its arbitrariness is captured under the coding-stability axiom. The load-bearing premises are methodological: observed session behavior represents real workflows, the two tested chatbots represent the AI QA paradigm, and the coding process is reliable.

assumptions (3)
  • domain assumption Self-report and think-aloud narration during a timed, task-structured session reflect how participants actually use AI QA systems in their real workflows.
    The findings in §4 rely on participants' descriptions of past workflows plus think-aloud behavior during 90-minute sessions. The mandated figure/table query and the unfamiliar-paper task may elicit behavior that differs from natural use.
  • domain assumption The behavior of ChatGPT (GPT-5) and Gemini (Gemini-2.5) at the time of the study is representative of AI-powered multimodal QA systems and their shortcomings.
    The tools are chosen 'at the time of the studies' (§3.2). Findings about vague descriptions and hallucinations could be specific to these model versions, which change on short timescales.
  • domain assumption The thematic analysis coding (18 codes → 8 themes → 3 higher-level themes) is stable enough to support the reported group differences.
    Per §3.3, the first author did primary coding; one external researcher reviewed only two of ten transcripts; no inter-rater reliability statistics are reported. The 49% vs. 29% and 56% percentages depend on the query categorization.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Querying Multimodal Scientific Papers with AI: Practices and Preferences Across Blind, Low-Vision, and Sighted Scientists." pith.science (2026). https://pith.science/paper/SRBFBGRL

@misc{pith2026260718514,
  author       = {Pith},
  title        = {Pith review of: Querying Multimodal Scientific Papers with AI: Practices and Preferences Across Blind, Low-Vision, and Sighted Scientists},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/SRBFBGRL}},
  note         = {Machine review of arXiv:2607.18514}
}
read the original abstract

Visual diagrams, figures, and tables are central to scientific papers, and convey information beyond what is captured in text. While blind or low-vision (BLV) scientists have traditionally relied on static alternative text to access figures in papers, the rise of artificial intelligence (AI) has made interactive question-answering (QA) a feasible paradigm for visual exploration; yet little is known about how scientists use visual QA in practice or how to improve its accessibility. In this work, we interview five BLV and five sighted scientists across different STEM fields to understand how they use two AI tools, ChatGPT and Gemini, to query multimodal scientific documents. Our findings characterize how scientists review multimodal content, including existing practices (along with accessibility workarounds) for engaging with visuals, and feedback on the suitability of AI-generated responses to multimodal queries. We further find that vague or incomplete image descriptions, as well as incorrect AI outputs more broadly, can cause both BLV and sighted scientists to abandon AI workflows. To support future research, we additionally contribute a dataset of 115 queries and responses from our participants' interactions with the AI tools for papers in their field. We close by discussing implications for AI-powered scientific QA systems, emphasizing considerations for access across abilities and domains.

Figures

Figures reproduced from arXiv: 2607.18514 by the authors.

Figure 1
Figure 1. We found three main categories of queries participants asked using QA systems—(1) General or Overview, (2) Figures [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. An example of a complex, multi-part figure from a neuroscience paper, “Low-dimensional criticality embedded in [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Ambiguity in AI descriptions of charts: the description for Figure 9 from the paper “Typing Haptically: Towards [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: AI descriptions provided visual information, but did not help with interpretation: BLV3 found no use from the [PITH_FULL_IMAGE:figures/full_fig_p009_4.png]
Figure 5
Figure 5. Figure 5: AI descriptions providing visual information but no interpretation also frustrated sighted participants. S2 expected [PITH_FULL_IMAGE:figures/full_fig_p010_5.png]
Figure 3
Figure 3. Figure 3: Include the origi￾nal caption and alt text (if any) and provide an overall description of the figure. 2 Familiar Gemini Typing Haptically: Towards Enabling Non￾auditory Smartphone Text Entry with Haptic Feedback for Blind and Low Vision Users http://dl.acm.org/doi/10. …
Figure 3
Figure 3. Figure 3: figure 3. Turnicate jargon [PITH_FULL_IMAGE:figures/full_fig_p022_3.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

59 extracted references · 10 canonical work pages

  1. [1]

    Introducing: Be My AI

    2024. Introducing: Be My AI. https://www.bemyeyes.com/blog/introducing-be- my-ai

  2. [2]

    Seeing AI - Talking Camera for the Blind

    2026. Seeing AI - Talking Camera for the Blind. https://www.seeingai.com/ ASSETS ’26, October 25–28, 2026, Vila Nova de Gaia, Portugal A. Chheda-Kothary, L. L. Wang, J. C. Chang, and J. Bragg

  3. [3]

    Storer, Antony Rishin Mukkath Roy, Ravi Kuber, and Stacy M

    Ali Abdolrahmani, Kevin M. Storer, Antony Rishin Mukkath Roy, Ravi Kuber, and Stacy M. Branham. 2020. Blind Leading the Sighted: Drawing Design Insights from Blind Users towards More Productivity-oriented Voice Interfaces.ACM Trans. Access. Comput.12, 4, Article 18 (Jan. 2020), 35 pages. doi:10.1145/3368426

  4. [4]

    Rahaf Alharbi, Pa Lor, Jaylin Herskovitz, Sarita Schoenebeck, and Robin N. Brewer

  5. [5]

    alphaXiv. 2026. alphaXiv: Explore. https://www.alphaxiv.org/

  6. [6]

    Anthropic. 2026. The AI for Problem Solvers | Claude by Anthropic. https: //claude.com/product/overview

  7. [7]

    1998.Contextual Design: Defining Customer- Centered Systems

    Hugh Beyer and Karen Holtzblatt. 1998.Contextual Design: Defining Customer- Centered Systems. Morgan Kaufmann, San Francisco, CA

  8. [8]

    Angela Glover Blackwell. 2016. The Curb-Cut Effect.Stanford Social Innovation Review15, 1 (2016), 28–33. doi:10.48558/YVMS-CC96

Show all 59 references
  1. [9]

    Virginia Braun and Victoria Clarke. 2006. Using thematic analysis in psychology. Qualitative Research in Psychology3, 2 (2006), 77–101. doi:10.1191/1478088706qp 063oa

  2. [11]

    Sanjana Shivani Chintalapati, Jonathan Bragg, and Lucy Lu Wang. 2022. A Dataset of Alt Texts from HCI Publications: Analyses and Uses Towards Producing More Descriptive Alt Texts of Data Visualizations in Scientific Papers. InProceedings of the 24th International ACM SIGACCESS...

  3. [12]

    Adriana Dutkiewicz, Dietmar Müller, John Cannon, Sioned Vaughan, and Sabin Zahirovic. 2018. Sequestration and subduction of deep-sea carbonate in the global ocean since the Early Cretaceous.Geology(12 2018). doi:10.1130/G45424.1

  4. [13]

    Anders Ericsson and Herbert A

    K. Anders Ericsson and Herbert A. Simon. 1993.Protocol Analysis: Verbal Reports as Data. MIT Press, Cambridge, MA

  5. [14]

    María Evagorou, Sibel Erduran, and Terhi Mäntylä. 2015. The role of visual repre- sentations in scientific practices: from conceptual understanding and knowledge generation to ‘seeing’ how science works.International Journal of STEM Education 2 (2015), 1–13. https://api.semant...

  6. [15]

    Fontenele, Joshua S

    Antonio J. Fontenele, Joshua S. Sooter, Vincent K. Norman, Sidharth H. Gautam, and Woodrow L. Shew. 2024. Low-dimensional criticality embedded in high- dimensional awake brain dynamics.Science Advances10, 17 (2024), eadj9303. doi:10.1126/sciadv.adj9303

  7. [16]

    Google. 2026. Learn about Gemini, the everyday AI assistant from Google. https://gemini.google/about/

  8. [17]

    Gorlewicz, Jennifer L

    Jenna L. Gorlewicz, Jennifer L. Tennison, P. Merlin Uesbeck, Margaret E. Richard, Hari P. Palani, Andreas Stefik, Derrick W. Smith, and Nicholas A. Giudice. 2020. Design Guidelines and Recommendations for Multimodal, Touchscreen-based Graphics.ACM Trans. Access. Comput.13, 3, ...

  9. [18]

    Joshua Gorniak, Yoon Kim, Donglai Wei, and Nam Wook Kim. 2024. VizAbility: Enhancing Chart Accessibility with LLM-based Conversational Interaction. In Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology(Pittsburgh, PA, USA)(UIST ’24). Associa...

  10. [19]

    Tingying He, Maggie McCracken, Daniel Hajas, Sarah Creem-Regehr, and Alexan- der Lex. 2026. Using Tactile Charts to Support Comprehension and Learn- ing of Complex Visualizations for Blind and Low-Vision Individuals.IEEE Transactions on Visualization and Computer Graphics32, 1...

  11. [20]

    Yugo Iwamoto, Muhammad Raees, Jamison Heard, and Garreth W. Tigwell. 2025. Exploring Generative AI to Support Disability Service Professionals in Writing Image Descriptions for HCI Science Figures. InProceedings of the 27th Interna- tional ACM SIGACCESS Conference on Computers...

  12. [21]

    Madhan Jeyaraman, Naveen Jeyaraman, Swaminathan Ramasubramanian, Ab- hishek Vaish, and Raju Vaishya. 2024. Decoding Research with a Glance: The Power of Graphical Abstracts and Infographics.Apollo Medicine22 (2024), 144 –

  13. [22]

    Wan Ju Kang, Eunki Kim, Na Min An, Sangryul Kim, Ha-Young Choi, Ki Hoon Kwak, and James Thorne. 2025. Sightation Counts: Leveraging Sighted User Feedback in Building a BLV-aligned Dataset of Diagram Descriptions.ArXiv abs/2503.13369 (2025). https://api.semanticscholar.org/Corp...

  14. [23]

    Jiho Kim, Arjun Srinivasan, Nam Wook Kim, and Yea-Seul Kim. 2023. Exploring Chart Question Answering for Blind and Low Vision Users. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems(Hamburg, Germany)(CHI ’23). Association for Computing Machinery, ...

  15. [24]

    Anukriti Kumar and Lucy Lu Wang. 2024. Uncovering the New Accessibility Crisis in Scholarly PDFs: Publishing Model and Platform Changes Contribute to Declining Scholarly Document Accessibility in the Last Decade. InProceedings of the 26th International ACM SIGACCESS Conference...

  16. [25]

    West, and Bill Howe

    Po-Shen Lee, Jevin D. West, and Bill Howe. 2016. Viziometrics: Analyzing Visual Information in the Scientific Literature.IEEE Transactions on Big Data4 (2016), 117–129. https://api.semanticscholar.org/CorpusID:3665638

  17. [26]

    I never realized sidewalks were a big deal

    Chu Li, Katrina Oi Yau Ma, Michael Saugstad, Kie Fujii, Molly Delaney, Yochai Eisenberg, Delphine Labbé, Judy L Shanley, Devon Snyder, Florian P P Thomas, and Jon E. Froehlich. 2024. “I never realized sidewalks were a big deal”: A Case Study of a Community-Driven Sidewalk Acce...

  18. [27]

    Froehlich

    Chu Li, Rock Yuren Pang, Arnavi Chheda-Kothary, Ather Sharif, Henok Assalif, Jeffrey Heer, and Jon E. Froehlich. 2026. GeoVisA11y: An AI- based Geovisualization Question-Answering System for Screen-Reader Users. arXiv:2603.07446 [cs.HC] https://arxiv.org/abs/2603.07446

  19. [28]

    Chuhan Li, Ziyao Shangguan, Yilun Zhao, Deyuan Li, Yixin Liu, and Arman Cohan. 2024. M3SciQA: A Multi-Modal Multi-Document Scientific QA Bench- mark for Evaluating Foundation Models.ArXivabs/2411.04075 (2024). https: //api.semanticscholar.org/CorpusID:273850071

  20. [29]

    Alan Lundgard and Arvind Satyanarayan. 2021. Accessible Visualization via Natural Language Descriptions: A Four-Level Model of Semantic Content.IEEE Transactions on Visualization and Computer GraphicsPP (2021), 1–1. https: //api.semanticscholar.org/CorpusID:238237335

  21. [31]

    Hayes, and Devva Kasnitz

    Jennifer Mankoff, Gillian R. Hayes, and Devva Kasnitz. 2010. Disability studies as a source of critical inquiry for the field of assistive technology. InProceedings of the 12th International ACM SIGACCESS Conference on Computers and Accessibility (Orlando, Florida, USA)(ASSETS...

  22. [32]

    Nora McDonald, Sarita Schoenebeck, and Andrea Forte. 2019. Reliability and Inter-rater Reliability in Qualitative Research: Norms and Guidelines for CSCW and HCI Practice.Proc. ACM Hum.-Comput. Interact.3, CSCW, Article 72 (Nov. 2019), 23 pages. doi:10.1145/3359174

  23. [33]

    Beverley Cherie Millar and Michelle Lim. 2022. The Role of Visual Abstracts in the Dissemination of Medical Research.The Ulster Medical Journal91 (2022), 67 – 78. https://api.semanticscholar.org/CorpusID:249821792

  24. [34]

    Accessibility People, You Go Work on That Thing of Yours over There

    Sanika Moharana, Cynthia L. Bennett, Erin Buehler, Michael Madaio, Vinita Tibdewal, and Shaun K. Kane. 2025. “Accessibility People, You Go Work on That Thing of Yours over There”: Addressing Disability Inclusion in AI Product Organizations.Proceedings of the AAAI/ACM Conferenc...

  25. [35]

    Bennett, and Edward Cutrell

    Meredith Ringel Morris, Jazette Johnson, Cynthia L. Bennett, and Edward Cutrell

  26. [36]

    Peya Mowar, Aaron Steinfeld, and Jeffrey P Bigham. 2026. iTagPDF: Towards Finally Automating PDF Accessibility. InProceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI ’26). Association for Computing Machinery, New York, NY, USA, Article 1116, 17 pa...

  27. [37]

    OpenAI. 2026. ChatGPT | AI Chatbot to Discover, Learn & Create. https: //chatgpt.com/overview/

  28. [38]

    Soya Park, Jonathan Bragg, Michael Chang, Kevin Larson, and Danielle Bragg

  29. [39]

    P. A. Patel, N. P. DeGroote, K. Jackson, T. Cash, S. M. Castellino, P. Jaggi, A. J. Es- benshade, and T. P. Miller. 2022. Infectious events in pediatric patients with acute lymphoblastic leukemia/lymphoma undergoing evaluation for fever without severe neutropenia.Cancer128, 23...

  30. [40]

    Gonzalez Penuela, Crescentia Jung, Sharon Y Lin, Ruiying Hu, and Shiri Azenkot

    Ricardo E. Gonzalez Penuela, Crescentia Jung, Sharon Y Lin, Ruiying Hu, and Shiri Azenkot. 2026. How Multimodal Large Language Models Support Access to Visual Information: A Diary Study With Blind and Low Vision People. https: //api.semanticscholar.org/CorpusID:285616065

  31. [41]

    Shraman Pramanick, Rama Chellappa, and Subhashini Venugopalan. 2024. SPIQA: A Dataset for Multimodal Question Answering on Scientific Papers. InAdvances Querying Multimodal Scientific Papers with AI ASSETS ’26, October 25–28, 2026, Vila Nova de Gaia, Portugal in Neural Informa...

  32. [42]

    Quora. 2026. Poe - Fast, Helpful AI Chat. https://poe.com/

  33. [43]

    Qwen. 2026. Qwen Chat. https://chat.qwen.ai/

  34. [44]

    Wang, Alida T

    Ather Sharif, Olivia H. Wang, Alida T. Muongchan, Katharina Reinecke, and Jacob O. Wobbrock. 2022. VoxLens: Making Online Data Visualizations Accessible with an Interactive JavaScript Plug-In. InProceedings of the 2022 CHI Conference on Human Factors in Computing Systems(New O...

  35. [45]

    Nikhil Singh, Lucy Lu Wang, and Jonathan Bragg. 2024. FigurA11y: AI Assistance for Writing Scientific Alt Text. InProceedings of the 29th International Conference on Intelligent User Interfaces(Greenville, SC, USA)(IUI ’24). Association for Com- puting Machinery, New York, NY,...

  36. [46]

    Sharon Spall. 1998. Peer Debriefing in Qualitative Research: Emerging Opera- tional Models. 4, 2 (Jan 1998), 280–292. doi:10.1177/107780049800400208

  37. [47]

    Fleischmann, Meredith Ringel Morris, and Danna Gurari

    Abigale Stangl, Nitin Verma, Kenneth R. Fleischmann, Meredith Ringel Morris, and Danna Gurari. 2021. Going Beyond One-Size-Fits-All Image Descriptions to Satisfy the Information Wants of People Who are Blind or Have Low Vision. In Proceedings of the 23rd International ACM SIGA...

  38. [48]

    Xinru Tang, Ali Abdolrahmani, Darren Gergle, and Anne Marie Piper. 2025. Everyday Uncertainty: How Blind People Use GenAI Tools for Information Access. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). Association for Computing Machinery...

  39. [50]

    B. N. Walker and L. M. Mauney. 2010. Universal Design of Auditory Graphs: A Comparison of Sonification Mappings for Visually Impaired and Sighted Listeners. ACM Trans. Access. Comput.2, 3, Article 12 (March 2010), 16 pages. doi:10.1145/ 1714458.1714459

  40. [51]

    Lucy Lu Wang, Isabel Cachola, Jonathan Bragg, Evie Yu-Yen Cheng, Chelsea Haupt, Matt Latzke, Bailey Kuehl, Madeleine N van Zuylen, Linda Wagner, and Daniel Weld. 2021. SciA11y: Converting Scientific Papers to Accessible HTML. In Proceedings of the 23rd International ACM SIGACC...

  41. [52]

    Wagner, and Daniel S

    Lucy Lu Wang, Isabel Cachola, Jonathan Bragg, Evie (Yu-Yen) Cheng, Chelsea Hess Haupt, Matt Latzke, Bailey Kuehl, Madeleine van Zuylen, Linda M. Wagner, and Daniel S. Weld. 2021. Improving the Accessibility of Scientific Doc- uments: Current State, User Needs, and a System Sol...

  42. [53]

    Candace Williams, Lilian de Greef, Ed Harris, Leah Findlater, Amy Pavel, and Cynthia Bennett. 2022. Toward supporting quality alt text in computing pub- lications. InProceedings of the 19th International Web for All Conference(Lyon, France)(W4A ’22). Association for Computing ...

  43. [54]

    Wobbrock, Shaun K

    Jacob O. Wobbrock, Shaun K. Kane, Krzysztof Z. Gajos, Susumu Harada, and Jon Froehlich. 2011. Ability-Based Design: Concept, Principles and Examples.ACM Trans. Access. Comput.3, 3, Article 9 (April 2011), 27 pages. doi:10.1145/1952383. 1952384

  44. [55]

    Jisu Yim, Donghyeon Ko, Taeho Kim, Taejun Kim, Jonggi Hong, and Geehyuk Lee. 2025. Typing Haptically: Towards Enabling Non-auditory Smartphone Text Entry with Haptic Feedback for Blind and Low Vision Users. InProceedings of the 38th Annual ACM Symposium on User Interface Softw...

  45. [56]

    Wai Yu and Stephen Brewster. 2003. Evaluation of multimodal graphs for blind people.Univers. Access Inf. Soc.2, 2 (June 2003), 105–124. doi:10.1007/s10209- 002-0042-6

  46. [57]

    Wobbrock, and Leah Findlater

    Lotus Zhang, Zhuohao (Jerry) Zhang, Gina Clepper, Franklin Mingzhe Li, Patrick Carrington, Jacob O. Wobbrock, and Leah Findlater. 2025. VizXpress: Towards Expressive Visual Content by Blind Creators Through AI Support. InProceedings of the 27th International ACM SIGACCESS Conf...

  47. [58]

    Renrui Zhang, Dongzhi Jiang, Yichi Zhang, Haokun Lin, Ziyu Guo, Pengshuo Qiu, Aojun Zhou, Pan Lu, Kai-Wei Chang, Peng Gao, and Hongsheng Li. 2024. MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?. InEuropean Conference on Computer Vision. ht...

  48. [59]

    minimally-guided meditation sessions

    Yilun Zhao, Chengye Wang, Chuhan Li, and Arman Cohan. 2025. Can Multimodal Foundation Models Understand Schematic Diagrams? An Empirical Study on Information-Seeking QA over Scientific Papers.ArXivabs/2507.10787 (2025). https://api.semanticscholar.org/CorpusID:280290601 ASSETS...

  49. [152]

    https://api.semanticscholar.org/CorpusID:272667870

  50. [2018]

    InPro- ceedings of the 2018 CHI Conference on Human Factors in Computing Systems (Montreal QC, Canada)(CHI ’18)

    Rich Representations of Visual Content for Screen Reader Users. InPro- ceedings of the 2018 CHI Conference on Human Factors in Computing Systems (Montreal QC, Canada)(CHI ’18). Association for Computing Machinery, New York, NY, USA, 1–11. doi:10.1145/3173574.3173633

  51. [2022]

    ACM Hum.-Comput

    Exploring Team-Sourced Hyperlinks to Address Navigation Challenges for Low-Vision Readers of Scientific Papers.Proc. ACM Hum.-Comput. Interact.6, CSCW2, Article 516 (Nov. 2022), 23 pages. doi:10.1145/3555629

Pith tools

Reviewed August 1, 2026 · model on record in the stance chip above.