REVIEW 4 major objections 4 minor 59 references
Querying Multimodal Scientific Papers with AI: Practices and Preferences Across Blind, Low-Vision, and Sighted Scientists
T0 review · 4 major / 4 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read Blind and low-vision scientists rely on AI to read figures and tables in papers, but vague or wrong descriptions drive both them and sighted scientists away.
desk verdict A useful empirical study of real scientist–AI queries; the qualitative findings hold, but the headline percentages don't survive a re-count from the paper's own appendix. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The study's load-bearing mechanism is a four-way query taxonomy — General/Overview, Figures/Tables, Methods/Findings, and Miscellaneous — applied to the 115 logged queries, which reveals sharply different distributions for BLV versus sighted users. The finding is explained by two roles: an 'access layer' for BLV scientists who cannot see the visual content, and an 'efficiency layer' for sighted scientists who want faster synthesis. The critical design insight is the observed 'validation asymmetry': sighted participants could return to the original paper to check an AI answer, while BLV participants had no accessible source to verify against, amplifying the cost of any hallucination or vague
What would settle it
Log real, unaided AI paper-QA sessions from BLV and sighted scientists — without any instruction to ask about figures or tables — and repeat the measurement after a model update; if the figure/table query split narrows to near equality or the 'vague descriptions' trust breakdown disappears, the paper's central claims would be features of the task design and model snapshot rather than of the underlying access gap.
Extended reading notes
Core claim
Based on semi-structured interviews with five BLV and five sighted scientists across STEM fields, the study watched participants query two commercial AI tools, ChatGPT and Gemini, against familiar and unfamiliar papers from their own domains, logging 115 queries. The discovery is that the systems play different roles for the two groups: for BLV scientists they are a primary access bridge to visual content, while for sighted scientists they are an efficiency layer for synthesis and verification. Across both groups, AI responses that described surface visual features (colors, line styles, chart types) without scientific meaning (axes, units, spatial locations, quantitative trends) led to loss
Load-bearing premise
The study assumes that the query behavior observed in the 90-minute sessions — including the instruction to make at least one query about a figure or table and the choice of two specific AI tools at one point in time — represents how scientists actually use AI paper-QA tools in their everyday work.
Editorial extensions
If this is right
- QA systems for scientific papers should prioritize interpretive, domain-aware figure descriptions — axes, units, trends, spatial references — over literal visual semantics like colors and line styles.
- AI tools should support verification without the original PDF, for example by citing exact figure panels or quoting source passages, because BLV scientists cannot independently open the paper.
- Systems should scaffold the observed 'funnel' strategy, returning an overview first and then offering progressive detail, so users do not have to manage the structure themselves.
- Transparency about whether a response is quoted, paraphrased, or inferred would help all scientists judge trustworthiness, a need that is acute for BLV users who cannot easily check hallucinations.
- The released dataset of 115 scientist-authored queries and responses offers a grounding resource for benchmarking multimodal QA systems in real user practice rather than synthetic test sets.
Reading between the lines
- We note that the study explicitly asked participants to make at least one figure/table query, which likely inflates the figure/table share in both groups; a naturalistic log study without such a prompt might show a smaller gap, though the opening interviews suggest the direction is real.
- The 'vague descriptions' failure is a snapshot of two specific AI assistants at one point in time; newer models may shift the failure mode from vagueness to overconfident but wrong interpretations, which would change the design priorities.
- One testable extension beyond the paper: QA systems that return answers with inline source citations (e.g., figure coordinates or quoted sentences) should differentially improve BLV scientists' trust and comprehension, and that improvement could be measured directly.
- The small sample (five BLV participants, all PhD students) limits the quantitative split; recruiting more senior BLV scientists could reveal different workarounds or an even stronger reliance on AI for visual access.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a qualitative study of ten scientists (five blind/low-vision, five sighted) using ChatGPT and Gemini to query multimodal scientific papers. It categorizes 115 queries and finds that BLV participants disproportionately ask about figures and tables (49%) while sighted participants focus on methods and findings (56%), with both groups following a shared "funnel" strategy. It also documents how vague or incorrect AI descriptions erode trust, and argues this is especially consequential for BLV scientists who cannot easily validate outputs against the source document. The paper contributes a dataset of 115 queries and responses, along with design implications for inclusive QA systems. The qualitative material is rich and transparently analyzed, but the quantitative headline contains internal inconsistencies and is vulnerable to task-design and sampling confounds.
Significance. If the findings stand, this is a valuable user-centered complement to benchmark-driven multimodal QA research. The dataset of 115 scientist-authored queries is a concrete resource for grounding future evaluation in real practices, and the thematic findings—domain-aware interpretation over surface description, progressive scaffolding, transparency about synthesis vs. quotation—are actionable. The study's methodological transparency (inductive/deductive coding, peer debriefing, affinity diagramming, candid limitations) is a strength. However, the central quantitative contrast in Fig. 1 currently rests on incorrect or unclearly denominated percentages and on an explicit study instruction that may inflate the very category being compared, so the headline claims need correction and scoping before they can be taken at face value.
major comments (4)
- [§4.2, Fig. 1, Appendix A] The reported sighted figure/table percentage (29%) and methods/findings percentage (56%) do not match Appendix A. Counting Table 2 gives sighted figure/table queries as 20/76 ≈ 26% and methods/findings as 46/76 ≈ 61% (the BLV 19/39 ≈ 49% is correct). If a different denominator was used (e.g., excluding General/Overview), the paper must state this; otherwise the percentages in the abstract and Figure 1 are inaccurate. The "almost double" contrast is qualitatively preserved under either count, but the specific numbers need correction.
- [§3.2, Fig. 1, §4.2] The instruction to "make at least one query about a figure or table" places a floor on the category being compared. The manuscript neither reports how many queries were attributable to this instruction nor gives per-participant query counts. Since BLV participants contributed 39 queries and sighted participants 76, even one mandated query per BLV participant accounts for ~13% of BLV queries versus ~6.6% of sighted queries, potentially inflating the 49% figure. Please report sensitivity analyses excluding the mandated queries, provide per-participant breakdowns, and soften the claim that QA systems function as "a primary access bridge" to visual content.
- [Table 1, §5.3] All five BLV participants are PhD students, whereas the sighted group includes industry researchers, a clinical researcher, and a program manager. The abstract and introduction generalize to "BLV scientists," but the observed query-type differences may partly reflect career stage or job tasks rather than vision status. §5.3 acknowledges the homogeneous BLV sample but does not connect it to the Fig. 1 contrast. Please scope the comparative claims to the studied population or provide evidence that role and experience do not confound the distribution.
- [§3.2, §4.3, §5.3] The trust/abandonment findings are observed with two specific model versions (GPT-5 and Gemini-2.5) at one point in time. The discussion and abstract sometimes refer to "AI systems" generically (e.g., §5.2). Please add an explicit limitation that the identified failure modes may be version-specific and that rapid model improvements could change the trust dynamics, rather than implying a timeless property of the paradigm.
minor comments (4)
- [Fig. 1 caption] State the within-group denominators (n=39 BLV, n=76 sighted) and clarify whether the percentages are out of all queries or a subset.
- [§4.2] Typo: "BLV1 noted that is responses become too cumbersome" should read "if responses become too cumbersome."
- [Appendix A] Add a column indicating which queries were explicitly mandated by the study instruction; this would directly address the sensitivity concern in Major Comment 2.
- [Abstract] The phrase "current AI systems" is time-sensitive; consider dating the observation or using "the systems studied here."
Circularity Check
No significant circularity: empirical interview study with self-contained findings; procedural constraints are validity concerns, not circular reductions.
full rationale
This paper is a qualitative interview study and does not claim a derivation, prediction, or first-principles result. The headline percentages (49% figures/tables for BLV participants, 56% methods/findings for sighted participants, Fig. 1, §4.2) are descriptive tabulations of the 115 logged queries, not parameters fitted to make a claim true. The procedure did instruct participants to "make at least one query about a figure or table" (§3.2), which could bias the marginal query-type distribution and weaken the 'primary access bridge' reading; however, that is a study-design validity threat, not a circular reduction—the counts are observed behavior, not defined by the claim. The paper's self-citations (GeoVisA11y [27], FigurA11y [45], SciA11y [51], and others) appear only as related work and as motivation for studying QA; none is invoked as an external theorem or as evidence for the new empirical findings. The paper also discloses relevant limitations in §5.3 (small sample, all BLV participants being PhD students, self-selected papers), which affects generalizability but does not create circularity. No step reduces to its own inputs by construction, and no fitted parameter is renamed as a prediction. Accordingly, the paper is self-contained against its own evidence base, and the circularity score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption Self-report and think-aloud narration during a timed, task-structured session reflect how participants actually use AI QA systems in their real workflows.
- domain assumption The behavior of ChatGPT (GPT-5) and Gemini (Gemini-2.5) at the time of the study is representative of AI-powered multimodal QA systems and their shortcomings.
- domain assumption The thematic analysis coding (18 codes → 8 themes → 3 higher-level themes) is stable enough to support the reported group differences.
Cite this review
Pith. "Pith review of Querying Multimodal Scientific Papers with AI: Practices and Preferences Across Blind, Low-Vision, and Sighted Scientists." pith.science (2026). https://pith.science/paper/SRBFBGRL
@misc{pith2026260718514,
author = {Pith},
title = {Pith review of: Querying Multimodal Scientific Papers with AI: Practices and Preferences Across Blind, Low-Vision, and Sighted Scientists},
year = {2026},
howpublished = {\url{https://pith.science/paper/SRBFBGRL}},
note = {Machine review of arXiv:2607.18514}
}
read the original abstract
Visual diagrams, figures, and tables are central to scientific papers, and convey information beyond what is captured in text. While blind or low-vision (BLV) scientists have traditionally relied on static alternative text to access figures in papers, the rise of artificial intelligence (AI) has made interactive question-answering (QA) a feasible paradigm for visual exploration; yet little is known about how scientists use visual QA in practice or how to improve its accessibility. In this work, we interview five BLV and five sighted scientists across different STEM fields to understand how they use two AI tools, ChatGPT and Gemini, to query multimodal scientific documents. Our findings characterize how scientists review multimodal content, including existing practices (along with accessibility workarounds) for engaging with visuals, and feedback on the suitability of AI-generated responses to multimodal queries. We further find that vague or incomplete image descriptions, as well as incorrect AI outputs more broadly, can cause both BLV and sighted scientists to abandon AI workflows. To support future research, we additionally contribute a dataset of 115 queries and responses from our participants' interactions with the AI tools for papers in their field. We close by discussing implications for AI-powered scientific QA systems, emphasizing considerations for access across abilities and domains.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Introducing: Be My AI
2024. Introducing: Be My AI. https://www.bemyeyes.com/blog/introducing-be- my-ai
2024
-
[2]
Seeing AI - Talking Camera for the Blind
2026. Seeing AI - Talking Camera for the Blind. https://www.seeingai.com/ ASSETS ’26, October 25–28, 2026, Vila Nova de Gaia, Portugal A. Chheda-Kothary, L. L. Wang, J. C. Chang, and J. Bragg
2026
-
[3]
Storer, Antony Rishin Mukkath Roy, Ravi Kuber, and Stacy M
Ali Abdolrahmani, Kevin M. Storer, Antony Rishin Mukkath Roy, Ravi Kuber, and Stacy M. Branham. 2020. Blind Leading the Sighted: Drawing Design Insights from Blind Users towards More Productivity-oriented Voice Interfaces.ACM Trans. Access. Comput.12, 4, Article 18 (Jan. 2020), 35 pages. doi:10.1145/3368426
doi:10.1145/3368426 2020
-
[4]
Rahaf Alharbi, Pa Lor, Jaylin Herskovitz, Sarita Schoenebeck, and Robin N. Brewer
-
[5]
alphaXiv. 2026. alphaXiv: Explore. https://www.alphaxiv.org/
2026
-
[6]
Anthropic. 2026. The AI for Problem Solvers | Claude by Anthropic. https: //claude.com/product/overview
2026
-
[7]
1998.Contextual Design: Defining Customer- Centered Systems
Hugh Beyer and Karen Holtzblatt. 1998.Contextual Design: Defining Customer- Centered Systems. Morgan Kaufmann, San Francisco, CA
1998
-
[8]
Angela Glover Blackwell. 2016. The Curb-Cut Effect.Stanford Social Innovation Review15, 1 (2016), 28–33. doi:10.48558/YVMS-CC96
Show all 59 references
-
[9]
Virginia Braun and Victoria Clarke. 2006. Using thematic analysis in psychology. Qualitative Research in Psychology3, 2 (2006), 77–101. doi:10.1191/1478088706qp 063oa
2006 doi
-
[11]
Sanjana Shivani Chintalapati, Jonathan Bragg, and Lucy Lu Wang. 2022. A Dataset of Alt Texts from HCI Publications: Analyses and Uses Towards Producing More Descriptive Alt Texts of Data Visualizations in Scientific Papers. InProceedings of the 24th International ACM SIGACCESS...
2022
-
[12]
Adriana Dutkiewicz, Dietmar Müller, John Cannon, Sioned Vaughan, and Sabin Zahirovic. 2018. Sequestration and subduction of deep-sea carbonate in the global ocean since the Early Cretaceous.Geology(12 2018). doi:10.1130/G45424.1
2018 doi
-
[13]
Anders Ericsson and Herbert A
K. Anders Ericsson and Herbert A. Simon. 1993.Protocol Analysis: Verbal Reports as Data. MIT Press, Cambridge, MA
1993
-
[14]
María Evagorou, Sibel Erduran, and Terhi Mäntylä. 2015. The role of visual repre- sentations in scientific practices: from conceptual understanding and knowledge generation to ‘seeing’ how science works.International Journal of STEM Education 2 (2015), 1–13. https://api.semant...
2015
-
[15]
Fontenele, Joshua S
Antonio J. Fontenele, Joshua S. Sooter, Vincent K. Norman, Sidharth H. Gautam, and Woodrow L. Shew. 2024. Low-dimensional criticality embedded in high- dimensional awake brain dynamics.Science Advances10, 17 (2024), eadj9303. doi:10.1126/sciadv.adj9303
2024 doi
-
[16]
Google. 2026. Learn about Gemini, the everyday AI assistant from Google. https://gemini.google/about/
2026
-
[17]
Gorlewicz, Jennifer L
Jenna L. Gorlewicz, Jennifer L. Tennison, P. Merlin Uesbeck, Margaret E. Richard, Hari P. Palani, Andreas Stefik, Derrick W. Smith, and Nicholas A. Giudice. 2020. Design Guidelines and Recommendations for Multimodal, Touchscreen-based Graphics.ACM Trans. Access. Comput.13, 3, ...
2020 doi
-
[18]
Joshua Gorniak, Yoon Kim, Donglai Wei, and Nam Wook Kim. 2024. VizAbility: Enhancing Chart Accessibility with LLM-based Conversational Interaction. In Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology(Pittsburgh, PA, USA)(UIST ’24). Associa...
2024
-
[19]
Tingying He, Maggie McCracken, Daniel Hajas, Sarah Creem-Regehr, and Alexan- der Lex. 2026. Using Tactile Charts to Support Comprehension and Learn- ing of Complex Visualizations for Blind and Low-Vision Individuals.IEEE Transactions on Visualization and Computer Graphics32, 1...
2026
-
[20]
Yugo Iwamoto, Muhammad Raees, Jamison Heard, and Garreth W. Tigwell. 2025. Exploring Generative AI to Support Disability Service Professionals in Writing Image Descriptions for HCI Science Figures. InProceedings of the 27th Interna- tional ACM SIGACCESS Conference on Computers...
2025
-
[21]
Madhan Jeyaraman, Naveen Jeyaraman, Swaminathan Ramasubramanian, Ab- hishek Vaish, and Raju Vaishya. 2024. Decoding Research with a Glance: The Power of Graphical Abstracts and Infographics.Apollo Medicine22 (2024), 144 –
2024
-
[22]
Wan Ju Kang, Eunki Kim, Na Min An, Sangryul Kim, Ha-Young Choi, Ki Hoon Kwak, and James Thorne. 2025. Sightation Counts: Leveraging Sighted User Feedback in Building a BLV-aligned Dataset of Diagram Descriptions.ArXiv abs/2503.13369 (2025). https://api.semanticscholar.org/Corp...
2025 arXiv
-
[23]
Jiho Kim, Arjun Srinivasan, Nam Wook Kim, and Yea-Seul Kim. 2023. Exploring Chart Question Answering for Blind and Low Vision Users. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems(Hamburg, Germany)(CHI ’23). Association for Computing Machinery, ...
2023
-
[24]
Anukriti Kumar and Lucy Lu Wang. 2024. Uncovering the New Accessibility Crisis in Scholarly PDFs: Publishing Model and Platform Changes Contribute to Declining Scholarly Document Accessibility in the Last Decade. InProceedings of the 26th International ACM SIGACCESS Conference...
2024
-
[25]
West, and Bill Howe
Po-Shen Lee, Jevin D. West, and Bill Howe. 2016. Viziometrics: Analyzing Visual Information in the Scientific Literature.IEEE Transactions on Big Data4 (2016), 117–129. https://api.semanticscholar.org/CorpusID:3665638
2016
-
[26]
I never realized sidewalks were a big deal
Chu Li, Katrina Oi Yau Ma, Michael Saugstad, Kie Fujii, Molly Delaney, Yochai Eisenberg, Delphine Labbé, Judy L Shanley, Devon Snyder, Florian P P Thomas, and Jon E. Froehlich. 2024. “I never realized sidewalks were a big deal”: A Case Study of a Community-Driven Sidewalk Acce...
2024
-
[27]
Froehlich
Chu Li, Rock Yuren Pang, Arnavi Chheda-Kothary, Ather Sharif, Henok Assalif, Jeffrey Heer, and Jon E. Froehlich. 2026. GeoVisA11y: An AI- based Geovisualization Question-Answering System for Screen-Reader Users. arXiv:2603.07446 [cs.HC] https://arxiv.org/abs/2603.07446
2026
-
[28]
Chuhan Li, Ziyao Shangguan, Yilun Zhao, Deyuan Li, Yixin Liu, and Arman Cohan. 2024. M3SciQA: A Multi-Modal Multi-Document Scientific QA Bench- mark for Evaluating Foundation Models.ArXivabs/2411.04075 (2024). https: //api.semanticscholar.org/CorpusID:273850071
2024 arXiv
-
[29]
Alan Lundgard and Arvind Satyanarayan. 2021. Accessible Visualization via Natural Language Descriptions: A Four-Level Model of Semantic Content.IEEE Transactions on Visualization and Computer GraphicsPP (2021), 1–1. https: //api.semanticscholar.org/CorpusID:238237335
2021
-
[31]
Hayes, and Devva Kasnitz
Jennifer Mankoff, Gillian R. Hayes, and Devva Kasnitz. 2010. Disability studies as a source of critical inquiry for the field of assistive technology. InProceedings of the 12th International ACM SIGACCESS Conference on Computers and Accessibility (Orlando, Florida, USA)(ASSETS...
2010
-
[32]
Nora McDonald, Sarita Schoenebeck, and Andrea Forte. 2019. Reliability and Inter-rater Reliability in Qualitative Research: Norms and Guidelines for CSCW and HCI Practice.Proc. ACM Hum.-Comput. Interact.3, CSCW, Article 72 (Nov. 2019), 23 pages. doi:10.1145/3359174
2019 doi
-
[33]
Beverley Cherie Millar and Michelle Lim. 2022. The Role of Visual Abstracts in the Dissemination of Medical Research.The Ulster Medical Journal91 (2022), 67 – 78. https://api.semanticscholar.org/CorpusID:249821792
2022
-
[34]
Accessibility People, You Go Work on That Thing of Yours over There
Sanika Moharana, Cynthia L. Bennett, Erin Buehler, Michael Madaio, Vinita Tibdewal, and Shaun K. Kane. 2025. “Accessibility People, You Go Work on That Thing of Yours over There”: Addressing Disability Inclusion in AI Product Organizations.Proceedings of the AAAI/ACM Conferenc...
2025 doi
-
[35]
Bennett, and Edward Cutrell
Meredith Ringel Morris, Jazette Johnson, Cynthia L. Bennett, and Edward Cutrell
-
[36]
Peya Mowar, Aaron Steinfeld, and Jeffrey P Bigham. 2026. iTagPDF: Towards Finally Automating PDF Accessibility. InProceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI ’26). Association for Computing Machinery, New York, NY, USA, Article 1116, 17 pa...
2026 doi
-
[37]
OpenAI. 2026. ChatGPT | AI Chatbot to Discover, Learn & Create. https: //chatgpt.com/overview/
2026
-
[38]
Soya Park, Jonathan Bragg, Michael Chang, Kevin Larson, and Danielle Bragg
-
[39]
P. A. Patel, N. P. DeGroote, K. Jackson, T. Cash, S. M. Castellino, P. Jaggi, A. J. Es- benshade, and T. P. Miller. 2022. Infectious events in pediatric patients with acute lymphoblastic leukemia/lymphoma undergoing evaluation for fever without severe neutropenia.Cancer128, 23...
2022 doi
-
[40]
Gonzalez Penuela, Crescentia Jung, Sharon Y Lin, Ruiying Hu, and Shiri Azenkot
Ricardo E. Gonzalez Penuela, Crescentia Jung, Sharon Y Lin, Ruiying Hu, and Shiri Azenkot. 2026. How Multimodal Large Language Models Support Access to Visual Information: A Diary Study With Blind and Low Vision People. https: //api.semanticscholar.org/CorpusID:285616065
2026
-
[41]
Shraman Pramanick, Rama Chellappa, and Subhashini Venugopalan. 2024. SPIQA: A Dataset for Multimodal Question Answering on Scientific Papers. InAdvances Querying Multimodal Scientific Papers with AI ASSETS ’26, October 25–28, 2026, Vila Nova de Gaia, Portugal in Neural Informa...
2024 doi
-
[42]
Quora. 2026. Poe - Fast, Helpful AI Chat. https://poe.com/
2026
-
[43]
Qwen. 2026. Qwen Chat. https://chat.qwen.ai/
2026
-
[44]
Wang, Alida T
Ather Sharif, Olivia H. Wang, Alida T. Muongchan, Katharina Reinecke, and Jacob O. Wobbrock. 2022. VoxLens: Making Online Data Visualizations Accessible with an Interactive JavaScript Plug-In. InProceedings of the 2022 CHI Conference on Human Factors in Computing Systems(New O...
2022
-
[45]
Nikhil Singh, Lucy Lu Wang, and Jonathan Bragg. 2024. FigurA11y: AI Assistance for Writing Scientific Alt Text. InProceedings of the 29th International Conference on Intelligent User Interfaces(Greenville, SC, USA)(IUI ’24). Association for Com- puting Machinery, New York, NY,...
2024
-
[46]
Sharon Spall. 1998. Peer Debriefing in Qualitative Research: Emerging Opera- tional Models. 4, 2 (Jan 1998), 280–292. doi:10.1177/107780049800400208
1998 doi
-
[47]
Fleischmann, Meredith Ringel Morris, and Danna Gurari
Abigale Stangl, Nitin Verma, Kenneth R. Fleischmann, Meredith Ringel Morris, and Danna Gurari. 2021. Going Beyond One-Size-Fits-All Image Descriptions to Satisfy the Information Wants of People Who are Blind or Have Low Vision. In Proceedings of the 23rd International ACM SIGA...
2021
-
[48]
Xinru Tang, Ali Abdolrahmani, Darren Gergle, and Anne Marie Piper. 2025. Everyday Uncertainty: How Blind People Use GenAI Tools for Information Access. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). Association for Computing Machinery...
2025
-
[50]
B. N. Walker and L. M. Mauney. 2010. Universal Design of Auditory Graphs: A Comparison of Sonification Mappings for Visually Impaired and Sighted Listeners. ACM Trans. Access. Comput.2, 3, Article 12 (March 2010), 16 pages. doi:10.1145/ 1714458.1714459
2010
-
[51]
Lucy Lu Wang, Isabel Cachola, Jonathan Bragg, Evie Yu-Yen Cheng, Chelsea Haupt, Matt Latzke, Bailey Kuehl, Madeleine N van Zuylen, Linda Wagner, and Daniel Weld. 2021. SciA11y: Converting Scientific Papers to Accessible HTML. In Proceedings of the 23rd International ACM SIGACC...
2021
-
[52]
Wagner, and Daniel S
Lucy Lu Wang, Isabel Cachola, Jonathan Bragg, Evie (Yu-Yen) Cheng, Chelsea Hess Haupt, Matt Latzke, Bailey Kuehl, Madeleine van Zuylen, Linda M. Wagner, and Daniel S. Weld. 2021. Improving the Accessibility of Scientific Doc- uments: Current State, User Needs, and a System Sol...
2021 arXiv
-
[53]
Candace Williams, Lilian de Greef, Ed Harris, Leah Findlater, Amy Pavel, and Cynthia Bennett. 2022. Toward supporting quality alt text in computing pub- lications. InProceedings of the 19th International Web for All Conference(Lyon, France)(W4A ’22). Association for Computing ...
2022
-
[54]
Wobbrock, Shaun K
Jacob O. Wobbrock, Shaun K. Kane, Krzysztof Z. Gajos, Susumu Harada, and Jon Froehlich. 2011. Ability-Based Design: Concept, Principles and Examples.ACM Trans. Access. Comput.3, 3, Article 9 (April 2011), 27 pages. doi:10.1145/1952383. 1952384
2011 doi
-
[55]
Jisu Yim, Donghyeon Ko, Taeho Kim, Taejun Kim, Jonggi Hong, and Geehyuk Lee. 2025. Typing Haptically: Towards Enabling Non-auditory Smartphone Text Entry with Haptic Feedback for Blind and Low Vision Users. InProceedings of the 38th Annual ACM Symposium on User Interface Softw...
2025
-
[56]
Wai Yu and Stephen Brewster. 2003. Evaluation of multimodal graphs for blind people.Univers. Access Inf. Soc.2, 2 (June 2003), 105–124. doi:10.1007/s10209- 002-0042-6
2003 doi
-
[57]
Wobbrock, and Leah Findlater
Lotus Zhang, Zhuohao (Jerry) Zhang, Gina Clepper, Franklin Mingzhe Li, Patrick Carrington, Jacob O. Wobbrock, and Leah Findlater. 2025. VizXpress: Towards Expressive Visual Content by Blind Creators Through AI Support. InProceedings of the 27th International ACM SIGACCESS Conf...
2025
-
[58]
Renrui Zhang, Dongzhi Jiang, Yichi Zhang, Haokun Lin, Ziyu Guo, Pengshuo Qiu, Aojun Zhou, Pan Lu, Kai-Wei Chang, Peng Gao, and Hongsheng Li. 2024. MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?. InEuropean Conference on Computer Vision. ht...
2024
-
[59]
minimally-guided meditation sessions
Yilun Zhao, Chengye Wang, Chuhan Li, and Arman Cohan. 2025. Can Multimodal Foundation Models Understand Schematic Diagrams? An Empirical Study on Information-Seeking QA over Scientific Papers.ArXivabs/2507.10787 (2025). https://api.semanticscholar.org/CorpusID:280290601 ASSETS...
2025 arXiv
-
[152]
https://api.semanticscholar.org/CorpusID:272667870
-
[2018]
InPro- ceedings of the 2018 CHI Conference on Human Factors in Computing Systems (Montreal QC, Canada)(CHI ’18)
Rich Representations of Visual Content for Screen Reader Users. InPro- ceedings of the 2018 CHI Conference on Human Factors in Computing Systems (Montreal QC, Canada)(CHI ’18). Association for Computing Machinery, New York, NY, USA, 1–11. doi:10.1145/3173574.3173633
2018
-
[2022]
ACM Hum.-Comput
Exploring Team-Sourced Hyperlinks to Address Navigation Challenges for Low-Vision Readers of Scientific Papers.Proc. ACM Hum.-Comput. Interact.6, CSCW2, Article 516 (Nov. 2022), 23 pages. doi:10.1145/3555629
2022 doi
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.