REVIEW 4 major objections 5 minor 97 references
Task Mode: Dynamic Filtering for Task-Specific Web Navigation using LLMs
T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Task Mode uses GPT-4o to hide irrelevant web content, cutting screen-reader task times from 211 to 102 seconds and shrinking the sighted/blind completion gap from 2x to 1.2x.
desk verdict Useful, original accessibility system with a strong prototype, but its central SRU time-gap claim rests on precomputed pages rather than the live extension, so treat the 1.2x figure as a proof-of-concept result, not a system measurement. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is an LLM-based relevance-scoring pipeline. The user's task is decomposed into five components (entity, constraints, actions, defaults, fallbacks), and GPT-4o scores each text element from 0 to 100 with explicit instructions to consider card structure, surrounding context, and repeated groups rather than isolated words. Images are scored by a weighted combination of alt-text similarity (0.3) and CLIP image-task similarity (0.7), while SVG icons are first labeled from their path data and then scored against the task. Scores are propagated upward in the DOM so parent containers of relevant children remain visible, and three rendering modes (color gradients, opacity fade, and threshold-based hiding with aria-hidden) translate the scores into visual and screen-reader-visible changes.
What would settle it
Have human raters mark the ground-truth set of elements necessary to answer each of the study's nine website-task pairs, then check whether Task Mode hides any element a rater marked necessary. If hiding removes needed content on a meaningful share of tasks, the time gains may partly reflect users answering with missing information; a follow-up study scoring answer correctness instead of only time would settle this.
Extended reading notes
Core claim
The central discovery is that task-specific, LLM-driven filtering of webpage elements improves usability for both vision users (VUs) and screen reader users (SRUs), with the largest gain for SRUs. The paper reports SRU mean task completion time dropping from 211 seconds (sigma = 49.1) to 102 seconds (sigma = 26.2) with p < 0.05, while VU time went from 107 seconds to 84.4 seconds without reaching significance; the SRU/VU completion-time gap shrank from about 2x to 1.2x. All six VUs and five of six SRUs said they would want to use Task Mode in the future, reporting less effort and fewer distractions. The system preserves the original page structure, supports adjustable relevance thresholds and multiple rendering modes, and maintains the task context across page transitions.
Load-bearing premise
GPT-4o's judgments of which page elements are relevant to the user's task are correct often enough that hiding every low-scoring element does not remove content the user needs.
Editorial extensions
If this is right
- Screen reader users can roughly halve task completion time on live commercial and information websites without site-specific retraining, because relevance is judged semantically from the DOM.
- The completion-time disparity between vision users and screen reader users can shrink from about 2x to about 1.2x on the evaluated task types, showing the usability gap is not fixed.
- Relevance-threshold adjustment gives users a controllable trade-off between a minimal filtered view and fuller context, which matters for tasks where missing content is costly.
- Because filtering complements rather than replaces screen-reader strategies like heading and landmark navigation, users keep their existing interaction habits while scanning fewer irrelevant regions.
- Maintaining the original page layout while hiding or de-emphasizing elements supports spatial orientation and builds trust, unlike agents that restructure or automate the page.
Reading between the lines
- If the per-element relevance scores are as accurate as the time results suggest, the same pipeline could plausibly extend to low-vision, motor-impaired, or attention-difficulty users who also benefit from a decluttered interface; the paper gestures at this but does not test those groups.
- A ground-truth evaluation of the scores (human-annotated sets of task-necessary elements) would settle whether hiding low-scoring elements ever removes content required to answer correctly; the current study measures time and self-report, not answer accuracy on the filtered view.
- The task-decomposition step could transfer to other media, such as highlighting relevant video segments or de-emphasizing irrelevant audio, as the authors suggest, but those extensions remain untested directions rather than demonstrated results.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Task Mode, a Chrome extension that uses GPT-4o-based relevance scoring to filter, fade, or highlight DOM elements according to a user-supplied task, with three rendering modes aimed at both vision users (VUs) and screen reader users (SRUs). The central empirical claim is that in a controlled user study with 12 participants, SRUs' task completion time dropped from a mean of 211 seconds to 102 seconds (p<0.05), VU time was not significantly changed, and the SRU/VU completion-time gap shrank from 2x to 1.2x. The paper also reports qualitative feedback on reduced effort, preserved page structure, trust, and future use cases, alongside technical measurements of parsing latency and cost.
Significance. If the claimed effect is real, the paper makes a useful contribution to accessible web navigation by showing that LLM-driven, task-specific filtering can be designed for both visual and non-visual access simultaneously. The system is described in sufficient detail to be re-implemented, the prompt templates are included, and the authors are transparent about several engineering limitations. The dual-user design and the focus on preserving page structure are genuine strengths. However, as detailed in the major comments, the headline quantitative result is not a measurement of the live system for SRUs, and the study does not directly verify task correctness, so the magnitude and even the direction of the reported effect need to be restated or re-measured before the paper's central claim can be accepted at face value.
major comments (4)
- [§5.1.2, §5.2, Table 4, §4.4] The headline SRU result is not a measurement of the live Task Mode system. Section 5.1.2 states that SRUs submitted their chosen websites and tasks in advance and that the team 'processed those pairs in advance and shared the Task Mode output with SRUs,' while VUs interacted with the live extension via Zoom Remote Control. Consequently, the 102.21 s mean in Table 4 excludes the 10.68 s per-page processing latency reported in §4.4, and the SRU/VU gap reduction from 2x to 1.2x is computed against a static filtered artifact rather than the dynamic system credited in the abstract and discussion. Re-adding one page's latency to the SRU mean yields roughly 113 s (a 1.34x gap), and for the two-page CNN tasks it yields roughly 123 s (a 1.46x gap) against the VU mean in Table 3, materially weakening the 1.2x claim. The authors should either collect SRU timings on the live system with end-to-end latency, or explicitly restate the claim as an evaluation of pre-rendered filtered output rather than of the interactive system.
- [§5.2, Table 4, §6.2.1] The study measures completion time and self-report but never measures task accuracy. Completion time alone cannot establish that Task Mode 'maintains performance,' because users could finish faster while missing content needed for correct answers. The paper reports no ground-truth comparison of final answers under Task Mode versus traditional browsing, and it does not analyze whether the filtered pages contained every element required for the given tasks. Section 6.2.1 acknowledges that useful elements can receive low scores and that omissions 'sometimes created doubt about whether something important was removed,' but this limitation is not reflected in the quantitative performance claim. The authors should add a correctness analysis (for example, logging final answers, verifying them against ground truth, and reporting task success rates by condition), or narrow the claim to efficiency only and explicitly caution that filtering accuracy was not independently validated.
- [§5.2, Table 3, Abstract] The claim that Task Mode 'maintained performance for VUs' is not supported by the reported statistics. VU completion time decreased from 106.87 s to 84.43 s, but the Wilcoxon test gives p=0.563 with N=6, which only shows that no significant difference was detected; it does not establish equivalence or the absence of a practically meaningful change. The abstract and discussion should either be rephrased to say that no statistically significant change was observed for VUs, or be supported by an equivalence test, confidence interval, or pre-specified non-inferiority criterion.
- [§4.2.2, §4.4] The relevance-scoring mechanism itself is not validated against any ground truth. The image relevance score is a heuristic 0.3/0.7 weighted combination of GPT-4o alt-text similarity and CLIP image similarity, with the ratio 'selected heuristically based on preliminary experiments' but no sensitivity analysis reported. Likewise, the client-side text-length threshold and the user-adjustable relevance threshold are not ablated or evaluated for their effect on filtering precision or recall. Section 4.4 reports latency, cost, and parse success on 20 websites, but none of these metrics measures whether the filtered output retains task-critical content. This is load-bearing because the user-study outcomes are attributed to the relevance-filtering pipeline, yet no component-level evidence connects the scoring choices to the observed efficiency gains.
minor comments (5)
- [§2.3] The phrase 'academic researh' should read 'academic research.'
- [Figure 4] Figure 4 shows relevance scores on a 0-to-1 scale (e.g., 0.89), while the pipeline text and prompts in §4.2.2 use a 0-to-100 scale; please align the notation to avoid confusion.
- [§5.2.3] Participant labels are inconsistent: some quotes are attributed to 'VU1' and 'VU3,' while others are attributed to 'P2,' 'P3,' 'P5,' and 'P6.' Please use a single consistent participant identifier scheme.
- [Tables 3 and 4] The tables report many Wilcoxon tests with no multiple-comparison correction or pre-registered primary outcome. At minimum, state which comparisons were planned as primary, or report adjusted p-values, so that readers can interpret the significant rows in context.
- [Table 5 caption] The caption refers to 'Part A' of the user study, while the procedure text in §5.1.3 uses 'Phase 1' and 'Phase 2'; please harmonize the terminology.
Circularity Check
No circularity: the reported times are measured user-study outcomes, not predictions derived from fitted parameters or self-citations.
full rationale
The paper's central quantitative claims—SRU completion time dropping from 211s to 102s, VU time staying roughly stable, and the SRU/VU gap shrinking from 2x to 1.2x—are the raw means of a within-subjects user study reported in Tables 3 and 4, with Wilcoxon tests. No equation, fitted parameter, or generative-model output is used to predict those times, so no result reduces to its input by construction. The only parameters that could be called fitted, such as the 0.3/0.7 image-score blend chosen 'heuristically based on preliminary experiments,' are engineering choices that do not force the completion-time outcome and are not renamed as predictions. The sole self-citation, reference [45], appears only in related-work context for prior image-description work and is not load-bearing for any claim in the derivation. The skeptic-flagged asymmetry—SRUs used precomputed filtered pages and the 10.68s/page latency from §4.4 was excluded from SRU timings—is a genuine measurement/construct-validity concern about what was timed, but it is not circularity: the reported means are still observed data, not values derived from the system's own scoring or from a prior self-citation. The limitation section §6.2.1 acknowledges that LLM relevance scoring can miss important elements, but that is an accuracy limitation, not a self-reference in the argument. I therefore find no circular step and assign score 0.
Assumptions & free parameters
free parameters (3)
- Image score weight ratio =
0.3 * Task-Alt Score + 0.7 * Task-Image Score
- Client-side text length threshold =
fewer than 3 characters
- Relevance threshold for 'Only Task-Specific' mode =
not stated for study (e.g., 60-75% elsewhere)
assumptions (4)
- domain assumption LLM relevance judgments are a valid proxy for task relevance.
- domain assumption DOM structure, alt text, and SVG path data provide sufficient information to judge element relevance.
- ad hoc to paper Precomputed filtered pages used in the SRU evaluation are representative of the live system's output.
- ad hoc to paper Completion time alone, without correctness measurement, is sufficient to infer performance improvements.
Cite this review
Pith. "Pith review of Task Mode: Dynamic Filtering for Task-Specific Web Navigation using LLMs." pith.science (2026). https://pith.science/paper/JJDGOERL
@misc{pith2026250714769,
author = {Pith},
title = {Pith review of: Task Mode: Dynamic Filtering for Task-Specific Web Navigation using LLMs},
year = {2026},
howpublished = {\url{https://pith.science/paper/JJDGOERL}},
note = {Machine review of arXiv:2507.14769}
}
read the original abstract
Modern web interfaces are unnecessarily complex to use as they overwhelm users with excessive text and visuals unrelated to their current goals. This problem particularly impacts screen reader users (SRUs), who navigate content sequentially and may spend minutes traversing irrelevant elements before reaching desired information compared to vision users (VUs) who visually skim in seconds. We present Task Mode, a system that dynamically filters web content based on user-specified goals using large language models to identify and prioritize relevant elements while minimizing distractions. Our approach preserves page structure while offering multiple viewing modes tailored to different access needs. Our user study with 12 participants (6 VUs, 6 SRUs) demonstrates that our approach reduced task completion time for SRUs while maintaining performance for VUs, decreasing the completion time gap between groups from 2x to 1.2x. 11 of 12 participants wanted to use Task Mode in the future, reporting that Task Mode supported completing tasks with less effort and fewer distractions. This work demonstrates how designing new interactions simultaneously for visual and non-visual access can reduce rather than reinforce accessibility disparities in future technology created by human-computer interaction researchers and practitioners.
Figures
Reference graph
Works this paper leans on
-
[1]
2025. Amazon. https://www.amazon.com/. Accessed: 2025
2025
-
[2]
Anthropic Computer Use
2025. Anthropic Computer Use. https://docs.anthropic.com/en/docs/agents-and- tools/computer-use. Accessed: 2025
2025
-
[3]
Apple Distraction Control
2025. Apple Distraction Control. https://support.apple.com/en-us/120682. Ac- cessed: 2025
2025
-
[4]
Best Buy
2025. Best Buy. https://www.bestbuy.com/. Accessed: 2025
2025
-
[5]
Breaking News, Latest News and Videos
2025. Breaking News, Latest News and Videos. https://www.cnn.com Accessed: 2025
2025
-
[6]
Browser Use
2025. Browser Use. https://browser-use.com/. Accessed: 2025
2025
-
[7]
2025. eBay. https://www.ebay.com/. Accessed: 2025
2025
-
[8]
2025. H-E-B. https://www.heb.com/ Accessed: 2025
2025
Show all 97 references
-
[9]
Introduction to Lighthouse
2025. Introduction to Lighthouse. https://developer.chrome.com/docs/lighthouse/ overview Accessed: 2025
2025
-
[10]
A Look at New Features in the JAWS 2023 Release
2025. A Look at New Features in the JAWS 2023 Release. https://afb.org/aw/23/ 11/18108 Accessed: 2025
2025
-
[11]
Nanobrowser
2025. Nanobrowser. https://github.com/nanobrowser/nanobrowser. Accessed: 2025
2025
-
[12]
OpenAI Operator
2025. OpenAI Operator. https://operator.chatgpt.com/. Accessed: 2025
2025
-
[13]
2025. React. https://react.dev/ Accessed: 2025
2025
-
[14]
SE Ranking â €” Robust SEO Software for Every Major Task
2025. SE Ranking â €” Robust SEO Software for Every Major Task. https: //seranking.com/ Accessed: 2025
2025
-
[15]
2025. Taxy AI. https://taxy.ai/. Accessed: 2025
2025
-
[16]
Understanding Success Criterion 3.2.3: Consistent Navigation | WAI | W3C
2025. Understanding Success Criterion 3.2.3: Consistent Navigation | WAI | W3C. https://www.w3.org/WAI/WCAG21/Understanding/consistent-navigation. html Accessed: 2025
2025
-
[17]
WCAG 2 Overview
2025. WCAG 2 Overview. https://www.w3.org/WAI/standards-guidelines/wcag/. Accessed: 2025
2025
-
[18]
WebAIM: The WebAIM Million - The 2025 report on the accessibility of the top 1,000,000 home pages
2025. WebAIM: The WebAIM Million - The 2025 report on the accessibility of the top 1,000,000 home pages. https://webaim.org/projects/million/ Accessed: 2025
2025
-
[19]
Wikipedia, the free encyclopedia
2025. Wikipedia, the free encyclopedia. https://www.wikipedia.org/ Accessed: 2025
2025
-
[20]
Hatim Alsayahani, Mohammed Alhamadi, Simon Harper, and Markel Vigo. 2025. The Effects of Customisation on the Usability of Visual Analytics Dashboards: the Good, the Bad, and the Ugly. InProceedings of the 30th International Conference on Intelligent User Interfaces. 1426–1439
2025
-
[21]
Erick Antonio Alves, Paula Christina Figueira Cardoso, and André Pimenta Freire. 2018. Automatically generated summaries as in-page web navigation accelerators for blind users. In Proceedings of the 17th Brazilian Symposium on Human Factors in Computing Systems. 1–8
2018
-
[22]
Apple Inc. 2024. Use Distraction Control in Safari to hide items on a webpage. Apple Support. https://support.apple.com/en-us/120682 Published on December 04, 2024
2024
-
[23]
Vikas Ashok, Yevgen Borodin, Yury Puzis, and IV Ramakrishnan. 2015. Capti- speak: a speech-enabled web screen reader. In Proceedings of the 12th International Web for All Conference. 1–10
2015
-
[24]
Vikas Ashok, Yury Puzis, Yevgen Borodin, and IV Ramakrishnan. 2017. Web screen reading automation assistance using semantic abstraction. InProceedings of the 22nd International Conference on Intelligent User Interfaces. 407–418
2017
-
[25]
Gilles Baechler, Srinivas Sunkara, Maria Wang, Fedir Zubach, Hassan Mansoor, Vincent Etter, Victor Cărbune, Jason Lin, Jindong Chen, and Abhanshu Sharma
-
[26]
Marcos Baez, Claudia Maria Cutrupi, Maristella Matera, Isabella Possaghi, Emanuele Pucci, Gianluca Spadone, Cinzia Cappiello, and Antonella Pasquale
-
[27]
Jeffrey P Bigham, Anna C Cavender, Jeremy T Brudvik, Jacob O Wobbrock, and Richard E Ladner. 2007. WebinSitu: a comparative analysis of blind and sighted browsing behavior. In Proceedings of the 9th International ACM SIGACCESS Conference on Computers and Accessibility. 51–58
2007
-
[28]
Andy Brown and Simon Harper. 2013. Dynamic injection of WAI-ARIA into web content. In Proceedings of the 10th International Cross-Disciplinary Conference on Web Accessibility. 1–4
2013
-
[29]
Yining Cao, Peiling Jiang, and Haijun Xia. 2025. Generative and Malleable User Interfaces with Generative and Evolving Task-Driven Data Model.arXiv preprint arXiv:2503.04084 (2025)
2025 arXiv
-
[30]
Jia Chen, Jiaxin Mao, Yiqun Liu, Fan Zhang, Min Zhang, and Shaoping Ma. 2021. Towards a better understanding of query reformulation behavior in web search. In Proceedings of the web conference 2021. 743–755
2021
-
[31]
Yifei Cheng, Yukang Yan, Xin Yi, Yuanchun Shi, and David Lindlbauer. 2021. SemanticAdapt: Optimization-based Adaptation of Mixed Reality Layouts Leveraging Virtual-Physical Semantic Connections. In The 34th Annual ACM Symposium on User Interface Software and Technology (Virtua...
2021
-
[32]
Arnavi Chheda-Kothary, David A Rios, Kynnedy Simone Smith, Avery Reyna, Cecilia Zhang, and Brian A Smith. 2023. Understanding Blind and Low Vi- sion Users’ Attitudes Towards Spatial Interactions in Desktop Screen Read- ers. In Proceedings of the 25th International ACM SIGACCES...
2023
-
[33]
Paul T Chiou, Ali S Alotaibi, and William GJ Halfond. 2023. Bagel: An ap- proach to automatically detect navigation-based web accessibility barriers for keyboard users. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–17
2023
-
[34]
Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Sam Stevens, Boshi Wang, Huan Sun, and Yu Su. 2023. Mind2web: Towards a generalist agent for the web. Advances in Neural Information Processing Systems 36 (2023), 28091–28114
2023
-
[35]
Ryan Fritz, Kim-Phuong L Vu, and Wayne E Dick. 2019. Customization: the path to a better and more accessible web experience. InHuman Interface and the Management of Information. Visual Information and Knowledge Management: Thematic Area, HIMI 2019, Held as Part of the 21st HCI...
2019
-
[36]
Krzysztof Gajos and Daniel S Weld. 2004. SUPPLE: automatically generating user interfaces. In Proceedings of the 9th international conference on Intelligent user interfaces. 93–100
2004
-
[37]
Gajos, Mary Czerwinski, Desney S
Krzysztof Z. Gajos, Mary Czerwinski, Desney S. Tan, and Daniel S. Weld
-
[38]
Krzysztof Z Gajos, Jing Jing Long, and Daniel S Weld. 2006. Automatically gener- ating custom user interfaces for users with physical disabilities. In Proceedings of the 8th International ACM SIGACCESS Conference on Computers and Accessibility. 243–244
2006
-
[39]
Krzysztof Z Gajos, Jacob O Wobbrock, and Daniel S Weld. 2007. Automatically generating user interfaces adapted to users’ motor and vision capabilities. In Proceedings of the 20th annual ACM symposium on User interface software and technology. 231–240
2007
-
[40]
Krzysztof Z Gajos, Jacob O Wobbrock, and Daniel S Weld. 2008. Improving the performance of motor-impaired users with automatically-generated, ability- based interfaces. In Proceedings of the SIGCHI conference on Human Factors in Computing Systems. 1257–1266
2008
-
[41]
getadblock.com. 2025. AdBlock — block ads across the web. Chrome Web Store. https://chromewebstore.google.com/detail/adblock-%E2%80%94-block- ads-acros/gighmmpiobklfepjocnamgkkbiglidom Version 6.18.0, Updated April 3, 2025
2025
-
[42]
Stéphanie Giraud, Pierre Thérouanne, and Dirk D Steiner. 2018. Web accessibility: Filtering redundant and irrelevant information improves website usability for blind users. International Journal of Human-Computer Studies 111 (2018), 23– 35
2018
-
[43]
Google. [n. d.]. Firebase Realtime Database Documentation. https://firebase. google.com/docs/database. Accessed: 2025-04-16
2025
-
[44]
Colin M Gray, Yubo Kou, Bryan Battles, Joseph Hoggatt, and Austin L Toombs
-
[45]
Ananya Gubbi Mohanbabu and Amy Pavel. 2024. Context-aware image de- scriptions for web accessibility. In Proceedings of the 26th International ACM SIGACCESS Conference on Computers and Accessibility. 1–17
2024
-
[46]
Simon Harper and Sean Bechhofer. 2007. SADIe: Structural semantics for ac- cessibility and device independence. ACM Transactions on Computer-Human Interaction (TOCHI) 14, 2 (2007), 10–es
2007
-
[47]
Sandra G Hart and Lowell E Staveland. 1988. Development of NASA-TLX (Task Load Index): Results of empirical and theoretical research. In Advances in psychology. Vol. 52. Elsevier, 139–183
1988
-
[48]
Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, and Dong Yu. 2024. WebVoyager: Building an end-to-end web agent with large multimodal models. arXiv preprint arXiv:2401.13919 (2024)
2024 arXiv
-
[49]
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi
-
[50]
Wenyi Hong, Weihan Wang, Qingsong Lv, Jiazheng Xu, Wenmeng Yu, Junhui Ji, Yan Wang, Zihan Wang, Yuxiao Dong, Ming Ding, et al. 2024. Cogagent: A visual language model for gui agents. In Proceedings of the IEEE/CVF Conference on Task Mode: Dynamic Filtering for Task-Specific We...
2024
-
[51]
Peter C Humphreys, David Raposo, Tobias Pohlen, Gregory Thornton, Rachita Chhaparia, Alistair Muldal, Josh Abramson, Petko Georgiev, Adam Santoro, and Timothy Lillicrap. 2022. A data-driven approach for learning to control computers. In International Conference on Machine Lear...
2022
-
[52]
Faria Huq, Zora Zhiruo Wang, Frank F Xu, Tianyue Ou, Shuyan Zhou, Jeffrey P Bigham, and Graham Neubig. 2025. CowPilot: A Framework for Autonomous and Human-Agent Collaborative Web Navigation. arXiv preprint arXiv:2501.16609 (2025)
2025
-
[53]
International Telecommunication Union. 2024. Measuring digital development: Facts and Figures 2024. Report. International Telecommunication Union. https:// www.itu.int/en/ITU-D/Statistics/Pages/facts/default.aspx Published in November 2024
2024
-
[54]
Clemens N Klokmose, James R Eagan, Siemen Baader, Wendy Mackay, and Michel Beaudouin-Lafon. 2015. Webstrates: shareable dynamic media. In Proceedings of the 28th Annual ACM Symposium on User Interface Software & Technology. 280–290
2015
-
[55]
Satwik Ram Kodandaram, Utku Uckun, Xiaojun Bi, IV Ramakrishnan, and Vikas Ashok. 2024. Enabling Uniform Computer Interaction Experience for Blind Users through Large Language Models. In Proceedings of the 26th International ACM SIGACCESS Conference on Computers and Accessibili...
2024
-
[56]
Geza Kovacs, Zhengxuan Wu, and Michael S Bernstein. 2021. Not now, ask later: users weaken their behavior change regimen over time, but expect to re- strengthen it imminently. InProceedings of the 2021 CHI Conference on Human Factors in Computing Systems. 1–14
2021
-
[57]
Hanyu Lai, Xiao Liu, Iat Long Iong, Shuntian Yao, Yuxuan Chen, Pengbo Shen, Hao Yu, Hanchen Zhang, Xiaohan Zhang, Yuxiao Dong, et al . 2024. AutoWe- bGLM: A Large Language Model-based Web Navigating Agent. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery ...
2024
-
[58]
HAO-PING HANK LEE, Yi-Shyuan Chiang, Lan Gao, Stephanie Yang, Philipp Winter, and Sauvik Das. 2025. Purpose Mode: Reducing Distraction Through Toggling Attention Capture Damaging Patterns on Social Media Websites. ACM Transactions on Computer-Human Interaction (2025)
2025
-
[59]
David Lindlbauer, Anna Maria Feit, and Otmar Hilliges. 2019. Context-Aware Online Adaptation of Mixed Reality Interfaces. InProceedings of the 32nd Annual ACM Symposium on User Interface Software and Technology (New Orleans, LA, USA) (UIST ’19). Association for Computing Machi...
2019
-
[60]
Evan Zheran Liu, Kelvin Guu, Panupong Pasupat, Tianlin Shi, and Percy Liang
-
[61]
Jiqun Liu, Matthew Mitsui, Nicholas J Belkin, and Chirag Shah. 2019. Task, information seeking intentions, and user behavior: Toward a multi-level un- derstanding of web search. In Proceedings of the 2019 conference on human information interaction and retrieval. 123–132
2019
-
[62]
Xing Han Lù, Zdeněk Kasner, and Siva Reddy. 2024. Weblinx: Real-world website navigation with multi-turn dialogue. arXiv preprint arXiv:2402.05930 (2024)
2024
-
[63]
Rongjun Ma, Henrik Lassila, Leysan Nurgalieva, and Janne Lindqvist. 2023. When browsing gets cluttered: exploring and modeling interactions of browsing clutter, browsing habits, and coping. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–29
2023
-
[64]
Eleni Michailidou, Simon Harper, and Sean Bechhofer. 2008. Investigating sighted users’ browsing behaviour to assist web accessibility. In Proceedings of the 10th international ACM SIGACCESS conference on Computers and accessibility. 121– 128
2008
-
[65]
Bryan Min, Matthew T Beaudouin-Lafon, Sangho Suh, and Haijun Xia. 2023. Demonstration of Masonview: Content-Driven Viewport Management. In Adjunct Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. 1–3
2023
-
[66]
arXiv preprint arXiv:1802.08802 (2018)
Reinforcement learning on web interfaces using workflow-guided explo- ration. arXiv preprint arXiv:1802.08802 (2018)
2018 arXiv
-
[67]
Nadine M Moacdieh and Nadine Sarter. 2017. Using eye tracking to detect the effects of clutter on visual search in real time. IEEE Transactions on Human-Machine Systems 47, 6 (2017), 896–902
2017
-
[68]
Alberto Monge Roffarello, Kai Lukoff, and Luigi De Russis. 2023. Defining and identifying attention capture deceptive designs in digital interfaces. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–19
2023
-
[69]
Vivian Genaro Motti and Jean Vanderdonckt. 2013. A computational frame- work for context-aware adaptation of user interfaces. In IEEE 7th International Conference on Research Challenges in Information Science (RCIS). 1–12. doi:10. 1109/RCIS.2013.6577709
2013
-
[70]
Alok Mysore and Philip J Guo. 2018. Porta: Profiling software tutorials using operating-system-wide activity tracing. In Proceedings of the 31st Annual ACM Symposium on User Interface Software and Technology. 201–212
2018
-
[71]
Peter O’Donovan, Aseem Agarwala, and Aaron Hertzmann. 2015. DesignScape: Design with Interactive Layout Suggestions. In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems (Seoul, Republic of Korea) (CHI ’15). Association for Computing Machinery...
2015
-
[72]
Bryan Min, Allen Chen, Yining Cao, and Haijun Xia. 2025. Malleable Overview- Detail Interfaces. arXiv preprint arXiv:2503.07782 (2025)
2025 arXiv
-
[73]
Yash Prakash, Akshay Kolgar Nayak, Mohan Sunkara, Sampath Jayarathna, Hae- Na Lee, and Vikas Ashok. 2024. All in One Place: Ensuring Usable Access to Online Shopping Items for Blind Users. Proceedings of the ACM on Human-Computer Interaction 8, EICS (2024), 1–25
2024
-
[74]
Yash Prakash, Mohan Sunkara, Hae-Na Lee, Sampath Jayarathna, and Vikas Ashok. 2023. AutoDesc: Facilitating Convenient Perusal of Web Data Items for Blind Users. In Proceedings of the 28th International Conference on Intelligent User Interfaces. 32–45
2023
-
[75]
Jokinen, Antti Oulasvirta, Zhenxin Wang, Chaklam Silpa- suwanchai, and Xiangshi Ren
Sayan Sarcar, Jussi P.P. Jokinen, Antti Oulasvirta, Zhenxin Wang, Chaklam Silpa- suwanchai, and Xiangshi Ren. 2018. Ability-Based Optimization of Touchscreen Interactions. IEEE Pervasive Computing 17, 1 (2018), 15–26. doi:10.1109/MPRV. 2018.011591058
2018
-
[76]
Tianlin Shi, Andrej Karpathy, Linxi Fan, Jonathan Hernandez, and Percy Liang
-
[77]
Jorge Sassaki Resende Silva, Paula Christina Figueira Cardoso, Raphael Winck- ler de Bettio, Daniela Cardoso Tavares, Carlos Alberto Silva, Willian Massami Watanabe, and André Pimenta Freire. 2024. In-Page Navigation Aids for Screen- Reader Users with Automatic Topicalisation ...
2024
-
[78]
Yichen Pan, Dehan Kong, Sida Zhou, Cheng Cui, Yifei Leng, Bing Jiang, Hangyu Liu, Yanyi Shang, Shuyan Zhou, Tongshuang Wu, et al. 2024. Webcanvas: Bench- marking web agents in online environments. arXiv preprint arXiv:2406.12373 (2024)
2024 arXiv
-
[79]
Hironobu Takagi, Chieko Asakawa, Kentarou Fukuda, and Junji Maeda. 2003. Accessibility designer: visualizing usability for the blind. ACM SIGACCESS accessibility and computing 77-78 (2003), 177–184
2003
-
[80]
Kashyap Todi, Jussi Jokinen, Kris Luyten, and Antti Oulasvirta. 2018. Familiari- sation: Restructuring layouts with visual learning models. In Proceedings of the 23rd International Conference on Intelligent User Interfaces. 547–558
2018
-
[81]
Ruolin Wang, Zixuan Chen, Mingrui Ray Zhang, Zhaoheng Li, Zhixiu Liu, Zihan Dang, Chun Yu, and Xiang’Anthony’ Chen. 2021. Revamp: Enhancing accessible information seeking experience of online shopping for blind or low vision users. In Proceedings of the 2021 CHI Conference on ...
2021
-
[82]
Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh J Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, et al . 2024. Os- world: Benchmarking multimodal agents for open-ended tasks in real computer environments. Advances in Neural Information Pro...
2024
-
[84]
Yaman Yu, Bektur Ryskeldiev, Ayaka Tsutsui, Matthew Gillingham, and Yang Wang. 2025. LLM-Driven Optimization of HTML Structure to Support Screen Reader Navigation. arXiv preprint arXiv:2502.18701 (2025)
2025 arXiv
-
[85]
Marc Sloan, Hui Yang, and Jun Wang. 2015. A term-based methodology for query reformulation understanding. Information Retrieval Journal 18 (2015), 145–165
2015
-
[86]
Yuhang Zhao, Elizabeth Kupferstein, Brenda Veronica Castro, Steven Feiner, and Shiri Azenkot. 2019. Designing AR visualizations to facilitate stair navigation for people with low vision. In Proceedings of the 32nd annual ACM symposium on user interface software and technology. 387–402
2019
-
[87]
Boyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun, and Yu Su. 2024. Gpt-4v (ision) is a generalist web agent, if grounded. arXiv preprint arXiv:2401.01614 (2024)
2024 arXiv
-
[88]
Boyuan Zheng, Boyu Gou, Scott Salisbury, Zheng Du, Huan Sun, and Yu Su. 2024. Webolympus: An open platform for web agents on live websites. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. 187–197
2024
-
[89]
Yu Zhong, TV Raman, Casey Burkhardt, Fadi Biadsy, and Jeffrey P Bigham. 2014. JustSpeak: enabling universal voice control on Android. In Proceedings of the 11th Web for All Conference. 1–4
2014
-
[90]
Dong Zhou, Ajay Chander, and Hiroshi Inamura. 2010. Optimizing user in- teraction for Web-based mobile tasks. In Proceedings of the 19th international conference on World wide web. 1333–1336
2010
-
[91]
entity":
Shuyan Zhou, Frank F Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, et al. 2023. Webarena: A realistic web environment for building autonomous agents. arXiv preprint arXiv:2307.13854 (2023). ASSETS ’25, October 26–29,...
2023 arXiv
-
[92]
Yuhang Zhao, Edward Cutrell, Christian Holz, Meredith Ringel Morris, Eyal Ofek, and Andrew D Wilson. 2019. SeeingVR: A set of tools to make virtual reality more accessible to people with low vision. In Proceedings of the 2019 CHI conference on human factors in computing systems. 1–14
2019
-
[2006]
In Proceedings of the WorkingConference on Advanced Visual Interfaces (Venezia, Italy) (AVI ’06)
Exploring the design space for adaptive graphical user interfaces. In Proceedings of the WorkingConference on Advanced Visual Interfaces (Venezia, Italy) (AVI ’06). Association for Computing Machinery, New York, NY, USA, 201–208. doi:10.1145/1133265.1133306
-
[2017]
In International Conference on Machine Learning
World of bits: An open-domain platform for web-based agents. In International Conference on Machine Learning. PMLR, 3135–3144
-
[2018]
In Proceedings of the 2018 CHI conference on human factors in computing systems
The dark (patterns) side of UX design. In Proceedings of the 2018 CHI conference on human factors in computing systems. 1–14
2018
-
[2021]
arXiv preprint arXiv:2104.08718 (2021)
Clipscore: A reference-free evaluation metric for image captioning. arXiv preprint arXiv:2104.08718 (2021)
2021 arXiv
-
[2022]
In CHI Conference on Human Factors in Computing Systems Extended Abstracts
Exploring challenges for conversational web browsing with blind and visually impaired users. In CHI Conference on Human Factors in Computing Systems Extended Abstracts. 1–7
-
[2024]
arXiv preprint arXiv:2402.04615 (2024)
Screenai: A vision-language model for ui and infographics understanding. arXiv preprint arXiv:2402.04615 (2024)
2024 arXiv
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.