Pith. sign in

REVIEW 3 major objections 5 minor 95 references

VeasyGuide: Personalized Visual Guidance for Low-vision Learners on Instructor Actions in Presentation Videos

T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The paper claims that VeasyGuide, a real-time motion-detection overlay that highlights and magnifies instructor pointing, marking, and sketching, lifts low-vision learners' action detection from 61% to 88% and cuts self-reported cognitive…

desk verdict A thoughtful cdesigned accessibility tool with real promise, but the headline detection gain rests on a between-video comparison that the paper's own analysis does not yet secure. read the letter →

arxiv 2507.21837 v2 pith:FKZL247U submitted 2025-07-29 cs.HC

classification cs.HC
keywords lowvisionpresentationvideosvideoaccessibilityvisualmotiondetectionco-designmagnificationuniversaldesign
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

VeasyGuide addresses a specific gap in educational presentation videos: instructors communicate by pointing, marking, and sketching on slides, but these visual actions are rarely described verbally, so low-vision learners must hunt for them or miss them. The paper claims that a real-time system which detects these actions from motion and renders a consistent, personalized highlight plus an optional auto-following zoom lets low-vision learners find what they are looking for, raising mean detection success from 61% to 88% and reducing self-reported cognitive load on every NASA-TLX dimension. The design was shaped by a co-design study with three low-vision participants and tested with eight low-vision and eight sighted viewers. The authors' position is that enhancing residual vision with guidance beats substituting for it with audio description, and that the same guidance can help sighted viewers stay attentive.

What carries the argument

The load-bearing mechanism is a two-stage activity recognition pipeline followed by a visualization module. The pipeline splits the video into shots with off-the-shelf shot detection, computes frame differences on one-third-second segments, keeps regions of change whose area exceeds $0.01\%$ of the frame, and builds a graph whose nodes are those regions, connected when they occur within 3 seconds and within 5% of the frame diagonal of each other. Edge weights come from Hu moments, seven shape-descriptor numbers that measure visual difference between regions, which lets the system merge or reject transient pointer trails. Connected components become activities, each rendered as a stationary, box-shaped, personalized highlight with a pointer icon and an optional auto-following zoom toggled with the Z key; a 1.5-second pre-activity trigger gives learners advance notice of where to look. The graph representation is what turns noisy per-frame motion into discrete, stable activities that can be highlighted and magnified.

What would settle it

Annotate a held-out set of presentation videos with the true locations of pointing, marking, and sketching actions, run VeasyGuide's pipeline on them, and compare the overlap of its activity boxes with the annotations; if low-vision users' success rate on those videos fails to beat the 61% baseline by a similar margin, the generalizability claim is not supported.

Watch

Extended reading notes

Core claim

The central claim is that the visual-search bottleneck for low-vision learners is not just the size of the content but knowing what to look for and where, and that an overlay can supply that knowledge in real time. On the paper's evidence, detection success rose from 61% to 88% (Welch's t-test $p<0.05$, Cohen's $d=1.30$), mean search time fell from 2.57s to 1.53s without reaching significance, and all NASA-TLX workload dimensions dropped significantly. In the guided condition no low-vision participant paused playback and fewer used screen magnification, suggesting the tool replaces part of the stop-and-search behavior of baseline viewing. The paper also reports that sighted viewers, already at ceiling on objective measures, subjectively preferred VeasyGuide for focus and comprehension, so the authors propose the mechanism as a broadly useful attention aid rather than only an accessibility tool.

Load-bearing premise

The load-bearing premise is that the fixed detection settings—how much motion counts as a region, how close regions must be in space and time to merge, and how short an activity can be—work on presentation videos beyond the six used in the study.

Editorial extensions

If this is right

  • A learner who misses roughly four in ten instructor actions could miss roughly one in ten with the default settings, a per-participant mean improvement of 74.5%.
  • The tool changes search strategy: in the guided condition no low-vision participant paused playback during the localization task, and screen-magnifier reliance dropped from six users to two.
  • The design implications—consistent familiar visuals, predictable spatial context including a 1.5-second pre-trigger, real-time personalization with immediate feedback, and preserving user agency—can guide future accessibility tools for visual search in video.
  • Because the pipeline uses lightweight motion detection and graph operations, the paper expects it to run on-device and to extend to live lectures and mainstream video platforms.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The localization task fixed highlight style to system defaults, so the 88% figure is a default-settings result; testing with participants' own personalized styles might change detection rates in either direction and would separate the effect of personalization from the effect of highlighting per se.
  • The pre-activity trigger likely explains part of the benefit; an ablation that removes the 1.5-second warning would isolate how much of the gain comes from anticipation versus visibility.
  • The graph representation is blind to semantic content, so the same machinery could be pointed at other moving targets, such as cursors in coding screencasts or moving regions in sports and remote collaboration, by retuning only the thresholds.
  • Since sighted viewers reported less clutter and better focus, the system may be a test bed for attention-as-accessibility design, where the same cue serves both perceptual and attentional functions.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper presents VeasyGuide, a web-based tool that uses motion detection to identify instructor actions (pointing, marking, sketching) in educational presentation videos and augments playback with personalized visual highlights and magnification. The system was developed through a co-design study with three low-vision (LV) participants. The evaluation involved 8 LV and 8 sighted participants in two tasks: a localization task (Task 1, videos V1–V4) measuring detection success and response time, and a viewing task (Task 2, videos V5–V6) measuring cognitive load and experience. The paper reports that VeasyGuide significantly improved LV users' detection of visual activities (61% to 88%), reduced NASA-TLX workload across all dimensions, and received positive feedback from sighted users, while detection time differences were not statistically significant. The main quantitative claim (RQ1) is based on an unpaired comparison of participants' success rates between conditions that used different video subsets, a design confound that the paper acknowledges but does not resolve.

Significance. If the reported effects hold, VeasyGuide addresses a real and under-explored accessibility barrier: LV learners missing instructor actions in slide-based videos. The co-design process and the derived design implications (DI1–DI4) are valuable and actionable, and the public web application makes the tool usable beyond the paper. The paper is honest about several limitations (Section 7), including the small sample and the lack of a systematic personalization evaluation in Task 1. The per-participant data in Table 7 are detailed and enable reanalysis, which is a strength. However, the headline success-rate claim currently rests on a between-video comparison, and the paper overstates the response-time finding, so the central contribution is not yet fully secured.

major comments (3)
  1. [§5.1.4, Table 7, Appendix E] The RQ1 success-rate analysis uses an unpaired Welch t-test on data that are paired at the participant level, but the baseline and VeasyGuide conditions used complementary subsets of videos V1–V4. Table 7 shows that the number of cued activities per condition differs within participants (e.g., L1: 11 vs 18; L7: 18 vs 11), and Appendix E reveals large per-video difficulty differences (V4 baseline success <25%). Consequently, the 61% vs 88% mean difference may be inflated or produced by which videos happened to appear in each condition rather than by the highlighting itself. Please reanalyze with a mixed-effects model that includes participant and video as random effects, report a stratified per-video comparison, or at minimum provide the paired per-participant difference and its confidence interval using the data in Table 7, and discuss how video difficulty is accounted for.
  2. [Abstract, §1, §5.2.2] The abstract and Section 1 claim that VeasyGuide yields 'faster response times' and 'improves the speed of detection,' but Section 5.2.2 states that the Mann-Whitney U test found no statistically significant difference (p > 0.05). Presenting the non-significant mean/median speed improvements as a demonstrated benefit overstates the evidence. Please either soften the speed claim to a non-significant trend or report a test that properly accounts for the paired design and video-level variation.
  3. [§4.1, §5.2.1] The activity recognition pipeline uses six hand-chosen thresholds (§4.1.1: RoC area 0.01% of frame, temporal closeness 3s, spatial closeness 5% of frame diagonal, Hu-moment merge threshold 0.5, minimum activity duration 5s, and 1.5s pre-trigger), but the paper reports no precision/recall evaluation of the detector on the six study videos. The user-success outcome in the VeasyGuide condition depends on the pipeline actually highlighting the cued activities; if recall is imperfect or the thresholds are implicitly tuned to the study videos, the 88% success rate is not a general property of the system. Please add a detection-accuracy evaluation (per-video recall/precision against the gold-standard activities) or a sensitivity analysis of the thresholds, and discuss the generalizability of these parameters to other presentation video styles.
minor comments (5)
  1. [§6.1.4] The statement that 'detection success gap reduced by 80.5%, and detection speed gap reduced by 27.3%' is not derived from the results in Section 5; please specify the formula (e.g., based on means or medians) or remove these numbers.
  2. [§5.1.3, Table 7] Please clarify how the total number of cued activities per condition in Table 7 was determined, including how V1–V4 were assigned to conditions for each participant and how activities shorter than one second were treated in the totals.
  3. [§5.1.1, Table 5] Two of the eight evaluation participants (L3 and L7) also participated in the co-design study; please acknowledge this overlap as a potential familiarity bias in the evaluation.
  4. [Figure 5] The horizontal axis of Figure 5 appears to be logarithmic; please label it as such or use a linear scale with clear units.
  5. [§4.2] The 1.5-second pre-activity highlight trigger is an important design element for predictability (DI3), but its isolated effect is not evaluated; please mention this explicitly in the limitations or future work.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: VeasyGuide's central claims rest on an empirical user study with an external video corpus, and no fitted parameter is renamed as a prediction.

full rationale

I walked the paper's claimed derivation chain and found no circular step that reduces a prediction to its inputs. The activity recognition pipeline uses pre-specified, hand-chosen thresholds (0.01% of frame area, 3 s temporal closeness, 5% of frame diagonal spatial closeness, Hu-moment merge threshold 0.5, 5 s minimum activity duration, 1.5 s pre-trigger); these constants are not fit to the user-study outcome, and the paper does not claim they were derived from the measured success rates. The RQ1 result (61% vs. 88% mean success) is an empirical between-condition comparison on six external videos from YouTube, DeepLearning.AI, and Khan Academy, with success measured by participant keypresses rather than by any equation in the paper. Likewise, the NASA-TLX and reaction-time results are measured outcomes, not constructions. The self-citation to Sechayk et al. [66] for the initial 'highlight any notable visual change' prototype is a design-inspiration citation; the evaluation does not depend on the correctness of that prior work, and no uniqueness or forced-choice argument is imported from it. The participation of L3 and L7 in both co-design and evaluation, and the fact that Task 1 split V1-V4 between conditions, are genuine internal-validity concerns about attribution and generalizability, but they are not circularity: they do not make the measured effect true by definition or by fitted-parameter renaming. The reported statistics are consistent with a real, if possibly confounded, empirical comparison. Under the hard rule that circularity must be exhibited as a specific reduction in the paper's own reasoning, the burden is not met.

Assumptions & free parameters 6 free parameters · 5 assumptions · 0 invented entities

The central evaluation rests on hand-chosen detection thresholds, standard motion-detection assumptions, and self-reported workload measures; none are fitted to the outcome, but all are load-bearing for generalization.

free parameters (6)
  • RoC area threshold = 0.01% of frame area
    Hand-chosen in Section 4.1.1 to filter compression artifacts; affects which motion regions become graph nodes.
  • Temporal closeness threshold = 3 seconds
    Hand-chosen in Section 4.1.1 to connect RoCs into activities; determines activity boundaries.
  • Spatial closeness threshold = 5% of frame diagonal
    Hand-chosen in Section 4.1.1; defines whether RoCs are spatially close enough to belong to the same activity.
  • Hu moments merge threshold = 0.5
    Hand-chosen in Section 4.1.1; controls which RoCs are merged to reduce transient noise.
  • Minimum activity duration = 5 seconds
    Set in the base prototype (Section 3.2.2) to define what counts as an activity for highlighting.
  • Pre-activity highlight trigger = 1.5 seconds
    Added from participant feedback (Section 3.2.2) to give advance notice; affects the timing of highlights in the study.
assumptions (5)
  • domain assumption Instructor actions such as pointing, marking, and sketching are frequent in educational presentation videos.
    Supported by the paper's own 300-video analysis (Appendix A), but that analysis used 6 raters with Likert ratings and no objective ground truth.
  • domain assumption Low-vision users prefer to use residual vision rather than audio description.
    Stated in Section 2.1 with citations to prior work; motivates the visual-enhancement approach.
  • domain assumption Frame differencing with contour detection captures instructor actions in screen-shared presentation videos.
    The core CV assumption in Section 4.1.1; not independently validated in this paper beyond the user study.
  • domain assumption NASA-TLX self-reports are a valid measure of cognitive load in this setting.
    Standard instrument used in Section 5.1.4, but self-reports are susceptible to demand characteristics in an accessibility tool evaluation.
  • domain assumption The co-design findings from 3 participants generalize to the broader low-vision population.
    The paper acknowledges the small co-design sample (Section 3) and that diverse visual abilities exist; the default settings may not suit everyone.

how reviews work

0 comments
Cite this review

Pith. "Pith review of VeasyGuide: Personalized Visual Guidance for Low-vision Learners on Instructor Actions in Presentation Videos." pith.science (2026). https://pith.science/paper/FKZL247U

@misc{pith2026250721837,
  author       = {Pith},
  title        = {Pith review of: VeasyGuide: Personalized Visual Guidance for Low-vision Learners on Instructor Actions in Presentation Videos},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FKZL247U}},
  note         = {Machine review of arXiv:2507.21837}
}
read the original abstract

Instructors often rely on visual actions such as pointing, marking, and sketching to convey information in educational presentation videos. These subtle visual cues often lack verbal descriptions, forcing low-vision (LV) learners to search for visual indicators or rely solely on audio, which can lead to missed information and increased cognitive load. To address this challenge, we conducted a co-design study with three LV participants and developed VeasyGuide, a tool that uses motion detection to identify instructor actions and dynamically highlight and magnify them. VeasyGuide produces familiar visual highlights that convey spatial context and adapt to diverse learners and content through extensive personalization and real-time visual feedback. VeasyGuide reduces visual search effort by clarifying what to look for and where to look. In an evaluation with 8 LV participants, learners demonstrated a significant improvement in detecting instructor actions, with faster response times and significantly reduced cognitive load. A separate evaluation with 8 sighted participants showed that VeasyGuide also enhanced engagement and attentiveness, suggesting its potential as a universally beneficial tool.

Figures

Figures reproduced from arXiv: 2507.21837 by the authors.

Figure 1
Figure 1. VeasyGuide’s interface includes: (a) a video player with highlighted areas (red rectangle and hand pointer), (a.1) zoom [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Overview of the VeasyGuide recognition pipeline. Input video segments are processed (a), and changes are detected [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Sample frames from some of the videos used in the [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (10 more)
Figure 4
Figure 4. Figure 4: Visual search success rates of LV participants (L1- [PITH_FULL_IMAGE:figures/full_fig_p009_4.png]
Figure 5
Figure 5. Figure 5: Visual search times of LV participants (L1-L8) in the [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]
Figure 6
Figure 6. Figure 6: Results of subjective evaluation of LV participants comparing the baseline to VeasyGuide. Blue tones indicate positive [PITH_FULL_IMAGE:figures/full_fig_p010_6.png]
Figure 7
Figure 7. Figure 7: Results of subjective evaluation of sighted users comparing the baseline to VeasyGuide. Blue tones indicate positive [PITH_FULL_IMAGE:figures/full_fig_p011_7.png]
Figure 8
Figure 8. Figure 8: The ratio of visual activities per domain. Any pre [PITH_FULL_IMAGE:figures/full_fig_p016_8.png]
Figure 9
Figure 9. Figure 9: The frequency of visual activities within videos [PITH_FULL_IMAGE:figures/full_fig_p016_9.png]
Figure 12
Figure 12. Figure 12: Normality of success rate distribution for LV par [PITH_FULL_IMAGE:figures/full_fig_p017_12.png]
Figure 13
Figure 13. Figure 13: Normality of detection speed distribution for LV [PITH_FULL_IMAGE:figures/full_fig_p017_13.png]
Figure 16
Figure 16. Figure 16: Detection speeds of LV participants in the local [PITH_FULL_IMAGE:figures/full_fig_p017_16.png]
Figure 14
Figure 14. Figure 14: Quiz completion times for both groups are shown, [PITH_FULL_IMAGE:figures/full_fig_p017_14.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

95 extracted references · 61 canonical work pages

  1. [1]

    Khan Academy. 2024. Elements and atoms | Atoms, compounds, and ions | Chemistry | Khan Academy. https://www.youtube.com/watch?v=IFKnq9QM6_A. Accessed: 2024-09-12

  2. [2]

    Khan Academy. 2024. England in the Age of Exploration. https://www.youtube. com/watch?v=DZxG2miqgZI. Accessed: 2024-09-12

  3. [3]

    Eler, Marouane Kessentini, and Ali Ouni

    Wajdi Aljedaani, Mohammed Alkahtani, Stephanie Ludi, Mohamed Wiem Mkaouer, Marcelo M. Eler, Marouane Kessentini, and Ali Ouni. 2023. The State of Accessibility in Blackboard: Survey and User Reviews Case Study. In Proceedings of the 20th International Web for All Conference (<conf-loc>, <city>Austin</city>, <state>TX</state>, <country>USA</country>, </con...

  4. [4]

    Ali Selman Aydin, Shirin Feiz, Vikas Ashok, and IV Ramakrishnan. 2020. Towards making videos accessible for low vision screen magnifier users. In Proceedings of the 25th International Conference on Intelligent User Interfaces (Cagliari, Italy) (IUI ’20) . Association for Computing Machinery, New York, NY, USA, 10–21. doi:10.1145/3377325.3377494

  5. [5]

    Patrick Baudisch, Desney Tan, Maxime Collomb, Dan Robbins, Ken Hinckley, Maneesh Agrawala, Shengdong Zhao, and Gonzalo Ramos. 2006. Phosphor: explaining transitions in the user interface using afterglow effects. InProceedings of the 19th annual ACM symposium on User interface software and technology . ACM, New York, NY, USA, 169–178

  6. [6]

    Porter, and I

    Syed Masum Billah, Vikas Ashok, Donald E. Porter, and I. V. Ramakrishnan

  7. [7]

    Christian Bognar. 2022. Naughty Dog’s Obsession With Yellow Explained . https://gamerant.com/naughty-dog-yellow-color-coding-environments- progression-design/ Accessed: 2025-04-17

  8. [8]

    Virginia Braun and Victoria Clarke. 2019. Reflecting on reflexive thematic analysis. Qualitative research in sport, exercise and health 11, 4 (2019), 589–597

Show all 95 references
  1. [9]

    Sheryl E Burgstahler and Rebecca C Cory. 2010. Universal design in higher education: From principles to practice . Harvard Education Press

  2. [10]

    Virgínia P Campos, Luiz MG Gonçalves, Wesnydy L Ribeiro, Tiago MU Araújo, Thaís G Do Rego, Pedro HV Figueiredo, Suanny FS Vieira, Thiago FS Costa, Caio C Moraes, Alexandre CS Cruz, et al. 2023. Machine generation of audio description for blind and visually impaired people.ACM ...

  3. [11]

    Ruei-Che Chang, Yuxuan Liu, and Anhong Guo. 2024. WorldScribe: Towards Context-Aware Live Visual Descriptions. In Proceedings of the 37th Annual ACM ASSETS ’25, October 26–29, 2025, Denver, CO, USA Sechayk et al. Symposium on User Interface Software and Technology . 1–18

  4. [12]

    Coursera. [n. d.]. Coursera. https://www.coursera.org/. Accessed: 2023-10-05

  5. [13]

    Cunningham, Joanne L

    Sheila J. Cunningham, Joanne L. Brebner, Francis Quinn, and David J. Turk. 2014. The Self-Reference Effect on Memory in Early Child- hood. Child Development 85, 2 (2014), 808–823. doi:10.1111/cdev.12144 arXiv:https://srcd.onlinelibrary.wiley.com/doi/pdf/10.1111/cdev.12144

  6. [14]

    and Adam D

    Robert D. and Adam D. 2024. PySceneDetect. https://github.com/Breakthrough/ PySceneDetect

  7. [15]

    DeepLearning.AI. [n. d.]. DeepLearning.AI. https://www.deeplearning.ai/. Ac- cessed: 2024-08-10

  8. [16]

    DeepLearningAI. 2024. #6 Machine Learning Specialization [Course 1, Week 1, Lesson 2]. https://www.youtube.com/watch?v=gG_wI_uGfIE. Accessed: 2024-09-12

  9. [17]

    Laurent Denoue, Scott Carter, Matthew Cooper, and John Adcock. 2013. Real-time direct manipulation of screen-based videos. InProceedings of the companion publi- cation of the 2013 international conference on Intelligent user interfaces companion . 43–44

  10. [18]

    NetworkX Developers. 2024. NetworkX. https://networkx.org/

  11. [19]

    Alfred T D’Agostino. 2021. Accessible teaching and learning in the undergraduate chemistry course and laboratory for blind and low-vision students. Journal of Chemical Education 99, 1 (2021), 140–147

  12. [20]

    doi:10.1145/3173574.3173594

  13. [21]

    Facebook

    Inc. Facebook. 2024. React - A JavaScript library for building user interfaces. https://reactjs.org. Accessed: 2024-09-12

  14. [22]

    Mirette Elias, Abi James, Edna Ruckhaus, Mari Carmen Suárez-Figueroa, Klaas An- dries De Graaf, Ali Khalili, Benjamin Wulff, Steffen Lohmann, and Sören Auer

  15. [23]

    In EC-TEL (Practitioner Proceedings)

    SlideWiki-Towards a Collaborative and Accessible Platform for Slide Pre- sentations.. In EC-TEL (Practitioner Proceedings). 1–3

  16. [24]

    Fox, Ahmad Ahmadzada, Clara T

    Dylan R. Fox, Ahmad Ahmadzada, Clara T. Friedman, Shiri Azenkot, Marlena A. Chu, Roberto Manduchi, and Emily A. Cooper. 2023. Using augmented reality to cue obstacles for people with low vision. Opt. Express 31, 4 (Feb 2023), 6827–6848. doi:10.1364/OE.479258

  17. [25]

    Tang, and Thomas Jaeger

    Danyang Fan, Sasa Junuzovic, John C. Tang, and Thomas Jaeger. 2023. Improv- ing the Accessibility of Screen-Shared Presentations by Enabling Concurrent Exploration. In Proceedings of the 25th International ACM SIGACCESS Conference on Computers and Accessibility, ASSETS 2023, N...

  18. [26]

    Sue Grice and Janet Hughes. 2009. Can music and animation improve the flow and attainment in online learning?Journal of Educational Multimedia and Hypermedia 18, 4 (2009), 385–403

  19. [27]

    Python Software Foundation. 2024. Python Programming Language. https: //www.python.org/

  20. [28]

    Taralynn Hartsell and Steve Chi-Yin Yuen. 2006. Video streaming in online learning. AACE Review (Formerly AACE Journal) 14, 1 (2006), 31–43

  21. [29]

    Sourojit Ghosh and Andrea Figueroa. 2023. Establishing TikTok as a Platform for Informal Learning: Evidence from Mixed-Methods Analysis of Creators and Viewers. In 56th Hawaii International Conference on System Sciences, HICSS 2023, Maui, Hawaii, USA, January 3-6, 2023, Tung X...

  22. [30]

    Maija Hirvonen, Marika Hakola, and Michael Klade. 2023. Co-translation, consul- tancy and joint authorship: User-centred translation and editing in collaborative audio description. Journal of Specialised Translation 39 (2023), 26–51

  23. [31]

    Tengda Han, Max Bain, Arsha Nagrani, Gül Varol, Weidi Xie, and Andrew Zisser- man. 2023. AutoAD: Movie description in context. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 18930–18940

  24. [32]

    Gwo-Jen Hwang, Li-Hsueh Yang, and Sheng-Yuan Wang. 2013. A concept map- embedded educational computer game for improving students’ learning perfor- mance in natural science courses. Computers & Education 69 (2013), 121–130

  25. [33]

    Alfonso Juan Hinojosa. 2015. Investigations on the impact of spatial ability and scientific reasoning of student comprehension in physics, state assessment tests, and STEM courses. The University of Texas at Arlington, Arlington, TX, USA

  26. [34]

    Julie A Jacko and Andrew Sears. 1998. Designing interfaces for an overlooked user group: Considering the visual profiles of partially sighted users. In Proceedings of the third international ACM conference on Assistive technologies . 75–77

  27. [35]

    Ming-Kuei Hu. 1962. Visual pattern recognition by moment invariants. IRE transactions on information theory 8, 2 (1962), 179–187

  28. [36]

    Ji, Brianna R

    Tiger F. Ji, Brianna R. Cochran, and Yuhang Zhao. 2022. VRBubble: Enhancing Peripheral Awareness of Avatars for People with Visual Impairments in Social Virtual Reality. In Proceedings of the 24th International ACM SIGACCESS Confer- ence on Computers and Accessibility, ASSETS ...

  29. [37]

    Touhidul Islam and Syed Masum Billah

    Md. Touhidul Islam and Syed Masum Billah. 2023. SpaceX Mag: An Automatic, Scalable, and Rapid Space Compactor for Optimizing Smartphone App Interfaces for Low-Vision Users. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 7, 2 (2023), 59:1–59:36. doi:10.1145/3596253

  30. [38]

    Lucy Jiang and Richard Ladner. 2022. Co-Designing Systems to Support Blind and Low Vision Audio Description Writers. In Proceedings of the 24th International ACM SIGACCESS Conference on Computers and Accessibility (Athens, Greece) (ASSETS ’22). Association for Computing Machin...

  31. [39]

    Gaurav Jain, Basel Hindi, Connor Courtien, Conrad Wyrick, Xin Yi Therese Xu, Michael C Malcolm, and Brian A. Smith. 2023. Towards Accessible Sports Broadcasts for Blind and Low-Vision Viewers. In Extended Abstracts of the 2023 CHI Conference on Human Factors in Computing Syste...

  32. [40]

    Hyeungshik Jung, Hijung Valentina Shin, and Juho Kim. 2018. Dynamicslide: Exploring the design space of reference-based interaction techniques for slide- based lecture videos. In Proceedings of the 2018 Workshop on Multimedia for Accessible Human Computer Interface . ACM, New ...

  33. [41]

    Lucy Jiang, Crescentia Jung, Mahika Phutane, Abigale Stangl, and Shiri Azenkot

  34. [42]

    Junhan Kong, Dena Sabha, Jeffrey P Bigham, Amy Pavel, and Anhong Guo. 2021. TutorialLens: authoring Interactive augmented reality tutorials through narration and demonstration. In Proceedings of the 2021 ACM Symposium on Spatial User Interaction. 1–11

  35. [43]

    Richard E Ladner and Kyle Rector. 2017. Making your presentation accessible. Interactions 24, 4 (2017), 56–59

  36. [44]

    Lucy Jiang, Mahika Phutane, and Shiri Azenkot. 2023. Beyond audio descrip- tion: Exploring 360 video accessibility with blind and low vision users through collaborative creation. In Proceedings of the 25th international ACM SIGACCESS conference on computers and accessibility . 1–17

  37. [45]

    Open Source Computer Vision Library. 2024. OpenCV. https://opencv.org/

  38. [46]

    Khan Academy. [n. d.]. Khan Academy. https://www.khanacademy.org/. Ac- cessed: 2024-08-10

  39. [47]

    Chuan-Yu Mo, Chengliang Wang, Jian Dai, and Peiqi Jin. 2022. Video playback speed influence on learning effect from the perspective of personalized adaptive learning: A study based on cognitive load theory. Frontiers in Psychology 13 (2022), 839982

  40. [48]

    Toni-Jan Keith Palma Monserrat, Shengdong Zhao, Kevin McGee, and An- shul Vikram Pandey. 2013. Notevideo: Facilitating navigation of blackboard- style lecture videos. In Proceedings of the SIGCHI conference on human factors in computing systems. 1139–1148

  41. [49]

    Susan J Leat, Gordon E Legge, and Mark A Bullimore. 1999. What is low vision? A re-evaluation of definitions. Optometry and Vision Science 76, 4 (1999), 198–211

  42. [50]

    Rosiana Natalie, Ruei-Che Chang, Smitha Sheshadri, Anhong Guo, and Kotaro Hara. 2024. Audio description customization. In Proceedings of the 26th Interna- tional ACM SIGACCESS Conference on Computers and Accessibility . 1–19

  43. [51]

    Elke Mattheiss, Georg Regal, David Sellitsch, and Manfred Tscheligi. 2017. User- centred design with visually impaired pupils: A case study of a game editor for orientation and mobility training. International Journal of Child-Computer Interaction 11 (2017), 12–18. doi:10.1016...

  44. [52]

    Evelyn Navarrete, Andreas Nehring, Sascha Schanze, Ralph Ewerth, and Anett Hoppe. 2023. A Closer Look into Recent Video-based Learning Research: A Comprehensive Review of Video Characteristics, Tools, Technologies, and Learn- ing Effectiveness. CoRR abs/2301.13617 (2023). doi:...

  45. [53]

    Zheng Ning, Brianna L Wimer, Kaiwen Jiang, Keyi Chen, Jerrick Ban, Yapeng Tian, Yuhang Zhao, and Toby Jia-Jun Li. 2024. SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision Viewers. In Proceedings of the CHI Conference on Hu...

  46. [54]

    Zahra J Muhsin, Rami Qahwaji, Faruque Ghanchi, and Majid Al-Taee. 2024. Review of substitutive assistive tools and technologies for people with visual impairments: recent advancements and prospects. Journal on Multimodal User Interfaces 18, 1 (2024), 135–156

  47. [55]

    Jaclyn Packer, Katie Vizenor, and Joshua A Miele. 2015. An overview of video description: history, benefits, and guidelines. Journal of Visual Impairment & Blindness 109, 2 (2015), 83–93

  48. [56]

    Rosiana Natalie, Jolene Loh, Huei Suen Tan, Joshua Tseng, Ian Luke Yi-Ren Chan, Ebrima H Jarjue, Hernisa Kacorri, and Kotaro Hara. 2021. The efficacy of collaborative authoring of video scene descriptions. In Proceedings of the 23rd International ACM SIGACCESS Conference on Co...

  49. [57]

    Amy Pavel, Gabriel Reyes, and Jeffrey P Bigham. 2020. Rescribe: Authoring and automatically editing audio descriptions. In Proceedings of the 33rd Annual ACM Symposium on User Interface Software and Technology . 747–759

  50. [59]

    American Council of the Blind. n.d.. Audio Description Project, Guidelines for Audio Describers. https://www.acb.org/adp/guidelines.html. Accessed: 2024-08- 10

  51. [60]

    Yash Prakash, Akshay Kolgar Nayak, Sampath Jayarathna, Hae-Na Lee, and Vikas Ashok. 2024. Understanding Low Vision Graphical Perception of Bar Charts. In Proceedings of the 26th International ACM SIGACCESS Conference on Computers and Accessibility. 1–10

  52. [61]

    Suraiya Parveen and Javeria Shah. 2021. A motion detection system in python and opencv. In 2021 third international conference on intelligent communication technologies and virtual mobile networks (ICICV) . IEEE, IEEE, Virtual Conference, VeasyGuide ASSETS ’25, October 26–29, ...

  53. [62]

    Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023. Robust speech recognition via large-scale weak supervision. In International conference on machine learning . PMLR, 28492–28518

  54. [63]

    Ashwin Ram, Han Xiao, Shengdong Zhao, and Chi-Wing Fu. 2023. VidAdapter: Adapting Blackboard-Style Videos for Ubiquitous Viewing. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 7, 3, Article 119 (sep 2023), 19 pages. doi:10. 1145/3610928

  55. [64]

    Yi-Hao Peng, JiWoong Jang, Jeffrey P Bigham, and Amy Pavel. 2021. Say it all: Feedback for improving non-visual presentation accessibility. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems . 1–12

  56. [65]

    Anastasia Schaadhardt, Alexis Hiniker, and Jacob O Wobbrock. 2021. Understand- ing blind screen-reader users’ experiences of digital artboards. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems . 1–19

  57. [66]

    Pallets Projects. 2024. Flask - A micro web framework for Python. https://flask. palletsprojects.com. Accessed: 2024-09-12

  58. [67]

    Joel Snyder. 2005. Audio description: The visual made verbal. In International congress series, Vol. 1282. Elsevier, 935–939

  59. [68]

    Abigale Stangl, Shasta Ihorn, Yue-Ting Siu, Aditya Bodi, Mar Castanon, Lothar D Narins, and Ilmi Yoon. 2023. The Potential of a Visual Dialogue Agent In a Tandem Automated Audio Description System for Videos. In Proceedings of the 25th International ACM SIGACCESS Conference on...

  60. [69]

    Andreas Sackl, Franziska Graf, Raimund Schatz, and Manfred Tscheligi. 2020. Ensuring accessibility: Individual video playback enhancements for low vision users. In Proceedings of the 22nd International ACM SIGACCESS Conference on Computers and Accessibility. 1–4

  61. [70]

    Lee Stearns, Leah Findlater, and Jon E Froehlich. 2018. Design of an augmented re- ality magnification aid for low vision users. InProceedings of the 20th international ACM SIGACCESS conference on computers and accessibility . 28–39

  62. [71]

    Yotam Sechayk, Ariel Shamir, and Takeo Igarashi. 2024. SmartLearn: Visual- Temporal Accessibility for Slide-based e-learning Videos. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, CHI EA 2024, Honolulu, HI, USA, May 11-16, 2024, Florian ’Flo...

  63. [72]

    Sarit Felicia Anais Szpiro, Shafeka Hashash, Yuhang Zhao, and Shiri Azenkot

  64. [73]

    Meini Tang, Roberto Manduchi, Susana Chung, and Raquel Prado. 2023. Screen Magnification for Readers with Low Vision: A Study on Usability and Perfor- mance. In Proceedings of the 25th International ACM SIGACCESS Conference on Computers and Accessibility. 1–15

  65. [74]

    Fleischmann, Meredith Ringel Morris, and Danna Gurari

    Abigale Stangl, Nitin Verma, Kenneth R. Fleischmann, Meredith Ringel Morris, and Danna Gurari. 2021. Going Beyond One-Size-Fits-All Image Descriptions to Satisfy the Information Wants of People Who are Blind or Have Low Vision. In Proceedings of the 23rd International ACM SIGA...

  66. [75]

    Chris Thompson

    Dr. Chris Thompson. 2024. Introduction to Neuroscience 2: Lecture 26: Neu- roethology. https://www.youtube.com/watch?v=7pQyw1rHEEg. Accessed: 2024-09-12

  67. [76]

    Satoshi Suzuki et al. 1985. Topological structural analysis of digitized binary images by border following. Computer vision, graphics, and image processing 30, 1 (1985), 32–46

  68. [77]

    The Organic Chemistry Tutor. 2024. Biology - Intro to Cell Structure - Quick Review! https://www.youtube.com/watch?v=vwAJ8ByQH2U. Accessed: 2024- 09-12

  69. [78]

    Ru Wang, Zach Potter, Yun Ho, Daniel Killough, Linxiu Zeng, Sanbrita Mondal, and Yuhang Zhao. 2024. GazePrompt: Enhancing Low Vision People’s Reading Experience with Gaze-Aware Augmentations. InProceedings of the CHI Conference on Human Factors in Computing Systems, CHI 2024, ...

  70. [79]

    Ru Wang, Linxiu Zeng, Xinyong Zhang, Sanbrita Mondal, and Yuhang Zhao. 2023. Understanding how low vision people read using eye tracking. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . ACM, New York, NY, USA, 1–17

  71. [80]

    ThatsEngineering. 2024. Frame Assignment For Robotic Manipulators - Direct Kinematics I. https://www.youtube.com/watch?v=fNIyNF87q9I. Accessed: 2024- 09-12

  72. [81]

    Yanan Wang, Yuhang Zhao, and Yea-Seul Kim. 2024. How Do Low-Vision Individ- uals Experience Information Visualization?. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–15

  73. [82]

    Ayaka Tsutsui, Kenta Yamamoto, Yinan Zhao, Ippei Suzuki, Kengo Tanaka, and Yoichi Ochiai. 2024. Low Vision Boxing: Participatory Design of Adap- tive Kickboxing Experiences with Low Vision Person. In The 26th Interna- tional ACM SIGACCESS Conference on Computers and Accessibil...

  74. [83]

    YouTube. [n. d.]. YouTube. https://www.youtube.com/. Accessed: 2024-08-10

  75. [84]

    Beste F Yuksel, Pooyan Fazli, Umang Mathur, Vaishali Bisht, Soo Jung Kim, Joshua Junhee Lee, Seung Jung Jin, Yue-Ting Siu, Joshua A Miele, and Ilmi Yoon

  76. [85]

    Yuhang Zhao, Elizabeth Kupferstein, Brenda Veronica Castro, Steven Feiner, and Shiri Azenkot. 2019. Designing AR Visualizations to Facilitate Stair Navigation for People with Low Vision. InProceedings of the 32nd Annual ACM Symposium on User Interface Software and Technology, ...

  77. [86]

    Yujia Wang, Wei Liang, Haikun Huang, Yongqi Zhang, Dingzeyu Li, and Lap-Fai Yu. 2021. Toward automatic audio description generation for accessible videos. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems . 1–12

  78. [87]

    Yuhang Zhao, Sarit Szpiro, Jonathan Knighten, and Shiri Azenkot. 2016. CueSee: exploring visual cues for people with low vision to facilitate a visual search task. In Proceedings of the 2016 ACM International Joint Conference on Pervasive and Ubiquitous Computing (Heidelberg, ...

  79. [88]

    Carmen Yip, Jie Mi Chong, Sin Yee Kwek, Yong Wang, and Kotaro Hara. 2021. Visionary Caption: Improving the Accessibility of Presentation Slides through Highlighting Visualization. In Proceedings of the 23rd International ACM SIGAC- CESS Conference on Computers and Accessibility . 1–4

  80. [89]

    Zoom. [n. d.]. Zoom. https://zoom.us/. Accessed: 2024-08-10. A Empirical Evaluation of Visual Activities in Presentation Videos To better understand how visual activities appear in educational presentation videos in the wild, we conducted an analysis of 300 YouTube [83] videos...

  81. [93]

    Yuhang Zhao, Elizabeth Kupferstein, Hathaitorn Rojnirun, Leah Findlater, and Shiri Azenkot. 2020. The Effectiveness of Visual and Audio Wayfinding Guid- ance on Smartglasses for People with Low Vision. In CHI ’20: CHI Confer- ence on Human Factors in Computing Systems, Honolul...

  82. [95]

    Yuhang Zhao, Sarit Szpiro, Lei Shi, and Shiri Azenkot. 2020. Designing and Evaluating a Customizable Head-mounted Vision Enhancement System for Peo- ple with Low Vision. ACM Trans. Access. Comput. 12, 4 (2020), 15:1–15:46. doi:10.1145/3361866

  83. [2016]

    InProceedings of the 18th International ACM SIGACCESS Conference on Computers and Accessibility

    How people with low vision access computing devices: Understanding chal- lenges and opportunities. InProceedings of the 18th International ACM SIGACCESS Conference on Computers and Accessibility . 171–180

  84. [2018]

    In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems, CHI 2018, Montreal, QC, Canada, April 21-26, 2018 , Regan L

    SteeringWheel: A Locality-Preserving Magnification Interface for Low Vision Web Browsing. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems, CHI 2018, Montreal, QC, Canada, April 21-26, 2018 , Regan L. Mandryk, Mark Hancock, Mark Perry, and Anna L...

  85. [2020]

    In Proceedings of the 2020 ACM Designing Interactive Systems Conference

    Human-in-the-loop machine learning to increase video accessibility for visually impaired and blind users. In Proceedings of the 2020 ACM Designing Interactive Systems Conference. 47–60

  86. [2023]

    doi:10.1145/3597638.3608411

    ACM, 44:1–44:16. doi:10.1145/3597638.3608411

  87. [2024]

    It’s Kind of Context Dependent

    "It’s Kind of Context Dependent": Understanding Blind and Low Vision People’s Video Accessibility Preferences Across Viewing Scenarios. In Proceed- ings of the CHI Conference on Human Factors in Computing Systems, CHI 2024, Honolulu, HI, USA, May 11-16, 2024, Florian ’Floyd’ M...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.