REVIEW 2 major objections 5 minor 108 references
Describe Now: User-Driven Audio Description for Blind and Low Vision Individuals
T0 review · 2 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Researchers find that user-driven AI audio descriptions give blind and low vision viewers a greater sense of control and active engagement, but impose higher cognitive workload and fear of missing out.
desk verdict Solid qualitative user study; the concise-vs-detailed preference numbers are confounded by description length, but the control and cognitive-load findings hold. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The prototype is an extended audio description interface: pressing C or D on the keyboard pauses the video, speaks a pre-generated concise (about 25 words) or detailed (about 100 words) description of the most recent second, and resumes playback. Descriptions were generated once per second for each of the seven videos by prompting GPT-4 Vision with 42 curated professional audio description guidelines, so that timing and content were under the viewer's control while the experiment held generation latency and nondeterminism constant. The per-second pre-generation and the two-level detail distinction are what allow clean measurement of activation frequency and type as dependent variables.
What would settle it
Record activation intervals for multiple videos per genre matched for pacing and speech density (or manipulate pacing of the same content). If average activation intervals no longer track genre labels once pace is controlled, the genre-specific frequency finding collapses; the qualitative control/load themes would remain unaffected.
Extended reading notes
Core claim
The central claim is that user-driven AI-generated audio descriptions are a viable and qualitatively different mode of video access for BLV people: viewers value the agency to decide when a description arrives and how detailed it is, likening the experience to having a live human describer. In a 20-participant study, activation intervals varied by genre, with Film and Animation receiving a description on average every 5.9 seconds versus every 12.3 seconds for Education, and concise descriptions were activated more often than detailed ones. Through thematic analysis of interviews, the paper identifies three themes: the sense of control and active engagement user-driven descriptions create, paired with cognitive load and FOMO; the dependence of format preference on video content, viewing context, and individual differences; and the benefits and drawbacks of AI-generated descriptions, including hallucinations, missing on-screen text in concise versions, and the desire for faster, customizable text-to-speech voices.
Load-bearing premise
The quantitative genre differences rest on the assumption that each of the seven chosen videos is representative of its entire genre, even though videos differed in length, pacing, and amount of speech.
Editorial extensions
If this is right
- If user-driven ADs are adopted, BLV viewers can skim, preview, and re-watch videos by requesting context only when they need it, expanding accessible content beyond professionally described films and TV.
- Differences in activation frequency across genres, if confirmed with more videos, imply that default description timing should adapt to content type, with faster-paced entertainment needing denser coverage than instructional or educational material.
- The observed cognitive load and FOMO suggest that a successful on-demand AD system should combine user control with cues that tell viewers when a description is available, and offer adjustable text-to-speech speed and voice.
- The findings imply a role shift for describers, from authoring every description to setting insertion timings, verifying accuracy, and removing hallucinations, while BLV users can become curators who save, edit, and share descriptions.
Reading between the lines
- The central UX trade-off (control vs. cognitive load) likely generalizes beyond this prototype to any interrupt-driven assistive system: whenever a user must decide to request information during a continuous task, the cost of monitoring and timing the request falls on the user. The authors' proposed audio/haptic cues are one concrete remedy, but a system that could predict when a viewer wants desc
- Because participants preferred different formats for the same genre depending on whether they were watching alone, with sighted people, or re-watching, a production AD platform may need per-viewer, per-context presets rather than one global setting; this is an unexplored design space that follows directly from the paper's context-dependency theme.
- A testable extension: compare activation times against an automatically computed "description demand" signal (scene-change rate, speech density, on-screen text presence) to see whether the genre differences in activation intervals are explained by low-level video properties instead of genre semantics; if so, the quantitative claim becomes a claim about pacing, not genre categories.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper introduces and evaluates user-driven AI-generated audio descriptions (ADs) for blind and low vision (BLV) viewers, in which the viewer decides when to receive a description and chooses between a concise (25-word) and a detailed (100-word) level. The authors pre-generated descriptions with GPT-4V for seven short videos spanning different genres, built an interface that pauses playback while a text-to-speech model reads the selected description, and conducted a study with 20 BLV participants. Quantitative measures include Likert ratings and logged activation counts and intervals; qualitative analysis of post-session interviews yielded three themes: an increased sense of control and active watching accompanied by higher cognitive load and fear of missing out (FOMO); preferences that depend on video content, viewing context, and individual differences; and positive and negative aspects of user-driven AI descriptions, including TTS considerations. The paper reports that concise descriptions were activated more often than detailed ones in all seven videos, that activation intervals varied by genre (e.g., Film & Animation mean 5.9s vs Education mean 12.3s), and that ratings for efficiency and effectiveness were medially positive while enjoyability was neutral. The authors position the work as an empirical exploration of an interactive AD paradigm and discuss implications for future AD platforms, the changing roles of describers and BLV users, and multisensory interaction.
Significance. If the findings hold, this is a useful empirical contribution to accessible video consumption: it provides a concrete alternative to pre-recorded AD in which BLV users control both timing and detail, with qualitative evidence about benefits (control, flexibility, applicability to diverse content) and costs (cognitive load, FOMO, disruption of video flow). The study has several strengths: two coders independently analyzed interview transcripts following a recognized thematic analysis approach; the participant sample spans diverse visual impairments, ages, and screen-reader experience; and descriptions were pre-generated to control for MLLM latency and non-determinism, which is a thoughtful design choice for a first-user study. The qualitative themes are well supported by participant quotes and are the strongest contribution. The quantitative findings are clearly labeled as summary statistics with explicit disclaimers about generalizability, which is appropriate given the one-video-per-genre design and single session.
major comments (2)
- [5.1.2, with 3.2 and 3.3] The comparison of concise versus detailed activation counts is presented as evidence about preferred level of detail, but the counts likely reflect interruption cost rather than a clean preference signal. Detailed descriptions have a 100-word cap and concise descriptions a 25-word cap (Section 3.2), and the interface pauses playback while the TTS reads the full description (Section 3.3). The listening time for a detailed description is therefore approximately four times that of a concise one, so each detailed activation substantially increases viewing time. Section 5.1.2 states that concise descriptions were activated more often (mean 5.42 vs 3.58) in all videos, and the paper frames this as 'preferred level of detail.' Yet Section 5.1.1 already notes that the extended presentation reduced enjoyability, and Theme 5.2.2 explicitly ties concise choices to 'minimally disrupting the video flow.' Section 6.4 lists limitations but does not mention this TTS-duration confound. Please reanalyze the data using activation counts per unit of description listening time, or report total time spent listening to each type, or reframe the claim so that it describes activation frequency rather than level-of-detail preference.
- [5.1.2 and 6.4] The genre-level differences in activation intervals (e.g., Film & Animation mean 5.9s vs Education mean 12.3s) are based on one researcher-selected video per genre, and those videos differ in length, amount of speech, and pacing. The paper correctly acknowledges this in Section 6.4 ('The pace of the video likely impacted the quantitative results'), so this is a generalizability limitation rather than an internal inconsistency. However, because the abstract and Section 6 state that there are differences 'for different videos' and Q2 asks about genre differences, readers may overgeneralize. Please make the single-video-per-genre limitation explicit at the first quantitative presentation (Section 5.1.2) and temper the language in the abstract and Section 6 accordingly.
minor comments (5)
- [5.1.1] The Friedman test is mentioned but no test statistic, degrees of freedom, or p-value is reported; either report the result in full or omit the mention of the test.
- [6, first paragraph] The sentence 'the time interval between descriptions varied significantly across genres' uses the word 'significantly' even though no inferential tests were run; consider replacing it with 'markedly' or 'substantially' to prevent a statistical misreading.
- [Figure 6] The figure caption labels participants as 'blind' and 'low-vision,' but the participant IDs (P4, P3, P11, P19) do not match the subscripted notation (P4_B, P3_LV, P11_LV, P19_B) used in Table 1; please align the notation for consistency.
- [5.2.3] The phrase 'implored too much background information' appears to be a word-choice error; 'included too much background information' or 'contained too much background information' would be clearer.
- [References] Reference [48] contains what appears to be a garbled URL fragment ('danielpatt321@gmail.com'); please verify and correct the URL.
Circularity Check
No significant circularity: the paper's claims are empirical measurements from a user study, not derived quantities.
full rationale
The central claims are empirical observations from a 20-participant user study: activation counts, activation intervals, Likert ratings, and thematic analysis of interviews. There is no fitted parameter renamed as a prediction, no equation that is equivalent to an input by construction, and no uniqueness theorem imported from the authors' prior work to force a choice. The quantitative genre comparisons (e.g., mean activation intervals in Figure 5) are summary statistics of logged key presses; the authors explicitly refrain from statistical tests and acknowledge in Section 6.4 that video pace likely impacted the results. The concise-vs-detailed activation comparison is an empirical finding that may be confounded by the longer TTS duration of detailed descriptions, but a confound is a validity threat, not circularity: the observed counts are not derived from the definition of concise/detailed. Self-citations such as VideoA11y [29] are used as background support for prompting MLLMs with AD guidelines and do not carry the central claim; the central user-experience findings stand on the study's own interview and log data. The paper is self-contained against external benchmarks in the sense that it reports new user data rather than deriving conclusions from prior outputs. Accordingly, no circular step can be exhibited with a specific reduction, and the correct verdict is no significant circularity.
Assumptions & free parameters
free parameters (2)
- Concise description word cap =
25 words
- Detailed description word cap =
100 words
assumptions (3)
- domain assumption One short video per genre is representative of that genre for AD timing and detail preferences.
- domain assumption Pre-generated GPT-4V descriptions are a valid stand-in for real-time user-driven AD generation.
- domain assumption Participant self-report (Likert ratings and interview responses) reflects actual experience.
Cite this review
Pith. "Pith review of Describe Now: User-Driven Audio Description for Blind and Low Vision Individuals." pith.science (2026). https://pith.science/paper/L3IVCJ7C
@misc{pith2026241111835,
author = {Pith},
title = {Pith review of: Describe Now: User-Driven Audio Description for Blind and Low Vision Individuals},
year = {2026},
howpublished = {\url{https://pith.science/paper/L3IVCJ7C}},
note = {Machine review of arXiv:2411.11835}
}
read the original abstract
Audio descriptions (AD) make videos accessible for blind and low vision (BLV) users by describing visual elements that cannot be understood from the main audio track. AD created by professionals or novice describers is time-consuming and offers little customization or control to BLV viewers on description length and content and when they receive it. To address this gap, we explore user-driven AI-generated descriptions, enabling BLV viewers to control both the timing and level of detail of the descriptions they receive. In a study, 20 BLV participants activated audio descriptions for seven different video genres with two levels of detail: concise and detailed. Our findings reveal differences in the preferred frequency and level of detail of ADs for different videos, participants' sense of control with this style of AD delivery, and its limitations. We discuss the implications of these findings for the development of future AD tools for BLV users.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
-
[1]
3Play Media. 2024. Audio Description: What It Is and How It Works. https: //www.3playmedia.com/learn/popular-topics/audio-description/
2024
-
[2]
Nayyer Aafaq, Ajmal Mian, Wei Liu, Syed Zulqarnain Gilani, and Mubarak Shah
-
[3]
Damien Ablart, Carlos Velasco, and Marianna Obrist. 2017. Integrating Mid- Air Haptics into Movie Experiences. In Proceedings of the ACM International Conference on Interactive Experiences for TV and Online Video (TVX) . 77–84
2017
-
[4]
American Council of the Blind. 2020. Beginner’s Guide to Audio Description. https://adp.acb.org/docs/Beginners-Guide-to-Audio-Description.pdf
2020
-
[5]
Anthropic. 2024. Meet Claude. https://www.anthropic.com/claude
2024
-
[6]
Aditya Bodi, Pooyan Fazli, Shasta Ihorn, Yue-Ting Siu, Andrew T Scott, Lothar Narins, Yash Kant, Abhishek Das, and Ilmi Yoon. 2021. Automated Video De- scription for Blind and Low Vision Users. In Proceedings of the ACM SIGCHI Conference Extended Abstracts on Human Factors in Computing Systems (CHI)
2021
-
[7]
Carmen Branje and Deborah Fels. 2012. LiveDescribe: Can Amateur Describers Create High-Quality Audio Description?Journal of Visual Impairment & Blindness 106 (03 2012), 154–165
2012
-
[8]
Virginia Campos, Tiago Araujo, Guido Souza Filho, and Luiz Gonçalves. 2020. CineAD: A System for Automated Audio Description Script Generation for the Visually Impaired. Universal Access in the Information Society 19 (03 2020)
2020
Show all 108 references
-
[9]
Luis Cavazos Quero, Jorge Iranzo Bartolomé, and Jundong Cho. 2021. Accessible Visual Artworks for Blind and Visually Impaired People: Comparing a Multimodal Approach with Tactile Graphics. Electronics 10, 3 (2021)
2021
-
[10]
Ruei-Che Chang, Yuxuan Liu, and Anhong Guo. 2024. WorldScribe: Towards Context-Aware Live Visual Descriptions. In Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology (Pittsburgh, PA, USA)(UIST ’24). Association for Computing Machinery, New Yo...
2024
-
[11]
Ruei-Che Chang, Chao-Hsien Ting, Chia-Sheng Hung, Wan-Chen Lee, Liang- Jin Chen, Yu-Tzu Chao, Bing-Yu Chen, and Anhong Guo. 2022. OmniScribe: Authoring Immersive Audio Descriptions for 360° Videos. In Proceedings of the ACM Symposium on User Interface Software and Technology (UIST)
2022
-
[12]
Agnieszka Chmiel and Iwona Mazur. 2022. A homogenous or heterogeneous audience? Audio description preferences of persons with congenital blindness, non-congenital blindness and low vision. Perspectives 30, 3 (2022), 552–567
2022
-
[13]
Cheng-Yu Chuang and Pooyan Fazli. 2023. ClearViD: Curriculum Learning for Video Description. arXiv:2311.04480 [cs.CV] (2023)
2023 arXiv
-
[14]
Victoria Clarke and Virginia Braun. 2021. Thematic Analysis: A Practical Guide . Sage Publications Ltd, Thousand Oaks, California, USA
2021
-
[15]
Google DeepMind. 2024. Gemini. https://deepmind.google/technologies/gemini/
2024
-
[16]
Described and Captioned Media Program (DCMP). 2024. Description Key for Educational Media. https://dcmp.org/learn/descriptionkey
2024
-
[17]
Benoît Encelle, Magali Ollagnier Beldame, and Yannick Prié. 2013. Towards the usage of pauses in audio-described videos. In Proceedings of the International Cross-Disciplinary Conference on Web Accessibility (W4A)
2013
-
[18]
Anita Fidyka and Anna Matamala. 2018. Audio description in 360º videos: Results from focus groups in Barcelona and Kraków. Translation Spaces 7, 2 (2018), 285– 303
2018
-
[19]
Nazaret Fresno, Judit Castellà, and Olga Soler-Vilageliu. 2016. ‘What should I say?’ Tentative criteria to prioritize information in the audio description of film characters. Researching Audio Description: New Approaches (2016), 143–167
2016
-
[20]
Mina Huh, Saelyne Yang, Yi-Hao Peng, Xiang ’Anthony’ Chen, Young-Ho Kim, and Amy Pavel. 2023. AVscript: Accessible Video Editing with Audio-Visual Scripts. In Proceedings of the ACM SIGCHI Conference on Human Factors in Com- puting Systems (CHI)
2023
-
[21]
Shasta Ihorn, Yue-Ting Siu, Aditya Bodi, Lothar Narins, Jose M Castanon, Yash Kant, Abhishek Das, Ilmi Yoon, and Pooyan Fazli. 2021. NarrationBot and InfoBot: A Hybrid System for Automated Video Description. arXiv:2111.03994 [cs.HC] (2021)
2021 arXiv
-
[22]
Jorge Iranzo Bartolome, Luis Cavazos Quero, Sunhee Kim, Myung-Yong Um, and Jundong Cho. 2019. Exploring Art with a Voice Controlled Multimodal Guide for Blind People. In Proceedings of the International Conference on Tangible, Embedded, and Embodied Interaction (Tempe, Arizona...
2019
-
[23]
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of Hallucination in Natural Language Generation. Comput. Surveys 55, 12 (March 2023), 1–38
2023
-
[24]
Lucy Jiang, Crescentia Jung, Mahika Phutane, Abigale Stangl, and Shiri Azenkot
-
[25]
Lucy Jiang and Richard Ladner. 2022. Co-Designing Systems to Support Blind and Low Vision Audio Description Writers. In Proceedings of the International ACM SIGACCESS Conference on Computers and Accessibility (ASSETS)
2022
-
[26]
Lucy Jiang, Mahika Phutane, and Shiri Azenkot. 2023. Beyond Audio Description: Exploring 360° Video Accessibility with Blind and Low Vision Users Through Collaborative Creation. In Proceedings of the International ACM SIGACCESS Con- ference on Computers and Accessibility (ASSETS)
2023
-
[27]
Georgina Kleege and Scott Wallin. 2015. Audio description as a pedagogical tool. Disability Studies Quarterly 35, 2 (2015)
2015
-
[28]
Masatomo Kobayashi, Kentarou Fukuda, Hironobu Takagi, and Chieko Asakawa
-
[29]
Chaoyu Li, Sid Padmanabhuni, Maryam Cheema, Hasti Seifi, and Pooyan Fazli
-
[30]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual In- struction Tuning. In Advances in Neural Information Processing Systems (NeurIPS) , A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.). Curran Associates, Inc., 34892–34916
2023
-
[31]
Xingyu Liu, Patrick Carrington, Xiang ’Anthony’ Chen, and Amy Pavel. 2021. What Makes Videos Accessible to Blind and Visually Impaired People?. In Pro- ceedings of the ACM SIGCHI Conference on Human Factors in Computing Systems (CHI). Article 272, 14 pages
2021
-
[32]
Xingyu "Bruce" Liu, Ruolin Wang, Dingzeyu Li, Xiang ’Anthony’ Chen, and Amy Pavel. 2022. CrossA11y: Identifying Video Accessibility Issues via Cross-modal Grounding. arXiv preprint arXiv:2208.11144 (2022)
2022 arXiv
-
[33]
Iwona Mazur and Agnieszka Chmiel. 2021. Audio Description Training: A Snap- shot of the Current Practices. The Interpreter and Translator Trainer 15, 1 (2021), 51–65
2021
-
[34]
Media Access Canada (MediAC). 2012. Described Video Best Practices Guidelines. http://www.mediac.ca/DVBPGDE_V2_28Feb2012.asp
2012
-
[35]
Morash, Yue-Ting Siu, Joshua A
Valerie S. Morash, Yue-Ting Siu, Joshua A. Miele, Lucia Hasty, and Steven Lan- dau. 2015. Guiding Novice Web Workers in Making Image Descriptions Using Templates. ACM Transactions on Accessible Computing (TACCESS) 7, 4, Article 12 (nov 2015), 21 pages
2015
-
[36]
Rosiana Natalie, Ruei-Che Chang, Smitha Sheshadri, Anhong Guo, and Kotaro Hara. 2024. Audio Description Customization. In Proceedings of the International ACM SIGACCESS Conference on Computers and Accessibility (ASSETS) . 1–19
2024
-
[37]
Rosiana Natalie, Jolene Loh, Huei Suen Tan, Joshua Tseng, Ian Luke Yi-Ren Chan, Ebrima H Jarjue, Hernisa Kacorri, and Kotaro Hara. 2021. The Efficacy of Collab- orative Authoring of Video Scene Descriptions. InProceedings of the International ACM SIGACCESS Conference on Comput...
2021
-
[38]
Rosiana Natalie, Joshua Tseng, Hernisa Kacorri, and Kotaro Hara. 2023. Support- ing Novices Author Audio Descriptions via Automatic Feedback. In Proceedings of the ACM SIGCHI Conference on Human Factors in Computing Systems (CHI)
2023
-
[39]
Netflix Studios. 2024. Audio Description Style Guide v2.5. https: //partnerhelp.netflixstudios.com/hc/en-us/articles/215510667-Audio- Description-Style-Guide-v2-5
2024
-
[40]
Josélia Neves. 2012. Multi-sensory Approaches to (Audio) Describing the Visual Arts. MonTI. Monografías de Traducción e Interpretación 4 (2012), 277–293
2012
-
[41]
Alexandre Nevsky, Timothy Neate, Elena Simperl, and Radu-Daniel Vatavu
-
[42]
Nguyen Nguyen, Jing Bi, Ali Vosoughi, Yapeng Tian, Pooyan Fazli, and Chenliang Xu. 2024. OSCaR: Object State Captioning and State Change Representation. In Findings of the Association for Computational Linguistics (NAACL)
2024
-
[43]
Zheng Ning, Brianna L Wimer, Kaiwen Jiang, Keyi Chen, Jerrick Ban, Yapeng Tian, Yuhang Zhao, and Toby Jia-Jun Li. 2024. SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision Viewers. In Proceedings of the ACM SIGCHI Conferenc...
2024
-
[44]
Ofcom. 2024. Provision of TV Access Services: Guidelines. https: //www.ofcom.org.uk/siteassets/resources/documents/consultations/category- 2-6-weeks/17danielpatt321@gmail.com8126-review-of-access-services- code/associated-documents/provision-of-tv-access-services-guidelines.pdf
2024
-
[45]
OpenAI. 2023. GPT-V System Card. Technical Report. OpenAI
2023
-
[46]
Amy Pavel, Gabriel Reyes, and Jeffrey P. Bigham. 2020. Rescribe: Authoring and Automatically Editing Audio Descriptions. In Proceedings of the ACM Symposium on User Interface Software and Technology (UIST) . 747–759
2020
-
[47]
Pitcher-Cooper, M
C. Pitcher-Cooper, M. Seth, B. Kao, J.M. Coughlan, and I. Yoon. 2023. You De- scribed, We Archived: A Rich Audio Description Dataset. Journal of Technology and Persons with Disabilities 11 (May 2023), 192–208
2023
-
[48]
Sonali Rai, Joan Greening, and Leen Petré. 2024. Provision of TV Access Services: Guidelines. https://www.ofcom.org.uk/siteassets/resources/documents/ consultations/category-2-6-weeks/17danielpatt321@gmail.com8126-review- of-access-services-code/associated-documents/provision-...
2024
-
[49]
Keita Saito, Akifumi Wachi, Koki Wataoka, and Youhei Akimoto. 2023. Verbosity Bias in Preference Labeling by Large Language Models. arXiv:2310.10076 [cs.CL]
2023 arXiv
-
[50]
Joel Snyder. 2005. Audio Description: The Visual Made Verbal. International Congress Series 1282 (09 2005), 935–939
2005
-
[51]
Hariharan Subramonyam, Roy Pea, Christopher Lawrence Pondoc, Maneesh Agrawala, and Colleen Seifert. 2024. Bridging the Gulf of Envisioning: Cognitive Design Challenges in LLM Interfaces. arXiv:2309.14459 [cs.HC]
2024 arXiv
-
[52]
Synthesia. 2023. Video Statistics in 2023. https://www.synthesia.io/post/video- statistics#:~:text=Video%20made%20up%2082%25%20of,10%25%20through% 20text
2023
-
[53]
Tess Van Daele, Akhil Iyer, Yuning Zhang, Jalyn C Derry, Mina Huh, and Amy Pavel. 2024. Making Short-Form Videos Accessible with Hierarchical Video Summaries. In Proceedings of the ACM SIGCHI Conference on Human Factors in Computing Systems (CHI)
2024
-
[54]
Lakshmie Narayan Viswanathan, Troy McDaniel, Sreekar Krishna, and Sethu- raman Panchanathan. 2010. Haptics in Audio Described Movies. In IEEE Inter- national Symposium on Haptic Audio Visual Environments and Games (HA VE) . 1–2
2010
-
[55]
Yujia Wang, Wei Liang, Haikun Huang, Yongqi Zhang, Dingzeyu Li, and Lap-Fai Yu. 2021. Toward Automatic Audio Description Generation for Accessible Videos. In Proceedings of the ACM SIGCHI Conference on Human Factors in Computing Systems (CHI)
2021
-
[56]
Lindsay Kolowich Wiegand. 2023. How Video Consumption is Changing in 2023 [New Research]. https://blog.hubspot.com/marketing/how-video-consumption- is-changing
2023
-
[57]
World Wide Web Consortium (W3C). 2008. Understanding Success Criterion 1.2.3: Audio Description or Media Alternative (Prerecorded). https://www.w3. org/TR/UNDERSTANDING-WCAG20/media-equiv-audio-desc-only.html
2008
-
[58]
YouDescribe. 2024. YouDescribe: Video Description for YouTube. https: //youdescribe.org/
2024
-
[59]
Yuksel, Pooyan Fazli, Umang Mathur, Vaishali Bisht, Soo Jung Kim, Joshua Junhee Lee, Seung Jung Jin, Yue-Ting Siu, Joshua A
Beste F. Yuksel, Pooyan Fazli, Umang Mathur, Vaishali Bisht, Soo Jung Kim, Joshua Junhee Lee, Seung Jung Jin, Yue-Ting Siu, Joshua A. Miele, and Ilmi Yoon
-
[60]
Lotus Zhang, Simon Sun, and Leah Findlater. 2023. Understanding Digital Content Creation Needs of Blind and Low Vision People. In Proceedings of International ACM SIGACCESS Conference on Computers and Accessibility (ASSETS) . DIS ’25, July 5–9, 2025, Funchal, Portugal Maryam C...
2023
-
[67]
Avoid over-describing — Do not include non-essential visual details. [39]
-
[68]
Description should not be opinionated unless content demands it. [39]
-
[69]
Choose level of detail based on plot relevance when describing scenes. [39]
-
[70]
Description should be informative and conversational, in present tense and third-person omniscient. [39]
-
[71]
Vocabulary used should ensure accuracy, clarity, and conciseness
The vocabulary should reflect the predominant language/accent of the program and should be consistent with the genre and tone of the content while also mindful of the target audience. Vocabulary used should ensure accuracy, clarity, and conciseness. [39]
-
[72]
Consider historical context and avoid words with negative connotations or bias. [39]
-
[73]
Pay attention to verbs — Choose vivid verbs over bland ones with adverbs. [39]
-
[74]
Use pronouns only when clear whom they refer to. [39]
-
[75]
Use comparisons for shapes and sizes with familiar and globally relevant objects. [39]
-
[76]
Maintain consistency in word choice, character qualities, and visual elements for all audio descriptions. [39]
-
[77]
Tone and vocabulary should match the target audience’s age range. [39]
-
[78]
Ensure no errors in word selection, pronunciation, diction, or enunciation. [16]
-
[79]
Start with general context, then add details. [16]
-
[80]
Describe shape, size, texture, or color as appropriate to the content. [16]
-
[81]
Use first-person narrative for engagement if required to engage the audience. [16]
-
[82]
Use articles appropriately to introduce or refer to subjects. [16]
-
[83]
Prefer formal speech over colloquialisms, except where appropriate. [16]
-
[84]
When introducing new terms, objects, or actions, label them first, and then follow with the definitions. [16]
-
[85]
Also, do not censor content
Describe objectively without personal interpretation or comment. Also, do not censor content. [16, 44]
-
[86]
Deliver narration steadily and impersonally (but not monotonously), matching the program’s tone. [44]
-
[87]
Adjust style for emotion and mood according to the program’s genre
It can be important to add emotion, excitement, lightness of touch at different points. Adjust style for emotion and mood according to the program’s genre. [44]
-
[88]
If it is children’s content, tailor language and pace for children’s TV, considering audience feedback. [44]
-
[89]
You should describe what you see
Do not alter, filter, or exclude content. You should describe what you see. Try to seek simplicity and succinctness in your description. [34]
-
[90]
Prioritize what is relevant when describing action as to not affect user experience. [39]
-
[91]
Include location, time, and weather conditions when relevant to the scene or plot. [39]
-
[92]
This is so that the intention of the program is conveyed
Focus on key content for learning and enjoyment when creating audio descriptions. This is so that the intention of the program is conveyed. [16]
-
[93]
When describing an instructional video/content, describe the sequence of activities first. [16]
-
[94]
For a dramatic production, include elements such as style, setting, focus, period, dress, facial features, objects, and aesthetics. [16]
-
[95]
Describe what is most essential for the viewer to know in order to follow, understand, and appreciate the intended learning outcomes of the video/content. [16]
-
[96]
Audio description should describe characters, locations, time and circumstances, on-screen action, and on-screen information. [44]
-
[97]
Describe only what a sighted viewer can see. [34]
-
[98]
Prioritize factual descriptions of traits like hair, skin, eyes, build, height, age, and visible disabilities
Describe main and key supporting characters’ visual aspects relevant to identity and personality. Prioritize factual descriptions of traits like hair, skin, eyes, build, height, age, and visible disabilities. Ensure consistency and avoid singling out characters for specific tr...
-
[99]
If unable to confirm or if not established in the plot, do not guess or assume racial, ethnic or gender identity. [39]
-
[100]
When naming characters for the first time, aim to include a descriptor before the name (e.g., a bearded man, Jack). [39]
-
[101]
Description should convey facial expressions, body language and reactions. [39]
-
[102]
When important to the meaning / intent of content, describe race using currently-accepted terminology. [16]
-
[103]
Avoid identifying characters solely by gender expression unless it offers unique insights not apparent otherwise to visually impaired viewers. [34]
-
[104]
Describe character clothing if it enhances characterization, plot, setting, or genre enjoyment. [34]
-
[105]
This may include making an announcement, such as ’Words appear’
If text on the screen is central to understanding, establish a pattern of on-screen words being read. This may include making an announcement, such as ’Words appear’. [34]
-
[106]
In the case of subtitles, the describer should read the translation after stating that a subtitle appears. [34]
-
[107]
When shot changes are critical to the understanding of the scene, indicate them by describing where the action is or where characters are present in the new shot. [39]
-
[108]
knee lifts
Provide description before the content rather than after. [39] Describe Now: User-Driven Audio Description for Blind and Low Vision Individuals DIS ’25, July 5–9, 2025, Funchal, Portugal B GPT-4V description prompt The following prompt was used to create concise and detailed d...
2025
-
[2009]
In Proceedings of the International ACM SIGACCESS Conference on Computers and Accessibility (ASSETS)
Providing Synthesized Audio Description for Online Videos. In Proceedings of the International ACM SIGACCESS Conference on Computers and Accessibility (ASSETS). 249–250
-
[2019]
Video Description: A Survey of Methods, Datasets, and Evaluation Metrics. Comput. Surveys 52, 6, Article 115 (oct 2019), 37 pages
2019
-
[2020]
In Proceedings of the ACM Conference on Designing Interactive Systems (DIS)
Human-in-the-loop Machine Learning to Increase Video Accessibility for Visually Impaired and Blind Users. In Proceedings of the ACM Conference on Designing Interactive Systems (DIS) . 47–60
-
[2023]
InProceedings of the ACM International Conference on Interactive Media Experiences (IMX)
Accessibility Research in Digital Audiovisual Media: What Has Been Achieved and What Should Be Done Next?. InProceedings of the ACM International Conference on Interactive Media Experiences (IMX) . 94–114
-
[2024]
It’s Kind of Context Dependent
“It’s Kind of Context Dependent”: Understanding Blind and Low Vision People’s Video Accessibility Preferences Across Viewing Scenarios. InProceedings of the ACM SIGCHI Conference on Human Factors in Computing Systems (CHI) . 1–20
-
[2025]
arXiv preprint arXiv:2502.20480 (2025)
VideoA11y: Method and Dataset for Accessible Video Description. arXiv preprint arXiv:2502.20480 (2025)
2025 arXiv
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.