Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

DanmuA11y: Making Time-Synced On-Screen Video Comments (Danmu) Accessible to Blind and Low Vision Users via Multi-Viewer Audio Discussions

T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read DanmuA11y claims that converting time-synced video comments into multi-viewer audio discussions with visual descriptions lets blind and low vision viewers understand Danmu, and reports significant gains in comprehension, smoothness, and…

desk verdict A genuinely useful accessibility system paper, but the evaluation glosses over nine abandoned baseline trials; worth peer review after a missing-data fix. read the letter →

arxiv 2501.15711 v1 pith:VMZM6DZM submitted 2025-01-27 cs.HC

classification cs.HC
keywords Danmuaccessibilityblindandlowvisionaudiodiscussionsspatialvideocommentslargelanguagemodelsco-watchingscreenreaderalternatives
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that Danmu—the scrolling, time-synced comments overlaid on online videos—excludes blind and low vision (BLV) viewers because it is visual, overlaps with video speech, and becomes disorganized when read aloud sequentially. To fix this, the authors built DanmuA11y, which turns Danmu into multi-viewer audio discussions: an AI narrator supplies visual context, curated comments are reorganized into dialogues, and spatial audio positions the speakers around the listener. In a within-subject study with twelve BLV viewers, the system significantly reduced Danmu confusion, made the viewing experience feel smoother, and strengthened the sense of co-watching with other people. The contribution is a demonstrated way to make time-synced social video accessible, with design implications for live-streaming and other commentary-heavy media.

What carries the argument

The central mechanism is a three-part pipeline anchored by a topic-quality and insertion-quality optimization. The pipeline first uses GPT-4o to group time-stamped Danmu comments into topics, filter redundant comments, and reorder them into dialogue-like sequences, and to generate a short visual description when a topic refers to something visible. It then schedules each topic at a non-speech segment or a speech break, solving an integer linear program that maximizes the weighted sum of topic quality (informativeness via sentence-embedding dissimilarity, creativity via a GPT-4o rating, and opinion diversity via sentiment labels) and insertion quality (language-model coherence minus a pause penalty). Finally, it renders each topic as a multi-viewer audio discussion using spatial audio and four alternating synthesized voices, with a shake gesture to access discussions on demand at speech breaks. This machinery lets the video's original speech remain uninterrupted while the audience commentary feels like a live conversation around the listener.

What would settle it

Run the pipeline on a fresh set of videos and measure whether the reported comprehension gains reproduce: if visual-description accuracy falls materially below 92.5% on videos with fast movement or unusual scenes, or if BLV participants' confusion rates with DanmuA11y rise toward the baseline rate, the central claim fails.

Watch

Extended reading notes

Core claim

DanmuA11y converts Danmu into audio discussions by three steps: grouping comments into topics and adding brief visual descriptions of what the comments refer to, scheduling each topic into a non-speech gap or speech break in the video via an integer-linear-programming optimization that balances topic quality (informativeness, creativity, opinion diversity) with insertion quality (coherence and pause length), and finally rendering the curated topics as dialogues spoken by an AI narrator and several human-voice virtual viewers placed around the listener with spatial audio. Compared with a baseline that simulated current screen-reader practice, the twelve BLV participants reported significantly fewer moments of Danmu confusion (nine versus fifty-two instances across thirty-six video views), wrote more detailed and more accurate video summaries, and rated the experience as more coherent, unobtrusive, and socially engaging. The authors present this as the first system to make Danmu itself accessible, rather than merely making a list of comments screen-reader friendly.

Load-bearing premise

The whole benefit depends on GPT-4o reliably turning comments into coherent topics, describing visuals accurately, and rating creativity; the paper's only direct check is a 92.5% accuracy spot-check of visual descriptions on 200 topics, and the authors note the model occasionally hallucinates or errs.

Editorial extensions

If this is right

  • BLV viewers can understand Danmu at levels comparable to the paper's measures: reported confusion fell from 52 to 9 instances, and video summaries contained fewer factual errors.
  • The same audio-discussion pattern can carry other time-synced commentary, such as live-stream chat, if the pipeline can run in real time.
  • Platform designers can preserve the co-watching feeling of Danmu for BLV audiences by pairing visual descriptions with curated, dialogue-organized, spatially separated voices.
  • Users can choose on-demand when to hear commentary, which reduces split attention while keeping the option to dive deeper.
  • Personalization is a direct next step: adjusting the weights of informativeness, creativity, and diversity, or adding user-defined filters, would tailor the experience.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same design could benefit sighted viewers in situations where the screen is unavailable or attention is split, since the audio-discussion format removes the need to read while watching.
  • Because the system's quality hinges on a large language model's vision-language ability, the approach will scale to other visual social media (memes, GIFs, short-form video) as that ability improves, rather than requiring per-platform redesign.
  • The optimization over topic placement is a general scheduling problem for accessibility; its reward function could be adapted to other constraints, such as minimizing interruption during key moments or maximizing discussion density for live events.
  • A testable extension is an ablation study removing spatial audio or visual descriptions; the current design bundles them, so their individual contributions to the reported social and comprehension gains are not yet isolated.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper presents DanmuA11y, a system that makes time-synced on-screen video comments (Danmu) accessible to blind and low vision (BLV) users by converting them into multi-viewer audio discussions. A formative study with eight BLV participants identified three challenges: missing visual context, speech interference between comments and video, and disorganized sequential access. The system addresses these with AI-generated visual descriptions, an optimization algorithm that places curated comment topics into non-speech segments and speech breaks, and spatial-audio rendering of multi-viewer dialogues. A within-subject evaluation with twelve BLV participants compared DanmuA11y against a baseline simulating current practice (an auto-scrolling list). The authors report significantly fewer confusion instances, smoother viewing, and higher social presence with DanmuA11y, alongside positive usability ratings. The paper contributes a novel accessible-Danmu system, a design rationale derived from a formative study, and an evaluation with the target user population.

Significance. If the reported results hold, this is a valuable contribution to accessibility research and to the growing literature on non-visual access to social video platforms. The work combines LLM-based content curation with spatial audio in a way that is well grounded in the target users' needs, and the evaluation directly involves BLV participants, which is a clear strength. The paper also provides detailed implementation information and a spot-checked accuracy of 92.5% for visual descriptions, supporting reproducibility. The main reservation is that the statistical comparison is clouded by incomplete baseline trials and an unstated treatment of missing data; this must be resolved before the central claims can be fully trusted.

major comments (3)
  1. [Section 6 and Table 5] Nine of 36 baseline trials were abandoned due to frustration, yet the paper reports confusion rates as '1.44 times per video' using a denominator of 36 (52 total instances). This treats incomplete trials as full exposure events and is inconsistent with the per-participant-per-video averaging shown in Figure 11. The manuscript does not state how the Wilcoxon signed-rank tests handled these missing values, nor how the video-summary word counts and factual-error counts were computed for participants who did not complete all baseline videos. Because attrition was caused by frustration with the baseline, the missing data are unlikely to be missing at random. The direction of any bias is not obvious: abandoned trials may have truncated the number of confusion reports (making the baseline look better) or may represent the most frustrated participants (making the baseline look worse). Please report a sensitivity analysis, e.g., analysis on complete pairs only, and a worst-case imputation for abandoned trials, and state explicitly how the denominators and paired tests were constructed.
  2. [Section 6.2.1] The paired t-test for summary length is reported with df=11 (i.e., 12 participants), but if nine baseline trials were abandoned, then nine baseline video summaries may be missing. The paper does not explain how missing summaries were handled: listwise deletion, imputation, or treating abandoned trials as zero-length summaries. As written, the reported t-test and the factual-error comparison (Z=-2.27, p<.05) cannot be verified. This is load-bearing because the comprehension claim depends on these pairwise comparisons, and the missingness is differential across the two conditions.
  3. [Section 4.7.4 and Section 7.5] The reliability of the pipeline is only partially validated. Visual descriptions were spot-checked on 200 topics (92.5% accuracy), and the creativity score distribution is reported, but the topic grouping, the dialogue reordering, and the coherence of inserted discussions are not evaluated against human judgment. The user study used only six videos for the controlled comparison, and Section 7.5 acknowledges that GPT-4o 'occasionally results in hallucinations and errors.' Without at least a sample-based agreement check on topic grouping and dialogue coherence, it is difficult to know whether the reported comprehension and social benefits generalise beyond the specific videos tested. The authors should either provide such validation or explicitly scope the claims to the studied video set.
minor comments (5)
  1. [Section 7.1] The heading 'Personalization of DanumA11y' contains a typo: 'DanumA11y' should be 'DanmuA11y'.
  2. [Section 7.3.1] The sentence 'the AI visual descriptions should provide provide new, complementary information' contains a duplicated word 'provide'.
  3. [Table 1] In Table 1, 'Bibibili' appears as a platform name; this should be 'Bilibili'.
  4. [Section 5.2.3 / Table 4] The paper reports 18 Wilcoxon tests on questionnaire items without any correction for multiple comparisons. Several effects are at p<.05 and might not survive a conservative correction; the main p<.01 findings are likely robust, but this should be acknowledged.
  5. [Section 4.5.1] The topic-quality weights (λ_i=λ_c=λ_d=1) and the insertion-quality weight (λ_p=0.25, λ=10) are empirical choices, but the paper does not report any sensitivity analysis for these parameters. A brief statement of how sensitive the results are to these weights would strengthen the system's credibility.

Circularity Check

0 steps flagged · score 1.0 of 10

No circular derivation: the evaluation outcomes are independent user measures, not reconstructed from the system's optimization objective.

full rationale

The paper's derivation chain is a standard HCI pipeline: formative study identifies challenges, the DanmuA11y system addresses them, and a within-subject user study measures outcomes. None of the reported outcomes are defined in terms of the system's internal scoring functions, and no parameter is fitted to the outcome data. The curation weights (lambda_i, lambda_c, lambda_d, and lambda_p) are hand-set before the study and are not derived from participant responses (Sections 4.5.1 and 4.5.2). Comprehension is measured by reported confusion instances and video-summary errors; viewing smoothness and social connection are measured by established questionnaires; these are independent of the system's topic-quality and insertion-quality objectives (Sections 5.2.2 and 6). GPT-4o is used as a processing tool with a reported 92.5% accuracy spot-check (Section 4.7.4), and the paper explicitly acknowledges its error modes (Section 7.5), which is a generalizability limitation rather than circular reasoning. The self-citations in the reference list support peripheral design suggestions and prior accessibility techniques, not the central claim. The nine uncompleted baseline trials (Section 6 and Table 5) raise a legitimate attrition and statistical-validity concern, but that is an empirical missing-data issue, not a case where a prediction reduces to its inputs by construction. Overall, no load-bearing circular step is present.

Assumptions & free parameters 9 free parameters · 6 assumptions · 0 invented entities

The central claim rests on the reliability of several AI components and on design parameters that are hand-set without sensitivity analysis. There is no formal derivation; the evidence is empirical and user-study based.

free parameters (9)
  • Topic quality weights λ_i, λ_c, λ_d = 1, 1, 1
    Equal weights for informativeness, creativity, and diversity in topic quality; chosen by hand with no sensitivity analysis.
  • Pause penalty weight λ_p = 0.25
    Set to allow pauses when necessary; hand-chosen in insertion quality formula.
  • Optimization balance weight λ = 10
    Empirically set to prioritize insertion coherence over topic quality in the ILP objective.
  • Non-speech gap duration threshold = 2 seconds
    Classifies gaps between sentences as non-speech; follows prior research but hand-set.
  • RMS volume threshold = 0.8
    Marks high-volume regions as inappropriate for insertion; hand-set.
  • Minimum speech segment length = 20 words
    Constraint on speech-break segmentation to avoid excessive pauses; hand-set.
  • Maximum pause duration at speech breaks = 10 seconds
    Normalizes pause score and caps inserted topic length; hand-set.
  • Shake acceleration threshold = 6 m/s²
    Trigger for on-demand access; hardware-specific hand-set value.
  • Speech rate for duration estimation = 3 Chinese words per second
    Assumed typical speech rate used to estimate topic duration and pause penalty.
assumptions (6)
  • domain assumption GPT-4o topic modeling and dialogue reorganization produce coherent, usable topics from raw Danmu comments.
    No quantitative evaluation of topic quality in the study; relies on prompt-based method from prior work.
  • domain assumption GPT-4o visual descriptions are sufficiently accurate and relevant.
    Spot check reported 92.5% accuracy on 200 sampled topics; errors acknowledged and discussed in Section 7.2.1.
  • domain assumption SceneDetect key frames represent the visual content of each segment adequately.
    Middle frame per shot used; fast-changing scenes (e.g., fight in V11) caused captioning errors.
  • domain assumption Universal Sentence Encoder cosine similarity captures informativeness of a topic relative to speech.
    Adopted from prior work without task-specific validation for Danmu content.
  • domain assumption davinci-002 log probability measures coherence between inserted topic and surrounding speech.
    Adopted from Rescribe [66]; no validation on Chinese Danmu and Mandarin video speech.
  • domain assumption The auto-scrolling list design probe is a valid baseline representing current BLV practices for accessing Danmu.
    Seven of eight formative participants identified Etong as the best existing tool; the probe replicates its interface.

how reviews work

0 comments
Cite this review

Pith. "Pith review of DanmuA11y: Making Time-Synced On-Screen Video Comments (Danmu) Accessible to Blind and Low Vision Users via Multi-Viewer Audio Discussions." pith.science (2026). https://pith.science/paper/VMZM6DZM

@misc{pith2026250115711,
  author       = {Pith},
  title        = {Pith review of: DanmuA11y: Making Time-Synced On-Screen Video Comments (Danmu) Accessible to Blind and Low Vision Users via Multi-Viewer Audio Discussions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VMZM6DZM}},
  note         = {Machine review of arXiv:2501.15711}
}
read the original abstract

By overlaying time-synced user comments on videos, Danmu creates a co-watching experience for online viewers. However, its visual-centric design poses significant challenges for blind and low vision (BLV) viewers. Our formative study identified three primary challenges that hinder BLV viewers' engagement with Danmu: the lack of visual context, the speech interference between comments and videos, and the disorganization of comments. To address these challenges, we present DanmuA11y, a system that makes Danmu accessible by transforming it into multi-viewer audio discussions. DanmuA11y incorporates three core features: (1) Augmenting Danmu with visual context, (2) Seamlessly integrating Danmu into videos, and (3) Presenting Danmu via multi-viewer discussions. Evaluation with twelve BLV viewers demonstrated that DanmuA11y significantly improved Danmu comprehension, provided smooth viewing experiences, and fostered social connections among viewers. We further highlight implications for enhancing commentary accessibility in video-based social media and live-streaming platforms.

Figures

Figures reproduced from arXiv: 2501.15711 by the authors.

Figure 1
Figure 1. (A) Danmu is a video-commenting feature that overlays time-synced user comments onto videos. It creates a co [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. The key characteristics of Danmu comments. [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. The design probe presents Danmu comments in an [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (9 more)
Figure 4
Figure 4. Figure 4: An example walk-through of DanmuA11y: (A) The system automatically plays multi-viewer discussions during non [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: The pipeline of DanmuA11y. 4.3 Video Segmentation DanmuA11y divides the video into two types of segments: non￾speech and speech segments. This speech-based segmentation al￾lows the system to identify proper insertion points for Danmu. 4.3.1 Non-Speech Segment. A non-sp…
Figure 6
Figure 6. Figure 6: DanmuA11y divides the video into non-speech and speech segments. The points between adjacent speech segments [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: (A) Danmu topics belong to each video segment. (B) Each topic includes a visual description and a list of comments. [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]
Figure 8
Figure 8. Figure 8: (A) The system places Danmu topics into insertion points. (B) It maximizes both topic quality and insertion quality. [PITH_FULL_IMAGE:figures/full_fig_p009_8.png]
Figure 9
Figure 9. Figure 9: DanmuA11y presents Danmu via multi-viewer discussions. (A) The system positions the AI narrator on the left side [PITH_FULL_IMAGE:figures/full_fig_p010_9.png]
Figure 10
Figure 10. Figure 10: Distributions of the ratings for DanmuA11y and the baseline (1=strongly negative, 7=strongly positive). The asterisks [PITH_FULL_IMAGE:figures/full_fig_p013_10.png]
Figure 11
Figure 11. Figure 11: Comprehension metrics in the evaluation study. These metrics were calculated as averages for each participant per [PITH_FULL_IMAGE:figures/full_fig_p013_11.png]
Figure 12
Figure 12. Figure 12: Screenshots from the 18-video dataset. Video creators’ IDs and human faces are obscured for privacy. [PITH_FULL_IMAGE:figures/full_fig_p021_12.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Vid2Coach: Transforming How-To Videos into Task Assistants

    cs.HC 2025-05 conditional novelty 6.0 of 10

    Vid2Coach converts how-to videos into a real-time, wearable task assistant that helps blind and low vision people cook with fewer errors.

Reference graph

Works this paper leans on

103 extracted references · 47 canonical work pages · cited by 1 Pith paper

  1. [1]

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Floren- cia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)

  2. [2]

    Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015. Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision . 2425–2433

  3. [3]

    Arthur Aron, Elaine N Aron, and Danny Smollan. 1992. Inclusion of other in the self scale and the structure of interpersonal closeness. Journal of personality and social psychology 63, 4 (1992), 596

  4. [4]

    Francesco Barbieri, Luis Espinosa Anke, and Jose Camacho-Collados. 2021. XLM- T: Multilingual language models in Twitter for sentiment analysis and beyond. arXiv preprint arXiv:2104.12250 (2021)

  5. [5]

    Virginia P Campos, Tiago MU de Araújo, Guido L de Souza Filho, and Luiz MG Gonçalves. 2020. CineAD: a system for automated audio description script generation for the visually impaired. Universal Access in the Information Society 19, 1 (2020), 99–111

  6. [6]

    Virgínia P Campos, Luiz MG Gonçalves, Wesnydy L Ribeiro, Tiago MU Araújo, Thaís G Do Rego, Pedro HV Figueiredo, Suanny FS Vieira, Thiago FS Costa, Caio C Moraes, Alexandre CS Cruz, et al. 2023. Machine generation of audio de- scription for blind and visually impaired people. ACM Transactions on Accessible Computing 16, 2 (2023), 1–28

  7. [7]

    Shuxian Cao, Dongliang Guo, Lina Cao, Shuo Li, Junlan Nie, Amit Kumar Singh, and Haibin Lv. 2023. VisDmk: visual analysis of massive emotional danmaku in online videos. The Visual Computer 39, 12 (2023), 6553–6570

  8. [8]

    Marina Ramos Caro. 2016. Testing audio narration: the emotional impact of language in audio description. Perspectives 24, 4 (2016), 606–634

Show all 103 references
  1. [9]

    D Cer. 2018. Universal sentence encoder. arXiv preprint arXiv:1803.11175 (2018)

  2. [10]

    Ruei-Che Chang, Yuxuan Liu, and Anhong Guo. 2024. WorldScribe: Towards Context-Aware Live Visual Descriptions. InProceedings of the 37th Annual ACM Symposium on User Interface Software and Technology . 1–18

  3. [11]

    Ruei-Che Chang, Chao-Hsien Ting, Chia-Sheng Hung, Wan-Chen Lee, Liang- Jin Chen, Yu-Tzu Chao, Bing-Yu Chen, and Anhong Guo. 2022. Omniscribe: Authoring immersive audio descriptions for 360 videos. In Proceedings of the 35th Annual ACM Symposium on User Interface Software and T...

  4. [12]

    Si Chen, Haocong Cheng, Jason Situ, Desirée Kirst, Suzy Su, Saumya Malhotra, Lawrence Angrave, Qi Wang, and Yun Huang. 2024. Towards Inclusive Video Commenting: Introducing Signmaku for the Deaf and Hard-of-Hearing. In Proceedings of the CHI Conference on Human Factors in Comp...

  5. [13]

    Shuai Chen, Sihang Li, Yanda Li, Junlin Zhu, Juanjuan Long, Siming Chen, Jiawan Zhang, and Xiaoru Yuan. 2022. DanmuVis: Visualizing danmu content dynamics and associated viewer behaviors in online videos. In Computer Graphics Forum, Vol. 41. Wiley Online Library, 429–440

  6. [14]

    I was afraid, but now I enjoy being a streamer!

    Xinyue Chen, Si Chen, Xu Wang, and Yun Huang. 2021. " I was afraid, but now I enjoy being a streamer!" Understanding the Challenges and Prospects of Using Live Streaming for Online Education. Proceedings of the ACM on Human-Computer Interaction 4, CSCW3 (2021), 1–32

  7. [15]

    Yue Chen, Qin Gao, and Pei-Luen Patrick Rau. 2017. Watching a movie alone yet together: understanding reasons for watching Danmaku videos. International Journal of Human–Computer Interaction 33, 9 (2017), 731–743

  8. [16]

    Tianheng Cheng, Lin Song, Yixiao Ge, Wenyu Liu, Xinggang Wang, and Ying Shan. 2024. Yolo-world: Real-time open-vocabulary object detection. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 16901–16911

  9. [17]

    Zesen Cheng, Sicong Leng, Hang Zhang, Yifei Xin, Xin Li, Guanzheng Chen, Yongxin Zhu, Wenqi Zhang, Ziyang Luo, Deli Zhao, et al. 2024. VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs. arXiv preprint arXiv:2406.07476 (2024)

  10. [18]

    Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre Défossez. 2024. Simple and controllable music generation. Advances in Neural Information Processing Systems 36 (2024)

  11. [19]

    Juliet Corbin and Anselm Strauss. 2015. Basics of qualitative research. Vol. 14. sage

  12. [20]

    Khang Dang and Sooyeon Lee. 2024. Musical Performances in Virtual Reality with Spatial and View-Dependent Audio Descriptions for Blind and Low-Vision Users. In The 26th International ACM SIGACCESS Conference on Computers and Accessibility. 1–5

  13. [21]

    Ginger Delmas, Philippe Weinzaepfel, Thomas Lucas, Francesc Moreno-Noguer, and Grégory Rogez. 2022. Posescript: 3d human poses from natural language. In European Conference on Computer Vision . Springer, 346–362

  14. [22]

    Cole Gleason, Patrick Carrington, Lydia B Chilton, Benjamin Gorman, Hernisa Kacorri, Andrés Monroy-Hernández, Meredith Ringel Morris, Garreth Tigwell, and Shaomei Wu. 2020. Future research directions for accessible social media. ACM SIGACCESS Accessibility and Computing 127 (2...

  15. [23]

    Cole Gleason, Patrick Carrington, Lydia B Chilton, Benjamin M Gorman, Hernisa Kacorri, Andrés Monroy-Hernández, Meredith Ringel Morris, Garreth W Tig- well, and Shaomei Wu. 2019. Addressing the accessibility of social media. In Companion Publication of the 2019 Conference on C...

  16. [24]

    Cole Gleason, Amy Pavel, Himalini Gururaj, Kris Kitani, and Jeffrey Bigham

  17. [25]

    Cole Gleason, Amy Pavel, Xingyu Liu, Patrick Carrington, Lydia B Chilton, and Jeffrey P Bigham. 2019. Making memes accessible. In Proceedings of the 21st International ACM SIGACCESS Conference on Computers and Accessibility . 367–376

  18. [26]

    Cole Gleason, Amy Pavel, Emma McCamey, Christina Low, Patrick Carrington, Kris M Kitani, and Jeffrey P Bigham. 2020. Twitter A11y: A browser extension to make Twitter images accessible. In Proceedings of the 2020 chi conference on human factors in computing systems . 1–12

  19. [27]

    Yizheng Gu, Chun Yu, Zhipeng Li, Weiqi Li, Shuchang Xu, Xiaoying Wei, and Yuanchun Shi. 2019. Accurate and low-latency sensing of touch contact on any surface with finger-worn IMU sensor. In Proceedings of the 32nd annual ACM symposium on user interface software and technology...

  20. [28]

    João Guerreiro, Yujin Kim, Rodrigo Nogueira, SeungA Chung, André Rodrigues, and Uran Oh. 2023. The design space of the auditory representation of objects and their behaviours in virtual reality for blind people. IEEE Transactions on Visualization and Computer Graphics 29, 5 (2...

  21. [29]

    Sandra G Hart. 2006. NASA-task load index (NASA-TLX); 20 years later. In Proceedings of the human factors and ergonomics society annual meeting , Vol. 50. Sage publications Sage CA: Los Angeles, CA, 904–908

  22. [30]

    Changyang He, Lu He, Tun Lu, and Bo Li. 2021. Beyond Entertainment: Un- packing Danmaku and Comments’ Role of Information Sharing and Sentiment Expression in Online Crisis Videos. Proceedings of the ACM on Human-Computer Interaction 5, CSCW2 (2021), 1–27

  23. [31]

    Ming He, Yong Ge, Enhong Chen, Qi Liu, and Xuesong Wang. 2017. Exploring the emerging type of comment for online videos: Danmu. ACM Transactions on the Web (TWEB) 12, 1 (2017), 1–33

  24. [32]

    Steven CH Hoi, Doyen Sahoo, Jing Lu, and Peilin Zhao. 2021. Online learning: A comprehensive survey. Neurocomputing 459 (2021), 249–289

  25. [33]

    Yushi Hu, Hang Hua, Zhengyuan Yang, Weijia Shi, Noah A Smith, and Jiebo Luo. 2023. Promptcap: Prompt-guided image captioning for vqa with gpt-3. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 2963–2975

  26. [34]

    Zeyu Huang, Xinyi Cao, Yuanhao Zhang, and Xiaojuan Ma. 2024. Sharing Frissons among Online Video Viewers: Exploring the Design of Affective Com- munication for Aesthetic Chills. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–19

  27. [35]

    Mina Huh, YunJung Lee, Dasom Choi, Haesoo Kim, Uran Oh, and Juho Kim. 2022. Cocomix: Utilizing Comments to Improve Non-Visual Webtoon Accessibility. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems . 1–18

  28. [36]

    Mina Huh, Yi-Hao Peng, and Amy Pavel. 2023. GenAssist: Making image generation accessible. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology . 1–17

  29. [37]

    Mina Huh, Saelyne Yang, Yi-Hao Peng, Xiang’Anthony’ Chen, Young-Ho Kim, and Amy Pavel. 2023. AVscript: Accessible Video Editing with Audio-Visual Scripts. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–17

  30. [38]

    Felix Immohr, Gareth Rendle, Christian Kehling, Anton Lammert, Steve Göring, Bernd Froehlich, and Alexander Raake. 2024. Subjective Evaluation of the Impact of Spatial Audio on Triadic Communication in Virtual Reality. In 2024 16th International Conference on Quality of Multim...

  31. [39]

    Felix Immohr, Gareth Rendle, Annika Neidhardt, Steve Göring, Rakesh Rao Ra- machandra Rao, Stephanie Arevalo Arboleda, Bernd Froehlich, and Alexander Raake. 2023. Proof-of-concept study to evaluate the impact of spatial audio on social presence and user behavior in multi-modal...

  32. [40]

    Gaurav Jain, Basel Hindi, Connor Courtien, Xin Yi Therese Xu, Conrad Wyrick, Michael Malcolm, and Brian A Smith. 2023. Front Row: Automatically Generat- ing Immersive Audio Representations of Tennis Broadcasts for Blind Viewers. In Proceedings of the 36th Annual ACM Symposium ...

  33. [41]

    Mohit Jain, Nirmalendu Diwakar, and Manohar Swaminathan. 2021. Smartphone usage by expert blind users. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems . 1–15

  34. [42]

    Shagun Jhaver, Quan Ze Chen, Detlef Knauss, and Amy X Zhang. 2022. Design- ing word filter tools for creator-led comment moderation. In Proceedings of the 2022 CHI conference on human factors in computing systems . 1–21

  35. [43]

    Lucy Jiang, Crescentia Jung, Mahika Phutane, Abigale Stangl, and Shiri Azenkot

  36. [44]

    Peng Jin, Ryuichi Takanobu, Wancai Zhang, Xiaochun Cao, and Li Yuan. 2024. Chat-univi: Unified visual representation empowers large language models with image and video understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 13700–13710

  37. [45]

    Joonyoung Jun, Woosuk Seo, Jihyeon Park, Subin Park, and Hyunggu Jung. 2021. Exploring the experiences of streamers with visual impairments. Proceedings of the ACM on Human-Computer Interaction 5, CSCW2 (2021), 1–23

  38. [46]

    Daniel Killough and Amy Pavel. 2023. Exploring Community-Driven Descrip- tions for Making Livestreams Accessible. In Proceedings of the 25th International ACM SIGACCESS Conference on Computers and Accessibility . 1–13

  39. [47]

    Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019. Natural questions: a benchmark for question answering research. Transactions of the Association for C...

  40. [48]

    Jaewook Lee, Yi-Hao Peng, Jaylin Herskovitz, and Anhong Guo. 2021. Image Explorer: Multi-layered touch exploration to make images accessible. In Pro- ceedings of the 23rd International ACM SIGACCESS Conference on Computers and Accessibility. 1–4

  41. [49]

    Yi-Chieh Lee, Wen-Chieh Lin, Fu-Yin Cherng, Hao-Chuan Wang, Ching-Ying Sung, and Jung-Tai King. 2015. Using Time-Anchored Peer Comments to En- hance Social Interaction in Online Educational Videos. In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing ...

  42. [50]

    James R Lewis. 2018. The system usability scale: past, present, and future. International Journal of Human–Computer Interaction 34, 7 (2018), 577–590

  43. [51]

    It Feels Like Taking a Gamble

    Franklin Mingzhe Li, Franchesca Spektor, Meng Xia, Mina Huh, Peter Cederberg, Yuqi Gong, Kristen Shinohara, and Patrick Carrington. 2022. “It Feels Like Taking a Gamble”: Exploring Perceptions, Practices, and Challenges of Using Makeup and Cosmetics for People with Visual Impa...

  44. [52]

    Juncheng Li, Wei Dai, Florian Metze, Shuhui Qu, and Samarjit Das. 2017. A comparison of deep learning methods for environmental sound detection. In 2017 IEEE International conference on acoustics, speech and signal processing (ICASSP). IEEE, 126–130

  45. [53]

    Chengzhong Liu, Shixu Zhou, Dingdong Liu, Junze Li, Zeyu Huang, and Xi- aojuan Ma. 2023. CoArgue: Fostering Lurkers’ Contribution to Collective Ar- guments in Community-based QA Platforms. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–17

  46. [54]

    Guanhong Liu, Tianyu Yu, Chun Yu, Haiqing Xu, Shuchang Xu, Ciyuan Yang, Feng Wang, Haipeng Mi, and Yuanchun Shi. 2021. Tactile Compass: Enabling Visually Impaired People to Follow a Path with Continuous Directional Feedback. In Proceedings of the 2021 CHI Conference on Human F...

  47. [55]

    Nian Liu, Ni Zhang, Kaiyuan Wan, Ling Shao, and Junwei Han. 2021. Visual saliency transformer. In Proceedings of the IEEE/CVF international conference on computer vision. 4722–4732

  48. [56]

    Xingyu Liu, Patrick Carrington, Xiang’Anthony’ Chen, and Amy Pavel. 2021. What makes videos accessible to blind and visually impaired people?. In Pro- ceedings of the 2021 CHI Conference on Human Factors in Computing Systems . 1–14

  49. [57]

    Xingyu" Bruce" Liu, Ruolin Wang, Dingzeyu Li, Xiang Anthony Chen, and Amy Pavel. 2022. CrossA11y: Identifying Video Accessibility Issues via Cross-modal Grounding. In Proceedings of the 35th Annual ACM Symposium on User Interface Software and Technology. 1–14

  50. [58]

    Zhicong Lu, Haijun Xia, Seongkook Heo, and Daniel Wigdor. 2018. You watch, you give, and you engage: a study of live streaming practices in China. In Proceedings of the 2018 CHI conference on human factors in computing systems . 1–13

  51. [59]

    Guangyi Lv, Kun Zhang, Le Wu, Enhong Chen, Tong Xu, Qi Liu, and Weidong He. 2019. Understanding the users and videos by mining a novel danmu dataset. IEEE Transactions on Big Data 8, 2 (2019), 535–551

  52. [60]

    Xiaojuan Ma and Nan Cao. 2017. Video-based evanescent, anonymous, asyn- chronous social interaction: Motivation and adaption to medium. In Proceedings of the 2017 ACM conference on computer supported cooperative work and social computing. 770–782

  53. [61]

    Jennifer Mankoff, Anind K Dey, Gary Hsieh, Julie Kientz, Scott Lederer, and Morgan Ames. 2003. Heuristic evaluation of ambient displays. In Proceedings of the SIGCHI conference on Human factors in computing systems . 169–176

  54. [62]

    Keenan R May, Brianna J Tomlinson, Xiaomeng Ma, Phillip Roberts, and Bruce N Walker. 2020. Spotlights and soundscapes: On the design of mixed reality auditory environments for persons with visual impairment. ACM Transactions on Accessible Computing (TACCESS) 13, 2 (2020), 1–47

  55. [63]

    Rosiana Natalie, Ruei-Che Chang, Smitha Sheshadri, Anhong Guo, and Kotaro Hara. 2024. Audio description customization. arXiv preprint arXiv:2408.11406 (2024)

  56. [64]

    Zheng Ning, Brianna L Wimer, Kaiwen Jiang, Keyi Chen, Jerrick Ban, Yapeng Tian, Yuhang Zhao, and Toby Jia-Jun Li. 2024. SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision Viewers. In Proceedings of the CHI Conference on Hu...

  57. [65]

    Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2023. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th annual acm symposium on user interface software and technology . 1–22

  58. [66]

    Amy Pavel, Gabriel Reyes, and Jeffrey P Bigham. 2020. Rescribe: Authoring and automatically editing audio descriptions. In Proceedings of the 33rd Annual ACM Symposium on User Interface Software and Technology . 747–759

  59. [67]

    Chau Minh Pham, Alexander Hoyle, Simeng Sun, and Mohit Iyyer. 2023. TopicGPT: A prompt-based topic modeling framework. arXiv preprint arXiv:2311.01449 (2023)

  60. [68]

    Javier Ramırez, José C Segura, Carmen Benıtez, Angel De La Torre, and Antonio Rubio. 2004. Efficient voice activity detection algorithms using long-term speech information. Speech communication 42, 3-4 (2004), 271–287

  61. [69]

    N Reimers. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT- Networks. arXiv preprint arXiv:1908.10084 (2019)

  62. [70]

    It Feels Like Being Locked in A Cage

    Ethan Z Rong, Mo Morgana Zhou, Zhicong Lu, and Mingming Fan. 2022. “It Feels Like Being Locked in A Cage”: Understanding Blind or Low Vision Streamers’ Perceptions of Content Curation Algorithms. In Proceedings of the 2022 ACM Designing Interactive Systems Conference . 571–585

  63. [71]

    Sanjiban Sekhar Roy, Akash Roy, Pijush Samui, Mostafa Gandomi, and Amir H Gandomi. 2023. Hateful sentiment detection in real-time tweets: An LSTM- based comparative approach. IEEE Transactions on Computational Social Systems (2023)

  64. [72]

    Woosuk Seo and Hyunggu Jung. 2021. Understanding the community of blind or visually impaired vloggers on YouTube. Universal Access in the Information Society 20 (2021), 31–44

  65. [73]

    Woosuk Seo and Hyunggu Jung. 2022. Challenges and opportunities to improve the accessibility of YouTube for people with visual impairments as content creators. Universal Access in the Information Society 21, 3 (2022), 767–770

  66. [74]

    Luning Sun, Hongyi Gu, Rebecca Myers, and Zheng Yuan. 2023. A New Dataset and Method for Creativity Assessment Using the Alternate Uses Task. In Bench- Council International Symposium on Intelligent Computers, Algorithms, and Ap- plications. Springer, 125–138

  67. [75]

    Zhida Sun, Mingfei Sun, Nan Cao, and Xiaojuan Ma. 2016. VideoForest: interac- tive visual summarization of video streams based on danmu data. In SIGGRAPH ASIA 2016 symposium on visualization . 1–8

  68. [76]

    S Gökhun Tanyer and Hamza Ozer. 2000. Voice activity detection in nonsta- tionary noise. IEEE Transactions on speech and audio processing 8, 4 (2000), 478–482

  69. [77]

    Garreth W Tigwell, Benjamin M Gorman, and Rachel Menzies. 2020. Emoji ac- cessibility for visually impaired people. InProceedings of the 2020 CHI Conference on Human Factors in Computing Systems . 1–14

  70. [78]

    Tess Van Daele, Akhil Iyer, Yuning Zhang, Jalyn C Derry, Mina Huh, and Amy Pavel. 2024. Making Short-Form Videos Accessible with Hierarchical Video Summaries. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–17

  71. [79]

    Ike Vayansky and Sathish AP Kumar. 2020. A review of topic modeling methods. Information Systems 94 (2020), 101582

  72. [80]

    Violeta Voykinska, Shiri Azenkot, Shaomei Wu, and Gilly Leshed. 2016. How blind people interact with visual content on social networking services. In Proceedings of the 19th acm conference on computer-supported cooperative work & social computing. 1584–1595

  73. [81]

    Agnieszka Walczak and Louise Fryer. 2017. Creative description: The impact of audio description style on presence in visually impaired audiences. British Journal of Visual Impairment 35, 1 (2017), 6–17

  74. [82]

    Ruolin Wang, Zixuan Chen, Mingrui Ray Zhang, Zhaoheng Li, Zhixiu Liu, Zihan Dang, Chun Yu, and Xiang’Anthony’ Chen. 2021. Revamp: Enhancing accessible information seeking experience of online shopping for blind or low vision users. In Proceedings of the 2021 CHI Conference on ...

  75. [83]

    Xiao Wang, Guangyao Chen, Guangwu Qian, Pengcheng Gao, Xiao-Yong Wei, Yaowei Wang, Yonghong Tian, and Wen Gao. 2023. Large-scale multi-modal pre-trained models: A comprehensive survey. Machine Intelligence Research 20, 4 (2023), 447–482

  76. [84]

    Yujia Wang, Wei Liang, Haikun Huang, Yongqi Zhang, Dingzeyu Li, and Lap- Fai Yu. 2021. Toward automatic audio description generation for accessible videos. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems. 1–12

  77. [85]

    Jonatas Wehrmann, Camila Kolling, and Rodrigo C Barros. 2020. Adaptive cross-modal embeddings for image-text alignment. In Proceedings of the AAAI conference on artificial intelligence , Vol. 34. 12313–12320

  78. [86]

    Frank Wilcoxon, S Katti, Roberta A Wilcox, et al . 1970. Critical values and probability levels for the Wilcoxon rank sum test and the Wilcoxon signed rank test. Selected tables in mathematical statistics 1 (1970), 171–259

  79. [87]

    Qunfang Wu, Yisi Sang, and Yun Huang. 2019. Danmaku: A new paradigm of social interaction via online videos. ACM Transactions on Social Computing 2, 2 (2019), 1–24

  80. [88]

    Qunfang Wu, Yisi Sang, Shan Zhang, and Yun Huang. 2018. Danmaku vs. forum comments: understanding user participation and knowledge sharing in online videos. In Proceedings of the 2018 ACM International Conference on Supporting Group Work. 209–218

  81. [89]

    Shaomei Wu, Jeffrey Wieland, Omid Farivar, and Julie Schiller. 2017. Automatic alt-text: Computer-generated image descriptions for blind users on a social network service. Inproceedings of the 2017 ACM conference on computer supported cooperative work and social computing . 1180–1192

  82. [90]

    Ying Xiang and Seong Wook Chae. 2022. Influence of perceived interactivity on continuous use intentions on the danmaku video sharing platform: Belonging- ness perspective. International Journal of Human–Computer Interaction 38, 6 (2022), 573–593

  83. [91]

    Shuchang Xu, Chang Chen, Zichen Liu, Xiaofu Jin, Lin-Ping Yuan, Yukang Yan, and Huamin Qu. 2024. Memory Reviver: Supporting Photo-Collection Reminiscence for People with Visual Impairment via a Proactive Chatbot. In Proceedings of the 37th Annual ACM Symposium on User Interfac...

  84. [92]

    Shuchang Xu, Ciyuan Yang, Wenhao Ge, Chun Yu, and Yuanchun Shi. 2020. Virtual Paving: Rendering a Smooth Path for People with Visual Impairment through Vibrotactile and Audio Feedback. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 4, 3, Article 99 (Sept. 2020), 25 page...

  85. [93]

    Xuhai Xu, Jun Gong, Carolina Brum, Lilian Liang, Bongsoo Suh, Shivam Kumar Gupta, Yash Agarwal, Laurence Lindsey, Runchang Kang, Behrooz Shahsavari, et al. 2022. Enabling hand gesture customization on wrist-worn devices. In Proceedings of the 2022 CHI Conference on Human Facto...

  86. [94]

    Yukang Yan, Yingtian Shi, Chun Yu, and Yuanchun Shi. 2020. Headcross: Ex- ploring head-based crossing selection on head-mounted displays. Proceedings of the ACM on interactive, mobile, wearable and ubiquitous technologies 4, 1 (2020), 1–22

  87. [95]

    Yukang Yan, Chun Yu, Xin Yi, and Yuanchun Shi. 2018. Headgesture: Hands-free input approach leveraging head movements for hmd devices. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 2, 4 (2018), 1–23

  88. [96]

    Ciyuan Yang, Shuchang Xu, Tianyu Yu, Guanhong Liu, Chun Yu, and Yuanchun Shi. 2021. LightGuide: Directing Visually Impaired People along a Path Using Light Cues. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 5, 2, Article 84 (June 2021), 27 pages. https://doi.org/10.11...

  89. [97]

    Sibei Yang, Guanbin Li, and Yizhou Yu. 2019. Cross-modal relationship inference for grounding referring expressions. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 4145–4154

  90. [98]

    Matin Yarmand, Dongwook Yoon, Samuel Dodson, Ido Roll, and Sidney S Fels

  91. [99]

    Mingrui Ray Zhang, Ruolin Wang, Xuhai Xu, Qisheng Li, Ather Sharif, and Jacob O Wobbrock. 2021. Voicemoji: Emoji entry using voice for visually im- paired people. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems. 1–18

  92. [100]

    egg dishes

    Mingrui Ray Zhang, Mingyuan Zhong, and Jacob O Wobbrock. 2022. Ga11y: An automated gif annotation system for visually impaired users. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems . 1–16. CHI ’25, April 26-May 1, 2025, Yokohama, Japan Shuchang...

  93. [2019]

    Can you believe [1: 21]?!

    " Can you believe [1: 21]?!" Content and Time-Based Reference Patterns in Video Comments. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems. 1–12

  94. [2020]

    Making GIFs Accessible. In Proceedings of the 22nd International ACM CHI ’25, April 26-May 1, 2025, Yokohama, Japan Shuchang Xu, Xiaofu Jin, Huamin Qu, and Yukang Yan SIGACCESS Conference on Computers and Accessibility . 1–10

  95. [2024]

    It’s Kind of Context Dependent

    “It’s Kind of Context Dependent”: Understanding Blind and Low Vision People’s Video Accessibility Preferences Across Viewing Scenarios. In Proceed- ings of the CHI Conference on Human Factors in Computing Systems . 1–20

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.