REVIEW 3 major objections 6 minor 104 references
Sonic Stage: Auto-Generating Interactive Spatial Soundscapes to Facilitate Dialogue Video Comprehension for Blind Viewers
T0 review · 3 major / 6 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read Sonic Stage shows that a coherent 3D soundscape—spatialized dialogue, diegetic sounds, and tap-for details—can carry the visual information that audio description cannot fit into speech-dense scenes, improving blind viewers' comprehension.
desk verdict A solid systems contribution with a real advance in scene-space audio; the comparative claim leans on an unvalidated baseline, but the core result is likely to hold. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the coherent 3D soundscape: a single reconstructed scene, built from sampled frames of the clip, into which all auditory cues are placed. The pipeline samples full and medium shots with moving characters, reconstructs a 3D point cloud of the space with a feed-forward model, tracks characters to recover their trajectories, and then optimizes the soundscape—listener at the characters' geometric center, left–right axis along the direction of greatest positional variance, logarithmic volume roll-off with distance, two-second trajectory smoothing, and a 70% spatial / 30% mono blend at speaker transitions. This shared reference frame is what keeps spatialized dialogue and diegetic
What would settle it
Take the same six user-study videos and re-run Sonic Stage with the dialogue artificially overlapped or with a scene change inserted mid-clip, then measure trajectory accuracy and recall. The paper's own stated assumptions predict a sharp drop: speaker labels would fail to link to on-screen characters, and the reconstructed 3D anchors would not exist. If comprehension accuracy for position and movement stayed near the reported 89% and 86% under those conditions, the boundary condition would be wrong; if it falls, the central claim is confirmed to hold only within the stated scope.
Extended reading notes
Core claim
On the authors' own terms, Sonic Stage's discovery is that the visual information audio description is forced to omit can be carried by the soundtrack itself. Three techniques work together: spatialized dialogue places each speaker's voice at a reconstructed 3D position so layout and movement are heard rather than described; diegetic sound effects render actions from the action's location; and interactive descriptions, invoked by a tap, supply dialogue-relevant details without pausing playback for long. The load-bearing design choice is scene-space anchoring: character trajectories and sound sources live in one reconstructed 3D scene, so camera cuts no longer jerk the audio around the way sc
Load-bearing premise
The system's benefit depends on the clip being a single physical space with non-overlapping speech and enough visual anchors for 3D reconstruction—the paper says so in its limitations—so scenes with overlapping dialogue, mid-scene cuts to new locations, or open environments would break the speaker linkage and spatial coherence that the user-study results rely on.
Editorial extensions
If this is right
- Spatialized dialogue alone appears to carry layout and movement: participants scored 89% on position and 86% on movement, versus 44% and 18% with the touch-exploration baseline, suggesting audio description could offload spatial information to the soundtrack.
- Diegetic sound with a half-second vibration made exploration self-timed: users explored more often with shorter pauses and yet watched no longer overall, evidence that cues can guide attention without stalling the narrative.
- Because cues live in scene space rather than screen space, the same mechanism should survive camera cuts, which the paper identifies as the main cause of confusion in screen-space interactive systems.
- The pipeline is compatible with existing audio description: it fills speech segments while audio description fills non-speech segments, and descriptions could be pruned to avoid redundancy with audio description.
- The system's stated scope is single-space dialogue scenes with non-overlapping speech; results do not yet speak to overlapping dialogue, scene changes, or open environments.
Reading between the lines
- If the 45-point and 68-point gaps on position and movement replicate, the result suggests a design principle beyond accessibility: spatial consistency is an information channel of its own, and interactive systems that present spatial data frame-by-frame in screen space are structurally disadvantaged.
- A testable extension: the same scene-space soundscape could be evaluated on live performances, where no camera cuts exist; the paper's participants already named concerts, opera, and dance as targets, but the system would need body-movement sonification that it does not currently attempt.
- The paper's own constraints imply a falsifiable boundary: on clips with overlapping speech or open scenes, the speaker-to-character linkage and anchor-based reconstruction should degrade, and comprehension gains should shrink; measuring that degradation would clarify how much of the benefit comes from the soundscape versus the selection of easy videos.
- Interactive description selection by dialogue relevance could generalize as a just-in-time audio design pattern: instead of describing everything, systems that surface the one detail tied to the current speech may reduce cognitive load in any audio-first interface.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents Sonic Stage, an automated system that converts dialogue-heavy videos into interactive 3D spatial soundscapes for blind and low-vision (BLV) viewers. Three techniques are combined: spatialized dialogue, diegetic sound effects, and on-demand interactive descriptions, all rendered in a stable 3D scene-space so that spatial cues remain coherent across camera cuts. The system is evaluated in a within-subject user study with 12 BLV participants against a baseline that the authors built and describe as 'modeled after SPICA.' The reported results show significantly higher objective recall for character position, movement, action, and visual detail, as well as higher spatial presence and narrative engagement, with no significant difference in dialogue comprehension. A separate technical evaluation reports 91.9% trajectory accuracy and 94.3% description accuracy on a 16-video dataset.
Significance. If the results hold, Sonic Stage addresses a real and important gap: audio description cannot convey visual actions during dense dialogue, and the paper's proposed auditory techniques are a plausible complement. The user study is a genuine contribution: it uses BLV participants, objective recall questions, a counterbalanced within-subject design, and both quantitative and qualitative analysis. The paper also ships a fully automated pipeline with concrete technical choices, including VGGT-based 3D reconstruction, which is a useful step toward scalable accessible video systems. However, the central comparative claim—that Sonic Stage improves comprehension over the state-of-the-art SPICA-style interaction—is only as strong as the fidelity of the author-implemented baseline, which the manuscript does not validate. The scope of the claim is also narrower than the abstract suggests, because the pipeline and evaluation are restricted to non-overlapping speech, single physical spaces, and dialogue-dominated scenes.
major comments (3)
- [§5.1.2, Appendix A.4, §2.1.4] The comparative claim is load-bearing and depends on an unvalidated baseline. The manuscript says the baseline is 'modeled after SPICA,' but no pilot study, expert audit, or feature-parity validation is reported to show that this implementation reproduces SPICA's spatial/temporal exploration behavior. Section 2.1.4 unqualifiedly refers to a 'comparative study between Sonic Stage and SPICA.' The baseline's movement recall of 18% is below the 33% chance level, which suggests that participants were not merely uninformed but actively confused by the screen-space interaction. This makes it difficult to attribute the large objective gaps (position 44% vs. 89%, movement 18% vs. 86%) to Sonic Stage's intrinsic benefits. The authors should either validate the baseline against SPICA, compare with the original SPICA system, or explicitly reframe the comparison as 'a touch-based screen-space baselin
- [§3.2, §7.5, Abstract] The central claim that Sonic Stage 'transforms dialogue videos' and significantly improves comprehension is bounded by assumptions that are stated only later: non-overlapping speech, a single physical space, sufficient visual anchors for 3D reconstruction, and high dialogue ratios (the user-study videos have 87–97% speech). These conditions hold for the selected clips but fail for much real film and TV, which includes overlapping dialogue, scene changes, and open or moving-camera scenes. The paper acknowledges these limits in §7.5, but the abstract and introduction present the system without these qualifications. The authors should state these boundary conditions prominently in the abstract and introduction, and ideally include an analysis of how the pipeline degrades when each assumption is violated, rather than leaving this to future work.
- [§4.2, §5.2.3] The statistical analysis is acceptable for a UIST-style user study, but the authors should address the multiple-comparison issue. Five per-category paired t-tests are reported without correction; while the effects for position, movement, action, and visual detail are individually significant at p<.01, a correction such as Bonferroni or Holm would make the evidence more robust and is standard for this number of tests. The technical evaluation also lacks an inter-rater reliability metric: Section 4.2 says one researcher labeled and a second 'reviewed the labels,' but no agreement score (e.g., Cohen's kappa) is reported. This should be added or explicitly justified.
minor comments (6)
- [§4.2] The claim that 'most inaccuracies were minor positional shifts below human auditory resolution' is not substantiated with measurements. If quantitative error magnitudes are available, reporting them would strengthen the argument that 91.9% trajectory accuracy is sufficient for the user experience.
- [§3.3.3] The soundscape parameters (V_max, V_min, D_near/D_far, spatial blend factor, smoothing window) are tuned with two BLV sound designers but no systematic sensitivity analysis is provided. A brief exploration of how the results change with these parameters would help establish that the user-study outcomes are not artifacts of a single hand-picked configuration.
- [§3.3.2] The character position is estimated as the mean of ten randomly sampled points from the projected segmentation mask. This is an arbitrary choice; a small sensitivity check (e.g., 5 vs. 20 points) or a rationale based on mask noise would be helpful.
- [Appendix A.4] The baseline description generation is said to use 'the method in the SPICA system' but no details are given about the object description model or prompts. Providing the exact generation pipeline would improve reproducibility.
- [§6.1, Figure 5] The figure labels for significance levels could be clearer: the text reports t-values and p-values, but the figure would benefit from explicit significance markers (e.g., asterisks) and error bars showing within-subject variability, not just between-subject standard errors.
- [§7.5] Minor typos: 'this approach that does not generalize well' should read 'this approach does not generalize well'; 'Future work could how to sonify' should read 'Future work could explore how to sonify.'
Circularity Check
No circularity found: the central claim is an empirical comparison, not a derivation from fitted inputs or self-citation.
full rationale
The paper's central assertion is empirical: a within-subject user study compares Sonic Stage with a baseline 'modeled after SPICA' and reports significantly better comprehension, spatial presence, and engagement. There is no mathematical derivation that reduces the reported outcome to the system's inputs. The technical evaluation uses manual labels and pretrained models (VGGT, YOLOv11, TalkNet, Gemini) against ground-truth video content; this is an external benchmark, not a fitted prediction. Hand-tuned parameters (Vmax, Vmin, Dnear, Dfar, smoothing window, spatial blend factor) were validated with two blind sound designers, but they were not fit to the comprehension outcome, so the headline improvements are not forced by those choices. The paper explicitly acknowledges scope limits—non-overlapping speech, single physical spaces, visual anchors for 3D reconstruction—which constrain when the system works but do not constitute circularity. The phrase in Section 2.1.4 calling the evaluation a 'comparative study between Sonic Stage and SPICA' is a loose description of a baseline modeled after SPICA, and the unvalidated fidelity of that baseline is a correctness/construct-validity risk, not a circularity pattern under the definitions used here. No equation is shown to equal another by construction, and no fitted parameter is renamed as a prediction. Self-citations appear but are not load-bearing for the main result.
Assumptions & free parameters
free parameters (5)
- Volume roll-off endpoints V_max, V_min =
1.0, 0.5
- Spatial blend factor =
0.7
- Trajectory smoothing window =
2 s
- Motion sampling threshold =
25% of bounding-box width/height per second
- Distance roll-off percentiles D_near/D_far =
median / 90th percentile of scene speaker-listener distances
assumptions (7)
- domain assumption Dialogue speech does not overlap between speakers
- domain assumption The scene is contained in a single physical space with enough visual anchors for 3D reconstruction
- domain assumption Stereo spatial audio with HRTF is an adequate channel for conveying spatial layout and movement to BLV viewers
- domain assumption Pretrained models (VGGT, YOLOv11, BoT-SORT, TalkNet, Gemini-2.5, ElevenLabs) are sufficiently accurate for detection, reconstruction, description, and sound generation
- ad hoc to paper A character's 3D position can be approximated by the mean of ten random points in the projected segmentation mask
- domain assumption The author-built baseline is a faithful implementation of SPICA
- domain assumption Manual labeling in the technical evaluation is reliable
Cite this review
Pith. "Pith review of Sonic Stage: Auto-Generating Interactive Spatial Soundscapes to Facilitate Dialogue Video Comprehension for Blind Viewers." pith.science (2026). https://pith.science/paper/KYUQJ7V4
@misc{pith2026260720835,
author = {Pith},
title = {Pith review of: Sonic Stage: Auto-Generating Interactive Spatial Soundscapes to Facilitate Dialogue Video Comprehension for Blind Viewers},
year = {2026},
howpublished = {\url{https://pith.science/paper/KYUQJ7V4}},
note = {Machine review of arXiv:2607.20835}
}
read the original abstract
Audio description (AD) makes film and television accessible to blind and low-vision (BLV) audiences by narrating characters' actions. However, in scenes with lots of dialogue, AD often omits important actions because it is constrained not to overlap with speech. It is not yet known how to convey characters' actions during dialogue. We present Sonic Stage, a system that transforms dialogue videos into interactive spatial soundscapes, enabling BLV audiences to intuitively understand characters' actions and movements through immersive auditory cues. Sonic Stage conveys essential visual information during dialogue through three auditory techniques: (1) spatialized dialogue to represent spatial layout, (2) diegetic sound to convey character actions, and (3) interactive descriptions to provide context-specific visual details. Evaluation with 12 BLV viewers showed that Sonic Stage significantly improved video comprehension, spatial presence, and narrative engagement. We highlight opportunities for enhancing video accessibility across diverse genres through immersive, interactive audio representations.
Figures
Figures from the paper (11 more)
Reference graph
Works this paper leans on
-
[1]
Scene Detect
2024. Scene Detect. https://www.scenedetect.com/. Accessed: 2025-06-11
2024
-
[2]
Unity: Spatial Blend
2025. Unity: Spatial Blend. https://docs.unity3d.com/ScriptReference/ AudioSource-spatialBlend.html. Accessed: 2025-06-11
2025
-
[3]
Nir Aharon, Roy Orfaig, and Ben-Zion Bobrovsky. 2022. BoT-SORT: Robust associations multi-pedestrian tracking.arXiv preprint arXiv:2206.14651(2022)
arXiv 2022
-
[4]
Sercan Arik, Jitong Chen, Kainan Peng, Wei Ping, and Yanqi Zhou. 2018. Neural voice cloning with a few samples.Advances in neural information processing systems31 (2018)
2018
-
[5]
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, et al . 2025. Qwen2. 5-vl technical report.arXiv preprint arXiv:2502.13923(2025)
arXiv 2025
-
[6]
Bela Balazs. 1985. Theory of the film: Sound. 116–125 pages
1985
-
[7]
Maxime Bleau, Camille van Acker, Natalina Martiniello, Joseph Paul Nemargut, and Maurice Ptito. 2023. Cognitive map formation in the blind is enhanced by three-dimensional tactile information.Scientific Reports13, 1 (2023), 9736
2023
-
[8]
James C Bliss, Michael H Katcher, Charles H Rogers, and Raymond P Shepard
Show all 104 references
-
[9]
2010.Immersed in media: Telepres- ence in everyday life
Cheryl Campanella Bracken and Paul Skalski. 2010.Immersed in media: Telepres- ence in everyday life. Routledge
2010
-
[10]
Virginia Braun and Victoria Clarke. 2019. Reflecting on reflexive thematic analysis. Qualitative research in sport, exercise and health11, 4 (2019), 589–597
2019
-
[11]
Rick Busselle and Helena Bilandzic. 2009. Measuring narrative engagement. Media psychology12, 4 (2009), 321–347
2009
-
[12]
Matthew Butler, Leona M Holloway, Samuel Reinders, Cagatay Goncu, and Kim Marriott. 2021. Technology developments in touch-based accessible graphics: A systematic review of research 2010-2020. InProceedings of the 2021 CHI conference on human factors in computing systems. 1–15...
2021
-
[13]
Ben Caldwell, Michael Cooper, Loretta Guarino Reid, Gregg Vanderheiden, Wendy Chisholm, John Slatin, and Jason White. 2008. Web content accessibility guide- lines (WCAG) 2.0.WWW Consortium (W3C)290, 1-34 (2008), 5–12
2008
-
[14]
Anil Çamcı, Kristine Lee, Cody J Roberts, and Angus G Forbes. 2017. INVISO: a cross-platform user interface for creating virtual sonic environments. InProceed- ings of the 30th Annual ACM Symposium on User Interface Software and Technology. 507–518
2017
-
[15]
Ruei-Che Chang, Chao-Hsien Ting, Chia-Sheng Hung, Wan-Chen Lee, Liang- Jin Chen, Yu-Tzu Chao, Bing-Yu Chen, and Anhong Guo. 2022. OmniScribe: Authoring Immersive Audio Descriptions for 360°Videos. InProceedings of the 35th Annual ACM Symposium on User Interface Software and Te...
2022
-
[16]
Maryam Cheema, Sina Elahimanesh, Samuel Martin, Pooyan Fazli, and Hasti Seifi
-
[17]
Maryam Cheema, Hasti Seifi, and Pooyan Fazli. 2025. Describe Now: User-Driven Audio Description for Blind and Low Vision Individuals. InProceedings of the 2025 ACM Designing Interactive Systems Conference. 458–474
2025
-
[18]
Boyuan Chen, Zhuo Xu, Sean Kirmani, Brain Ichter, Dorsa Sadigh, Leonidas Guibas, and Fei Xia. 2024. Spatialvlm: Endowing vision-language models with spatial reasoning capabilities. InProceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition. 14455–14465
2024
-
[19]
Chang Chen, Sicheng Song, Shuchang Xu, Zhicheng Li, Huamin Qu, and Yanna Lin. 2025. RhythmTA: A Visual-Aided Interactive System for ESL Rhythm Train- ing via Dubbing Practice. InProceedings of the 38th Annual ACM Symposium on User Interface Software and Technology. 1–15
2025
-
[20]
Shi Chen, Jingao Zhang, Suqi Lou, Xiaodong Wang, Wei Xiang, and Lingyun Sun. 2025. Voice by the Non-sighted: Practices and Challenges of Audiobook Voice Actors with Blind and Low Vision in China. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems. 1–19
2025
-
[21]
Arnavi Chheda-Kothary, Ather Sharif, David Angel Rios, and Brian A Smith
-
[22]
2019.Audio-vision: sound on screen
Michel Chion. 2019.Audio-vision: sound on screen. Columbia University Press
2019
-
[23]
Hyunsung Cho, Alexander Wang, Divya Kartik, Emily Liying Xie, Yukang Yan, and David Lindlbauer. 2024. Auptimize: Optimal placement of spatial audio cues for extended reality. InProceedings of the 37th Annual ACM Symposium on User Interface Software and Technology. 1–14
2024
-
[24]
It Brought Me Joy
" It Brought Me Joy": Opportunities for Spatial Browsing in Desktop Screen Readers. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems. 1–18
2025
-
[25]
Hainan Cui, Xiang Gao, Shuhan Shen, and Zhanyi Hu. 2017. HSfM: Hybrid structure-from-motion. InProceedings of the IEEE conference on computer vision and pattern recognition. 1212–1221
2017
-
[26]
Ginger Delmas, Philippe Weinzaepfel, Thomas Lucas, Francesc Moreno-Noguer, and Grégory Rogez. 2022. Posescript: 3d human poses from natural language. In European Conference on Computer Vision. Springer, 346–362
2022
-
[27]
Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, et al. 2025. Gemini 2.5: Pushing the frontier with advanced reasoning, multi- modality, long context, and next generation agentic...
2025 arXiv
-
[28]
Alessandro Favero, Luca Zancato, Matthew Trager, Siddharth Choudhary, Pra- muditha Perera, Alessandro Achille, Ashwin Swaminathan, and Stefano Soatto
-
[29]
Gaudio. 2025. AI stem splitter. https://www.gaudiolab.com/gaudio-studio/
2025
-
[30]
Hilko Donker, Palle Klante, and Peter Gorny. 2002. The design of auditory user interfaces for blind users. InProceedings of the second Nordic conference on Human-computer interaction. 149–156
2002
-
[31]
Google. 2025. Gemini Audio Understanding. https://ai.google.dev/gemini-api/ docs/audio
2025
-
[32]
Google. 2025. Gemini Video Understanding. https://ai.google.dev/gemini-api/ docs/video-understanding
2025
-
[33]
João Guerreiro, Yujin Kim, Rodrigo Nogueira, SeungA Chung, André Rodrigues, and Uran Oh. 2023. The design space of the auditory representation of objects and their behaviours in virtual reality for blind people.IEEE Transactions on visualization and computer graphics29, 5 (202...
2023
-
[34]
Rohit Girmaji, Bhav Beri, Ramanathan Subramanian, and Vineet Gandhi. 2025. EditIQ: Automated Cinematic Editing of Static Wide-Angle Videos via Dialogue Interpretation and Saliency Cues. InProceedings of the 30th International Confer- ence on Intelligent User Interfaces. 609–623
2025
-
[35]
Tilo Hartmann, Werner Wirth, Holger Schramm, Christoph Klimmt, Peter Vorderer, André Gysbers, Saskia Böcking, Niklas Ravaja, Jari Laarni, Timo Saari, et al. 2015. The spatial presence experience scale (SPES).Journal of Media Psychology(2015)
2015
-
[36]
Leona Holloway, Kim Marriott, and Matthew Butler. 2018. Accessible maps for the blind: Comparing 3D printed models with tactile graphics. InProceedings of the 2018 CHI conference on human factors in computing systems. 1–13
2018
-
[37]
Gaurav Jain, Basel Hindi, Connor Courtien, Xin Yi Therese Xu, Conrad Wyrick, Michael Malcolm, and Brian A. Smith. 2023. Front Row: Automatically Generating Immersive Audio Representations of Tennis Broadcasts for Blind Viewers. In Proceedings of the 36th Annual ACM Symposium o...
2023
-
[38]
2003.Multiple view geometry in computer vision
Richard Hartley and Andrew Zisserman. 2003.Multiple view geometry in computer vision. Cambridge university press
2003
-
[39]
Lucy Jiang, Crescentia Jung, Mahika Phutane, Abigale Stangl, and Shiri Azenkot
-
[40]
Lucy Jiang, Mahika Phutane, and Shiri Azenkot. 2023. Beyond Audio Description: Exploring 360°Video Accessibility with Blind and Low Vision Users Through Collaborative Creation. InProceedings of the 25th International ACM SIGACCESS Conference on Computers and Accessibility(New ...
2023
-
[41]
Ziyi Jiang, Mengjie Jian, Jiajia Hu, Hongze Zhao, Huamin Qu, Shuchang Xu, and Guanhong Liu. 2025. From Audio Description to Movement: Challenges and Design Principles in Video-based Body Exercise for Blind and Low Vision Users. InCompanion Publication of the 2025 Conference on...
2025
-
[42]
Chutian Jiang, Emily Kuang, and Mingming Fan. 2025. How can haptic feedback assist people with blind and low vision (BLV): A systematic literature review. ACM Transactions on Accessible Computing18, 1 (2025), 1–57
2025
-
[43]
Rahima Khanam and Muhammad Hussain. 2024. Yolov11: An overview of the key architectural enhancements.arXiv preprint arXiv:2410.17725(2024)
2024 arXiv
-
[44]
It’s Kind of Context Dependent
“It’s Kind of Context Dependent”: Understanding Blind and Low Vision People’s Video Accessibility Preferences Across Viewing Scenarios. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems(Honolulu, HI, USA)(CHI ’24). Association for Computing Machine...
2024
-
[45]
Mackenzie Leake, Abe Davis, Anh Truong, and Maneesh Agrawala. 2017. Com- putational video editing for dialogue-driven scenes.ACM Trans. Graph.36, 4 (2017), 130–1
2017
-
[46]
James R Lewis. 2018. The system usability scale: past, present, and future.Inter- national Journal of Human–Computer Interaction34, 7 (2018), 577–590
2018
-
[47]
Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis
-
[48]
David Chuan-En Lin, Anastasis Germanidis, Cristóbal Valenzuela, Yining Shi, and Nikolas Martelaro. 2023. Soundify: Matching sound effects to video. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. 1–13
2023
-
[49]
Guanhong Liu, Tianyu Yu, Chun Yu, Haiqing Xu, Shuchang Xu, Ciyuan Yang, Feng Wang, Haipeng Mi, and Yuanchun Shi. 2021. Tactile compass: Enabling visually impaired people to follow a path with continuous directional feedback. InProceedings of the 2021 CHI Conference on Human Fa...
2021
-
[50]
Daniel Killough and Amy Pavel. 2023. Exploring Community-Driven Descriptions for Making Livestreams Accessible. InProceedings of the 25th International ACM SIGACCESS Conference on Computers and Accessibility. 1–13
2023
-
[51]
Sheng Liu, Xiaohan Nie, and Raffay Hamid. 2022. Depth-guided sparse structure- from-motion for movies and tv shows. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 15980–15989
2022
-
[52]
Xingyu Liu, Patrick Carrington, Xiang’Anthony’ Chen, and Amy Pavel. 2021. What makes videos accessible to blind and visually impaired people?. InPro- ceedings of the 2021 CHI Conference on Human Factors in Computing Systems. 1–14
2021
-
[53]
Chaoyu Li, Sid Padmanabhuni, Maryam S Cheema, Hasti Seifi, and Pooyan Fazli
-
[54]
In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25)
VideoA11y: Method and Dataset for Accessible Video Description. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). Association for Computing Machinery, New York, NY, USA, Article 1055, 29 pages. doi:10.1145/3706598.3714096
2025
-
[55]
Mariana Lopez, Gavin Kearney, and Krisztián Hofstädter. 2022. Seeing films through sound: Sound design, spatial audio, and accessibility for visually impaired audiences.British Journal of Visual Impairment40, 2 (2022), 117–144. Sonic Stage UIST ’26, November 02–05, 2026, Detro...
2022
-
[56]
Mariana Julieta Lopez, Gavin Kearney, and Krisztian Hofstadter. 2021. Enhancing audio description: Inclusive cinematic experiences through sound design.Journal of Audiovisual Translation4, 1 (2021), 157–182
2021
-
[57]
Hongbo Liu, Jingwen He, Yi Jin, Dian Zheng, Yuhao Dong, Fan Zhang, Ziqi Huang, Yinan He, Yangguang Li, Weichao Chen, et al. 2025. ShotBench: Expert- Level Cinematic Understanding in Vision-Language Models.arXiv preprint arXiv:2506.21356(2025)
2025
-
[58]
Troy McDaniel, Lakshmie Narayan Viswanathan, and Sethuraman Panchanathan
-
[59]
2012.Acoustics: sound fields and transducers
Tim Mellow. 2012.Acoustics: sound fields and transducers. Academic Press
2012
-
[60]
Xingyu Liu, Biao Wang, Wayne Zhang, Ziqian Liao, Ziwen Li, Amy Pavel, Xi- ang’Anthony’ Chen, et al. 2025. CoSight: Exploring Viewer Contributions to Online Video Accessibility Through Descriptive Commenting.arXiv preprint arXiv:2508.08582(2025)
2025 arXiv
-
[61]
Xingyu" Bruce" Liu, Ruolin Wang, Dingzeyu Li, Xiang Anthony Chen, and Amy Pavel. 2022. CrossA11y: Identifying Video Accessibility Issues via Cross-modal Grounding. InProceedings of the 35th Annual ACM Symposium on User Interface Software and Technology. 1–14
2022
-
[62]
Rosiana Natalie, Joshua Tseng, Hernisa Kacorri, and Kotaro Hara. 2023. Support- ing novices author audio descriptions via automatic feedback. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–18
2023
-
[63]
Netflix. 2024. Audio Description Style Guide v2.5. https://partnerhelp. netflixstudios.com/hc/en-us/articles/215510667-Audio-Description-Style- Guide-v2-5. Accessed: 2025-06-29
2024
-
[64]
Michał Maćkowski, Piotr Brzoza, Mateusz Kawulok, Rafał Meisel, and Dominik Spinczyk. 2023. Multimodal presentation of interactive audio-tactile graphics supporting the perception of visual information by blind people.ACM Transac- tions on Multimedia Computing, Communications a...
2023
-
[65]
Zheng Ning, Zheng Zhang, Jerrick Ban, Kaiwen Jiang, Ruohong Gan, Yapeng Tian, and Toby Jia-Jun Li. 2024. MIMOSA: Human-AI Co-Creation of Computational Spatial Audio Effects on Videos. InProceedings of the 16th Conference on Creativity & Cognition(Chicago, IL, USA)(C&C ’24). As...
2024
-
[66]
Ofcom. 2024. Guidelines on Providing TV and On-Demand Access Services. https://www.ofcom.org.uk/siteassets/resources/documents/tv-radio-and-on- demand/broadcast-codes/other-codes/ofcoms-guidelines-on-providing-tv- and-on-demand-access-services.pdf. Accessed: 2025-06-29
2024
-
[67]
Amy Pavel, Gabriel Reyes, and Jeffrey P. Bigham. 2020. Rescribe: Authoring and Automatically Editing Audio Descriptions. InProceedings of the 33rd Annual ACM Symposium on User Interface Software and Technology(Virtual Event, USA) (UIST ’20). Association for Computing Machinery...
2020
-
[68]
Pedro Morgado, Nuno Nvasconcelos, Timothy Langlois, and Oliver Wang. 2018. Self-supervised generation of spatial audio for 360 video.Advances in neural information processing systems31 (2018)
2018
-
[69]
Rosiana Natalie, Ruei-Che Chang, Smitha Sheshadri, Anhong Guo, and Kotaro Hara. 2024. Audio Description Customization. InProceedings of the 26th Inter- national ACM SIGACCESS Conference on Computers and Accessibility(St. John’s, NL, Canada)(ASSETS ’24). Association for Computi...
2024
-
[70]
SC Lannom. 2025. Guide to Camera Shots: Every Shot Size Explained. https: //www.studiobinder.com/blog/types-of-camera-shots-sizes-in-film/
2025
-
[71]
Johannes L Schonberger and Jan-Michael Frahm. 2016. Structure-from-motion revisited. InProceedings of the IEEE conference on computer vision and pattern recognition. 4104–4113
2016
-
[72]
Zheng Ning, Brianna L Wimer, Kaiwen Jiang, Keyi Chen, Jerrick Ban, Yapeng Tian, Yuhang Zhao, and Toby Jia-Jun Li. 2024. SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision Viewers. InProceedings of the 2024 CHI Conference o...
2024
-
[73]
Ruijie Tao, Zexu Pan, Rohan Kumar Das, Xinyuan Qian, Mike Zheng Shou, and Haizhou Li. 2021. Is someone speaking? exploring long-term temporal features for audio-visual active speaker detection. InProceedings of the 29th ACM international conference on multimedia. 3927–3935
2021
-
[74]
Tencent. 2025. Tencent Transcription. https://cloud.tencent.com/product/asr/
2025
-
[75]
Dimitrios Tzovaras, Georgios Nikolakis, Georgios Fergadis, Stratos Malasiotis, and Modestos Stavrakis. 2004. Design and implementation of haptic virtual environments for the training of the visually impaired.IEEE Transactions on Neural Systems and Rehabilitation Engineering12,...
2004
-
[76]
Yi-Hao Peng, Jeffrey P Bigham, and Amy Pavel. 2021. Slidecho: Flexible non- visual exploration of presentation videos. InProceedings of the 23rd International ACM SIGACCESS Conference on Computers and Accessibility. 1–12
2021
-
[77]
Minghan Qin, Wanhua Li, Jiawei Zhou, Haoqian Wang, and Hanspeter Pfister
-
[78]
InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Langsplat: 3d language gaussian splatting. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 20051–20060
-
[79]
Lakshmie Narayan Viswanathan, Troy McDaniel, Sreekar Krishna, and Sethu- raman Panchanathan. 2010. Haptics in audio described movies. In2010 IEEE International Symposium on Haptic Audio Visual Environments and Games. IEEE, 1–2
2010
-
[80]
Volcengine. 2025. Volcengine Text-to-Speech. https://www.volcengine.com/ product/tts
2025
-
[81]
Abigale Stangl, Shasta Ihorn, Yue-Ting Siu, Aditya Bodi, Mar Castanon, Lothar D Narins, and Ilmi Yoon. 2023. The potential of a visual dialogue agent in a tan- dem automated audio description system for videos. InProceedings of the 25th International ACM SIGACCESS Conference o...
2023
-
[82]
Jianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi, Christian Rup- precht, and David Novotny. 2025. Vggt: Visual geometry grounded transformer. InProceedings of the Computer Vision and Pattern Recognition Conference. 5294– 5306
2025
-
[83]
Yujia Wang, Wei Liang, Haikun Huang, Yongqi Zhang, Dingzeyu Li, and Lap-Fai Yu. 2021. Toward Automatic Audio Description Generation for Accessible Videos. InProceedings of the 2021 CHI Conference on Human Factors in Computing Systems (Yokohama, Japan)(CHI ’21). Association for...
2021
-
[84]
Frank Wilcoxon, S Katti, Roberta A Wilcox, et al . 1970. Critical values and probability levels for the Wilcoxon rank sum test and the Wilcoxon signed rank test.Selected tables in mathematical statistics1 (1970), 171–259
1970
-
[85]
Valve Corporation. 2025. Steam Audio. https://valvesoftware.github.io/steam- audio
2025
-
[86]
Tess Van Daele, Akhil Iyer, Yuning Zhang, Jalyn C Derry, Mina Huh, and Amy Pavel. 2024. Making Short-Form Videos Accessible with Hierarchical Video Summaries. InProceedings of the CHI Conference on Human Factors in Computing Systems. 1–17
2024
-
[87]
Valentijn T Visch, Ed S Tan, and Dylan Molenaar. 2010. The emotional and cognitive effect of immersion in film viewing.Cognition and Emotion24, 8 (2010), 1439–1445
2010
-
[88]
Shuchang Xu, Xiaofu Jin, Huamin Qu, and Yukang Yan. 2025. DanmuA11y: Making Time-Synced On-Screen Video Comments (Danmu) Accessible to Blind and Low Vision Users via Multi-Viewer Audio Discussions. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems ...
2025
-
[89]
Shuchang Xu, Xiaofu Jin, Wenshuo Zhang, Huamin Qu, and Yukang Yan. 2025. Branch Explorer: Leveraging Branching Narratives to Support Interactive 360° Video Viewing for Blind and Low Vision Users.arXiv preprint arXiv:2507.09959 (2025)
2025 arXiv
-
[90]
Jindu Wang, Runze Cai, Shuchang Xu, Tianrui Hu, Huamin Qu, Shengdong Zhao, and Lin-Ping Yuan. 2026. Wearable AR for Restorative Breaks: How Interactive Narrative Experiences Support Relaxation for Young People. InProceedings of the 2026 CHI Conference on Human Factors in Compu...
2026
-
[91]
Xudong Xu, Hang Zhou, Ziwei Liu, Bo Dai, Xiaogang Wang, and Dahua Lin
-
[92]
Baiqiao Yin, Qineng Wang, Pingyue Zhang, Jianshu Zhang, Kangrui Wang, Zihan Wang, Jieyu Zhang, Keshigeyan Chandrasegaran, Han Liu, Ranjay Kr- ishna, et al. 2025. Spatial Mental Modeling from Limited Views.arXiv preprint arXiv:2506.21458(2025)
2025
-
[93]
Yuhang Zhao, Cynthia L Bennett, Hrvoje Benko, Edward Cutrell, Christian Holz, Meredith Ringel Morris, and Mike Sinclair. 2018. Enabling people with visual impairments to navigate virtual reality with a haptic and auditory cane simulation. InProceedings of the 2018 CHI conferen...
2018
-
[94]
2013.Head-related transfer function and virtual auditory display
Bosun Xie. 2013.Head-related transfer function and virtual auditory display. J. Ross Publishing
2013
-
[95]
Shuchang Xu, Chang Chen, Zichen Liu, Xiaofu Jin, Lin-Ping Yuan, Yukang Yan, and Huamin Qu. 2024. Memory reviver: supporting photo-collection reminiscence for people with visual impairment via a proactive Chatbot. InProceedings of the 37th Annual ACM Symposium on User Interface...
2024
-
[96]
Smith, and Yukang Yan
Shuchang Xu, Xiaofu Jin, Gaurav Jain, Wenshuo Zhang, Huamin Qu, Brian A. Smith, and Yukang Yan. 2026. Sonic Stage: Automatically Generating an In- teractive Spatial Soundscape to Facilitate Dialogue Video Comprehension for Blind and Low Vision Viewers. InProceedings of the Ext...
2026
-
[99]
Shuchang Xu, Ciyuan Yang, Wenhao Ge, Chun Yu, and Yuanchun Shi. 2020. Virtual Paving: Rendering a smooth path for people with visual impairment through vibrotactile and audio feedback.Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies4, 3 (2020), 1–25
2020
-
[104]
Hang Zhou, Xudong Xu, Dahua Lin, Xiaogang Wang, and Ziwei Liu. 2020. Sep- stereo: Visually guided stereophonic audio generation by associating source separation. InEuropean Conference on Computer Vision. Springer, 52–69. UIST ’26, November 02–05, 2026, Detroit, MI, USA Xu et a...
2020
-
[2007]
Optical-to-tactile image conversion for the blind.IEEE Transactions on Man-Machine Systems11, 1 (2007), 58–65
2007
-
[2013]
In2013 IEEE International Conference on Multimedia and Expo (ICME)
An evaluation of haptic descriptions for audio described films for individuals who are blind. In2013 IEEE International Conference on Multimedia and Expo (ICME). IEEE, 1–6
-
[2021]
In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Visually informed binaural audio generation without binaural audios. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 15485–15494
-
[2023]
Graph.42, 4 (2023), 139–1
3D Gaussian splatting for real-time radiance field rendering.ACM Trans. Graph.42, 4 (2023), 139–1
2023
-
[2024]
InPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Multi-modal hallucination control by visual information grounding. InPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 14303–14312
-
[2025]
arXiv preprint arXiv:2508.01092(2025)
DescribePro: Collaborative Audio Description with Human-AI Interaction. arXiv preprint arXiv:2508.01092(2025)
2025 arXiv
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.