Pith. sign in

REVIEW 3 major objections 6 minor 1 cited by

Branch Explorer: Leveraging Branching Narratives to Support Interactive 360{\deg} Video Viewing for Blind and Low Vision Users

T0 review · 3 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Converting 360° videos into branching narratives restores genuine interactivity for blind and low vision viewers, a 12-participant study reports.

desk verdict A solid accessibility systems paper with a real user study; the agency findings are credible, but the paper must close an internal-validity gap on how baseline materials were corrected. read the letter →

arxiv 2507.09959 v1 pith:RD6IKHFE submitted 2025-07-14 cs.HC

classification cs.HC
keywords BlindLowVisionVisualImpairment360°VideoAudioDescriptionInteractiveStorytellingBranchingNarrativesAccessibility
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Branch Explorer claims that the reason blind and low vision (BLV) viewers are excluded from 360° video is not the medium itself but the linear audio-description model bolted onto it. The paper proposes replacing that model with branching narratives—stories that pause at detected decision points and fork according to the viewer's choice—and shows a fully automated pipeline can build such stories from any 360° video. In a within-subject study with 12 BLV viewers, the system produced significantly stronger feelings of freedom, meaningful choice, and narrative presence than a pause-and-explore baseline. If the result holds, interactive 360° video no longer requires a sighted audio-describer to hand-author every path, which matters because the medium's core promise is choosing your own viewing path.

What carries the argument

The central object is the branching narrative: a story graph whose nodes are branching points and whose edges are viewport paths through the 360° footage. Branch Explorer constructs this graph automatically. A saliency prediction model supplies frame-level viewing directions, which are clustered and linked into candidate paths; a weighted diversity score combining angular separation, caption-semantic distance, and aggregated saliency selects which paths become branches; and a continuation-writing language-model step generates second-person narrations and navigation cues. The interaction layer maps three gestures to three exploration levels—shake for the two suggested branches, swipe to browse the full branch list, and head-turn for object descriptions during pauses—so that choice never requires leaving the narrative flow.

What would settle it

Re-run the 12-viewer comparison with a baseline that fully reproduces the spatialized object descriptions and authoring workflow of the prior pause-and-explore system; if the agency and narrative-presence differences disappear, the headline result is an artifact of a weakened baseline.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that a branching-narrative structure—not richer descriptions of a static paused frame—is what unlocks interactive 360° video for BLV audiences. The authors show that branching points can be detected automatically by avoiding speech and loud audio, aligning with scene transitions, and enforcing a 30-second minimum interval; that branch options can be curated to maximize spatial, semantic, and social diversity; and that coherent second-person narration plus lightweight gestures (shake to choose, swipe to browse, head-turn to explore) keeps the experience immersive. Evaluated against a baseline that stripped out all branching features, Branch Explorer scored significantly higher on user agency (freedom and meaningful choice, both $p < .01$), narrative presence ($p < .01$), and willingness to re-watch ($p < .01$), while comprehension questions were answered correctly 95.8% versus 87.5% of the time. The authors also report that users developed their own exploration strategies, such as switching branches to synthesize perspectives or pausing to track spatial changes over time.

Load-bearing premise

The headline gains rest on the pause-and-explore baseline faithfully representing the state-of-the-art alternative, and on description corrections being applied symmetrically to both study conditions.

Editorial extensions

If this is right

  • If the effect is real, BLV viewers can experience the signature interactivity of 360° video—choosing where to look and which story to follow—without hand-authored audio description.
  • The same pipeline could be applied to existing 360° catalogs, since it needs only the video itself and generic models for saliency, captioning, and narration.
  • The three diversity axes give system designers concrete knobs for personalization: spatial, semantic, and social weights could be tuned per user or per viewing session.
  • Because comprehension improved alongside agency, the branching structure appears to support rather than undermine understanding of the main storyline.
  • The identified exploration strategies, such as branch-switching to synthesize and pausing to track layout change, point to new interaction features like slow-motion playback or context-aware cueing.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the evaluation compares a full system against a control that removes only branching features, the study does not isolate how much of the agency gain comes from the diversity-optimization pipeline versus the simple fact of offering choices; a minimal two-branch version might capture most of the effect.
  • The 78% narration accuracy suggests fully automatic branching is viable for static scenes but will likely require object tracking or human-in-the-loop correction for fast-moving cinematic content.
  • The saliency-plus-diversity machinery could be inverted: instead of selecting branches for a viewer, it could predict which directions sighted viewers will diverge toward, informing viewport-aware streaming or automatic shot selection.
  • The swipe-to-browse hierarchy points toward a scalable navigation paradigm—branch lists could become a searchable index over an entire 360° library rather than just one video's decision points.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This paper presents Branch Explorer, a system that converts 360° videos into accessible branching narratives for blind and low vision (BLV) viewers. The pipeline detects branch points from scene, speech, and audio constraints, generates viewing-direction branches from saliency maps, optimizes branch sets for spatial, semantic, and social diversity, produces coherent second-person narrations and navigation cues with GPT-4o/VideoA11y, and supports interaction through phone shakes, swipes, and head turns. A formative study with eight BLV participants motivated three design considerations: diverse branch options, smooth story progression, and immersive navigation. A technical evaluation on 18 videos reports branch-timing agreement comparable to two professional describers, a diversity rating of 6.65/7, and description accuracies of 78% for branch narrations, 96% for navigation cues, and 89% for object descriptions. A within-subject user study with 12 BLV participants compared Branch Explorer with a baseline that removed all branching features; the authors report significant improvements in agency, narrative presence, enjoyment, and willingness to use, as well as a borderline improvement in comprehension (p = .05). The paper concludes with design implications, personalization directions, and limitations.

Significance. If the results hold, Branch Explorer is a meaningful step toward making interactive 360° video accessible to BLV audiences. The work is notable for combining a fully automatic branch-generation pipeline with a user study involving BLV participants, and for grounding design decisions in a formative study. The technical evaluation uses professional audio describers and reports not only positive outcomes but also null results (e.g., spatial presence), which increases confidence in the balance of the reporting. The within-subject design is counterbalanced, and the main agency and immersion effects are large and consistent across several measures. The paper does not release code or study materials, but the pipeline description is detailed enough to support replication. The main weaknesses are reporting gaps around the fairness of the baseline and the manual description-correction procedure; these are fixable in revision and do not undermine the overall contribution.

major comments (3)
  1. [6.1.4] The manual-correction procedure is ambiguous about symmetry. The text says 'we manually corrected errors in descriptions only when they contradicted the visual content,' but it does not state that the same corrections were applied to the Branch Explorer materials and the baseline materials. Since Section 6.1.2 defines the baseline by removing branching features, both conditions draw on the same branch-narration pipeline; Table 1 shows that branch narrations are the error-prone content (mean 0.60 errors per item vs. 0.04 and 0.11 for navigation cues and object descriptions). If corrections were applied only to the experimental condition, the significant differences in Section 7.2 (freedom, meaningful choices, narrative presence) could be attributable to description quality rather than to branching interaction. Please state explicitly that the identical correction procedure was applied to both conditions, report the number of corrected descriptions per condition, and confirm that the baseline narration, selected as the highest-social-diversity branch, received the same treatment.
  2. [6.1.2] The baseline is described as 'simulated' pause-and-explore, but the paper does not justify that this simulation faithfully represents the cited state of the art [11]. The baseline retains the pipeline's narration, navigation cues, object descriptions, vibrations, and timeline controls, while the narration is taken from the branch with the highest social diversity score. It is unclear whether this yields the same amount, timing, and richness of description as the pause-and-explore interactions in OmniScribe, or whether the baseline is in effect a no-branching control. Because Sections 7.1-7.3 compare all outcomes against this baseline, the authors should provide a concrete mapping between baseline features and the cited system, and either demonstrate content equivalence (e.g., identical narration clips along the default path) or reinterpret the comparison as Branch Explorer vs. a stripped-down control. This clarification directly affects the strength of the 'state-of-the-art' framing in the abstract and introduction.
  3. [7.2, 7.3] The inferential reporting is incomplete for the number of tests performed. The paper reports a long series of Wilcoxon signed-rank tests (e.g., freedom, meaningful choices, interesting actions, narrative presence, enjoyment, smoothness, willingness to re-watch) and a paired t-test for comprehension, with p-values ranging from .05 to less than .01, but applies no multiple-comparison correction and reports no effect sizes. With N=12, the borderline p=.05 for comprehension (Section 7.3) and the absence of an effect on spatial presence (p=.08 and .71) should be interpreted cautiously. Please report exact p-values, effect sizes (e.g., r or Cohen's d), and the total number of planned tests, and state whether the headline agency and immersion effects are robust to a conservative correction or should be considered exploratory. This matters because the abstract states the agency result as a single unqualified conclusion.
minor comments (6)
  1. [5.1.2] The diversity rating is based on one professional describer's judgment of twenty randomly sampled branch sets; please clarify how many branch sets were sampled per video and whether the rater was blind to the research questions.
  2. [5.1.1] The Jaccard agreement rates (51.5%-61.1%) are described as 'comparable' to inter-expert agreement (57.6%); given the small number of videos (V1-V4), please report the raw counts of branch points per video and confidence intervals or a measure of uncertainty around these rates.
  3. [6.1.5] The comprehension questions are described as answerable regardless of chosen narrative path, but the questions themselves are not reported; including the full set in supplementary materials would help readers assess whether counting 'I'm not sure' as incorrect biases the comprehension comparison.
  4. [4.4.2] The branch-selection stopping criterion uses lambda = 0.75 and the option limit is five branches, but no sensitivity analysis is provided; please state whether the reported diversity ratings and user-level results are sensitive to these thresholds.
  5. [6.1.4] It is not stated whether the manual correction of descriptions occurred before or after the technical accuracy measurements in Section 5.2; if corrections were applied before the accuracy evaluation, Table 1 would overstate the raw pipeline accuracy.
  6. [Figure 11] The figure reports distributions for many subjective items, but the exact questionnaire items and their sources are not listed in the main text; several constructs such as 'freedom' and 'meaningful choices' appear to be measured with single items, and single-item measures need reliability evidence or should be reported as exploratory.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the central claims rest on an external user study and externally anchored technical checks, not on self-referential derivation.

full rationale

The paper's headline claims are outcomes of a controlled within-subject user study (N=12) in which Branch Explorer is compared to a baseline that removes branching-related features. The reported dependent measures—agency, meaningful choice, narrative presence, comprehension accuracy, summary word count, and willingness to use—are self-report or behavioral data collected from participants, not quantities computed from the system's own diversity or narration formulas. None of the equations in Sections 4.4–4.5 (diversity scores D_spa, D_sem, D_soc, D, and the greedy branch selection with λ=0.75) are defined in terms of the evaluation outcomes in Section 7. Design parameters (30-second branching interval, λ=0.75, equal diversity weights, maximum of five options) are grounded in the formative study or prior literature, not fitted to the evaluation results. The technical evaluation is externally anchored: branching timing is compared to two independent professional describers via Jaccard similarity, branch diversity is rated by an external professional describer, and description accuracy is checked against video content with two-researcher review. These are independent checks, not self-fulfilling definitions. The authors' self-citations, such as [87] for numbered navigation identifiers and [88] for the shaking gesture, are design precedents rather than load-bearing proof of the central claim; no uniqueness theorem or forced-choice argument is imported from prior work. The baseline [11] is external prior work, and the baseline condition is described as a simulation of that approach rather than being claimed to reproduce it perfectly. The manual correction of descriptions (Section 6.1.4) introduces a potential internal-validity confound if corrections were applied asymmetrically across conditions, but this is a study-design concern, not a circular derivation: the correction rule (fix errors that contradict visual content) is not defined in terms of the user-study outcomes, and those outcomes are not inputs to the correction step. Per the review rules, such concerns belong under correctness risk rather than circularity. Overall, no step in the claimed derivation reduces to its own inputs, and no load-bearing self-citation chain is present.

Assumptions & free parameters 8 free parameters · 6 assumptions · 0 invented entities

The central claim rests on a stack of tooling assumptions rather than on new mathematics. The pipeline leans on pretrained models (ATSal for saliency, BLIP-2 and Universal Sentence Encoder for semantic similarity, GPT-4o via the VideoA11y pipeline for narration, SceneDetect for shot boundaries) and on a handful of hand-set thresholds (equal diversity weights, lambda = 0.75, 30-second minimum branching interval, 30-degree clustering threshold, 120x90 viewport, 5-option cap, 0.8 RMS loudness threshold). None of these is fitted to the evaluation outcome, and several are grounded in the formative study or prior literature, which keeps the circularity burden low. The most fragile assumptions are semantic: that saliency peaks mark narratively meaningful viewing directions for BLV viewers, that caption similarity approximates semantic diversity, and that aggregated saliency approximates social interest; the authors explicitly flag the last one as a stand-in for real-user viewing data. No new entities are posited; the branching narrative concept is borrowed from interactive storytelling.

free parameters (8)
  • lambda (diversity stop threshold) = 0.75
    Section 4.4.2; branches are added until adding one drops the diversity score below 0.75 times the previous value; hand-set without ablation.
  • diversity weights w_spa, w_sem, w_soc = 1/3 each
    Section 4.4.1; equal weights chosen by hand, stated as personalizable but not varied in the evaluation.
  • minimum branching interval = 30 seconds
    Sections 3.2.2 and 4.3.1; all formative-study participants preferred 30s over 15s/45s; adopted as the system minimum.
  • viewport size = 120 degrees horizontal x 90 degrees vertical
    Section 4.3.2; taken from prior streaming literature [43]; defines what content is captioned and compared.
  • maximum branch options per branching point = 5
    Section 4.4.2; from Hick's law and Miller's 7 +/- 2 references [24, 55]; caps the choice set.
  • loud-music RMS threshold = 0.8
    Section 4.3.1; one-second sliding window flags loud regions as unsuitable for branching; hand-set.
  • clustering merge threshold = 30 degrees
    Section 4.3.2; agglomerative centroid linkage merges viewing directions within 30 degrees; hand-set.
  • branching agreement window = 5 seconds
    Section 5.1.1; branching points within 5 seconds of each other counted as equivalent when computing Jaccard agreement with experts; affects the reported agreement rates.
assumptions (6)
  • domain assumption ATSal saliency peaks identify viewing directions that BLV users will find meaningful.
    Section 4.3.2 selects branch directions from ATSal saliency centroids; ATSal was trained on sighted viewers' gaze and is never validated against BLV exploration preferences.
  • domain assumption BLIP-2 caption similarity approximates the semantic diversity of branch content.
    Section 4.4.1 uses caption cosine similarity (BLIP-2 plus Universal Sentence Encoder) as the semantic diversity metric.
  • domain assumption Aggregated saliency within a viewport approximates social interest.
    Section 4.4.1 computes social diversity from ATSal saliency sums; the paper itself says real-user viewing data would be more accurate 'in the future'.
  • domain assumption GPT-4o via the VideoA11y pipeline produces narrations that are coherent and accurate enough for the studied experience.
    Section 4.5.1; measured at 78% accuracy overall in Section 5.2.2 and lower (71-77%) in dynamic scenes, so the experience depends on an error-prone generator.
  • domain assumption SceneDetect HSV boundaries align with narrative scene transitions.
    Section 4.3.1 uses SceneDetect shot boundaries as branching-point candidates; agreement with professional describers is only 51.5-61.1% (Jaccard), comparable to inter-expert agreement.
  • standard math Standard clustering and linking procedures behave as intended on spherical coordinates.
    Section 4.3.2 uses agglomerative clustering with a 30-degree threshold and angular-difference minimization; this is background method, not the core risk.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Branch Explorer: Leveraging Branching Narratives to Support Interactive 360{\deg} Video Viewing for Blind and Low Vision Users." pith.science (2026). https://pith.science/paper/RD6IKHFE

@misc{pith2026250709959,
  author       = {Pith},
  title        = {Pith review of: Branch Explorer: Leveraging Branching Narratives to Support Interactive 360\deg Video Viewing for Blind and Low Vision Users},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/RD6IKHFE}},
  note         = {Machine review of arXiv:2507.09959}
}
read the original abstract

360{\deg} videos enable users to freely choose their viewing paths, but blind and low vision (BLV) users are often excluded from this interactive experience. To bridge this gap, we present Branch Explorer, a system that transforms 360{\deg} videos into branching narratives -- stories that dynamically unfold based on viewer choices -- to support interactive viewing for BLV audiences. Our formative study identified three key considerations for accessible branching narratives: providing diverse branch options, ensuring coherent story progression, and enabling immersive navigation among branches. To address these needs, Branch Explorer employs a multi-modal machine learning pipeline to generate diverse narrative paths, allowing users to flexibly make choices at detected branching points and seamlessly engage with each storyline through immersive audio guidance. Evaluation with 12 BLV viewers showed that Branch Explorer significantly enhanced user agency and engagement in 360{\deg} video viewing. Users also developed personalized strategies for exploring 360{\deg} content. We further highlight implications for supporting accessible exploration of videos and virtual environments.

Figures

Figures reproduced from arXiv: 2507.09959 by the authors.

Figure 1
Figure 1. Branch Explorer transforms 360° videos into branching narratives—stories that dynamically unfold based on viewer choices—to create an engaging experience for blind and low vision (BLV) users. It employs a multi-modal machine learning pipeline to generate diverse narrative paths, enabling BLV users to make choices at key branching points and explore each story￾line through immersive audio guidance. The figure shows t… view at source ↗
Figure 2
Figure 2. Comparison of the pause-and-explore approach, linear audio descriptions, and branching narratives for 360° videos. community has investigated new methods to make 360° videos accessible while preserving their interactive nature [11, 30]. Several studies have examined BLV viewers’ preferences for describing 360° videos, revealing the need for interactive descrip￾tions. For example, Fidyka et al. [18–20] reported that … view at source ↗
Figure 3
Figure 3. The design probe includes a branch option menu, [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (9 more)
Figure 4
Figure 4. Figure 4: Three key design considerations for accessible branching narratives in 360 [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: An example walk-through of Branch Explorer: (A) The system narrates key visual elements in the second person [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 6
Figure 6. Figure 6: The pipeline of Branch Explorer. The first module employs an optimization approach to curate branches with high [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 7
Figure 7. Figure 7: (A) Branch Explorer generates branches by linking [PITH_FULL_IMAGE:figures/full_fig_p007_7.png]
Figure 8
Figure 8. Figure 8: The system generates scene titles, branch titles, and branch narrations to ensure narrative coherence. This figure illustrates their hierarchy, generation strategy, and display timing (i.e., when they are provided to users) [PITH_FULL_IMAGE:figures/full_fig_p008_8.png]
Figure 9
Figure 9. Figure 9: Branch Explorer provides three exploration fea [PITH_FULL_IMAGE:figures/full_fig_p009_9.png]
Figure 10
Figure 10. Figure 10: The four videos used in the evaluation study. [PITH_FULL_IMAGE:figures/full_fig_p011_10.png]
Figure 11
Figure 11. Figure 11: Distributions of participant ratings for both systems (1=strongly negative, 7=strongly positive). Asterisks denote [PITH_FULL_IMAGE:figures/full_fig_p012_11.png]
Figure 12
Figure 12. Figure 12: Video comprehension metrics from the evaluation study. These values were computed as averages per participant [PITH_FULL_IMAGE:figures/full_fig_p013_12.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sonic Stage: Auto-Generating Interactive Spatial Soundscapes to Facilitate Dialogue Video Comprehension for Blind Viewers

    cs.HC 2026-07 conditional novelty 6.0 of 10

    A system that spatializes dialogue, adds diegetic sound, and offers tap-to-hear descriptions improved blind viewers' understanding of character position, movement, actions, and visual details in dialogue-only scenes.

Reference graph

Works this paper leans on

100 extracted references · 44 canonical work pages · cited by 1 Pith paper

  1. [11]

    Ruei-Che Chang, Chao-Hsien Ting, Chia-Sheng Hung, Wan-Chen Lee, Liang- Jin Chen, Yu-Tzu Chao, Bing-Yu Chen, and Anhong Guo. 2022. OmniScribe: Authoring Immersive Audio Descriptions for 360° Videos. In Proceedings of the 35th Annual ACM Symposium on User Interface Software and Technology (Bend, OR, USA) (UIST ’22). Association for Computing Machinery, New ...

  2. [1]

    Erik M Altmann and J Gregory Trafton. 2004. Task interruption: Resumption lag and the role of cues. In Proceedings of the Annual Meeting of the Cognitive Science Society, Vol. 26

  3. [2]

    Daniel Andrews and Chris Baber. 2014. Visualizing interactive narratives: em- ploying a branching comic to tell a story and show its readings. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Toronto, Ontario, Canada) (CHI ’14). Association for Computing Machinery, New York, NY, USA, 1895–1904. doi:10.1145/2556288.2557296

  4. [3]

    Alan Baddeley. 2020. Working memory. Memory (2020), 71–111

  5. [4]

    Youtube Official Blog. 2017. Hot and Cold: Heatmaps in VR. https://blog.youtube/ news-and-events/hot-and-cold-heatmaps-in-vr/ Accessed: 2025-01-20

  6. [5]

    Rick Busselle and Helena Bilandzic. 2009. Measuring narrative engagement. Media psychology 12, 4 (2009), 321–347

  7. [6]

    Janelynn Camingue, Elin Carstensdottir, and Edward F. Melcer. 2021. What is a Visual Novel? Proc. ACM Hum.-Comput. Interact. 5, CHI PLAY, Article 285 (Oct. 2021), 18 pages. doi:10.1145/3474712

  8. [7]

    Daniel Cer, Yinfei Yang, Sheng-yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar, et al

Show all 100 references
  1. [8]

    Seunghoon Cha, Jungjin Lee, Seunghwa Jeong, Younghui Kim, and Junyong Noh

  2. [9]

    Ruei-Che Chang, Chia-Sheng Hung, Bing-Yu Chen, Dhruv Jain, and Anhong Guo

  3. [12]

    The Guide Has Your Back

    Jazmin Collins, Crescentia Jung, Yeonju Jang, Danielle Montour, Andrea Steven- son Won, and Shiri Azenkot. 2023. “The Guide Has Your Back”: Exploring How Sighted Guides Can Enhance Accessibility in Social Virtual Reality for Blind and Low Vision People. In Proceedings of the 2...

  4. [13]

    Valve Corporation. 2011. Official Portal 2 Website. https://www.thinkwithportals. com/ Accessed: 2025-07-12

  5. [14]

    Yasser Dahou, Marouane Tliba, Kevin McGuinness, and Noel O’Connor. 2021. Atsal: an attention based architecture for saliency prediction in 360◦ videos. In International Conference on Pattern Recognition . Springer, 305–320

  6. [15]

    Khang Dang and Sooyeon Lee. 2024. Musical Performances in Virtual Reality with Spatial and View-Dependent Audio Descriptions for Blind and Low-Vision Users. In The 26th International ACM SIGACCESS Conference on Computers and Accessibility. 1–5

  7. [16]

    Ahmed Elmezeny, Nina Edenhofer, and Jeffrey Wimmer. 2018. Immersive sto- rytelling in 360-degree videos: An analysis of interplay between narrative and technical immersion. Journal For Virtual Worlds Research 11, 1 (2018)

  8. [17]

    Equal Entry. 2020. Audio Descriptions for 360 Degree Video: Best Practices. https://www.youtube.com/watch?v=jOX6gxUZq8w Accessed: 2025-01-24

  9. [18]

    Anita Fidyka and Anna Matamala. 2018. Audio description in 360º videos: Results from focus groups in Barcelona and Kraków. Translation Spaces 7, 2 (2018), 285– 303

  10. [19]

    Anita Fidyka and Anna Matamala. 2021. Retelling narrative in 360 videos: Impli- cations for audio description. Translation Studies 14, 3 (2021), 298–312

  11. [20]

    Anita Fidyka, Anna Matamala, Olga Soler Vilageliu, and Blanca Arias-Badia. 2021. Audio description in 360 content: results from a reception study. Skase Journal of Translation and Interpretation 14, 1 (2021), 14–32

  12. [21]

    Jan Gugenheimer, Dennis Wolf, Gabriel Haas, Sebastian Krebs, and Enrico Rukzio

  13. [22]

    Tilo Hartmann, Werner Wirth, Holger Schramm, Christoph Klimmt, Peter Vorderer, André Gysbers, Saskia Böcking, Niklas Ravaja, Jari Laarni, Timo Saari, et al. 2015. The spatial presence experience scale (SPES). Journal of Media Psychology (2015)

  14. [23]

    Shuting He, Hao Luo, Pichao Wang, Fan Wang, Hao Li, and Wei Jiang. 2021. Tran- sreid: Transformer-based object re-identification. In Proceedings of the IEEE/CVF international conference on computer vision . 15013–15022

  15. [24]

    William E Hick. 1952. On the rate of gain of information. Quarterly Journal of experimental psychology 4, 1 (1952), 11–26

  16. [25]

    Steven CH Hoi, Doyen Sahoo, Jing Lu, and Peilin Zhao. 2021. Online learning: A comprehensive survey. Neurocomputing 459 (2021), 249–289

  17. [26]

    Business Research Insights. 2024. VR and 360◦ Video Market Size, Share, Growth, and Industry Analysis. https://www.businessresearchinsights.com/market- reports/vr-and-360-video-market-118394 Accessed: 2025-02-24

  18. [27]

    Gaurav Jain, Basel Hindi, Connor Courtien, Xin Yi Therese Xu, Conrad Wyrick, Michael Malcolm, and Brian A. Smith. 2023. Front Row: Automatically Generating Immersive Audio Representations of Tennis Broadcasts for Blind Viewers. In Proceedings of the 36th Annual ACM Symposium o...

  19. [28]

    Chandrika Jayant, Hanjie Ji, Samuel White, and Jeffrey P Bigham. 2011. Sup- porting blind photography. In The proceedings of the 13th international ACM SIGACCESS conference on Computers and accessibility . 203–210

  20. [29]

    Lucy Jiang, Crescentia Jung, Mahika Phutane, Abigale Stangl, and Shiri Azenkot

  21. [30]

    Lucy Jiang, Mahika Phutane, and Shiri Azenkot. 2023. Beyond Audio Description: Exploring 360° Video Accessibility with Blind and Low Vision Users Through Collaborative Creation. In Proceedings of the 25th International ACM SIGACCESS Conference on Computers and Accessibility (N...

  22. [31]

    Yue Jiang, Zixin Guo, Hamed Rezazadegan Tavakoli, Luis A Leiva, and Antti Oulasvirta. 2024. EyeFormer: predicting personalized scanpaths with transformer- guided reinforcement learning. InProceedings of the 37th Annual ACM Symposium on User Interface Software and Technology . 1–15

  23. [33]

    Yili Jin, Junhua Liu, Fangxin Wang, and Shuguang Cui. 2022. Where are you looking? A large-scale dataset of head and gaze behavior for 360-degree videos and a pilot study. In Proceedings of the 30th ACM International Conference on Multimedia. 1025–1034

  24. [34]

    It’s Kind of Context Dependent

    “It’s Kind of Context Dependent”: Understanding Blind and Low Vision People’s Video Accessibility Preferences Across Viewing Scenarios. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24). Association for Computing Machi...

  25. [35]

    Daniel Kepplinger, Günter Wallner, Simone Kriglstein, and Michael Lankes. 2020. See, Feel, Move: player behaviour analysis through combined visualization of gaze, emotions, and movement. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems . 1–14

  26. [36]

    Daniel Killough and Amy Pavel. 2023. Exploring Community-Driven Descriptions for Making Livestreams Accessible. In Proceedings of the 25th International ACM SIGACCESS Conference on Computers and Accessibility . 1–13

  27. [37]

    Anatole Lécuyer, Pascal Mobuchon, Christine Mégard, Jérôme Perret, Claude Andriot, and J-P Colinot. 2003. HOMERE: a multimodal system for visually impaired people to explore virtual environments. In IEEE Virtual Reality, 2003. Proceedings. IEEE, 251–258

  28. [38]

    Chaeeun Lee, Jinwook Kim, Hyeonbeom Yi, and Woohun Lee. 2024. Viewer2Explorer: Designing a Map Interface for Spatial Navigation in Linear 360 Museum Exhibition Video. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24)....

  29. [39]

    Kyoungkook Kang and Sunghyun Cho. 2019. Interactive and automatic navigation for 360° video playback. ACM Trans. Graph. 38, 4, Article 108 (July 2019), 11 pages. doi:10.1145/3306346.3323046

  30. [40]

    Min Seok Lee, WooSeok Shin, and Sung Won Han. 2022. Tracer: Extreme attention guided salient object tracing network (student abstract). In Proceedings of the AAAI conference on artificial intelligence , Vol. 36. 12993–12994

  31. [41]

    James R Lewis. 2018. The system usability scale: past, present, and future. Inter- national Journal of Human–Computer Interaction 34, 7 (2018), 577–590

  32. [42]

    Chaoyu Li, Sid Padmanabhuni, Maryam Cheema, Hasti Seifi, and Pooyan Fazli

  33. [43]

    Chenge Li, Weixi Zhang, Yong Liu, and Yao Wang. 2019. Very long term field of view prediction for 360-degree video streaming. In 2019 IEEE Conference on Multimedia Information Processing and Retrieval (MIPR) . IEEE, 297–302

  34. [44]

    Jaewook Lee, Yi-Hao Peng, Jaylin Herskovitz, and Anhong Guo. 2021. Image Explorer: Multi-Layered Touch Exploration to Make Images Accessible. In Pro- ceedings of the 23rd International ACM SIGACCESS Conference on Computers and Accessibility (Virtual Event, USA) (ASSETS ’21). A...

  35. [45]

    Yen-Chen Lin, Yung-Ju Chang, Hou-Ning Hu, Hsien-Tzu Cheng, Chi-Wen Huang, and Min Sun. 2017. Tell Me Where to Look: Investigating Ways for Assisting Focus in 360° Video. In Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems (Denver, Colorado, USA)(CHI...

  36. [47]

    Guanhong Liu, Tianyu Yu, Chun Yu, Haiqing Xu, Shuchang Xu, Ciyuan Yang, Feng Wang, Haipeng Mi, and Yuanchun Shi. 2021. Tactile Compass: Enabling Visually Impaired People to Follow a Path with Continuous Directional Feedback. In Proceedings of the 2021 CHI Conference on Human F...

  37. [48]

    Hanchao Liu, Wenyuan Xue, Yifei Chen, Dapeng Chen, Xiutian Zhao, Ke Wang, Liping Hou, Rongjun Li, and Wei Peng. 2024. A survey on hallucination in large vision-language models. arXiv preprint arXiv:2402.00253 (2024)

  38. [49]

    Kun Liu, Mengxue Qu, Yang Liu, Yunchao Wei, Wenming Zhe, Yao Zhao, and Wu Liu. 2024. Single-frame supervision for spatio-temporal video grounding. IEEE Transactions on Pattern Analysis and Machine Intelligence (2024)

  39. [50]

    Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. BLIP-2: Boot- strapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models. In Proceedings of the 40th International Conference on Ma- chine Learning (Proceedings of Machine Learning Res...

  40. [51]

    Liu, Rorik Henrikson, Tovi Grossman, Michael Glueck, and Mark Parent

    Sean J. Liu, Rorik Henrikson, Tovi Grossman, Michael Glueck, and Mark Parent

  41. [52]

    Xingyu" Bruce" Liu, Ruolin Wang, Dingzeyu Li, Xiang Anthony Chen, and Amy Pavel. 2022. CrossA11y: Identifying Video Accessibility Issues via Cross-modal Grounding. In Proceedings of the 35th Annual ACM Symposium on User Interface Software and Technology. 1–14

  42. [53]

    Sandy Louchart and Ruth Aylett. 2003. Solving the narrative paradox in VEs– lessons from RPGs. InInternational workshop on intelligent virtual agents. Springer, 244–248

  43. [54]

    Keenan R May, Brianna J Tomlinson, Xiaomeng Ma, Phillip Roberts, and Bruce N Walker. 2020. Spotlights and soundscapes: On the design of mixed reality auditory environments for persons with visual impairment.ACM Transactions on Accessible Computing (TACCESS) 13, 2 (2020), 1–47

  44. [55]

    George A Miller. 1956. The magical number seven, plus or minus two: Some limits on our capacity for processing information. Psychological review 63, 2 (1956), 81

  45. [56]

    Liu, Maneesh Agrawala, Stephen DiVerdi, and Aaron Hertzmann

    Sean J. Liu, Maneesh Agrawala, Stephen DiVerdi, and Aaron Hertzmann. 2019. View-Dependent Video Textures for 360° Video. In Proceedings of the 32nd Annual ACM Symposium on User Interface Software and Technology (New Orleans, LA, USA) (UIST ’19). Association for Computing Machi...

  46. [57]

    Christopher Moser and Xiaowen Fang. 2015. Narrative structure and player experience in role-playing games. International Journal of Human-Computer Interaction 31, 2 (2015), 146–156

  47. [58]

    Fionn Murtagh and Pedro Contreras. 2012. Algorithms for hierarchical cluster- ing: an overview. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery 2, 1 (2012), 86–97

  48. [59]

    Vishnu Nair, Jay L Karp, Samuel Silverman, Mohar Kalra, Hollis Lehv, Faizan Jamil, and Brian A. Smith. 2021. NavStick: Making Video Games Blind-Accessible via the Ability to Look Around. In The 34th Annual ACM Symposium on User UIST ’25, September 28-October 1, 2025, Busan, Re...

  49. [60]

    Vishnu Nair, Hanxiu ’Hazel’ Zhu, and Brian A. Smith. 2023. ImageAssist: Tools for Enhancing Touchscreen-Based Image Exploration Systems for Blind and Low Vision Users. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23)....

  50. [61]

    Vishnu Nair, Hanxiu ’Hazel’ Zhu, Peize Song, Jizhong Wang, and Brian A. Smith

  51. [62]

    Rosiana Natalie, Ruei-Che Chang, Smitha Sheshadri, Anhong Guo, and Kotaro Hara. 2024. Audio Description Customization. In Proceedings of the 26th Inter- national ACM SIGACCESS Conference on Computers and Accessibility (St. John’s, NL, Canada) (ASSETS ’24). Association for Comp...

  52. [63]

    Motor Ability

    Claire L Mitchell and Jacob O Wobbrock. 2024. Characterizing" Motor Ability" for Ability-Based Design. In Proceedings of the 26th International ACM SIGACCESS Conference on Computers and Accessibility . 1–15

  53. [64]

    Zheng Ning, Zheng Zhang, Jerrick Ban, Kaiwen Jiang, Ruohong Gan, Yapeng Tian, and Toby Jia-Jun Li. 2024. MIMOSA: Human-AI Co-Creation of Computational Spatial Audio Effects on Videos. InProceedings of the 16th Conference on Creativity & Cognition (Chicago, IL, USA) (C&C ’24). ...

  54. [65]

    Amy Pavel, Björn Hartmann, and Maneesh Agrawala. 2017. Shot Orientation Controls for Interactive Cinematography with 360 Video. In Proceedings of the 30th Annual ACM Symposium on User Interface Software and Technology (Québec City, QC, Canada) (UIST ’17). Association for Compu...

  55. [66]

    Amy Pavel, Gabriel Reyes, and Jeffrey P. Bigham. 2020. Rescribe: Authoring and Automatically Editing Audio Descriptions. In Proceedings of the 33rd Annual ACM Symposium on User Interface Software and Technology (Virtual Event, USA) (UIST ’20). Association for Computing Machine...

  56. [67]

    Pope, Robert Dawes, Florian Schweiger, and Alia Sheikh

    Vanessa C. Pope, Robert Dawes, Florian Schweiger, and Alia Sheikh. 2017. The Geometry of Storytelling: Theatrical Use of Space for 360-degree Videos and Virtual Reality. In Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems (Denver, Colorado, USA)(CHI...

  57. [68]

    Sima Rahimizhian, Ali Ozturen, and Mustafa Ilkan. 2020. Emerging realm of 360-degree technology to promote tourism destination. Technology in Society 63 (2020), 101411

  58. [70]

    Sylvia Rothe and Heinrich Hußmann. 2018. Guiding the viewer in cinematic vir- tual reality by diegetic cues. In Augmented Reality, Virtual Reality, and Computer Graphics: 5th International Conference, A VR 2018, Otranto, Italy, June 24–27, 2018, Proceedings, Part I 5 . Springe...

  59. [71]

    Zheng Ning, Brianna L Wimer, Kaiwen Jiang, Keyi Chen, Jerrick Ban, Yapeng Tian, Yuhang Zhao, and Toby Jia-Jun Li. 2024. SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision Viewers. In Proceedings of the 2024 CHI Conference ...

  60. [72]

    Alexa F Siu, Mike Sinclair, Robert Kovacs, Eyal Ofek, Christian Holz, and Edward Cutrell. 2020. Virtual reality without vision: A haptic and auditory white cane to navigate complex virtual worlds. In Proceedings of the 2020 CHI conference on human factors in computing systems . 1–13

  61. [73]

    Filip Škola, Selma Rizvić, Marco Cozza, Loris Barbieri, Fabio Bruno, Dimitrios Skarlatos, and Fotis Liarokapis. 2020. Virtual reality with 360-video storytelling in cultural heritage: Study of presence, engagement, and immersion. Sensors 20, 20 (2020), 5851

  62. [74]

    Chareen Snelson and Yu-Chang Hsu. 2020. Educational 360-degree videos in virtual reality: A scoping review of the emerging research. TechTrends 64, 3 (2020), 404–412

  63. [75]

    Abigale Stangl, Shasta Ihorn, Yue-Ting Siu, Aditya Bodi, Mar Castanon, Lothar D Narins, and Ilmi Yoon. 2023. The potential of a visual dialogue agent in a tan- dem automated audio description system for videos. In Proceedings of the 25th International ACM SIGACCESS Conference ...

  64. [76]

    Sophia C Steinhaeusser, Sebastian Oberdörfer, Sebastian von Mammen, Marc Erich Latoschik, and Birgit Lugrin. 2022. Joyful adventures and fright- ening places–designing emotion-inducing virtual environments. Frontiers in Virtual Reality 3 (2022), 919163

  65. [77]

    Anyi Rao, Linning Xu, and Dahua Lin. 2022. Shoot360: Normal View Video Creation from City Panorama Footage. In ACM SIGGRAPH 2022 Conference Pro- ceedings (Vancouver, BC, Canada) (SIGGRAPH ’22). Association for Computing Machinery, New York, NY, USA, Article 13, 9 pages. doi:10...

  66. [78]

    Yu-Chuan Su, Dinesh Jayaraman, and Kristen Grauman. 2016. Pano2vid: Auto- matic cinematography for watching 360 videos. In Asian Conference on Computer Vision. Springer, 154–171

  67. [79]

    Richard M Ryan, C Scott Rigby, and Andrew Przybylski. 2006. The motivational pull of video games: A self-determination theory approach. Motivation and emotion 30 (2006), 344–360

  68. [80]

    Tess Van Daele, Akhil Iyer, Yuning Zhang, Jalyn C Derry, Mina Huh, and Amy Pavel. 2024. Making Short-Form Videos Accessible with Hierarchical Video Summaries. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–17

  69. [81]

    Sarah Van Der Land, Alexander P Schouten, Frans Feldberg, Bart van Den Hooff, and Marleen Huysman. 2013. Lost in space? Cognitive fit and cognitive load in 3D virtual environments. Computers in Human Behavior 29, 3 (2013), 1054–1064

  70. [82]

    Miao Wang, Yi-Jun Li, Wen-Xuan Zhang, Christian Richardt, and Shi-Min Hu

  71. [83]

    Yujia Wang, Wei Liang, Haikun Huang, Yongqi Zhang, Dingzeyu Li, and Lap-Fai Yu. 2021. Toward Automatic Audio Description Generation for Accessible Videos. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (Yokohama, Japan) (CHI ’21). Association f...

  72. [84]

    Spencer Whitehead, Heng Ji, Mohit Bansal, Shih-Fu Chang, and Clare Voss

  73. [85]

    Yu-Chuan Su and Kristen Grauman. 2017. Making 360 ° Video Watchable in 2D: Learning Videography for Click Free Viewing. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 1368–1376. doi:10.1109/CVPR. 2017.150

  74. [86]

    Chenglei Wu, Zhihao Tan, Zhi Wang, and Shiqiang Yang. 2017. A dataset for exploring user behaviors in VR spherical video streaming. In Proceedings of the 8th ACM on Multimedia Systems Conference . 193–198

  75. [87]

    Anh Truong, Sara Chen, Ersin Yumer, David Salesin, and Wilmot Li. 2018. Ex- tracting regular fov shots from 360 event footage. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems . 1–11

  76. [88]

    Shuchang Xu, Xiaofu Jin, Huamin Qu, and Yukang Yan. 2025. DanmuA11y: Making Time-Synced On-Screen Video Comments (Danmu) Accessible to Blind and Low Vision Users via Multi-Viewer Audio Discussions. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems...

  77. [89]

    Shuchang Xu, Ciyuan Yang, Wenhao Ge, Chun Yu, and Yuanchun Shi. 2020. Virtual Paving: Rendering a Smooth Path for People with Visual Impairment through Vibrotactile and Audio Feedback. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 4, 3, Article 99 (Sept. 2020), 25 page...

  78. [90]

    Ciyuan Yang, Shuchang Xu, Tianyu Yu, Guanhong Liu, Chun Yu, and Yuanchun Shi. 2021. LightGuide: Directing Visually Impaired People along a Path Using Light Cues. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 5, 2, Article 84 (June 2021), 27 pages. doi:10.1145/3463524

  79. [91]

    In 2020 IEEE International Symposium on Mixed and Augmented Reality (ISMAR)

    Transitioning360: Content-aware NFoV Virtual Camera Paths for 360 ° Video Playback. In 2020 IEEE International Symposium on Mixed and Augmented Reality (ISMAR). 185–194. doi:10.1109/ISMAR50242.2020.00040

  80. [92]

    Yifu Zhang, Peize Sun, Yi Jiang, Dongdong Yu, Fucheng Weng, Zehuan Yuan, Ping Luo, Wenyu Liu, and Xinggang Wang. 2022. Bytetrack: Multi-object tracking by associating every detection box. In European conference on computer vision . Springer, 1–21

  81. [93]

    Bennett, Hrvoje Benko, Edward Cutrell, Christian Holz, Meredith Ringel Morris, and Mike Sinclair

    Yuhang Zhao, Cynthia L. Bennett, Hrvoje Benko, Edward Cutrell, Christian Holz, Meredith Ringel Morris, and Mike Sinclair. 2018. Enabling People with Visual Impairments to Navigate Virtual Reality with a Haptic and Auditory Cane Simulation. In Proceedings of the 2018 CHI Confer...

  82. [94]

    In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing

    Incorporating background knowledge into video description generation. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 3992–4001

  83. [95]

    Frank Wilcoxon, S Katti, Roberta A Wilcox, et al . 1970. Critical values and probability levels for the Wilcoxon rank sum test and the Wilcoxon signed rank test. Selected tables in mathematical statistics 1 (1970), 171–259

  84. [97]

    Shuchang Xu, Chang Chen, Zichen Liu, Xiaofu Jin, Lin-Ping Yuan, Yukang Yan, and Huamin Qu. 2024. Memory Reviver: Supporting Photo-Collection Reminis- cence for People with Visual Impairment via a Proactive Chatbot. In Proceedings of the 37th Annual ACM Symposium on User Interf...

  85. [101]

    Zhengyuan Yang, Linjie Li, Kevin Lin, Jianfeng Wang, Chung-Ching Lin, Zicheng Liu, and Lijuan Wang. 2023. The dawn of lmms: Preliminary explorations with gpt-4v (ision). arXiv preprint arXiv:2309.17421 9, 1 (2023), 1

  86. [104]

    Yuhang Zhao, Edward Cutrell, Christian Holz, Meredith Ringel Morris, Eyal Ofek, and Andrew D. Wilson. 2019. SeeingVR: A Set of Tools to Make Virtual Reality More Accessible to People with Low Vision. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Syste...

  87. [2016]

    SwiVRChair: A Motorized Swivel Chair to Nudge Users’ Orientation for 360 Degree Storytelling in Virtual Reality. In Proceedings of the 2016 CHI Branch Explorer: Interactive 360 ° Video Viewing for Blind and Low Vision Users UIST ’25, September 28-October 1, 2025, Busan, Republ...

  88. [2018]

    arXiv preprint arXiv:1803.11175 (2018)

    Universal sentence encoder. arXiv preprint arXiv:1803.11175 (2018)

  89. [2020]

    ACM Trans

    Enhanced Interactive 360 ° Viewing via Automatic Guidance. ACM Trans. Graph. 39, 5, Article 154 (May 2020), 15 pages. doi:10.1145/3183794

  90. [2023]

    In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (San Francisco, CA, USA) (UIST ’23)

    RadarVR: Exploring Spatiotemporal Visual Guidance in Cinematic VR. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (San Francisco, CA, USA) (UIST ’23). Association for Computing Machinery, New York, NY, USA, Article 86, 14 pages. doi:1...

  91. [2024]

    In Proceedings of the 2024 ACM Designing Interactive Systems Confer- ence (Copenhagen, Denmark) (DIS ’24)

    SoundShift: Exploring Sound Manipulations for Accessible Mixed-Reality Awareness. In Proceedings of the 2024 ACM Designing Interactive Systems Confer- ence (Copenhagen, Denmark) (DIS ’24). Association for Computing Machinery, New York, NY, USA, 116–132. doi:10.1145/3643834.3661556

  92. [2025]

    In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems

    VideoA11y: Method and Dataset for Accessible Video Description. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems . 1–29

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.