{"id":"b748162f-d8e9-4470-a111-d08101587469","arxiv_id":"2501.15711","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":9,"one_line_summary":"DanmuA11y converts visual Danmu comments into curated, spatial-audio multi-viewer discussions, improving comprehension and social presence for blind and low vision viewers in a 12-person study.","lead":"The paper presents DanmuA11y, a system that reads time-synced video comments (Danmu) aloud as a multi-viewer audio discussion for blind and low vision viewers. In a study with 12 BLV participants, users understood comments better, watched more smoothly, and felt more social connection than with a baseline tool.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Nine of 36 baseline trials were abandoned due to frustration, yet the analysis reports rates and tests as if all trials were complete; this attrition bias is the most load-bearing threat to the claimed comparison.","rationale":"The reader's weakest assumption was that GPT-4o outputs would generalize across videos, with only a 92.5% spot check. That is a legitimate external-validity concern, but it affects generalization rather than the internal validity of the controlled comparison. A more load-bearing threat is the differential attrition in the baseline condition: nine of 36 baseline trials were abandoned because participants found the baseline frustrating, and the analysis does not account for this. Reported confusion counts are divided by 36 even though nine baseline trials were incomplete, which suggests the missing trials may have been counted as zero-confusion trials or excluded from numerators while kept in denominators; either way, the baseline's failure cases are systematically underrepresented. The same issue can influence the overall subjective ratings, because participants who abandoned baseline videos due to frustration may anchor their post-study questionnaires on that negative experience, making DanmuA11y look better than a fair comparison would. This concern is concrete, tied to the paper's own Table 5 and Section 6, and testable with a sensitivity analysis. I therefore agree with the CONDITIONAL verdict, but for a different primary reason than the reader emphasized; the reader did mention incomplete trials in the rationale, so this is partial agreement rather than disagreement. The paper's contribution remains valuable, and the GPT-4o reliance is a genuine secondary concern, but the attrition bias should be the first checkpoint before accepting the empirical claims as robust.","tokens_in":30627,"tokens_out":2910,"duration_ms":28980,"concrete_test":"Recompute every primary comparison after assigning the nine incomplete baseline trials the worst observed baseline outcome on each metric (e.g., maximum confusion count per video, most negative Likert rating, and a failed video summary), preserving the paired structure per participant. If any of the reported p-values (confusion instances, coherence, unobtrusiveness, social presence, closeness) crosses 0.05, or if the confusion-rate gap drops below the reported 1.44 vs. 0.25 per video, the headline benefits cannot be separated from attrition bias. Also report a complete-case analysis using only the five participants who finished all baseline trials, and state whether the paired Wilcoxon results survive on that subset.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 6 states that all 36 DanmuA11y trials were completed, but seven participants did not complete nine baseline trials because the baseline was frustrating. The paper then reports confusion rates and Wilcoxon tests without any sensitivity analysis or imputation for these missing trials. Abandoned baseline trials are unlikely to be missing at random: they represent the worst baseline experiences, so removing them (or implicitly treating them as non-events in rate calculations) inflates DanmuA11y's apparent advantage on comprehension, viewing smoothness, and social connection. The central claim depends on a paired comparison, but the pairing is incomplete for 25% of baseline trials, and the reported denominator of 36 for confusion instances suggests the incomplete trials may have been counted as zero-confusion observations, which would directly bias the p-values. Because the paper's headline results are all statistically significant at p<.01, this concern is not merely a statistical technicality: it concerns whether the effect is real or an artifact of differential attrition between the two arms.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents DanmuA11y, a system that makes time-synced on-screen video comments (Danmu) accessible to blind and low vision (BLV) users by converting them into multi-viewer audio discussions. A formative study with eight BLV participants identified three challenges: missing visual context, speech interference between comments and video, and disorganized sequential access. The system addresses these with AI-generated visual descriptions, an optimization algorithm that places curated comment topics into non-speech segments and speech breaks, and spatial-audio rendering of multi-viewer dialogues. A within-subject evaluation with twelve BLV participants compared DanmuA11y against a baseline simulating current practice (an auto-scrolling list). The authors report significantly fewer confusion instances, smoother viewing, and higher social presence with DanmuA11y, alongside positive usability ratings. The paper contributes a novel accessible-Danmu system, a design rationale derived from a formative study, and an evaluation with the target user population.","tokens_in":30788,"tokens_out":5721,"duration_ms":57917,"significance":"If the reported results hold, this is a valuable contribution to accessibility research and to the growing literature on non-visual access to social video platforms. The work combines LLM-based content curation with spatial audio in a way that is well grounded in the target users' needs, and the evaluation directly involves BLV participants, which is a clear strength. The paper also provides detailed implementation information and a spot-checked accuracy of 92.5% for visual descriptions, supporting reproducibility. The main reservation is that the statistical comparison is clouded by incomplete baseline trials and an unstated treatment of missing data; this must be resolved before the central claims can be fully trusted.","major_comments":[{"comment":"Nine of 36 baseline trials were abandoned due to frustration, yet the paper reports confusion rates as '1.44 times per video' using a denominator of 36 (52 total instances). This treats incomplete trials as full exposure events and is inconsistent with the per-participant-per-video averaging shown in Figure 11. The manuscript does not state how the Wilcoxon signed-rank tests handled these missing values, nor how the video-summary word counts and factual-error counts were computed for participants who did not complete all baseline videos. Because attrition was caused by frustration with the baseline, the missing data are unlikely to be missing at random. The direction of any bias is not obvious: abandoned trials may have truncated the number of confusion reports (making the baseline look better) or may represent the most frustrated participants (making the baseline look worse). Please report a sensitivity analysis, e.g., analysis on complete pairs only, and a worst-case imputation for abandoned trials, and state explicitly how the denominators and paired tests were constructed.","section":"Section 6 and Table 5"},{"comment":"The paired t-test for summary length is reported with df=11 (i.e., 12 participants), but if nine baseline trials were abandoned, then nine baseline video summaries may be missing. The paper does not explain how missing summaries were handled: listwise deletion, imputation, or treating abandoned trials as zero-length summaries. As written, the reported t-test and the factual-error comparison (Z=-2.27, p<.05) cannot be verified. This is load-bearing because the comprehension claim depends on these pairwise comparisons, and the missingness is differential across the two conditions.","section":"Section 6.2.1"},{"comment":"The reliability of the pipeline is only partially validated. Visual descriptions were spot-checked on 200 topics (92.5% accuracy), and the creativity score distribution is reported, but the topic grouping, the dialogue reordering, and the coherence of inserted discussions are not evaluated against human judgment. The user study used only six videos for the controlled comparison, and Section 7.5 acknowledges that GPT-4o 'occasionally results in hallucinations and errors.' Without at least a sample-based agreement check on topic grouping and dialogue coherence, it is difficult to know whether the reported comprehension and social benefits generalise beyond the specific videos tested. The authors should either provide such validation or explicitly scope the claims to the studied video set.","section":"Section 4.7.4 and Section 7.5"}],"minor_comments":[{"comment":"The heading 'Personalization of DanumA11y' contains a typo: 'DanumA11y' should be 'DanmuA11y'.","section":"Section 7.1"},{"comment":"The sentence 'the AI visual descriptions should provide provide new, complementary information' contains a duplicated word 'provide'.","section":"Section 7.3.1"},{"comment":"In Table 1, 'Bibibili' appears as a platform name; this should be 'Bilibili'.","section":"Table 1"},{"comment":"The paper reports 18 Wilcoxon tests on questionnaire items without any correction for multiple comparisons. Several effects are at p<.05 and might not survive a conservative correction; the main p<.01 findings are likely robust, but this should be acknowledged.","section":"Section 5.2.3 / Table 4"},{"comment":"The topic-quality weights (λ_i=λ_c=λ_d=1) and the insertion-quality weight (λ_p=0.25, λ=10) are empirical choices, but the paper does not report any sensitivity analysis for these parameters. A brief statement of how sensitive the results are to these weights would strengthen the system's credibility.","section":"Section 4.5.1"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses an important problem and the user study is a genuine strength, but the incomplete baseline trials and the lack of transparency about missing-data handling are the most pressing issues. The reported effect sizes are large, so a sensitivity analysis may well preserve the conclusions, but the current presentation does not allow a reader to verify the claims. I would also direct the authors to the baseline validity concern: the comparison is against a custom probe rather than an existing deployed tool, which should be discussed more carefully. If the missing-data analysis is satisfying, I would be willing to support acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"DanmuA11y is a competent, caring piece of systems work: it identifies a real gap—BLV viewers can't meaningfully use Danmu—and builds a plausible three-part pipeline (LLM topic curation, visual descriptions, speech-aware insertion, spatial-audio multi-viewer discussion). The formative study is well done, and the 12-participant evaluation is a genuine effort with large effects. The authors also show good faith: they report a 92.5% accuracy spot check on visual descriptions and admit GPT-4o can hallucinate.\n\nThe soft spot is the missing-data handling in Section 6. The paper reports 36 trials per system, but seven participants abandoned nine baseline trials. Those missing trials are not missing at random—they are the worst baseline experiences. The paper then reports confusion rates per video and Wilcoxon tests without any sensitivity analysis or imputation. The numbers imply the missing trials were counted as zero-confusion observations, which biases the comparison in favor of DanmuA11y. That said, the raw effect is large (52 vs 9 confusion instances), so I suspect the main conclusions would survive a proper analysis—but the paper doesn't show that. This is a statistical sloppiness that a sharp reviewer will catch.\n\nTwo smaller issues: the baseline is a custom probe simulating current practice, not an existing tool, which limits the generality of the comparison; and LLM topic/creativity reliability is spot-checked rather than systematically validated. No code or data is released, which is common for CHI but limits replication.\n\nOverall, this deserves a serious peer review. A good referee should ask for a sensitivity analysis on the incomplete trials, a clear statement of denominators, and ideally a worst-case imputation. I'd send it to review and let the authors fix it. I'd bring it to a reading group if you want a good example of an accessibility systems paper with a subtle but fixable data issue.","headline":"A genuinely useful accessibility system paper, but the evaluation glosses over nine abandoned baseline trials; worth peer review after a missing-data fix.","tokens_in":31379,"tokens_out":2892,"would_cite":true,"duration_ms":26846,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"DanmuA11y claims that converting time-synced video comments into multi-viewer audio discussions with visual descriptions lets blind and low vision viewers understand Danmu, and reports significant gains in comprehension, smoothness, and…","keywords":["Danmu accessibility","blind and low vision","audio discussions","spatial audio","video comments","large language models","co-watching","screen reader alternatives"],"falsifier":"Run the pipeline on a fresh set of videos and measure whether the reported comprehension gains reproduce: if visual-description accuracy falls materially below 92.5% on videos with fast movement or unusual scenes, or if BLV participants' confusion rates with DanmuA11y rise toward the baseline rate, the central claim fails.","tokens_in":30369,"feed_emoji":"🎧","tokens_out":5293,"duration_ms":46175,"temperature":0.7,"pith_summary":"The paper argues that Danmu—the scrolling, time-synced comments overlaid on online videos—excludes blind and low vision (BLV) viewers because it is visual, overlaps with video speech, and becomes disorganized when read aloud sequentially. To fix this, the authors built DanmuA11y, which turns Danmu into multi-viewer audio discussions: an AI narrator supplies visual context, curated comments are reorganized into dialogues, and spatial audio positions the speakers around the listener. In a within-subject study with twelve BLV viewers, the system significantly reduced Danmu confusion, made the viewing experience feel smoother, and strengthened the sense of co-watching with other people. The contribution is a demonstrated way to make time-synced social video accessible, with design implications for live-streaming and other commentary-heavy media.","feed_headline":"Audio conversations make Danmu accessible to blind viewers","feed_subtitle":"In a 12-viewer study, the system cut confusion from 52 to 9 instances and strengthened co-watching.","key_machinery":"The central mechanism is a three-part pipeline anchored by a topic-quality and insertion-quality optimization. The pipeline first uses GPT-4o to group time-stamped Danmu comments into topics, filter redundant comments, and reorder them into dialogue-like sequences, and to generate a short visual description when a topic refers to something visible. It then schedules each topic at a non-speech segment or a speech break, solving an integer linear program that maximizes the weighted sum of topic quality (informativeness via sentence-embedding dissimilarity, creativity via a GPT-4o rating, and opinion diversity via sentiment labels) and insertion quality (language-model coherence minus a pause penalty). Finally, it renders each topic as a multi-viewer audio discussion using spatial audio and four alternating synthesized voices, with a shake gesture to access discussions on demand at speech breaks. This machinery lets the video's original speech remain uninterrupted while the audience commentary feels like a live conversation around the listener.","core_discovery":"DanmuA11y converts Danmu into audio discussions by three steps: grouping comments into topics and adding brief visual descriptions of what the comments refer to, scheduling each topic into a non-speech gap or speech break in the video via an integer-linear-programming optimization that balances topic quality (informativeness, creativity, opinion diversity) with insertion quality (coherence and pause length), and finally rendering the curated topics as dialogues spoken by an AI narrator and several human-voice virtual viewers placed around the listener with spatial audio. Compared with a baseline that simulated current screen-reader practice, the twelve BLV participants reported significantly fewer moments of Danmu confusion (nine versus fifty-two instances across thirty-six video views), wrote more detailed and more accurate video summaries, and rated the experience as more coherent, unobtrusive, and socially engaging. The authors present this as the first system to make Danmu itself accessible, rather than merely making a list of comments screen-reader friendly.","pith_inferences":["The same design could benefit sighted viewers in situations where the screen is unavailable or attention is split, since the audio-discussion format removes the need to read while watching.","Because the system's quality hinges on a large language model's vision-language ability, the approach will scale to other visual social media (memes, GIFs, short-form video) as that ability improves, rather than requiring per-platform redesign.","The optimization over topic placement is a general scheduling problem for accessibility; its reward function could be adapted to other constraints, such as minimizing interruption during key moments or maximizing discussion density for live events.","A testable extension is an ablation study removing spatial audio or visual descriptions; the current design bundles them, so their individual contributions to the reported social and comprehension gains are not yet isolated."],"forward_implications":["BLV viewers can understand Danmu at levels comparable to the paper's measures: reported confusion fell from 52 to 9 instances, and video summaries contained fewer factual errors.","The same audio-discussion pattern can carry other time-synced commentary, such as live-stream chat, if the pipeline can run in real time.","Platform designers can preserve the co-watching feeling of Danmu for BLV audiences by pairing visual descriptions with curated, dialogue-organized, spatially separated voices.","Users can choose on-demand when to hear commentary, which reduces split attention while keeping the option to dive deeper.","Personalization is a direct next step: adjusting the weights of informativeness, creativity, and diversity, or adding user-defined filters, would tailor the experience."],"supporting_citations":[{"why":"Establishes that Danmu creates a co-watching experience and motivates making it accessible.","marker":"[15]"},{"why":"Supplies the method for detecting non-speech segments via transcript gaps and volume.","marker":"[57]"},{"why":"Provides the coherence-scoring technique and audio-description placement heuristics reused by DanmuA11y.","marker":"[66]"},{"why":"Provides the informativeness metric and the automatic audio-description pipeline extended here.","marker":"[84]"},{"why":"Provides the prompt-based topic modeling method used to group and summarize Danmu comments.","marker":"[67]"},{"why":"Supplies the GPT-based subjective creativity-rating method used to score topics.","marker":"[65]"},{"why":"Defines the Inclusion of Other in the Self scale used to measure social closeness.","marker":"[3]"},{"why":"Supports the claim that spatial audio increases social presence, the basis for the multi-viewer layout.","marker":"[39]"}],"fun_headline_variants":["Audio discussions turn Danmu into co-listening for blind users","DanmuA11y converts video comments into spatial audio talks","Blind viewers hear Danmu as multi-voice audio discussions","Study: audio Danmu cuts confusion from 52 to 9 instances","Making Danmu accessible: audio dialogues for low-vision viewers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole benefit depends on GPT-4o reliably turning comments into coherent topics, describing visuals accurately, and rating creativity; the paper's only direct check is a 92.5% accuracy spot-check of visual descriptions on 200 topics, and the authors note the model occasionally hallucinates or errs.","fun_headline_variants_meta":{"raw":{"variants":["Audio discussions turn Danmu into co-listening for blind users","DanmuA11y converts video comments into spatial audio talks","Blind viewers hear Danmu as multi-voice audio discussions","Study: audio Danmu cuts confusion from 52 to 9 instances","Making Danmu accessible: audio dialogues for low-vision viewers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000227,"raw_usage":{"total_tokens":1468,"prompt_tokens":939,"completion_tokens":529,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":555,"completion_tokens_details":{"reasoning_tokens":441}},"tokens_in":555,"tokens_out":529,"duration_ms":5507,"temperature":1.0,"reasoning_tokens":441,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T14:00:30.732313+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the pipeline on a fresh set of videos and measure whether the reported comprehension gains reproduce: if visual-description accuracy falls materially below 92.5% on videos with fast movement or unusual scenes, or if BLV participants' confusion rates with DanmuA11y rise toward the baseline rate, the central claim fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports the claim that spatial audio increases social presence, the basis for the multi-viewer layout."}],"review_version":1}