REVIEW 4 major objections 5 minor 15 references
Closed ASL Interpreting for Online Videos
T0 review · 4 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Deaf and hard-of-hearing viewers prefer a user-adjustable, movable ASL interpreter overlay over static and tracked interpreter layouts.
desk verdict A genuinely new accessibility interface idea, but the paper's own statistics contradict its headline preference ranking and the design confounds content with condition. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is the three closed-interpreting layouts implemented in HTML: static (lecture and interpreter side by side, immovable), tracked (interpreter window manually repositioned to align with the area of the lecture being discussed), and customizable (drag-to-move, drag-to-resize, an opacity slider from 0 to 100, plus pause/play and hide/show). The evaluation uses Likert-scale ratings (1–5) on satisfaction, understanding, and ease of viewing for all three conditions, extra feature ratings for the customizable condition, open-ended comments, and Mann-Whitney U tests for pairwise comparisons.
What would settle it
Run the same three-layout comparison with the identical lecture content in every condition (or fully counterbalancing video-to-layout assignment), and add a factual comprehension check after viewing; if the customizable layout no longer outranks static and tracked, the preference finding collapses.
Extended reading notes
Core claim
The central claim is that when deaf viewers watch interpreted online videos, they prefer an interpreter window they can reposition and resize over a fixed side-by-side layout, and they prefer that flexible overlay even over one that automatically tracks the lecture content. In the authors' comparison, the customizable implementation received the highest mean ratings on satisfaction, understanding, and ease of viewing (4.36, 4.58, and 4.42 on 1–5 scales), and the static implementation received the lowest (3.16, 3.89, and 3). The Mann-Whitney U tests showed static versus custom differences were statistically significant, and tracked versus static was significant for satisfaction and ease; tracked versus custom differences did not reach significance. Open-ended comments support the ranking, with viewers complaining about losing track of content in the static layout and praising the control offered by the customizable one.
Load-bearing premise
The three layouts were paired with different lecture videos and only the viewing order was randomized; if the videos differed in difficulty or visual density, the rating gaps could reflect video content rather than interpreter layout, and the study relied on self-reported understanding with no objective comprehension check.
Editorial extensions
If this is right
- If confirmed, video platforms could offer closed interpreting the way they offer closed captions, letting viewers place the interpreter over or beside the content and adjust it to their needs.
- The higher ratings for customizable and tracked layouts over static support designs that reduce the distance between interpreter and referenced lecture content.
- Resizing and repositioning should be foregrounded in interface design; the lower transparency rating suggests opacity is a secondary control.
- The preference data motivate automated tracking, since manual tracking produced gains but moved too abruptly for some viewers.
- Future work can target smooth tracking and interpreter backgrounds that blend with lecture content.
Reading between the lines
- If the preference ranking reflects genuine comprehension differences, giving viewers control over interpreter placement could shrink part of the information-access gap for deaf viewers in online lectures without requiring video creators to produce separate interpreted and uninterpreted versions.
- The pattern is consistent with a split-attention explanation: moving the interpreter nearer the referenced content reduces gaze shifts, which may be why tracked and customizable outperformed static; this predicts that a comprehension benefit would persist under eye-tracking measures, not just self-report.
- A testable extension: compare a smooth automatic tracking mode against manual drag placement to see whether users prefer control or convenience once tracking is polished.
- The adjustable overlay could generalize beyond ASL to caption placement, subtitles, or picture-in-picture for any video where viewers juggle two visual streams.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces "closed interpreting," a user-adjustable overlay for presenting an American Sign Language (ASL) interpreter alongside online videos. The tool allows the viewer to move, resize, and change the transparency of the interpreter video, analogous to closed captioning but with dynamic control. The authors report a within-subjects study with 19 deaf and hard-of-hearing participants who compared three implementations: static side-by-side interpreting, tracked interpreting (where the interpreter window moves to follow the lecture content), and customizable interpreting. Participants rated satisfaction, understanding, and ease of viewing on Likert scales, answered open-ended questions, and rated the value of the adjustable features. The paper reports descriptive means and Mann-Whitney U tests, and closes by claiming that viewers preferred tracked over static and customizable over both. The paper also reports that resizing and relocating were more valued than transparency.
Significance. If the findings were valid, the paper would make a useful contribution to the accessibility community: it offers a concrete tool design, gives empirical evidence about deaf viewers' preferences for adjustable interpreter placement, and provides a design guideline that resizing and repositioning may matter more than transparency. The qualitative comments are informative and consistent with the stated preferences. The tool itself is a plausible step toward better online video accessibility. However, the load-bearing statistical and experimental-design problems described below mean that the central preference claim is not established; the contribution is therefore currently more suggestive than confirmatory.
major comments (4)
- [Section 4 and Section 5] The conclusion contradicts the reported results. Section 4 states that "tracked and custom were not statistically significant," yet Section 5 claims "a noticeable and significant increase in satisfaction, understanding, and ease of viewing ... when comparing the tracked implementation to the customizable implementation." The final sentence of the paper, "Our study indicates that people preferred ... the customizable interpreting over both the tracked and static interpreting views," is not supported by the statistical tests reported in Table 1. The authors must either correct the conclusion to align with the data or provide a defensible statistical analysis that actually supports the claim.
- [Section 3.1] The three conditions were paired with three different YouTube lecture videos, and only the presentation order was randomized; assignment of video content to condition was not counterbalanced. Any differences in video difficulty, visual density, or instructor behavior are therefore aliased with the interface condition. This is a fundamental confound: even the significant static-versus-custom comparisons cannot be cleanly attributed to the interpreting layout. This issue cannot be repaired by reanalysis and would require a new experiment with counterbalanced content-to-condition assignment.
- [Section 4, Table 1] The Mann-Whitney U test is a between-subjects test, but the study used a within-subjects design: each participant rated all three implementations. Applying this test treats repeated observations as independent, inflates the effective sample size, and makes the p-values invalid. A paired test (e.g., Wilcoxon signed-rank test) or a repeated-measures model should be used. Additionally, nine comparisons are reported without any correction for multiple testing, so even the static-versus-custom differences may be false positives.
- [Section 4, Table 1] The statistical reporting is incomplete: no p-values, test statistics, or effect sizes are given for the Mann-Whitney U tests. The text only says that results are or are not statistically significant, so the reader cannot verify the magnitude or reliability of the differences. Full test results should be reported, along with confidence intervals or effect sizes for the mean differences.
minor comments (5)
- [Section 3.1] The last paragraph says participants "were asked to rate their understanding and experience with each of the two videos," but there were three videos; this should say three.
- [Section 3.1] The sentence "The participants watched videos in different order to reduce the bias that comes with watching a certain video, first or watching a certain video before or after another one" is grammatically awkward and should be revised for clarity.
- [Section 2] In the Related Work, the sentence beginning "Deaf students learn less than their hearing peers research on accessible views..." appears to be missing a word or punctuation; it should be rewritten.
- [Figures 4 and 5] The bar charts show means without error bars. Given the small sample size and the descriptive nature of the comparisons, adding error bars (e.g., standard error or confidence intervals) would help the reader assess variability.
- [Section 4] The term "custom" is used interchangeably with "customizable" in the text and Table 1. The notation should be consistent to avoid confusion.
Circularity Check
No circularity: the paper's conclusions rest on direct survey responses, not on fitted parameters, derived quantities, or load-bearing self-citations.
full rationale
This paper is an empirical user study comparing three closed-interpreting interface conditions. Its central claims are based on Likert-scale self-reports and open-ended qualitative comments, not on any mathematical derivation, fitted parameter, or model whose outputs are recycled as predictions. The related-work citations to the authors' earlier papers ([3], [4], [5]) are contextual background and are not used to justify the study's conclusions; no uniqueness theorem or prior result is invoked to force the interface design. The 'closed interpreting' concept is an analogy to closed captioning rather than a renamed known result, and the paper's contribution is an implementation and evaluation rather than a formal derivation. The most serious issues with the paper are methodological and inferential, not circular: the conditions were paired with different lecture videos while only viewing order was randomized (Section 3.1), so video content is confounded with condition; Mann-Whitney U tests were applied to paired within-subject ratings as though independent; and the text in Section 5 asserts a significant increase from tracked to customizable even though the paper's own Table 1 says those two conditions were not statistically significant. These are threats to validity and internal consistency, but they are not instances of the paper defining its conclusion into its premises, fitting a parameter and then calling it a prediction, or importing a conclusion from a self-citation. Under the strict circularity criteria, nothing in the paper reduces by construction to its own inputs, so the appropriate score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption The three lecture videos assigned to static, tracked, and customizable conditions are comparable in difficulty, so rating differences are attributable to the interface.
- domain assumption Self-reported Likert ratings of satisfaction, understanding, and ease of viewing are valid proxies for actual comprehension and usability.
- domain assumption A sample of 19 self-selected deaf and hard of hearing volunteers is representative enough to support the stated preference claims.
- domain assumption The human-tracked interpreter placement represents what a usable or automatic tracking system would provide.
Cite this review
Pith. "Pith review of Closed ASL Interpreting for Online Videos." pith.science (2026). https://pith.science/paper/3J765FXV
@misc{pith2026190902538,
author = {Pith},
title = {Pith review of: Closed ASL Interpreting for Online Videos},
year = {2026},
howpublished = {\url{https://pith.science/paper/3J765FXV}},
note = {Machine review of arXiv:1909.02538}
}
read the original abstract
Deaf individuals face great challenges in today's society. It can be very difficult to be able to understand different forms of media without a sense of hearing. Many videos and movies found online today are not captioned, and even fewer have a supporting video with an interpreter. Also, even with a supporting interpreter video provided, information is still lost due to the inability to look at both the video and the interpreter simultaneously. To alleviate this issue, we came up with a tool called closed interpreting. Similar to closed captioning, it will be displayed with an online video and can be toggled on and off. However, the closed interpreter is also user-adjustable. Settings, such as interpreter size, transparency, and location, can be adjusted. Our goal with this study is to find out what deaf and hard of hearing viewers like about videos that come with interpreters, and whether the adjustability is beneficial.
Figures
Reference graph
Works this paper leans on
-
[1]
Introduction It can be difficult for people with sensory disabilities, to obtain equal access to information. Deaf and hard of hearing viewers usually need visual access to aural information in online videos as they cannot understand the aural information. The majority of videos online are no t captioned and even fewer show sign language translation via a...
-
[2]
Related Work Prior research work has shown that deaf and hard of hearing viewers do not get equal access to videos, compared with their hearing peers. In general, hearing viewers are able to listen to the verbal information and attend to the visual information simultaneously [7]. On the other hand, deaf viewers usually cannot view the v isual translation ...
-
[3]
Methodology To implement closed interpreting, we used a lecture video that taught mathematics sequences. We also obtained an interpreter video that corresponds to the lecture video and both videos were embedded in an HTML browser. In the static closed interpreting implementation, both the interpreter and lecture videos are side by side and immovable. An e...
-
[4]
Experimental Results After conducting the experiments, we collected the Likert scale results, answers to open -ended questions, and other feedback. Each of the three implementations had three Likert scale questions in common, relating to user satisfaction, understanding of the lecture video contents, and how easy it was to view the lecture and interpreter...
-
[5]
Conclusions We found that many people did indeed like the tracked and customizable closed interpreting implementation over the static one. While there were several survey results and feedback that disagreed with each other, in general, there was a noticeable and significant increase in satisfaction, understanding, and ease of viewing when comparing the st...
-
[6]
Future Work Tracked video could be done automatically, by programming it t o locate the most optimal place for interpreter . Other possibilities include addressing open-ended feedback, such as the background of the interpreter and the lecture vid eo contrasting too sharply in the customizable implementation
-
[7]
ACKNOWLEDGMENTS This work was supported by the Na tional Science Foundation Award IIS-1460894
-
[8]
Cavender, A.C., Bigham, J.P. and Ladner, R.E. 2009. ClassInFocus: Enabling improved visual attention strategies for deaf and hard of hearing students. Proceedings of the 11th International ACM SIGACCESS Conference on Computers and Accessibility - ASSETS ’09 (New York, New York, USA, 2009), 67–74
work page 2009
Show all 15 references
-
[9]
Jensema, C. 1996. Closed -captioned television presentation speed and vocabulary. American Annals of the Deaf. 141, 4 (1996), 284–292
1996
-
[10]
and Kushalnagar, P
Kushalnagar, R. and Kushalnagar, P. 2014. Collaborative Gaze Cues and Replay for Deaf and Hard of Hearing Students. Computers Helping People with Special Needs (2014), 415–422
2014
-
[11]
and Steed, D.S
Kushalnagar, R.S., Rivera, J.J., Yu, W. and Steed, D.S
-
[13]
and Kolash, N.A
Kushalnagar, R.S., Trager, B.P., Beiter, K.B. and Kolash, N.A. 2013. American Sign Language: Maximum Live Replay Speed. Rehabilitation Engineering and Assistive Technology Society of North America (Seattle, WA, 2013), 1–4
2013
-
[14]
and Seewagen, R
Marschark, M., Sapere, P., Convertino, C. and Seewagen, R. 2005. Access to postsecondary education through sign language interpreting. Journal of Deaf Studies and Deaf Education. 10, 1 (Jan. 2005), 38–50
2005
-
[15]
and Sweller, J
Mousavi, S.Y., Low, R. and Sweller, J. 1995. Reducing cognitive load by mixing auditory and visual presentation modes. Journal of Educational Psychology. 87, 2 (1995), 319–334
1995
-
[2014]
Proceedings of the 16th international ACM SIGACCESS conference on Computers & accessibility - ASSETS ’14 (New York, New York, USA, 2014), 287–288
AVD -LV. Proceedings of the 16th international ACM SIGACCESS conference on Computers & accessibility - ASSETS ’14 (New York, New York, USA, 2014), 287–288
2014
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.