{"id":"8078f1a7-6046-4177-82c6-b8f834e04b2d","arxiv_id":"1909.02538","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In a small user study, deaf and hard of hearing participants rated customizable closed interpreting higher than static or tracked interpreting for satisfaction, understanding, and ease of viewing.","lead":"Deaf and hard of hearing viewers in a 19-person study preferred a movable, resizable sign-language interpreter overlay over a fixed side-by-side layout for online videos. The paper is an early, useful test of an accessibility feature that could give deaf viewers caption-like control over interpreter placement.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's own analysis contradicts its main claim: Table 1 reports tracked vs. custom as not statistically significant, yet Section 5 calls the increase significant, and video content is confounded with condition.","rationale":"I agree with the reader's conditional verdict and only partially with the identified weakest assumption. The reader flagged the video-content confound, which is real and load-bearing. My stress-test finds a more direct internal problem: Table 1 explicitly reports that tracked and custom were not statistically significant, while Section 5 claims a significant increase between tracked and customizable. This contradiction means the paper's central comparative claim is unsupported as written. In addition, the Mann-Whitney U test is applied to paired within-subject ratings as if the observations were independent, so the reported significance levels are not trustworthy for this design; a paired test or mixed model is required. The condition-video confound compounds the issue by preventing a clean attribution of even the significant static-vs-custom difference to the interface. These concerns do not change the overall recommendation: the paper remains a plausible preliminary exploration that could be conditionally accepted if the claims are softened, the statistical analysis is corrected, and the confounding limitation is acknowledged. Since the reader already recommended CONDITIONAL, I leave the verdict unchanged.","tokens_in":5241,"tokens_out":4584,"duration_ms":49976,"concrete_test":"Obtain the raw Likert data from the 19 participants and re-analyze it as a repeated-measures comparison: a paired Wilcoxon signed-rank test for custom vs. tracked on each question, and a cumulative-link mixed model with participant random intercept plus condition and lecture-video identity as predictors. If custom vs. tracked remains non-significant, or if the custom vs. static effect disappears when video content is included as a covariate, the Section 5 conclusion should be withdrawn or softened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, that customizable interpreting is preferred over both tracked and static, is not supported by the paper's own results. Section 4 states that \"tracked and custom were not statistically significant,\" but Section 5 asserts a \"significant increase\" when comparing tracked to customizable. This is an internal contradiction: the headline ranking rests on a contrast the reported tests did not find significant. Even setting this aside, Section 3.1 paired each condition with a different YouTube lecture video and randomized only viewing order, not the assignment of video content to condition. Any difference in video difficulty, visual density, or instructor behavior is thereby aliased with interface condition, so even the significant static-vs-custom comparisons cannot be cleanly attributed to the interface. The use of Mann-Whitney U tests on paired within-subject ratings likewise treats repeated observations as independent, making the reported p-values invalid for this design. The qualitative comments are consistent with the means, but they cannot supply statistical support for a comparison the inferential analysis fails to establish.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces \"closed interpreting,\" a user-adjustable overlay for presenting an American Sign Language (ASL) interpreter alongside online videos. The tool allows the viewer to move, resize, and change the transparency of the interpreter video, analogous to closed captioning but with dynamic control. The authors report a within-subjects study with 19 deaf and hard-of-hearing participants who compared three implementations: static side-by-side interpreting, tracked interpreting (where the interpreter window moves to follow the lecture content), and customizable interpreting. Participants rated satisfaction, understanding, and ease of viewing on Likert scales, answered open-ended questions, and rated the value of the adjustable features. The paper reports descriptive means and Mann-Whitney U tests, and closes by claiming that viewers preferred tracked over static and customizable over both. The paper also reports that resizing and relocating were more valued than transparency.","tokens_in":5363,"tokens_out":4361,"duration_ms":43957,"significance":"If the findings were valid, the paper would make a useful contribution to the accessibility community: it offers a concrete tool design, gives empirical evidence about deaf viewers' preferences for adjustable interpreter placement, and provides a design guideline that resizing and repositioning may matter more than transparency. The qualitative comments are informative and consistent with the stated preferences. The tool itself is a plausible step toward better online video accessibility. However, the load-bearing statistical and experimental-design problems described below mean that the central preference claim is not established; the contribution is therefore currently more suggestive than confirmatory.","major_comments":[{"comment":"The conclusion contradicts the reported results. Section 4 states that \"tracked and custom were not statistically significant,\" yet Section 5 claims \"a noticeable and significant increase in satisfaction, understanding, and ease of viewing ... when comparing the tracked implementation to the customizable implementation.\" The final sentence of the paper, \"Our study indicates that people preferred ... the customizable interpreting over both the tracked and static interpreting views,\" is not supported by the statistical tests reported in Table 1. The authors must either correct the conclusion to align with the data or provide a defensible statistical analysis that actually supports the claim.","section":"Section 4 and Section 5"},{"comment":"The three conditions were paired with three different YouTube lecture videos, and only the presentation order was randomized; assignment of video content to condition was not counterbalanced. Any differences in video difficulty, visual density, or instructor behavior are therefore aliased with the interface condition. This is a fundamental confound: even the significant static-versus-custom comparisons cannot be cleanly attributed to the interpreting layout. This issue cannot be repaired by reanalysis and would require a new experiment with counterbalanced content-to-condition assignment.","section":"Section 3.1"},{"comment":"The Mann-Whitney U test is a between-subjects test, but the study used a within-subjects design: each participant rated all three implementations. Applying this test treats repeated observations as independent, inflates the effective sample size, and makes the p-values invalid. A paired test (e.g., Wilcoxon signed-rank test) or a repeated-measures model should be used. Additionally, nine comparisons are reported without any correction for multiple testing, so even the static-versus-custom differences may be false positives.","section":"Section 4, Table 1"},{"comment":"The statistical reporting is incomplete: no p-values, test statistics, or effect sizes are given for the Mann-Whitney U tests. The text only says that results are or are not statistically significant, so the reader cannot verify the magnitude or reliability of the differences. Full test results should be reported, along with confidence intervals or effect sizes for the mean differences.","section":"Section 4, Table 1"}],"minor_comments":[{"comment":"The last paragraph says participants \"were asked to rate their understanding and experience with each of the two videos,\" but there were three videos; this should say three.","section":"Section 3.1"},{"comment":"The sentence \"The participants watched videos in different order to reduce the bias that comes with watching a certain video, first or watching a certain video before or after another one\" is grammatically awkward and should be revised for clarity.","section":"Section 3.1"},{"comment":"In the Related Work, the sentence beginning \"Deaf students learn less than their hearing peers research on accessible views...\" appears to be missing a word or punctuation; it should be rewritten.","section":"Section 2"},{"comment":"The bar charts show means without error bars. Given the small sample size and the descriptive nature of the comparisons, adding error bars (e.g., standard error or confidence intervals) would help the reader assess variability.","section":"Figures 4 and 5"},{"comment":"The term \"custom\" is used interchangeably with \"customizable\" in the text and Table 1. The notation should be consistent to avoid confusion.","section":"Section 4"}],"recommendation":"reject","confidential_remarks":"This is a short paper with a useful tool idea and some valuable qualitative feedback, but the central quantitative claim is undermined by an internal contradiction between the results and the conclusion, an uncontrolled confound between condition and video content, and an inappropriate statistical test for the within-subjects design. The confound in particular cannot be fixed by revision; it requires a new experiment. I recommend rejection, although a redesigned study with counterbalancing and correct paired statistics could be a strong future contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper introduces \"closed interpreting\" — a toggleable, user-adjustable ASL interpreter overlay for online video — and reports a 19-participant preference study. That interface idea is genuinely new as far as the cited prior work goes, and the open-ended comments give a useful, concrete picture of what deaf and hard of hearing viewers want from an interpreter overlay. The feature-level ratings (move 4.42, resize 4.31, transparency 3.42) are plausible and directionally consistent.\n\nThe problem is that the paper's central claim — customizable interpreting preferred over both tracked and static — is not supported by its own results. Section 4 explicitly says tracked and custom were \"not statistically significant,\" yet Section 5 asserts a \"significant increase\" for that same comparison. That is an internal contradiction. Table 1 does not report p-values or effect sizes, only the significance verdict, so a reader cannot even check the contested comparison.\n\nThere is also a design confound: each condition was paired with a different YouTube lecture video, and only viewing order was randomized, not the assignment of content to condition. Any difference in lecture difficulty, visual density, or instructor behavior is aliased with the interface condition. So even the significant static-vs-custom comparisons cannot be cleanly attributed to the interface. On top of that, Mann-Whitney U tests on within-subject Likert ratings treat repeated observations as independent; the appropriate paired test would be Wilcoxon signed-rank. The reported p-values are therefore invalid as computed.\n\nThe qualitative comments are consistent with the means, and I don't doubt the general direction. But the paper overreaches in its conclusion, and the statistical reporting is too thin for the claims it makes. This is a promising preliminary finding, not a definitive result.\n\nWho is this for? Researchers working on accessibility interfaces for deaf users, and possibly video platforms thinking about interpreter placement. They would get a useful concept and some suggestive data, but they should not cite the specific ranking without checking a revised version. I would not desk reject it — the idea deserves referee time — but I would send it back with a clear request for paired tests, content counterbalancing or an explicit caveat, and softened conclusions. If this is for a workshop or poster venue, it is close to acceptable as-is; if it is for a full archival venue, it needs substantial revision before it can be trusted.","headline":"A genuinely new accessibility interface idea, but the paper's own statistics contradict its headline preference ranking and the design confounds content with condition.","tokens_in":5871,"tokens_out":1540,"would_cite":false,"duration_ms":18466,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Deaf and hard-of-hearing viewers prefer a user-adjustable, movable ASL interpreter overlay over static and tracked interpreter layouts.","keywords":["closed interpreting","American Sign Language","ASL interpreter overlay","deaf and hard of hearing","video accessibility","user preferences","adjustable video interfaces","online video"],"falsifier":"Run the same three-layout comparison with the identical lecture content in every condition (or fully counterbalancing video-to-layout assignment), and add a factual comprehension check after viewing; if the customizable layout no longer outranks static and tracked, the preference finding collapses.","tokens_in":5026,"feed_emoji":"🧏","tokens_out":5260,"duration_ms":50636,"temperature":0.7,"pith_summary":"This paper proposes \"closed interpreting\": an optional American Sign Language interpreter video that overlays an online lecture and can be toggled on and off, moved, resized, and made transparent, just as closed captions can be toggled. The authors report a study with 19 deaf and hard-of-hearing participants who watched lecture videos under three layouts: static side-by-side, tracked (the interpreter window follows the instructor's content), and customizable. Ratings for satisfaction, understanding, and ease of viewing were highest for the customizable layout, and the difference between customizable and static was statistically significant; tracked also beat static on satisfaction and ease. The paper concludes that viewers prefer adjustable interpreter placement, and that resizing and relocating mattered more than transparency.","feed_headline":"Adjustable ASL interpreter beats static and tracked layouts","feed_subtitle":"Deaf viewers in a 19-person study favored a movable, resizable ASL interpreter.","key_machinery":"The carrying mechanism is the three closed-interpreting layouts implemented in HTML: static (lecture and interpreter side by side, immovable), tracked (interpreter window manually repositioned to align with the area of the lecture being discussed), and customizable (drag-to-move, drag-to-resize, an opacity slider from 0 to 100, plus pause/play and hide/show). The evaluation uses Likert-scale ratings (1–5) on satisfaction, understanding, and ease of viewing for all three conditions, extra feature ratings for the customizable condition, open-ended comments, and Mann-Whitney U tests for pairwise comparisons.","core_discovery":"The central claim is that when deaf viewers watch interpreted online videos, they prefer an interpreter window they can reposition and resize over a fixed side-by-side layout, and they prefer that flexible overlay even over one that automatically tracks the lecture content. In the authors' comparison, the customizable implementation received the highest mean ratings on satisfaction, understanding, and ease of viewing (4.36, 4.58, and 4.42 on 1–5 scales), and the static implementation received the lowest (3.16, 3.89, and 3). The Mann-Whitney U tests showed static versus custom differences were statistically significant, and tracked versus static was significant for satisfaction and ease; tracked versus custom differences did not reach significance. Open-ended comments support the ranking, with viewers complaining about losing track of content in the static layout and praising the control offered by the customizable one.","pith_inferences":["If the preference ranking reflects genuine comprehension differences, giving viewers control over interpreter placement could shrink part of the information-access gap for deaf viewers in online lectures without requiring video creators to produce separate interpreted and uninterpreted versions.","The pattern is consistent with a split-attention explanation: moving the interpreter nearer the referenced content reduces gaze shifts, which may be why tracked and customizable outperformed static; this predicts that a comprehension benefit would persist under eye-tracking measures, not just self-report.","A testable extension: compare a smooth automatic tracking mode against manual drag placement to see whether users prefer control or convenience once tracking is polished.","The adjustable overlay could generalize beyond ASL to caption placement, subtitles, or picture-in-picture for any video where viewers juggle two visual streams."],"forward_implications":["If confirmed, video platforms could offer closed interpreting the way they offer closed captions, letting viewers place the interpreter over or beside the content and adjust it to their needs.","The higher ratings for customizable and tracked layouts over static support designs that reduce the distance between interpreter and referenced lecture content.","Resizing and repositioning should be foregrounded in interface design; the lower transparency rating suggests opacity is a secondary control.","The preference data motivate automated tracking, since manual tracking produced gains but moved too abruptly for some viewers.","Future work can target smooth tracking and interpreter backgrounds that blend with lecture content."],"supporting_citations":[{"why":"Supplies the earlier finding on preferred screen configuration for interpreter video, which the static and customizable layouts build on.","marker":"[1]"},{"why":"Documents caption speed exceeding reading ability, motivating why viewers may prefer interpreters over captions.","marker":"[2]"},{"why":"Establishes the ~2x replay-speed limit for ASL, informing the playback features in the tool.","marker":"[5]"},{"why":"Explains why deaf students learn less from on-screen text, supporting the argument for interpreter-based access.","marker":"[6]"},{"why":"Evaluates a pausing/resuming caption interface that the closed-interpreting controls extend.","marker":"[4]"},{"why":"Shows value of gaze cues for switching attention between interpreter and slides, related to tracked interpreting.","marker":"[3]"}],"fun_headline_variants":["Adjustable ASL overlay beats static and tracked layouts","Deaf viewers prefer customizable sign-language overlay","Movable ASL interpreter wins over fixed and auto-tracking","Resizable ASL window tops static and tracked in study","User-adjustable ASL interpreter gets highest ratings"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The three layouts were paired with different lecture videos and only the viewing order was randomized; if the videos differed in difficulty or visual density, the rating gaps could reflect video content rather than interpreter layout, and the study relied on self-reported understanding with no objective comprehension check.","fun_headline_variants_meta":{"raw":{"variants":["Adjustable ASL overlay beats static and tracked layouts","Deaf viewers prefer customizable sign-language overlay","Movable ASL interpreter wins over fixed and auto-tracking","Resizable ASL window tops static and tracked in study","User-adjustable ASL interpreter gets highest ratings"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000339,"raw_usage":{"total_tokens":1830,"prompt_tokens":861,"completion_tokens":969,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":477,"completion_tokens_details":{"reasoning_tokens":892}},"tokens_in":477,"tokens_out":969,"duration_ms":7935,"temperature":1.0,"reasoning_tokens":892,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:46:38.273785+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same three-layout comparison with the identical lecture content in every condition (or fully counterbalancing video-to-layout assignment), and add a factual comprehension check after viewing; if the customizable layout no longer outranks static and tracked, the preference finding collapses.","supporting_citations":[{"cited_title":"closed interpreting","cited_arxiv_id":null,"evidence_quote":"Supplies the earlier finding on preferred screen configuration for interpreter video, which the static and customizable layouts build on."},{"cited_title":"In general, hearing viewers are able to listen to the verbal information and attend to the visual information simultaneously [7]","cited_arxiv_id":null,"evidence_quote":"Documents caption speed exceeding reading ability, motivating why viewers may prefer interpreters over captions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the ~2x replay-speed limit for ASL, informing the playback features in the tool."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Explains why deaf students learn less from on-screen text, supporting the argument for interpreter-based access."},{"cited_title":"One problem is when I looked away from interpreter video to read math, I missed what interpreter said","cited_arxiv_id":null,"evidence_quote":"Evaluates a pausing/resuming caption interface that the closed-interpreting controls extend."},{"cited_title":"tracking","cited_arxiv_id":null,"evidence_quote":"Shows value of gaze cues for switching attention between interpreter and slides, related to tracked interpreting."}],"review_version":1}