{"id":"b2be2275-7b6c-4f3c-a346-7a4863f34fd4","arxiv_id":"2508.19971","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"CapTune lets caption creators set bounds and viewers tune non-speech caption text; a 19-person qualitative evaluation reported greater engagement and retained creative control.","lead":"CapTune lets caption creators set safe boundaries, then deaf and hard of hearing viewers adjust non-speech captions, such as sound effects and music descriptions, to their own taste using AI. A small qualitative study with 7 creators and 12 DHH viewers suggests this can raise emotional engagement while keeping creator control.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unverified LLM adherence to creator anchors is the load-bearing risk: no metric shows transformed captions stay in bounds or preserve meaning, yet the abstract claims it preserves creator intent.","rationale":"The reader's weakest assumption is exactly the load-bearing risk I identify: GPT-4o's transformation reliability is unmeasured, yet the abstract's 'preserving creator intent' and 'enhancing emotional engagement' depend on it. This is not a disagreement with the HCI community's qualitative methods; it is a pointed gap between the system's stated guarantee and the evidence provided. The paper itself flags this limitation in Sec. 7.7, and participant quotes in Secs. 5.3.2 and 6.4.6 show observable failures of semantic fidelity and consistency. The qualitative themes of creative control and emotional engagement are still valuable and may survive additional scrutiny, so a full rejection is unwarranted. However, the conditional verdict is appropriate: the abstract's strong causal language should be softened and the LLM reliability assumption should be measured. My concrete test would settle whether the concern lands by quantifying adherence and semantic fidelity. The reader's verdict remains CONDITIONAL, so no change to the verdict is needed.","tokens_in":23507,"tokens_out":3804,"duration_ms":42798,"concrete_test":"Collect a corpus of N >= 100 NSI captions spanning multiple sound types and video genres. For each caption, define lower/upper anchors via the Creator Tool, then generate transformed captions at a grid of (level of detail, expressiveness) values using the Sec. 4.3.3 prompt. Have independent human raters (at least 3 per item) judge: (a) whether each output falls within the creator-defined bounds on the detail and expressiveness scales; (b) whether it preserves the source sound event's meaning; (c) whether identical input captions yield consistent transformations. Pre-register thresholds, e.g., >=90% within bounds and >=95% semantic fidelity. If adherence falls below threshold, the 'preserving creator intent' claim is unsupported and the abstract's causal language must be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that CapTune 'preserves creator intent' depends entirely on GPT-4o producing transformed captions that stay within the creator-defined anchor space and preserve the semantic content of each sound event. The 'anchored' mechanism in Sec. 4.3.3 is prompt engineering: it feeds anchor captions and interpolation ratios to GPT-4o, but nothing checks that the output actually lies within the anchor bounds or retains the original meaning. The authors concede in Sec. 7.7 that outputs 'can still produce inconsistent or semantically inaccurate outputs' and that they collected 'no systematic metrics of caption transformation accuracy or consistency.' Participant data already show concrete failures: C2 (Sec. 5.3.2) flagged an interpretation as 'overly specific'; P2 (Sec. 6.4.6) observed identical [Dolphin whistles] transformed inconsistently; P3 and P8 questioned semantic overreach. Without an adherence/fidelity measurement, the abstract's 'preserving creator intent' is an unsupported assertion. Moreover, the qualitative 'creative control' findings cannot be attributed to the anchor mechanism rather than to generic LLM prompting, because no condition isolates the anchor-based prompt from unconstrained generation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces CapTune, a system for customizing non-speech captions (NSI) in video for deaf and hard of hearing (DHH) viewers. Caption creators define a two-dimensional transformation space using anchor points for Level of Detail and Expressiveness; viewers then select preferences within that space, optionally toggling sound representation method and genre alignment. Transformations are generated by GPT-4o using prompts that encode interpolation ratios (Eqs. 2 and 3) relative to the creator's anchors and the current caption. The system was evaluated in two qualitative studies: seven caption creators used the Creator Tool, and twelve DHH participants used the Viewer Client. The paper reports that creators felt they retained creative control and that DHH viewers experienced enhanced emotional engagement, while also surfacing trade-offs between information richness and cognitive load, tensions between interpretive and descriptive sound representation, and context-dependent preferences. The authors acknowledge in Section 7.7 that the system can produce inconsistent or semantically inaccurate outputs and that no systematic metrics of caption transformation accuracy were collected.","tokens_in":23664,"tokens_out":2806,"duration_ms":30890,"significance":"If the central claims hold, CapTune is a useful contribution to accessible media: it moves beyond visual styling of captions to personalized transformation of caption text itself, grounded in a qualitative analysis of DHH viewers' expressed preferences. The system is open-sourced, and the evaluation is substantial for an HCI paper, including a three-level codebook with 168 third-level codes and interrater reliability of 0.74, which is a concrete strength. The design space (level of detail, expressiveness, sound representation, genre alignment) is well motivated and the findings on context-dependent preferences and cognitive load are valuable for future captioning systems. However, the two headline claims—'preserving creator intent' and 'enhancing viewers' emotional engagement'—are not equally supported. The first depends on an unverified assumption that GPT-4o transformations stay within creator-defined bounds and preserve sound-event semantics; the second rests entirely on self-report without a control or baseline. These gaps are acknowledged in the paper itself, but they are load-bearing for the abstract's claims.","major_comments":[{"comment":"The claim that CapTune 'preserves creator intent' is not supported by the evidence presented. Equations (2) and (3) only compute interpolation ratios; they do not constrain the LLM output. The transformation is a prompted call to GPT-4o with no verification that the output lies within the anchor-defined bounds or preserves the meaning of the original sound event. The authors concede in §7.7 that outputs 'can still produce inconsistent or semantically inaccurate outputs' and that no systematic metrics of caption transformation accuracy or consistency were collected. Participant data already show concrete violations: C2 (Section 5.3.2) flagged 'Anna exhales with a sigh of relief' as overly specific; P2 (Section 6.4.6) observed identical [Dolphin whistles] transformed inconsistently; P3 and P8 questioned interpretive overreach. The abstract's 'preserving creator intent' is therefore an over","section":"§4.3.3 and §7.7"},{"comment":"The claim that CapTune 'enhanced viewers' emotional engagement' rests on self-report from nine of twelve participants during a single-session, no-baseline, no-control qualitative study. Participants were introduced to the system and asked to explore it while watching short clips; there was no comparison with the original unmodified captions, no alternative condition (e.g., unconstrained LLM transformation, or a non-anchored personalization interface), and no measurement of engagement beyond interview statements. This design cannot rule out novelty effects or demand characteristics, and it does not support the directional claim of 'enhancing' engagement. A controlled comparison or at minimum a pre/post self-report with the original caption track as baseline would be needed to substantiate the wording in the abstract.","section":"§6.4.1 and §6.2"},{"comment":"The qualitative findings about creators' 'creative control' cannot be attributed specifically to the anchored transformation mechanism, because no condition isolates it from generic LLM prompting. Creators interacted with a full interface that included sliders, previews, manual editing, and locking; any of these could produce the sense of agency that participants described. C2's and C6's concerns about semantic accuracy are evidence that the anchor mechanism did not reliably prevent over-interpretation. Without an ablation or a side-by-side comparison with unconstrained GPT-4o generation, the paper should not imply that anchoring is the mechanism responsible for the reported creative-control benefits; it can only claim that creators felt control in this particular system configuration.","section":"§4.3.3, §5.3.2, and §7.2"}],"minor_comments":[{"comment":"The text 'visualized in Figure 4.2' (Section 4.2.3) appears to reference a figure number incorrectly; check whether it should be a numbered figure, likely Figure 1 or a dedicated anchor-space figure.","section":"Figure references"},{"comment":"There is a typo in C6's entry: 'Professioal ads' should be 'Professional ads.'","section":"Table 2"},{"comment":"The description of how GPT-4o determines baseline Level of Detail and Expressiveness values is underspecified. State whether this is a one-time analysis of the whole caption file or per caption, and provide the exact prompt used, since the baseline affects all subsequent transformations.","section":"§4.2.2"},{"comment":"The prompt template refers to 'the [lower-anchor captions]' and 'the [upper-anchor captions]' without clarifying how multiple captions at an anchor are sampled or represented. If anchors are per-caption, clarify; if anchors are global, explain how a single pair of exemplars is chosen for each transformation.","section":"§4.3.3"},{"comment":"The Reddit-derived dataset is small (51 posts from 13 unique threads) and may be subject to self-selection. This is acceptable for a formative analysis, but the paper should acknowledge more explicitly that the four design opportunities are drawn from a narrow online sample.","section":"§3.1"},{"comment":"The inconsistency example of [Dolphin whistles] is an important data point, but the text does not state which parameter settings produced the differing outputs. Including that context would strengthen the finding and help future work reproduce or address the issue.","section":"§6.4.6"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid ASSETS contribution in terms of system design and qualitative evaluation, but the abstract overclaims what is demonstrated. The author team is well-known in this area and the open-sourced code is a positive. The main risk is that reviewers or readers take 'preserving creator intent' as a verified property of the anchoring method when it is, by the authors' own admission, an aspiration. I recommend major revision focused on claim calibration and, ideally, adding a lightweight adherence check or a comparative condition. I would not reject: the core system and user-study findings are valuable and the limitations are stated in the paper, even if the abstract does not reflect them."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"CapTune is a legitimate new system contribution: it turns static non-speech captions into a viewer-adjustable parameter space, using creator-set anchors as guardrails and letting viewers tune detail, expressiveness, sound representation, and genre alignment. The work builds sensibly on prior caption-styling and LLM-control literature, and the co-produced creator/viewer workflow is genuinely new for this domain. The qualitative evaluation is careful—168 third-level codes, IRR 0.74, transparent quotes that show both strengths and failures—and the authors do not hide the system's unreliability. The Reddit analysis grounding the four design dimensions is modest in scale (51 posts) but used appropriately as formative input, not as a population-level claim. Credit is due for the open-sourcing intent and the clear writing.\n\nThe soft spots are real but proportionate. The abstract says CapTune \"preserves creator intent\" and \"enhanced viewers' emotional engagement,\" but the engagement claim rests on self-report with no control or baseline, and the preservation claim depends entirely on GPT-4o staying within anchor bounds and keeping semantics intact. No systematic accuracy or consistency metric is reported. The stress-test note is correct: Sec. 7.7 admits outputs \"can still produce inconsistent or semantically inaccurate outputs,\" and participant quotes in Secs. 5.3.2 and 6.4.6 show concrete failures—C2's \"overly specific\" flag, P2's inconsistent [Dolphin whistles], P3 and P8's interpretive-overreach concerns. Because no condition isolates anchor-based prompting from unconstrained generation, the observed \"creative control\" cannot be attributed specifically to the anchoring mechanism. That is a load-bearing limitation, but it is a limitation, not a fatal flaw; the qualitative findings about preferences, tension between information richness and cognitive load, and context-dependence all survive. The small inconsistencies (slider recalibration description vs. Figure 3, GitHub link without commit hash or data) are minor and fixable. I also note the paper cites the authors' own prior work where relevant, which is fine here since those systems are concrete and checkable.\n\nWho is this for? Accessibility researchers, captioning tool designers, and anyone working on controllable LLM generation for end-user customization. It deserves a serious referee—the system is novel enough, the evaluation is honest enough, and the limitations are stated clearly enough that peer review should engage with the claims rather than desk-reject. My verdict would be major revision: soften the abstract to match the evidence, add a simple fidelity measure of anchor adherence on a sample of transformations, and isolate the anchor mechanism in a follow-up. I would bring it to a reading group and would cite it as the current state of the art in personalized NSI captioning, with the caveat about unmeasured reliability noted.","headline":"CapTune is a solid, clearly described HCI systems paper whose main risk—unverified LLM adherence to creator anchors—is real but openly acknowledged; the qualitative findings hold up, though the abstract overstates what was measured.","tokens_in":24237,"tokens_out":1167,"would_cite":true,"duration_ms":15428,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"CapTune claims non-speech captions can be personalized for deaf and hard-of-hearing viewers—tuning detail, expressiveness, sound style, and genre fit—while anchored, creator-set bounds preserve editorial control.","keywords":["non-speech captions","deaf and hard of hearing","caption personalization","generative AI","large language models","anchored transformation","accessibility","video captions"],"falsifier":"A systematic accuracy and consistency audit: take a fixed caption set across genres, apply the transformation at every grid point in the creator-defined space, and have independent raters (including DHH viewers) judge whether each output stays within the anchor bounds and preserves the sound event's meaning. The paper's own reported failures—[Ice freezing sound] becoming [Chill ice cracking], and identical [Dolphin whistles] captions transformed differently in different scenes—are concrete starting points; if bound violations or meaning shifts are frequent even at moderate settings, the claim","tokens_in":23334,"feed_emoji":"🎬","tokens_out":9332,"duration_ms":93718,"temperature":0.7,"pith_summary":"The paper sets out to replace the \"one-size-fits-all\" non-speech caption track with a dynamic one: captions that deaf and hard-of-hearing viewers can tune to their own needs without overriding what the caption author intended. CapTune does this with two tools—a Creator Tool in which authors mark a two-dimensional safe space for transformations, bounded by anchor captions along Level of Detail and Expressiveness, and a Viewer Client in which each viewer picks settings inside that space, plus a sound-representation style and optional genre alignment. The rewriting is done by GPT-4o, prompted with interpolation ratios that place the viewer's choice relative to the anchors and the current caption, together with audio-visual scene context extracted by a video language model. Evaluations with seven caption creators and twelve DHH participants found that creators felt they kept editorial control and viewers reported greater emotional and narrative engagement; the study also surfaced tensions between information richness and cognitive load, between interpretive and descriptive sound language, and the strongly context-dependent nature of caption preferences.","feed_headline":"Tune sound captions per viewer, within creator-set bounds","feed_subtitle":"Deaf and hard-of-hearing viewers adjust detail, expressiveness, and sound style; creators keep editorial control.","key_machinery":"The anchored transformation space: a two-dimensional space whose axes are Level of Detail and Expressiveness (each a 1–10 semantic scale), with its bounds set by creator-chosen lower and upper anchor captions. Transformation requests are converted into two ratios per axis—r (Eq. 2), the requested setting's position between the anchors, and δ (Eq. 3), the signed magnitude of change from the current caption—embedded in a structured prompt that instructs GPT-4o to interpolate among the original caption, the lower-anchor caption, and the upper-anchor caption, using audio-visual scene descriptions extracted by VideoLLaMA2 as context. This interpolation-between-examples mechanism is what keeps gen","core_discovery":"CapTune's central claim is that the text of non-speech captions—not just their visual styling—can be safely personalized by anchoring generative transformations to creator-defined boundaries. A caption author sets two anchor points: a lower anchor for the most minimal acceptable captions and an upper anchor for the most elaborate ones, on the Level of Detail and Expressiveness axes. The Viewer Client exposes only cells inside this space; when a viewer picks a setting, the system computes an interpolation ratio (Eq. 2) locating the choice between the two anchors and a change ratio (Eq. 3) measuring the shift from the current caption, using both to instruct GPT-4o to interpolate between the an","pith_inferences":["The two-anchor, ratio-driven prompting recipe is domain-neutral: the same \"human sets bounds with concrete examples, model interpolates between them\" pattern could extend to audio description, simplified subtitles for language learners, or other constrained rewriting tasks—though the paper does not claim this.","The reported inconsistency of identical sound sources (e.g., dolphin whistles) across scenes hints at a testable invariant: at fixed parameter settings, a given sound event should yield the same transformed caption regardless of context; adding such a consistency check would strengthen the pipeline.","The interpretive-versus-descriptive tension participants voiced suggests a fifth axis or an explicit marking of inferred content (e.g., \"warm purr\" as interpretation), which the paper mentions as future work only in passing.","Short clips (2–8 minutes) leave open whether emotional-engagement gains persist across feature-length content, where fatigue and cross-scene narrative consistency become dominant factors."],"forward_implications":["Non-speech captions can be treated as co-authored media: the creator sets the safe range, the viewer personalizes within it, and the language model does the rewriting—offering a template for accessibility content that is neither fixed nor unconstrained.","Caption-customization interfaces can be built around interpolation between concrete anchor examples rather than abstract style rules, which creators in the study found intuitive.","Viewer preference is context-dependent (genre, scene pacing, viewing intent), so future systems should support scene-level or context-aware caption adaptation rather than a single global setting.","Design requirements follow directly: user profiles that retain preferences, preview-and-compare views, explainable transformation logic, and vocabulary simplification for ASL-first or non-native-English viewers.","The four-parameter scheme makes explicit a core trade-off: richer, more expressive captions heighten emotional engagement but raise cognitive load and can crowd out the viewer's own interpretation."],"supporting_citations":[{"why":"Alonzo et al.'s captioning-and-visualization system for non-speech sounds is the primary prior approach CapTune extends beyond visual styling.","marker":"[13]"},{"why":"Caption Royale's typographic affective captions motivate text-level personalization while leaving caption wording itself unchanged.","marker":"[21]"},{"why":"May et al.'s critical design of NSI caption augmentation frames the push toward richer textual presentation of sounds.","marker":"[49]"},{"why":"The Stranger Things captions interview supplies the running example of expressive, genre-aligned NSI that viewers cite as aspirational.","marker":"[58]"},{"why":"Kumar et al.'s constrained text generation is the framework the anchored transformation model adapts into accessibility.","marker":"[41]"},{"why":"Rosen's study of sound in American Deaf literature grounds the claim that DHH viewers perceive sound through alternate modalities.","marker":"[57]"},{"why":"One of the Reddit threads analyzed in Section 3; it supplies first-person evidence that DHH caption preferences vary (concise versus detailed).","marker":"[29]"},{"why":"Guest et al.'s Applied Thematic Analysis is the coding method used to analyze both the Reddit corpus and the two evaluations.","marker":"[30]"}],"fun_headline_variants":["Anchored AI lets Deaf viewers tailor sound captions","Caption personalization within creator-set guardrails","Per-viewer sound captions, bounded by creators","Custom captions: viewers pick detail, creators set limits","Safe caption tuning: anchors preserve creator intent"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The entire system rests on one assumption: that GPT-4o, given the anchor captions, the interpolation ratios, and the scene context, actually rewrites each caption so it stays within the creator's bounds and preserves the meaning of the original sound event—and the paper concedes in its limitations that outputs can be inconsistent or semantically inaccurate.","fun_headline_variants_meta":{"raw":{"variants":["Anchored AI lets Deaf viewers tailor sound captions","Caption personalization within creator-set guardrails","Per-viewer sound captions, bounded by creators","Custom captions: viewers pick detail, creators set limits","Safe caption tuning: anchors preserve creator intent"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000126,"raw_usage":{"total_tokens":908,"prompt_tokens":668,"completion_tokens":240,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":412,"completion_tokens_details":{"reasoning_tokens":180}},"tokens_in":412,"tokens_out":240,"duration_ms":3077,"temperature":1.0,"reasoning_tokens":180,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T15:18:48.838139+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A systematic accuracy and consistency audit: take a fixed caption set across genres, apply the transformation at every grid point in the creator-defined space, and have independent raters (including DHH viewers) judge whether each output stays within the anchor bounds and preserves the sound event's meaning. The paper's own reported failures—[Ice freezing sound] becoming [Chill ice cracking], and identical [Dolphin whistles] captions transformed differently in different scenes—are concrete starting points; if bound violations or meaning shifts are frequent even at moderate settings, the claim","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Alonzo et al.'s captioning-and-visualization system for non-speech sounds is the primary prior approach CapTune extends beyond visual styling."},{"cited_title":"LEE, DEBORAH I","cited_arxiv_id":null,"evidence_quote":"May et al.'s critical design of NSI caption augmentation frames the push toward richer textual presentation of sounds."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Rosen's study of sound in American Deaf literature grounds the claim that DHH viewers perceive sound through alternate modalities."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"One of the Reddit threads analyzed in Section 3; it supplies first-person evidence that DHH caption preferences vary (concise versus detailed)."}],"review_version":1}