{"id":"a665d902-1993-4e04-bd95-e68b0fd4d607","arxiv_id":"2411.17015","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Trinity, a hybrid mobile-plus-PC presentation support system, generates synchronized verbal, nonverbal, and visual delivery prompts and was rated significantly more helpful than two baselines in a user study.","lead":"Trinity is a smartphone-and-laptop system that helps English-as-a-foreign-language students deliver academic presentations by synchronizing their spoken words, body language, and slides in real time. A controlled study with EFL presenters found the system was rated more helpful than two existing presentation-support tools, though it also increased cognitive load.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract's 'without excessive cognitive load' claim is contradicted by reported data: Trinity showed significantly higher cognitive load (M=5.68, H=12.365, p=0.004) than both baselines.","rationale":"The reader's weakest_assumption was audience blindness, which is a genuine and honestly acknowledged threat to the comparative helpfulness claim. I agree this is important. However, the most load-bearing concern about the central claim as stated is the direct contradiction between the abstract's 'without excessive cognitive load' and the paper's own significant cognitive-load results. This is an internal inconsistency, not a speculation about unmeasured bias. The reported statistics show Trinity had significantly higher cognitive load and workload than both baselines, with a mean of 5.68 on a 7-point scale. Unless 'excessive' is defined with a threshold that accommodates this value, the claim is unsupported. The participant-count mismatch (33 in the abstract vs 62 in §5.2) compounds the concern about reporting accuracy, but the cognitive-load contradiction alone is enough to require revision. The core helpfulness finding may survive, so I do not recommend rejection; a conditional acceptance with mandatory correction of the abstract, clarification of the cognitive-load threshold, and release of anonymized data is appropriate. This is consistent with the reader's CONDITIONAL verdict, hence I mark the verdict as CONDITIONAL rather than UNCHANGED, though the difference is one of emphasis rather than outcome.","tokens_in":33478,"tokens_out":4570,"duration_ms":43045,"concrete_test":"Request the raw questionnaire data and re-run the cognitive-load and workload analyses with a pre-registered non-inferiority margin (e.g., Trinity mean within +0.5 of the higher baseline on the 7-point scale). If the 95% CI for the difference exceeds the margin, the abstract's 'without excessive cognitive load' statement is falsified and must be removed or qualified.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim in the abstract is that Trinity supports AOP delivery 'without excessive cognitive load.' The paper's own results in §6.2 iii contradict this: Kruskal-Wallis tests show significant differences in workload (H=9.395, p=0.010), cognitive load (H=12.365, p=0.004), and self-perceived performance (H=11.545, p=0.010). Post-hoc tests indicate Trinity incurs significantly higher cognitive load (M=5.68, SD=1.24; p=0.036 vs IntelliPrompter, p=0.007 vs OfficeRemote) and higher workload (M=4.54, SD=0.83; p=0.025, p=0.034). The authors explicitly state that 'the baselines impose less workload and cognitive load as they only involve script reading without additional operations.' Thus, the abstract's 'without excessive cognitive load' is not supported by the reported statistics. No definition of 'excessive' is provided, and a mean of 5.68 on a 7-point scale is well above the midpoint. This is not a minor wording issue: it is a direct internal inconsistency in the headline finding. A secondary reporting inconsistency—abstract states 33 presenters and 21 audience members while §5.2 reports 62 presenters and 25 audience members—further weakens confidence in the data presentation, but the cognitive-load contradiction is the more load-bearing issue because it falsifies an explicit part of the central claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents Trinity, a hybrid mobile-centric system that supports EFL students' academic oral presentations by synchronizing verbal, nonverbal, and visual delivery channels. The system consists of a PowerPoint add-in that refines scripts and generates customizable delivery prompts using GPT-4, and a smartphone app that provides remote slide control, speech pace modulation, and integrated emoji-based delivery prompts. The authors report a formative study (survey, design study, and expert interview) and a controlled between-subject user study comparing Trinity with IntelliPrompter and OfficeRemote, with presenter self-reports, audience ratings, interaction logs, and interviews. The abstract claims that Trinity effectively supports AOP delivery and is perceived as significantly more helpful than baselines, without excessive cognitive load.","tokens_in":33737,"tokens_out":2818,"duration_ms":26753,"significance":"If the results hold, Trinity makes a useful contribution to presentation-support systems for EFL students by targeting the synchronization of multiple communication channels in real time, an aspect largely underexplored compared with single-channel training or teleprompter-style tools. The formative study is carefully conducted, and the design goals are grounded in multiple stakeholder perspectives. The paper also transparently reports several limitations, including the difficulty of blinding the audience and the challenges of technical stability. However, the headline claims are currently weakened by internal inconsistencies between the abstract and the reported statistics, an unmet definition of 'excessive cognitive load,' and overstatements of pairwise comparison results.","major_comments":[{"comment":"The abstract and Section 5.2 report inconsistent participant numbers. The abstract (and Section 1) state the user study involved 33 EFL student presenters and 21 audience members, while Section 5.2 reports 65 recruited presenters, 62 after dropouts (21 Trinity, 21 IntelliPrompter, 20 OfficeRemote), and 25 audience members. This discrepancy undermines confidence in the data reporting and needs to be resolved, as the statistics in Section 6 use the 62-presenter sample.","section":"Abstract and §6.2 iii"},{"comment":"Given that the central effectiveness claim relies heavily on audience ratings, the admitted violation of audience blindness in Section 5.5 is a load-bearing confound. The authors acknowledge that 'the audience may still infer some condition-related information from presenters' usage patterns (i.e., how they interact with devices), potentially influencing their ratings.' Since presenters using Trinity physically interact with a phone, while IntelliPrompter users interact with a laptop, audience members could plausibly identify the condition and bias their ratings of eye contact, gestures, and composure. The paper should either provide evidence that this did not occur (e.g., a manipulation check or analysis of ratings by audience members' awareness) or substantially temper the causal claims drawn from the audience data.","section":"§6.1 ii"}],"minor_comments":[{"comment":"There is an inconsistency in the description of the dropout: the text says 'After three dropouts, Trinity and IntelliPrompter conditions had 21 presenters, and OfficeRemote conditions had 20,' but the initial allocation is described as 'three groups of 22 for the Trinity and IntelliPrompter conditions, and 21 for the OfficeRemote condition'; a three-person dropout from a 22/22/21 split should yield 21/21/20 only if the dropout came from the OfficeRemote group and one from each of the other groups, which should be clarified.","section":"§5.2"},{"comment":"The phrase 'ease-to-use' is a typo for 'ease of use,' and the quote from S10 contains a grammatical error ('how to it works') that should be corrected.","section":"§6.1 i"},{"comment":"The questionnaire table lists 'How satisfing was the system?' which should read 'How satisfying was the system?'","section":"Table 6"},{"comment":"The description of conditions uses the phrase 'without𝑄&𝐴 sessions' with a mathematical symbol; this appears to be a formatting error and should read 'without Q&A sessions.'","section":"§5.1"},{"comment":"The pairwise p-values are inconsistently formatted (e.g., some use 'p_TI' and others use 'p<0.1'); standardizing the notation and explicitly listing which comparisons are non-significant would improve readability and accuracy.","section":"§6.1 ii"}],"recommendation":"major_revision","confidential_remarks":"The paper has a solid design study and system implementation, but the abstract overstates the empirical findings in ways that the paper's own statistics contradict. The sample-size discrepancy and the acknowledged audience-blindness confound further weaken the central claims. I believe the issues are fixable with a careful revision that aligns the abstract with the data, redefines or removes the 'without excessive cognitive load' claim, and presents pairwise comparisons with appropriate hedging. I would not recommend rejection, but the revision needs to be substantive rather than cosmetic."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Trinity is a genuinely integrated system: it combines mobile slide control, LLM-based script polishing, and emoji-coded delivery prompts into one on-the-fly workflow, and I don't know of prior work that synchronizes all three channels this way. The formative work is a real strength—survey, design study, expert interviews, and the design goals follow from the findings. The system itself is thoughtfully built, and the qualitative results (audience observations, presenter interviews) give plausible evidence that the multichannel support changes delivery.\n\nThe soft spots are real but mostly fixable. The participant-count discrepancy is glaring: the abstract and introduction say 33 presenters and 21 audience members, while §5.2 reports 65 recruited, 62 completing, and 25 audience. That needs a correction, not an explanation. More serious is the cognitive-load contradiction. The abstract claims the system works 'without excessive cognitive load,' but the paper's own Kruskal-Wallis results show Trinity had significantly higher cognitive load than both baselines (M=5.68 on a 7-point scale, p=0.036 vs IntelliPrompter, p=0.007 vs OfficeRemote). The authors even state the baselines impose less load because they only involve script reading. That is an internal inconsistency in the headline finding, and 'excessive' is never defined. The claim that Trinity 'outperformed both baselines in all aspects' is also overbroad: several pairwise comparisons are nonsignificant (eye contact vs IntelliPrompter p=0.18, visual control vs OfficeRemote p=0.316). The audience-blindness limitation is honestly acknowledged, and it is a genuine threat to the comparative ratings—but the authors flag it themselves, and the presenter self-report data are less affected.\n\nOn balance, the central design contribution holds up. The evaluation supports the conclusion that Trinity is helpful and that presenters want to use it; the load data complicate the story but don't sink the system. What the current write-up needs is an accurate abstract, a transparent discussion of the cognitive-load tradeoff, and a release of the prototype and anonymized data. This is the kind of paper a serious referee can work with.\n\nMy take: send it to review, but the reviewers should push hard on the quantitative reporting before publication. I'd cite the system design and formative study in related work; I would not cite the evaluation as evidence until the numbers are corrected.","headline":"A well-built integrated presentation-support system with a thorough formative study, but the reported evaluation contains internal contradictions (participant counts, the 'without excessive cognitive load' claim) that need fixing before the paper's central claims can be trusted.","tokens_in":34320,"tokens_out":854,"would_cite":true,"duration_ms":10369,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Trinity, a mobile-plus-AI system, claims to improve EFL students' academic oral presentations by synchronizing verbal, nonverbal, and visual delivery cues on the fly.","keywords":["Multichannel communication","Academic oral presentation","Delivery support","EFL students","On-the-fly feedback","Large language models","Mobile prompter","User study"],"falsifier":"Run the same three-condition comparison with audience members who cannot see the presenter's device—for example, watching through a one-way mirror or hearing only audio with slides—and check whether Trinity still receives significantly higher ratings. If the advantage disappears, the perceived helpfulness is an artifact of visible device use.","tokens_in":33228,"feed_emoji":"🎤","tokens_out":5193,"duration_ms":49385,"temperature":0.7,"pith_summary":"This paper argues that English-as-a-foreign-language students give monotonous, disjointed academic oral presentations because their verbal, nonverbal, and visual channels are not coordinated, and that a hybrid support system can fix this during the live talk. It introduces Trinity, a PowerPoint add-in paired with a smartphone prompter: GPT-4 polishes the script and inserts delivery cues, while the app tracks speech pace, controls slides remotely, and displays emoji prompts for gestures, eye contact, volume, and facial expression. A controlled between-subject study with 33 presenters and 21 audience members reports that Trinity was perceived as significantly more helpful than two baselines and improved audience ratings on eye contact, gesture, vocal variety, slide effectiveness, and channel consistency. The paper also reports that workload and cognitive-load scores were higher for Trinity, though the abstract characterizes the added load as not excessive.","feed_headline":"Trinity syncs voice, body, and slides for better talks","feed_subtitle":"In a user study, presenters rated it significantly more helpful than PC-only or phone-only baselines.","key_machinery":"The load-bearing object is Trinity's augmented prompter: a smartphone app that holds the LLM-polished script, the emoji-encoded delivery prompts, and the remote slide controls, while a server connects it to a PowerPoint add-in. The synchronization mechanism is a dual-layered speech-pace display (global progress bars plus sentence-level underpainting), an emoji lookup table that converts GPT-4's textual prompts into glanceable nonverbal cues, and a speech-recognition pipeline using BM25 string matching to scroll the script in step with the speaker's words. Together these let one device carry verbal guidance, nonverbal reminders, and visual control so the presenter can move, gesture, and make eye contact instead of staying anchored to a laptop.","core_discovery":"The central claim is that on-the-fly synchronization of the three communication channels is what makes delivery support effective, and that a mobile-centric system can deliver that synchronization in real time. Trinity combines an LLM-refined script with live prompting: in-line emoji cues encode verbal and nonverbal modulations, progress bars and underpainting regulate speech pace, and thumbnails plus tapping let the presenter drive slides from the phone. In the user study, Trinity significantly outperformed IntelliPrompter and OfficeRemote on perceived helpfulness in supporting delivery ($H = 9.471$, $p = 0.007$) and on likelihood of future use ($H = 10.382$, $p = 0.005$), and audience ratings significantly favored Trinity on composure, gesture, eye contact, facial expression, vocal pitch, speech rate, volume, slide effectiveness, and speech-behavior-visual consistency. The paper presents this as evidence that integrated multichannel guidance, not just script reading or slide navigation, drives perceived presentation quality.","pith_inferences":["Editor's inference: if the helpfulness advantage survives a fully blinded replication, the same three-channel synchronization design could generalize to conference talks, teaching, or public-speaking coaching, where the core problem is also coordinating voice, body, and slides in real time.","Editor's inference: the emoji-prompt scheme offers a cheap, testable way to transfer expert delivery advice into live cues; a natural next experiment would remove the prompts while keeping the polished script to isolate which component caused the audience-visible gains.","Editor's inference: the higher cognitive-load scores suggest an adaptive prompt-density mechanism—fewer cues for familiar content or later presentations—could preserve Trinity's benefits while lowering the reported load."],"forward_implications":["Presenters using Trinity were rated by audiences as significantly better on eye contact, facial expression, composure, gesture, vocal pitch, speech rate, volume, and consistency among speech, behavior, and slides.","In-line emoji prompts can cue nonverbal behavior without raising attentional load beyond what script reading already requires.","LLM-based script polishing can save preparation time and improve fluency, but presenters need to review and edit the output to keep their own speech style.","System malfunctions and speech-recognition jitters sharply reduce trust and increase cognitive load, so robustness is a prerequisite for any on-the-fly delivery support.","Moving delivery support to a smartphone frees presenters from the lectern and appears to encourage gestures and audience engagement that PC-only tools do not."],"supporting_citations":[{"why":"IntelliPrompter is the PC-based speech-tracking baseline that Trinity is compared against in the user study.","marker":"[5]"},{"why":"Rhema supplies the prior real-time pace-and-volume cue approach that motivates on-the-fly support but covers only limited channels.","marker":"[79]"},{"why":"This standardized rubric is used by audience members to rate verbal, nonverbal, visual, consistency, and content dimensions.","marker":"[66]"},{"why":"This work informs the questionnaire and weighted-importance measurement methods used in the formative study and user study.","marker":"[55]"},{"why":"NASA-TLX is the workload measure used to test the claim that Trinity does not impose excessive cognitive load.","marker":"[34]"},{"why":"BM25 is the string-matching algorithm used to track speech progress and synchronize script scrolling on the phone.","marker":"[72]"},{"why":"This source supplies the evidence that large language models can understand and generate human language, justifying the GPT-4-based script refinement and prompt generation.","marker":"[96]"}],"fun_headline_variants":["Trinity aligns speech, gestures, and slides in real time","Study: Trinity's integrated prompts beat script readers and slide remotes","Mobile app syncs voice, body, and slides for EFL talks","Trinity: one app to sync your talk's voice, body, and visuals","Real-time sync of voice, gestures, and slides enhances presentations"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The main results assume audience members could not tell which support tool a presenter was using; the paper admits that presenters' visible phone use may have revealed the condition and biased ratings.","fun_headline_variants_meta":{"raw":{"variants":["Trinity aligns speech, gestures, and slides in real time","Study: Trinity's integrated prompts beat script readers and slide remotes","Mobile app syncs voice, body, and slides for EFL talks","Trinity: one app to sync your talk's voice, body, and visuals","Real-time sync of voice, gestures, and slides enhances presentations"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000813,"raw_usage":{"total_tokens":3557,"prompt_tokens":929,"completion_tokens":2628,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":545,"completion_tokens_details":{"reasoning_tokens":2536}},"tokens_in":545,"tokens_out":2628,"duration_ms":17681,"temperature":1.0,"reasoning_tokens":2536,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T12:36:33.513816+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same three-condition comparison with audience members who cannot see the presenter's device—for example, watching through a one-way mirror or hearing only audio with slides—and check whether Trinity still receives significantly higher ratings. If the advantage disappears, the perceived helpfulness is an artifact of visible device use.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Rhema supplies the prior real-time pace-and-volume cue approach that motivates on-the-fly support but covers only limited channels."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"This standardized rubric is used by audience members to rate verbal, nonverbal, visual, consistency, and content dimensions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"This work informs the questionnaire and weighted-importance measurement methods used in the formative study and user study."},{"cited_title":"Does Conceptual Representation Require Embodiment? Insights From Large Language Models","cited_arxiv_id":"2305.19103","evidence_quote":"This source supplies the evidence that large language models can understand and generate human language, justifying the GPT-4-based script refinement and prompt generation."}],"review_version":1}