{"id":"9e4023f9-c90d-450f-a835-9d42fbc478ae","arxiv_id":"2504.20365","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"An n=297 online experiment found that AI text presentation speed and order affect perceived comfort, quality, humanness, and trustworthiness, with medium speed rated best.","lead":"AI writing tools that stream text at different speeds change what users think of the tool and its output. A medium streaming speed was rated most comfortable and highest quality, while backwards or random text lowered trust and perceived quality.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Repeated-measures data are analyzed as independent observations: the reported df (4,2965) ignore participant clustering, so the 'medium is best' quality and comfort comparisons need a mixed-model reanalysis.","rationale":"The reader's weakest assumption (imagined versus real writing) is acknowledged in Sec. 6.3 and primarily affects external validity. The more load-bearing issue is internal: the statistical model treats repeated measures as independent. Sec. 4.2 gives N = 297 and 10 observations per participant, while Sec. 4.3.1 reports df of 4,2965 for style, which is consistent only with a between-subjects one-way ANOVA on 2970 trials. No Error(participant) term or mixed model is reported. This is not a stylistic preference; it affects the p-values used to accept H1, H2, H3, and H5. Because H2's effect is small (eta^2_p = 0.02) and the medium-fast gap is only 0.14 scale points, a proper subject-level error term could change pairwise conclusions even if the omnibus test remains significant. The ART robustness check also requires a within-subject error structure to be valid, and the reported df indicate the parametric ANOVA is the primary analysis. I therefore recommend retaining the CONDITIONAL verdict: the qualitative finding is plausible and the large comfort and humanness effects are probably robust, but the paper should be required to supply the mixed-model reanalysis and data before the perceived-quality claim is accepted. This concern is about the reported error term, not about author integrity or intent.","tokens_in":28824,"tokens_out":7936,"duration_ms":93145,"concrete_test":"Obtain the trial-level data and reanalyze with a linear mixed model: fixed effects for text style, genre, and their interaction; random intercept for participant (with random slopes for style if identifiable); Satterthwaite degrees of freedom. Apply the same Bonferroni-corrected threshold (p < 0.00625) to Tukey-adjusted contrasts for medium versus fast, medium versus slow, and forward versus non-forward styles on the H1, H2, H3, and H5 outcomes. If the medium-versus-fast contrast for perceived quality (H2) loses significance under this model, the abstract's quality claim should be downgraded; if all key contrasts remain significant, the central finding survives with corrected statistics.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Sec. 4.1 describes a fully within-subjects design: each of the 297 participants saw all five text appearance styles in both genres, contributing 10 trials. Yet Sec. 4.3.1 reports 'independent multi-way ANOVAs' with error df of 4,2965 for style. That df equals 2970 total trials minus 1 minus 4, which treats each trial as an independent observation; a repeated-measures analysis with participant as a random effect would have roughly 296 × 4 = 1184 error df for the within-subject style factor. Ignoring this clustering inflates the reported evidence, and the problem matters most for the smallest headline effect: perceived quality (H2, eta^2_p = 0.02, medium-minus-fast mean gap only 0.14 on a -2 to 2 scale). The large comfort and humanness effects (F = 170.61 and F = 57.97) would likely survive a proper analysis, so I am not claiming the qualitative finding is fabricated; but the exact p-values, post-hoc letter groupings, and the abstract's 'perceived quality' claim are not trustworthy as reported until the data are reanalyzed with subject-level error.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper examines how the presentation style of AI-generated text (speed and character order) affects users' perceptions of an AI writing tool. In an online study with 297 participants, each participant saw five text appearance conditions (slow, medium, fast, backwards, random) in two genres (creative, professional) and rated reading comfort, perceived text quality, humanness, trust, respect, liking, and intention to use. The authors report that medium speed is most comfortable and yields the highest perceived quality, slow and medium are perceived as most human-like, forward presentation is more trustworthy than backwards or random, and that genre shows no consistent interaction. The paper includes qualitative thematic analysis and an exploratory appendix.","tokens_in":29052,"tokens_out":5727,"duration_ms":61830,"significance":"If the reported effects are valid, this work makes a meaningful contribution to HCI and AI interface design by demonstrating that a seemingly cosmetic design choice—the speed and order of streamed AI text—can change users' judgments of output quality and trustworthiness. The study's strengths include pre-registered-style hypothesis finalization one month before data collection, fixed texts, randomized speed-text pairs, a within-subjects design, Bonferroni correction, and an ART robustness check. The large comfort effect (η²p = 0.19) is substantial and the qualitative findings are rich. However, a serious statistical flaw—the analysis treats repeated-measures data as independent observations—undermines the reported p-values and post-hoc groupings, particularly for the small effects such as perceived quality. The central qualitative narrative is likely defensible, but the quantitative evidence needs to be re-analyzed before the paper can be accepted.","major_comments":[{"comment":"The statistical analysis treats each of the 2970 trials (297 participants × 10 within-subject trials) as independent observations. The reported error df of 4,2965 for presentation style corresponds to N = 2970 independent observations, not to the repeated-measures structure described in Section 4.1, where every participant saw all five styles in both genres. This inflates the test statistics and invalidates the exact p-values, effect sizes, and Tukey post-hoc letter groupings. This is load-bearing for the paper's smaller effects, especially H2 (perceived quality: η²p = 0.02, medium-minus-fast mean gap 0.14 on a -2 to 2 scale) and for H4–H8. Please re-analyze with a mixed-effects model (participant as a random effect, with scenario/genre as appropriate) or a repeated-measures ANOVA, and report corrected F, p, effect sizes, and post-hoc comparisons. The Aligned Rank Transform robustness check must also account for the within-subjects design. The large comfort and humanness effects may survive, but the exact numbers and the abstract's claim about perceived quality are not trustworthy as reported.","section":"Section 4.3.1 / Section 5.1–5.4"},{"comment":"The statement that medium presentation 'resulted in the highest perceived quality' is a post-hoc finding from pairwise comparisons after a null-hypothesis ANOVA (H2 predicted no effect). Once the repeated-measures analysis is correctly performed, the post-hoc comparisons must be recomputed with the proper error terms. The abstract and conclusion should reflect the corrected results, and any claim that speed is 'correlated with' quality should be expressed as a perceptual effect conditional on the corrected analysis.","section":"Section 5.2 (H2) and Section 7"},{"comment":"The hypotheses H4, H6, and H8 explicitly involve genre-by-presentation interactions (e.g., fast for professional, slow/medium for creative), yet the results sections do not report the interaction term's F and p values. The conclusion states 'we do not find evidence of a consistent interaction between text appearance speed and genre system,' but the basis for this claim is not presented. Please report the interaction F, df, and p for each dependent variable, or explicitly state that interactions were not tested and adjust the interpretation accordingly.","section":"Sections 3.3 and 5.4"}],"minor_comments":[{"comment":"The abstract says 'speed is correlated with perceived humanness and trustworthiness of the AI tool, as well as the perceived quality of the generated text.' This is imprecise because the manipulated factor is presentation style, which includes character order (backwards, random) as well as speed; 'correlated' is also weaker than the experimental design allows—suggest 'affected' or 'influenced'.","section":"Abstract"},{"comment":"Medium is described as 'slightly faster than average reading speed,' but the cited reference in Section 2.2 gives average reading speeds of 200–400 wpm; 600 wpm is 1.5 to 3 times that range. Please revise the justification for the medium anchor.","section":"Section 3.1.1 (Medium)"},{"comment":"The description 'random insertion via insertion-sort' is unclear, because insertion sort is a deterministic ordering algorithm, not a random insertion process. Please clarify the exact character placement procedure.","section":"Section 3.1.1 (Random)"},{"comment":"The phrase 'independent multi-way ANOVAs' is ambiguous and could be read as 'independent-observations ANOVAs,' which is precisely the problem described above. Consider renaming to 'separate ANOVAs' and explicitly noting the need for a repeated-measures approach.","section":"Section 4.3.1"},{"comment":"For the genre rows in H2, H4, H5, and H8, significant p-values are shown as '***' but no compact letter display is provided; the text says post-hoc comparisons were conducted for significant ANOVAs. Please add the corresponding letter groupings or explain why they are omitted.","section":"Table 2"},{"comment":"The figure contains placeholder text such as '/gid00035' that appears to be a rendering artifact; please replace these strings with proper textual labels.","section":"Figure 2"},{"comment":"The limitation about 'imagining writing does not create the same experience as writing' is appropriately acknowledged. Consider explicitly tying this to the quantitative results—e.g., effects might differ when users are actively composing rather than reading prefilled text.","section":"Section 6.3"}],"recommendation":"major_revision","confidential_remarks":"The core finding—that text presentation style affects perceptions of AI writing tools—is interesting and likely correct in direction, given the very large comfort effect and the broad qualitative support. The main obstruction is the statistical analysis: the reported ANOVAs ignore the within-subjects structure, so the exact p-values and post-hoc groupings are not valid as reported. This is fixable with a re-analysis, so I am not recommending rejection. I would also encourage the editor to ensure that the revised manuscript explicitly reports interaction effects for the genre hypotheses and that the abstract's quality claim is calibrated to the corrected analysis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dave—the core of this paper is new and worth your time: a within-subjects comparison of five text presentation styles (slow, medium, fast, backwards, random) in AI co-writing, crossing two genres with 297 people. The design is genuinely careful: fixed texts, pre-specified hypotheses, randomized speed-text pairings, Bonferroni correction, and an ART robustness check. The qualitative analysis is also thoughtful, showing how people read along with the stream and attribute humanness, thoughtfulness, and even creativity to the presentation rather than the content. That central idea—streaming speed is a design variable that shapes trust and perceived quality—is plausible and well motivated by prior work on speech rate and typing.\n\nNow the soft spots, in proportion. The stress-test note is correct and matters. Section 4.1 describes a fully within-subjects design: each participant saw all five styles in both genres, so 10 trials per person. Yet the reported df, 4,2965, treat those 2,970 trials as independent observations. A proper repeated-measures analysis would use subject-level error, giving roughly 1,184 df for the within-subject style factor. Ignoring the clustering inflates the F statistics and p-values. The large effects (comfort F=170.6, humanness F=58.0) will almost certainly survive a mixed-model reanalysis, so I do not think the qualitative story is fabricated. But the smallest headline effect—perceived quality, eta-squared 0.02, medium-minus-fast gap of only 0.14 on a -2 to 2 scale—might not survive. That is exactly the claim the abstract leads with, and it is not trustworthy as reported until the data are reanalyzed with participant as a random effect. The post-hoc letter groupings and the 'correlated' language in the abstract also need to be reconsidered; the pattern is non-monotonic (medium best, fast and slow similar), not a simple correlation.\n\nThe other issues are minor. The authors acknowledge that participants imagined writing rather than actually writing, which limits generalizability, and they do that honestly. They do not provide data or analysis scripts, which makes the reanalysis request harder to evaluate. Single-item measures for comfort and quality are weak but acceptable for a first study. Citation pattern looks fine; no self-citation inflation.\n\nWho is this for? Researchers building LLM writing interfaces or studying human-AI perception. It deserves a serious referee, not a desk reject, because the design space is new and the central phenomenon is likely real. My recommendation: send it to review with a clear request for a repeated-measures reanalysis and for the data and scripts to be posted. The revision may shrink the quality claim, but the paper will still be a useful contribution. If the reanalysis kills the quality effect, that is a publishable finding too—just a narrower one.","headline":"A well-designed study of how text streaming speed shapes perceptions of AI writing tools, but the analysis ignores the within-subjects structure and the small 'perceived quality' effect needs a reanalysis before that claim is taken seriously.","tokens_in":29573,"tokens_out":1617,"would_cite":true,"duration_ms":19893,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI text speed shifts users' trust and quality judgments.","keywords":["AI writing tools","text presentation speed","text streaming","perceived quality","trustworthiness","anthropomorphism","creative writing","user study"],"falsifier":"A replication in which participants actually write their own text and receive unique AI completions, with the same speed conditions, would falsify the generalization if the medium-speed advantage on quality and trust disappears. A simpler check: have readers rate the final complete texts without seeing the animation; if speed-blind ratings still differ by condition, the effect would be in the content, not the presentation.","tokens_in":28630,"feed_emoji":"✍️","tokens_out":4488,"duration_ms":42751,"temperature":0.7,"pith_summary":"This paper asks whether the way AI-generated text appears on screen—typed out slowly, quickly, in reverse, or in random order—changes how readers judge the tool and its output. Across 297 online participants imagining themselves co-writing in creative and professional scenarios, the authors find that medium-speed presentation (600 words per minute, close to reading speed) is the most comfortable and earns the highest perceived text quality. Slower and medium speeds make the AI seem more human-like, while backwards and random presentations are rated less trustworthy. The authors conclude that text appearance speed is not a neutral interface detail: it influences perceptions of quality and trust, which can shape whether users accept or reject AI-generated text.","feed_headline":"AI typing speed shifts users' trust and quality judgments","feed_subtitle":"Medium-speed AI text came across as most human and trustworthy to 297 readers.","key_machinery":"The central object is the text presentation style, operationalized as five within-subject conditions: slow (160 wpm), medium (600 wpm, near average reading speed), fast (6,000 wpm, approximating large language model token generation), backwards (characters appear in reverse), and random (characters inserted in random order), with the content held fixed. These conditions isolate the perceptual contribution of appearance speed and order from the meaning of the text. The mechanism is that users read along with the appearing text, so the pace and order either match or disrupt their reading process, producing comfort or discomfort and cueing human-like versus machine-like attributions.","core_discovery":"The central discovery is that the speed and order of text appearance affect users' perceptions of an AI writing tool independently of the text content. In the experiment, the identical text presented at medium speed was rated highest on comfort and quality; slow and medium speeds were perceived as more human-like; and the two deliberately non-anthropomorphic styles—backwards and random character order—were rated as less trustworthy. The authors report that users read along with the generation, attribute human-like qualities such as thoughtfulness to slower appearance, and show divided preferences tied to their writing values rather than a consistent genre effect. Their conclusion is that interface presentation decisions influence judgments of system quality and trustworthiness, and thereby influence how generated text is used.","pith_inferences":["The authors' setup fixes the text content; an untested extension is whether the same perceptual effects appear when users write their own text and receive personalized completions, where involvement and ownership may override presentation cues.","If the trust effect holds in real use, one testable design response is to decouple display speed from model computation speed and let users set the pace, which would break the current coupling between latency optimization and perception.","The findings suggest a concrete hypothesis for neighboring domains: adding human-like pauses or backspacing to AI output, as some tools already do, may increase perceived thoughtfulness and thereby raise acceptance rates of fallible content.","A direct follow-up prediction: blind evaluators who rate the same final texts without seeing the animation should show no speed effect, confirming that the effect lives in presentation rather than content."],"forward_implications":["Medium-speed text appearance, near average reading speed, produces the highest reading comfort and perceived text quality; designers who want favorable quality judgments should not default to maximum speed.","Backwards and random text presentation reduce perceived trustworthiness and quality, so non-anthropomorphic streaming styles carry a perception cost even when the final text is identical.","Slow and medium speeds make the AI seem more human-like; users may accept more output from tools that appear thoughtful, a consequence the authors flag as potentially unintentional manipulation.","Text presentation effects on comfort, quality, humanness, and trust were not substantially moderated by genre, indicating the effects are not confined to one writing context.","Because perceptions influence use, tools that display text too fast or too slowly may change how much users scrutinize suggestions, influencing acceptance and rejection decisions."],"supporting_citations":[{"why":"Earlier VDU study showing that medium text-presentation rates yield higher reading comprehension, the basis for choosing the speed anchors.","marker":"[115]"},{"why":"Supplies the average reading speed range used to set the medium condition near 600 wpm.","marker":"[92]"},{"why":"Speech-rate research showing that speaking pace affects perceived competence and personality, the analog motivating text-speed perception.","marker":"[111]"},{"why":"Co-writing study where inserting and deleting words is interpreted as struggling, informing the interpretation of text appearance as a process cue.","marker":"[119]"},{"why":"Shows that co-writing with opinionated language models affects users' views, evidence that tool presentation can influence writers.","marker":"[54]"},{"why":"Provides the ownership and adoption measures adapted for the survey and the exploratory analysis.","marker":"[31]"},{"why":"Nonparametric analysis method used to confirm the parametric ANOVA results.","marker":"[121]"}],"fun_headline_variants":["Text pacing alters trust in AI writing tools","AI output speed shapes perceived quality and trust","Medium-speed AI text reads as most human","Character order and speed drive AI trust judgments"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The study's load-bearing premise is that imagining oneself as a co-writer produces the same perceptions as actually writing; the authors note that imagining does not create the same experience as writing, so the measured effects may not extend to real writing sessions.","fun_headline_variants_meta":{"raw":{"variants":["Text pacing alters trust in AI writing tools","AI output speed shapes perceived quality and trust","Medium-speed AI text reads as most human","Character order and speed drive AI trust judgments"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000282,"raw_usage":{"total_tokens":1600,"prompt_tokens":807,"completion_tokens":793,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":423,"completion_tokens_details":{"reasoning_tokens":738}},"tokens_in":423,"tokens_out":793,"duration_ms":8846,"temperature":1.0,"reasoning_tokens":738,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:30:31.147506+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication in which participants actually write their own text and receive unique AI completions, with the same speed conditions, would falsify the generalization if the medium-speed advantage on quality and trust disappears. A simpler check: have readers rate the final complete texts without seeing the animation; if speed-blind ratings still differ by condition, the effect would be in the content, not the presentation.","supporting_citations":[{"cited_title":"Tombaugh, Michael D","cited_arxiv_id":null,"evidence_quote":"Earlier VDU study showing that medium text-presentation rates yield higher reading comprehension, the basis for choosing the speed anchors."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Speech-rate research showing that speaking pace affects perceived competence and personality, the analog motivating text-speed perception."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Co-writing study where inserting and deleting words is interpreted as struggling, informing the interpretation of text appearance as a process cue."},{"cited_title":"Wobbrock, Leah Findlater, Darren Gergle, and James J","cited_arxiv_id":null,"evidence_quote":"Nonparametric analysis method used to confirm the parametric ANOVA results."}],"review_version":1}