{"id":"978cd4c0-4c10-4240-9555-17a4150bd7b6","arxiv_id":"2501.18103","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"An LLM chatbot that overlaps with users' typing was perceived as more communicative and immersive than a turn-taking chatbot in a small user study.","lead":"This paper builds a chatbot, OverlapBot, that lets people and an AI type and respond while the other is still typing, instead of taking strict turns. In a user study, people found the overlap-capable bot more natural, communicative, and fast, which suggests text chat with AI could feel more like human conversation.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Study 2 control changes both the interface and the model: OverlapBot used a finetuned Llama3-8B while the baseline was unfinetuned, so perceived benefits may not be caused by overlap itself.","rationale":"The reader's weakest assumption is the same confound I identify: the comparison changes both the overlap interface and the underlying model, so the observed preference cannot be cleanly attributed to overlap. This is load-bearing because the paper's novelty claim rests on the comparative user study; the formative study and Appendix A show the model can perform overlap, but neither isolates overlap from model quality or response brevity. A matched-model control is feasible and would directly settle the attribution. The qualitative findings are coherent and in themselves valuable, and the paper acknowledges some limitations, so a conditional verdict is appropriate: with the added control the central claim could be accepted, but without it the effect remains underdetermined. I therefore recommend no change to the reader's CONDITIONAL verdict.","tokens_in":23059,"tokens_out":4465,"duration_ms":49313,"concrete_test":"Run a within-subject study with three conditions: (A) OverlapBot as implemented; (B) the same finetuned Llama3-8B model in the conventional turn-taking UI, with no real-time typing visibility, no preemptive responses or backchanneling during typing, and no interruption deletion; (C) the original unfinetuned Llama3-8B baseline. Keep the tutorial, task, and session length identical, and instruct models to produce comparable response lengths. If perceived communicativeness, immersion, and naturalness ratings for (B) are not significantly lower than (A), then overlap is not the active ingredient; if (A) still beats (B) with the model held constant, the overlap claim survives.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim attributes users' preference to the overlap capability, but the Study 2 comparison (Section 5) varies more than overlap. OverlapBot runs a finetuned Llama3-8B trained for overlap (Appendix A), while the control uses 'the basic, unfinetuned Llama3-8B model.' The quantitative table also shows OverlapBot's responses are much shorter (133.40 vs 177.64 characters) and its user turns are shorter (43.18 vs 62.36 characters). Participants' 'Brief Response' theme (Section 5.1.1) explicitly credits brevity as a perceived advantage. Therefore the observed ratings of 'communicative' and 'immersive' could reflect the finetuned model's more attentive, concise, chat-style responses rather than the overlap interaction itself. The paper acknowledges the brevity is 'likely influenced by a conversation dataset' and 'not a problem of overlapping itself,' but it provides no control condition that holds the model and response style constant while toggling overlap. Without such a control, the central claim about overlap as the active ingredient is underdetermined.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes replacing strict turn-taking in text-based human-LLM chat with an interface that supports simultaneous typing and overlapping messages. Study 1 (a formative probe with seven dyads) identifies preemptive answering, backchanneling, and deletion as natural overlap behaviors. The authors then build OverlapBot, a Llama3-8B model fine-tuned with Switchboard and instruction-tuning data to reproduce these behaviors. Study 2 (a within-subject user study with 18 participants) compares OverlapBot to a conventional turn-taking chatbot and reports that OverlapBot produced shorter messages, more turns, and qualitative perceptions of being more communicative, immersive, and fast. The paper concludes that overlap capability fosters faster and more natural interactions and offers design insights for overlap-capable text-based AI systems.","tokens_in":23304,"tokens_out":6283,"duration_ms":67777,"significance":"If the causal attribution is accepted, the paper opens a useful design space for text-based human-LLM interaction, extending prior real-time human-human messaging results to LLM agents. The work has genuine strengths: the formative study is well motivated, the prototype is concretely described, the fine-tuning strategy uses public datasets and is evaluated automatically in Appendix A, and the user study includes behavioral logs and rich interview excerpts. However, the central claim is currently underdetermined because the Study 2 comparison varies both the interface and the underlying model, and because the quantitative evidence lacks inferential statistics. These issues directly affect the abstract's claim that overlap itself fosters faster and more natural interactions.","major_comments":[{"comment":"The central comparison is confounded: OverlapBot runs a Llama3-8B fine-tuned on Switchboard and instruction data, while the 'conventional chat system' uses 'the basic, unfinetuned Llama3-8B model.' The OverlapBot condition also adds real-time typing display, backchanneling, preemptive answers, and deletion. Consequently, the perceived differences in 'communicative' and 'immersive' interactions reported in Section 5.1.1 cannot be uniquely attributed to the overlapping capability; they may reflect the fine-tuned model's more conversational tone, its shorter responses, or the real-time streaming display. The abstract and Section 7 state the conclusion as if overlap is the active ingredient. Please add a condition that holds the model and response style fixed while toggling overlap, or substantially reframe the contribution as an evaluation of the OverlapBot system as a whole rather than a demonstration that overlap per se drives the effect.","section":"Section 5, first two paragraphs; Appendix A"},{"comment":"The paper's own 'Brief Response' theme undercuts the speed attribution. Table 1 shows OverlapBot produces notably shorter chatbot messages (133.40 vs. 177.64 characters) and shorter user messages (43.18 vs. 62.36), and participants explicitly praised speed while also criticizing brevity and lack of detail. The text says the brevity 'was likely influenced by a conversation dataset ... not a problem of overlapping itself,' but no control condition separates response brevity from overlap. A parsimonious alternative explanation is that the 'speedy' perception is driven by the fine-tuned model's concise, chat-style outputs rather than by overlap. At minimum, the analysis should examine whether perceived speed tracks the measured overlap rate or message length, and the discussion should acknowledge that the interface and the response style are not separated.","section":"Section 5.1.1, 'Brief Response'; Table 1"},{"comment":"The quantitative comparison in Table 1 reports only means and standard deviations for each condition, with no paired significance tests, confidence intervals, or effect sizes, despite the within-subject design. Statements such as OverlapBot 'facilitated shorter message lengths and a higher number of turns' and that the chatbot sent messages more frequently, 'indicating its ability to provide more information within the same timeframe,' are therefore descriptive rather than statistically supported. Please add appropriate paired inferential analyses (e.g., paired t-tests or Wilcoxon signed-rank tests with effect sizes), or explicitly restrict the quantitative section to descriptive observations and avoid drawing causal conclusions from it.","section":"Table 1"}],"minor_comments":[{"comment":"The paper states that participants were asked which interface they preferred, but it does not report the number who preferred each condition or quote a preference statistic; adding this information would strengthen the qualitative preference claims.","section":"Section 5"},{"comment":"Thematic analysis is described as being conducted by three authors, but the paper does not report a codebook, coding procedure, or inter-rater agreement; a brief description and per-theme frequencies would aid reproducibility and interpretation.","section":"Section 5.1"},{"comment":"The 130-character response truncation threshold is justified only by 'empirical testing'; please report the test results or describe the threshold as a preliminary design parameter requiring further validation.","section":"Section 4.3"},{"comment":"Reference [53] duplicates reference [52] (both are Skantze 2021, 'Turn-taking in Conversational Systems and Human-Robot Interaction'); the duplicate should be removed and subsequent numbering corrected.","section":"References"},{"comment":"The fine-tuning and evaluation section does not report decoding parameters such as temperature, max tokens, or the exact prompts used for generation; including these would improve reproducibility of the automatic evaluation in Table 3.","section":"Appendix A"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the HCI venue's scope and is likely to interest the community, but the main confound is load-bearing: the user study varies the model, the interface, and the response style simultaneously. A rigorous fix would require either a new control condition with the same fine-tuned model in a turn-taking interface or a careful reframing of the contribution. The editor may wish to weigh whether the authors can gather additional data or whether a revised, more modest claim would be acceptable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Know this paper as a solid design exploration with a confounded central comparison. The genuinely new thing is the prototype: a text chat where both the bot and the user can overlap by typing concurrently, with the bot doing backchanneling and preemptive answers and deleting/regenerating when interrupted. The formative study gives real evidence that people spontaneously do this in text, and the three observed interaction patterns (preemptive response as active listening, backchanneling acknowledged but not answered, short interruption commands) are credible and worth reporting. The qualitative analysis is thoughtful and the quotes support the claims.\n\nThe soft spot is exactly what you flagged: Study 2 changes interface and model at the same time. OverlapBot is a finetuned Llama3-8B that produces shorter, chat-style responses; the control is the unfinetuned base model. Participants explicitly credit brevity as an advantage in the 'Brief Response' theme, and the paper itself says this is 'likely influenced by a conversation dataset.' So the 'more communicative and immersive' ratings cannot be cleanly attributed to overlap per se. The quantitative table also has no inferential statistics, just means and SDs, so the turn-rate and message-length differences are descriptive. The authors acknowledge the brevity issue but never run a control that holds the model and response style constant while toggling overlap.\n\nA smaller issue: the Appendix A evaluation of the finetuned model compares against GPT baselines on the authors' own tag classification task. The task is derived from the same data engineering that created the model, so it reads more as a sanity check than an independent benchmark.\n\nNone of this is fatal. The qualitative findings would likely survive a cleaner control, and the design insights (customizable overlap, cultural adaptivity, overlap suited to interactive/multiparty tasks) are sensible and grounded. The paper is also honest about its limitations, including the single-thread delay and the homogeneous participant pool.\n\nWho is this for? HCI researchers and conversational AI designers interested in interaction paradigms beyond strict turn-taking. It deserves a serious referee. I would send it to review, but I'd push for a revised study with a matched baseline — same model, same response style, overlap toggled — plus some significance testing or effect sizes. If the authors do that, the central claim becomes actually testable.","headline":"A credible design exploration of text-based overlap in human-LLM chat, but the main user study is confounded: the overlap interface and the finetuned model change together, so the perceived benefits can't be cleanly pinned on overlap itself.","tokens_in":23792,"tokens_out":1690,"would_cite":true,"duration_ms":17404,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that text-based human-LLM chat need not follow strict turn-taking: a chatbot that overlaps with the user's typing by preemptive answering and backchanneling was perceived as more communicative, immersive, faster, and more…","keywords":["human-AI interaction","text-based chatbot","overlapping messages","turn-taking","large language model","backchanneling","real-time typing","interruption"],"falsifier":"Run the same 18-participant design with two conditions that differ only in whether typing is visible and overlap is allowed, holding the model, response style, and response length identical; if users do not rate the overlap condition as more communicative and immersive, the central claim fails.","tokens_in":22902,"feed_emoji":"💬","tokens_out":5143,"duration_ms":50618,"temperature":0.7,"pith_summary":"The paper argues that strict turn-taking is an unnecessary constraint in text-based human-LLM conversations. Drawing on a formative study in which pairs of people chatted with visible real-time typing, it identifies three overlap behaviors—preemptive answers, backchanneling, and deletion—and builds OverlapBot, an LLM chatbot that performs the first two and reacts to interruptions by deleting and regenerating. In a within-subject study with 18 users, OverlapBot was perceived as more communicative and immersive than a conventional turn-taking chatbot, with faster exchanges and shorter messages. If correct, this opens a design space in which text chat with LLMs mirrors the fluidity of spoken conversation rather than a chess-like alternation of turns.","feed_headline":"Overlapping chatbot feels more human than turn-taking chat","feed_subtitle":"A prototype that answers and reacts while you type was rated more communicative and immersive than a turn-taking bot.","key_machinery":"The central object is OverlapBot, a prototype web chatbot built on Llama3-8B finetuned with customized datasets. The mechanism has three parts: real-time typing display (every keystroke visible, chatbot output streamed character by character); a finetuned model that decides at each point whether to [Await] or [Overlap], and if overlapping whether to produce [Understanding] (backchannel) or [Answer] (preemptive response); and interruption handling, where the user's overlap causes the chatbot to delete its prior text and regenerate, with a 130-character threshold that leaves '...' to signal continuation. This machinery is what translates observed human overlap behaviors into an LLM policy.","core_discovery":"The central claim is that text-based overlap is not only possible in human-LLM interaction but is instinctively adopted and positively valued. OverlapBot lets both parties type simultaneously: the user sees the chatbot's words appear in real time and can interrupt it, and the chatbot backchannels ('yeah') or gives preemptive answers before the user finishes a sentence, deleting its own unfinished response when interrupted. In the user study, participants produced more turns per minute, sent shorter messages, and described the experience as like talking to a real person; they read preemptive answers as evidence of listening, acknowledged backchanneling without replying, and used brief commands like 'stop' or 'okay' to cut the chatbot off. The paper concludes that overlap-capable interfaces make text-based human-LLM conversation more natural and efficient than strict turn-taking.","pith_inferences":["If overlap reduces the need for complete prompts, prompt engineering may become less central: users can start half-formed questions and let the chatbot's early responses guide them, a shift the paper hints at but does not test directly.","The same overlap mechanisms could be transplanted to voice agents or multimodal interfaces, where timing and interruption cues already exist; the paper's deletion-based resolution suggests text offers a unique repair channel that speech lacks.","A testable extension would measure whether overlap benefits persist in task-oriented settings like summarization or time-critical information seeking, where the paper's design discussion predicts moderate benefits at best.","Overlap could change trust dynamics: if users read preemptive answers as listening, then confidently wrong early guesses might be treated as more credible than they should be—a risk the paper does not address."],"forward_implications":["Text-based LLM interfaces can be built so that typing visibility alone creates overlap opportunities, and users will use them without explicit instruction.","Conversations with overlap-capable chatbots move faster: the study measured higher turns per minute and shorter messages for both the user and the chatbot.","Users interpret a chatbot's mid-typing responses as active listening, which drives the perceived humanness and immersion.","Interruptions can replace stop buttons: users naturally type short commands to halt the chatbot, and the chatbot can regenerate based on the interruption.","Overlap designs need user control, since some users find frequent or poorly timed overlaps intrusive; adjustable frequency and typing visibility are implied design requirements."],"supporting_citations":[{"why":"The cited real-time text messaging study supplies the key precedent that live typing reduces silence and enables instant backchanneling and deletion, which the formative study extends to human-LLM chat.","marker":"[27]"},{"why":"The Switchboard dialogue-act corpus provides the conversational data whose overlapping signals were consolidated into [Understanding] tags for finetuning.","marker":"[18]"},{"why":"The turn-taking review supplies the distinction between cooperative and competitive overlap and the interruption-resolution mechanisms that the paper applies to text.","marker":"[52]"},{"why":"Defines backchanneling as strategic active-listening cues, the behavior OverlapBot reproduces with its [Understanding] responses.","marker":"[20]"},{"why":"Provides LoRA, the parameter-efficient finetuning method used to train OverlapBot on the customized overlap dataset.","marker":"[21]"},{"why":"Supports the research-probe methodology used in the formative study to observe natural overlap behavior in pairs.","marker":"[60]"},{"why":"The duplex-model work is the closest technical precedent for generating while receiving input, which the paper contrasts with its interaction-design approach.","marker":"[66]"}],"fun_headline_variants":["Overlapping AI chat feels more human than turn-taking","Chatbot that interrupts and backchannels wins user study","Text overlap: chatbot that talks while you type","Interruptible AI: overlapping messages feel more natural","OverlapBot makes text chat with AI more fluid"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the perceived benefits come from the overlapping capability itself and not from other differences between the two chatbots, since OverlapBot was finetuned and produced shorter responses than the unfinetuned baseline it was compared against.","fun_headline_variants_meta":{"raw":{"variants":["Overlapping AI chat feels more human than turn-taking","Chatbot that interrupts and backchannels wins user study","Text overlap: chatbot that talks while you type","Interruptible AI: overlapping messages feel more natural","OverlapBot makes text chat with AI more fluid"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000225,"raw_usage":{"total_tokens":1412,"prompt_tokens":839,"completion_tokens":573,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":455,"completion_tokens_details":{"reasoning_tokens":497}},"tokens_in":455,"tokens_out":573,"duration_ms":6871,"temperature":1.0,"reasoning_tokens":497,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T00:38:37.553522+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 18-participant design with two conditions that differ only in whether typing is visible and overlap is allowed, holding the model, response style, and response length identical; if users do not rate the overlap condition as more communicative and immersive, the central claim fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The Switchboard dialogue-act corpus provides the conversational data whose overlapping signals were consolidated into [Understanding] tags for finetuning."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines backchanneling as strategic active-listening cues, the behavior OverlapBot reproduces with its [Understanding] responses."}],"review_version":1}