{"id":"3c13f470-8dab-4851-9803-b052851b5cc4","arxiv_id":"2504.12943","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Adults with social loneliness customize LLM chatbots with personas, voices, and avatars to serve varied emotional needs, from comfort and self-reflection to confronting stressful figures, based on a one-week field study with 22 participants.","lead":"Researchers built ChatLab, a tool for making custom chatbots, and watched 22 lonely adults create personalized digital companions for about a week. The study shows how people use voices, avatars, and backstories to shape chatbots into friends, mirrors, rivals, and therapists for emotional support.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The field protocol explicitly encouraged daily customization—tutorial, reminders, required diary entries, structured template—so the five-purpose persona taxonomy may be probe-elicited rather than natural user behavior, and the paper does not analyze this confound.","rationale":"The reader's weakest assumption identifies the same core risk: the protocol actively encouraged customization without analyzing that influence, so the findings may not generalize to natural behavior. I agree. This is the most load-bearing concern because the strongest claim is precisely about how individuals construct and interact with chatbots, and the main evidence comes from a deployment where customization was mandated, reminded daily, and scaffolded by a template. The paper explicitly states in Section 3.1.3 that active customization had to be encouraged to collect diverse samples, and Section 4.2.3 imposed diary and reminder requirements. Section 7.1 further admits the template nudged participants. These admissions strengthen the concern that the five-purpose taxonomy and the enrichment practices are partly artifacts of the probe. I do not think this warrants rejection: qualitative HCI and Research through Design studies often use such probes to elicit design-relevant behavior, and the themes are supported by extensive quotes, systematic coding, and concrete customization logs. However, the external-validity claim is overstated relative to the protocol, so the work should remain CONDITIONAL rather than being read as evidence of natural real-world customization practice. I also considered the mismatch between P5's reported diagnosed depression and the stated exclusion criterion, and the absence of released transcripts/code. The P5 issue is concerning for reporting consistency and safety but does not directly undermine the strongest descriptive claim as much as the protocol confound does. The lack of raw data limits reproducibility but is common in qualitative work and is not the central argument's weakest point. A targeted re-analysis of existing logs or a minimal no-encouragement replication could settle whether the concern actually lands; until then, the conditional verdict is appropriate.","tokens_in":31644,"tokens_out":4649,"duration_ms":51621,"concrete_test":"Re-analyze the timestamped ChatLab logs and diary entries to test whether customization was protocol-driven: (1) plot persona creation/customization events relative to daily reminders and the five-diary-entry requirement, and (2) compare the pilot tutorial's prompting examples and the interface hints/template text against the five persona categories in Table 3. If customization concentrates immediately after reminders or before diary deadlines, or if the tutorial examples already contain the same categories (e.g., philosopher, mirror-self, ex-partner), then the taxonomy is substantially probe-elicited. A cleaner check would be a small follow-up deployment without the tutorial customization examples, daily reminders, diary requirement, or structured hints; if persona diversity and customization depth collapse, the field-study findings were inflated by the protocol.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central descriptive claim is that participants constructed chatbot personas for five distinct purposes and actively enriched them with voices, avatars, and role-play scenarios. But the customization behavior was elicited by the study protocol, not merely observed. Section 3.1.3 (DR3) states that the design had to 'encourage active customization throughout the study period'; Section 4.2.2 says participants were 'encouraged to be creative and explore various ways to construct a conversation partner'; Section 4.2.3 required daily use, at least five diary entries, and daily reminders; and Section 3.2.2 describes a structured template with hints and an AI-polish feature. The diary form itself asked about 'the customized setting of the chatbot and reasons of the settings.' Section 7.1 even concedes that the template 'explicitly nudged participants to enter information about themselves and the chatbot.' Under these conditions, the five-category taxonomy in Table 3 may reflect the probe's demand characteristics as much as participants' genuine real-world motives. The paper frames the results as answering RQ1 ('how individuals construct and interact') and as 'real-world practices,' yet it never analyzes the extent to which the protocol induced the observed customization, such as whether persona creation clustered around reminders or diary deadlines. This is load-bearing because if customization was substantially protocol-elicited, the descriptive findings describe behavior inside the artificial deployment rather than how individuals would customize in everyday life.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a Research through Design (RtD) study of ChatLab, a prototype platform that lets users construct custom LLM-powered chatbots for emotional support by specifying persona descriptions, avatars, voices, model choice, and temperature. Twenty-two Chinese participants experiencing moderate to high social loneliness used ChatLab for seven to ten days, completed diary entries, and took part in interviews and design activities. The authors identify five categories of persona construction (emotional reliance, confronting stressors, intellectual discourse, self-discovery, and therapeutic support), describe how participants enriched personas with voice/avatar choices and relationship dynamics, and synthesize design implications for future emotional-support chatbots. The main claimed contributions are an empirical account of real-world customization practices and design implications for individualized emotional support in the age of generative AI.","tokens_in":31866,"tokens_out":3388,"duration_ms":36775,"significance":"If the descriptive findings are taken as evidence about how lay users customize LLM chatbots, the paper fills a real gap: prior work emphasized surface-level customization or prompt engineering, while this study captures rich, longitudinal data on persona construction and multimodal enrichment. The paper's strengths include a systematic thematic analysis grounded in 693 codes, abundant participant quotes, conversation logs, diary entries, and transparent reporting of the prototype and procedure. The authors also deserve credit for acknowledging in Section 7.1 that the customization template nudged participants to reflect on themselves and their chatbot. The all-Chinese sample is acknowledged in Section 8. However, the central claim about 'real-world practices' is weakened by a protocol that actively encouraged customization, and the manuscript does not analyze this confound; this is the main correctness risk.","major_comments":[{"comment":"The protocol actively prompted the very behavior the paper claims to observe. Section 3.1.3 states that the design needed to 'encourage active customization throughout the study period'; Section 4.2.2 reports that participants were explicitly 'encouraged to be creative and explore various ways to construct a conversation partner'; and Section 4.2.3 required daily use, at least five diary entries, and daily reminders. The diary form itself asked about customization settings and reasons, and Section 7.1 concedes that the template 'explicitly nudged participants to enter information about themselves and the chatbot.' The paper frames its findings as answering RQ1 ('how individuals construct and interact') and as 'real-world practices' in the contribution statement in Section 1, yet it never analyzes whether the observed customization was induced by these study pressures. This is load-bearing because the five-category taxonomy in Table 3 may reflect demand characteristics of the probe rather than naturalistic user behavior. I ask the authors to either provide evidence that persona creation was not clustered around reminders or diary deadlines (e.g., time-stamped logs relative to reminders), compare early versus late days of the study, or explicitly reframe the contribution as 'practices within a customization-encouraging probe' and discuss the consequence for generalizability. Section 8 does not currently address this issue.","section":"Sections 3.1.3, 4.2.2, 4.2.3, and 7.1"},{"comment":"The paper reports a wide range of persona counts (1 to 13 per participant) and engagement levels (total rounds from 31 to 148), but it does not connect this variability to the protocol's encouragements. If the goal is to describe how individuals construct personas, the analysis should distinguish personas that arose from the participant's own initiative from those that were produced to satisfy the minimum diary requirement or after a reminder. The authors have the conversation logs and diary timestamps needed to perform this check; without it, the claim that participants 'actively constructed' personas for the five purposes in Table 3 is not fully supported. A targeted analysis of persona-creation timing and diary-entry timing would substantially strengthen the central descriptive claim.","section":"Table 2 and Section 5.1"}],"minor_comments":[{"comment":"Typo: 'custimization' should be 'customization' in the heading 'Opportunities for Individualized Emotional Support.'","section":"Section 7.2"},{"comment":"Typo: 'custmoize' should be 'customize' in the description of the design activity.","section":"Section 4.2.4"},{"comment":"The voice name 'Onxy' should be 'Onyx' for consistency with OpenAI's naming.","section":"Table 2, footnote c"},{"comment":"The philosopher 'Jean-Paul Sartre' is misspelled as 'Satre' in the 'Philosophical figure' row and in the text near Section 5.4.1; please correct this.","section":"Table 3"},{"comment":"The abstract lists four persona purposes and then 'etc.', while the paper identifies five categories; consider listing all five in the abstract for accuracy.","section":"Abstract and Section 5.1"},{"comment":"There is a tension between the design goal of an open, non-leading customization interface (DR1) and the tutorial's explicit encouragement to 'be creative and explore various ways to construct a conversation partner'; the paper should reconcile these descriptions or explain the intended balance.","section":"Section 3.2.2 and Section 4.2.2"}],"recommendation":"major_revision","confidential_remarks":"This is a competent qualitative RtD study with rich data, but the reviewer's main concern is that the protocol's explicit encouragement of customization conflates probe-induced behavior with naturalistic practice. The paper already partially acknowledges this in Section 7.1, which is a positive sign, but the limitation is not carried through to the central claims or the limitations section. If the authors can provide a timing-based or otherwise quantitative demonstration that customization was not primarily an artifact of reminders/diary demands, or if they reframe the contribution as a probe-elicited design-driven account, the paper would be publishable. I would lean toward accepting after such a revision. The paper also sits at the boundary of empirical HCI and design research; the authors should ensure the framing matches the target venue's expectations."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a solid qualitative study with a genuinely useful taxonomy and a real deployment. If you work on mental wellbeing tools or LLM personas, it's worth reading. The main thing to know is that the protocol encouraged customization in explicit ways, so the findings are as much about the probe as about natural user behavior.\n\nWhat's new: prior work (Clochat, Song et al., Li et al.) focused on language-only customization or interviews about existing chatbots. Here they built a working multimodal prototype with prompt, voice, and avatar customization, ran a week-long field study with 22 participants, and produced concrete practices: confronting stressors, mirror selves, avatar-based mood expression, and role-play narratives. The thematic analysis is systematic—multiple coders, 693 codes, rich quotes and diary entries. The design implications in RQ2 are grounded in participant drawings and suggestions, not just the authors' speculations. I found Table 2 and Table 3 particularly informative.\n\nSoft spots, in order of significance. First, the demand-characteristics issue is real. DR3 says they deliberately encouraged active customization; the tutorial urged creativity; participants got daily reminders, a minimum of five diary entries, a structured template with hints, and an AI polish feature. The paper concedes the template nudged them (Section 7.1). Given RQ1 is framed as 'how individuals construct and interact,' the taxonomy in Table 3 likely over-represents deliberate, multi-persona construction compared with what people would do organically. The authors don't analyze temporal patterns (e.g., whether persona creation clustered around reminders). This is a moderate limitation, not fatal, because the RtD framing explicitly treats ChatLab as a probe, and the paper could fix it by tempering claims of 'real-world practices' and discussing the probe's influence. Second, Table 1 lists P5 with diagnosed depression, which appears to conflict with the exclusion criterion for formal diagnosis. This needs clarification. Third, the sample is small, all Chinese, self-selected, and no raw transcripts or code are released. These are normal for qualitative HCI, but they limit reproducibility and generalizability. Finally, the study doesn't measure changes in wellbeing, which the authors acknowledge.\n\nOverall, the central descriptive claims are plausible and well-supported by the data. The weaknesses are manageable in revision. I'd send this to peer review and encourage the authors to (a) analyze or at least discuss the protocol's nudge, (b) reconcile P5, and (c) soften the 'real-world' language.","headline":"A useful RtD study of LLM persona customization, but the protocol's explicit nudge to customize means the taxonomy partly reflects the probe's demands.","tokens_in":32411,"tokens_out":2996,"would_cite":true,"duration_ms":30268,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper reports that, given a customizable LLM chatbot, socially lonely adults construct personas for emotional reliance, confronting stressors, intellectual discourse, self-discovery, and therapeutic support, then enrich them with…","keywords":["Emotional support","Chatbot","Wellbeing","Large language model","Prompt","Customization","Human-computer interaction","Persona"],"falsifier":"Run the same deployment with the same customization tool but remove the daily reminders, the five-diary-entry minimum, and the tutorial's encouragement to be creative, and compare whether participants still construct multiple personas and actively adjust voices and avatars; if most users settle into one default persona and rarely touch voice or avatar settings, the claim that such construction is a natural practice would be contradicted.","tokens_in":31420,"feed_emoji":"💬","tokens_out":7227,"duration_ms":73428,"temperature":0.7,"pith_summary":"This paper reports that when people are given a simple way to customize a large language model (LLM) powered chatbot, they construct personas for five distinct emotional purposes: emotional reliance, confronting stressors, connecting to intellectual discourse, fostering self-discovery, and requesting therapeutic support. The claim is grounded in a week-long field deployment of a research prototype, ChatLab, with 22 socially lonely adults, followed by interviews and design activities. The paper's central argument is that customization is not just a technical tuning step; participants actively enrich their personas with voices, avatars, and role-play scenarios to shape the relationship dynamics and encourage more open and honest conversations. A careful reader would care because this recasts personalization as a reflective practice in which users articulate their own needs, rather than a feature the system performs for them.","feed_headline":"Lonely users craft chatbots as pets, philosophers, and therapists","feed_subtitle":"A week-long field study shows persona-building and voice choices shape more open support conversations.","key_machinery":"The carrying object is ChatLab, a research prototype built as a design probe. Its customization interface combines a structured template of text boxes in which users describe themselves and their desired chatbot persona, with hints and an optional AI-polish feature that composes the entries into a prompt; an editable full prompt sent to the language model; a library of 70 Mandarin voices and 77 emoji avatars; and model and temperature controls. This machinery matters because it lets verbal and non-verbal cues be customized in the same loop, which is what revealed the practices of persona enrichment and relationship shaping.","core_discovery":"The central discovery is that lay users treat an LLM chatbot's customization surface as a staging ground for conversations, not merely an output setting. Participants wrote themselves and their chatbot into roles—a beloved pet, a crush, an irritable advisor, an ex-partner, a philosopher, a mirror of the self, a psychologist—and matched those roles with voices and avatars to make the interaction feel coherent and emotionally charged. The paper argues these construction practices serve real emotional work: confronting unresolved stressors, exploring existential questions, experiencing alternative perspectives through role play, and pushing the chatbot past polite neutrality toward candor. It also documents that attempts to make chatbots deliberately 'bad,' angry, or intense often failed because the models stayed too neutral, which the paper treats as a limitation of current LLM behavior in this context.","pith_inferences":["If the customization template itself is what prompts reflection, then a structured persona-builder could serve as a lightweight emotional articulation tool independent of the chatbot's response quality; a testable extension would compare well-being outcomes between users who write detailed personas and users who pick preset personas.","The repeated failure of 'too neutral' responses suggests a design space for user-controlled emotional intensity and safety constraints, in which users could authorize the chatbot to be harsh or provocative in a bounded role-play setting.","The emphasis on local accents and well-known regional voices in this Chinese sample implies dialect-localized voice libraries may be a decisive customization feature for non-English emotional-support chatbots, rather than a peripheral nicety.","Participants' varied memory preferences point toward 'blockable' rather than deletable memory as a mechanism that respects evolving emotional states; this could be studied by testing whether blocked memories resurface in later sessions."],"forward_implications":["Chatbot customization for emotional support will be used for purposes beyond getting comfort, including confronting people or situations that cause stress and exploring philosophical questions.","Voice and avatar choices function as social cues that users align with persona identity and use to express their own mood, so design should treat them as first-class customization options rather than decorations.","Users will write personal anecdotes and role-play instructions into prompts to steer conversational dynamics, asking for autonomy, emotional intensity, and even profanity to escape neutral assistant tones.","Current LLMs respond too neutrally to sustain deliberately negative or confrontational personas, so users' attempts at emotionally intense role play often fall flat.","Future support tools should offer proactive learning from user traces, adjustable memory retention and blocking, AI-assisted persona construction, and community sharing of personas."],"supporting_citations":[{"why":"The CloChat platform study is the direct antecedent for persona customization that ChatLab extends with open text, voice, and avatar options.","marker":"[37]"},{"why":"Social media analysis of how people prompt GPT for mental-health support; it motivates studying the motivations behind customization.","marker":"[51]"},{"why":"Interviews about using LLM chatbots for mental health support; this frames the gap in understanding customization practices and motivations.","marker":"[87]"},{"why":"Experimental evidence that speech speed and formality choices affect perceptions of conversational agents; supports voice as a meaningful customization channel.","marker":"[99]"},{"why":"Taxonomy of social cues for conversational agents; justifies treating avatar and voice as non-verbal cues that enrich personas.","marker":"[27]"},{"why":"Longitudinal study showing that customization in mental-health apps increases user autonomy; grounds the design rationale for encouraging active customization.","marker":"[105]"},{"why":"UCLA Loneliness Scale used to screen and select participants with moderate to high social loneliness.","marker":"[76]"},{"why":"Thematic analysis method used to code interviews, diary entries, and design artifacts.","marker":"[11]"}],"fun_headline_variants":["Chatbot builders stage roles: pets, philosophers, mirrors","Custom chatbots become confidants, mirrors, and sparring partners","Role-playing with chatbots aids emotional work in interactions","Users script chatbot personas to foster candid conversations","Personalized chatbots: users cast roles for honest dialogue"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the personas and interaction choices observed reflect the participants' own needs and creativity, rather than being substantially manufactured by the study's daily reminders, required diary entries, tutorial urging creativity, and template hints with an AI polish feature.","fun_headline_variants_meta":{"raw":{"variants":["Chatbot builders stage roles: pets, philosophers, mirrors","Custom chatbots become confidants, mirrors, and sparring partners","Role-playing with chatbots aids emotional work in interactions","Users script chatbot personas to foster candid conversations","Personalized chatbots: users cast roles for honest dialogue"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000387,"raw_usage":{"total_tokens":1998,"prompt_tokens":858,"completion_tokens":1140,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":474,"completion_tokens_details":{"reasoning_tokens":1063}},"tokens_in":474,"tokens_out":1140,"duration_ms":11407,"temperature":1.0,"reasoning_tokens":1063,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:18:29.328702+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same deployment with the same customization tool but remove the daily reminders, the five-diary-entry minimum, and the tutorial's encouragement to be creative, and compare whether participants still construct multiple personas and actively adjust voices and avatars; if most users settle into one default persona and rarely touch voice or avatar settings, the claim that such construction is a natural practice would be contradicted.","supporting_citations":[{"cited_title":"Personality and Individual Differences 133 (2018), 109–114","cited_arxiv_id":null,"evidence_quote":"Longitudinal study showing that customization in mental-health apps increases user autonomy; grounds the design rationale for encouraging active customization."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Thematic analysis method used to code interviews, diary entries, and design artifacts."}],"review_version":1}