{"id":"8efb72fb-f595-466f-84b1-9ac21629935e","arxiv_id":"2505.08894","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A six-month deployment of a WhatsApp-based LLM chatbot with 97 active users found that factual and health questions dominated, and that daily push questions and a leaderboard were associated with higher engagement.","lead":"A team from Tufts built WaLLM, a free AI chatbot that answers questions over WhatsApp, and ran it for six months with about 100 users in Pakistan, Sudan, and the US diaspora. The study reports what people asked (mostly health and factual questions) and which features kept them engaged, offering early design lessons for AI services in developing regions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'two-thirds of activity within 24h of topQ' statistic is a calendar artifact because topQ was sent daily for 62 days; the abstract's causal-sounding engagement claim is not supported by the reported analysis.","rationale":"The reader's weakest assumption concerns external validity: the recruited sample (direct invitations and snowballing; diaspora users as surrogates) may not represent developing-region populations. I agree that is a genuine limitation, and it is partially acknowledged in §7. However, the more immediately verifiable and more damaging problem is internal: the topQ engagement claim is based on a statistic that is largely a calendar artifact. If topQ is sent daily for 62 days, then any day in that window is 'within 24 hours of a topQ,' so the 62.93% figure mostly tells us where in the deployment timeline users happened to be active. This is not an external-validity dispute; even a perfectly representative sample would not rescue the causal reading. The leaderboard '3x' claim is similarly correlational, but the paper at least reports a within-session Wilcoxon test for leaderboard sessions, whereas the topQ analysis relies on the confounded day-level comparison and the overlap metric. The descriptive findings (55% factual, 28% health, topic distribution) are useful despite lacking inter-rater reliability statistics; the reader's CONDITIONAL verdict already asks for reliability checks and corrected claims. I would keep the verdict CONDITIONAL and add the placebo test as a required revision, since the concern does not invalidate the study's descriptive contributions but does undermine a central causal claim in the abstract.","tokens_in":23259,"tokens_out":6145,"duration_ms":61261,"concrete_test":"Placebo test: reassign the 62 actual topQ dates to a shifted 62-day block (e.g., 30 days earlier or later) that contains no real topQ sends, and recompute for the same 47-user sample the average fraction of active days falling within 24 hours of a placebo topQ date, conditional on the user being registered. If the placebo fraction is statistically indistinguishable from 62.93%, the topQ engagement effect is a calendar artifact. A secondary check: repeat the 'days with topQ vs days without' comparison using only days inside the topQ period, randomly labeling half the days as pseudo-topQ; a significant ratio under this null would support a real nudge effect.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline engagement claim—'Two-thirds of users' activity occurred within 24 hours of the daily top question' (abstract) and 'this feature has encouraged users to engage more' (§1)—is not established by the analysis in §5.2.1. TopQ messages were sent daily for 62 consecutive days beginning three weeks after launch (§3.1.4). Therefore, during that entire window, every day is within 24 hours of a topQ message. The reported 62.93% figure measures the fraction of a user's active days that fell inside the 62-day topQ window, not the fraction attributable to receiving a nudge. The supporting comparison—'days when a topQ was sent had twice as many active users compared to days without a topQ'—is confounded with calendar time and user tenure: non-topQ days all fall outside the topQ period, which is also later in the deployment when churn is high. No control group, no difference-in-differences, and no placebo test is provided. The §7 limitations do not mention this confound. Because the nudge result is a central contribution of the paper, this overreach is load-bearing; if the metric is a calendar artifact, the push-nudge conclusion is unsupported even if the descriptive topic findings survive.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper describes the design, deployment, and log analysis of WaLLM, a WhatsApp-based LLM chatbot targeted at users in Pakistan, Sudan, and the US diaspora. Over a six-month deployment, 112 registered users generated approximately 14.7K interactions, which the authors analyze with a mixed-methods approach: quantitative analysis of sessions and interactive features, plus qualitative coding of a 10% sample of freeform queries. The paper reports that 55% of user queries seek factual information, that health and well-being is the most popular topic (28%), that two-thirds of user activity occurred within 24 hours of the daily Top Question message, and that users who accessed the leaderboard interacted with the service three times more than those who did not. The authors present design implications about trust, cultural customization, and user interface for generative AI in developing regions.","tokens_in":23518,"tokens_out":6095,"duration_ms":58382,"significance":"The deployment itself is a valuable design probe: it provides rare in-the-wild log data on LLM use over a familiar chat interface in developing-region contexts, and the appendix includes the actual prompts, giving the work a reproducible implementation core. The descriptive findings about query topics, use cases, and accuracy are plausible and potentially useful to ICTD and HCI researchers. However, the headline engagement claims are not supported by the reported analysis: the TopQ statistic is vulnerable to a calendar-time confound, and the leaderboard comparison is based on a post hoc subset rather than a comparison against non-users. The paper's main contribution is therefore best framed as descriptive log-based insights and design lessons, not causal evidence for nudging or gamification effects. If the claims are revised to match the evidence, the paper could make a solid contribution; in its current form, the abstract and central narrative overstate what the data show.","major_comments":[{"comment":"The headline engagement claim—'Two-thirds of users' activity occurred within 24 hours of the daily top question'—is not established by the reported analysis. Because TopQ messages were sent daily for 62 consecutive days (§3.1.4), every day in that two-month window is within 24 hours of a TopQ message; the 62.93% figure therefore measures the fraction of users' active days that fell inside the TopQ period, not an activity response to a nudge. The supporting comparison between days with and without a TopQ is confounded with calendar time and tenure, since non-TopQ days fall outside the TopQ period and are later in the deployment when churn is high. No difference-in-differences, placebo test, or pre/post comparison is provided. The abstract and §1 causal wording should be softened or the analysis redesigned before this finding is presented as evidence that the daily nudge increased engagement.","section":"§5.2.1 and Abstract"},{"comment":"The abstract's claim that 'Users who accessed the Leaderboard interacted with WaLLM 3x as those who did not' is not supported by Table 4. Table 4 compares frequent leaderboard users (n=5, selected by repeated access) with occasional users (n=9), not with non-users. The 3:1 ratio in the text refers to frequent users' sessions with Leaderboard access versus their sessions without it (10.8 vs 3.8 interactions per session), a within-user correlation that is subject to self-selection. A direct comparison between users who ever used the leaderboard and matched non-users is missing, so the group-level '3x' claim should be removed or re-derived from a suitable comparison.","section":"§5.2.2, Table 4, and Abstract"},{"comment":"The topic and intent percentages, including the headline 55% factual and 28% health figures, rest on a qualitative coding of a 10% sample (360 freeform queries) by the authors without reported inter-rater reliability, a codebook, or a disagreement-resolution procedure. Treating these thematic labels as quantitative population estimates is therefore fragile. Please add reliability statistics (e.g., Cohen's kappa on a subsample), describe the coding process in more detail, and report the percentages as sample-based descriptive statistics rather than as population estimates.","section":"§4.3.2, §5.1.1"}],"minor_comments":[{"comment":"The reported 'Chi-square Test: U = 165.78' is mislabeled; U conventionally denotes a Mann-Whitney statistic, while a chi-square test should report χ² with degrees of freedom and the contingency table on which it is based.","section":"§5.2.1"},{"comment":"Similarly, the 'Wilcoxon Signed-Rank Test: U = 0.0' should report W (or V) rather than U, and should state the number of pairs on which the signed-rank test is computed.","section":"§5.2.2"},{"comment":"The user categories are defined by session counts, so statements such as 'Regular users have maintained 10 times the number of active days' are partly definitional; these should be presented as descriptive consequences of the category definitions rather than as independent behavioral discoveries.","section":"§4.3.1"},{"comment":"The limitations around self-selected, English-fluent, diaspora-inclusive recruitment are acknowledged in §7 but should be reflected in the abstract and conclusion whenever the topic-mix results are stated for 'users in developing regions,' to avoid overgeneralization.","section":"§4.1, §7, Abstract"},{"comment":"The text says 'approximately 22%' of freeform queries need additional information, but the two stated components are 8% and 13%, which sum to 21%; please reconcile these numbers.","section":"§5.1.3"},{"comment":"The percentages reported in the text (76%, 17%, 7%) do not match Table 1 counts (17, 74, 6 out of 97 users, i.e., approximately 17.5%, 76.3%, 6.2%); please correct the percentages and their ordering.","section":"§4.3.1 and Table 1"}],"recommendation":"major_revision","confidential_remarks":"In my view, the manuscript's core value is the deployment dataset and the qualitative/descriptive insights, not the causal engagement effect sizes. The TopQ and leaderboard claims in the abstract should be brought in line with the descriptive analysis; if the authors revise to focus on log-based insights and design lessons, the paper can make a solid contribution. I also note that reference [7], a 2009 CHI extended abstract about a different context (Liberia TRC), is weak support for the claim that diaspora users are surrogates for home-country users; the authors should either strengthen this justification or explicitly weaken the corresponding claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know before you spend time on this. It is one of the few real deployments of a general-purpose LLM chatbot over WhatsApp in the Global South, with six months of logs and ~97 active users. The descriptive findings—what people ask, the topic mix, the way they use buttons—are genuinely new and worth reading. The causal-sounding engagement numbers in the abstract, though, do not hold up. The stress-test note is correct, and the paper's headline topQ claim is a calendar artifact.\n\nTopQ was sent daily for 62 days starting three weeks into a six-month deployment. During that window, every active day is within 24 hours of a topQ message. So \"62.93% of active days fell within 24 hours of a topQ\" is mostly a statement about where the window sits, not about nudging. The supporting comparison—days with topQ had twice as many active users—is confounded with time and tenure: non-topQ days are later in the deployment, after the service had lost its early users. There is no control, no difference-in-differences, no placebo. That is a load-bearing problem because engagement and nudging is one of the paper's central contributions.\n\nThe leaderboard claim in the abstract is also overreached. The body never compares users who accessed the leaderboard with users who did not. It splits leaderboard users into occasional and frequent, then compares sessions with and without leaderboard access within those groups. Self-selection is an obvious confound: people who are already active are the ones looking at the leaderboard.\n\nWhat's good: the topic analysis is plausible and useful. 55% factual, health at 28%, with concrete examples that show real need (nutrition, disease, personal advice). The response-accuracy check—9% inaccurate, two-thirds uncontested—is an honest, interesting piece of evidence. The prompts are in an appendix, and the limitations section is candid about the diaspora surrogate assumption and the English-only interface. The statistical reporting is sloppy (they call a chi-square U, and a Wilcoxon U), and the qualitative sample has no inter-rater reliability, but those are fixable.\n\nWho's the reader? ICTD and HCI people studying LLM access, plus anyone who wants a case study in how easy it is to overclaim from observational log data. The descriptive parts deserve a serious referee; the causal claims need to be rewritten as descriptive observations or dropped. Send it to review—it should not be desk rejected, but it needs major revision.","headline":"A useful deployment study whose headline engagement numbers are not supported by the analysis; the descriptive topic findings are the real contribution.","tokens_in":24055,"tokens_out":3719,"would_cite":true,"duration_ms":35101,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A WhatsApp-based LLM chatbot, run for six months in developing-region communities, draws mostly factual and health-related queries, with daily push messages and leaderboard access tied to higher engagement.","keywords":["WhatsApp chatbot","large language models","developing regions","user engagement","health information seeking","gamification","digital divide","deployment study"],"falsifier":"A direct test: run the same WhatsApp chatbot for six months with a sample whose in-country location is verified rather than inferred from phone numbers, recruited through community organizations instead of personal networks, and compare the topic mix and the size of the post-push activity spike; if health falls well below 28% or the 24-hour spike disappears, the paper's central generalizations fail.","tokens_in":23072,"feed_emoji":"💬","tokens_out":6239,"duration_ms":63069,"temperature":0.7,"pith_summary":"This paper argues that a general-purpose LLM chatbot delivered over WhatsApp can serve as a practical information and engagement platform in developing regions, and that deployment logs can reveal what users actually seek. Across six months and roughly 100 users in Pakistan, Sudan, and the US diaspora, it reports that 55% of freely typed queries request factual information, that health and well-being is the leading topic at 28%, and that two-thirds of user activity falls within 24 hours of a daily pushed \"Top Question\" message. The stakes, if true, are that a familiar chat interface plus low-cost nudges and simple gamification can measurably increase sustained interaction with AI services where app stores and web interfaces are barriers.","feed_headline":"WhatsApp AI chatbot logs: health is the top question","feed_subtitle":"A six-month, 14.7K-query deployment in Pakistan, Sudan, and the US diaspora shows what people actually ask.","key_machinery":"The operating mechanism is a WhatsApp-native chatbot that wraps off-the-shelf LLMs behind plain text plus button-based interactive messages. The features doing the argumentative work are the daily \"Top Question of the Day\" push, curated \"Trending\" and \"Recent\" query lists, AI-generated suggested follow-ups, and a points-and-leaderboard reward system. These features turn a one-shot Q&A tool into a recurring engagement loop, and the timestamps captured by the WhatsApp Business API let the paper attribute activity spikes to specific nudges, such as the sharp rise in sessions within the first hour after the daily push.","core_discovery":"The paper's central claim is that deploying an LLM assistant inside WhatsApp produces a usable picture of what nonexpert users in developing-region communities want from generative AI, and that design choices measurably shape how they engage. The data show factual information-seeking dominates (55% of freeform queries), health and well-being is the most common topic (28%), users treat the chatbot as a trusted source for nutrition and disease questions, daily push messages are followed by a statistically significant doubling of active users, and users who repeatedly accessed the leaderboard interacted with the service about three times as much as those who did not. The paper also reports that roughly 9% of sampled responses had accuracy issues and that about two-thirds of those inaccurate responses were not contested by users.","pith_inferences":["Inference: The 28% health share likely understates health-related use, since advice and nutrition queries classified outside \"health\" still concern bodily and medical matters; combining topic and intent labels could reveal an even stronger health orientation.","Inference: The study cannot fully separate whether the daily push causes activity or arrives when activity would occur anyway; a randomized A/B design on push timing would distinguish a reminder effect from selection.","Inference: If leaderboard access is largely a marker of already-active users rather than a driver, gamification features should be evaluated causally before being credited with engagement gains.","Inference: The findings suggest a testable extension: adding local-language buttons and retrieval-grounded health answers should increase both factual accuracy and trust, and that can be measured directly in a follow-up deployment."],"forward_implications":["If these patterns hold, WhatsApp is a viable channel for general-purpose AI access in low-bandwidth, low-literacy settings, reducing the need to build separate apps.","Daily push messages produce measurable engagement, so a cheap, non-intrusive reminder may be enough to sustain use over months.","A health-dominated query mix implies that LLM services aimed at developing regions should prioritize accurate, localizable health content rather than assuming use will be mostly entertainment or casual chat.","The roughly 3x interaction gap for leaderboard users, even if partly self-selected, suggests that social visibility and gamification can be an effective engagement lever for a subset of users.","Because about two-thirds of inaccurate responses go uncontested, trust calibration and hallucination awareness are immediate design responsibilities for any such deployment."],"supporting_citations":[{"why":"Supplies the precedent that diaspora users have usage patterns similar to home-country users, grounding the recruitment assumption for the US-based sample.","marker":"[7]"},{"why":"Provides the intent taxonomy used to classify freeform queries as factual, non-factual, advice, or other intent categories.","marker":"[10]"},{"why":"Shows that WhatsApp is preferred and more usable than complex custom platforms in low-digital-literacy contexts, motivating the channel choice.","marker":"[17]"},{"why":"Previous evidence that reply-suggestion buttons increase chatbot usage, supporting the design of button-based navigation and follow-up suggestions.","marker":"[31]"},{"why":"A prior WhatsApp-based LLM chatbot for community health workers, used as the task-specific comparison point that WaLLM extends toward general-purpose use.","marker":"[45]"},{"why":"Introduces the idea of leading conversational search by suggesting useful questions, which underlies the suggested-follow-ups and nudge framing.","marker":"[47]"},{"why":"A log analysis of LLM queries in the wild that supplies the session-based methodology and comparable observations about UI-button friction.","marker":"[54]"}],"fun_headline_variants":["WhatsApp AI logs: health queries top the list","AI chatbot on WhatsApp: users ask about health most","Leaderboard triples AI chatbot engagement on WhatsApp","Daily push messages double active users of WhatsApp AI","14.7K queries reveal WhatsApp AI trust in health info"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the recruited sample—people who accepted direct invitations or snowball referrals, were fluent in English, and included many US-based diaspora members—behaves like the broader developing-region population the deployment is meant to inform.","fun_headline_variants_meta":{"raw":{"variants":["WhatsApp AI logs: health queries top the list","AI chatbot on WhatsApp: users ask about health most","Leaderboard triples AI chatbot engagement on WhatsApp","Daily push messages double active users of WhatsApp AI","14.7K queries reveal WhatsApp AI trust in health info"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000389,"raw_usage":{"total_tokens":2048,"prompt_tokens":942,"completion_tokens":1106,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":558,"completion_tokens_details":{"reasoning_tokens":1030}},"tokens_in":558,"tokens_out":1106,"duration_ms":9486,"temperature":1.0,"reasoning_tokens":1030,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:45:25.809618+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct test: run the same WhatsApp chatbot for six months with a sample whose in-country location is verified rather than inferred from phone numbers, recruited through community organizations instead of personal networks, and compare the topic mix and the size of the post-push activity spike; if health falls well below 28% or the 24-hour spike disappears, the paper's central generalizations fail.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows that WhatsApp is preferred and more usable than complex custom platforms in low-digital-literacy contexts, motivating the channel choice."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Previous evidence that reply-suggestion buttons increase chatbot usage, supporting the design of button-based navigation and follow-up suggestions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"A log analysis of LLM queries in the wild that supplies the session-based methodology and comparable observations about UI-button friction."}],"review_version":1}