{"id":"c9a110b8-e421-4ce8-a558-abde329c243d","arxiv_id":"2607.15626","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Chinese users treat LLM fortunetelling as entertainment and emotional support; it subtly shifts confidence and timing but rarely reverses decisions.","lead":"The paper studies how people in China use large language models like DeepSeek for fortunetelling, analyzing 1,045 social-media posts and interviewing 20 users. It finds users mainly seek emotional comfort rather than accurate predictions, and the results nudge confidence and timing of decisions rather than changing their direction.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's 'rarely change beliefs/decisions' is undercut by the paper's own Table 4: 15/20 participants shifted pre/post confidence, and the reported mean change (9.53) does not match the table (9.35). 'Change' is never operationalized.","rationale":"The paper is a useful first qualitative map of LLM fortunetelling: the social media taxonomy is well-sourced, Cohen's kappa is reported, and the interview findings are internally coherent. The strongest_claim is believable as a description of how users narrate their experience. But the scientific claim 'rarely change beliefs/decisions' is the central contribution, and it rests on a distinction between direction and magnitude that the paper never makes explicit. The Table 4 mismatch is a concrete reproducibility problem in the one quantitative element of Study 2. This does not make the qualitative insights worthless, but it means the headline claim needs an operational definition and a corrected analysis before it can be accepted as stated. A control condition would further strengthen causal language, though the paper's wording 'associated with' is appropriately cautious. I therefore do not change the reader's CONDITIONAL verdict.","tokens_in":18938,"tokens_out":6356,"duration_ms":63747,"concrete_test":"Independently re-analyze the Phase 2 pre/post data in Table 4: compute the absolute change for each participant, report counts exceeding thresholds of 0, 5, 10, and 20 points, and verify the mean (9.53 vs. 9.35). Then separately classify decision-related ratings for direction reversal (e.g., pre above 50 to post below 50, or vice versa). If 15/20 participants have nonzero shifts, the sentence 'results rarely change beliefs' must be revised to specify 'rarely reverse the direction of beliefs,' or the abstract's claim fails as stated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing concern is that the abstract's central claim—'results rarely change users' initial beliefs or decisions'—is not operationalized, and the paper's own quantitative data appear to contradict it. Phase 2 pre/post confidence ratings in Table 4 show nonzero absolute changes for 15 of 20 participants (75%), with shifts up to 30 points. The text acknowledges only that 'five participants showed no change' and reports 'average change of 9.53 (SD=8.84)' for the remaining participants. Recomputing from Table 4 gives a mean absolute change of 9.35 over all 20 (or 12.47 over the 15 nonzero-change participants), so the stated mean is not reproducible from the published table. More importantly, whether this constitutes 'rarely change' depends entirely on an unstated threshold: if any nonzero shift counts, 75% changed; if only direction reversals count, the claim may hold, but no such analysis is reported. The sample-size/recruitment limitation is acknowledged and would weaken generalizability, but even a perfect sample would not resolve this definitional ambiguity. The think-aloud trial also lacks a no-LLM control, so the pre/post shifts cannot be attributed to LLM fortunetelling rather than regression to the mean or task engagement.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates LLM-based fortunetelling in China through two qualitative studies: a descriptive content analysis of 1,045 Chinese social media posts from RedNote and Weibo, and semi-structured interviews with 20 users of LLM fortunetelling, including a think-aloud trial with pre/post confidence ratings. The social media analysis yields taxonomies of foretold topics, prompt structures, general attitudes, emotional reactions, perceived credibility, and perceived practicality. The interview study characterizes motivations (decision support, emotional comfort, entertainment, social substitution), perceptions (vague or ingratiating outputs, partial conviction), and impacts on beliefs and decisions (reinforcing existing inclinations, adjusting the timing of actions rather than reversing decision directions). The abstract's central claim is that users treat LLM fortunetelling more as emotional support than accurate prediction, and that results rarely change users' initial beliefs or decisions, while being associated with subtle mindset shifts and small behavioral adjustments.","tokens_in":19135,"tokens_out":4490,"duration_ms":48968,"significance":"If the central claim is properly supported, this is a timely contribution to HCI research on AI-mediated decision-making in self-referential, emotionally loaded everyday contexts, and to cultural studies of digital spirituality and platform-mediated uncertainty. The study has notable strengths: a two-study mixed qualitative design, inter-coder agreement reported (mean Cohen's kappa = 0.845, lowest 0.712), a transparent codebook with platform-stratified statistics, an interview protocol that includes pre/post confidence ratings rather than purely retrospective accounts, and candid acknowledgment of sample and generalizability limitations. However, the paper's central quantitative support for the abstract's 'rarely change' claim is not robust: the operationalization of 'change' is unspecified, the reported mean change is not reproducible from Table 4, and the same table shows nonzero pre/post shifts for 15 of 20 participants. These issues are local and addressable by reanalysis and reframing, but they are load-bearing for the main conclusion.","major_comments":[{"comment":"The central claim that LLM fortunetelling 'rarely changes users' initial beliefs or decisions' is not operationalized, and the paper's own quantitative data point the other way. Table 4 shows nonzero absolute pre/post shifts for 15 of 20 participants (75%). Recomputing from Table 4 gives a mean absolute change of 9.35 points over all 20 participants (12.47 over the 15 nonzero shifters), not the reported 9.53 (SD = 8.84) for 'the remaining' participants. Also, the text says 17 participants rated confidence in the experiment, but Table 4 lists pre/post scores for all 20. If 'change' means any shift in confidence, 75% of the sample changed; if it means a direction reversal of a decision, a different analysis is needed. The authors should report direction-specific outcomes and define a threshold, or revise the abstract to describe 'modest confidence calibration' rather than 'rarely change.'","section":"Abstract; §3.2 Findings (RQ3); Table 4"},{"comment":"The think-aloud trial lacks any no-LLM control or repeated baseline measurement, so the pre/post confidence shifts cannot be attributed to LLM fortunetelling specifically; regression to the mean, task engagement, or the act of verbalizing thoughts could produce similar changes. Given that RQ3 asks how users' beliefs and decisions are 'affected' by LLM fortunetelling, the causal language in the findings and abstract should be softened, or an appropriate control/comparison condition should be added.","section":"Phase 2 Procedure; Findings (RQ3)"},{"comment":"The interview sample is small (N = 20), heavily female (15/20), aged 20–43, and partly recruited through the authors' WeChat networks; only 7 of 67 invited RedNote creators participated. The limitations section acknowledges this, but the abstract's generalizing claim that LLM fortunetelling 'rarely changes' users' beliefs/decisions draws on this sample. At minimum, the claims should be qualified to the recruited sample, and recruitment response rates should be reported in the Participants section rather than only implied in the text.","section":"Participants; Limitations"}],"minor_comments":[{"comment":"Typo: 'has boost a new accessible form' should be 'has boosted a new accessible form.'","section":"Introduction"},{"comment":"The citation 'Huizi, Yiwei et al. 2022' is incomplete and nonstandard; please provide the full author list and venue.","section":"Related Work / References"},{"comment":"The column header 'Confidence in Decision' is ambiguous because the text distinguishes 'Belief' confidence from 'Decision' likelihood. Rename to 'Pre' and 'Post' with explicit units or clarify the header.","section":"Table 4"},{"comment":"The statement that '17 participants chose a topic to predict and rated their confidence' is inconsistent with Table 4, which shows pre/post ratings for all 20 participants. Please reconcile the count.","section":"Findings (RQ3)"},{"comment":"The interpretation in terms of motivated reasoning, confirmation bias, and the Barnum effect is plausible, but these constructs were not directly coded or measured; the discussion should present them as an interpretive lens rather than as empirical findings of the study.","section":"Discussion"}],"recommendation":"major_revision","confidential_remarks":"The paper's qualitative contributions are solid and the topic is timely. The main concern is the mismatch between the abstract's 'rarely change' claim and the paper's own Table 4, plus the unreproducible mean value. These are fixable with a reanalysis and a more careful operationalization of 'change'; I would be willing to review a revised version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nWorth a look if you care about how ordinary people use LLMs for personal, self-referential decisions. This is the first systematic description of LLM fortunetelling in China—1,045 social media posts coded with a careful taxonomy, plus 20 interviews with a think-aloud trial. The two-study design is appropriate for a first map, and the inter-coder agreement (mean kappa 0.845, lowest 0.712) is fine for this kind of open coding.\n\nWhat's genuinely useful: the taxonomy of topics (romance, career, general fortune dominate), prompt structures (Bazi dominates), and perceptions (skepticism 36%, trust 28%, practicality 50%). The interview findings—users treat this less as prediction, more as emotional support and self-reflection; behavioral effects are mostly timing adjustments (delay resignation, postpone a competition), not direction reversals—are plausible and well illustrated with quotes. That is the paper's real contribution.\n\nThe soft spots are real but not fatal. The abstract's \"rarely change users' initial beliefs or decisions\" is not operationalized, and Table 4 shows 15/20 participants moved their pre/post confidence scores, with a mean absolute change of 9.35 points—not the 9.53 reported. If \"rarely change\" means \"no shift at all,\" the data contradict it; if it means \"no direction reversal,\" the paper never reports that analysis. That needs fixing. The interview sample is small, female-heavy, and partly recruited through the authors' WeChat networks; the limitations section acknowledges this, but doesn't fix the over-reach in the abstract. The think-aloud trial has no no-LLM control, so the pre/post shifts could partly be regression to the mean or task engagement. And there is no released codebook, data, or transcripts, which limits reproducibility.\n\nNone of this sinks the central qualitative insight. The paper is a descriptive first step, not a causal study, and it mostly knows that. The 9.53/9.35 mismatch and the undefined \"rarely change\" should be corrected before publication, and the claim scaled back to \"confidence shifts modestly, direction rarely reverses.\"\n\nWho's this for? HCI/ICWSM readers interested in AI-mediated decision-making, belief formation, and playful or spiritual uses of LLMs. I'd send it to peer review, and I'd cite it if I were writing about AI and belief. Bring to reading group if you want a live case of an abstract outrunning its own data.\n\nRecommendation: engage, send to reviewers, ask for a revision that reconciles the numbers and sharpens the \"change\" definition.","headline":"First systematic map of LLM fortunetelling in China; the qualitative core is solid, but the abstract's 'rarely change beliefs' claim is undercut by the paper's own pre/post numbers.","tokens_in":19751,"tokens_out":3089,"would_cite":true,"duration_ms":27750,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"People turn to LLM fortunetelling for emotional support and self-reflection, not for accurate prediction, and it rarely changes the direction of their decisions.","keywords":["LLM fortunetelling","China","emotional support","belief change","decision-making","Barnum effect","confirmation bias","qualitative study"],"falsifier":"A longitudinal study with a representative sample tracking the same users' decisions over several months, or an experiment that randomly assigns users to positive, negative, and neutral forecasts and finds that the direction of their stated plans changes systematically with the forecast, would settle whether the claimed 'reinforcement, not redirection' effect is real or an artefact of a small, self-selected sample.","tokens_in":18723,"feed_emoji":"🔮","tokens_out":4111,"duration_ms":46083,"temperature":0.7,"pith_summary":"People in China are increasingly asking large language models to tell their fortunes, and this paper asks what that practice does to them. Analysing 1,045 social media posts and interviewing 20 users, it finds that people consult LLM fortune-tellers mainly for emotional comfort and self-reflection rather than for accurate prediction. The results rarely reverse a user's existing belief or decision; instead they nudge confidence levels and, occasionally, the timing of an action, such as delaying a resignation or postponing a competition. Users tend to accept the model's flattering, vague answers through confirmation bias and the Barnum effect, treating the exchange as a reassuring ritual. The paper matters because it positions AI fortunetelling as a normal everyday interaction that shapes how uncertainty is felt and expressed, even when it does not change the decision itself.","feed_headline":"LLM fortune-tellers shift confidence more than decisions","feed_subtitle":"1,045 posts and 20 interviews show AI divination mainly reinforces existing beliefs.","key_machinery":"The argument rests on a two-part qualitative apparatus. First, a manually coded taxonomy of 1,045 social media posts categorises foretold topics, emotions, credibility judgements, and perceived practicality. Second, a think-aloud interview study with 20 users includes a live trial in which each participant rates their confidence in a decision or predicted outcome before and after consulting the LLM; the pre/post scores supply the paper's quantitative handle on belief change. The interpretive machinery is the trio of motivated reasoning, confirmation bias, and the Barnum effect, which explains why vague, flattering outputs feel personally accurate and why they reinforce rather than redirect e","core_discovery":"The central discovery is that LLM fortunetelling functions less as a predictive tool and more as a vehicle for emotional reflection and meaning-making. Across the two studies, users approached the LLM with structured prompts rooted in traditional systems such as Bazi, asked about romance, career, academics, and even lost objects, and generally came away positive or amused. When predictions matched reality, users felt surprise and took it as validation; when outputs were vague, users projected their own situations onto them. In the interview trial, 15 of 20 participants shifted their pre/post confidence or decision-likelihood scores, by an average of about 9.5 points, yet the direction of the","pith_inferences":["A direct consequence the paper leaves implicit: if LLM fortunetelling mainly reinforces pre-existing inclinations, its practical risk may be less about bad advice and more about quietly closing off options the user had not already considered; a testable extension would measure whether the set of options users consider narrows after a session.","The claim that results 'rarely change decisions' rests on short-term, self-reported measures; an implication is that small confidence shifts, repeated over weeks, could compound into larger belief drift, which a longitudinal diary study could detect.","Because the mechanism is the Barnum effect, the same interaction pattern should appear with any fluent text generator, not just LLMs; a comparison condition using a scripted pseudo-random 'fortune' generator could isolate whether the LLM's persona matters at all.","The interview sample skewed female (15 of 20); if fortunetelling is a culturally gendered practice, the emotional-support finding may generalise mainly to women, and a mixed-gender replication could change the conclusions."],"forward_implications":["LLM fortunetelling should be evaluated as an emotional-support or reflection tool, not by predictive accuracy, since that is how users themselves judge it.","Design interventions such as reflective prompts, entertainment-versus-support modes, and warnings about sensitive data may reduce over-reliance without destroying the comfort users seek.","Social media amplification of coincidental 'accurate' predictions can skew public perception of LLM reliability, even when individual users remain skeptical.","Over repeated use, vague and ingratiating outputs may weaken critical scrutiny, so the same practice that reassures users could also make them underestimate long-term risks.","Traditional Chinese divination frameworks such as Bazi structure how users prompt the LLM, meaning the model inherits the cultural authority of those systems."],"fun_headline_variants":["AI fortune-telling shifts confidence, not choices","LLM divination: emotional crutch, not crystal ball","Why AI fortune-tellers rarely change decisions","LLM fortunetelling boosts feelings, not decisions","AI fortune-tellers: support over prophecy"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The claim that LLM fortunetelling 'rarely changes' beliefs and decisions rests on 20 volunteer interviewees, mostly young women recruited partly through the authors' own networks, plus a definition of 'change' that treats direction as the test while counting confidence shifts as mere calibration.","fun_headline_variants_meta":{"raw":{"variants":["AI fortune-telling shifts confidence, not choices","LLM divination: emotional crutch, not crystal ball","Why AI fortune-tellers rarely change decisions","LLM fortunetelling boosts feelings, not decisions","AI fortune-tellers: support over prophecy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000102,"raw_usage":{"total_tokens":841,"prompt_tokens":706,"completion_tokens":135,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":450,"completion_tokens_details":{"reasoning_tokens":62}},"tokens_in":450,"tokens_out":135,"duration_ms":2434,"temperature":1.0,"reasoning_tokens":62,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T22:43:14.428241+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A longitudinal study with a representative sample tracking the same users' decisions over several months, or an experiment that randomly assigns users to positive, negative, and neutral forecasts and finds that the direction of their stated plans changes systematically with the forecast, would settle whether the claimed 'reinforcement, not redirection' effect is real or an artefact of a small, self-selected sample.","supporting_citations":[],"review_version":1}