{"id":"dd84cf32-6e13-428b-9774-3631b6270124","arxiv_id":"2505.05786","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"An LLM-driven interactive fiction that puts players in the role of a stigmatized worker raised self-reported understanding and empathy, but stigma reduction itself is not directly measured or controlled.","lead":"This paper tests an AI-powered text game in which players spend a simulated day as a janitor, firefighter, police officer, or caregiver, and asks whether the experience reduces stigma toward those jobs. In a 100-person study, self-reported understanding of the jobs rose and participants reported feeling empathy, but the design lacks a control group and a direct stigma measure, so the headline claim is not yet established.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Appendix A.3 custom scale measures knowledge and perceived value, not stigma; without a validated stigma outcome or control arm, Section 5.1's claim that the LLM-based IF 'demonstrates' stigma reduction is unsupported.","rationale":"The paper has real strengths: the qualitative strand is credible, participants independently reported both empathy and unintended stereotype reinforcement, the realism check by incumbent workers (Table 2) is a useful external anchor, and Section 5.4 candidly acknowledges self-report and comparison limitations. However, these strengths do not close the gap between what was measured and what is claimed. The strongest claim is specifically about stigma reduction, and the quantitative instrument in Appendix A.3 is a knowledge and familiarity scale, not an attitudes scale. Even the large pre/post gains in 'Professional Knowledge' and 'Occupational Stress and Risk' are consistent with simple information provision; the one potentially attitudinal dimension, 'Occupational Value Cognition,' was unchanged for police officers and firefighters (Tables 6 and 8). The post-only empathy and IOS scores cannot evidence change, and the lack of any control arm leaves demand characteristics, testing effects, and the face-swapping procedure as confounds. The qualitative quotes support increased understanding and self-reported empathy, while P1's quote directly illustrates stereotype amplification. Therefore Section 5.1's 'demonstrate... reducing stigma' overstates what the data can show. This is not a dispute about the intervention's plausibility; it is a gap between construct and outcome. A revised version that measures stigma directly, with baseline and comparison conditions, and tempers the conclusion would satisfy the concern. Since this is exactly the condition the reader placed, the verdict remains conditional.","tokens_in":27875,"tokens_out":4784,"duration_ms":54607,"concrete_test":"Conduct a pre/post study with random assignment to (a) the LLM-based IF, (b) an attention-matched control that reads a static first-person narrative of the same occupation, and (c) a no-intervention control, administering both the Appendix A.3 scale and a validated stigma measure (e.g., the Social Distance Scale or an occupational stereotype semantic differential) before and after. If the IF arm does not significantly outperform both controls on the validated stigma measure, Section 5.1's 'demonstrate... reducing stigma' claim is unsupported; if the Appendix A.3 gains are not mirrored by the validated measure, the custom scale is measuring knowledge, not stigma.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Section 5.1: 'Our findings demonstrate that LLM-based perspective-taking IF is an effective approach to reducing occupational stigma') rests on an unsecured measurement premise. The quantitative outcome is the custom scale in Appendix A.3, whose ten items ask about familiarity with tasks, skills, working hours, stress, risks, and perceived social contribution. These items capture knowledge and perceived value, not the stigmatizing attitudes the paper itself defines in Section 2.1: negative labeling, stereotyping, and social exclusion. No item asks about social distance, discomfort, willingness to interact, endorsement of stereotypes, or discriminatory intentions. The pre/post gains therefore support 'understanding increased,' not 'stigma decreased.' The Empathy, Distress, and IOS scores in Table 4 are post-only, so they cannot show a change from baseline, and the absence of any control condition leaves demand effects, repeated-testing familiarity, and the AI face-swapping step (Section 3.2) as alternative explanations for any observed shift. The qualitative results do not rescue the claim: P1 explicitly reported that dramatizing firefighting danger 'might intensify the public's stereotypes' and could make people more reluctant to enter the profession, which is the opposite direction for stigma. Either the paper must add a validated stigma instrument with pre/post measurement and a control arm, or its conclusions must be narrowed to increased understanding and self-reported empathy.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces an LLM-powered interactive fiction (IF) framework for perspective-taking and reports an experiment (n=100) in which participants were randomly assigned to one of four 'dirty work' occupations (janitors, firefighters, police officers, caregivers). The authors measure pre-post change on a custom 10-item scale, post-intervention empathy, personal distress, and Inclusion of Other in the Self (IOS), and conduct 15 semi-structured interviews. They report significant increases in understanding, high empathy and closeness, and qualitative themes of immersion, emotional resonance, and respect. The paper concludes that LLM-based perspective-taking IF reduces occupational stigma toward dirty work, while also acknowledging limitations including stereotype reinforcement and absence of comparison conditions.","tokens_in":28042,"tokens_out":4814,"duration_ms":50027,"significance":"This is a potentially valuable contribution to the growing area of LLM-based prosocial interventions. The interactive fiction framework is novel in allowing dynamic, personalized perspective-taking relative to static video/text interventions, and the authors took the unusual and laudable step of having professional evaluators rate the realism of the scenarios (Appendix A.1). If the causal claim were supported, the approach would be a scalable, accessible tool for mitigating occupational stigma. However, the evidence as presented is not sufficient to support the paper's central claim: the outcome measure does not assess stigma, the design lacks a control arm and baseline empathy/IOS, and qualitative data include explicit reports of stereotype reinforcement. The work is better positioned as a feasibility study of LLM-based IF for increasing understanding and empathy; the present framing overstates what the data can show.","major_comments":[{"comment":"The central claim that 'LLM-based perspective-taking IF is an effective approach to reducing occupational stigma' is not supported by the measurement instrument. The custom scale in Appendix A.3 contains items about familiarity with tasks, skills, working hours, stress, risks, and perceived social contribution. These capture knowledge and perceived value, not the stigma constructs (negative labeling, stereotyping, social exclusion) defined in §2.1. No item asks about social distance, discomfort, willingness to interact, stereotype endorsement, or discriminatory intention. The pre/post gains therefore demonstrate increased self-rated understanding, not reduced stigma. The authors must either add a validated stigma outcome administered pre/post, or substantially narrow the claim (including the title and RQ1) to 'increased understanding and empathy' and present the work as exploratory.","section":"§5.1 and Appendix A.3"},{"comment":"The post-only measurement of Empathy, Personal Distress, and IOS cannot support any claim of change caused by the intervention. Without baseline scores on these measures or a control/comparison condition, high scores are uninterpretable with respect to the intervention's effect. The lack of any control arm leaves demand effects, repeated-testing familiarity, and the co-administered AI face-swapping step (described in §3.2) as equally plausible explanations for the observed pre-post shifts on the custom scale. Causal language throughout §5.1 and the conclusion should be removed or the design must be supplemented with a control condition.","section":"§3.2 and Table 4"},{"comment":"The qualitative evidence undercuts the stigma-reduction conclusion. Participant P1's statement that the simulation 'might intensify the public's stereotypes' and could make people 'more reluctant to become firefighters' is an explicit path toward increased stigma and avoidance behavior, which the authors themselves interpret as the 'paradox of empathy.' Yet this limitation is relegated to a challenge theme, while §5.1 nonetheless asserts that the method 'demonstrates' stigma reduction. The paper needs to either reconcile this finding with the conclusion (e.g., by restricting the claim to attitude warmth without behavioral avoidance) or provide evidence that stereotype reinforcement is outweighed by other components.","section":"§4.3.2 and P1 quote"}],"minor_comments":[{"comment":"In the first paragraph of Section 5, 'the mechanisms through which this approach operation' should read 'operates'.","section":"Section 5"},{"comment":"The response anchors (1=Not at all familiar, 7=Extremely familiar) are not appropriate for the 'Perceived Value of the Profession' items, which ask for evaluation rather than familiarity; use consistent Likert anchors or separate instructions.","section":"Appendix A.3"},{"comment":"Table 9 would benefit from reporting adjusted p-values and confidence intervals consistently; several rows only report P with no effect size, and it is unclear whether the values are already Bonferroni-adjusted.","section":"Table 9"},{"comment":"The qualitative coding section does not report inter-coder reliability or an audit trail; please add a statement on coding reliability.","section":"Section 4.2"},{"comment":"Eligibility required English proficiency while all participants were native Chinese speakers; clarify the reason for this requirement and how the interviews in Chinese were verified against translated responses.","section":"Section 3.2"},{"comment":"The Beliefs about Empathy scale is referenced but not described or cited with validation information; please provide the items or a citation.","section":"Section 3.2"}],"recommendation":"major_revision","confidential_remarks":"The paper appears to be a FAccT submission; the mismatch between the abstract (which is appropriately narrow) and Section 5.1 (which overclaims) suggests the authors are aware of the evidentiary gap. If the authors can be induced to reframe the paper as a pilot/proof-of-concept for LLM-based IF to increase understanding, with explicitly exploratory language, the work could become publishable; otherwise the causal claim cannot stand. I would suggest the editor communicate this clearly. Section 5.4 implicitly admits the lack of control and comparison, yet the discussion still makes causal claims, so the revision should either add new data or systematically soften all causal assertions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a clearly-written, honest exploratory study, but the evidence supports only the narrow claim that an LLM-based interactive fiction increased participants' self-reported understanding of four stigmatized occupations. Section 5.1's assertion that the findings \"demonstrate\" stigma reduction overreaches, and the conclusion repeats that overreach.\n\nWhat's new and good: as far as the paper's own cited literature goes, this is the first application of LLM-based interactive fiction to occupational stigma, with a four-occupation experiment and mixed methods. The qualitative strand is the most credible part—participants describe immersion, empathy, and also unintended stereotype reinforcement, and P1's comment about firefighting danger is exactly the kind of nuance that makes the interviews feel real. The incumbent-worker realism check in Table 2 is a useful external anchor and suggests the authors did careful prompt refinement. The paper is also transparent about its limitations in Section 5.4, which is more than many FAccT submissions do.\n\nSoft spots: the load-bearing problem is measurement. The custom scale in Appendix A.3 asks about familiarity with tasks, stress, risks, and perceived social contribution—knowledge and valuation, not the stigma constructs the paper itself defines in Section 2.1 (negative labeling, stereotyping, social exclusion). Pre/post gains on that scale show understanding increased, not that stigma decreased. Empathy, distress, and IOS scores are post-only, so there is no baseline, and the absence of any control arm leaves demand effects, repeated-testing familiarity, and the AI face-swapping step as alternative explanations. The abstract is more careful than Section 5.1, but the conclusion falls back into the stronger language. The qualitative results do not rescue the claim; P1 explicitly notes the intervention could make people more reluctant to enter the profession, which is the wrong direction for stigma. None of this is fatal to the exploratory contribution, but the central claim needs to be scaled back or the design needs a validated stigma instrument, pre/post baselines, and a comparison condition.\n\nCitation pattern looks reasonable. No code or data are shipped, which limits independent replication. The paper is worth a serious referee—it addresses a real problem and the application is novel—but it should be a conditional accept with major revision, not a direct accept. I'd bring it to a reading group focused on AI-for-social-good or HCI methods.","headline":"Honest exploratory study showing LLM-based interactive fiction can increase self-reported understanding of dirty-work occupations, but the evidence does not support the paper's stronger claim of stigma reduction.","tokens_in":28655,"tokens_out":1565,"would_cite":false,"duration_ms":17576,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI interactive fiction can soften stigma toward 'dirty work' jobs, study claims — role-playing a day as janitor, firefighter, police officer, or caregiver raised understanding and empathy.","keywords":["dirty work","stigma reduction","perspective-taking","interactive fiction","large language models","empathy","occupational bias","GPT-4o simulation"],"falsifier":"A randomized three-arm trial comparing the LLM-based interactive fiction against a static first-person written narrative of the same occupation and against a no-intervention control, using an established stigma outcome such as social distance, hiring intentions, or an implicit association test. If the static narrative matches the interactive fiction's gains, dynamic interactivity is not the active ingredient; if the control arm also gains, the effect is a measurement artefact; if participants' reluctance to enter the profession rises along with their empathy, the stereotype-amplification risk dominates the intended effect.","tokens_in":27536,"feed_emoji":"🎭","tokens_out":9671,"duration_ms":95480,"temperature":0.7,"pith_summary":"This paper claims that a text-based interactive fiction powered by a large language model can reduce public stigma toward 'dirty work' occupations — janitors, firefighters, police officers, and caregivers — by letting participants live a day in the worker's role. In a study of 100 participants, self-rated understanding of the assigned profession rose significantly from before to after the session, and post-session scores on empathy, personal distress, and perceived closeness to the workers were high. Interviews with 15 participants point to immersion, guided reflection, and emotional resonance as the active ingredients, but they also surface a risk: generated storylines can unintentionally amplify the very stereotypes the intervention targets, for example by foregrounding danger. The finding matters because a scalable, low-cost text intervention could extend perspective-taking and stigma reduction beyond lab settings and support occupational equity for essential but stigmatized workers.","feed_headline":"Role-playing a day in an AI game shifts views of dirty work","feed_subtitle":"After playing a day as the worker, 100 participants reported sharply higher understanding and empathy.","key_machinery":"The load-bearing object is an LLM-driven interactive fiction deployed on a commercial chatbot platform with GPT-4o as the base model. The experience has three phases: an introductory guide to the simulated occupation; a branching scenario phase in which the user chooses among suggested options or types free responses and receives real-time consequences that advance the narrative; and a closing summary that recaps the day and poses reflective questions. Scenario design follows the Ease of Self-Simulation heuristic, giving detailed background context so participants can imagine themselves in the role, and draws on Social Cognitive Theory by including scenes of social judgment and family pressure. A co-intervention administered before the session — merging the participant's own photo into an image of the assigned role via AI face-swapping — was meant to deepen immersion. The measurement side rests on a custom ten-item occupation scale given before and after, the Empathy and Personal Distress scales, and the Inclusion of Other in the Self scale, all self-report.","core_discovery":"The paper's central claim is that LLM-based perspective-taking interactive fiction is effective at reducing stigma toward dirty work (its RQ1). Participants who role-played a day as a janitor, firefighter, police officer, or caregiver showed significant pre-to-post gains on a custom ten-item scale covering knowledge of the profession's tasks, stress, risks, and social value, and they reported high empathy, personal distress, and self-other closeness immediately afterward, with no significant differences between the four occupation groups. The qualitative interviews suggest the effect runs through several channels: concrete insight into job responsibilities and environments, appreciation of occupational stress and family conflict, emotional resonance with relatable scenarios, a closing summary that prompts reflection, and a felt sense of professional fulfillment. The authors also report a specific backfire mode: some storylines reinforced existing stereotypes, and greater awareness of danger made one participant more reluctant to see people enter the profession, a tension the paper treats as a design challenge rather than a refutation of the approach.","pith_inferences":["The paper's quantitative evidence is strongest for gains in self-reported understanding and empathy; 'stigma reduction' is an inference from these proxies, and a reader should not treat discriminatory attitudes as measured until a study uses social-distance, hiring-intention, or implicit-association outcomes.","A control arm comparing the LLM-IF to a static first-person narrative of the same occupation would isolate whether real-time adaptive interactivity is the active ingredient or whether ordinary perspective-taking content suffices.","The firefighter finding implies empathy and stigma reduction can diverge: respect and fear can rise together. A testable extension is to measure whether post-session reluctance to enter the profession is correlated with higher empathy scores.","The AI face-swapping step is entangled with the interactive fiction in the current design; a factorial design with and without the face-swap would clarify which component drives the immersion effect."],"forward_implications":["The approach offers a scalable, low-cost alternative to VR and other media-based perspective-taking, deliverable through any chatbot interface.","Because scenarios are simulated, the method protects the privacy of real workers, including those in sensitive roles like policing.","The effect held across four distinct occupations spanning physical and social dirty work, with no significant differences between groups, suggesting the framework transfers without redesign.","The same prompt-driven framework could be pointed at other stigmatized or marginalized roles without building new narrative materials from scratch.","Effective use requires balancing empathy generation against stereotype reinforcement, since danger-heavy storylines can raise respect while also raising reluctance to enter the profession."],"supporting_citations":[{"why":"Defines dirty work and its physical, social, and moral categories; supplies the stigma framework and the four target occupations the study uses.","marker":"[4]"},{"why":"Establishes the mechanism the whole design leans on: feeling empathy for a member of a stigmatized group improves attitudes toward the group.","marker":"[10]"},{"why":"The large-scale VR perspective-taking study this work adapts, including the choice of the Inclusion of Other in the Self scale as its closeness measure.","marker":"[40]"},{"why":"Supplies the Empathy and Personal Distress scales administered as post-test measures of the intervention's emotional effect.","marker":"[9]"},{"why":"The Interpersonal Reactivity Index used to verify that baseline empathy traits were balanced across the four occupation groups.","marker":"[22]"},{"why":"The Inclusion of Other in the Self scale, the closeness measure the study uses to capture self-other overlap with dirty workers.","marker":"[3]"},{"why":"The suppression account the paper invokes to explain why storylines can inadvertently amplify the stereotypes they aim to reduce.","marker":"[20]"},{"why":"The dirty-work conceptualization and measurement review used to justify the Chinese-context framing and to explain null effects on police and firefighter value perception.","marker":"[86]"}],"fun_headline_variants":["LLM game role-play cuts stigma for dirty work jobs","Interactive fiction with LLMs boosts empathy for stigmatized roles","Playing a day as a worker shifts views on dirty work","AI perspective-taking game reduces stigma, but risks stereotypes","Role-play simulations raise support for essential stigmatized jobs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The study assumes that the pre-to-post rise on its custom understanding scale, together with high post-only empathy and closeness scores, measures stigma reduction caused by the interactive fiction itself — but with no control group and no direct measure of stigmatizing attitudes, the same numbers could come from demand effects, social desirability, or the simple experience of learning new facts about a job.","fun_headline_variants_meta":{"raw":{"variants":["LLM game role-play cuts stigma for dirty work jobs","Interactive fiction with LLMs boosts empathy for stigmatized roles","Playing a day as a worker shifts views on dirty work","AI perspective-taking game reduces stigma, but risks stereotypes","Role-play simulations raise support for essential stigmatized jobs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000122,"raw_usage":{"total_tokens":1096,"prompt_tokens":947,"completion_tokens":149,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":563,"completion_tokens_details":{"reasoning_tokens":70}},"tokens_in":563,"tokens_out":149,"duration_ms":2597,"temperature":1.0,"reasoning_tokens":70,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:57:32.913473+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A randomized three-arm trial comparing the LLM-based interactive fiction against a static first-person written narrative of the same occupation and against a no-intervention control, using an established stigma outcome such as social distance, hiring intentions, or an implicit association test. If the static narrative matches the interactive fiction's gains, dynamic interactivity is not the active ingredient; if the control arm also gains, the effect is a measurement artefact; if participants' reluctance to enter the profession rises along with their empathy, the stereotype-amplification risk dominates the intended effect.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The Interpersonal Reactivity Index used to verify that baseline empathy traits were balanced across the four occupation groups."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The suppression account the paper invokes to explain why storylines can inadvertently amplify the stereotypes they aim to reduce."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The dirty-work conceptualization and measurement review used to justify the Chinese-context framing and to explain null effects on police and firefighter value perception."}],"review_version":1}