{"id":"d1fd2d93-2d60-4f4d-b4d7-dae038bf114e","arxiv_id":"2505.22418","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In interviews with 10 SNAP applicants, trust in an LLM assistant was found to shape whether AI reduces or creates new administrative burdens.","lead":"This study interviewed 10 people rejected for US food benefits, had them try a GPT-4o-based assistant, and mapped how trust in the AI changed the emotional, learning, and paperwork costs they felt. It offers an early framework for designing government AI assistants that reduce administrative burdens instead of adding new ones.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's 'new burdens' are mostly anticipated costs from scenario/preference questions, not measured effects of using SNAP-LLM; the central 'reshaping' claim overstates what the interviews can support.","rationale":"The reader identified sample self-selection, small size, and the controlled setting as the weakest assumption. My concern is closely related but more specifically targets the construct being measured: the 'new burdens' are largely anticipatory projections elicited by scenario and preference questions, not observed consequences of using the system. Even within the existing sample, the interview design does not cleanly separate actual interaction experiences from hypothetical speculation, so the central claim that trust 'reshapes' administrative burdens goes beyond what the data can establish. This does not warrant rejection because the paper is explicitly exploratory, candidly acknowledges controlled-setting and brief-interaction limitations in Section 6.4, and the anticipated-cost framework remains useful for design. However, the Discussion and Conclusion should consistently hedge with 'anticipated' or 'expected' costs. The reader's CONDITIONAL verdict already captures the need for tempering, so I leave the verdict unchanged. My concrete test would settle whether the internal-validity concern lands: if the new-cost themes are grounded primarily in scenario/preference responses, then the overclaim is real and the paper should be revised; if a substantial portion arose during direct interaction with SNAP-LLM, the concern is mitigated. Coding the data by interview stage is a feasible, non-destructive check that the authors can perform with their existing transcripts.","tokens_in":18191,"tokens_out":2728,"duration_ms":35043,"concrete_test":"Re-code the interview data, tagging each quote or finding by which interview stage produced it: (a) actual interaction with SNAP-LLM (step 4), (b) scenario-writing about future/imagined use (step 5), or (c) preference/comparison questions without direct interaction (step 6). For each new-cost theme listed in Table 1 (e.g., fear of stricter AI denials, loss of accompaniment, privacy learning burden, verification burden), record how many supporting quotes come from each stage. If fewer than one-third of the new-cost themes are grounded in direct interaction with SNAP-LLM rather than hypothetical scenarios and preferences, then the central claim should be reframed as 'anticipated trust-burden dynamics' rather than 'reshaping administrative burdens.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, stated in the Discussion and Conclusion, is that LLM-assisted benefits systems 'can alleviate traditional burdens but also generate new psychological, learning, and compliance costs' and that users' trust in LLMs 'shapes these emerging costs.' The evidence, however, comes from interview steps that blend actual interaction with SNAP-LLM (step 4) with scenario-writing and preference questions about hypothetical use (steps 5 and 6). Many of the reported new costs—fear of AI denial, loss of human advocacy, privacy-related learning burdens, information-verification compliance costs—are participants' projections about what they would feel or do, not experiences observed during a real interaction. Section 5.6.1 reports that n=5 'readily accepted the generated information without additional verification,' but in a one-hour controlled session there may have been no real-world verification opportunity, so this is an anticipated behavior, not an incurred cost. The Discussion and Conclusion drop the conditional framing and state that these systems 'generate' new costs, overclaiming experienced effects. This is an internal-validity issue: the trust-burden framework rests on anticipated burden, a legitimate but weaker construct than the 'reshaping administrative burdens' claim. It is distinct from the reader's external-validity concern about the sample, and it applies even to the 10 participants studied.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a qualitative interview study with 10 SNAP applicants who had experienced at least one rejection. Participants discussed their prior experiences with SNAP, interacted with a GPT-4o-based prototype called SNAP-LLM, wrote hypothetical usage scenarios, and stated preferences between human caseworkers and LLM-based support. The authors use thematic analysis to organize findings around the three administrative-burden cost categories (learning, psychological, compliance) and the three trust dimensions (competence, integrity, benevolence). The paper claims that LLM-assisted benefits systems can alleviate traditional burdens but also introduce new psychological, learning, and compliance costs, and that users' trust in LLMs shapes these emerging costs. It concludes with design implications for calibrating trust and disclosing evidence-based performance information.","tokens_in":18442,"tokens_out":4255,"duration_ms":53228,"significance":"If the central claim holds, the paper makes a useful contribution by connecting the administrative-burden framework to citizen-facing LLM systems and by showing how trust dimensions may moderate those burdens. The use of a concrete prototype rather than purely abstract vignettes, the detailed interview protocol, and the inclusion of numerous participant quotes are strengths. The authors also state explicit limitations in Section 6.4, acknowledge that the quantitative measures in Section 5.7 are only descriptive, and frame the work as exploratory theory building. However, the central claim is stronger than the evidence: much of the reported 'new burden' evidence comes from hypothetical scenario responses rather than observed interaction effects, and the small, self-selected, high-digital-literacy sample limits generalizability. These issues make the framework plausible and interesting but not yet fully supported as stated.","major_comments":[{"comment":"The evidence for several newly introduced costs is based on anticipated or hypothetical behavior rather than observed behavior during the SNAP-LLM interaction. For example, Section 5.6.1 states that n=5 'readily accepted the generated information without additional verification,' but in a one-hour controlled Zoom session there was no real-world verification opportunity; the data can support at most an expressed willingness not to verify. Steps 5 and 6 of the interview asked participants to write scenarios and state preferences, which elicits projections about hypothetical use. The Discussion (Section 6) and Conclusion (Section 7) nonetheless state that these systems 'generate' new costs. This conflates anticipated burdens with experienced burdens and is load-bearing for the central 'reshaping' claim. I recommend explicitly coding and reporting which findings reflect actual interaction versus projected use, and revising the central claim to refer to 'reported potential for' or 'anticipated' new costs unless additional evidence is supplied.","section":"Section 4.4 and Section 5.6.1"},{"comment":"The claim in Section 4 that a sample of 10 is 'sufficient for us to achieve analytic generalization, reader generalization, and saturation' is not substantiated. Participants were self-selected Reddit users with at least one SNAP rejection and a mean self-reported digital literacy of 4 out of 5, so the sample is unlikely to represent the full spectrum of the SNAP applicant population. Section 6.4 acknowledges some of these limitations, but the abstract and Discussion still present the findings as a general framework. I recommend reframing the contribution as an exploratory, hypothesis-generating model and either removing the saturation assertion or supporting it with a theme-by-participant matrix and a more precise definition of what saturation means in this context.","section":"Section 4 and Section 6.4"},{"comment":"The coding process pivoted after initial coding revealed trust as a recurring theme, and the codes were then reorganized around established AI trust dimensions. Because the findings are presented through those same dimensions, there is a risk of interpretive circularity: the categories used to organize results were derived from, and then applied back to, the same data. Please describe what steps were taken to guard against forcing data into the trust framework, such as negative-case analysis, explicit comparison with alternative organizing schemes, or independent coding with discussion of disagreements. This is important because the paper's framework rests on the mapping between trust dimensions and burden types.","section":"Section 4.5"},{"comment":"The SNAP-LLM prototype was piloted with two domain experts but was not formally evaluated for output accuracy against the policy manual. Participants' trust judgments and reported burdens were formed in response to the prototype's actual answers, so any inaccuracies or hallucinated policy statements could directly influence the reported trust-burden dynamics. Section 6.4 acknowledges this as a limitation, but the body of the paper, especially Section 5.7 where quantitative trust ratings are reported, should state more clearly that the prototype's accuracy was not independently verified and that the ratings therefore characterize perceived rather than validated system performance.","section":"Section 4.1 and Section 5.7"}],"minor_comments":[{"comment":"The phrase '41 million federally determined low-income applicants' is unclear; 'federally determined' appears to mean income-eligible under federal rules, but the wording should be revised for precision.","section":"Abstract and Section 1"},{"comment":"The interview structure includes scenario-writing and preference questions after the live interaction; please clarify in the protocol description whether the interviewer explicitly distinguished 'what you experienced with SNAP-LLM' from 'what you imagine would happen in a real situation,' since this distinction is central to interpreting the findings.","section":"Section 4.4"},{"comment":"Please state whether the post-interaction survey items were self-administered or read aloud by the interviewer, and note again in this section that the quantitative results are descriptive only, as the Methods section states.","section":"Section 5.7"},{"comment":"Table 1 is dense and not all rows are explicitly discussed in the text; consider adding a pointer in Section 6 that tells readers which rows correspond to which subsections of Section 5, and consider tightening the row labels so they are self-contained.","section":"Table 1"},{"comment":"Some reference formatting is nonstandard, for example reference [65] lists 'the U.S. Digital Service, the Centers on Medicare, and Medicaid Services' as the author with a duplicated 'the'; please check the reference list against the venue's citation style.","section":"References"},{"comment":"There is a typographical oddity in 'un( )intended' in Section 5.4.3, and the same P10 quote about punitive AI appears in both Section 5.4.1 and Section 5.6.1; consider consolidating repeated quotes to avoid redundancy.","section":"Section 5.4.3 and Section 5.6.1"},{"comment":"The paper does not describe how disagreements between the two coders were resolved; if the team used consensus discussion or a specific reflexive thematic analysis approach, please state this explicitly.","section":"Section 4.5"}],"recommendation":"major_revision","confidential_remarks":"This is an exploratory qualitative study with acknowledged limitations, and I would not reject it solely because of the small sample. The main concern is that the central claim about 'reshaping' administrative burdens is worded more strongly than the anticipated-burden evidence supports. If the authors revise the framing to distinguish anticipated from experienced costs and moderate the generalization claims, the paper could make a useful contribution to FAccT on trust and administrative burden in LLM-based benefits systems."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper. First, it is a competent, honest qualitative study: ten SNAP applicants interacted with a prompt-engineered GPT-4o prototype and then talked about trust and burden. The authors are transparent about sample size, recruitment, and the controlled setting. Second, the central claim as stated in the Discussion and Conclusion outruns the evidence. The paper says LLM-assisted systems “generate” new psychological, learning, and compliance costs, but most of the supporting material comes from scenario-writing, preference questions, and projections about hypothetical use, not from observed costs during the actual interaction. That is a real internal-validity problem, and it is distinct from the external-validity limits the authors acknowledge. The stress-test note is correct here.\n\nWhat is genuinely new: the empirical mapping of Mayer's competence/integrity/benevolence trust dimensions onto Herd and Moynihan's three burden categories, in an LLM-assisted benefits context. That mapping is clearly organized (Table 1 is useful), and the interview quotes do support the reported themes. The design implications—source attribution, cognitive forcing functions, hybrid human-AI workflows, transparency about institutional affiliation—are reasonable and grounded in participant language. The topic matters: SNAP applicants are a vulnerable, understudied population in the LLM-in-government literature. For an exploratory qualitative paper, the method is adequately described and the limitations section is candid. I also credit the authors for flagging overtrust as a concern, even if the n=5 “accepted without verification” finding is better described as anticipated behavior than as an incurred cost.\n\nSoft spots, in order of severity. The Discussion and Conclusion drop the conditional framing that the analysis actually supports. The authors should say participants anticipated or feared these costs, not that the systems generate them. The coding pivot in Step 3—reorganizing LLM-related codes around trust dimensions after trust emerged as a theme—is reported transparently, but it creates a mild circularity risk; there is no intercoder agreement metric, which is acceptable in thematic analysis but should be reported or explained. The prototype was never formally evaluated for accuracy, so participants' competence judgments are based on a black box; that is fine for studying perceptions, but the paper should not imply the system's actual reliability is known. None of this is fatal. The sample is small and self-selected, but the authors are honest about that, and for theory-building qualitative work it is defensible.\n\nWho is this for: HCI and public-administration researchers working on trust in AI or administrative burden. It is a plausible first step, not a settled framework. It deserves a serious referee. I would send it to review with the expectation of major revision: reframe the burden findings as anticipatory and preference-based, temper the novelty claims, and either add a formal evaluation of the prototype or explicitly scope the contribution to perceived effects.","headline":"A careful exploratory interview study worth refereeing, but the Discussion overstates actual effects: much of the 'new burden' evidence is participants' anticipated or hypothetical costs, not measured effects of using the prototype.","tokens_in":18928,"tokens_out":1718,"would_cite":true,"duration_ms":26259,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that trust in an LLM along competence, integrity, and benevolence determines whether AI-assisted benefits systems like SNAP reduce administrative burdens or introduce new ones.","keywords":["administrative burden","trust in AI","large language models","SNAP","public benefits","trust calibration","human-AI interaction","qualitative interviews"],"falsifier":"A study with a larger, demographically representative sample of SNAP applicants who use a deployed LLM assistant over several weeks could falsify the framework by showing that variations in competence, integrity, and benevolence trust do not line up with the reported learning, psychological, and compliance costs—for example, if high benevolence trust coincides with high psychological costs, or if users with low competence trust report no additional verification burden.","tokens_in":17994,"feed_emoji":"🤖","tokens_out":4179,"duration_ms":40692,"temperature":0.7,"pith_summary":"The paper argues that LLM-powered assistants for public benefits like SNAP do not simply reduce the hassle of applying for benefits; they change what counts as hassle. Drawing on interviews with ten SNAP applicants who had faced rejection, it proposes that trust in the AI along three lines—competence, integrity, and benevolence—determines whether the system lowers the familiar learning, psychological, and compliance costs of benefits administration or creates new ones. Some new costs arise from mistrust, such as extra effort spent double-checking answers; others arise from overtrust, such as accepting faulty guidance without verification. If this trust-burden dynamic holds, designing such systems is less about maximizing trust and more about calibrating it to the task and the evidence.","feed_headline":"AI trust decides whether LLM benefits tools add or cut red tape","feed_subtitle":"Interviews with SNAP applicants show how AI trust reshapes learning, psychological, and compliance costs.","key_machinery":"The central object is the trust-burden framework that maps three trust dimensions—competence (perceived ability to perform), integrity (perceived adherence to principles), and benevolence (perceived good intentions toward the user)—onto the three administrative burden types from the literature on administrative burdens: learning costs (understanding rules and procedures), psychological costs (stress, stigma, loss of autonomy), and compliance costs (time and money spent meeting requirements). The framework is built from a prompt-engineered GPT-4o prototype, SNAP-LLM, grounded in Indiana's SNAP/TANF policy manual, which interview participants used briefly under controlled conditions; their open-ended responses were coded thematically, with trust emerging as the mediator that reorganizes the burden categories.","core_discovery":"The central claim is that an LLM-assisted benefits system should be understood as a trust-burden mediator: users' beliefs about the system's ability (competence), honesty (integrity), and goodwill (benevolence) shape each of the three administrative burden categories. The authors identify new burden subtypes introduced by LLM use—for example, learning costs about data security protocols, psychological costs from the perceived loss of human accompaniment, and compliance costs from verification of AI answers—and show that the same trust beliefs that reduce traditional burdens can produce these new ones. The paper frames this as an initial framework for trust-burden dynamics in AI-assisted administration and concludes that agencies should disclose evidence-based performance information rather than let applicants form preferences from beliefs.","pith_inferences":["The trust-burden mapping suggests a testable mechanism: interventions that increase one trust dimension (e.g., showing integrity through source citations) might reduce specific burden subtypes while leaving others unchanged, allowing agencies to target design features to the cost they most need to cut.","The authors' finding that participants conflated the informational LLM with a decision-making system implies that deployed systems may need explicit role disclosure (e.g., 'this system advises, it does not decide') to avoid misplaced integrity concerns—an extension not tested in the interviews.","If overtrust is as widespread as the five participants who skipped verification, then the framework predicts that purely user-side transparency tools may fail; burden reduction may require system-side safeguards such as automated verification of outputs against the policy manual.","The sample's high digital literacy suggests the framework may need recalibration for populations with lower digital fluency, where the learning burdens of using an LLM interface could dominate the traditional learning costs they replace."],"forward_implications":["LLM assistants can lower traditional barriers—clearer policy explanations, 24/7 access, and reduced stigma—but agencies should plan for new burden types that appear only when AI is introduced.","Trust calibration, not trust maximization, is the design goal: features such as source attribution, currency indicators, and cognitive forcing functions can reduce overtrust-driven compliance costs.","Hybrid designs that let users choose human caseworkers for sensitive stages and AI for routine information tasks can preserve the psychological value of accompaniment while capturing efficiency gains.","Evidence-based information disclosure—comparing accuracy and consistency of human versus AI sources—could support more informed choices, though beliefs may override disclosed metrics.","The framework implies that evaluations of AI in public services should measure burden changes alongside accuracy, not accuracy alone."],"supporting_citations":[{"why":"Supplies the three-cost administrative burden framework (learning, psychological, compliance) that the paper extends to LLM-assisted systems.","marker":"[32]"},{"why":"Defines learning, psychological, and compliance costs in citizen-state interactions and grounds the burden categories used in the analysis.","marker":"[54]"},{"why":"Provides the three-dimensional trust model (competence, integrity, benevolence) that the paper applies to users' trust in AI.","marker":"[50]"},{"why":"Motivates trust as a key factor in user engagement with AI systems and frames the responsible-trust design goal.","marker":"[45]"},{"why":"Supplies the adapted survey items used to measure benevolence, integrity, and competence trust in SNAP-LLM.","marker":"[52]"},{"why":"Grounds the application of AI trust dimensions to end-user perceptions in real-world system use.","marker":"[41]"},{"why":"Documents stigma and eligibility misperception as traditional burden barriers that the interview findings echo.","marker":"[6]"},{"why":"Supplies items for measuring experienced administrative burden in the post-interaction survey.","marker":"[35]"}],"fun_headline_variants":["Trust in AI: the hidden cost of benefits red tape","SNAP applicants show how AI trust reshapes bureaucracy","LLM benefits: new burdens from misplaced trust","Calibrating trust in AI to lighten administrative load","Evidence-based trust could reduce SNAP red tape"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The framework rests on ten self-selected Reddit users who had been rejected for SNAP at least once, reported above-average digital literacy, and used a non-deployed prototype for a brief, researcher-facilitated session; if these users are not representative of typical SNAP applicants, the claimed trust-burden dynamics may not generalize.","fun_headline_variants_meta":{"raw":{"variants":["Trust in AI: the hidden cost of benefits red tape","SNAP applicants show how AI trust reshapes bureaucracy","LLM benefits: new burdens from misplaced trust","Calibrating trust in AI to lighten administrative load","Evidence-based trust could reduce SNAP red tape"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000596,"raw_usage":{"total_tokens":2743,"prompt_tokens":855,"completion_tokens":1888,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":471,"completion_tokens_details":{"reasoning_tokens":1812}},"tokens_in":471,"tokens_out":1888,"duration_ms":15500,"temperature":1.0,"reasoning_tokens":1812,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:07:36.331936+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A study with a larger, demographically representative sample of SNAP applicants who use a deployed LLM assistant over several weeks could falsify the framework by showing that variations in competence, integrity, and benevolence trust do not line up with the reported learning, psychological, and compliance costs—for example, if high benevolence trust coincides with high psychological costs, or if users with low competence trust report no additional verification burden.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the adapted survey items used to measure benevolence, integrity, and competence trust in SNAP-LLM."},{"cited_title":"2019.Administrative burden: Policymaking by other means","cited_arxiv_id":null,"evidence_quote":"Supplies the three-cost administrative burden framework (learning, psychological, compliance) that the paper extends to LLM-assisted systems."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines learning, psychological, and compliance costs in citizen-state interactions and grounds the burden categories used in the analysis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the three-dimensional trust model (competence, integrity, benevolence) that the paper applies to users' trust in AI."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Motivates trust as a key factor in user engagement with AI systems and frames the responsible-trust design goal."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Grounds the application of AI trust dimensions to end-user perceptions in real-world system use."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents stigma and eligibility misperception as traditional burden barriers that the interview findings echo."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies items for measuring experienced administrative burden in the post-interaction survey."}],"review_version":1}