{"id":"046caff6-73cf-46d4-8483-c573b915fdb3","arxiv_id":"2504.21032","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"An adapted survey scale can partially rank LLMs on perceived explanation quality for tax refund decisions, but only the best versus worst model differences were statistically significant.","lead":"This paper tests whether a survey scale developed earlier can tell e-government providers which large language model writes the best explanations for tax refund decisions. With 128 respondents, the scale separated the weakest model from the rest, but could not clearly rank the top three.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Pairwise contrasts show only the worst model is separable, so the method does not determine a ranking among the top three LLMs as claimed.","rationale":"The reader's weakest assumption focused on the validity of the researcher-written ground truth and the un-revalidated adapted scale. I agree those are genuine validity threats. However, the more load-bearing concern for the central claim is that the paper's own inferential statistics do not deliver a ranking: the top three models are statistically indistinguishable on both fidelity and interpretability. This is not a speculation about hidden bias; it is directly visible in the reported pairwise contrasts (Section VI) and in the descriptive statistics (Table I). Even if the scale and ground truth are perfectly valid, the experiment is underpowered to rank the top three candidates, so the hypothesis as stated is only partially tested. The reader did note this issue in the rationale, but did not make it the primary weakest assumption; hence 'partial' agreement. The verdict of CONDITIONAL remains appropriate, because the paper honestly reports its limitations and provides data, but the conclusion must be tempered or the study design must be strengthened to actually separate the candidate LLMs. No ad hominem or theatrical framing is intended; this is an internal statistical gap between the hypothesis and the results.","tokens_in":10376,"tokens_out":2509,"duration_ms":24625,"concrete_test":"Re-run the pairwise contrast analysis on the MANCOVA adjusted means in Table II with a multiple-comparison correction (e.g., Tukey HSD) and report all significant pairwise differences; also compute the minimum detectable effect at 80% power with the reported group sizes (n=52–76) and standard deviations (1.14–1.38). If only contrasts involving flan-ul2-20b are significant and the minimum detectable effect exceeds the observed differences among the top three models (~0.2 points), then the study cannot support a full ranking, and the central claim must be narrowed to 'separating the worst model from the others.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"Hypothesis 1 asserts that perceived quality can be empirically elicited to determine the ranking among a set of LLM model types. The paper's own results in Section VI undermine this: pairwise contrast analysis reveals significant differences only between the best and worst model for fidelity (p<.001) and only a mildly significant difference between best and worst for interpretability (p<.1). The adjusted means for granite-3-8b-instruct, llama-3-1-70b, and GPT-4o differ by about 0.1–0.2 on a 7-point scale, and no pairwise contrast among these three is reported as significant. Thus the evidence supports only a partial ordering that separates flan-ul2-20b from the other three; it cannot determine a full ranking among the set. The conclusion that the scale enables selecting 'the most appropriate LLM' overstates the statistical resolution of the experiment. This is an internal inconsistency between the stated hypothesis and the reported inferential results, independent of whether the adapted scale or the researcher-written ground truth narratives are valid. The method, as demonstrated, can only reject the worst model, not rank the remaining candidates.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript adapts a previously developed six-construct scale (fidelity and interpretability subscales) to the e-government domain of tax-refund explanations, and conducts an online survey in which 128 participants rate explanations generated by four LLMs across three inquiry cases. The authors report MANCOVA results, pairwise contrasts, and exploratory attempts to automate human ratings via regression on embeddings and LLM-as-a-judge prompting. The central claim is that perceived explanation quality can be empirically elicited to rank alternative LLM model types, and that the resulting comparison provides a basis for choosing an LLM for e-government explanation generation.","tokens_in":10531,"tokens_out":5649,"duration_ms":62150,"significance":"If the statistical and measurement concerns are resolved, the paper would offer a concrete, openly documented template for comparing LLM-generated explanations in a public-service context. Its strengths include the use of a scale with prior validation in business-process explanation research, public availability of survey materials, prompts, and data on Zenodo, and transparent reporting of the limited predictive power of the automation attempts. The study also demonstrates large overall differences in perceived fidelity and interpretability across LLM types. However, the paper's inferential results currently support only a partial ordering, and the experiment's statistical analysis does not fully account for its mixed design, so the central ranking claim requires revision or additional evidence.","major_comments":[{"comment":"The results do not support the full ranking stated in Hypothesis 1. The pairwise contrast analysis reported in Section VI shows significant differences only between the best and worst model for fidelity (p < .001) and only a mildly significant best-versus-worst difference for interpretability (p < .1). The adjusted means for granite-3-8b-instruct, llama-3-1-70b, and GPT-4o differ by roughly 0.1–0.2 on a 7-point scale, and no pairwise contrast among these three is reported as significant. The data therefore establish at most that flan-ul2-20b is perceived as worse on fidelity, not a full ranking among the four models. The authors should either revise Hypothesis 1 and the conclusions to state the achievable resolution as a partial ordering, or supply additional evidence such as equivalence tests, Bayesian model ranking, or a power analysis showing that the design could detect smaller differences.","section":"Section IV, Section VI, Table II"},{"comment":"The MANCOVA appears to treat the 256 ratings as independent observations (residual df = 248), although each of the 128 participants contributed two ratings for two explanations generated by different LLMs. Under a mixed design, ratings from the same participant are likely correlated, so the reported F-tests and pairwise contrasts may have biased standard errors. The authors should re-analyze the data with a repeated-measures or mixed-effects model that includes participant-level random effects, or at least report cluster-robust standard errors, and indicate whether the pattern of pairwise comparisons changes.","section":"Section V-B, Table II"},{"comment":"Fidelity is operationalized as agreement with a researcher-written 'ground truth' narrative, and the six-construct scale from prior work [7] is adapted with only wording changes. The manuscript does not report re-validation evidence for the adapted scale in the tax-refund domain, such as internal consistency, factor structure, or measurement invariance, nor any check on the adequacy of the ground-truth narratives, such as expert review or inter-rater agreement. Since the entire model comparison rests on these measurement decisions, the paper should provide such evidence or explicitly temper the claim that the scale is directly transferable to new e-government contexts.","section":"Section V-C and Section V-B"},{"comment":"The sample is not representative of the general taxpayer population: participants were recruited through AI4GOV project partners and their networks, 76.6% hold graduate degrees, and 93% rated their digital literacy at 5 or higher on a 7-point scale. The statement in Section V-A that the sample's representativeness is 'ensured' by the fact that 97.7% are taxpayers is not sufficient. This limitation should be stated explicitly in the discussion, and the authors should consider a sensitivity analysis or an explicit argument about how selection on education and digital literacy could affect the LLM ranking.","section":"Section V-A, Section VI"}],"minor_comments":[{"comment":"The text refers to 'interoperability' where 'interpretability' is meant; this typo also recurs in Section VII and should be corrected throughout.","section":"Section VI, Figure 3"},{"comment":"The table is labeled 'MANCOVA' but reports univariate F-tests for fidelity and interpretability only; the multivariate test statistics (e.g., Pillai's trace, Wilks' lambda) should be reported, or the table relabeled as ANCOVA.","section":"Table II"},{"comment":"The sentence reporting 'overall predictive power ... (i.e., 0.34 and 0.37, respectively)' is ambiguous because these values do not match the preceding R2 values of 0.132 and 0.117; please clarify what statistic these numbers represent.","section":"Section VII-B"},{"comment":"The phrase 'mildly significant' for the interpretability contrast is non-standard; please report the exact p-value and a confidence interval for the best-versus-worst comparison.","section":"Section VI"},{"comment":"The prompt example contains the phrase 'national task refund process', which appears to be a typo for 'national tax refund process'.","section":"Figure 2"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper applies your group's earlier explanation-quality scale to a real e-government use case (tax refunds), with four LLMs and 128 respondents. The data and prompts are on Zenodo, which is good practice. What is genuinely new here is the demonstration that the adapted 24-statement scale can be deployed in a new domain and yield statistically interpretable comparisons. The MANCOVA analysis is competently done, and the authors are unusually honest about the exploratory nature of the predictive and LLM-as-judge parts—those results are weak and they say so.\n\nThe soft spot is the central claim. Hypothesis 1 says the method can determine a ranking among a set of LLMs. The pairwise contrasts show significant differences only between the worst model (flan-ul2-20b) and the others on fidelity, and only a mild difference on interpretability. The top three—granite, llama, GPT-4o—are statistically indistinguishable. That means the method, as demonstrated, separates the worst from the rest, not a full ranking. The conclusion that the scale enables selecting the most appropriate LLM overstates the resolution of the experiment. This is not a fatal flaw, but the paper should either soften the claim to 'reject clearly unsuitable models' or acknowledge that a ranking requires larger samples or additional criteria.\n\nTwo other concerns, both real but smaller. The sample is highly educated and digitally literate, so the numeric ratings may not generalize to the general taxpayer population. And the 'ground truth' narratives were written by the researchers; if those are biased, the fidelity ratings inherit that bias. Both are acknowledged in passing but not deeply discussed.\n\nOverall: the paper is a solid, workmanlike application of an existing scale, with transparent reporting and a reproducible artifact. It does not break new theoretical ground, and the main empirical finding is more modest than the title and conclusion imply. For a researcher working on eGovernment LLM deployment, it is worth a read. I would not cite it as evidence that ranking LLMs by perceived quality works; I would cite it as an example of how to run such a comparison and what the statistical limits look like.\n\nRecommendation: send it to peer review, but with a major-revision request. The authors need to align their claims with the actual statistical resolution, and ideally add a sensitivity analysis around the ground-truth construction. It deserves referee time; it is not a desk reject.","headline":"A useful applied comparison of LLMs for eGov explanations, but the data only separate the worst model, so the promised ranking method is not yet demonstrated.","tokens_in":700,"tokens_out":1310,"would_cite":false,"duration_ms":24645,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The perceived quality of LLM-generated explanations can be measured and used to rank alternative LLMs for e-government services.","keywords":["large language models","e-government","explainability","explanation quality","fidelity and interpretability","tax refund process","user study","LLM selection"],"falsifier":"Have an independent team write different ground-truth explanations for the same three tax cases and repeat the 128-person survey; if the resulting ranking of the four LLMs changes materially, the elicited scores are an artifact of the reference narrative rather than a stable property of the models.","tokens_in":10175,"feed_emoji":"📊","tokens_out":7579,"duration_ms":68994,"temperature":0.7,"pith_summary":"Governments want to automate explanations of decisions like tax refunds, but choosing among the many available LLMs has been largely a matter of guesswork. This paper claims that the choice can be grounded in data: it adapts a previously validated six-construct scale for explanation quality and uses it in a 128-participant study of LLM-generated tax-refund explanations. A multivariate analysis of covariance found statistically significant differences among four LLM types on both fidelity and interpretability, though pairwise tests separated only the best from the worst model on fidelity. The paper concludes that the scale offers ordinal, reproducible guidance for LLM selection, with the final choice depending on how a provider weights fidelity versus interpretability.","feed_headline":"A 24-statement survey ranks LLMs for e-government explanations","feed_subtitle":"Trust in e-government depends on clear explanations; this study turns LLM choice into a measurable comparison.","key_machinery":"The load-bearing instrument is the adapted six-construct explanation-quality scale: six self-reported 1–7 Likert items grouped into fidelity (completeness, soundness, causability) and interpretability (clarity, compactness, comprehensibility), with trust and curiosity included as covariates after the original scale development. The scale converts a subjective impression of an explanation into a numeric vector that can be compared across LLMs through a multivariate analysis of covariance, turning 'which model explains best' into a testable difference in means. Explanations were generated with prompts embedding the tax-refund process model, causal execution dependencies, and, for an in-flight case, the executed trace log, so the comparison holds the generation context fixed and isolates the linguistic output of each model.","core_discovery":"The central claim, stated as Hypothesis 1, is that the perceived quality of LLM-generated explanations can be empirically elicited to determine the ranking among a set of alternative LLM model types. To test it, the authors adapted their earlier explanation-quality scale, rewording 24 statements to the tax-refund context, and had 128 citizens rate explanations generated by granite-3-8b-instruct, llama-3-1-70b, GPT-4o, and flan-ul2-20b for three inquiry cases. Controlling for trust, curiosity, digital literacy, and business-process expertise, the analysis showed a significant effect of LLM type on both fidelity and interpretability; contrast analysis found a significant gap only between the top and bottom models on fidelity and a mildly significant gap on interpretability. The authors therefore present the scale as a practical instrument for ranking LLMs, while noting that non-extreme pairs may not be clearly separable and that the final choice depends on a provider's weighting of the two quality dimensions.","pith_inferences":["Editorial extension: the reported contrasts suggest the strongest practical role of the scale is screening out a clearly weaker LLM, while ties among strong models would still require a provider's own weighting of fidelity versus interpretability.","Editorial extension: the paper changes only the wording when adapting the scale, so a natural follow-up is to re-validate the six-construct structure on a different e-government service, such as a benefits or permit decision, where the stakes and vocabulary differ.","Editorial extension: the weak LLM-as-a-judge correlations may still support a two-stage pipeline—machine scoring to shortlist candidates, human panels to calibrate finalists—though the current sparse data cannot establish that pipeline."],"forward_implications":["An e-government provider can use the 24-statement survey to identify which LLM produces explanations citizens find most faithful and understandable, at least at the level of separating the weakest model from the stronger ones.","The scale transfers to other e-government processes by rewording the statements, so a public authority that already runs citizen surveys does not need to build a new questionnaire from scratch.","Trust and curiosity, not digital literacy or BPM expertise, moderated perceived quality, meaning providers should monitor citizens' general trust in AI rather than their technical skills.","The final choice among closely rated models is an explicit trade-off between fidelity and interpretability; the scale makes that trade-off visible instead of forcing a single winner.","Automated replacement of the survey is not yet reliable: on the current dataset, embeddings-based regression explained little variance and LLM-as-a-judge ratings correlated only weakly with human scores."],"supporting_citations":[{"why":"Supplies the six-construct explanation-quality scale and its original validation, which the paper adapts by rewording statements for the tax-refund domain.","marker":"[7]"},{"why":"Provides the causal execution dependencies used inside the prompts, so the generated explanations are grounded in the process structure rather than produced from free-form text.","marker":"[8]"},{"why":"Makes the survey forms, prompts, and collected data publicly available, grounding the reported comparison in downloadable artifacts for replication.","marker":"[19]"}],"fun_headline_variants":["Survey ranks LLMs for e-gov explanations","128 raters judge LLM explanations for tax refunds","New scale picks best LLM for e-gov explanations","LLM choice for e-gov: a 24-statement survey","Fidelity vs interpretability: ranking LLM explainers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison depends on the researchers' own 'ground truth' narratives being an unbiased reference standard for fidelity, and on the reused 24-statement scale still measuring the same six constructs after only wording changes.","fun_headline_variants_meta":{"raw":{"variants":["Survey ranks LLMs for e-gov explanations","128 raters judge LLM explanations for tax refunds","New scale picks best LLM for e-gov explanations","LLM choice for e-gov: a 24-statement survey","Fidelity vs interpretability: ranking LLM explainers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0002,"raw_usage":{"total_tokens":1382,"prompt_tokens":958,"completion_tokens":424,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":574,"completion_tokens_details":{"reasoning_tokens":342}},"tokens_in":574,"tokens_out":424,"duration_ms":4158,"temperature":1.0,"reasoning_tokens":342,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:59:11.422320+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have an independent team write different ground-truth explanations for the same three tax cases and repeat the 128-person survey; if the resulting ranking of the four LLMs changes materially, the elicited scores are an artifact of the reference narrative rather than a stable property of the models.","supporting_citations":[{"cited_title":"How well can large language models explain business processes as perceived by users?","cited_arxiv_id":null,"evidence_quote":"Supplies the six-construct explanation-quality scale and its original validation, which the paper adapts by rewording statements for the tax-refund domain."},{"cited_title":"The WHY in Business Processes: Discovery of Causal Execution Dependencies,","cited_arxiv_id":null,"evidence_quote":"Provides the causal execution dependencies used inside the prompts, so the generated explanations are grounded in the process structure rather than produced from free-form text."},{"cited_title":"Explanation Quality Survey - the Tax Refund Case,","cited_arxiv_id":null,"evidence_quote":"Makes the survey forms, prompts, and collected data publicly available, grounding the reported comparison in downloadable artifacts for replication."}],"review_version":1}