Pith. sign in

REVIEW 4 major objections 6 minor 93 references

This paper proposes a two-step mixed-methods framework that first measures what clinicians, patients, policymakers, and other stakeholders need from explanations of health simulations, then steers LLMs to write summaries matching each group

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-05 05:55 UTC pith:IBEYBLQW

load-bearing objection A clear, honest vision paper for stakeholder-tailored LLM summaries of health simulations; the construct-validity gap between 'information needs' and the empathy measure is the main thing to fix before this becomes a validated pipeline. the 4 major comments →

arxiv 2509.04646 v1 pith:IBEYBLQW submitted 2025-09-04 cs.AI cs.ET

Towards Personalized Explanations for Health Simulations: A Mixed-Methods Framework for Stakeholder-Centric Summarization

classification cs.AI cs.ET
keywords personalized summarizationlarge language modelshealth simulationagent-based modelsstakeholder preferencesmixed-methods frameworkempathy measurementmodel-to-text generation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Health simulations are complex, and the people who could use them—clinicians, policymakers, patients, caregivers, and advocates—need different information in different styles. This paper argues that current LLM-generated explanations of models and simulations are one-size-fits-all and therefore miss the mark. It proposes a step-by-step mixed-methods framework: generate technically correct candidate summaries using a 2x2 factorial design over content and style, measure stakeholder reactions with validated empathy questionnaires, analyze the results per stakeholder group, and steer an LLM through preference optimization to produce tailored summaries. A factuality check by modelers sits between generation and elicitation. If the framework works, it gives a repeatable process for turning any health model and its simulation outputs into summaries that are correct, preference-aligned, and actionable for each audience.

Core claim

The paper's central claim is that the barrier to using health simulations is not only technical but communicative: we lack a systematic account of what different stakeholders need from an explanation. The authors' main contribution is a vision and framework for eliciting, incorporating, and evaluating those needs. The framework has two broad steps. First, the model structure is decomposed into small structured pieces (such as RDF triplets) and translated into text, while simulation outputs are summarized via statistical analysis or multimodal LLMs; candidate summaries are then generated by varying content and style in a designed experiment, checked for factuality by modelers, and piloted wit

What carries the argument

The mechanism is a closed loop: model decomposition into structured triplets small enough for an LLM; factorial design generating candidate summaries that vary controllable attributes such as length, tone, and topic coverage; a modeler factuality check covering knowledge, reasoning, and relevance errors; validated empathy instruments (the Toronto Empathy Questionnaire and the State Empathy Scale) capturing reader reactions; factorial analysis with effect sizes, power analysis, and repeated-measures ANOVA identifying group preferences; preference optimization (DPO and similar alignment methods) steering the LLM; and automatic LLM-based evaluation metrics that filter candidates before improved

Load-bearing premise

The framework assumes that stakeholder groups have stable, distinct preferences for summary content and style that can be elicited reliably with a 2x2 factorial design plus empathy questionnaires, and then used to steer an LLM—an assumption the seven-person pilot does not test.

What would settle it

Run the full loop with two stakeholder groups on the same health simulation. If the factorial analysis finds no significant between-group differences in empathy scores across the four content-by-style summaries, or if a retest one month later reverses each group's preferred combination, the central premise of stable group-level tailoring fails.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • Different stakeholder groups can receive summaries with the same factual core but different content coverage and style, replacing the single generic text now produced.
  • Factuality becomes a gate: no candidate summary reaches stakeholders until modelers have checked it for knowledge, reasoning, and relevance errors.
  • Because preferences are analyzed per group and the LLM is steered accordingly, the approach extends existing model-to-text pipelines without retraining from scratch; few-shot examples can suffice for domain adaptation.
  • Evaluation shifts from text quality alone to whether tailored summaries change downstream decisions, ideally tested by a randomized controlled trial comparing generic versus tailored summaries.
  • The same pipeline can be re-run for new models or new stakeholder groups, making the process repeatable rather than a one-off customization.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If preferences are not stable across time or contexts, group-level tailoring may need to become individual-level or dynamically updated; the seven-person pilot does not yet rule this out.
  • Empathy is the only reaction dimension the protocol measures; trust, perceived accuracy, cognitive load, and actionability may also drive preferences and would not be captured directly.
  • A stronger test than preference alignment is behavioral: tailored summaries should change the decisions stakeholders actually make. A vignette study comparing choices after generic versus tailored summaries would be a cheap precursor to the named randomized controlled trial.
  • The reliance on empathy instruments assumes textual summaries are evaluated emotionally; administrative or technical audiences may prefer a low-empathy executive style, which would complicate the factorial interpretation.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper argues that existing LLM-based explanations of health simulations are one-size-fits-all and proposes a two-step framework to elicit stakeholder-specific content and style preferences and to steer LLM summarization accordingly. Step 1 decomposes simulation models, generates candidate summaries via a 2x2 factorial design over content and style, checks factuality with modelers, and pilots the summaries with stakeholders using empathy questionnaires. Step 2 analyzes factorial data per stakeholder group, steers LLM generation via preference optimization, evaluates with LLM-based and human metrics, and shares optimized summaries back to participants. The authors explicitly position the contribution as a vision and framework, and they defer empirical validation to future work, including ablation studies and a downstream randomized controlled trial.

Significance. If the framework works as intended, it would address a real gap by making health simulation outputs accessible to diverse stakeholders in a tailored way. The paper's strengths are its clear articulation of a repeatable pipeline, the inclusion of a modeler factuality gate, the use of designed experiments for preference elicitation, and an unusually candid discussion of limitations, including an explicit call for ablation studies and downstream decision metrics. However, the claimed ability to elicit content and style needs is not demonstrated, and the key measurement choice (empathy) lacks construct-validity evidence. As submitted, the paper is best read as a research proposal rather than a validated method; its value will depend on the planned empirical studies.

major comments (4)
  1. [Step 1: Identify Information Needs and Preferred Styles From Different Stakeholder Groups] The central claim is that the framework 'elicit[s] ... style and content needs'; however, the only outcome measure specified for the stakeholder pilot is empathy (TEQ and State Empathy Scale). The paper does not justify why maximizing state empathy is equivalent to satisfying a stakeholder's information needs; indeed, the hospital-administrator example (executive summaries, bullet points, business-oriented language) suggests a summary can be appropriate without being maximally empathy-inducing. Because Step 2's factorial analysis and LLM steering optimize the measured outcome, this construct-validity gap is load-bearing. Please either add direct measures of perceived usefulness, comprehension, or decision quality, or reposition the framework as one for empathy-oriented summaries, and justify the use of a single affective outcome across heterogeneous stakeholder groups.
  2. [Step 1: Identify Information Needs and Preferred Styles From Different Stakeholder Groups] The text states: 'Since perceiving a narrative as immersive and compelling depends on the mental state of the reader, we recommend using the validated Toronto Empathy Questionnaire (TEQ; 16 items) to obtain multiple empathy measures.' TEQ is a trait empathy scale, not a state measure of narrative response; using it in this role is likely a misapplication. The State Empathy Scale is the appropriate state measure. Please clarify that TEQ is a baseline covariate and remove the implication that TEQ measures the reader's state in response to a summary.
  3. [Step 2: Optimize the Alignment of Language Models and Stakeholder Communication Needs] The optimization loop is underspecified. The text says 'An optimization process is involved, as the new summaries should be automatically assessed and the architecture adjusted if the scores are insufficient,' but no objective function, score thresholds, or adjustment rules are given. For a framework that claims to be a repeatable process, this prevents replication. Specify at a conceptual level what is optimized (e.g., which evaluation metrics serve as rewards), what 'insufficient' means, and which architectural parameters (RAG parameters, temperature, or others) are adjusted.
  4. [Discussion: On Participatory AI] The paper itself concedes that 'The scientific basis for this framework will be strengthened by collecting experimental data, performing ablation studies...' and that 'the ultimate demonstration that the pipeline works lies in its ability to affect decisions.' I agree. As submitted, the only empirical result is a pilot with n=7 reporting a median response time of 19.16 minutes, which does not test preference validity, summary quality, or decision impact. The authors should either explicitly scope the paper as a position/vision paper with no empirical claims, making the title and abstract match, or include a proof-of-concept with a small but substantive evaluation.
minor comments (6)
  1. [Proposed Framework, Step 1] The factorial design is referred to as '22 factorial design'; this should be typeset as 2^2 factorial design.
  2. [Proposed Framework, Step 2] Typo: 'repeated measures ANOV A' should read 'repeated measures ANOVA'.
  3. [Author affiliations] There is a spacing issue in 'Old Dominion University / 1030 University Blvd, Suffolk, V A 23435, USA' — 'V A' should be 'VA'.
  4. [References] The reference 'Ahrweiler et al. 2019' has 'policy advise'; likely should be 'policy advice'.
  5. [Figure 3] The feedback loop between 'automated assessment' and 'architecture adjustment' would be easier to follow if the figure annotated the specific metrics and parameters involved.
  6. [Proposed Framework, Step 1] The internal loop that asks participants for their preferences and then evaluates whether summaries match those preferences is not inherently circular, because modeler factuality checks and the proposed downstream RCT provide independent anchors. The paper would benefit from stating this explicitly to preempt concerns about self-reference.

Circularity Check

0 steps flagged

No significant circularity: the paper is a framework proposal with an iterative feedback loop and an external RCT anchor; self-citations are used as ordinary building blocks, not as load-bearing self-justification.

full rationale

The paper does not derive quantitative predictions or present equations that reduce to fitted parameters; it proposes a mixed-methods framework. The central loop—generating candidate summaries, measuring participant reactions, factorial-analyzing those reactions, steering the LLM, and re-assessing—is an iterative user-centered design process, not a claim that a prediction is validated by its own inputs. The only pilot (n=7) is reported as response time, not as evidence that empathy scores validate content/style preferences, and the paper explicitly defers causal validation to a future randomized controlled trial measuring downstream decisions, providing an external anchor. Self-citations (e.g., Gandee and Giabbanelli 2024 for model-to-text decomposition; Giabbanelli et al. 2024a for few-shot saturation; Giabbanelli et al. 2025 for empathy-based storytelling) are used as prior empirical building blocks with alternatives listed (e.g., RDF Walks), so they are not load-bearing uniqueness arguments. The construct-validity concern—that empathy questionnaires are used as a proxy for information needs—is a substantive measurement assumption, but it is not circular in the logical sense: 'preferred' is operationalized via empathy, but this is an explicit modeling choice, not a derivation that makes the framework's output equivalent to its input. The paper itself acknowledges its limitations, stating that 'The scientific basis for this framework will be strengthened by collecting experimental data, performing ablation studies...' (Discussion). Thus no circular step meets the evidentiary bar required by the review rules.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 0 invented entities

No free parameters or invented entities; four domain assumptions carry the framework, all untested. The most fragile is that stakeholder preferences are stable and elicitable with the chosen instruments.

axioms (4)
  • domain assumption Distinct stakeholder groups have stable, divergent preferences for summary content and style that can be elicited and generalized.
    The framework's Step 1 and Step 2 depend on stable group-level preferences; the paper motivates this with the hospital layout example but provides no data.
  • domain assumption The Toronto Empathy Questionnaire and State Empathy Scale are valid proxies for whether a summary is understandable and actionable.
    Step 1 recommends TEQ and State Empathy Scale as time-efficient instruments; their predictive validity for summary-driven decision-making is assumed.
  • ad hoc to paper A 2x2 factorial design over content and style captures the variation in stakeholder needs.
    Step 1 states 'if each controllable aspect is simplified by two options, then we have 22 factorial design'; this restricts the preference space to two binary attributes.
  • domain assumption LLMs can produce factually correct summaries if model decomposition, few-shot fine-tuning, and modeler review are applied.
    The framework leans on prior model-to-text work by the authors and others; factuality is checked after generation, not guaranteed.

pith-pipeline@v1.4.0-alltime-deepseek-medium · 16808 in / 11776 out tokens · 113009 ms · 2026-08-05T05:55:53.255483+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of Towards Personalized Explanations for Health Simulations: A Mixed-Methods Framework for Stakeholder-Centric Summarization." pith.science (2026). https://pith.science/paper/IBEYBLQW

@misc{pith2026250904646,
  author       = {Pith},
  title        = {Pith review of: Towards Personalized Explanations for Health Simulations: A Mixed-Methods Framework for Stakeholder-Centric Summarization},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IBEYBLQW}},
  note         = {Machine review of arXiv:2509.04646}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Modeling & Simulation (M&S) approaches such as agent-based models hold significant potential to support decision-making activities in health, with recent examples including the adoption of vaccines, and a vast literature on healthy eating behaviors and physical activity behaviors. These models are potentially usable by different stakeholder groups, as they support policy-makers to estimate the consequences of potential interventions and they can guide individuals in making healthy choices in complex environments. However, this potential may not be fully realized because of the models' complexity, which makes them inaccessible to the stakeholders who could benefit the most. While Large Language Models (LLMs) can translate simulation outputs and the design of models into text, current approaches typically rely on one-size-fits-all summaries that fail to reflect the varied informational needs and stylistic preferences of clinicians, policymakers, patients, caregivers, and health advocates. This limitation stems from a fundamental gap: we lack a systematic understanding of what these stakeholders need from explanations and how to tailor them accordingly. To address this gap, we present a step-by-step framework to identify stakeholder needs and guide LLMs in generating tailored explanations of health simulations. Our procedure uses a mixed-methods design by first eliciting the explanation needs and stylistic preferences of diverse health stakeholders, then optimizing the ability of LLMs to generate tailored outputs (e.g., via controllable attribute tuning), and then evaluating through a comprehensive range of metrics to further improve the tailored generation of summaries.

Figures

Figures reproduced from arXiv: 2509.04646 by Ameeta Agrawal, Philippe J. Giabbanelli.

Figure 1
Figure 1. Figure 1: A model consists of elements and interrelationships from the problem domain, exemplified here as suicide prevention. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: A simulation model has (1) a static structure, which can be decomposed and transformed into text using existing [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Our second step analyses the survey data to find what each group needs in a summary. Then, we steer LLMs in pro [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

93 extracted references · 70 canonical work pages · 4 internal anchors

  1. [1]

    Ahmed, R.; and Hemanth, D. J. 2025. Hybrid text summarization: Integrating extractive and abstractive models for enhanced cross-domain summarization. Intelligent Decision Technologies, 18724981251322745

  2. [2]

    M.; Capellas, B

    Ahrweiler, P.; Sp \"a th, E.; Siqueiros Garc \' a, J. M.; Capellas, B. L.; and Wurster, D. 2025. Inclusive technology co-design for participatory AI. Participatory Artificial Intelligence in Public Social Services: From Bias to Fairness in Assessing Beneficiaries, 35--62

  3. [3]

    Ahrweiler, P.; et al. 2019. Co-designing social simulation models for policy advise: lessons learned from the INFSO-SKIN study. In 2019 Spring simulation conference (SpringSim), 1--12. IEEE

  4. [4]

    Aminpour, P.; Schwermer, H.; and Gray, S. 2021. Do social identity and cognitive diversity correlate in environmental stakeholders? A novel approach to measuring cognitive distance within and between groups. Plos one, 16(11): e0244907

  5. [5]

    D.; Papandrianos, N

    Apostolopoulos, I. D.; Papandrianos, N. I.; Papathanasiou, N. D.; and Papageorgiou, E. I. 2024. Fuzzy cognitive map applications in medicine over the last two decades: A review study. Bioengineering, 11(2): 139

  6. [6]

    L.; and Giabbanelli, P

    Baniukiewicz, M.; Dick, Z. L.; and Giabbanelli, P. J. 2018. Capturing the fast-food landscape in England using large-scale network analysis. EPJ Data Science, 7(1): 39

  7. [7]

    A.; Wornow, M.; Swaminathan, A.; Lehmann, L

    Bedi, S.; Liu, Y.; Orr-Ewing, L.; Dash, D.; Koyejo, S.; Callahan, A.; Fries, J. A.; Wornow, M.; Swaminathan, A.; Lehmann, L. S.; et al. 2025. Testing and evaluation of health care applications of large language models: a systematic review. Jama

  8. [8]

    K.; Zaghir, J.; Zheng, Y.; Bensahla, A.; Bjelogrlic, M.; and Lovis, C

    Bednarczyk, L.; Reichenpfader, D.; Gaudet-Blavignac, C.; Ette, A. K.; Zaghir, J.; Zheng, Y.; Bensahla, A.; Bjelogrlic, M.; and Lovis, C. 2025. Scientific evidence for clinical text summarization using large language models: scoping review. Journal of Medical Internet Research, 27: e68998

  9. [9]

    Belfrage, M.; et al. 2024. Simulating change: A systematic literature review of agent-based models for policy-making. In 2024 Annual Modeling and Simulation Conference (ANNSIM), 1--13. IEEE

  10. [10]

    J.; and Tobias, A

    Brooks, R. J.; and Tobias, A. M. 1996. Choosing the best model: Level of detail, complexity, and model performance. Mathematical and computer modelling, 24(4): 1--14

  11. [11]

    H.; Kader, R.; Ortiz-Prado, E.; Makowski, M

    Busch, F.; Hoffmann, L.; Rueger, C.; van Dijk, E. H.; Kader, R.; Ortiz-Prado, E.; Makowski, M. R.; Saba, L.; Hadamitzky, M.; Kather, J. N.; et al. 2025. Current applications and challenges in large language models for patient care: a systematic review. Communications Medicine, 5(1): 26

  12. [12]

    J.; and Gotz, D

    Caban, J. J.; and Gotz, D. 2015. Visual analytics in healthcare--opportunities and research challenges. Journal of the American Medical Informatics Association, 22(2): 260--262

  13. [13]

    Y.; et al

    Chu, S. Y.; et al. 2025. Think together and work better: Combining humans' and LLMs' think-aloud outcomes for effective text evaluation. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, 1--23

  14. [14]

    N.; and Buchmann, R

    Dolha, D. N.; and Buchmann, R. A. 2024. Generative AI for BPMN process analysis: experiments with multi-modal process representations. In International Conference on Business Informatics Research, 19--35. Springer

  15. [15]

    Dong, H.; Xiong, W.; Goyal, D.; Zhang, Y.; Chow, W.; Pan, R.; Diao, S.; Zhang, J.; Shum, K.; and Zhang, T. 2023. Raft: Reward ranked finetuning for generative foundation model alignment. arXiv preprint arXiv:2304.06767

  16. [16]

    C.; et al

    Dos Santos, V. C.; et al. 2025. Enhancing healthcare operations: a systematic literature review on approaches for hospital facility layout planning. Journal of Health Organization and Management, 39(1): 22--45

  17. [17]

    Fahland, D.; Fournier, F.; Limonad, L.; Skarbovsky, I.; and Swevels, A. J. 2024. How well can large language models explain business processes? arXiv preprint arXiv:2401.12846

  18. [18]

    Fedeli, A.; and Manrique Negrin, D. A. 2024. Towards a collaborative approach for Digital Twin simulation models comprehension. In Proceedings of the ACM/IEEE 27th International Conference on Model Driven Engineering Languages and Systems, MODELS Companion '24, 660–664. New York, NY, USA: Association for Computing Machinery. ISBN 9798400706226

  19. [19]

    Ferrand, N.; Hassenforder, E.; and Girard, S. 2024. Engineering participation: Preparing and designing a participatory process. Transformative Participation for Socio-Ecological Sustainability-Around the CoOPLAGE pathways, 109--121

  20. [20]

    W.; Triplett, Z.; Asif, N.; Susanto, A.; Chowdhury, A.; Azcoaga Lorenzo, A.; Dras, M.; and Berkovsky, S

    Fraile Navarro, D.; Coiera, E.; Hambly, T. W.; Triplett, Z.; Asif, N.; Susanto, A.; Chowdhury, A.; Azcoaga Lorenzo, A.; Dras, M.; and Berkovsky, S. 2025. Expert evaluation of large language models for clinical dialogue summarization. Scientific reports, 15(1): 1195

  21. [21]

    J.; and Giabbanelli, P

    Gandee, T. J.; and Giabbanelli, P. J. 2024. Combining natural language generation and graph algorithms to explain causal maps through meaningful paragraphs. In International Conference on Conceptual Modeling, 359--376. Springer

  22. [22]

    J.; et al

    Gandee, T. J.; et al. 2024. A Visual Analytics Environment for Navigating Large Conceptual Models by Leveraging Generative Artificial Intelligence. Mathematics, 12(13): 1946

  23. [23]

    Gao, M.; Ruan, J.; Sun, R.; Yin, X.; Yang, S.; and Wan, X. 2023. Human-like summarization evaluation with chatgpt. arXiv preprint arXiv:2304.02554

  24. [24]

    Ghaffarzadegan, N.; Lyneis, J.; and Richardson, G. P. 2011. How small system dynamics models can help the public policy process. System Dynamics Review, 27(1): 22--44

  25. [25]

    Giabbanelli, P.; Phatak, A.; Mago, V.; and Agrawal, A. 2024 a . Narrating Causal Graphs with Large Language Models. In Hawaii International Conference on System Sciences 2024 (HICSS-57)

  26. [26]

    J.; and Baniukiewicz, M

    Giabbanelli, P. J.; and Baniukiewicz, M. 2018. Navigating complex systems for policymaking using simple software tools. In Advanced data analytics in health, 21--40. Springer

  27. [27]

    J.; and Baniukiewicz, M

    Giabbanelli, P. J.; and Baniukiewicz, M. 2019. Visual analytics to identify temporal patterns and variability in simulations from cellular automata. ACM Transactions on Modeling and Computer Simulation (TOMACS), 29(1): 1--26

  28. [28]

    J.; Daumas, C.; Flandre, N

    Giabbanelli, P. J.; Daumas, C.; Flandre, N. Y.; Pitkar, A.; and Vazquez-Estrada, J. 2025. Promoting empathy in decision-making by turning agent-based models into stories using large-language models. Journal of Simulation

  29. [29]

    J.; and Vesuvala, C

    Giabbanelli, P. J.; and Vesuvala, C. X. 2023. Human factors in leveraging systems science to shape public policy for obesity: A usability study. Information, 14(3): 196

  30. [30]

    J.; et al

    Giabbanelli, P. J.; et al. 2024 b . Broadening Access to Simulations for End-Users via Large Language Models: Challenges and Opportunities. In 2024 winter simulation conference (wsc), 2535--2546. IEEE

  31. [31]

    Gierend, K.; Kr \"u ger, F.; Genehr, S.; Hartmann, F.; Siegel, F.; Waltemath, D.; Ganslandt, T.; and Zeleke, A. A. 2024. Provenance information for biomedical data and workflows: Scoping review. Journal of medical Internet research, 26: e51297

  32. [32]

    Gravel, J.; D’Amours-Gravel, M.; and Osmanlliu, E. 2023. Learning to fake it: limited responses and fabricated references provided by ChatGPT for medical questions. Mayo Clinic Proceedings: Digital Health, 1(3): 226--234

  33. [33]

    A.; Gray, S.; Cox, L

    Gray, S. A.; Gray, S.; Cox, L. J.; and Henly-Shepard, S. 2013. Mental modeler: a fuzzy-logic cognitive mapping modeling tool for adaptive environmental management. In 2013 46th Hawaii international conference on system sciences, 965--973. IEEE

  34. [34]

    Haddad, E.; and Bugarin, K. 2020. Crisis Control: The Use of Simulations for Policy Decisionmaking. Policy Brief, PB 20, 38

  35. [35]

    a m \"a l \

    H \"a m \"a l \"a inen, R. P.; Luoma, J.; and Saarinen, E. 2013. On the importance of behavioral operational research: The case of understanding and communicating about dynamic systems. European Journal of Operational Research, 228(3): 623--634

  36. [36]

    Hassan, S.; Thompson, C.; Adams, J.; Chang, M.; Derbyshire, D.; Keeble, M.; Liu, B.; Mytton, O.; Rahilly, J.; Savory, B.; et al. 2024. The adoption and implementation of local government planning policy to manage hot food takeaways near schools in England: A qualitative process evaluation. Social Science & Medicine, 362: 117431

  37. [37]

    Hayashi, H.; Budania, P.; Wang, P.; Ackerson, C.; Neervannan, R.; and Neubig, G. 2021. Wikiasp: A dataset for multi-domain aspect-based summarization. Transactions of the Association for Computational Linguistics, 9: 211--225

  38. [38]

    He, J.; Yang, Y.; Long, W.; Xiong, D.; Gutierrez-Basulto, V.; and Pan, J. Z. 2025. Evaluating and Improving Graph to Text Generation with Large Language Models. arXiv preprint arXiv:2501.14497

  39. [39]

    Hinrichs, M.; Wang, J.; Roe, C.; and Johnston, E. W. 2025. AI Integration in Mental Health Services: Examining Trends in the USA and Peoria, Illinois. In Participatory Artificial Intelligence in Public Social Services: From Bias to Fairness in Assessing Beneficiaries, 255--275. Springer Nature Switzerland Cham

  40. [40]

    C.; Ghumrawi, K

    Huddleston, J.; Galgoczy, M. C.; Ghumrawi, K. A.; Giabbanelli, P. J.; Rice, K. L.; Nataraj, N.; Brown, M. M.; Harper, C. R.; and Florence, C. S. 2022. Design and Deployment of a Simulation Platform: Case Study of an Agent-Based Model for Youth Suicide Prevention. In 2022 Winter Simulation Conference (WSC), 2582--2593. IEEE

  41. [41]

    Janssen, M.; and Helbig, N. 2018. Innovating and changing the policy-cycle: Policy-makers be prepared! Government Information Quarterly, 35(4): S99--S105

  42. [42]

    Jeong, D.; Aggarwal, S.; Robinson, J.; Kumar, N.; Spearot, A.; and Park, D. S. 2023. Exhaustive or exhausting? Evidence on respondent fatigue in long surveys. Journal of Development Economics, 161: 102992

  43. [43]

    Kammler, C.; et al. 2023. Towards a Social Simulation Interaction Tool for Policy Makers—A New Research Agenda to Enable Usage of More Complex Social Simulations. In Conference of the European Social Simulation Association, 163--176. Springer

  44. [44]

    Keeble, M.; Adams, J.; Amies-Cull, B.; Chang, M.; Cummins, S.; Derbyshire, D.; Hammond, D.; Hassan, S.; Liu, B.; Medina-Lara, A.; et al. 2024. Public acceptability of proposals to manage new takeaway food outlets near schools: cross-sectional analysis of the 2021 International Food Policy Study. Cities & Health, 8(6): 1094--1107

  45. [45]

    J.; Timmons, S.; Luo, C.; and Shi, L

    Khademi, A.; Zhang, D.; Giabbanelli, P. J.; Timmons, S.; Luo, C.; and Shi, L. 2018. An agent-based model of healthy eating with applications to hypertension. In Advanced Data Analytics in Health, 43--58. Springer

  46. [46]

    P.; Gipp, B.; and Ruas, T

    Kirstein, F.; Wahle, J. P.; Gipp, B.; and Ruas, T. 2025. Cads: A systematic literature review on the challenges of abstractive dialogue summarization. Journal of Artificial Intelligence Research, 82: 313--365

  47. [47]

    A.; Giabbanelli, P

    Lavin, E. A.; Giabbanelli, P. J.; Stefanik, A. T.; Gray, S. A.; and Arlinghaus, R. 2018. Should we simulate mental models to assess whether they agree? In Proceedings of the annual simulation symposium, 1--12

  48. [48]

    Lee, H.; Phatale, S.; Mansoor, H.; Mesnard, T.; Ferret, J.; Lu, K.; Bishop, C.; Hall, E.; Carbune, V.; Rastogi, A.; et al. 2023. Rlaif vs. rlhf: Scaling reinforcement learning from human feedback with ai feedback. arXiv preprint arXiv:2309.00267

  49. [49]

    Lima, F. F. d.; and Os \'o rio, F. d. L. 2021. Empathy: assessment instruments and psychometric quality--a systematic literature review with a meta-analysis of the past ten years. Frontiers in psychology, 12: 781346

  50. [50]

    Liu, Y.; Iter, D.; Xu, Y.; Wang, S.; Xu, R.; and Zhu, C. 2023. G-eval: NLG evaluation using gpt-4 with better human alignment. arXiv preprint arXiv:2303.16634

  51. [51]

    F.; Martinez-Moyano, I

    Luna-Reyes, L. F.; Martinez-Moyano, I. J.; Pardo, T. A.; Cresswell, A. M.; Andersen, D. F.; and Richardson, G. P. 2006. Anatomy of a group model-building intervention: Building dynamic theory from case study research. System Dynamics Review: The Journal of the System Dynamics Society, 22(4): 291--320

  52. [52]

    Manellanga, R.; and David, I. 2024. Participatory and collaborative modeling of sustainable systems: A systematic review. In Proceedings of the ACM/IEEE 27th international conference on model driven engineering languages and systems, 645--654

  53. [53]

    Mavridis, A.; Tegos, S.; Anastasiou, C.; Papoutsoglou, M.; and Meditskos, G. 2025. Large language models for intelligent RDF knowledge graph construction: results from medical ontology mapping. Frontiers in Artificial Intelligence, 8: 1546179

  54. [54]

    Montibeller, G. 2018. Behavioral challenges in policy analysis with conflicting objectives. In Recent advances in optimization and modeling of contemporary problems, 85--108. INFORMS

  55. [55]

    Mussa, O.; Rana, O.; Goossens, B.; Orozco-terWengel, P.; and Perera, C. 2024. Towards Enhancing Linked Data Retrieval in Conversational UIs Using Large Language Models. In International Conference on Web Information Systems Engineering, 246--261. Springer

  56. [56]

    Nezhad, B.; et al. 2025. Fair Summarization: Bridging Quality and Diversity in Extractive Summaries. In Prabhakaran, V.; Dev, S.; Benotti, L.; Hershcovich, D.; Cao, Y.; Zhou, L.; Cabello, L.; and Adebara, I., eds., Proceedings of the 3rd Workshop on Cross-Cultural Considerations in NLP (C3NLP 2025), 22--34. Albuquerque, New Mexico: Association for Computa...

  57. [57]

    Olabisi, O.; and Agrawal, A. 2024. Understanding Position Bias Effects on Fairness in Social Multi-Document Summarization. In Scherrer, Y.; Jauhiainen, T.; Ljube s i \'c , N.; Zampieri, M.; Nakov, P.; and Tiedemann, J., eds., Proceedings of the Eleventh Workshop on NLP for Similar Languages, Varieties, and Dialects (VarDial 2024), 117--129. Mexico City, M...

  58. [58]

    U.; Soroush, A.; Sakhuja, A.; Freeman, R.; Horowitz, C

    Omar, M.; Sorin, V.; Agbareia, R.; Apakama, D. U.; Soroush, A.; Sakhuja, A.; Freeman, R.; Horowitz, C. R.; Richardson, L. D.; Nadkarni, G. N.; et al. 2025. Evaluating and addressing demographic disparities in medical large language models: a systematic review. International Journal for Equity in Health, 24(1): 57

  59. [59]

    C.; Burgoine, T.; Sharp, S

    Patterson, R.; Ogilvie, D.; Hoenink, J. C.; Burgoine, T.; Sharp, S. J.; Hajna, S.; and Panter, J. 2025. Combined associations of takeaway food availability and walkability with adiposity: Cross-sectional and longitudinal analyses. Health & Place, 91: 103405

  60. [60]

    Peters, U.; and Chin-Yee, B. 2025. Generalization Bias in Large Language Model Summarization of Scientific Research. arXiv:2504.00025

  61. [61]

    D.; Mancino, M

    Ponzo, V.; Goitre, I.; Favaro, E.; Merlo, F. D.; Mancino, M. V.; Riso, S.; and Bo, S. 2024. Is ChatGPT an effective tool for providing dietary advice? Nutrients, 16(4): 469

  62. [62]

    D.; and Finn, C

    Rafailov, R.; Sharma, A.; Mitchell, E.; Ermon, S.; Manning, C. D.; and Finn, C. 2023. Direct preference optimization: your language model is secretly a reward model. In Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS '23. Red Hook, NY, USA: Curran Associates Inc

  63. [63]

    F.; Bansal, M.; and Dreyer, M

    Ribeiro, L. F.; Bansal, M.; and Dreyer, M. 2023. Generating summaries with controllable readability levels. arXiv preprint arXiv:2310.10623

  64. [64]

    Robinson, S.; and Brooks, R. 2024. Assumptions and simplifications in discrete-event simulation modelling. Journal of Simulation, 1--18

  65. [65]

    Z.; et al

    Sarmiento, I.; Cockcroft, A.; Dion, A.; Belaid, L.; Silver, H.; Pizarro, K.; Pimentel, J.; Tratt, E.; Skerritt, L.; Ghadirian, M. Z.; et al. 2024. Fuzzy cognitive mapping in participatory research and decision making: a practice review. Archives of Public Health, 82(1): 76

  66. [66]

    Savory, B.; Thompson, C.; Hassan, S.; Adams, J.; Amies-Cull, B.; Chang, M.; Derbyshire, D.; Keeble, M.; Liu, B.; Medina-Lara, A.; et al. 2025. ``It does help but there's a limit...'': Young people's perspectives on policies to manage hot food takeaways opening near schools. Social Science & Medicine, 368: 117810

  67. [67]

    Schaaff, K.; Reinig, C.; and Schlippe, T. 2023. Exploring ChatGPT’s empathic abilities. In 2023 11th international conference on affective computing and intelligent interaction (ACII), 1--8. IEEE

  68. [68]

    B.; Zhao, Z.; Sayin, B.; Flek, L.; and Rosso, P

    Schlicht, I. B.; Zhao, Z.; Sayin, B.; Flek, L.; and Rosso, P. 2025. Do LLMs provide consistent answers to health-related questions across languages? In European Conference on Information Retrieval, 314--322. Springer

  69. [69]

    HEART-felt Narratives: Tracing Empathy and Narrative Style in Personal Stories with LLMs

    Shen, J.; Mire, J.; Park, H. W.; Breazeal, C.; and Sap, M. 2024. Heart-felt narratives: Tracing empathy and narrative style in personal stories with llms. arXiv preprint arXiv:2405.17633

  70. [70]

    Shool, S.; Adimi, S.; Saboori Amleshi, R.; Bitaraf, E.; Golpira, R.; and Tara, M. 2025. A systematic review of large language model (LLM) evaluations in clinical medicine. BMC Medical Informatics and Decision Making, 25(1): 117

  71. [71]

    A.; and Giabbanelli, P

    Shrestha, A.; Mielke, K.; Nguyen, T. A.; and Giabbanelli, P. J. 2022. Automatically explaining a model: Using deep neural networks to generate text from causal maps. In 2022 Winter simulation conference (WSC), 2629--2640. IEEE

  72. [72]

    R.; Cole-Lewis, H.; et al

    Singhal, K.; Tu, T.; Gottweis, J.; Sayres, R.; Wulczyn, E.; Amin, M.; Hou, L.; Clark, K.; Pfohl, S. R.; Cole-Lewis, H.; et al. 2025. Toward expert-level medical question answering with large language models. Nature Medicine, 31(3): 943--950

  73. [73]

    N.; McKinnon, M

    Spreng, R. N.; McKinnon, M. C.; Mar, R. A.; and Levine, B. 2009. The Toronto Empathy Questionnaire: Scale development and initial validation of a factor-analytic solution to multiple empathy measures. Journal of personality assessment, 91(1): 62--71

  74. [74]

    St-Aubin, B.; Wainer, G.; and Loor, F. 2023. A survey of visualization capabilities for simulation environments. In 2023 Annual Modeling and Simulation Conference (ANNSIM), 13--24. IEEE Computer Society

  75. [75]

    D.; Lauf, S.; Magliocca, N

    Sun, Z.; Lorscheid, I.; Millington, J. D.; Lauf, S.; Magliocca, N. R.; Groeneveld, J.; Balbi, S.; Nolzen, H.; M \"u ller, B.; Schulze, J.; et al. 2016. Simple or complicated agent-based models? A complicated issue. Environmental Modelling & Software, 86: 56--67

  76. [76]

    Tran, D.; Dolgun, A.; and Demirhan, H. 2020. Weighted inter-rater agreement measures for ordinal outcomes. Communications in Statistics-Simulation and Computation, 49(4): 989--1003

  77. [77]

    Urlana, A.; Mishra, P.; Roy, T.; and Mishra, R. 2023. Controllable Text Summarization: Unraveling Challenges, Approaches, and Prospects--A Survey. arXiv preprint arXiv:2311.09212

  78. [78]

    van der Zee, D.-J. 2017. Approaches for simulation model simplification. In 2017 Winter Simulation Conference (WSC), 4197--4208. IEEE

  79. [79]

    H.; and Scheffer, M

    Van Nes, E. H.; and Scheffer, M. 2005. A strategy to improve the contribution of complex simulation models to ecological theory. Ecological modelling, 185(2-4): 153--164

  80. [80]

    P.; Seehofnerov \'a , A.; et al

    Van Veen, D.; Van Uden, C.; Blankemeier, L.; Delbrouck, J.-B.; Aali, A.; Bluethgen, C.; Pareek, A.; Polacin, M.; Reis, E. P.; Seehofnerov \'a , A.; et al. 2024. Adapted large language models can outperform medical experts in clinical text summarization. Nature medicine, 30(4): 1134--1142

Showing first 80 references.