REVIEW 4 major objections 6 minor 93 references
This paper proposes a two-step mixed-methods framework that first measures what clinicians, patients, policymakers, and other stakeholders need from explanations of health simulations, then steers LLMs to write summaries matching each group
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-05 05:55 UTC pith:IBEYBLQW
load-bearing objection A clear, honest vision paper for stakeholder-tailored LLM summaries of health simulations; the construct-validity gap between 'information needs' and the empathy measure is the main thing to fix before this becomes a validated pipeline. the 4 major comments →
Towards Personalized Explanations for Health Simulations: A Mixed-Methods Framework for Stakeholder-Centric Summarization
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that the barrier to using health simulations is not only technical but communicative: we lack a systematic account of what different stakeholders need from an explanation. The authors' main contribution is a vision and framework for eliciting, incorporating, and evaluating those needs. The framework has two broad steps. First, the model structure is decomposed into small structured pieces (such as RDF triplets) and translated into text, while simulation outputs are summarized via statistical analysis or multimodal LLMs; candidate summaries are then generated by varying content and style in a designed experiment, checked for factuality by modelers, and piloted wit
What carries the argument
The mechanism is a closed loop: model decomposition into structured triplets small enough for an LLM; factorial design generating candidate summaries that vary controllable attributes such as length, tone, and topic coverage; a modeler factuality check covering knowledge, reasoning, and relevance errors; validated empathy instruments (the Toronto Empathy Questionnaire and the State Empathy Scale) capturing reader reactions; factorial analysis with effect sizes, power analysis, and repeated-measures ANOVA identifying group preferences; preference optimization (DPO and similar alignment methods) steering the LLM; and automatic LLM-based evaluation metrics that filter candidates before improved
Load-bearing premise
The framework assumes that stakeholder groups have stable, distinct preferences for summary content and style that can be elicited reliably with a 2x2 factorial design plus empathy questionnaires, and then used to steer an LLM—an assumption the seven-person pilot does not test.
What would settle it
Run the full loop with two stakeholder groups on the same health simulation. If the factorial analysis finds no significant between-group differences in empathy scores across the four content-by-style summaries, or if a retest one month later reverses each group's preferred combination, the central premise of stable group-level tailoring fails.
If this is right
- Different stakeholder groups can receive summaries with the same factual core but different content coverage and style, replacing the single generic text now produced.
- Factuality becomes a gate: no candidate summary reaches stakeholders until modelers have checked it for knowledge, reasoning, and relevance errors.
- Because preferences are analyzed per group and the LLM is steered accordingly, the approach extends existing model-to-text pipelines without retraining from scratch; few-shot examples can suffice for domain adaptation.
- Evaluation shifts from text quality alone to whether tailored summaries change downstream decisions, ideally tested by a randomized controlled trial comparing generic versus tailored summaries.
- The same pipeline can be re-run for new models or new stakeholder groups, making the process repeatable rather than a one-off customization.
Where Pith is reading between the lines
- If preferences are not stable across time or contexts, group-level tailoring may need to become individual-level or dynamically updated; the seven-person pilot does not yet rule this out.
- Empathy is the only reaction dimension the protocol measures; trust, perceived accuracy, cognitive load, and actionability may also drive preferences and would not be captured directly.
- A stronger test than preference alignment is behavioral: tailored summaries should change the decisions stakeholders actually make. A vignette study comparing choices after generic versus tailored summaries would be a cheap precursor to the named randomized controlled trial.
- The reliance on empathy instruments assumes textual summaries are evaluated emotionally; administrative or technical audiences may prefer a low-empathy executive style, which would complicate the factorial interpretation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that existing LLM-based explanations of health simulations are one-size-fits-all and proposes a two-step framework to elicit stakeholder-specific content and style preferences and to steer LLM summarization accordingly. Step 1 decomposes simulation models, generates candidate summaries via a 2x2 factorial design over content and style, checks factuality with modelers, and pilots the summaries with stakeholders using empathy questionnaires. Step 2 analyzes factorial data per stakeholder group, steers LLM generation via preference optimization, evaluates with LLM-based and human metrics, and shares optimized summaries back to participants. The authors explicitly position the contribution as a vision and framework, and they defer empirical validation to future work, including ablation studies and a downstream randomized controlled trial.
Significance. If the framework works as intended, it would address a real gap by making health simulation outputs accessible to diverse stakeholders in a tailored way. The paper's strengths are its clear articulation of a repeatable pipeline, the inclusion of a modeler factuality gate, the use of designed experiments for preference elicitation, and an unusually candid discussion of limitations, including an explicit call for ablation studies and downstream decision metrics. However, the claimed ability to elicit content and style needs is not demonstrated, and the key measurement choice (empathy) lacks construct-validity evidence. As submitted, the paper is best read as a research proposal rather than a validated method; its value will depend on the planned empirical studies.
major comments (4)
- [Step 1: Identify Information Needs and Preferred Styles From Different Stakeholder Groups] The central claim is that the framework 'elicit[s] ... style and content needs'; however, the only outcome measure specified for the stakeholder pilot is empathy (TEQ and State Empathy Scale). The paper does not justify why maximizing state empathy is equivalent to satisfying a stakeholder's information needs; indeed, the hospital-administrator example (executive summaries, bullet points, business-oriented language) suggests a summary can be appropriate without being maximally empathy-inducing. Because Step 2's factorial analysis and LLM steering optimize the measured outcome, this construct-validity gap is load-bearing. Please either add direct measures of perceived usefulness, comprehension, or decision quality, or reposition the framework as one for empathy-oriented summaries, and justify the use of a single affective outcome across heterogeneous stakeholder groups.
- [Step 1: Identify Information Needs and Preferred Styles From Different Stakeholder Groups] The text states: 'Since perceiving a narrative as immersive and compelling depends on the mental state of the reader, we recommend using the validated Toronto Empathy Questionnaire (TEQ; 16 items) to obtain multiple empathy measures.' TEQ is a trait empathy scale, not a state measure of narrative response; using it in this role is likely a misapplication. The State Empathy Scale is the appropriate state measure. Please clarify that TEQ is a baseline covariate and remove the implication that TEQ measures the reader's state in response to a summary.
- [Step 2: Optimize the Alignment of Language Models and Stakeholder Communication Needs] The optimization loop is underspecified. The text says 'An optimization process is involved, as the new summaries should be automatically assessed and the architecture adjusted if the scores are insufficient,' but no objective function, score thresholds, or adjustment rules are given. For a framework that claims to be a repeatable process, this prevents replication. Specify at a conceptual level what is optimized (e.g., which evaluation metrics serve as rewards), what 'insufficient' means, and which architectural parameters (RAG parameters, temperature, or others) are adjusted.
- [Discussion: On Participatory AI] The paper itself concedes that 'The scientific basis for this framework will be strengthened by collecting experimental data, performing ablation studies...' and that 'the ultimate demonstration that the pipeline works lies in its ability to affect decisions.' I agree. As submitted, the only empirical result is a pilot with n=7 reporting a median response time of 19.16 minutes, which does not test preference validity, summary quality, or decision impact. The authors should either explicitly scope the paper as a position/vision paper with no empirical claims, making the title and abstract match, or include a proof-of-concept with a small but substantive evaluation.
minor comments (6)
- [Proposed Framework, Step 1] The factorial design is referred to as '22 factorial design'; this should be typeset as 2^2 factorial design.
- [Proposed Framework, Step 2] Typo: 'repeated measures ANOV A' should read 'repeated measures ANOVA'.
- [Author affiliations] There is a spacing issue in 'Old Dominion University / 1030 University Blvd, Suffolk, V A 23435, USA' — 'V A' should be 'VA'.
- [References] The reference 'Ahrweiler et al. 2019' has 'policy advise'; likely should be 'policy advice'.
- [Figure 3] The feedback loop between 'automated assessment' and 'architecture adjustment' would be easier to follow if the figure annotated the specific metrics and parameters involved.
- [Proposed Framework, Step 1] The internal loop that asks participants for their preferences and then evaluates whether summaries match those preferences is not inherently circular, because modeler factuality checks and the proposed downstream RCT provide independent anchors. The paper would benefit from stating this explicitly to preempt concerns about self-reference.
Circularity Check
No significant circularity: the paper is a framework proposal with an iterative feedback loop and an external RCT anchor; self-citations are used as ordinary building blocks, not as load-bearing self-justification.
full rationale
The paper does not derive quantitative predictions or present equations that reduce to fitted parameters; it proposes a mixed-methods framework. The central loop—generating candidate summaries, measuring participant reactions, factorial-analyzing those reactions, steering the LLM, and re-assessing—is an iterative user-centered design process, not a claim that a prediction is validated by its own inputs. The only pilot (n=7) is reported as response time, not as evidence that empathy scores validate content/style preferences, and the paper explicitly defers causal validation to a future randomized controlled trial measuring downstream decisions, providing an external anchor. Self-citations (e.g., Gandee and Giabbanelli 2024 for model-to-text decomposition; Giabbanelli et al. 2024a for few-shot saturation; Giabbanelli et al. 2025 for empathy-based storytelling) are used as prior empirical building blocks with alternatives listed (e.g., RDF Walks), so they are not load-bearing uniqueness arguments. The construct-validity concern—that empathy questionnaires are used as a proxy for information needs—is a substantive measurement assumption, but it is not circular in the logical sense: 'preferred' is operationalized via empathy, but this is an explicit modeling choice, not a derivation that makes the framework's output equivalent to its input. The paper itself acknowledges its limitations, stating that 'The scientific basis for this framework will be strengthened by collecting experimental data, performing ablation studies...' (Discussion). Thus no circular step meets the evidentiary bar required by the review rules.
Axiom & Free-Parameter Ledger
axioms (4)
- domain assumption Distinct stakeholder groups have stable, divergent preferences for summary content and style that can be elicited and generalized.
- domain assumption The Toronto Empathy Questionnaire and State Empathy Scale are valid proxies for whether a summary is understandable and actionable.
- ad hoc to paper A 2x2 factorial design over content and style captures the variation in stakeholder needs.
- domain assumption LLMs can produce factually correct summaries if model decomposition, few-shot fine-tuning, and modeler review are applied.
Cite this review
Pith. "Pith review of Towards Personalized Explanations for Health Simulations: A Mixed-Methods Framework for Stakeholder-Centric Summarization." pith.science (2026). https://pith.science/paper/IBEYBLQW
@misc{pith2026250904646,
author = {Pith},
title = {Pith review of: Towards Personalized Explanations for Health Simulations: A Mixed-Methods Framework for Stakeholder-Centric Summarization},
year = {2026},
howpublished = {\url{https://pith.science/paper/IBEYBLQW}},
note = {Machine review of arXiv:2509.04646}
}
read the original abstract
Modeling & Simulation (M&S) approaches such as agent-based models hold significant potential to support decision-making activities in health, with recent examples including the adoption of vaccines, and a vast literature on healthy eating behaviors and physical activity behaviors. These models are potentially usable by different stakeholder groups, as they support policy-makers to estimate the consequences of potential interventions and they can guide individuals in making healthy choices in complex environments. However, this potential may not be fully realized because of the models' complexity, which makes them inaccessible to the stakeholders who could benefit the most. While Large Language Models (LLMs) can translate simulation outputs and the design of models into text, current approaches typically rely on one-size-fits-all summaries that fail to reflect the varied informational needs and stylistic preferences of clinicians, policymakers, patients, caregivers, and health advocates. This limitation stems from a fundamental gap: we lack a systematic understanding of what these stakeholders need from explanations and how to tailor them accordingly. To address this gap, we present a step-by-step framework to identify stakeholder needs and guide LLMs in generating tailored explanations of health simulations. Our procedure uses a mixed-methods design by first eliciting the explanation needs and stylistic preferences of diverse health stakeholders, then optimizing the ability of LLMs to generate tailored outputs (e.g., via controllable attribute tuning), and then evaluating through a comprehensive range of metrics to further improve the tailored generation of summaries.
Figures
Reference graph
Works this paper leans on
-
[1]
Ahmed, R.; and Hemanth, D. J. 2025. Hybrid text summarization: Integrating extractive and abstractive models for enhanced cross-domain summarization. Intelligent Decision Technologies, 18724981251322745
2025
-
[2]
M.; Capellas, B
Ahrweiler, P.; Sp \"a th, E.; Siqueiros Garc \' a, J. M.; Capellas, B. L.; and Wurster, D. 2025. Inclusive technology co-design for participatory AI. Participatory Artificial Intelligence in Public Social Services: From Bias to Fairness in Assessing Beneficiaries, 35--62
2025
-
[3]
Ahrweiler, P.; et al. 2019. Co-designing social simulation models for policy advise: lessons learned from the INFSO-SKIN study. In 2019 Spring simulation conference (SpringSim), 1--12. IEEE
2019
-
[4]
Aminpour, P.; Schwermer, H.; and Gray, S. 2021. Do social identity and cognitive diversity correlate in environmental stakeholders? A novel approach to measuring cognitive distance within and between groups. Plos one, 16(11): e0244907
2021
-
[5]
D.; Papandrianos, N
Apostolopoulos, I. D.; Papandrianos, N. I.; Papathanasiou, N. D.; and Papageorgiou, E. I. 2024. Fuzzy cognitive map applications in medicine over the last two decades: A review study. Bioengineering, 11(2): 139
2024
-
[6]
L.; and Giabbanelli, P
Baniukiewicz, M.; Dick, Z. L.; and Giabbanelli, P. J. 2018. Capturing the fast-food landscape in England using large-scale network analysis. EPJ Data Science, 7(1): 39
2018
-
[7]
A.; Wornow, M.; Swaminathan, A.; Lehmann, L
Bedi, S.; Liu, Y.; Orr-Ewing, L.; Dash, D.; Koyejo, S.; Callahan, A.; Fries, J. A.; Wornow, M.; Swaminathan, A.; Lehmann, L. S.; et al. 2025. Testing and evaluation of health care applications of large language models: a systematic review. Jama
2025
-
[8]
K.; Zaghir, J.; Zheng, Y.; Bensahla, A.; Bjelogrlic, M.; and Lovis, C
Bednarczyk, L.; Reichenpfader, D.; Gaudet-Blavignac, C.; Ette, A. K.; Zaghir, J.; Zheng, Y.; Bensahla, A.; Bjelogrlic, M.; and Lovis, C. 2025. Scientific evidence for clinical text summarization using large language models: scoping review. Journal of Medical Internet Research, 27: e68998
2025
-
[9]
Belfrage, M.; et al. 2024. Simulating change: A systematic literature review of agent-based models for policy-making. In 2024 Annual Modeling and Simulation Conference (ANNSIM), 1--13. IEEE
2024
-
[10]
J.; and Tobias, A
Brooks, R. J.; and Tobias, A. M. 1996. Choosing the best model: Level of detail, complexity, and model performance. Mathematical and computer modelling, 24(4): 1--14
1996
-
[11]
H.; Kader, R.; Ortiz-Prado, E.; Makowski, M
Busch, F.; Hoffmann, L.; Rueger, C.; van Dijk, E. H.; Kader, R.; Ortiz-Prado, E.; Makowski, M. R.; Saba, L.; Hadamitzky, M.; Kather, J. N.; et al. 2025. Current applications and challenges in large language models for patient care: a systematic review. Communications Medicine, 5(1): 26
2025
-
[12]
J.; and Gotz, D
Caban, J. J.; and Gotz, D. 2015. Visual analytics in healthcare--opportunities and research challenges. Journal of the American Medical Informatics Association, 22(2): 260--262
2015
-
[13]
Y.; et al
Chu, S. Y.; et al. 2025. Think together and work better: Combining humans' and LLMs' think-aloud outcomes for effective text evaluation. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, 1--23
2025
-
[14]
Dolha, D. N.; and Buchmann, R. A. 2024. Generative AI for BPMN process analysis: experiments with multi-modal process representations. In International Conference on Business Informatics Research, 19--35. Springer
work page 2024
-
[15]
Dong, H.; Xiong, W.; Goyal, D.; Zhang, Y.; Chow, W.; Pan, R.; Diao, S.; Zhang, J.; Shum, K.; and Zhang, T. 2023. Raft: Reward ranked finetuning for generative foundation model alignment. arXiv preprint arXiv:2304.06767
Pith/arXiv arXiv 2023
- [16]
-
[17]
Fahland, D.; Fournier, F.; Limonad, L.; Skarbovsky, I.; and Swevels, A. J. 2024. How well can large language models explain business processes? arXiv preprint arXiv:2401.12846
Pith/arXiv arXiv 2024
-
[18]
Fedeli, A.; and Manrique Negrin, D. A. 2024. Towards a collaborative approach for Digital Twin simulation models comprehension. In Proceedings of the ACM/IEEE 27th International Conference on Model Driven Engineering Languages and Systems, MODELS Companion '24, 660–664. New York, NY, USA: Association for Computing Machinery. ISBN 9798400706226
work page 2024
-
[19]
Ferrand, N.; Hassenforder, E.; and Girard, S. 2024. Engineering participation: Preparing and designing a participatory process. Transformative Participation for Socio-Ecological Sustainability-Around the CoOPLAGE pathways, 109--121
work page 2024
-
[20]
Fraile Navarro, D.; Coiera, E.; Hambly, T. W.; Triplett, Z.; Asif, N.; Susanto, A.; Chowdhury, A.; Azcoaga Lorenzo, A.; Dras, M.; and Berkovsky, S. 2025. Expert evaluation of large language models for clinical dialogue summarization. Scientific reports, 15(1): 1195
work page 2025
-
[21]
Gandee, T. J.; and Giabbanelli, P. J. 2024. Combining natural language generation and graph algorithms to explain causal maps through meaningful paragraphs. In International Conference on Conceptual Modeling, 359--376. Springer
work page 2024
- [22]
-
[23]
Gao, M.; Ruan, J.; Sun, R.; Yin, X.; Yang, S.; and Wan, X. 2023. Human-like summarization evaluation with chatgpt. arXiv preprint arXiv:2304.02554
Pith/arXiv arXiv 2023
-
[24]
Ghaffarzadegan, N.; Lyneis, J.; and Richardson, G. P. 2011. How small system dynamics models can help the public policy process. System Dynamics Review, 27(1): 22--44
work page 2011
-
[25]
Giabbanelli, P.; Phatak, A.; Mago, V.; and Agrawal, A. 2024 a . Narrating Causal Graphs with Large Language Models. In Hawaii International Conference on System Sciences 2024 (HICSS-57)
work page 2024
-
[26]
Giabbanelli, P. J.; and Baniukiewicz, M. 2018. Navigating complex systems for policymaking using simple software tools. In Advanced data analytics in health, 21--40. Springer
work page 2018
-
[27]
Giabbanelli, P. J.; and Baniukiewicz, M. 2019. Visual analytics to identify temporal patterns and variability in simulations from cellular automata. ACM Transactions on Modeling and Computer Simulation (TOMACS), 29(1): 1--26
work page 2019
-
[28]
Giabbanelli, P. J.; Daumas, C.; Flandre, N. Y.; Pitkar, A.; and Vazquez-Estrada, J. 2025. Promoting empathy in decision-making by turning agent-based models into stories using large-language models. Journal of Simulation
work page 2025
-
[29]
Giabbanelli, P. J.; and Vesuvala, C. X. 2023. Human factors in leveraging systems science to shape public policy for obesity: A usability study. Information, 14(3): 196
work page 2023
- [30]
-
[31]
Gierend, K.; Kr \"u ger, F.; Genehr, S.; Hartmann, F.; Siegel, F.; Waltemath, D.; Ganslandt, T.; and Zeleke, A. A. 2024. Provenance information for biomedical data and workflows: Scoping review. Journal of medical Internet research, 26: e51297
work page 2024
-
[32]
Gravel, J.; D’Amours-Gravel, M.; and Osmanlliu, E. 2023. Learning to fake it: limited responses and fabricated references provided by ChatGPT for medical questions. Mayo Clinic Proceedings: Digital Health, 1(3): 226--234
work page 2023
-
[33]
Gray, S. A.; Gray, S.; Cox, L. J.; and Henly-Shepard, S. 2013. Mental modeler: a fuzzy-logic cognitive mapping modeling tool for adaptive environmental management. In 2013 46th Hawaii international conference on system sciences, 965--973. IEEE
work page 2013
-
[34]
Haddad, E.; and Bugarin, K. 2020. Crisis Control: The Use of Simulations for Policy Decisionmaking. Policy Brief, PB 20, 38
work page 2020
-
[35]
H \"a m \"a l \"a inen, R. P.; Luoma, J.; and Saarinen, E. 2013. On the importance of behavioral operational research: The case of understanding and communicating about dynamic systems. European Journal of Operational Research, 228(3): 623--634
work page 2013
-
[36]
Hassan, S.; Thompson, C.; Adams, J.; Chang, M.; Derbyshire, D.; Keeble, M.; Liu, B.; Mytton, O.; Rahilly, J.; Savory, B.; et al. 2024. The adoption and implementation of local government planning policy to manage hot food takeaways near schools in England: A qualitative process evaluation. Social Science & Medicine, 362: 117431
work page 2024
-
[37]
Hayashi, H.; Budania, P.; Wang, P.; Ackerson, C.; Neervannan, R.; and Neubig, G. 2021. Wikiasp: A dataset for multi-domain aspect-based summarization. Transactions of the Association for Computational Linguistics, 9: 211--225
work page 2021
-
[38]
He, J.; Yang, Y.; Long, W.; Xiong, D.; Gutierrez-Basulto, V.; and Pan, J. Z. 2025. Evaluating and Improving Graph to Text Generation with Large Language Models. arXiv preprint arXiv:2501.14497
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[39]
Hinrichs, M.; Wang, J.; Roe, C.; and Johnston, E. W. 2025. AI Integration in Mental Health Services: Examining Trends in the USA and Peoria, Illinois. In Participatory Artificial Intelligence in Public Social Services: From Bias to Fairness in Assessing Beneficiaries, 255--275. Springer Nature Switzerland Cham
work page 2025
-
[40]
Huddleston, J.; Galgoczy, M. C.; Ghumrawi, K. A.; Giabbanelli, P. J.; Rice, K. L.; Nataraj, N.; Brown, M. M.; Harper, C. R.; and Florence, C. S. 2022. Design and Deployment of a Simulation Platform: Case Study of an Agent-Based Model for Youth Suicide Prevention. In 2022 Winter Simulation Conference (WSC), 2582--2593. IEEE
work page 2022
-
[41]
Janssen, M.; and Helbig, N. 2018. Innovating and changing the policy-cycle: Policy-makers be prepared! Government Information Quarterly, 35(4): S99--S105
work page 2018
-
[42]
Jeong, D.; Aggarwal, S.; Robinson, J.; Kumar, N.; Spearot, A.; and Park, D. S. 2023. Exhaustive or exhausting? Evidence on respondent fatigue in long surveys. Journal of Development Economics, 161: 102992
work page 2023
-
[43]
Kammler, C.; et al. 2023. Towards a Social Simulation Interaction Tool for Policy Makers—A New Research Agenda to Enable Usage of More Complex Social Simulations. In Conference of the European Social Simulation Association, 163--176. Springer
work page 2023
-
[44]
Keeble, M.; Adams, J.; Amies-Cull, B.; Chang, M.; Cummins, S.; Derbyshire, D.; Hammond, D.; Hassan, S.; Liu, B.; Medina-Lara, A.; et al. 2024. Public acceptability of proposals to manage new takeaway food outlets near schools: cross-sectional analysis of the 2021 International Food Policy Study. Cities & Health, 8(6): 1094--1107
work page 2024
-
[45]
J.; Timmons, S.; Luo, C.; and Shi, L
Khademi, A.; Zhang, D.; Giabbanelli, P. J.; Timmons, S.; Luo, C.; and Shi, L. 2018. An agent-based model of healthy eating with applications to hypertension. In Advanced Data Analytics in Health, 43--58. Springer
work page 2018
-
[46]
Kirstein, F.; Wahle, J. P.; Gipp, B.; and Ruas, T. 2025. Cads: A systematic literature review on the challenges of abstractive dialogue summarization. Journal of Artificial Intelligence Research, 82: 313--365
work page 2025
-
[47]
Lavin, E. A.; Giabbanelli, P. J.; Stefanik, A. T.; Gray, S. A.; and Arlinghaus, R. 2018. Should we simulate mental models to assess whether they agree? In Proceedings of the annual simulation symposium, 1--12
work page 2018
-
[48]
Lee, H.; Phatale, S.; Mansoor, H.; Mesnard, T.; Ferret, J.; Lu, K.; Bishop, C.; Hall, E.; Carbune, V.; Rastogi, A.; et al. 2023. Rlaif vs. rlhf: Scaling reinforcement learning from human feedback with ai feedback. arXiv preprint arXiv:2309.00267
Pith/arXiv arXiv 2023
-
[49]
Lima, F. F. d.; and Os \'o rio, F. d. L. 2021. Empathy: assessment instruments and psychometric quality--a systematic literature review with a meta-analysis of the past ten years. Frontiers in psychology, 12: 781346
work page 2021
-
[50]
Liu, Y.; Iter, D.; Xu, Y.; Wang, S.; Xu, R.; and Zhu, C. 2023. G-eval: NLG evaluation using gpt-4 with better human alignment. arXiv preprint arXiv:2303.16634
Pith/arXiv arXiv 2023
-
[51]
Luna-Reyes, L. F.; Martinez-Moyano, I. J.; Pardo, T. A.; Cresswell, A. M.; Andersen, D. F.; and Richardson, G. P. 2006. Anatomy of a group model-building intervention: Building dynamic theory from case study research. System Dynamics Review: The Journal of the System Dynamics Society, 22(4): 291--320
work page 2006
-
[52]
Manellanga, R.; and David, I. 2024. Participatory and collaborative modeling of sustainable systems: A systematic review. In Proceedings of the ACM/IEEE 27th international conference on model driven engineering languages and systems, 645--654
work page 2024
-
[53]
Mavridis, A.; Tegos, S.; Anastasiou, C.; Papoutsoglou, M.; and Meditskos, G. 2025. Large language models for intelligent RDF knowledge graph construction: results from medical ontology mapping. Frontiers in Artificial Intelligence, 8: 1546179
work page 2025
-
[54]
Montibeller, G. 2018. Behavioral challenges in policy analysis with conflicting objectives. In Recent advances in optimization and modeling of contemporary problems, 85--108. INFORMS
work page 2018
-
[55]
Mussa, O.; Rana, O.; Goossens, B.; Orozco-terWengel, P.; and Perera, C. 2024. Towards Enhancing Linked Data Retrieval in Conversational UIs Using Large Language Models. In International Conference on Web Information Systems Engineering, 246--261. Springer
work page 2024
-
[56]
Nezhad, B.; et al. 2025. Fair Summarization: Bridging Quality and Diversity in Extractive Summaries. In Prabhakaran, V.; Dev, S.; Benotti, L.; Hershcovich, D.; Cao, Y.; Zhou, L.; Cabello, L.; and Adebara, I., eds., Proceedings of the 3rd Workshop on Cross-Cultural Considerations in NLP (C3NLP 2025), 22--34. Albuquerque, New Mexico: Association for Computa...
work page 2025
-
[57]
Olabisi, O.; and Agrawal, A. 2024. Understanding Position Bias Effects on Fairness in Social Multi-Document Summarization. In Scherrer, Y.; Jauhiainen, T.; Ljube s i \'c , N.; Zampieri, M.; Nakov, P.; and Tiedemann, J., eds., Proceedings of the Eleventh Workshop on NLP for Similar Languages, Varieties, and Dialects (VarDial 2024), 117--129. Mexico City, M...
work page 2024
-
[58]
U.; Soroush, A.; Sakhuja, A.; Freeman, R.; Horowitz, C
Omar, M.; Sorin, V.; Agbareia, R.; Apakama, D. U.; Soroush, A.; Sakhuja, A.; Freeman, R.; Horowitz, C. R.; Richardson, L. D.; Nadkarni, G. N.; et al. 2025. Evaluating and addressing demographic disparities in medical large language models: a systematic review. International Journal for Equity in Health, 24(1): 57
work page 2025
-
[59]
Patterson, R.; Ogilvie, D.; Hoenink, J. C.; Burgoine, T.; Sharp, S. J.; Hajna, S.; and Panter, J. 2025. Combined associations of takeaway food availability and walkability with adiposity: Cross-sectional and longitudinal analyses. Health & Place, 91: 103405
work page 2025
-
[60]
Peters, U.; and Chin-Yee, B. 2025. Generalization Bias in Large Language Model Summarization of Scientific Research. arXiv:2504.00025
Pith/arXiv arXiv 2025
-
[61]
Ponzo, V.; Goitre, I.; Favaro, E.; Merlo, F. D.; Mancino, M. V.; Riso, S.; and Bo, S. 2024. Is ChatGPT an effective tool for providing dietary advice? Nutrients, 16(4): 469
work page 2024
-
[62]
Rafailov, R.; Sharma, A.; Mitchell, E.; Ermon, S.; Manning, C. D.; and Finn, C. 2023. Direct preference optimization: your language model is secretly a reward model. In Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS '23. Red Hook, NY, USA: Curran Associates Inc
work page 2023
-
[63]
Ribeiro, L. F.; Bansal, M.; and Dreyer, M. 2023. Generating summaries with controllable readability levels. arXiv preprint arXiv:2310.10623
Pith/arXiv arXiv 2023
-
[64]
Robinson, S.; and Brooks, R. 2024. Assumptions and simplifications in discrete-event simulation modelling. Journal of Simulation, 1--18
work page 2024
- [65]
-
[66]
Savory, B.; Thompson, C.; Hassan, S.; Adams, J.; Amies-Cull, B.; Chang, M.; Derbyshire, D.; Keeble, M.; Liu, B.; Medina-Lara, A.; et al. 2025. ``It does help but there's a limit...'': Young people's perspectives on policies to manage hot food takeaways opening near schools. Social Science & Medicine, 368: 117810
work page 2025
-
[67]
Schaaff, K.; Reinig, C.; and Schlippe, T. 2023. Exploring ChatGPT’s empathic abilities. In 2023 11th international conference on affective computing and intelligent interaction (ACII), 1--8. IEEE
work page 2023
-
[68]
B.; Zhao, Z.; Sayin, B.; Flek, L.; and Rosso, P
Schlicht, I. B.; Zhao, Z.; Sayin, B.; Flek, L.; and Rosso, P. 2025. Do LLMs provide consistent answers to health-related questions across languages? In European Conference on Information Retrieval, 314--322. Springer
work page 2025
-
[69]
HEART-felt Narratives: Tracing Empathy and Narrative Style in Personal Stories with LLMs
Shen, J.; Mire, J.; Park, H. W.; Breazeal, C.; and Sap, M. 2024. Heart-felt narratives: Tracing empathy and narrative style in personal stories with llms. arXiv preprint arXiv:2405.17633
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[70]
Shool, S.; Adimi, S.; Saboori Amleshi, R.; Bitaraf, E.; Golpira, R.; and Tara, M. 2025. A systematic review of large language model (LLM) evaluations in clinical medicine. BMC Medical Informatics and Decision Making, 25(1): 117
work page 2025
-
[71]
Shrestha, A.; Mielke, K.; Nguyen, T. A.; and Giabbanelli, P. J. 2022. Automatically explaining a model: Using deep neural networks to generate text from causal maps. In 2022 Winter simulation conference (WSC), 2629--2640. IEEE
work page 2022
-
[72]
Singhal, K.; Tu, T.; Gottweis, J.; Sayres, R.; Wulczyn, E.; Amin, M.; Hou, L.; Clark, K.; Pfohl, S. R.; Cole-Lewis, H.; et al. 2025. Toward expert-level medical question answering with large language models. Nature Medicine, 31(3): 943--950
work page 2025
-
[73]
Spreng, R. N.; McKinnon, M. C.; Mar, R. A.; and Levine, B. 2009. The Toronto Empathy Questionnaire: Scale development and initial validation of a factor-analytic solution to multiple empathy measures. Journal of personality assessment, 91(1): 62--71
work page 2009
-
[74]
St-Aubin, B.; Wainer, G.; and Loor, F. 2023. A survey of visualization capabilities for simulation environments. In 2023 Annual Modeling and Simulation Conference (ANNSIM), 13--24. IEEE Computer Society
work page 2023
-
[75]
Sun, Z.; Lorscheid, I.; Millington, J. D.; Lauf, S.; Magliocca, N. R.; Groeneveld, J.; Balbi, S.; Nolzen, H.; M \"u ller, B.; Schulze, J.; et al. 2016. Simple or complicated agent-based models? A complicated issue. Environmental Modelling & Software, 86: 56--67
work page 2016
-
[76]
Tran, D.; Dolgun, A.; and Demirhan, H. 2020. Weighted inter-rater agreement measures for ordinal outcomes. Communications in Statistics-Simulation and Computation, 49(4): 989--1003
work page 2020
-
[77]
Urlana, A.; Mishra, P.; Roy, T.; and Mishra, R. 2023. Controllable Text Summarization: Unraveling Challenges, Approaches, and Prospects--A Survey. arXiv preprint arXiv:2311.09212
work page internal anchor Pith review Pith/arXiv arXiv 2023
-
[78]
van der Zee, D.-J. 2017. Approaches for simulation model simplification. In 2017 Winter Simulation Conference (WSC), 4197--4208. IEEE
work page 2017
-
[79]
Van Nes, E. H.; and Scheffer, M. 2005. A strategy to improve the contribution of complex simulation models to ecological theory. Ecological modelling, 185(2-4): 153--164
work page 2005
-
[80]
P.; Seehofnerov \'a , A.; et al
Van Veen, D.; Van Uden, C.; Blankemeier, L.; Delbrouck, J.-B.; Aali, A.; Bluethgen, C.; Pareek, A.; Polacin, M.; Reis, E. P.; Seehofnerov \'a , A.; et al. 2024. Adapted large language models can outperform medical experts in clinical text summarization. Nature medicine, 30(4): 1134--1142
work page 2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.