REVIEW 3 major objections 6 minor 1 cited by
Towards Experience-Centered AI: A Framework for Integrating Lived Experience in Design and Development
T0 review · 3 major / 6 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read The paper argues that integrating users' lived experience into every stage of AI development produces models that reflect how people actually think, remember, and feel.
desk verdict LEAF is a genuinely useful synthesis and a real gap-filler for AI ethics/HCI, but the taxonomy miscount and the retrospective case studies need fixing before it can be more than a well-positioned proposal. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the LEAF (Lived Experience Centered AI Framework), a conceptual scaffold with two parts. One part is a taxonomy of lived-experience dimensions—sense of self, health, social and cultural identity, and learning—assembled from philosophy, psychology, education, healthcare, and human-computer interaction. The other part is a lifecycle view of AI development, with named stages where those dimensions can enter: problem definition, data curation, model design, evaluation, and post-deployment monitoring. The framework's work is translation: it converts first-person experiential knowledge into concrete design and evaluation requirements.
What would settle it
Build two versions of a clinical or educational AI system, one developed with lived-experience integration at every pipeline stage and one without, and compare user-reported trust, empathy, perceived safety, and task success in a randomized study. If the lived-experience version shows no advantage—or if documented harms such as decontextualized medical advice persist—the framework's central claim fails.
Extended reading notes
Core claim
The paper's central claim is that lived experience should be a foundational input to AI design and evaluation, not an optional ethical overlay. To make that actionable, it introduces LEAF, which maps thematic dimensions of lived experience onto five stages of the AI development pipeline: problem definition, data curation and annotation, model design, evaluation and testing, and post-deployment monitoring. The framework argues that real-world system failures arise when the experiential layer is omitted—an autograder whose false positives discourage learning, a clinical chatbot that misses practitioner expertise, or an NLP model that treats sacred texts as neutral corpora. Centering lived expe
Load-bearing premise
The load-bearing premise is that a taxonomy of lived experience assembled from psychology, education, healthcare, and social policy transfers to AI systems in its current form; the paper itself concedes in Section 5.4 that broader applicability and completeness remain to be validated.
Editorial extensions
If this is right
- Evaluation in high-stakes domains would include experiential axes—empathy, trust-building, relationship-building, and communication efficacy—alongside accuracy benchmarks.
- Autograders and educational AI would be designed with student and instructor feedback loops so grading errors do not silently shape learning behavior.
- Post-deployment monitoring would rely on community reporting tools, ethnographic studies, and user diaries, not only technical performance metrics.
- Cultural or sacred data would be treated as embedded in tradition and community meaning, not as neutral training corpora.
Reading between the lines
- A testable consequence the paper leaves implicit: systems built with lived-experience integration should show smaller performance and trust gaps across demographic groups, because the framework is explicitly aimed at surfacing harms invisible to developers.
- Operationalizing 'lived experience' will require a governance answer to who legitimately represents a community's experience; the framework recommends methods such as co-design and participatory evaluation but does not resolve this question.
- The framework could be turned into an audit checklist: for each pipeline stage, a reviewer would ask which experiential dimensions were included and how, making the proposal testable and falsifiable in practice.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes LEAF, a conceptual framework for integrating lived experience into the design, evaluation, and deployment of AI systems. It defines lived experience through philosophy, psychosocial literature, and HCI, organizes it into thematic dimensions, and maps those dimensions onto stages of an AI development pipeline. The manuscript illustrates the framework with four case studies (an autograder in education, the AMIE clinical conversational agent, sacred-text use in NLP, and multimodal task instruction) and concludes with recommendations for more empathetic, context-aware AI. The central claim is that centering lived experience can lead to AI models that more accurately reflect the retrospective, emotional, and contextual dimensions of human cognition. The contribution is conceptual rather than empirical; the authors explicitly defer large-scale validation to future work.
Significance. If the framework were coherent and operationalized, it could provide a useful vocabulary and structure for human-AI alignment and human-centered design research. The paper's interdisciplinary synthesis, its attention to pipeline stages, and its explicit acknowledgment of limitations in §§5.2–5.4 are strengths. It also correctly distinguishes the proposed lived-experience lens from purely psychological taxonomies or static checklists. However, the paper's principal contribution is a conceptual proposal, and its internal inconsistency in the core taxonomy (seven vs. four dimensions in §2.3) currently limits its scientific uptake. The case studies are retrospective reinterpretations rather than prospective demonstrations, and the abstract's claim that centered lived experience 'can lead to models that more accurately reflect' human cognition is not yet supported by evidence or a concrete evaluative procedure. These issues are fixable within the manuscript's scope, but they require substantive revision.
major comments (3)
- [§2.3, §1] The framework's central taxonomy is internally inconsistent. §2.3 states that the authors 'identify seven interrelated thematic dimensions' and will outline each, but the bulleted list that follows defines only four: Sense of Self, Health, Social and Cultural, and Learning. The subsequent paragraph on moral alignment is not presented as a dimension, and Figure 1 is not verifiable from the manuscript text, so the reader cannot determine what the other three dimensions are. §1, by contrast, describes the framework as structured around four key dimensions. Because this taxonomy is the operational core of LEAF, the framework's coverage of lived experience is indeterminate as written. The authors should either define the missing dimensions explicitly or correct the count in §2.3 and align Figure 1 and §1 accordingly.
- [§4, §5.4] Section 4 claims that the case studies 'show how integrating lived experiences during the problem definition and development stages of AI leads to more context-aware, ethical, and socially aligned systems.' However, each case study is a post-hoc reinterpretation of prior work (Li et al. 2023; Tu et al. 2025; Hutchinson 2024; Nguyen et al. 2025), not a prospective application of LEAF with a comparison condition. The 'Role of Lived Experience' subsections state what LEAF would add, but no evidence is presented that applying LEAF changes model behavior, outputs, or downstream outcomes. Section 5.4 itself concedes that broader applicability 'remains to be discussed and validated.' This is load-bearing because the manuscript's central claim is that centering lived experience can lead to more accurate reflection of human cognition. I recommend reframing the case studies as illustrative rather
- [§3, §5.4] The LEAF framework is presented as actionable, but §3 does not specify how the proposed dimensions map onto the five pipeline stages. For example, §3.3 recommends participatory design and human-in-the-loop methods, but it does not say which lived-experience dimension should be elicited at which stage, how to resolve conflicts among stakeholders' experiences, or what criteria would indicate successful integration. Without such operationalization, the framework risks becoming a general call for user involvement rather than a distinct 'experience-centered' framework. The paper itself notes in §5.4 that metrics and validation are absent; at minimum, a worked instantiation—one domain with explicit dimension-to-stage mappings and success measures—would make the framework actionable and testable.
minor comments (6)
- [§1] Typo: 'the the view from inside' should read 'the view from inside.'
- [§4.2] The text refers to 'AIME outperformed primary care physicians'; the system is named AMIE (Articulate Medical Intelligence Explorer).
- [§3.5] Typo: 'lived experiments into post deployment monitoring' should be 'lived experiences into post-deployment monitoring.'
- [Abstract, §4] The abstract lists three application domains (education, healthcare, cultural alignment), but §4 presents four case studies, including 'AI for Task Instruction.' Align the counts or list all four domains.
- [§2.3] The paragraph on 'moral alignment' reads as an unnumbered fifth dimension. Either integrate it explicitly into one of the four dimensions or remove it to avoid further ambiguity about the taxonomy's structure.
- [Figure 1] Figure 1 is referenced but not reproducible from the manuscript text; ensure the figure is included and that its content matches the enumerated dimensions.
Circularity Check
LEAF is a non-circular synthesis of external literature; only a minor definitional overlap where the claimed benefit restates the paper's own definition of lived experience.
-
self definitional
[Abstract; §1 Introduction; §2.3 Dimensions]
"Abstract: 'arguing that centering lived experience can lead to models that more accurately reflect the retrospective, emotional, and contextual dimensions of human cognition.' §1: 'lived experience — defined as retrospective, emotional, and contextual — as foundational to AI system design and evaluation.'"
The central benefit (models that reflect retrospective, emotional, and contextual dimensions) is the same triple used in the paper's own definition of lived experience. The claimed outcome thus restates the input definition rather than being established by independent evidence. This is mild because the framework's actual content (taxonomy, pipeline stages, case studies) is synthesized from external literature and does not depend on this overlap. No equations, fitted parameters, or empirical predictions occur, so this is the only circular element.
full rationale
No significant circularity. The paper is a conceptual position/framework proposal, not a derivation: there are no equations, no fitted parameters, and no quantitative predictions whose values are forced by inputs. Its central construct—lived experience—is grounded in external literature (phenomenology, psychology, education, healthcare), and the LEAF framework explicitly builds on prior independent frameworks (participatory design, value-sensitive design, model cards, etc.) rather than importing its conclusion from the authors' own prior claims. The several self-citations (Chandra et al. 2025a,b; Song et al. 2024; Jin et al. 2024) are used as empirical evidence of AI harms and limitations; none assumes LEAF or the paper's thesis, so per the rules they count as independent support and do not raise the score. Section 5.4 honestly concedes that broader applicability 'remains to be discussed and validated,' which undercuts any charge that the framework overclaims. Two weaknesses are real but are not circularity: §2.3 promises 'seven interrelated thematic dimensions' yet defines only four (Sense of Self, Health, Social and Cultural, Learning), leaving the taxonomy's coverage indeterminate; and the case studies in §4 are retrospective reinterpretations of prior work rather than prospective tests of LEAF. Both are correctness/validity risks, not input-equals-output reductions. The only mild circular step is the definitional overlap flagged above: the paper defines lived experience as 'retrospective, emotional, and contextual' and then states that centering it yields models reflecting precisely those dimensions, making that headline benefit true by construction. Because the framework's substance is externally sourced and the claim is not used to justify the framework itself, the overall circularity score is 1.
Assumptions & free parameters
assumptions (3)
- domain assumption Lived experience is a well-defined construct with universal relevance to AI design.
- domain assumption Including lived experience in the AI lifecycle will improve alignment, trust, and wellbeing.
- domain assumption The five pipeline stages are the right and sufficient insertion points for lived experience.
Cite this review
Pith. "Pith review of Towards Experience-Centered AI: A Framework for Integrating Lived Experience in Design and Development." pith.science (2026). https://pith.science/paper/EKGVU4GK
@misc{pith2026250806849,
author = {Pith},
title = {Pith review of: Towards Experience-Centered AI: A Framework for Integrating Lived Experience in Design and Development},
year = {2026},
howpublished = {\url{https://pith.science/paper/EKGVU4GK}},
note = {Machine review of arXiv:2508.06849}
}
read the original abstract
Lived experiences fundamentally shape how individuals interact with AI systems, influencing perceptions of safety, trust, and usability. While prior research has focused on developing techniques to emulate human preferences, and proposed taxonomies to categorize risks (such as psychological harms and algorithmic biases), these efforts have provided limited systematic understanding of lived human experiences or actionable strategies for embedding them meaningfully into the AI development lifecycle. This work proposes a framework for meaningfully integrating lived experience into the design and evaluation of AI systems. We synthesize interdisciplinary literature across lived experience philosophy, human-centered design, and human-AI interaction, arguing that centering lived experience can lead to models that more accurately reflect the retrospective, emotional, and contextual dimensions of human cognition. Drawing from a wide body of work across psychology, education, healthcare, and social policy, we present a targeted taxonomy of lived experiences with specific applicability to AI systems. To ground our framework, we examine three application domains (i) education, (ii) healthcare, and (iii) cultural alignment, illustrating how lived experience informs user goals, system expectations, and ethical considerations in each context. We further incorporate insights from AI system operators and human-AI partnerships to highlight challenges in responsibility allocation, mental model calibration, and long-term system adaptation. We conclude with actionable recommendations for developing experience-centered AI systems that are not only technically robust but also empathetic, context-aware, and aligned with human realities. This work offers a foundation for future research that bridges technical development with the lived experiences of those impacted by AI systems.
Figures
Forward citations
Cited by 1 Pith paper
-
The Consensus Trap: Dissecting Subjectivity and the "Ground Truth" Illusion in Data Annotation
A literature review concludes that pursuing consensus in data annotation creates biased AI by dismissing subjective disagreements and enforcing geographic hegemony, and proposes mapping diversity instead.
Reference graph
Works this paper leans on
-
[5]
Journal of Medical Internet Research, 27: e67485
AI for IMPACTS Framework for Evaluating the Long-Term Real-World Impacts of AI-Powered Clinician Tools: Systematic Review and Narrative Synthesis. Journal of Medical Internet Research, 27: e67485. Jin, Y .; Chandra, M.; Verma, G.; Hu, Y .; De Choudhury, M.; and Kumar, S. 2024. Better to ask in english: Cross-lingual evaluation of large language models for...
work page 2024
-
[8]
Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464): 447–453. Olawade, D. B.; Wada, O. Z.; Odetayo, A.; David-Olawade, A. C.; Asaolu, F.; and Eberhardt, J. 2024. Enhancing mental health with Artificial Intelligence: Current trends and future prospects. Journal of medicine, surgery, and public health, 100099....
work page 2024
-
[9]
arXiv preprint arXiv:2501.01056
Risks of Cultural Erasure in Large Language Models. arXiv preprint arXiv:2501.01056. Raees, M.; Meijerink, I.; Lykourentzou, I.; Khan, V .-J.; and Papangelis, K. 2024. From explainable to interactive AI: A literature review on current trends in human-AI interaction. International Journal of Human-Computer Studies, 103301. Raji, I. D.; Smart, A.; White, R....
arXiv 2024
-
[14]
arXiv preprint arXiv:2401.14362
The typing cure: Experiences with large language model chatbots for mental health support. arXiv preprint arXiv:2401.14362. Soon, Y . E.; Murray, C. M.; Aguilar, A.; and Boshoff, K
-
[15]
International Journal of Nursing Studies, 109: 103619
Consumer involvement in university education pro- grams in the nursing, midwifery, and allied health profes- sions: a systematic scoping review. International Journal of Nursing Studies, 109: 103619. Spiegelberg, E. 2012. The phenomenological movement: A historical introduction, volume 5. Springer Science & Busi- ness Media. Staff, M. 2016. In-Depth: How ...
arXiv 2012
-
[16]
Towards conversational diagnostic artificial intelli- gence. Nature, 1–9. Turchin, A. 2019. AI alignment problem:“human values” don’t actually exist. arXiv. Umbrello, S.; and De Bellis, A. F. 2018. A value-sensitive design approach to intelligent agents. In Artificial intel- ligence safety and security , 395–409. Chapman and Hal- l/CRC. van der Maden, W.;...
work page 2019
-
[17]
Sharing of cultural values and heritage through story- telling in the digital age. Frontiers in Psychology, 14. Sec- tion: Educational Psychology
-
[1086]
CRC Press. Russon, J. 2003. Human experience: Philosophy, neurosis, and the elements of everyday life. SUNY Press. Sartor, C. 2023. Mental health and lived experience: The value of lived experience expertise in global mental health. Cambridge Prisms: Global Mental Health, 10: e38. Schaub, F.; K ¨onings, B.; and Weber, M. 2015. Context- adaptive privacy: L...
work page 2003
Show all 17 references
-
[2014]
InProceedings of the SIGCHI Conference on Human Factors in Computing Systems, CHI ’14, 1163–1172
Personal tracking as lived informatics. InProceedings of the SIGCHI Conference on Human Factors in Computing Systems, CHI ’14, 1163–1172. New York, NY , USA: Asso- ciation for Computing Machinery. ISBN 9781450324731. Rosson, M. B.; and Carroll, J. M. 2007. Scenario-based de- s...
2007
-
[2017]
Human–Computer Interac- tion, 32(5–6): 197–207
Introduction to This Special Issue on the Lived Expe- rience of Personal Informatics. Human–Computer Interac- tion, 32(5–6): 197–207. Costanza-Chock, S. 2020. Design justice: Community-led practices to build the worlds we need. The MIT Press. De, A.; Kanthawala, S.; and Maddox...
2020 arXiv
-
[2018]
It happened to be the perfect thing
Safe Spaces and Safe Places: Unpacking Technology- Mediated Experiences of Safety and Harm with Transgender People. Proc. ACM Hum.-Comput. Interact., 2(CSCW). Sch¨on, D. A. 1979. The reflective practitioner. New York. Schraw, G.; and Dennison, R. S. 1994. Assessing metacog- ni...
1979
-
[2019]
In Proceedings of the conference on fairness, accountability, and trans- parency, 220–229
Model cards for model reporting. In Proceedings of the conference on fairness, accountability, and trans- parency, 220–229. Mohan, S.; Ramea, K.; Price, B.; Shreve, M.; Eldardiry, H.; and Nelson, L. 2019. Building Jarvis-A Learner-Aware Con- versational Trainer. In IUI Worksho...
2019 arXiv
-
[2020]
Lived Ex- perience
Closing the AI accountability gap: Defining an end- to-end framework for internal algorithmic auditing. In Pro- ceedings of the 2020 conference on fairness, accountability, and transparency, 33–44. Reed, R. 2021. AI in Religion, AI for Religion, AI and Re- ligion: Towards a th...
2020
-
[2021]
Frontiers in ar- tificial intelligence, 4: 622364
Human-versus artificial intelligence. Frontiers in ar- tificial intelligence, 4: 622364. Kruks, S. 1992. Gender and subjectivity: Simone de Beau- voir and contemporary feminism. Signs: Journal of Women in Culture and Society, 18(1): 89–110. Kruks, S. 2014. Women’s ‘lived exper...
1992 arXiv
-
[2023]
Journal of Biomedical Informatics, 137: 104274
Benchmark datasets driving artificial intelligence de- velopment fail to capture the needs of medical professionals. Journal of Biomedical Informatics, 137: 104274. Boud, D.; Keogh, R.; and Walker, D. 2013.Reflection: Turn- ing experience into learning. Routledge. Bravansky, M...
2013 arXiv
-
[2024]
Antony, M
Do Large Language Models Discriminate in Hir- ing Decisions on the Basis of Race, Ethnicity, and Gender? arXiv preprint arXiv:2406.10486. Antony, M. V . 2001. Is ‘consciousness’ ambiguous?Journal of Consciousness Studies, 8(2): 19–44. Armutat, S.; Wattenberg, M.; and Mauritz, ...
2001 arXiv
-
[2025]
Nature Medicine, 31(3): 743–744
Medical large language model for diagnostic reason- ing across specialties. Nature Medicine, 31(3): 743–744. Adeofe, L. 1995. Artificial intelligence and subjective expe- rience. In Proceedings of Southcon ’95, 403–408. Afroogh, S.; Akbari, A.; Malone, E.; Kargar, M.; and Alam...
1995
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.