Pith. sign in

REVIEW 2 major objections 7 minor 12 references

Towards culturally-appropriate conversational AI for health in the majority world: An exploratory study with citizens and professionals in Latin America

T0 review · 2 major / 7 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Health chatbots in Latin America will need to model families and material realities, not just language and values, this participatory study argues.

desk verdict A carefully reported exploratory study whose Pluriversal CAI for Health framework is a genuinely useful extension of the TCE; the 'culture loses meaning' claim is oversold as a finding but the entanglement it asserts is visible in the participant data. read the letter →

arxiv 2507.01719 v1 pith:EED7QNXJ submitted 2025-07-02 cs.HC cs.AI

classification cs.HCcs.AI
keywords conversationalAIhealthchatbotsculturalappropriatenessLatinAmericapluriversaldesignparticipatoryworkshopsthematicanalysismajorityworld
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that making conversational AI for health truly appropriate for Latin America cannot be accomplished by adding more cultural data or fine-tuning models alone. Drawing on eight participatory workshops with 108 Spanish-speaking citizens and health and AI professionals in Peru, Argentina, and UK-based Latin American migrant communities, the authors identify sites where health chatbots are likely to misalign with local realities. Their central finding is that at ground level, culture is inseparable from economics, politics, geography, family structure, and local logistics, so a chatbot that ignores these factors will feel culturally wrong even if its language and values are tuned. To guide designers, they propose a Pluriversal Conversational AI for Health framework that maps individual, relational, and ecosystem influences and suggests that more relationality and tolerance, rather than just more data, may be needed.

What carries the argument

The load-bearing device is the Pluriversal CAI for Health framework, a concentric-circle map of what a health chatbot must hold in view: the Individual (demographics, emotional state, medical history, knowledge and information needs), Relationships (family interdependence, kinship structure, responsibilities and reciprocity, community dynamics, friends and workplace), and Ecosystems (healthcare system, natural environment, mobility and infrastructure, economic and political systems, history and traditions), plus cross-cutting cultural elements adapted from an existing taxonomy and the added theme of power dynamics and discrimination. The framework is generated from qualitative workshop data—projective storytelling worksheets in which participants described third-person health scenarios, followed by role-play conversations with a persona-prompted GPT-3.5 Turbo chatbot—and analyzed through reflexive inductive thematic analysis. It does the work of converting scattered stories about health conversations into a checklist of where cultural misalignment can arise and where future CAI evaluation and benchmarking should look.

What would settle it

A comparative user study in which Latin American users rate responses from a framework-informed health chatbot versus a conventional chatbot on identical health queries would settle it: if the framework-informed responses are not judged more culturally appropriate and trustworthy, the claim that family dynamics and material constraints are load-bearing for cultural fit fails.

Watch

Extended reading notes

Core claim

The study's central claim is that culturally appropriate health CAI in the majority world requires a holistic framework because academic boundaries around 'culture' do not hold in lived health experience. Based on thematic analysis of workshop stories and discussions, the authors find that health conversations in Latin America are shaped by entities beyond the individual user: the family as a decision-making unit, community dynamics, traditional and alternative health practices, misinformation, financial constraints, crime, transport, regional disease, and food availability. They conclude that a chatbot that recommends an electronic air filter to someone without window glass and with regular power outages will be experienced as culturally incompetent, even though the failure is economic and material. The proposed Pluriversal CAI for Health framework therefore includes three realms—the Individual, Relationships, and Ecosystems—with cross-cutting themes of artefacts and technologies, concepts, norms and morals, values and beliefs, and power dynamics, and it incorporates plural notions of conviviality and relationality rather than treating cultural fit as a fixed set of trainable elements.

Load-bearing premise

The framework's generalizability rests on the assumption that eight workshops with 108 Spanish-speaking participants, recruited mostly through existing professional and community networks in Peru, Argentina, and UK migrant communities, surface themes stable enough to represent Latin American and majority-world health experiences.

Editorial extensions

If this is right

  • Health chatbots deployed in Latin America should treat the patient and their family as a single decision-making unit, accounting for how health decisions affect and depend on relatives.
  • Culturally appropriate CAI must incorporate local logistics and material conditions such as medicine availability, transport, costs, and food access, or its advice will read as foreign even if the language is perfect.
  • A credible chatbot could act as an on-ramp to formal care and an information bridge between appointments, but only if it is explicitly designed not to replace human care, which participants feared.
  • Because participants reported that patients may be more honest with a computer, CAI has a potential role in reducing stigma-related nondisclosure and guiding users toward appropriate services.
  • Multimodal, multimedia communication such as images, videos, stories, and WhatsApp-style delivery is likely necessary to serve mixed literacy levels and learning preferences.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the framework implies a practical data architecture in which chatbots maintain separate context slots for individual, relational, and ecosystem information, and can flag when relational or material context is missing rather than assuming a lone user.
  • Editorial inference: the relational emphasis suggests a testable design: in collectivist settings, interventions that also address the older patient's adult children or in-laws may outperform one-to-one personalized coaching, and this could be measured in a randomized trial.
  • Editorial inference: because all participants were Spanish-speaking and mostly reached through existing networks, the strongest test of pluriversality is replication with Indigenous-language communities and non-mestizo populations; until then, the framework's Latin America-wide reach is provisional.
  • Editorial inference: the paper's 'humility and tolerance' direction points to an alternative benchmark—measuring whether a chatbot that asks preference-eliciting questions and hedges assumptions is rated more appropriate in unfamiliar cultural contexts than one that tries to mirror a detected culture.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 7 minor

Summary. This paper reports an exploratory qualitative study with 108 Spanish-speaking participants across eight participatory workshops in Peru, Argentina, and Latin American migrant communities in the UK. Using projective storytelling, consequence scanning, and a customized GPT-3.5-based chatbot, the authors identify themes around health conversations, barriers to healthcare access, and opportunities/risks of conversational AI. The central claim is that academic boundaries around 'culture' lose meaning at the ground level because cultural experience is entangled with economic, political, geographic, and logistical systems; the paper proposes a 'Pluriversal Conversational AI for Health' framework that encompasses three realms (ecosystems, relationships, individual) plus cross-cutting themes, extending Liu et al.'s Taxonomy of Cultural Elements (TCE).

Significance. If the findings and framework hold, this is a timely contribution to culturally-appropriate conversational AI for health, representing an understudied region and a bottom-up, participatory approach. Strengths include the use of COREQ reporting, explicit researcher positionality, multiple workshop sites and stakeholders, a detailed code tree in the appendix, and a thoughtful discussion linking the data to pluriversal design and conviviality. The paper also raises valuable open technical questions. However, the central empirical claim that culture is inextricably entangled with material systems is presented as a direct finding yet lacks systematic within-narrative evidence, and the framework is positioned as correcting existing taxonomies without a demonstrated coding comparison.

major comments (2)
  1. [§5.1, §6.2.1] The paper's central claim, stated in the Abstract ('academic boundaries on notions of culture lose meaning at the ground level') and in §6.2.1 ('our data revealed that notions of culture are so entangled...'), is presented as an empirical finding. However, Section 5.1 reports the thematic analysis in separate categories: 'Socio-cultural ecosystem' and 'Constraints to healthcare access' are distinct categories, and no within-narrative or unit-of-analysis analysis is reported that would demonstrate co-occurrence of cultural and material/geographic/logistical themes within the same participant stories. The illustrative quotes in §5.1.6 (e.g., the Cardiff 'vacunadas' quote and the Huancayo discrimination quote) do show such co-occurrence, but the reader is not told how representative these are across the 682 excerpts. The Discussion illustrates the entanglement claim with a hypothetical Cuban asthma example rather than with participant data. To make this claim load-bearing, the authors should either re-analyze the transcripts at the story level and report the frequency and patterns of theme co-occurrence, or explicitly relabel the claim as a conceptual interpretation synthesized from the data rather than a direct empirical result.
  2. [§6.2.2] The framework is described as incorporating 'all elements of the TCE' and adding categories; it is therefore a superset. The paper's added value, however, rests on the claim that a strict TCE-based analysis would be insufficient ('a strict and exclusive definition loses meaning on the ground'). This insufficiency is asserted in §6.2.1 and §6.2.2 but never demonstrated empirically. For example, the authors do not report a coding comparison in which TCE elements were applied to the data and shown to leave material constraints or relational dynamics as residual categories. Without such evidence, the framework is a plausible conceptual proposal, but the claim that it addresses a failure of existing taxonomies is not supported. The authors should either provide a systematic mapping (e.g., a table showing which TCE elements were insufficient and what data fell outside them) or revise the framing to present the framework as a domain-specific extension that is complementary to, rather than corrective of, the TCE.
minor comments (7)
  1. [§4.3.1] The analysis is described as performed by one coder (Peters) with feedback from Da Re and Calvo; given that the paper follows reflexive thematic analysis, the single-coder approach is defensible, but the rationale should be stated explicitly, and the description of the coding rounds (62 initial codes, 682 excerpts, 107 final codes) could be expanded to clarify how the feedback from other researchers changed the code tree.
  2. [Abstract, §7] The abstract states 'Our findings show...' without qualification, although Section 7 appropriately acknowledges that results cannot represent the rest of Latin America or even the two countries entirely; consider adding a qualifier such as 'in this sample' to the abstract and framing the implications as hypotheses to be tested in other settings.
  3. [Table 2] Table 2 is difficult to parse because the Realm column values (Ecosystems, Relationships, Individual, Cross-cutting) are not visually distinguished from the theme names in the second column, and the caption uses 'Spheres' while the text uses 'Realms'; please reformat the table and align terminology, and ensure Figure 2 is legible in the final version.
  4. [§2.1] The citation of Frankfurt (1971) as an example of documented cultural biases in LLMs is puzzling, since the cited work is a philosophy paper on free will; please verify that this citation is correct and place it in the appropriate context.
  5. [§5.1.7] The quote in which a participant obtains a Cipro dosage after claiming to be a doctor demonstrates a breakdown of medical guardrails; the authors mention the female gendering of the assistant but do not comment on the safety vulnerability revealed here, which is directly relevant to the 'boundaries and limits' theme.
  6. [§6.2.2] In §6.2.2, 'Table 3' is referenced but the table shown is numbered Table 2; please renumber consistently.
  7. [Whole manuscript] There are several typographical and spelling errors, including 're highlighted' in §6.2.2, 'Karusula' in §6.3, 'Hovey' in §6.2.1, and 'Caros Paz' in §5.1.5; a careful proofread is needed.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the Pluriversal framework is a transparent interpretive synthesis of qualitative data and prior taxonomies, not a fitted prediction or a self-citation-dependent derivation.

full rationale

The paper's derivation chain is qualitative: workshop data were coded through reflexive inductive thematic analysis (Sections 4.3 and 5.1), and the Pluriversal CAI framework was then assembled in the Discussion to organize those themes alongside Liu et al.'s Taxonomy of Cultural Elements (Section 6.2.2). The framework is explicitly presented as a first-iteration organizing device ('we present a first iteration that we ultimately found helpful in our own work for: 1. Organising the data from our study...'), not as a prediction or a fitted parameter. The central claim that cultural and material systems are entangled is an interpretive synthesis of the coded data and prior literature; whatever its evidentiary strength, that is a question of support and generalizability, not circularity. The paper does not define culture in terms of the framework, nor does it invoke a uniqueness theorem or a self-authored result to force its conclusions. Self-citations (e.g., Maina et al. 2024, Fort et al. 2024) appear only as background context and are not load-bearing for the framework. The limitations section (Section 7) also explicitly acknowledges that results cannot represent the whole continent, further indicating that the authors are not asserting the framework as a self-validating necessity.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper introduces no fitted numerical parameters or invented physical entities. Its assumptions are qualitative: that participatory storytelling reflects lived experience, that one researcher's coding with feedback is trustworthy, and that the TCE is a sound conceptual base.

assumptions (3)
  • domain assumption Participants' projective stories and group discussions provide valid windows onto lived experiences and cultural factors in health conversations.
    The entire thematic analysis treats participant narratives as evidence about real-world health communication, but projective methods are indirect and may not map cleanly to actual behavior.
  • domain assumption Reflexive thematic analysis performed primarily by one researcher with feedback from two co-authors yields trustworthy themes.
    No inter-rater reliability or independent coding was reported; the analysis is inherently interpretative (Section 4.3.1).
  • domain assumption Existing taxonomy of cultural elements (TCE, Liu et al. 2024) is a valid foundation for organizing culture in CAI.
    The framework incorporates all TCE elements and adapts several; if TCE is flawed, the framework inherits that limitation (Section 6.2.2).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Towards culturally-appropriate conversational AI for health in the majority world: An exploratory study with citizens and professionals in Latin America." pith.science (2026). https://pith.science/paper/EED7QNXJ

@misc{pith2026250701719,
  author       = {Pith},
  title        = {Pith review of: Towards culturally-appropriate conversational AI for health in the majority world: An exploratory study with citizens and professionals in Latin America},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/EED7QNXJ}},
  note         = {Machine review of arXiv:2507.01719}
}
read the original abstract

There is justifiable interest in leveraging conversational AI (CAI) for health across the majority world, but to be effective, CAI must respond appropriately within culturally and linguistically diverse contexts. Therefore, we need ways to address the fact that current LLMs exclude many lived experiences globally. Various advances are underway which focus on top-down approaches and increasing training data. In this paper, we aim to complement these with a bottom-up locally-grounded approach based on qualitative data collected during participatory workshops in Latin America. Our goal is to construct a rich and human-centred understanding of: a) potential areas of cultural misalignment in digital health; b) regional perspectives on chatbots for health and c)strategies for creating culturally-appropriate CAI; with a focus on the understudied Latin American context. Our findings show that academic boundaries on notions of culture lose meaning at the ground level and technologies will need to engage with a broader framework; one that encapsulates the way economics, politics, geography and local logistics are entangled in cultural experience. To this end, we introduce a framework for 'Pluriversal Conversational AI for Health' which allows for the possibility that more relationality and tolerance, rather than just more data, may be called for.

Figures

Figures reproduced from arXiv: 2507.01719 by the authors.

Figure 1
Figure 1. Two of the visuals provided to workshop participants to prompt generation of health stories. (Images created by Fernanda Espinoza and adapted from Till et al., 2022) [PITH_FULL_IMAGE:figures/full_fig_p009_1.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

12 extracted references · 11 canonical work pages

  1. [1]

    Health conversation facilitators, barriers and triggers – triggers and elements influencing the quality of a health conversation, such as honesty, pacing and tone

  2. [2]

    Individual characteristics that impact health conversations – Contextual and demographic elements of an individual that impact health conversation such as medical history, emotional state and access needs

  3. [3]

    Communication ecosystem - interlocuters and components of the health conversation ecosystem such as doctors, insurance providers and the media

  4. [4]

    Socio-cultural ecosystem – Relationships and cultural systems that influence a health conversation such as family, spirituality and regional politics

  5. [5]

    Constraints to healthcare access – Obstacles to access such as mobility limitations, reluctance to access formal care, crime and corruption

  6. [6]

    the problem of fake news and about the news in the newspapers, in which they can spread poorly interpreted research…people start to lose trust in medicine itself

    CAI for health - ideas and preferences with respect to CAI use for healthcare such as perceived risks and opportunities. 5.1.1 Themes overview A complete code-tree is included as an appendix. Table 1 shows example themes and sub-themes with associated excerpts. Herein we focus on the themes that have direct implications for Conversational AI design. For c...

  7. [8]

    and beliefs

    values and beliefs (we added the words “and beliefs” to “values” to more clearly highlight the beliefs component of this element). We also added one additional cross-cutting theme not already in the TCE: “Power dynamics and discrimination” as this theme represents a major concern within work on culturally-appropriate CAI and we believe that it is importan...

  8. [9]

    Relationship

    “Relationship” is replaced by the top-level category “Relationships” in our framework which itself breaks into a series of important themes

Show all 12 references
  1. [10]

    Demographics

    “Demographics” (e.g. income, gender, nationality), didn’t capture our data that surfaced other individual characteristics not typically considered demographics (e.g. medical history, personality, lifestyle factors). We therefore replaced this with the theme “Individual charact...

  2. [11]

    Knowledge

    “Knowledge” was expanded to “Knowledge and information needs” which considers information preferences as well, for example, for receiving more versus less detail about a medical issue

  3. [12]

    linguistic containers

    “Context” in the TCE is defined as: “the ‘containers’ of communications which can be linguistic such as surrounding sentences or extra-linguistic including social settings, non- verbal cues (e.g., gesture), or historical contexts (e.g., colonization).” Based on the importance ...

  4. [2024]

    relationship

    which incorporates previous taxonomies from computational linguistics but also draws on cultural research in sociology, anthropology and other disciplines. As such, we opted to work with the TCE to see if we might situate our findings within it. Of course, our code tree and th...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.