Pith. sign in

REVIEW 15 cited by

The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.16019 v2 pith:B5IILBLL submitted 2024-04-24 cs.CL

classification cs.CL
keywords feedbackalignmentprismwhatcountriesdatasethumanindividualised
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Human feedback is central to the alignment of Large Language Models (LLMs). However, open questions remain about methods (how), domains (where), people (who) and objectives (to what end) of feedback processes. To navigate these questions, we introduce PRISM, a dataset that maps the sociodemographics and stated preferences of 1,500 diverse participants from 75 countries, to their contextual preferences and fine-grained feedback in 8,011 live conversations with 21 LLMs. With PRISM, we contribute (i) wider geographic and demographic participation in feedback; (ii) census-representative samples for two countries (UK, US); and (iii) individualised ratings that link to detailed participant profiles, permitting personalisation and attribution of sample artefacts. We target subjective and multicultural perspectives on value-laden and controversial issues, where we expect interpersonal and cross-cultural disagreement. We use PRISM in three case studies to demonstrate the need for careful consideration of which humans provide what alignment data.

Discussion (0). Sign in to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models

    cs.LG 2025-07 conditional novelty 7.0 of 10

    A demographically diverse annotation dataset shows that safety perceptions for text-to-image outputs vary by rater identity and that conventional safety classifiers under-detect bias harms flagged by minority-group raters.

  2. Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Small hyperbolic models (146M–3B) report 100% creative-seed preference, 90.7% compliance-gap detection, and a selective-gating skeleton–wallpaper memory pilot as a companion-AI stack.

  3. CuMA: Aligning LLMs with Sparse Cultural Values via Demographic-Aware Mixture of Adapters

    cs.CL 2026-01 conditional novelty 6.0 of 10

    A demographic-conditioned mixture of LoRA experts improves LLM cultural alignment and reduces the tendency of dense models to produce generic, averaged responses.

  4. The Alignment Veto: How Safety Training Suppresses Cultural Knowledge in LLMs

    cs.CL 2025-10 reject novelty 6.0 of 10

    The full text builds the MENA Values benchmark (864 questions, 7 models) and reports that LLM cultural answers shift with language, decline with reasoning prompts, and hide strong internal preferences behind refusals—...

  5. The Personalization Trap: How User Memory Alters Emotional Reasoning in LLMs

    cs.AI 2025-10 conditional novelty 6.0 of 10

    Adding user memory to LLMs degrades their emotional-intelligence test scores and systematically disadvantages marginalized user profiles.

  6. CultureSynth: A Hierarchical Taxonomy-Guided and Retrieval-Augmented Framework for Cultural Question-Answer Synthesis

    cs.CL 2025-09 conditional novelty 6.0 of 10

    A taxonomy-guided retrieval-augmented framework generates CultureSynth-7, a multilingual cultural QA benchmark, and its evaluation of 14 LLMs suggests cultural competence emerges around 3B parameters.

  7. Toward Socially Aware Vision-Language Models: Evaluating Cultural Competence Through Multimodal Story Generation

    cs.CL 2025-08 conditional novelty 6.0 of 10

    An evaluation of five VLMs on culturally-prompted multimodal story generation finds measurable cultural adaptation alongside metric bias and inverse alignment in some models.

  8. From Seed to Harvest: Augmenting Human Creativity with AI for Red-teaming Text-to-Image Models

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Seed2Harvest expands 1,000 human adversarial prompts into 27,650 LLM-generated variants that keep roughly comparable unsafe-image trigger rates and add hundreds of new geographic contexts.

  9. CALMA: A Process for Deriving Context-aligned Axes for Language Model Alignment

    cs.CY 2025-07 conditional novelty 6.0 of 10

    CALMA is a grounded-theory, participatory method for deriving community-specific language model alignment axes from open-ended user interactions and group discussion, piloted with two small groups.

  10. ModelCitizens: Representing Community Voices in Online Safety

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A community-annotated toxicity dataset with conversational context shows that models trained on ingroup labels outperform state-of-the-art moderation APIs.

  11. Amulet: Putting Complex Multi-Turn Conversations on the Stand with LLM Juries

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Dialog-act and maxim-aware prompting improves LLM judge accuracy on multi-turn preference data by up to 8 points, with further gains from jury-style voting.

  12. Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A new benchmark elicits AI models' value priorities from choices in 3,000 AI-risk dilemmas and reports correlations between those priorities and risky behaviors, including on the external HarmBench.

  13. A Taxonomy of Linguistic Expressions That Contribute To Anthropomorphism of Language Technologies

    cs.HC 2025-02 conditional novelty 6.0 of 10

    A taxonomy of 19 types of linguistic expressions and 5 guiding lenses for identifying when language technology outputs may contribute to anthropomorphism.

  14. Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge?

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Tool-augmented LLM annotators improve agreement with ground-truth preferences on long-form factual and coding tasks, with mixed results on math, compared to standard LLM-as-a-judge baselines.

  15. Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges

    cs.AI 2025-07 reject novelty 1.0 of 10

    A broad survey of LLM alignment that catalogs objectives, benchmarks, SFT/RLHF/DPO methods, and safety challenges, without contributing new experimental or theoretical results.

Pith tools