Pith. sign in

REVIEW 5 cited by

Value FULCRA: Mapping Large Language Models to the Multidimensional Spectrum of Basic Human Values

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.10766 v1 pith:KCUJYTXV submitted 2023-11-15 cs.CL cs.AI

classification cs.CLcs.AI
keywords basicvaluevaluesalignmentfulcrallmsbehaviorsexisting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The rapid advancement of Large Language Models (LLMs) has attracted much attention to value alignment for their responsible development. However, how to define values in this context remains a largely unexplored question. Existing work mainly follows the Helpful, Honest, Harmless principle and specifies values as risk criteria formulated in the AI community, e.g., fairness and privacy protection, suffering from poor clarity, adaptability and transparency. Inspired by basic values in humanity and social science across cultures, this work proposes a novel basic value alignment paradigm and introduces a value space spanned by basic value dimensions. All LLMs' behaviors can be mapped into the space by identifying the underlying values, possessing the potential to address the three challenges. To foster future research, we apply the representative Schwartz's Theory of Basic Values as an initialized example and construct FULCRA, a dataset consisting of 5k (LLM output, value vector) pairs. Our extensive analysis of FULCRA reveals the underlying relation between basic values and LLMs' behaviors, demonstrating that our approach not only covers existing mainstream risks but also anticipates possibly unidentified ones. Additionally, we present an initial implementation of the basic value evaluation and alignment, paving the way for future research in this line.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluating the Prompt Steerability of Large Language Models

    cs.CL 2024-11 conditional novelty 6.0 of 10

    A formal benchmark with steerability indices shows that six open-weight LLMs are only partially steerable by prompting, with strong baseline skew and directional asymmetry.

  2. Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values

    cs.AI 2025-01 conditional novelty 5.0 of 10

    Value Compass Benchmarks is a live, self-evolving platform that scores 33 LLMs across 27 value dimensions from four value systems, aiming to reveal true behavioral alignment with human values.

  3. Are the Values of LLMs Structurally Aligned with Humans? A Causal Perspective

    cs.CL 2024-12 reject novelty 5.0 of 10

    A dependency graph of 17 values learned from two LLMs predicts side effects of role and SAE steering, but the causal and human-alignment claims are unsupported.

  4. A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy

    cs.AI 2025-01 conditional novelty 4.0 of 10

    A survey that organizes responsible-LLM research into five risk dimensions and four intervention phases, reviewing privacy, hallucination, value, toxicity, and jailbreak mitigation.

  5. Large language models for artificial general intelligence (AGI): A survey of foundational principles and approaches

    cs.AI 2025-01 conditional novelty 3.0 of 10

    This survey argues that embodiment, symbol grounding, causality, and memory are the foundational principles needed to make large language models achieve artificial general intelligence.

Pith tools