{"id":"cc3ec357-98f5-4a98-9edc-8fee60c0d49e","arxiv_id":"2505.01015","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Value Portrait links LLM value scores to human PVQ-validated items from real conversations, finding that LLMs emphasize Benevolence, Security, and Self-Direction while downplaying Tradition, Power, and Achievement.","lead":"The authors built a new benchmark, Value Portrait, that measures what values large language models hold by asking people to rate how much AI-generated responses match their own thinking. The benchmark is designed to reflect real user-AI conversations and to avoid the biases of earlier annotation methods.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The benchmark's item validity is established on human raters but assumed to transfer to LLM self-reports; no external criterion shows LLM scores measure the intended value constructs.","rationale":"I read the paper as a serious attempt to ground LLM value benchmarks in psychometric practice; the construction pipeline, the correlation-based item selection, and the transparency about limitations are genuine strengths. The single most load-bearing step is the unstated equivalence between human self-report and LLM self-report. The reader's weakest_assumption identifies exactly this step, and I agree it is the point where the central claim is least secure. My attack adds two pieces of internal evidence that make the concern concrete rather than hypothetical: Appendix D shows that even GPT-4o's own value-targeted text generation does not reliably express the intended values when judged by human correlations, and Appendix I shows that prompting Benevolence increases Security more than Benevolence, a dissociation between the model's stated value and its behavior under the benchmark's own scoring. These observations do not refute the benchmark, but they show that the assumption 'a similarity rating is a direct measure of the model's value orientation' can fail in controlled conditions. The proposed forced-choice test would settle whether the benchmark's normalized scores predict actual value-relevant choices across models. Because the paper's contribution as a benchmark resource remains valuable and the required validation is well-defined, the conditional verdict stands; no change is needed.","tokens_in":40690,"tokens_out":3932,"duration_ms":42761,"concrete_test":"Run a convergent-validity test: for each of the 10 value dimensions, select a sample of Value Portrait queries and pair responses with strong positive and strong negative correlations for that value (matched for length and style). Present these pairs to the same 44 LLMs asking \"Which response is closer to your own thoughts?\" and compute the fraction of positive-value choices per model. Correlate this forced-choice preference with the model's normalized Value Portrait score for that dimension. If the per-dimension correlations are not consistently positive (e.g., mean r < 0.3) across models, the benchmark scores do not predict value-relevant preferences and the human-to-LLM transfer assumption fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central inference is in Section 3.4.1: items are retained because human similarity ratings correlate |r| >= 0.3 with PVQ-21 scores, and those same items are then used to score LLMs from the prompt \"How similar is this response to your own thoughts?\" The validity established in humans is not re-established for models. Cronbach's alpha in Appendix F only shows that LLM responses are internally consistent; it does not show the scores track the Schwartz value constructs. The paper's own data supply evidence that LLM self-reports can dissociate from value-relevant behavior: Appendix D reports that GPT-4o responses explicitly generated to express one value aligned with the intended PVQ dimension only 11.25% of the time, and Appendix I shows that steering GPT-4o toward Benevolence raised Security (+1.08) more than Benevolence (+0.11). If LLM similarity ratings instead reflect prompt compliance, safety-aligned response style, or training-data statistics, then the reported value profiles in Section 4.1 and the demographic bias comparisons in Section 4.2 do not measure model values. The transfer assumption is the load-bearing step and it is currently unverified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Value Portrait, a benchmark for evaluating the value orientations of large language models. Items are query-response pairs sampled from human-LLM interactions (ShareGPT, LMSYS) and human-human advisory contexts (Reddit Scruples, Dear Abby), with responses generated by GPT-4o. Each item was rated by human Prolific participants for similarity to their own thoughts, and Spearman correlations were computed between these ratings and participants' PVQ-21 (and BFI-10) scores; items with |r| >= 0.3 and p < 0.05 were retained. The benchmark is then administered to 44 LLMs, which rate how similar each item is to their own thoughts, and value orientations are computed from the averaged, normalized ratings. The authors report that LLMs prioritize Benevolence, Security, and Self-Direction and de-emphasize Tradition, Power, and Achievement, and they analyze demographic persona biases and value-steering effects for GPT-4o. The paper claims the benchmark is psychometrically and ecologically valid for LLM value assessment.","tokens_in":40890,"tokens_out":7586,"duration_ms":73671,"significance":"If the construct-validity assumptions are established, Value Portrait would be a valuable resource: it uses ecologically sourced queries, provides a large 44-model evaluation, and is accompanied by released code and data. The manuscript is transparent in reporting the failure of value-targeted generation (Appendix D) and unexpected steering interactions (Appendix I), which are useful negative results for the field. However, the central load-bearing step—transferring human item validity to LLM self-reports—is unverified, and the paper's own appendices supply evidence of dissociation between LLM stated similarity ratings and value-relevant behavior. The benchmark's usefulness as a measure of model values therefore remains conditional on additional validation or on a more cautious reframing of the claims.","major_comments":[{"comment":"The central construct-validity argument transfers item validity from human raters to LLMs without direct evidence. Correlations computed on human participants (Section 3.4.1) justify retaining items for humans; Cronbach's alpha in Appendix F is computed on LLM responses and only shows internal consistency, not that the scale measures the intended Schwartz constructs in models. Appendix D reports that GPT-4o responses explicitly generated to express a value aligned with the intended PVQ dimension only 11.25% of the time, and Appendix I shows that steering GPT-4o toward Benevolence raised Security by +1.08 versus +0.11 for Benevolence; both results suggest LLM similarity ratings can dissociate from the intended value constructs. Appendix L also explicitly frames the key assumption as an unsupported 'assuming outputs reflect the model's internal preferences.' Because the value profiles in Section 4.1 and the bias comparisons depend on this assumption, the benchmark's central claim is not yet established.","section":"§3.4.1–3.4.2, Appendices D, F, I, L"},{"comment":"The item-selection procedure is vulnerable to multiple-testing and underpowered detection. The paper computes thousands of correlations (520 items × 10 value dimensions, plus 5 BFI traits) and retains all correlations with |r| ≥ 0.3 and p < 0.05 without any multiple-comparison correction; under the null, roughly 5% of several thousand tests will pass the p-threshold, so many retained items are likely false positives. Moreover, the claim that the average 46 participants per item yields power 0.8 to detect r = 0.3 at p < 0.05 is not consistent with standard sample-size calculations: with N = 46 the approximate power for a two-sided Pearson test of r = 0.3 is about 0.5, and likely lower for Spearman. The authors should provide a power analysis and either correct for multiplicity (e.g., FDR) or justify the threshold selection on independent grounds.","section":"§3.2.3 and §3.4.1"},{"comment":"The demographic-bias comparisons are purely descriptive and lack inferential support. The text states that GPT-4o 'significantly exaggerates' gender differences and 'amplifies political value differences,' but no confidence intervals, standard errors, or significance tests are reported for the GPT-4o persona-condition scores or for the comparison with ESS human data. The ESS samples are weighted and clustered; the analysis in Appendix H reports unweighted relative differences with no within-country or within-group variances. The bias claims in the abstract and introduction therefore currently exceed what the data support. Please report uncertainty quantification (e.g., bootstrapped CIs across items and prompts) and, where possible, test the human-model differences.","section":"§4.2, Tables 9–11"}],"minor_comments":[{"comment":"The evaluation of prior datasets uses different response formats from the original annotation tasks (binary Yes/No for ValueNet and Likert for FULCRA), and samples only 20 items per dataset; the reported 5% and 10% alignment rates should be interpreted as a pilot and the protocol mismatch acknowledged.","section":"§3.1, Appendix B"},{"comment":"The filtering criterion in Figure 3 reads 'At least One Corr > 0.3' but the text specifies |r| ≥ 0.3 with p < 0.05; please make the figure consistent.","section":"§3.4.1, Figure 3"},{"comment":"The released dataset should be checked against the LMSYS-Chat-1M license quoted in Appendix K, which prohibits distribution, copying, and transfer to third parties; the paper does not state whether the dataset has been approved for redistribution.","section":"Appendix K"},{"comment":"The persona prompts used in Appendix H.1 are extremely short (e.g., 'Your gender is male.') and no pilot or manipulation check is reported to confirm that these prompts actually instantiate the intended persona in GPT-4o; consider reporting the full set of prompts and including a check (e.g., self-reported persona adherence).","section":"§4.2, Appendix H.1"},{"comment":"There are several minor typographical issues (e.g., 'Y ohan Jo' in the author list, inconsistent superscript markers in Table 5); a careful proofread is recommended.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The main concern is that the authors' own appendices (D, I, L) supply evidence that undercuts the transfer assumption; if the authors can either validate the LLM measure against an external behavioral criterion (e.g., a separate value-relevant task) or substantially soften the claims to 'self-reported similarity judgments' rather than 'values,' the paper could be publishable. Also worth checking: the LMSYS license issue if the dataset is publicly released. The paper is within scope for cs.CL, but the psychometric framing invites measurement-validity scrutiny."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Value Portrait is a worthwhile new resource: 104 realistic queries, 520 query-response pairs, 7,800 human similarity ratings tied to PVQ-21 and BFI-10, and a 44-model evaluation, with code and data released. The correlation-based item selection is a genuine step beyond perceived-value annotation, and the paper's own control analyses (cross-loading distances consistent with Schwartz's circular structure, and the honest reporting of failed value-targeted generation) are informative. The demographic bias comparisons are creative.\n\nThe soft spot is exactly where the stress-test note points. Item validity is established in human raters, then transferred to LLMs by asking the model the same self-report question. That is an assumption, not a demonstrated fact. Cronbach's alpha on LLM responses tells you only that LLM answers are internally consistent, not that they track Schwartz constructs. The paper's Appendix D (only 11.25% of value-targeted GPT-4o responses aligned) and Appendix I (Benevolence steering raised Security more than Benevolence) are not disqualifying, but they are exactly the kind of dissociation that makes the transfer assumption non-trivial. The Limitations section and Appendix L acknowledge general self-report issues but never directly address the human-to-LLM transfer. That should be addressed head-on.\n\nThe multiple comparisons issue is real but secondary: 7,800 correlations at p<0.05 with n≈46 will produce many expected false positives around r=0.3. The cross-loading analysis suggests signal, but a corrected threshold or split-half replication would make the item selection much more convincing.\n\nNone of this kills the paper. The core methodological point—correlation with actual value scores is better than perceived-value labeling—holds. The benchmark is reproducible and useful regardless of whether the specific value profiles are final. The headline value priorities (Benevolence, Security, Self-Direction over Tradition, Power, Achievement) replicate prior work, so they are plausible; the demographic bias numbers should be treated as suggestive until the measurement issue is resolved.\n\nThis paper deserves a serious referee. I would send it to review with a request to either soften the validity claims or add an external LLM-behavior validation, and to correct for multiple comparisons. The benchmark is worth having in the literature.","headline":"A genuinely useful, reproducible benchmark that improves on perceived-value annotation, but its central assumption that human-validated items measure LLM values is unverified, so the value profiles should be read as suggestive.","tokens_in":712,"tokens_out":868,"would_cite":true,"duration_ms":49392,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that LLM value orientations can be measured with psychometrically validated items, and that 44 evaluated models uniformly prioritize Benevolence, Security, and Self-Direction while de-emphasizing Tradition, Power, and…","keywords":["value orientation","large language models","Schwartz theory of basic values","psychometric benchmark","PVQ-21","demographic bias","value alignment","persona prompting"],"falsifier":"Run the same 520 items on an unaligned base model (no instruction tuning or RLHF) or on a model instructed to role-play a maximally unhelpful assistant: if the profile still shows high Benevolence, Security, and Self-Direction, the instrument is reading training-text priors rather than model values. A complementary check is behavioral: have each model produce free-form answers to the same 104 queries, then test whether models that self-report higher Benevolence on Value Portrait also generate measurably more benevolent text; if self-reports and behavior diverge, the self-report reading fails.","tokens_in":40465,"feed_emoji":"🧭","tokens_out":9701,"duration_ms":91913,"temperature":0.7,"pith_summary":"The paper argues that existing value benchmarks for language models mislabel responses because annotators guess which values a text expresses, importing their own biases; it proposes instead to validate each test item psychometrically, by whether people who actually hold a value identify with the response. The result is Value Portrait, a benchmark of 520 query-response pairs drawn from real human-LLM interactions (ShareGPT, LMSYS) and human advisory exchanges (Reddit's AITA, Dear Abby), each tagged with the Schwartz values (and Big Five traits) whose scores correlate with human raters' \"how similar is this to your own thoughts?\" judgments ($|r_s| \\geq 0.3$, $p<0.05$). Used to evaluate 44 models, the benchmark yields a consistent cross-model profile, with Benevolence, Security, and Self-Direction prioritized and Tradition, Power, and Achievement de-emphasized, and it shows that persona-prompted GPT-4o exaggerates gender, age, and political value differences relative to European Social Survey data. If the framework is right, it gives developers and auditors a contamination-resistant way to measure what values models actually express, turning value claims about LLMs from opinion into something testable.","feed_headline":"44 LLMs put benevolence first, tradition and power last","feed_subtitle":"Psychometrically validated quiz items reveal one shared model profile — and exaggerated group gaps.","key_machinery":"The load-bearing device is the value-correlation item. For every query-response pair, the paper computes Spearman correlations between crowdworkers' six-point similarity ratings and their scores on the official Portrait Values Questionnaire (PVQ-21, ten Schwartz value dimensions) and the BFI-10 personality inventory; items whose correlation with a dimension passes $|r_s| \\geq 0.3$ with $p<0.05$ become items for that dimension. LLM evaluation then reuses the same human self-report prompt, \"How similar is this response to your own thoughts?\", averaging responses across six prompt variants, and converts raw scores into relative value priorities by subtracting each model's mean item response, following Schwartz's ipsatization convention for human value measurement. The Schwartz ten-value circle supplies the taxonomy, and the same correlation pipeline extends the benchmark to the Big Five personality traits.","core_discovery":"Value Portrait claims to be a psychometrically and ecologically valid instrument for measuring LLM value orientations. Each of its 520 items is a real-world query paired with a response generated by GPT-4o, and each item's link to a value is established empirically: roughly 46 crowdworkers per item rated how similar the response was to their own thoughts, and those ratings were Spearman-correlated with the same raters' PVQ-21 value (and BFI-10 personality) scores. Items with $|r_s| \\geq 0.3$ and $p<0.05$ for a dimension became that dimension's items, leaving 549 value correlations and 287 trait correlations. Administered to 44 LLMs with the same six-point \"how similar to your own thoughts?\" question, averaged over six prompt variants and mean-centered following Schwartz's scoring method, the benchmark shows high internal consistency (Cronbach's $\\alpha$ between 0.76 and 0.96) and reveals a stable profile across models: Benevolence, Security, and Self-Direction rank highest, while Tradition, Power, and Achievement rank lowest. The authors further report that reasoning models amplify Benevolence, larger models differentiate values more sharply, and persona-prompted GPT-4o overstates demographic differences, for example a male-female Conformity gap of 0.51 versus 0.02 in human data and a Left-Right Hedonism gap of 0.74 versus 0.03, while steering experiments show most values respond to prompts but Benevolence does not, since steering toward it raised Security by 1.08.","pith_inferences":["Because Schwartz values form a correlated circular structure, the reported single-dimension scores will partly reflect neighboring values; the paper's own cross-loading analysis (same-direction value pairs average circular distance 1.59, opposite-direction pairs 3.54) means the ten scores should be read as one profile, not ten independent measures.","A testable extension the paper leaves implicit: compare each model's Value Portrait self-ratings with its free-form responses to the same 104 queries; if self-reported high-Benevolence models do not generate more benevolent texts, the instrument captures stated preference rather than expressed behavior.","The demographic-bias result is pinned to the European Social Survey as the human baseline; rerunning the persona audit against non-Western value datasets would show whether GPT-4o's amplification of group differences is a general property or specific to the values and demographics the ESS measures.","If LLM similarity ratings are driven by training-text statistics rather than a stable self, Value Portrait's absolute profile may partly reflect the moral vocabulary dominant in alignment data; the benchmark would still reliably rank models relative to one another, a use that survives even if the absolute profile does not."],"forward_implications":["Any new LLM can be audited directly: run the released items, average over the six prompts, and read off mean-centered value scores that are internally consistent ($\\alpha \\geq 0.76$).","The uniform profile across 44 models implies that instruction tuning and safety alignment have pushed the field's models toward a shared value posture, so claims of distinctive brand-specific values become deviations to be explained rather than defaults.","Persona auditing against representative human survey data becomes a standard bias check: GPT-4o's exaggerated gender, age, and political gaps are concrete, measurable risks for anyone generating synthetic demographic data with LLMs.","Prompt steering is not a reliable value-control knob: while Universalism, Power, Hedonism, and Self-Direction respond to steering, Benevolence does not, so alignment interventions need measurement rather than assumed prompt effects.","The identical pipeline yields Big Five trait scores, so one benchmark covers both value and personality assessment without re-annotation."],"supporting_citations":[{"why":"Supplies the Portrait Values Questionnaire (PVQ-21), the instrument whose dimension scores anchor the correlation-based item annotation.","marker":"(Schwartz, 2003)"},{"why":"Provides the ten-dimension theory of basic values that forms the benchmark's value taxonomy and its predicted circular structure.","marker":"(Schwartz, 1992)"},{"why":"The psychometric tradition of validating new items by correlating them with established value measures, which the annotation method explicitly follows.","marker":"(Davidov et al., 2008)"},{"why":"Source of real-world human-LLM queries (LMSYS-Chat-1M) used to build ecologically valid items.","marker":"(Zheng et al., 2024)"},{"why":"Source of human-human advisory queries (Scruples, from Reddit's AITA forum) used to broaden value-laden scenarios.","marker":"(Lourie et al., 2021)"},{"why":"Supplies BFI-10, used to extend the correlation annotation to the Big Five personality traits.","marker":"(Rammstedt and John, 2007)"},{"why":"Justifies the $|r| \\geq 0.3$ threshold as a moderate correlation for retaining items.","marker":"(Cohen, 1988)"},{"why":"Supplies the persona steering method used to probe GPT-4o's demographic value perceptions.","marker":"(Hu and Collier, 2024)"},{"why":"Provides the alpha coefficient used to report the benchmark's internal-consistency reliability.","marker":"(Cronbach, 1951)"},{"why":"Human baseline data from 37,498 participants for comparing persona-prompted value profiles against real demographic differences.","marker":"European Social Survey (ESS)"}],"fun_headline_variants":["LLMs put benevolence, security, self-direction first; tradition, power last","Validated quiz shows LLMs value benevolence, security, self-direction most","LLMs share a value hierarchy: benevolence top, tradition bottom","Psychometrics reveal LLM value profile: benevolence, security, self-direction","LLMs' values measured: benevolence first, power nearly last"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"That an LLM's answer to \"How similar is this response to your own thoughts?\" registers the model's actual value orientation the same way the identical question registers a human's, rather than capturing prompt compliance, socially desirable mimicry, or statistics of the model's training text instead; the paper's own appendix concedes this assumes outputs reflect the model's internal preferences.","fun_headline_variants_meta":{"raw":{"variants":["LLMs put benevolence, security, self-direction first; tradition, power last","Validated quiz shows LLMs value benevolence, security, self-direction most","LLMs share a value hierarchy: benevolence top, tradition bottom","Psychometrics reveal LLM value profile: benevolence, security, self-direction","LLMs' values measured: benevolence first, power nearly last"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00082,"raw_usage":{"total_tokens":3664,"prompt_tokens":1097,"completion_tokens":2567,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":713,"completion_tokens_details":{"reasoning_tokens":2464}},"tokens_in":713,"tokens_out":2567,"duration_ms":15408,"temperature":1.0,"reasoning_tokens":2464,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:28:27.601498+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 520 items on an unaligned base model (no instruction tuning or RLHF) or on a model instructed to role-play a maximally unhelpful assistant: if the profile still shows high Benevolence, Security, and Self-Direction, the instrument is reading training-text priors rather than model values. A complementary check is behavioral: have each model produce free-form answers to the same 104 queries, then test whether models that self-report higher Benevolence on Value Portrait also generate measurably more benevolent text; if self-reports and behavior diverge, the self-report reading fails.","supporting_citations":[],"review_version":1}