{"id":"1b6856d1-5f46-42e4-9aca-95ab4a6138ca","arxiv_id":"2501.07071","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Value Compass Benchmarks is a live, self-evolving platform that scores 33 LLMs across 27 value dimensions from four value systems, aiming to reveal true behavioral alignment with human values.","lead":"This paper introduces Value Compass Benchmarks, an online platform that evaluates the values and safety risks of 33 large language models using four value systems and dynamically generated test questions. It aims to measure whether AI models actually behave according to human values, not just whether they know the right answers, which matters for alignment and cultural fit.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The value scores and item selection both depend on the unvalidated CLAVE recognizer (Eq. 1 vs. final scoring), creating a closed loop; without human-label agreement on generated items, the platform's validity claim is not established.","rationale":"The paper's central claim is that the platform provides valid and informative behavioral value scores for 33 LLMs across 27 dimensions. For that claim to hold, the value recognizer F must be a valid measurement instrument: it must identify values in open-ended responses the way humans would, or at least in a way that has been independently shown to track human judgments. The paper does not supply that evidence. Worse, the formalism makes the same recognizer doubly load-bearing: it is used inside the item-selection objective in Eq. (1) and again in the final score computation. This creates a closed loop whose outputs can look informative even if F is systematically wrong, because the generator is explicitly optimized to produce items on which the recognizer yields divergent value readings across models. Without an external anchor, the reported disparities among models and the case-study 'value-behavior correlations' may be artifacts of CLAVE's inductive biases rather than properties of the LLMs. The direction of the work is sound—generative, dynamic evaluation is a reasonable response to contamination and ceiling effects—and the platform's interactive and cultural-visualization features are useful. But those strengths do not establish measurement validity. The reader identified precisely this closed-loop risk; I see no additional load-bearing concern that would change the conditional verdict. The proposed human-labeling check is the most direct way to settle whether the reported scores correspond to human-recognizable value conformity.","tokens_in":14478,"tokens_out":2707,"duration_ms":28467,"concrete_test":"Have 3 trained annotators label which value dimensions are reflected in a stratified sample of, say, 200 generated items × 2 responses per item from 10 of the 33 models, using the published dimension definitions. Compute agreement between majority human labels and CLAVE's labels (e.g., Cohen's kappa per dimension). Then recompute value scores for those models using only human labels on that sample and rank-correlate (Spearman) with the reported CLAVE-based scores. If per-dimension agreement falls below a pre-registered threshold (e.g., kappa < 0.6) or the rank correlation is low (< 0.7), the closed-loop pipeline is producing scores that do not reflect human-recognizable value conformity; the validity claim would need to be weakened and CLAVE recalibrated or replaced.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.2 defines the recognizer F (CLAVE) and Eq. (1). The item generator is optimized with p_i(v|x) = E_{y~p_i(y|x)}[p_F(v|x,y)]; the final scores are then s^v_i = E_{x~X_v, y~p_i}[F(x,y)]. Thus the same recognizer both constructs the test items and measures the outcome. Any systematic misreading by CLAVE—e.g., over-attributing certain values to certain response styles—will be amplified by item selection and then reported as a genuine value difference. The paper reports no human agreement study for CLAVE on the generated items, no independent recognizer check, and no end-to-end comparison of CLAVE-based scores against human judgments for the 33 models. Figures 4 and 5 are illustrative; they do not establish criterion validity. The user study (n=15) asks about perceived usefulness, not score accuracy. So the central claim of 'valid and informative behavioral value scores' rests on an unvalidated closed loop.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents Value Compass Benchmarks, an online platform for evaluating the values of 33 LLMs across four value systems (Schwartz Theory of Basic Values, Moral Foundation Theory, an LLM-specific value system, and a safety taxonomy). Its main technical contribution is a generative self-evolving evaluation paradigm: an item generator is optimized via Eq. (1) to produce test scenarios that maximize divergence in the value distributions recognized by CLAVE, and value scores are computed by feeding model responses to the same recognizer. The platform provides fine-grained scores, customized comparisons, and cultural-alignment visualizations. Section 3 reports quantitative comparisons, case studies, and a 15-participant user study.","tokens_in":14698,"tokens_out":5027,"duration_ms":48630,"significance":"The paper addresses a genuine gap: existing value benchmarks are mostly static and discriminative, and the proposed combination of pluralistic value systems, dynamically generated scenarios, and an interactive dashboard is a useful engineering contribution. If the scores were shown to be valid, the platform would be valuable for monitoring alignment and cultural fit of LLMs. Strengths include a public online leaderboard covering 33 recent models, transparent multi-dimensional reporting, external grounding in social-science value surveys for cultural alignment, and a fully specified pipeline that is easy to reproduce from Eq. (1). The main weakness is that the validity evidence is currently not sufficient.","major_comments":[{"comment":"The same recognizer F (CLAVE) is used both to select test items and to compute final value scores. Specifically, p_i(v|x) in Eq. (1) is estimated as E_{y~p_i(y|x)}[p_F(v|x,y)], and the reported score s^v_i is E_{x~X_v, y~p_i(y|x)}[F(x,y)]. The manuscript contains no human agreement study for CLAVE on the generated items, no comparison against an independent recognizer, and no end-to-end comparison of the platform's scores with human judgments for the 33 evaluated models. Without such validation, systematic misreadings by CLAVE are amplified by item selection and then reported as genuine value differences; the central validity claim is therefore not established.","section":"Section 2.2, Eq. (1) and the scoring formula"},{"comment":"The evidence that generative evaluation is more valid than discriminative evaluation is not statistically supported. The figure shows qualitative differences in scores and the text interprets them as 'overestimation' and 'vulnerabilities', but lower or more dispersed scores do not by themselves show that the generative scores are more accurate; they could reflect recognizer bias or differences in prompt difficulty. No error bars, confidence intervals, or significance tests are reported, and the comparisons use only two models in Fig. 4(a) and four in Fig. 4(b). The claim that the platform yields 'valid' scores for 33 LLMs would need a criterion-validity analysis, for example against human ratings on a sample of generated items.","section":"Section 3, Fig. 4"},{"comment":"The 15-participant study measures perceived usefulness, informativeness, and usability (SUS), not whether the value scores are correct. The participants were not asked to judge the accuracy of the scores on held-out items, and the results are reported as means without confidence intervals. The user study can support a claim about usability, but it should not be cited as evidence for the validity of the value measurements; a validation study with human labels and more participants is needed.","section":"Section 3, user study"}],"minor_comments":[{"comment":"The abstract and the full text present inconsistent title variants: the reader-facing title mentions 'Fundamental and Validated Evaluation', while the full text begins with 'A Comprehensive, Generative and Self-Evolving Platform'. Please align the title across versions.","section":"Title and abstract"},{"comment":"The phrase 'the first dynamic, online and interactive platform' is an overclaim: dynamic evaluation platforms and leaderboards exist (e.g., DyVal, LMSYS Chatbot Arena). Suggest softening to 'a dynamic, online and interactive platform'.","section":"Section 1"},{"comment":"Equation (1) depends on a hyperparameter α and on the number of response samples per item, but no sensitivity analysis or default-value justification is given. Please report how α was set and whether the ranking is stable under reasonable choices.","section":"Section 2.2, Eq. (1)"},{"comment":"Minor typos: 'we presents' in the abstract, 'demdominance' and 'encomapesses' in Appendix A. The Power definition should read 'dominance'.","section":"Appendix A"},{"comment":"Figure 4 would benefit from error bars and a description of how many items and samples were used; without that, the visual gap between static and evolving items cannot be assessed.","section":"Section 3, Fig. 4"}],"recommendation":"major_revision","confidential_remarks":"The reader's concern about the closed-loop validity is well founded. The paper is closer to a systems/demo contribution than a measurement-validity study; as a systems paper it is strong, but the current validation section does not support the validity claim. I would encourage the authors to add a human-annotation study of CLAVE on generated items, an independent recognizer check, and statistical reporting before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper is worth a look, but the headline claim needs to be read with caution. The Value Compass Benchmarks is a genuinely new artifact: an online platform that scores 33 LLMs across 27 value dimensions drawn from four value systems, with customizable weighting and cultural alignment analysis. The integration and deployment are real contributions, and the move from static discriminative benchmarks to generative, self-evolving items is a sensible direction. I'd trust this as a useful tool for model selection and for structuring discussions of value pluralism.\n\nThe problem is that the paper's validity argument does not close. The item generator is optimized using CLAVE, the group's own value recognizer, to find items that make models look different; the final value scores are then computed using the same CLAVE on those very items. If CLAVE misreads response styles, both the item selection and the scores are biased in the same direction. The paper reports no human agreement study for CLAVE on the generated items, no independent recognizer check, and no end-to-end validation of the resulting scores against human judgments. Figures 4 and 5 are illustrative, not evidential, and the user study (n=15) measures perceived usefulness, not score accuracy. The absence of error bars in Fig. 4 makes the discriminative-vs-generative comparison hard to interpret.\n\nI don't think this is a fatal flaw in the platform's utility—the scores are still informative for relative comparisons within the system's own logic. But the claim of \"valid and informative behavioral value scores\" (Section 1) is not established. The authors need to show that CLAVE's labels on generated items agree with human raters, and ideally that the scores predict known cultural differences or independent human judgments.\n\nThat said, this deserves serious peer review. It's a substantial, well-written systems paper with a clear practical purpose. The closed-loop concern is exactly the kind of thing referees should probe, and the authors have the tools to address it. I'd send it to review with the expectation of major revision.\n\nReading group: probably yes if you want to discuss what counts as validation for generative benchmarks. Cite? I wouldn't cite it in my own work in the next year, but I'd point to it as an example of the genre.","headline":"A genuinely useful platform for LLM value comparison, but the paper's validity claim rests on an unvalidated closed loop that needs human-judgment evidence before the scores can be taken at face value.","tokens_in":15246,"tokens_out":2228,"would_cite":false,"duration_ms":22271,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Value Compass Benchmarks is a dynamic, online platform for diagnosing LLM values through generative, self-evolving tests.","keywords":["LLM value evaluation","value alignment","generative evaluation","self-evolving benchmark","value pluralism","cultural alignment","basic human values","leaderboard"],"falsifier":"Sample a set of open-ended responses from the 33 evaluated LLMs to freshly generated items, have independent human annotators label the value dimensions expressed, and compare their labels with CLAVE's; if agreement is low or if replacing CLAVE with a human-validated recognizer changes the model rankings, the platform's validity claim fails.","tokens_in":14298,"feed_emoji":"🧭","tokens_out":13023,"duration_ms":102830,"temperature":0.7,"pith_summary":"The paper argues that existing LLM value benchmarks fail on validity grounds: static, discriminative questions measure what models know about values rather than how they behave, and they become uninformative as models improve. To fix this, it presents Value Compass Benchmarks, an online platform that automatically generates fresh test items, scores open-ended responses through a value recognizer, and updates the items as new models are released. The platform reports fine-grained, behavior-based value scores for 33 leading LLMs across 27 dimensions and adds tools to compare models, weight scores by personal priorities, and see cultural alignment. If the approach works, value evaluation becomes a living diagnostic of behavioral conformity rather than a one-dimensional knowledge test.","feed_headline":"Live platform scores LLM values from behavior, not quizzes","feed_subtitle":"Fresh generated scenarios, not static quizzes, reveal how 33 models behave across 27 value dimensions.","key_machinery":"The load-bearing mechanism is the two-term item-generation objective in Eq. (1), which trains the generator to produce test items that both elicit values, by maximizing mutual information between responses and the value-dimension vector, and remain informative, by maximizing divergence among the evaluated models' predicted value distributions. The recognizer CLAVE supplies the value-probability estimates inside that objective and again when scoring responses, so the same function carries both item selection and the final behavioral score. A model's score on a value dimension is the expected recognizer output over generated items and sampled responses, which is what makes the evaluation behavioral rather than knowledge-based.","core_discovery":"Value Compass Benchmarks is presented as the first dynamic, online and interactive platform devoted to comprehensive value diagnosis of LLMs, and the paper's central claim is that generative, self-evolving evaluation is more valid than static discriminative benchmarks. Instead of asking models to select the value-aligned answer, the platform generates novel, value-evoking scenarios, samples each model's open-ended response, and uses a hybrid value recognizer (CLAVE) to estimate the value distribution expressed by that behavior. The value score for a model on a dimension is the expected recognizer output over generated items, and the item generator is re-optimized at Eq. (1) whenever newer models arrive, balancing value elicitation against informativeness. The platform reports fine-grained scores across 27 dimensions from four value systems, supports user-weighted aggregation through social welfare functions, and maps each model onto cultural value vectors for alignment analysis.","pith_inferences":["Inference: Because the same recognizer CLAVE both selects test items and computes final scores, re-running the platform with a recognizer whose labels are validated against human judgments would be a direct robustness check; the reported model differences are partly a property of the recognizer.","Inference: The self-evolving generative design could transfer to other pluralistic targets, such as political values, professional ethics codes, or organization-specific value statements, where static questionnaires face the same contamination and ceiling problems.","Inference: Cultural alignment maps built from survey-reported value vectors could double as a training-data diagnostic, since a model trained largely on one region's text should approximate that region's value profile unless alignment training shifts it.","Inference: The case studies imply value scores track observable behavioral differences; adding high-stakes or adversarial scenarios would test whether those differences persist when a model is under pressure."],"forward_implications":["A model that can recite the value-aligned answer but does not act on it in realistic scenarios will receive a low behavioral score, closing the knowledge-behavior gap.","The benchmark regenerates test items as new LLMs appear, so evaluations stay informative and contaminated or saturated items are replaced.","Users can supply their own weights over value dimensions through social welfare functions, making the best model relative to a person's priorities instead of a single average.","Cultural alignment analysis reveals which documented cultural value vectors each model resembles, giving developers a concrete target for cultural adaptation.","Fine-grained scores across 27 dimensions support case-level diagnosis of specific misalignments, rather than only an overall safety ranking."],"supporting_citations":[{"why":"Supplies the self-evolving item generator and the two-term optimization objective in Eq. (1).","marker":"(Jiang et al., 2024)"},{"why":"Provides CLAVE, the hybrid value recognizer used for both item selection and final scoring.","marker":"(Yao et al., 2024)"},{"why":"Introduces the generative evaluation schema that scores values from open-ended behavior rather than predefined answers.","marker":"(Duan et al., 2023)"},{"why":"Contributes the dynamic evaluation framing that motivates co-evolving test items.","marker":"(Zhu et al., 2023)"},{"why":"Defines the ten basic human values used as one of the four value systems.","marker":"(Schwartz, 2012)"},{"why":"Defines Moral Foundation Theory with its five moral foundations.","marker":"(Graham et al., 2013)"},{"why":"Provides the LLM-specific value system with three core dimensions.","marker":"(Biedma et al., 2024)"},{"why":"Supplies the hierarchical safety taxonomy used for the safety value system.","marker":"(Li et al., 2024)"},{"why":"Provides the social welfare function framework used for customized score aggregation.","marker":"(Arrow, 2012)"}],"fun_headline_variants":["Behavior-based scoring reveals LLM values instead of quiz knowledge","Adaptive platform tests LLM values through evolving scenarios","The platform scores 33 models on 27 value dimensions via generated behaviors","Value Compass scores LLMs by what they do, not what they know","Pluralistic weighting lets users tailor LLM value scores"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire pipeline assumes that the value recognizer CLAVE labels open-ended responses as reliably as a human would, because the same recognizer is used both to choose which test items to generate and to compute the final value scores.","fun_headline_variants_meta":{"raw":{"variants":["Behavior-based scoring reveals LLM values instead of quiz knowledge","Adaptive platform tests LLM values through evolving scenarios","The platform scores 33 models on 27 value dimensions via generated behaviors","Value Compass scores LLMs by what they do, not what they know","Pluralistic weighting lets users tailor LLM value scores"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000719,"raw_usage":{"total_tokens":3239,"prompt_tokens":967,"completion_tokens":2272,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":583,"completion_tokens_details":{"reasoning_tokens":2185}},"tokens_in":583,"tokens_out":2272,"duration_ms":16956,"temperature":1.0,"reasoning_tokens":2185,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:48:58.405347+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Sample a set of open-ended responses from the 33 evaluated LLMs to freshly generated items, have independent human annotators label the value dimensions expressed, and compare their labels with CLAVE's; if agreement is low or if replacing CLAVE with a human-validated recognizer changes the model rankings, the platform's validity claim fails.","supporting_citations":[{"cited_title":"Denevil: Towards Deciphering and Navigating the Ethical Values of Large Language Models via Instruction Learning","cited_arxiv_id":"2310.11053","evidence_quote":"Introduces the generative evaluation schema that scores values from open-ended behavior rather than predefined answers."}],"review_version":1}