{"id":"473353dc-9fed-4577-8573-09420252d735","arxiv_id":"2608.10327","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A systematic annotation of 94 AI alignment papers shows the field largely equates human values with measurable preferences, rarely defines values, and is increasingly removing humans from alignment evaluation.","lead":"This paper analyzed 94 AI alignment papers and found that most never define human values, instead treating them as preferences to be optimized, often with AI raters standing in for people. It matters because governments and companies are spending heavily on so-called value-aligned AI, and the field's unexamined assumptions may be shrinking what counts as a human value.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Field-level percentages rest on a citation snowball from four preference-RLHF seeds; the sample is a preference-paradigm citation neighborhood, not the field, and the paper's own Limitations section concedes no completeness claim.","rationale":"I read the paper in good faith first. The qualitative core is credible: the quoted passages and examples strongly support the observation that many influential alignment papers use 'preferences' without defining 'values,' and the philosophical critique of utility maximization is coherent and well-sourced. The load-bearing problem is not the qualitative observation but the quantitative field-level inference built on it. The 82% preferences-as-values and 87% utility-maximization figures are presented as properties of the AI value alignment field, yet the sampling frame is a snowball from four seeds that all belong to the preference/RLHF paradigm. That makes the sample a citation neighborhood centered on the very framework the paper critiques. The paper's own Limitations section concedes that no completeness or comprehensiveness is claimed, which is honest but does not cure the mismatch between the sample and the field-level language in the abstract and conclusion. The proposed diversified-seed sensitivity analysis would settle whether this bias is material: if a broader, independently seeded corpus yields similar percentages, the concern largely evaporates; if it shifts the percentages substantially, the headline statistics should be reframed. I agree with the reader's weakest assumption and recommend no change to the CONDITIONAL verdict; I would not escalate to rejection because the qualitative evidence stands independently and the authors are transparent about their limitations.","tokens_in":21501,"tokens_out":5191,"duration_ms":52574,"concrete_test":"Reconstruct the corpus with a diversified seed set: add at least four non-preference seeds, e.g., Gabriel 2020 (philosophical survey), Bai et al. 2022 Constitutional AI, Ji et al. 2023 comprehensive survey, Kasirzadeh 2024a or Sorensen et al. 2024 (pluralistic alignment), and Zhi-Xuan et al. 2024 (Beyond Preferences). Use the same Semantic Scholar API, the same 2016-2024 window, the same citation-count truncation (or, better, keep all unique papers), and re-annotate with the same rubric, ideally releasing the full annotated dataset. Then compare the six key prevalence statistics (defines value, preferences-as-values, values measurable, utility maximization, individual vs. collective, static/dynamic) between the original sample and the diversified sample.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim is a statement about the field: most alignment research does not define values and defaults to preferences and utility maximization. The evidence base is 94 papers obtained by snowballing citations to Christiano et al. 2017, Askell et al. 2021, Bai et al. 2022, and Ouyang et al. 2022 (Methods, Snowball sampling). All four seeds are canonical preference/RLHF-style papers. Snowball sampling from these seeds yields the set of papers that cite them, which is a citation neighborhood centered on the preference paradigm, not a sample of value-alignment research as a whole. Papers that do not engage with those four seeds, such as constitutional AI built on AI feedback, principle-based or normative-ethics alignment, and human-rights frameworks, can enter the sample only if they happen to cite one of the preference-RLHF seeds, and are then subject to a citation-count truncation that further weights toward the same paradigm. The annotation filter ('discarded if the term value alignment was used only in a cursory sense') is applied after sampling and cannot repair the frame. The paper's own Limitations section states 'we make no claims to the completeness or comprehensiveness of the sampled papers,' yet the abstract and conclusion report 82% and 87% as properties of the field. If the sampling frame omits a substantial non-preferential strand of alignment work, the quantitative headline is an artifact of seed choice rather than a measurement of the field. This does not undermine the qualitative observation that many influential papers leave 'values' undefined, but it does undercut the specific percentages.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper annotates 94 AI value alignment papers, sampled by snowballing citations from four canonical preference/RLHF seed papers, using a 13-question rubric drawn from philosophy and anthropology. It reports that most sampled papers do not define values, commonly equate values with preferences, treat values as measurable, static, and individual, and rely on utility maximization; it argues that the field has implicitly adopted an economic theory of value, and it discusses alternative anthropological and psychological conceptions. The central qualitative claim is that the value alignment literature rarely engages critically with what human values are.","tokens_in":21908,"tokens_out":5791,"duration_ms":53257,"significance":"If the quantitative claims were properly scoped to the sample, the paper would be a useful critical intervention. The qualitative observation—that many influential alignment papers use 'values' and 'preferences' interchangeably without definition—is well supported by quoted examples and is robust to the authors' own criteria. The rubric is a transparent instrument for interrogating the philosophical commitments of alignment research, and the paper usefully brings anthropology and social philosophy into a debate that is usually conducted in technical terms. However, the field-level percentages are not established by the sampling design, and several rubric categories have near-chance inter-annotator agreement.","major_comments":[{"comment":"The field-level claim that 87% of papers implicitly or explicitly use utility maximization, and the closely related 82% figure for preferences-as-values, rest on a sample created by snowballing citations from four seed papers that all belong to the preference/RLHF paradigm (Christiano et al. 2017; Askell et al. 2021; Bai et al. 2022; Ouyang et al. 2022). Snowball sampling returns a citation neighborhood centered on those seeds, not a random or representative sample of AI value alignment research; principle-based, constitutional, or human-rights-based alignment work enters the sample only if it happens to cite one of the four preference seeds. The annotation filter described in Methods ('discarded if the term value alignment was used only in a cursory sense') is applied after sampling and cannot repair this frame. The Limitations section explicitly disclaims completeness, yet the Abstract and Conclusion present the percentages as properties of the field. Please either re-run the annotation on an expanded seed set that includes non-preference paradigms and report sensitivity, or reframe all quantitative findings as descriptive of the preference-RLHF citation neighborhood only.","section":"Methods, Snowball sampling; Discussion, Utility maximization; Abstract"},{"comment":"The sentence in the Annotation process section, 'we achieved a high degree of inter-annotator agreement, which averages overall around 85%,' is not supported by Table 1. The raw percentage agreement in Table 1 ranges from 51.0% to 88.8%, and several rubric categories that feed the headline claims are at or near chance: 'Thin vs Thick' at 52.8% (Krippendorff's alpha 0.073), 'Monism vs Pluralism' at 54.0%, 'Values Static or Dynamic' at 51.0%, and 'Should AI follow human values?' at 56.9% (alpha -0.015). This is a load-bearing problem for the quantitative prevalence claims that depend on these categories. Please report chance-corrected agreement for each rubric item instead of an aggregate raw-agreement figure, and temper the claims that rely on the least reliable categories.","section":"Methods, Annotation process; Table 1"}],"minor_comments":[{"comment":"The second contribution bullet says 'snowballed sample of 100 commonly cited AI alignment papers,' but Methods and Findings report 94 annotated papers after discarding cursory mentions; make the numbers consistent.","section":"Statement of Contributions"},{"comment":"The numbered rubric lists 13 questions, but Table 1 contains 15 rows (for example, 'Sources for determining values' and 'Which humans' values?'); align the table with the rubric or explain the additional items.","section":"Methods, Annotation rubric; Table 1"},{"comment":"Figures 1 and 2 are referenced in the text but are not included in the submission; ensure the final version contains the figures and their captions.","section":"Findings, Figures 1 and 2"},{"comment":"There are several typos and wording slips: 'we conducted a analysis' (Introduction), 'a rise harms' (Abstract), 'The question ofwhat moral values are' (Related Work), and 'Think vs Thick' in the Table 1 header should be 'Thin vs Thick.'","section":"Throughout"},{"comment":"The sentence 'We sorted these from highest to lowest Google Scholar citations. We sorted the papers by Google Scholar citation count' repeats the same information; keep one formulation.","section":"Methods, Snowball sampling"},{"comment":"Several citations to web resources (Anthropic 2025, OpenAI 2025, Institute 2025) lack URLs and access dates; please complete these entries for reproducibility.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper's method and contribution are closer to a qualitative social-science study than to a typical empirical cs.AI submission; this is not a flaw if the venue welcomes critical and interdisciplinary work, but the fit should be assessed. The authors should be encouraged to release the annotated dataset and qualitative notes so that other researchers can audit the rubric scores, especially for the high-stakes percentage claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nWorth a read if you care about the alignment critique literature. The new thing is empirical: the authors annotated 94 papers with a rubric drawn from philosophy and anthropology, documenting how rarely 'values' is defined, how often preferences are substituted, and how utility maximization shows up as the technical frame. The qualitative spine is solid. They quote enough papers to show the elision between values and preferences is real, and the rubric distinctions (thick/thin, monism/pluralism, measurable vs immeasurable) give a useful vocabulary for talking about what alignment research actually assumes.\n\nCredit where due: they engage the existing critical literature (Gabriel, Zhi-Xuan, Lindström, Arzberger, Casper) rather than ignoring it. They also flag the turn to autoraters and synthetic feedback as a further step away from contested human values, which is a genuinely important observation.\n\nNow the soft spots. The headline percentages—79%, 82%, 87%—are reported as if they describe the field, but the sample is a citation snowball from four seed papers that all sit in the RLHF preference paradigm. That makes the 94 papers a preference-centered citation neighborhood, not a representative slice of value alignment research. The Limitations section explicitly disclaims completeness, but the abstract and conclusion do not carry that caveat. If the aim is to characterize the field, the claims need reframing as \"among papers in this citation network,\" with an explicit acknowledgment that the sampling frame biases toward the very paradigm being criticized.\n\nThe inter-rater reliability is also weaker than the \"around 85%\" claim suggests. Several key rubric items have Cohen's kappa near or below 0.3, including \"defines value\" (0.22) and \"static/dynamic\" (0.13). Raw agreement is inflated by skewed distributions. For a rubric meant to produce prevalence estimates, that is a real problem—some numbers likely reflect the annotators' interpretive frames as much as the papers.\n\nAnd the annotated dataset is not released, so nobody can re-analyze or check the coding. That limits the paper's value as an evidence base.\n\nNone of this kills the central qualitative point. I'm fairly convinced that much alignment research runs on an unexamined preference/utility theory of value. But I would treat the specific percentages as indicative, not measured, until the sampling and reliability issues are addressed.\n\nWould I send it to review? Yes—a serious referee will help the authors sharpen the claims and release the data. It deserves engagement, but it needs revision before its numbers should be quoted.","headline":"A useful critical meta-analysis whose qualitative core holds up, but the headline percentages should not be read as field-level measurements because the snowball sample is seeded entirely inside the RLHF preference paradigm.","tokens_in":22502,"tokens_out":2506,"would_cite":true,"duration_ms":26593,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI alignment research mostly skips defining values and defaults to utility maximization, this review finds.","keywords":["AI alignment","value alignment","human values","preferences","utility maximization","reinforcement learning from human feedback","value pluralism","anthropological theory of value"],"falsifier":"Apply the same rubric to a new 94-paper sample built from seed papers that are not preference-based—for example, Constitutional AI, human-rights-based alignment, or normative-ethics work—and count how many define values and use utility maximization; if most of those papers define values and substantially fewer than 87% use utility maximization, the field-level claim is falsified and the finding stands only for the preference-based paradigm.","tokens_in":23059,"feed_emoji":"⚖️","tokens_out":9692,"duration_ms":87185,"temperature":0.7,"pith_summary":"This paper argues that AI value alignment research has built an entire technical program without first answering what human values are. The authors annotated 94 highly cited alignment papers, drawn by snowballing from four canonical preference-based papers, and found that most never define 'values': 79% give no definition or description, 82% treat preferences as a stand-in for values, and 87% implicitly or explicitly embed utility maximization in their alignment framework. The paper claims this is not neutral engineering but the adoption of a specific economic worldview—values as measurable, individual, static, utility-maximizing preferences—while the field presents itself as aligning with human values generally. If the claim is right, the dominant alignment paradigm is silently foreclosing other conceptions of value, and 'value-aligned' claims made to governments and the public rest on unexamined philosophical commitments.","feed_headline":"94 alignment papers rarely define values, study finds","feed_subtitle":"Most equate values with preferences and 87% use utility maximization—an economic worldview, not neutral engineering.","key_machinery":"The load-bearing instrument is the annotation rubric: thirteen forced-choice questions, each with binary options or 'not applicable,' built from distinctions in philosophy and anthropology—monism vs. pluralism, measurable vs. immeasurable, revealed preferences vs. prescribed principles, abstract vs. concrete, individual vs. collective, dynamic vs. static, thin vs. thick, whether values can emerge autonomously in AI, whether the paper names which humans' values are targeted, whether rater pool size is reported, and whether utility maximization is used. The rubric translates an implicit philosophical position into a countable label, so that the prevalence of each commitment across the 94 papers can be estimated and reported as field-level percentages. The snowball-sampled corpus, seeded by four preference-based papers, supplies the cases to which the rubric is applied.","core_discovery":"The paper reports a systematic annotation of 94 highly cited AI value alignment papers, selected by snowball sampling from four canonical preference-based papers. Its central claim is that this literature has no explicit theory of human values: 79% of the sampled papers neither define nor describe what they mean by 'values,' 82% treat preferences as a stand-in for values, 90% treat values as measurable, 53% treat values as static, 66% give 'thin' characterizations, and 91% never specify which humans' values are being aligned. The paper further finds that 87% of the papers implicitly or explicitly use utility maximization as part of the technical framework, which it reads as an unacknowledged commitment to a specific economic worldview: values are individual, rational, monist, static, utility-maximizing preferences. The conclusion is that value alignment is not philosophically neutral; it is enacting a particular theory of value while presenting itself as aligning with human values in general.","pith_inferences":["Because all four starting papers belong to the same 'learn from human preferences' school, the headline percentages may describe that school's citation ecology more than the whole alignment field; a snowball seeded from constitutional or human-rights alignment work could produce different rates.","A direct test: apply the same rubric to a sample seeded from non-preference papers and compare rates; a large drop in utility-maximization use would reframe the finding from a property of 'the field' to a property of the dominant paradigm.","The observed 'alignment without humans' trend implies a governance consequence the paper leaves implicit: when models rate models, accountability for whose values are encoded becomes diffuse, and audits of training data may need to treat simulated raters as a distinct category from human raters.","The rubric could be repurposed as a disclosure tool: developers could state, for each deployed system, whether its theory of value is monist or pluralist, static or dynamic, preference-based or principle-based, before claiming 'value alignment.'"],"forward_implications":["If the field's implicit theory of value is economic preference satisfaction, then claims that models are 'aligned with human values' overstate what reinforcement learning from human feedback (RLHF) and direct preference optimization actually encode: at best they encode the preferences of small, usually undocumented rater pools.","The shift from human annotators to LLM-as-a-judge and synthetic preference data means values are increasingly enacted without any human input in the loop, which the paper argues closes off alternative methods for contesting and enacting values in foundation models.","Because 91% of the sampled papers never identify which humans' values they are aligning to, the field's default is a universalist, culture-free picture of value that hides the situated nature of the values actually being encoded.","Making these commitments explicit turns 'whose values?' from a background assumption into a design question that researchers, deployers, and regulators can ask of any alignment system.","The paper's reading of the 'pluralist turn' implies that adding diverse preferences to a reward model does not by itself escape the preference-based theory of value; pluralism is still being expressed inside the utility-maximizing framework."],"supporting_citations":[{"why":"Seed paper for the snowball sample; it establishes preference-based RLHF as the paradigm the corpus is drawn from.","marker":"(Christiano et al. 2017b)"},{"why":"Seed paper introducing the helpful/honest/harmless framing; one of four citation roots of the 94-paper set.","marker":"(Askell et al. 2021)"},{"why":"Seed paper for training a helpful and harmless assistant with RLHF; root of the sample.","marker":"(Bai et al. 2022)"},{"why":"Seed paper for instruction-following via human feedback (InstructGPT); root of the sample.","marker":"(Ouyang et al. 2022)"},{"why":"Supplies the 'preferentist' framing and the three assumptions about preferences that the rubric operationalizes.","marker":"(Zhi-Xuan et al. 2024)"},{"why":"Evidence for the paper's premise that alignment leaves 'what moral values are' undefined.","marker":"(Gabriel 2020)"},{"why":"Canonical statement of economic utility and revealed preference theory that the paper identifies as alignment's implicit theory of value.","marker":"(Becker 1976)"},{"why":"Anthropological theory of value used as the alternative against which the field's thin economic conception is assessed.","marker":"(Graeber 2001)"}],"fun_headline_variants":["Most AI alignment papers never define values, study finds","87% of alignment papers use utility maximization implicitly","Value alignment research often equates values with preferences","Study: 94 alignment papers lack explicit theory of value","Alignment papers' hidden economic assumptions revealed in study"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The entire statistical picture depends on the assumption that the four seed papers, all from the same preference-learning school, open a fair window onto the whole value alignment field, so the percentages computed from the 94 papers that cite them are treated as properties of the field rather than of that school.","fun_headline_variants_meta":{"raw":{"variants":["Most AI alignment papers never define values, study finds","87% of alignment papers use utility maximization implicitly","Value alignment research often equates values with preferences","Study: 94 alignment papers lack explicit theory of value","Alignment papers' hidden economic assumptions revealed in study"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000767,"raw_usage":{"total_tokens":3228,"prompt_tokens":976,"completion_tokens":2252,"prompt_tokens_details":{"cached_tokens":0},"prompt_cache_hit_tokens":0,"prompt_cache_miss_tokens":976,"completion_tokens_details":{"reasoning_tokens":2177}},"tokens_in":976,"tokens_out":2252,"duration_ms":15280,"temperature":1.0,"reasoning_tokens":2177,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:08:26.552274+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Apply the same rubric to a new 94-paper sample built from seed papers that are not preference-based—for example, Constitutional AI, human-rights-based alignment, or normative-ethics work—and count how many define values and use utility maximization; if most of those papers define values and substantially fewer than 87% use utility maximization, the field-level claim is falsified and the finding stands only for the preference-based paradigm.","supporting_citations":[],"review_version":1}