{"id":"4baef143-2fc6-410d-b9f3-3473c93f46f6","arxiv_id":"2502.10441","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"The paper formalizes alignment discretion and shows empirically that annotators and models exercise substantial, often arbitrary, and mutually divergent discretion when applying alignment principles.","lead":"The paper defines 'alignment discretion' as the unexamined freedom annotators have to decide which AI outputs are better or safer, and proposes metrics to measure when and how it is exercised. Applying these metrics to safety alignment datasets reveals high levels of arbitrary human judgments and large discrepancies between human and algorithmic discretion, suggesting RLHF may not transfer human values to models.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The empirical headline depends on GPT-4o as the principle oracle, which is also an audited annotator; this feedback loop can manufacture consensus and near-zero arbitrariness for GPT-4o, so the reported magnitudes are not robust.","rationale":"The reader's weakest assumption is exactly the load-bearing concern: Def. 5 assumes a perfect oracle for principle-specific preferences, and the paper instantiates that oracle with GPT-4o, the same model whose LLM preferences are then audited. The near-zero arbitrariness of GPT-4o in Tab. 1 is a visible symptom of the feedback loop. If GPT-4o's principle judgments are biased or noisy, every downstream classification and metric inherits that bias. This is not fatal to the paper's conceptual contribution—the formalization of discretion is principle-agnostic, and the qualitative point that vague principles leave room for unexamined discretion is independently plausible. But the quantitative claims about how much discretion exists, how arbitrary humans are, and how far models diverge from human discretion are contingent on one proprietary oracle, and no code is released to re-run the computations. A re-run with an independent oracle is the single check that would settle whether the empirical magnitudes hold. I therefore keep the reader's CONDITIONAL verdict; the concern is real but does not overturn the paper's main qualitative conclusion.","tokens_in":37120,"tokens_out":4292,"duration_ms":42439,"concrete_test":"Recompute Fig. 3, Tab. 1, and Tab. 2 with Claude 3.5 Sonnet (or a panel of human experts) as the principle oracle while keeping the audited annotator set identical, and report per-principle oracle agreement on Prefc. If the consensus/conflict/indifference labels change on more than ~10% of pairs, or if DA for humans or DD for RLHF models moves outside the published bootstrap CIs, then the reported magnitudes are oracle-dependent and should be re-stated as conditional on GPT-4o.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central numbers—Fig. 3, Tab. 1, Tab. 2, and the supremacy/priority figures—are all computed from GPT-4o's principle-specific preferences Prefc (Def. 5, Sec. B.3). The paper's own Tab. 1 exposes the circularity: since GPT-4o is both oracle and audited LLM, its arbitrariness is 0.65% on HH, far below every other annotator. That is expected if the oracle and the annotator share the same judgment model, but it means 'consensus' is not an independent ground truth; it is GPT-4o's own opinion of what the 21 principles require. Human 'arbitrariness' (28.9% on HH) is therefore disagreement with GPT-4o's interpretation, not necessarily irrationality, especially because the authors concede in Sec. 8 that judging even 'reject cruelty' requires discretion. A secondary issue is that HH labels come from different crowdworkers per item; pooling them into one 'human annotator' (Sec. 6, Fig. 4) mixes inter-annotator pluralism with individual discretion. The formal framework is principle-agnostic and survives this, but the empirical magnitude claims do not.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces the concept of 'alignment discretion' to describe the latitude that annotators, human or algorithmic, have in deciding which model outputs are 'better' or 'safer' during preference-based AI alignment. Drawing on legal theory, the authors formalize principle-specific preferences and define several metrics: principle consensus/conflict/indifference, discretion arbitrariness, principle supremacy, ELO-style principle priority weights, and discretion discrepancy. They apply these metrics to the HH-RLHF and PKU-SafeRLHF datasets, using GPT-4o as an oracle for 21 principles from Collective Constitutional AI, and compare human, reward-model, and LLM annotators. The paper reports, among other results, 28.9% human discretion arbitrariness on HH, reward-model discrepancies around 14-20%, and RLHF-tuned LLM discrepancies up to 71.2%, and concludes that current alignment processes allow excessive and unexamined discretion.","tokens_in":37305,"tokens_out":6291,"duration_ms":54847,"significance":"The formal framework is a useful step: the definitions are crisp, largely principle-agnostic, and the paper carefully reports bootstrap standard errors, controls for positional bias by swapping response order, and uses separate train/test splits. The legal analogy and the proposed metrics could provide a common vocabulary for discussing annotator latitude and principle prioritization in alignment research. However, the empirical magnitudes that carry the paper's central claims are computed through a single GPT-4o oracle that is also one of the audited annotators, and the paper's own Sec. 8 acknowledges the resulting feedback loop. For that reason, the headline numbers should be read as conditional on the oracle's interpretation of the 21 selected principles, not as an independent measure of arbitrariness or of failure to transfer human discretion. The framework itself survives this concern, but the empirical support for the broad conclusions needs substantial reworking.","major_comments":[{"comment":"The load-bearing empirical quantities are not independent of the oracle. Def. 5 sets Prefc = Preforacle for every principle, and Sec. B.3 instantiates the oracle with GPT-4o; all consensus/conflict/indifference classifications (Fig. 3), arbitrariness rates (Tab. 1), supremacies (Fig. 4), priorities (Fig. 5), and discrepancies (Tab. 2) are computed from GPT-4o's judgments. Tab. 1 accordingly reports GPT-4o's arbitrariness as 0.65% on HH, which the paper itself attributes to the oracle/annotator overlap. This is not merely a residual limitation acknowledged in Sec. 8; it changes the interpretation of every headline magnitude. A human label that disagrees with the GPT-4o-based consensus is counted as 'arbitrary' even when it reflects a defensible interpretation of an abstract principle, and the paper concedes in Sec. 8 that even 'reject cruelty' requires discretion. I therefore cannot read 28.9% (HH human) or 71.2% (Llama-3 fine-tuned) as estimates of excessive discretion in alignment; they are estimates of disagreement with one model's interpretation of 21 fixed principles. The framework survives, but the empirical claims need to be reframed as conditional on the oracle, or better, re-run with independent oracles and/or human principle-specific labels; at minimum, the sensitivity of Fig. 3 and Tab. 1 to the choice of oracle should be reported.","section":"Sec. 6.1 / Def. 5 / Sec. B.3 / Tab. 1"},{"comment":"The analysis treats the entire HH preference set as if it were produced by a single 'human annotator', even though HH labels come from different crowdworkers per item and the paper's own Sec. 2 cites Anthropic's reliance on crowdworker diversity. For principle supremacy and priority, the metric pools all these labels into one Bernoulli estimate per principle pair. This conflates inter-annotator pluralism with the discretion of a single decision-maker, which is exactly the distinction the paper needs to make to support claims about 'annotators' having excessive discretion. Reporting annotator-level analyses, or at least clarifying that the human baseline is the aggregated dataset preference, is necessary before the claim that human annotators frequently use their power of discretion arbitrarily (Sec. 7) is supported.","section":"Sec. 6.2 / Fig. 4 / Def. 8"},{"comment":"The claim that RLHF fails to transfer human discretion to LLMs is based on an apples-to-oranges comparison. Reward-model preferences are defined as the sign of a scalar reward difference (Def. 2), while LLM preferences are elicited through a textual template (Def. 3). Tab. 2 shows fine-tuned reward models with DD around 14-20% but the corresponding RLHF-tuned LLMs at 40-70%; before interpreting this gap as a fundamental limitation of RLHF, the authors need to control for the evaluation modality, for example by scoring the fine-tuned policy's outputs with the reward model used in training, or by eliciting reward-model preferences with the same textual template. With only two base models, one reward model per dataset, and no random-seed variation, the claim that translating human discretion from reward models to LLMs is an open problem (Sec. 6.2 and abstract) is stronger than the evidence.","section":"Sec. 6.1 / Def. 2-3 / Tab. 2"},{"comment":"Several priority estimates rest on very small conflict counts. In Fig. 4, many cells report conflict counts below 10, and the priority weights in Eq. (13) are fitted to such sparse supremacies. The bootstrap standard errors in Tab. 2 appear to capture resampling over dataset items only; they do not propagate oracle-instance variability or the choice of principle set, both of which are substantial because a different oracle can reclassify a pair as consensus versus conflict. Reporting the stability of w* and DD across oracle models and principle subsets would make the quantitative comparisons in Tab. 2 interpretable.","section":"Def. 9 / Fig. 4 / Tab. 2"}],"minor_comments":[{"comment":"The phrase 'human and algorihmic discretion' contains a typo ('algorihmic' should be 'algorithmic').","section":"Sec. 3"},{"comment":"The caption refers to 'Def 5.1a', 'Def 5.1b', and 'Def 5.1c', which do not match any numbered definition; these should likely be Def. 6a, 6b, and 6c.","section":"Fig. 2 caption"},{"comment":"Some cells display non-zero percentages with '(0)' conflict counts or otherwise inconsistent counts (e.g., '40% (0)' and '80% (0)'); these should be reconciled with the stated totals or the notation should be explained.","section":"Figs. 4 and 14"},{"comment":"The RLHF training section mentions hyperparameter sweeps but does not report the chosen values for learning rate, batch size, KL coefficient, or number of PPO epochs; concrete configurations are needed for reproducibility, especially because the paper does not provide a code or data release link.","section":"Appendix B.5"}],"recommendation":"major_revision","confidential_remarks":"The paper's formal contribution is real and the topic is timely, but the empirical sections are currently framed as if they measure discretion directly, whereas they measure disagreement with a single GPT-4o oracle. I would advise the editor that the manuscript is suitable for major revision rather than rejection, provided the authors rework the empirical claims to be explicitly oracle-conditional and add sensitivity analyses. A policy-oriented reading of the current abstract could easily overstate the findings, so the framing of the headline numbers should be corrected before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth reading for the formalism, not for the headline numbers. The paper gives a clean vocabulary and measurement framework for something real: the latitude annotators have when principles conflict, are silent, or leave room for judgment. The consensus/conflict/indifference trichotomy is simple but illuminating, and the supremacy/priority/arbitrariness/discrepancy metrics are well-defined and principle-agnostic. The legal analogy is not window-dressing; it motivates the distinction between arbitrary and principled discretion. I gave credit for the authors explicitly flagging the oracle limitation in Sec. 8, and for the bootstrap standard errors and positional-bias checks. That is honest work.\n\nThe soft spot is exactly what the stress-test note says: GPT-4o is both the oracle that defines principle-specific preferences and one of the audited annotators. Its 0.65% arbitrariness on HH is essentially tautological, and the comparison of other models' arbitrariness against a GPT-4o-defined consensus conflates substantive disagreement with noise or alternative interpretations. The authors acknowledge this, but the abstract and conclusion still push the excessive-discretion claim without consistently carrying the caveat. Also, no code or data is released, so I could not verify the computations. The 'human annotator' in HH-RLHF is an aggregation of many crowdworkers, which mixes inter-annotator pluralism with individual discretion; worth stating as a limitation. And the phrase 'challenges the purpose of having any principles' overshoots: the results show that principles are incomplete and applied inconsistently, not that principles are pointless.\n\nThat said, the central idea survives. Discretion in alignment datasets is real, unmeasured, and under-studied. The framework is a useful starting point, and the qualitative conclusion would not be overturned by re-running with a separate oracle (or excluding GPT-4o from the audited set). The empirical magnitudes are not fixed measurements; they are one plausible instantiation.\n\nI would bring this to reading group and cite the formal definitions in future work. The paper deserves a serious referee: the conceptual contribution is solid, the flaws are fixable, and a good revision could turn it into a reference for discretion metrics. Send it to peer review.","headline":"A genuinely useful formalization of discretion in alignment, but the empirical headline numbers rest on a single-model oracle that also sits in the audited set.","tokens_in":37893,"tokens_out":1722,"would_cite":true,"duration_ms":638003,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that AI alignment hides an excessive, largely unexamined form of discretion: annotators and models, not principles, decide which outputs are 'better' or 'safer', and those choices are often arbitrary.","keywords":["AI alignment","AI safety","alignment discretion","human feedback","RLHF","principle prioritization","reward models","judicial discretion"],"falsifier":"Take a random subsample of the same two preference datasets, have at least two independent oracle systems or a panel of human judges label which response better adheres to each of the 21 principles, and recompute discretion arbitrariness and discretion discrepancy. If the human 28.9% arbitrariness rate or the priority rankings move by more than the reported bootstrap standard errors, the paper's diagnosis is an artifact of the chosen oracle; if they are stable, the existence of large, arbitrary discretion is confirmed.","tokens_in":1865,"feed_emoji":"⚖️","tokens_out":2379,"duration_ms":106410,"temperature":0.7,"pith_summary":"AI alignment is built on pairwise choices: annotators and models decide which response is 'better' or 'safer.' The paper argues that the latitude granted by those choices, which it calls alignment discretion, is excessive, mostly unexamined, and frequently arbitrary, since abstract principles routinely conflict or give no guidance in the cases that matter. To make the phenomenon measurable, the paper defines when discretion is required (principle conflict or indifference) and how it is exercised (whether it contradicts a principle consensus, and which principles win when they clash). Across two standard safety-preference datasets, human annotators disagreed with a unanimous verdict of all principles 28.9% of the time on one dataset and 14–20% on the other, while an RLHF-fine-tuned model's principle ranking diverged from human annotators' by up to 71.2%. If the claim holds, feedback-based alignment encodes unexamined value judgments and should not be treated as a transparent grounding for safety.","feed_headline":"28.9% of AI safety labels defy agreed principles","feed_subtitle":"Audit of two safety datasets finds 28.9% human arbitrariness and up to 71.2% model divergence.","key_machinery":"The central object is alignment discretion, operationalized through four linked measurements built on ternary preference functions. A preference function returns +1, −1, or 0 for each response pair; principle-specific preference functions use an assumed oracle to score how well each response adheres to one principle at a time. Given all principle votes on a pair, the taxonomy classifies the pair as principle consensus (all non-indifferent principles agree), principle conflict (principles disagree), or principle indifference (all abstain). Discretion arbitrariness is the frequency with which an annotator picks the response opposed to a consensus; principle supremacy is the empirical probability that one principle wins over another when they clash; principle priority fits an ELO-style logistic model to those pairwise win frequencies to produce a single ranking per annotator; and discretion discrepancy is the normalized Kendall-tau distance between two annotators' rankings. This machinery turns the legal notion of discretion—when it is required, how it is exercised, and whether it is consistent across decision-makers—into numbers that can be computed for any preference dataset.","core_discovery":"On the authors' own terms, the central discovery is that alignment discretion is both necessary and currently out of control: annotators must exercise judgment precisely because principles conflict or are indecisive, yet the field has no systematic record of how that judgment is used. The paper formalizes the three situations principles can be in—consensus, conflict, and indifference—and then measures, per annotator, how often they contradict a consensus (discretion arbitrariness), which principles win when two conflict (principle supremacy), the one-dimensional priority ordering implied by those wins, and how far any two annotators' orderings are apart (discretion discrepancy). Human labels in the first dataset contradicted a unanimous principle verdict 28.9% of the time, and an RLHF-fine-tuned model diverged from human principle priorities by 71.2%; reward models fine-tuned on the same preference data stayed within roughly 15–20% ranking discrepancy, while off-the-shelf models sat between 16% and 53%. The authors conclude that there is currently an excessive amount of discretion in the hands of model developers and annotators, that principles alone underdetermine aligned behavior, and that algorithms develop their own forms of discretion rather than inheriting human discretion. They also note that the oracle model they use for principle judgments shows near-zero arbitrariness, a self-confirmation effect they flag rather than count as evidence of ideal alignment.","pith_inferences":["If the paper's measurements are sound, the natural next product is a standard 'discretion card' for preference datasets and aligned models, reporting arbitrariness, conflict rate, and priority rankings alongside behavior-benchmark scores, so discretionary latitude becomes auditable instead of incidental.","A sensitivity test the authors did not run: replace the single principle-judging oracle with several independent judges, including human panels and different LLMs, on a random subsample; if arbitrariness and discrepancy rates shift with the oracle, part of what the paper labels 'discretion' is actually evaluator noise, and the field needs a disambiguation protocol.","The legal analogy suggests a mechanism the paper does not develop: appellate review. A practical implementation would record an annotator's principle-supremacy profile at annotation time and flag decisions that deviate from that annotator's own prior profile, much as courts check whether a decision departs from precedent.","One implicit consequence for pluralistic alignment is that if different communities genuinely rank principles differently, measured 'discrepancy' is not always a defect; the metric could double as a diagnostic for whose values a model is aligned to, turning a limitation into a governance tool."],"forward_implications":["Human preference datasets already encode implicit hierarchies of principles, so reusing or fine-tuning on a dataset means inheriting those latent priorities along with the preference labels.","Reward models can partly learn human discretion, staying within roughly 15–20% ranking discrepancy after fine-tuning, but translating that discretion into an RLHF-tuned policy fails, with discrepancies rising to roughly 40–71%; transferring discretion from a reward model to an LLM is an open problem.","Because roughly 80–85% of response pairs in both datasets are consensus or indifference, most of what customizes an aligned model happens in the 15–20% of conflicted cases where principles underdetermine the answer.","Off-the-shelf models do not mirror human annotators' principle priorities, with discrepancies between 16% and 53%, so using them as de facto arbiters of what is 'better' shifts alignment away from the humans whose preferences the datasets record.","Alignment frameworks that declare a set of principles without documenting how conflicts are resolved will keep producing systems whose behavior is shaped by unrecorded, unreviewed discretion."],"supporting_citations":[{"why":"Supplies the first human pairwise-preference dataset on which the arbitrariness and priority results are computed.","marker":"[5]"},{"why":"Defines the principle-based alignment-from-AI-feedback approach whose stochastic principle application motivates measuring discretion.","marker":"[6]"},{"why":"Provides the Bradley-Terry-Luce model that turns reward scores into preference probabilities in the RLHF formalism.","marker":"[13]"},{"why":"The Elo rating system inspires the logistic likelihood used to derive one-dimensional principle priorities from pairwise supremacy frequencies.","marker":"[35]"},{"why":"Supplies the 21 seed statements used as the principle set C in all experiments.","marker":"[48]"},{"why":"Supplies the second preference dataset with separate helpfulness and safety labels used to validate the metrics.","marker":"[54]"},{"why":"Defines generalized distances between rankings, the normalized Kendall tau used by the discretion discrepancy metric.","marker":"[62]"},{"why":"Establishes the RLHF-from-human-feedback pipeline whose annotator instructions leave 'which output is better' open, the process this paper audits.","marker":"[80]"}],"fun_headline_variants":["28.9% of AI safety labels are arbitrary discretion","AI alignment discretion: 28.9% arbitrary, 71.2% divergence","AI labels: 28.9% arbitrary; models diverge 71.2%","Human AI labelers: 28.9% arbitrary; models diverge 71.2%"],"cache_read_input_tokens":40064,"weakest_assumption_plain":"The paper's measurements all rest on a single zero-shot LLM as the oracle that decides, for each principle, which response adheres to it better; if that oracle's principle-wise judgments are biased or noisy, every downstream number—consensus rates, arbitrariness, supremacy, priorities, and discrepancies—inherits the error.","fun_headline_variants_meta":{"raw":{"variants":["28.9% of AI safety labels are arbitrary discretion","AI alignment discretion: 28.9% arbitrary, 71.2% divergence","AI labels: 28.9% arbitrary; models diverge 71.2%","Human AI labelers: 28.9% arbitrary; models diverge 71.2%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001028,"raw_usage":{"total_tokens":4400,"prompt_tokens":1084,"completion_tokens":3316,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":700,"completion_tokens_details":{"reasoning_tokens":3227}},"tokens_in":700,"tokens_out":3316,"duration_ms":21149,"temperature":1.0,"reasoning_tokens":3227,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T15:59:23.941527+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random subsample of the same two preference datasets, have at least two independent oracle systems or a panel of human judges label which response better adheres to each of the 21 principles, and recompute discretion arbitrariness and discretion discrepancy. If the human 28.9% arbitrariness rate or the priority rankings move by more than the reported bootstrap standard errors, the paper's diagnosis is an artifact of the chosen oracle; if they are stable, the existence of large, arbitrary discretion is confirmed.","supporting_citations":[{"cited_title":"Collective constitutional ai: Aligning a language model with public input","cited_arxiv_id":null,"evidence_quote":"Supplies the 21 seed statements used as the principle set C in all experiments."},{"cited_title":"Generalized distances between rankings","cited_arxiv_id":null,"evidence_quote":"Defines generalized distances between rankings, the normalized Kendall tau used by the discretion discrepancy metric."}],"review_version":1}