{"id":"cd9d1b30-b4bb-482a-8f02-a8ea0ca43ca5","arxiv_id":"2508.19922","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"HEAL evaluates preference optimization by measuring ranking accuracy and strength correlation between model likelihoods and proxy reward scores over multi-response hypothesis spaces.","lead":"This paper introduces HEAL, an evaluation framework that tests how well LLM alignment methods (DPO, SimPO, ORPO) learn preferences by comparing their response rankings against a proxy reward model's rankings. It also releases UniHypoBench, a benchmark with multiple responses per prompt, and reports that current methods capture proxy preferences only partially, with SimPO generalizing best.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim rests on same-proxy evaluation; Table 4 shows no measurable human-preference capture, so 'effectively capture preferences' is unsupported beyond ArmoRM-specific patterns.","rationale":"The reader's weakest_assumption correctly identifies the same-proxy training/evaluation loop as the most load-bearing vulnerability. The paper's own results support this concern: Table 4 shows that against human annotations, the optimized policies perform at or below chance and even the base model has negative PSC, so the claimed 'capture' does not extend to human preferences. Table 1's cross-proxy (GRM) results show no consistent gain, reinforcing that the learned signal is tied to ArmoRM. This does not make the framework useless—HEAL is still a valid diagnostic for proxy-specific preference fitting—but it does mean the abstract's unqualified 'effectively capture preferences' overstates the evidence. Since the reader already reached CONDITIONAL and this concern matches their weakest assumption, no verdict change is needed; the paper should be revised to scope the central claim explicitly to the training proxy and to report an independent gold-standard check.","tokens_in":16782,"tokens_out":3518,"duration_ms":42830,"concrete_test":"Re-run the HEAL evaluation of LLaMA-3-8B-Instruct + DPO/SimPO/ORPO (Table 1 models) on the same UniHypoBench/HelpSteer2/UltraFeedback hypothesis sets, but replace ArmoRM with an independent gold standard that did not generate the training labels—e.g., human preference judgments collected on the sampled response sets, or a different reward model such as a GPT-4-based annotator—and report RA/PSC with bootstrap 95% CIs and the base-model delta. If the optimized models' improvement over the base model is not significantly positive under this independent gold standard, the central claim is confined to ArmoRM-specific preference fitting and the abstract's 'effectively capture preferences' should be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central claim ('current preference learning methods can effectively capture preferences provided by proxy models while simultaneously suppressing negative samples') is evaluated with ArmoRM-Llama-3-8B-v0.1 as both the training-label source and the gold-standard evaluator (Sec. 4.1: 'ensuring consistency between training and evaluation preference distributions'). Under this design, a policy that memorizes ArmoRM's pairwise labels can score well on RA without learning generalizable preferences. The circularity is not merely formal: when the same trained policies are scored against the original human annotations in HelpSteer2 (Table 4), RA is near chance (46.61–48.18) and PSC is negative (-0.068 to -0.104); optimized models do not beat the base model on human judgments. The GRM 'different distribution' rows in Table 1 similarly show no consistent improvement over base on UniHypo/HelpSteer2, confirming the learned signal is proxy-specific. Therefore the headline conclusion should be read as 'policies can fit the training proxy's rankings,' not as evidence that preference learning captures preferences the field cares about. The paper's Limitations section acknowledges a restricted model set but does not flag this same-proxy evaluation as a boundary on the central claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes HEAL, an evaluation framework that recasts preference alignment as re-ranking in a hypothesis space and introduces two metrics: ranking accuracy (RA, based on Kendall's tau-b) and preference strength correlation (PSC, based on Pearson correlation). To support the framework, the authors construct UniHypoBench, a multi-response benchmark. They evaluate DPO, SimPO, and ORPO on several LLaMA models, using ArmoRM as both the training-label annotation source and the primary gold-standard evaluator, GRM as a second proxy, and HelpSteer2 human annotations as an additional gold standard. The paper's central claim is that current preference learning methods 'effectively capture preferences provided by proxy models while simultaneously suppressing negative samples'; secondary claims concern model-specific preference patterns and SimPO's stronger out-of-distribution generalization.","tokens_in":17103,"tokens_out":6210,"duration_ms":71918,"significance":"The framework is a useful conceptual contribution: evaluating alignment through ranked hypothesis spaces rather than single sampled responses is a genuine improvement in diagnostic methodology, and the proposed RA/PSC pair addresses an under-explored distinction between ordinal preference capture and continuous preference-strength calibration. The UniHypoBench resource, with 10+ responses per prompt, is potentially valuable for future work. The experimental effort is substantial, spanning multiple models, proxies, and ablations, and the authors provide code and data. If the central empirical claim were supported, the paper would be a solid contribution to preference-learning evaluation. However, the same-proxy evaluation design substantially limits the claim that can be drawn, and several claims are stronger than the presented evidence. The framework itself remains defensible and useful, but the empirical conclusions need either re-scoping or additional independent validation.","major_comments":[{"comment":"The primary evaluation is circular with respect to the central claim. Section 4.1 states that ArmoRM is used both to annotate the preference training data and as the gold-standard evaluator, 'ensuring consistency between training and evaluation preference distributions.' Since DPO, SimPO, and ORPO all directly optimize the trained policy's likelihoods toward ArmoRM-labeled preferred responses, high RA/PSC in the same-distribution rows of Table 1 can largely reflect fitting to the training labeler rather than 'effectively captur[ing] preferences.' Table 4 strengthens this concern: against the original HelpSteer2 human annotations, RA is near chance (44.79-48.18) and PSC is negative (-0.037 to -0.104); optimized models do not beat the SFT base. Table 1's GRM 'different preference distribution' rows also show no consistent improvement over the base model. The abstract and Section 4.2 should","section":"Section 4.1, 4.2; Tables 1 and 4"},{"comment":"No error bars, confidence intervals, or significance tests are reported for RA or PSC, which are dataset-level means over prompts. Many differences in Table 1 are under one percentage point (e.g., LLaMA-3.2-3B base vs +DPO on UniHypo: RA 54.64 vs 54.72; PSC 0.152 vs 0.154), and Appendix B's additional-backbone results show no movement at all for several models (e.g., Mistral-7B-it-v0.3 rows are nearly identical across methods). The claim that 'preference optimization effectively captures preference information' across the tested models is not statistically grounded. Please report bootstrap confidence intervals, paired tests over prompts, or otherwise justify that the observed differences are not noise.","section":"Section 4.2, Tables 1 and 7"},{"comment":"The text in Section 4.3 is contradicted by the table it cites. The text says 'the results reveal that performance shows no significant improvement even when evaluated on the model's own preference distribution' and 'we observe performance degradation in some cases, particularly for the SimPO-aligned model.' Table 2, however, shows SimPO improving RA from 55.93 to 61.50 and PSC from 0.164 to 0.301 (w/o length normalization), and from 54.12 to 59.70 and 0.100 to 0.252 (w/ length normalization). This inconsistency undermines the narrative in that subsection and needs to be corrected or the analysis re-described.","section":"Section 4.3, Table 2"},{"comment":"The proposed 'ranking accuracy' metric is defined as (tau_b + 1)/2. This mapping does not equal the proportion of concordant pairs when ties are present, because tau_b's denominator includes tie-correction terms. The statement in Definition 3 that the mapping 'is equal to assigning a zero-valued weight to the discordant pairs' is therefore inaccurate. Since RA is one of the two central metrics, the definition should be made mathematically precise, or the metric should be replaced with a true concordance proportion (or otherwise explicitly justified).","section":"Section 3.2, Eqs. (6)-(7)"}],"minor_comments":[{"comment":"The column header 'RS.' should presumably be 'RA.'; elsewhere the paper consistently uses RA for ranking accuracy.","section":"Table 4"},{"comment":"The notation for the hypothesis space is confusing: Yx is defined as a set but simultaneously assigned an ordering constraint. A ranked list or tuple notation (y1, y2, ...) with a separate definition of the ordering would be clearer.","section":"Eqs. (4)-(5)"},{"comment":"The text says 'We evaluated our approach using three models, including LLaMA-3.2-3B-Instruct and LLaMA-3-8B-Instruct,' but only two are named. Please either list all three or correct the count.","section":"Section 4.1"},{"comment":"The citation for Mistral in Appendix B points to Jiang et al. (2023), 'From CLIP to DINO,' which is not the Mistral model paper. Please correct the reference.","section":"Appendix B / References"},{"comment":"Several table cells in Table 1 run together (e.g., '54.640.152'), making the table hard to read. Consistent spacing would help.","section":"General"},{"comment":"The sentence 'our findings reveal that preference learning algorithms are notably adept at capturing most of these sub-dimensions' is difficult to reconcile with Table 3, where helpfulness RA is near chance and complexity/verbosity PSC are strongly negative. Please qualify this statement or revise it to match the data.","section":"Section 5"}],"recommendation":"major_revision","confidential_remarks":"The framework and benchmark are worth keeping, but the central empirical claim is overstated because of the same-proxy training/evaluation design. The main revision should either (a) add an independent gold-standard evaluation as primary evidence, or (b) explicitly rescope the claim to 'methods fit the training proxy's rankings' and carefully discuss what the human-annotation results imply for the field-level interpretation. Statistical support for the reported differences is also needed. With those changes, the paper could be suitable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe paper is worth your time, but the abstract's central claim is stronger than what the evidence supports. The framework itself is genuinely new and useful: HEAL evaluates alignment by comparing the policy's ranking of multiple responses against a gold-standard preference ordering, using Kendall's tau for ranking accuracy and Pearson correlation for preference strength. UniHypoBench, with 3k prompts and 8+ responses each from diverse LLMs, is a real contribution, and the finding that SimPO transfers better out-of-distribution than DPO/ORPO is interesting. The upset-plot analysis is a nice diagnostic tool.\n\nThe soft spot is load-bearing. The main evaluation uses ArmoRM as both the source of training labels and the gold-standard evaluator, which makes it unsurprising that policies partially reproduce ArmoRM's rankings. It does not establish that they \"effectively capture preferences\" in any general sense. The paper's own Table 4 shows that against the original HelpSteer2 human annotations, ranking accuracy is near chance (46–48%) and preference strength correlation is negative, and optimized models do not beat the base. The GRM different-distribution rows show no consistent gains. So the evidence says the learned signal is proxy-specific, not preference-general. The paper includes the human-annotation experiment, which is honest, but the discussion spins it as evidence that reward models can serve as human proxies — that is not supported.\n\nOther soft spots: no error bars or significance tests, and many reported differences are under 1%. The formal definition of the hypothesis space (an infinite ordered set) does not match the actual finite sampled spaces used in experiments. The Limitations section does not flag the same-proxy constraint, which is the most important boundary on the central claim.\n\nNone of this sinks the framework — HEAL can be used with an independent gold standard, and the benchmark is a useful resource. But the abstract and discussion need reining in, and the evaluation needs a held-out estimator and variance reporting. The thinking is clear, but the inference is overreached.\n\nI'd send this to peer review with a request for major revision, and I'd bring it to a reading group to discuss the evaluation design.\n\nBest.","headline":"Useful diagnostic framework, but the 'effectively capture preferences' claim collapses once the gold standard is the same proxy that wrote the training labels.","tokens_in":17573,"tokens_out":2890,"would_cite":true,"duration_ms":32093,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that DPO, SimPO, and ORPO learn the proxy reward model's ranking only partially—suppressing negative responses works, but fine-grained preference strength is missed, and only SimPO keeps the ranking on out-of-distribution p","keywords":["preference optimization","LLM alignment","hypothesis space","ranking accuracy","preference strength correlation","proxy reward model","DPO","SimPO"],"falsifier":"Score every UniHypoBench prompt's candidate responses with held-out human annotations (or a second reward model never used to build training pairs) and compute HEAL's ranking accuracy for DPO, SimPO, and ORPO against that gold standard. If ranking accuracy stays near the 50 percent chance level while ArmoRM-based accuracy is high, the paper's central claim would hold only for one proxy, not for preference capture generally; if SimPO's out-of-distribution advantage disappears under a human gold standard, the generalization result is an artifact of proxy similarity rather than true preference le","tokens_in":16712,"feed_emoji":"🎯","tokens_out":8294,"duration_ms":82221,"temperature":0.7,"pith_summary":"This paper argues that the usual way of evaluating preference-aligned LLMs—sampling one response and comparing it with a reference—misses how the model behaves across the full set of responses it could generate. It introduces HEAL, which treats alignment as a re-ranking problem: for each prompt, many candidate responses are ordered by the policy's generation likelihoods and by a gold-standard proxy reward model, and the two orderings are compared with ranking accuracy (Tau-b) and preference strength correlation (Pearson). Applying HEAL to DPO, SimPO, and ORPO, the paper concludes that these methods do capture proxy preferences, mostly by downweighting disliked responses, but the capture is incomplete: ranking accuracy stays modest, preference strength correlation is weak, learned preferences are specific to the proxy, and only SimPO generalizes to out-of-distribution data. The contribution is mainly diagnostic: a sampling-free, multi-candidate evaluation lens plus UniHypoBench, a benchmark with many candidate responses per prompt, for locating exactly where alignment methods fail.","feed_headline":"Preference tuning learns the proxy ranking only partway","feed_subtitle":"New multi-candidate evaluation shows DPO and ORPO lose the ranking out-of-distribution; SimPO holds.","key_machinery":"The ranked hypothesis space: for each prompt, the set of plausible responses is ordered by an indicator function, with generation likelihood for the policy and reward score for the gold-standard proxy. HEAL measures alignment as the agreement between these two orderings using ranking accuracy, a rescaling of Tau-b, and preference strength correlation, an expectation-based Pearson correlation between the two indicator values. UniHypoBench supplies the multi-candidate hypothesis spaces, with 2,985 prompts each containing at least eight responses generated by different LLMs, so the two metrics can be computed without sampling the aligned model.","core_discovery":"The paper's central claim is that current preference optimization methods effectively capture the preferences encoded by a proxy reward model while simultaneously suppressing negative samples. Using HEAL, the paper shows this capture is real but partial: DPO, SimPO, and ORPO raise ranking accuracy and preference strength correlation relative to the base model on in-distribution data, but ranking accuracy rarely passes 67%, preference strength correlation usually stays below 0.3, and performance drops markedly when the gold-standard proxy is changed. The exception is SimPO, which shows substantial generalization to out-of-distribution conditions. The paper interprets this pattern as evidence","pith_inferences":["Editorial inference: Because the same ArmoRM proxy scored the training pairs and defines the gold standard, HEAL's 'capture' numbers are at least partly a measure of fit to the training signal; re-running HEAL against independently collected human rankings on UniHypoBench would separate proxy-fitting from genuine preference learning.","Editorial inference: The weak preference-strength correlations point to a concrete failure mode in applications that threshold on likelihood or reward—such as best-of-n reranking—where ordinal correctness on coarse pairs can coexist with badly miscalibrated scores.","Editorial inference: UniHypoBench's many-candidate structure could be reused as a training objective, not just an evaluation set; directly maximizing ranking accuracy or strength correlation may yield alignment methods that overcome the partial-capture limit the paper documents.","Editorial inference: If personalized alignment is the goal, the framework's Appendix A.3 assumption that annotators or proxy models 'accurately reflect the target preferences' is the hard part, since HEAL measures alignment only relative to whatever preference distribution the reference encodes."],"forward_implications":["If HEAL's measurements are right, pairwise win-rate and RewardBench-style accuracy overstate alignment quality, because models can pass single-pair checks while failing on multi-candidate rankings and strength calibration.","SimPO's out-of-distribution robustness, if it holds, makes length-normalized reference-free objectives the most promising current direction for alignment that survives distribution shift.","Proxy-specific preference signatures imply that a single alignment score is not meaningful without specifying which preference distribution was used for training and evaluation.","The finding that methods suppress negatives without reaching a bimodal separation suggests future preference losses should target discriminative calibration, not just pairwise margins.","HEAL's deterministic, sampling-free procedure can serve as a low-cost complement or alternative to LLM-as-a-Judge for routine alignment checks."],"supporting_citations":[{"why":"Introduces DPO, the primary preference-optimization objective the paper trains and evaluates.","marker":"(Rafailov et al., 2024)"},{"why":"Introduces SimPO and supplies the pre-optimized LLaMA-3-8B weights used for evaluation.","marker":"(Meng et al., 2024)"},{"why":"Introduces ORPO, the monolithic reference-free preference method tested in the study.","marker":"(Hong et al., 2024)"},{"why":"ArmoRM-LLaMA-3-8B-v0.1, the proxy reward model that annotates preference data and serves as the gold standard for HEAL evaluation.","marker":"(Wang et al., 2024b)"},{"why":"RewardBench, the source of the instruction set expanded into UniHypoBench and the ranking-based evaluation idea HEAL builds on.","marker":"(Lambert et al., 2024)"},{"why":"UltraFeedback, the large preference dataset used for training and in-distribution evaluation.","marker":"(Cui et al., 2023)"},{"why":"HelpSteer2-Preference, the high-quality dataset with human preference dimensions used for out-of-distribution and multidimensional evaluation.","marker":"(Wang et al., 2024c)"},{"why":"Prior finding that preference learning algorithms do not learn preference rankings, which the paper's hypothesis-space results extend to a multi-candidate setting.","marker":"(Chen et al., 2024)"},{"why":"UpSet plots, the visualization technique used to analyze preference intersections across methods and proxies.","marker":"(Lex et al., 2014)"}],"fun_headline_variants":["Preference tuning hits ranking accuracy cap near 67%","SimPO wins out-of-distribution, DPO and ORPO fall","HEAL shows preference methods only partially learn proxy","Preference alignment: ranking stalls, SimPO generalizes","New eval finds preference tuning partial, SimPO robust"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The load-bearing premise is that the proxy reward model's scores are a trustworthy gold standard for alignment quality; the same proxy annotated the training pairs, and the paper's own comparison shows much weaker agreement with human annotations, so the central conclusion is about capturing that proxy's preferences rather than human preferences.","fun_headline_variants_meta":{"raw":{"variants":["Preference tuning hits ranking accuracy cap near 67%","SimPO wins out-of-distribution, DPO and ORPO fall","HEAL shows preference methods only partially learn proxy","Preference alignment: ranking stalls, SimPO generalizes","New eval finds preference tuning partial, SimPO robust"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000169,"raw_usage":{"total_tokens":1097,"prompt_tokens":734,"completion_tokens":363,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":478,"completion_tokens_details":{"reasoning_tokens":282}},"tokens_in":478,"tokens_out":363,"duration_ms":4787,"temperature":1.0,"reasoning_tokens":282,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T15:20:24.193876+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Score every UniHypoBench prompt's candidate responses with held-out human annotations (or a second reward model never used to build training pairs) and compute HEAL's ranking accuracy for DPO, SimPO, and ORPO against that gold standard. If ranking accuracy stays near the 50 percent chance level while ArmoRM-based accuracy is high, the paper's central claim would hold only for one proxy, not for preference capture generally; if SimPO's out-of-distribution advantage disappears under a human gold standard, the generalization result is an artifact of proxy similarity rather than true preference le","supporting_citations":[],"review_version":1}