{"id":"4e3b5b2e-14b5-4b3f-928a-56ab33c5d40c","arxiv_id":"2508.08509","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":3,"one_line_summary":"Few-shot comparative regression steers LLM choices to individual user preferences across fine-grained attributes, evaluated on two new benchmarks adapted from MIC and HelpSteer2.","lead":"This paper proposes a way to make chatbot answers adapt to each individual user's values: the model is shown a few examples of the user's preferences, plus a list of fine-grained attributes, and compares candidate answers against those values. It also introduces two new test sets, adapted from existing corpora, for measuring this kind of personalized alignment.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Benchmark adaptation may be circular if labels come from the same LLM/attribute machinery as the method; the central outperformance claim cannot be evaluated from the unreadable text.","rationale":"The concern is load-bearing because the paper's empirical claim rests entirely on benchmarks the authors designed; the abstract explicitly touts these benchmarks as evidence. Without access to the label-generation process, a claim of state-of-the-art performance is not independently testable. I agree partially with the reader's weakest_assumption: the reader lists both attribute sufficiency and benchmark independence; I focus on benchmark independence because it is more directly checkable and, if violated, undermines the central 'outperforming' claim. Given the unreadable full text, I do not propose changing the verdict; UNVERDICTED remains correct until the text is recovered. If the text is recovered and the labels are human-derived, my concern would be resolved. If labels are LLM-generated with the same attributes, the verdict should move to REJECT and the claims reduced.","tokens_in":21810,"tokens_out":3455,"duration_ms":38707,"concrete_test":"Obtain the source .tex or read the HTML version of the paper to locate the benchmark adaptation description (likely §4). Determine the label generation procedure: if MIC/HelpSteer2 adapted labels are created by LLM prompting that uses the same attribute list as the method, then run the proposed method and a simple attribute-free in-context preference baseline on (a) a random sample of the original HelpSteer2 human ratings and (b) a human-annotated MIC validation set. Compute pairwise preference accuracy. If the proposed method's margin over the baseline is smaller than the margin reported on the adapted benchmarks, the claimed superiority is an artifact of the adaptation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim is that the few-shot comparative regression approach 'outperforms' baselines on two proposed benchmarks. The abstract does not specify how the MIC and HelpSteer2 adaptations generate target labels. If those labels are produced by prompting an LLM (e.g., GPT-4) with the same fine-grained attribute taxonomy and comparative format that the method itself uses, then the method is being scored against predictions derived from the same inductive bias. This is a circularity risk: 'pluralistic alignment' would be measured by the LLM's own attribute decomposition rather than by independent human preferences. The full text is mojibake, so the adaptation protocol is unverifiable. This is not an accusation of misconduct; it is a concrete gap that the abstract leaves open. The benchmark validity is the load-bearing pillar of the outperformance claim; if it falls, the central claim has no independent evidence.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a training-free, few-shot comparative regression framework for pluralistic alignment. The idea is to represent user preferences as weighted judgments over a fine-grained attribute taxonomy, infer a user's preference function from a few in-context comparison demonstrations, and then rank response candidates accordingly. The authors also introduce two new benchmarks derived from the Moral Integrity Corpus (MIC) and HelpSteer2, and claim that their approach is interpretable, compatible with different attributes and LLMs, and outperforms multiple baselines and state-of-the-art methods. The abstract emphasizes the framework's ability to steer LLMs toward individual user preferences without retraining.","tokens_in":21921,"tokens_out":4317,"duration_ms":54749,"significance":"If the empirical claims hold, the contribution is potentially valuable: in-context, few-shot steering of LLM choices via an interpretable attribute-based comparison mechanism would be a practical step toward pluralistic alignment, and the two proposed benchmark adaptations could provide useful evaluation resources. The paper also promises cross-LLM compatibility and no training, which would be meaningful advances. However, the significance cannot currently be assessed. The provided full text is almost entirely unreadable due to character corruption, so the method, benchmark construction, experimental protocol, and results tables are not inspectable. The benchmark-validity concern raised by the reader is real and unresolved: if the adapted benchmark labels were produced using the same attribute taxonomy and comparison-prompt machinery as the method, the 'outperforming' result could be self-confirming rather than evidence of alignment to human preferences. The core idea deserves consideration, but this submission does not yet provide the evidence needed to evaluate it.","major_comments":[{"comment":"The central claim rests on two new benchmarks, but the adaptation protocol is not recoverable from this submission. If the target labels are generated by prompting an LLM with the same fine-grained attribute taxonomy and comparative prompt format used by the method, then evaluating the method against those labels measures agreement with the LLM's own inductive bias, not true user preferences. Please provide the full benchmark-construction pipeline, specify the source of labels (original human labels from MIC/HelpSteer2 vs. newly generated LLM labels), and show that the attribute taxonomy was not used to filter or relabel the evaluation data. Without independent human labels, the outperformance claim is potentially circular.","section":"Abstract / Benchmark construction (MIC- and HelpSteer2-derived)"},{"comment":"The provided manuscript text is corrupted; the results tables are unreadable and no metric definitions, baselines, number of runs, error bars, or significance tests are visible. The abstract's claim of 'outperforming multiple baseline and state-of-the-art methods' is therefore not supported by inspectable evidence. Please supply a readable version with: (i) exact evaluation metrics, (ii) comparison methods and their configurations, (iii) variance and statistical significance, and (iv) ablations over number of demonstrations and attribute set. This is a load-bearing requirement, not a stylistic point.","section":"Results tables (full text)"},{"comment":"The method's load-bearing assumption is that a hand-chosen set of fine-grained attributes spans the user-preference space relevant to alignment. The manuscript does not provide evidence for this: no analysis of attribute coverage, no sensitivity to attribute set, and no user study. Since the proposed benchmarks are built around the same attribute lens, a positive result could be an artifact of the taxonomy rather than evidence of steerable pluralism. Please add (i) a description of how the attribute set was chosen, (ii) an ablation replacing or removing attributes, and (iii) at least a small human-preference validation showing that attribute-level comparisons map onto actual user choices.","section":"Method / attribute taxonomy"}],"minor_comments":[{"comment":"The term 'comparative regression' is not defined. Please specify formally what is regressed, on what inputs, and how individual attribute ratings are aggregated into a final ranking.","section":"Abstract / Method"},{"comment":"The paper's contribution is explicitly 'few-shot'; please report the exact number of demonstrations used in the main experiments and whether performance is sensitive to this hyperparameter.","section":"Experiments"},{"comment":"The full-text rendering is garbled, making it impossible to check references, appendices, and limitations. If this is a submission-file encoding issue, a cleanly rendered version should be provided. A limitations section discussing the attribute-taxonomy assumption would also strengthen the paper.","section":"Full text"}],"recommendation":"major_revision","confidential_remarks":"The current electronic version is unreadable beyond the abstract and fragments, so I have based my assessment on the abstract and the reader's report. The circularity concern about benchmark adaptation is not resolved and is central to the claim. I recommend requiring a properly rendered, complete manuscript with a fully specified evaluation protocol and independent benchmark-label provenance before any further editorial decision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: I can't evaluate the body of this paper because the provided text is mostly mojibake, so anything beyond the abstract is guesswork. That said, the abstract describes something worth a referee's time even on its face: a method for steering a language model to an individual user's preferences using a few in-context comparisons over a fine-grained attribute set, with no retraining. The problem it targets—RLHF averaging over diverse preferences—is real and current.\n\nThe payoff if it works: interpretable, model-agnostic steerability, plus two new benchmarks adapted from MIC and HelpSteer2. The attribute list visible in the fragments is substantive (helpfulness, truthfulness, fairness, harmlessness, etc.), so 'pluralism' isn't just a label. The comparative-regression framing is a sensible way to get preference scores without scalar rewards. The novelty is real but not radical—few-shot in-context and comparative prompting are known; the contribution is the combination and the benchmarks.\n\nThe gaps, in order. The central 'outperforms multiple baselines' claim is uncheckable here: the results tables are unreadable and no evaluation protocol is visible. The circularity worry from the stress test is legitimate and unresolved: if the benchmark labels come from an LLM prompted with the same attribute taxonomy and comparative format the method uses, the gains could be self-confirming. The abstract does not say where the labels come from, and the adaptation protocol is invisible. I'd also want ablations and error bars for the free parameters (taxonomy, demonstration count, aggregation rule). The citation pattern is impossible to assess in this artifact. None of this is evidence of fraud; it's just what we can't see.\n\nWho should read it: anyone working on pluralistic alignment, reward modeling, or benchmark construction. Even in this fragmentary state, it would spark a good reading-group discussion about method design and evaluation hygiene.\n\nRecommendation: get a clean copy of the actual PDF. The idea is coherent and important enough that a serious editor should send it to review rather than desk-reject. If the body matches the abstract, it has a real shot; if the benchmark adaptation turns out to be circular, that's exactly what reviewers should catch.","headline":"Unreadable text, but the abstract describes a real problem and a plausible method; deserves a clean-copy referee.","tokens_in":22526,"tokens_out":7329,"would_cite":false,"duration_ms":69356,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that an LLM can be steered to individual user preferences at inference time with a few comparative examples grounded in fine-grained attributes, needing no retraining and beating prior alignment methods on two new benchmar","keywords":["pluralistic alignment","few-shot comparative regression","in-context learning","fine-grained attributes","value alignment","reward modeling","Moral Integrity Corpus","HelpSteer2"],"falsifier":"Collect held-out pairwise comparisons from real users on the same prompts used in the benchmarks and run the method with each user's demonstrations; if its predicted preferences are no more accurate than a baseline that ignores the demonstrations (or than a model using the average preference), the central claim is false. A second check: perturb the attribute taxonomy—remove or add one attribute—and measure whether the method's rankings change; if they do not change for attributes a user explicitly says matter, the attribute grounding is not doing the claimed work.","tokens_in":21573,"feed_emoji":"🎯","tokens_out":4826,"duration_ms":56147,"temperature":0.7,"pith_summary":"The paper tries to show that alignment can be made pluralistic—tailored to an individual rather than averaged over everyone—by doing the adaptation inside the prompt instead of in the weights. Its method, few-shot comparative regression, gives an LLM a list of fine-grained attributes and a few examples of how a particular user compares responses, then has the model compare candidate answers along those attributes and choose accordingly, with no training. To test this, the paper builds two new benchmarks by adapting the Moral Integrity Corpus and HelpSteer2, and reports that the method outperforms several baselines and prior methods while staying interpretable and model-agnostic. If the claim holds, personalized value alignment becomes a prompting problem rather than a retraining problem.","feed_headline":"No retraining: a few comparisons tailor an LLM to your values","feed_subtitle":"Attribute-based regression in the prompt beats average-preference alignment on two new pluralistic benchmarks, the paper reports.","key_machinery":"The central object is few-shot comparative regression: rather than score a single response with a scalar reward, the model receives a taxonomy of fine-grained attributes relevant to the task and a few demonstration comparisons that express a user's priorities, then performs pairwise comparisons of candidate responses along those attributes and aggregates them into a ranked choice. The attribute taxonomy is what carries individual preferences; the demonstrations are what condition the model on a particular user; no weights are updated. The same machinery is intended to work with different attribute sets and different base LLMs.","core_discovery":"Pluralistic alignment can be achieved without retraining: an LLM is shown a small set of comparison examples that encode one user's preferences, together with a fine-grained attribute taxonomy, and asked to score or rank candidate responses by comparing them along those attributes. The paper calls this few-shot comparative regression and argues that it captures individual preference functions that scalar-reward RLHF averages away. On two new benchmarks constructed from the Moral Integrity Corpus and HelpSteer2—one for value-aligned decision-making and one for reward modeling—the approach is reported to outperform multiple baseline and existing alignment methods while remaining interpretable","pith_inferences":["A natural extension not explored in the paper is to let the attribute taxonomy itself be generated or selected per user or per query, since the method's ceiling is set by whether the fixed taxonomy covers what actually drives a user's choices.","The benchmark labels are adapted from existing corpora, so a direct test against live human pairwise judgments—especially disagreements between users—would be the clearest check on whether the comparisons generalize beyond the demonstrations.","Because the method is training-free, it could also serve as a preference elicitation probe: the attributes that most change choices under different demonstrations reveal which dimensions dominate a user's utility, a diagnostic that scalar reward models obscure."],"forward_implications":["A deployed LLM can be redirected to a new user's values by supplying a handful of example comparisons, with no training run.","The same base model can serve divergent preference profiles by swapping the attribute list and demonstrations, making pluralistic behavior a deployment-time property.","Alignment choices become inspectable: the comparison scores over fine-grained attributes explain why one response was preferred over another.","The two proposed benchmarks give a common test bed for comparing steerable alignment methods on value-aligned decision-making and reward modeling.","If the reported gains hold, scalar-reward RLHF is not required for personalization; in-context preference regression can match or beat it on these tasks."],"supporting_citations":[],"fun_headline_variants":["A few comparisons steer LLMs to your values","Few-shot comparative regression personalizes LLM alignment","Align LLMs to your preferences without retraining","Pluralistic alignment via in-context comparisons"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The approach assumes that a hand-chosen list of fine-grained attributes plus a few example comparisons is enough to capture the particular preference function that drives a real user's choices; if the taxonomy misses what matters or the demonstrations do not generalize, the steerability claim collapses.","fun_headline_variants_meta":{"raw":{"variants":["A few comparisons steer LLMs to your values","Few-shot comparative regression personalizes LLM alignment","Align LLMs to your preferences without retraining","Pluralistic alignment via in-context comparisons"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000245,"raw_usage":{"total_tokens":1359,"prompt_tokens":716,"completion_tokens":643,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":460,"completion_tokens_details":{"reasoning_tokens":584}},"tokens_in":460,"tokens_out":643,"duration_ms":8023,"temperature":1.0,"reasoning_tokens":584,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T21:31:17.882521+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect held-out pairwise comparisons from real users on the same prompts used in the benchmarks and run the method with each user's demonstrations; if its predicted preferences are no more accurate than a baseline that ignores the demonstrations (or than a model using the average preference), the central claim is false. A second check: perturb the attribute taxonomy—remove or add one attribute—and measure whether the method's rankings change; if they do not change for attributes a user explicitly says matter, the attribute grounding is not doing the claimed work.","supporting_citations":[],"review_version":1}