{"id":"a2b27910-40dd-48bf-8585-84f69695c686","arxiv_id":"2607.05742","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"low","formal_verification":"none","parameter_count":2,"one_line_summary":"Evaluator-specific demonstrations with retrospective reasoning improve LLM simulation of individual preference judges by up to 9.9 points over a non-personalized base judge, while interface telemetry often degrades accuracy.","lead":"LLMs can better simulate a specific human evaluator's preference judgments when given that person's past labels plus their own reasoning traces. The work shows reasoning helps a lot, cheap click telemetry often hurts, and some people are systematically harder to simulate.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the reader's already-flagged post-hoc reasoning fidelity caveat.","rationale":"The paper's central empirical claim is modest and carefully scoped: evaluator-specific multi-facet demonstrations, especially retrospective reasoning, improve three-class simulation of individual judges over a zero-shot Base Judge, and the improvement is personalization rather than generic ICL benefit. The 4\times4\times4 design, disjoint demo/validation splits, cross-evaluator control, majority-class baselines, and deviation-item metrics (Appendix I) collectively support that claim without circularity. The only soft premise is the fidelity of post-hoc reasoning, which the authors already flag in Limitations and which the Reader correctly elevated. Because that premise is acknowledged and the reported gains remain directionally robust even under conservative averaging, no stronger load-bearing attack is warranted. Verdict stays CONDITIONAL (pending data release and broader populations); no adjustment needed.","tokens_in":27078,"tokens_out":509,"duration_ms":6087,"concrete_test":"Re-run the recommended configuration (Claude-3.5-Sonnet, 8-shot J+RR) after replacing Stage-2 think-alouds with a concurrent verbal protocol collected during Stage 1 on a held-out subset of ~200 items; if the accuracy lift over Base Judge falls below ~5 pp or loses significance under the same Wilcoxon test, the post-hoc fidelity concern lands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader's weakest_assumption correctly isolates the softest premise: Stage-2 retrospective think-alouds (cued by interaction replay) may partially rationalize rather than report the original decision process (Limitations; Ericsson & Simon). That caveat is already stated by the authors and does not undermine the empirical claim as written. The strongest claim—up to +9.9 pp over the same-model Base Judge under Claude-3.5 8-shot J+RR, with cross-evaluator control confirming personalization—is supported by the factorial results (Tables 1–2, Fig. 2), the matched-subset control (+0.028 / +0.044), and the deviation-item analysis (Appendix I). No hidden circularity, derivation failure, or unacknowledged confound appears load-bearing. Telemetry's negative effect and the modest absolute accuracy (still near per-evaluator majority-class) are reported honestly and do not reverse the directional claim.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes PERSONAJUDGE, an ICL framework that simulates an individual evaluator’s three-class preference judgment (Prefer A / Neutral / Prefer B) by conditioning an LLM on that evaluator’s prior categorical labels plus optional interface telemetry and retrospective reasoning traces. Using a 4×4×4 factorial design over 32 trained annotators and 4,200 HH-style judgments (helpfulness and harmlessness), the authors report that evaluator-specific demonstrations improve three-class accuracy over the same model’s zero-shot Base Judge by up to 9.9 pp (Claude-3.5-Sonnet, 8-shot J+RR on Harmlessness), that retrospective reasoning is the most useful complementary signal while event-level telemetry often hurts, and that simulation difficulty is systematic—predicted by neutral usage and divergence from consensus—with neutral usage a stable cross-task trait (r=0.728). Controls include a cross-evaluator demonstration control, oracle/majority baselines, and a deviation-item analysis showing modest but genuine individual capture.","tokens_in":27340,"tokens_out":801,"duration_ms":8555,"significance":"If the results hold, the work supplies a concrete, carefully controlled methodology for moving LLM-as-Judge pipelines from consensus simulation toward individual-aware evaluation. The multi-facet data collection protocol, two-round cascade, factorial design, cross-evaluator personalization control, and deviation-item analysis are reusable contributions for the field. The honest reporting of modest absolute accuracy (near per-evaluator majority-class), the negative telemetry effect, and the cost–benefit asymmetry between reasoning and telemetry are themselves useful methodological findings for scaling personalized assessment and for reward modeling under heterogeneous preferences. Strengths include transparent baselines, non-parametric significance testing with multiple-comparison correction, and explicit limitations on post-hoc reasoning fidelity.","major_comments":[{"comment":"§5.1.1–5.1.2 and Table 2: the headline “up to 9.9 pp” gain is configuration-specific (Claude-3.5, 8-shot J+RR on Harmlessness). After FDR correction over the 64 conditions (Appendix H.4), only 3 Harmlessness and 0 Helpfulness configurations remain significant; the recommended configuration is significant only as a planned comparison. The abstract and main claims should state more clearly that average gains are small (+1.4 / +2.8 pp) and that most of the 64 cells do not survive family-wise correction, so that readers do not over-generalize the peak number.","section":null},{"comment":"§5.1.3 and Appendix I: PERSONAJUDGE does not significantly exceed the per-evaluator majority-class baseline (∆ = −0.019, p=0.95 Harmlessness; +0.042, p=0.14 Helpfulness). The deviation-item analysis shows genuine individual capture (accuracy ~0.36 on items where consensus predictors score 0 by construction), but the absolute individual signal remains modest. The paper’s framing of “individual evaluator simulation” should more explicitly position the method as a complement to, rather than a replacement for, simple per-person predictors, and discuss what additional signal would be needed to clear that bar.","section":null},{"comment":"Limitations and §3.2.2 / Stage-2 protocol: the largest gains rest on retrospective think-alouds cued by interaction replay. The authors correctly note possible rationalization (Ericsson & Simon), but provide no quantitative check (e.g., inter-rater agreement on criteria extracted from traces, or correlation of trace content with Stage-1 dwell/revisit patterns). A short validation or sensitivity analysis would strengthen the claim that J+RR gains reflect decision criteria rather than post-hoc narrative.","section":null}],"minor_comments":[],"recommendation":"minor_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a solid empirical HCI/ML paper on individual-aware LLM judges. The new piece is not the idea of personalization, but the multi-facet demonstration setup (categorical labels + retrospective reasoning + event-level interface telemetry) run through a clean 4×4×4 design on 32 trained annotators and 4,200 HH judgments, with the right controls.\n\nWhat they do well: disjoint demo/validation splits, two-round cascade that keeps Neutral as a real class, position-bias controls, a cross-evaluator control that isolates personalization from “any demos help,” oracle/majority baselines, and deviation-item analysis showing they recover some genuine individual signal (not just the group answer). The headline numbers hold: Claude-3.5 8-shot J+RR reaches 0.581 vs 0.482 Base on Harmlessness (+9.9 pp); J+RR is consistently best; event-level telemetry often hurts; gains are modest and still sit near per-evaluator majority-class. Neutral-usage rate predicts difficulty (especially Helpfulness) and is the stable cross-task trait (r=0.728), while simulatability itself does not transfer. They report the cost asymmetry (~5× for reasoning) without overselling.\n\nSoft spots, in proportion: the largest gains rest on Stage-2 post-hoc think-alouds cued by replay; the authors flag the Ericsson & Simon rationalization risk themselves, and it is the real load-bearing caveat. Absolute accuracy remains limited; data/code are not public; n=32 trained annotators and pairwise HH only. None of that overturns the directional claims or the negative telemetry finding.\n\nMath and stats look appropriate (Friedman/Wilcoxon, Holm/FDR, paired evaluator-level tests). Citations cover LLM-as-judge, RLHF, process tracing, and personalization without obvious gaps or circular self-citation forcing the result.\n\nThis is for people building personalized judges, reward models under heterogeneous preferences, or evaluation-fairness audits. Worth a serious referee. I would engage with it and cite the signal-comparison and simulatability results.","headline":"Careful factorial study showing that evaluator-specific ICL (especially retrospective reasoning) can beat a same-model base judge by up to ~10 points, with honest negative telemetry results and systematic predictors of who is hard to simulate.","tokens_in":27933,"tokens_out":553,"would_cite":true,"duration_ms":7412,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Evaluator-specific reasoning traces let LLMs simulate individual preference judges better than consensus-only baselines.","keywords":["LLM-as-judge","individual preference simulation","in-context learning","retrospective reasoning","interface telemetry","evaluator variation","personalized evaluation","Helpful and Harmless"],"falsifier":"A delayed re-test of the same evaluators on held-out items, or a side-by-side comparison of concurrent versus retrospective verbal reports, that shows the reasoning traces fail to improve simulation once rationalization is controlled for.","tokens_in":27998,"feed_emoji":"🧑‍⚖️","tokens_out":580,"duration_ms":6718,"temperature":0.7,"pith_summary":"Most LLM-as-judge pipelines learn crowd consensus and erase how different people actually decide. PERSONAJUDGE instead feeds an LLM each evaluator's own past labels plus optional process data—retrospective reasoning and interface telemetry—and asks the model to simulate that person's next three-way judgment (prefer A, neutral, prefer B). Across 32 trained annotators and 4,200 judgments on helpfulness and harmlessness tasks, the best configuration improves accuracy by as much as 9.9 points over the same model with no personal demonstrations. Retrospective reasoning supplies the largest lift; raw interface telemetry often hurts. Simulation difficulty is systematic: people who use Neutral a lot or diverge from consensus are harder to mimic, and the Neutral-usage tendency itself is stable across tasks even when simulatability is not. The work shows that individual-aware evaluation is feasible, but only when the right kind of process data is collected.","feed_headline":"Reasoning traces beat telemetry for personalizing LLM judges","feed_subtitle":"Up to 9.9 points better than zero-shot; cheaper click logs often hurt simulation accuracy","key_machinery":"PERSONAJUDGE: a two-round in-context learning cascade that first predicts whether the target evaluator will express any preference, then (if needed) predicts its direction, using demonstrations that can include labels, interface telemetry, and retrospective reasoning traces.","core_discovery":"Conditioning an LLM on an evaluator's own multi-facet demonstrations—especially categorical judgments paired with retrospective reasoning—raises three-class simulation accuracy over a zero-shot Base Judge by up to 9.9 percentage points, and the gain is personalization rather than generic demonstration benefit.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Reasoning traces lift LLM evaluator simulation by up to 9.9 points","Reasoning demos beat telemetry for personalizing individual LLM judges","Click telemetry often hurts LLM judge personalization; reasoning helps","Evaluator reasoning traces outperform telemetry in preference simulation","Personal LLM judges gain most from reasoning traces not interface logs"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The post-hoc think-alouds collected after replaying each judgment are assumed to be faithful enough accounts of the original decision criteria rather than after-the-fact rationalizations.","fun_headline_variants_meta":{"raw":{"variants":["Reasoning traces lift LLM evaluator simulation by up to 9.9 points","Reasoning demos beat telemetry for personalizing individual LLM judges","Click telemetry often hurts LLM judge personalization; reasoning helps","Evaluator reasoning traces outperform telemetry in preference simulation","Personal LLM judges gain most from reasoning traces not interface logs"]},"model":"grok-4.5","effort":"low","cost_usd":0.004476,"raw_usage":{"total_tokens":1246,"prompt_tokens":750,"num_sources_used":0,"completion_tokens":85,"cost_in_usd_ticks":44760000,"prompt_tokens_details":{"text_tokens":750,"audio_tokens":0,"image_tokens":0,"cached_tokens":0},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":411,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":750,"tokens_out":85,"duration_ms":5881,"temperature":1.0,"reasoning_tokens":411,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T02:41:16.776202+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A delayed re-test of the same evaluators on held-out items, or a side-by-side comparison of concurrent versus retrospective verbal reports, that shows the reasoning traces fail to improve simulation once rationalization is controlled for.","supporting_citations":[],"review_version":1}