{"id":"6a3af76f-9a6d-49ab-a5d7-c2665e3d95b9","arxiv_id":"2508.12790","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":1,"one_line_summary":"Rubric-based rewards extend reinforcement learning to open-ended text generation, yielding a 30B model that outperforms a 671B model on humanities-style benchmarks.","lead":"The authors extend reinforcement learning from verifiable rewards to open-ended tasks by using 10,000+ rubric-based scoring criteria as training signals. A 30B-parameter model trained this way beats a 671B frontier model on several open-ended benchmarks, hinting that large-scale rubric rewards could replace costly human preference labels.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The load-bearing premise is that rubric rewards track genuine quality; the supplied text never demonstrates this, so the +5.2% gain could be rubric-scorer overlap with benchmarks rather than real improvement.","rationale":"I read the abstract in good faith: the paper extends RLVR to open-ended tasks via rubric rewards, reports large gains with few samples, and open-sources a model. That would be valuable if true. The single most load-bearing condition is the validity of the rubric reward. If the reward does not track human judgments, RL maximizes a proxy that can diverge from quality, especially with only 5K samples and 10K rubrics. The abstract gives no quantitative validation of this condition, and the corrupted full text prevents checking whether such validation exists. The reader's weakest assumption matches this concern (reward validity / benchmark overlap), so I agree. I am not raising a disagreement with consensus; I am asking for a standard RLHF-style check that is routine for reward-signal claims. Since the provided text cannot support a verdict either way, the reader's UNVERDICTED status remains unchanged; if the recovered full text lacks the human-preference check, the verdict should become CONDITIONAL pending that evidence.","tokens_in":21159,"tokens_out":3058,"duration_ms":32488,"concrete_test":"Recover the original PDF and run a blinded human-preference evaluation on a held-out sample from the open-ended benchmarks, comparing rubric-RL Qwen-30B-A3B, base Qwen-30B-A3B, and DeepSeek-V3; also compute the Spearman correlation between rubric scores and human quality ratings on responses not used in training. If pairwise human preference for the rubric-RL model over the base model is not significant, or if rubric-human Spearman is below ~0.4, the reward signal is not valid and the headline gain is likely an artifact of rubric/evaluator overlap.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Abstract: 'With only 5K+ samples, our system improves by +5.2% on open-ended benchmarks... outperforming a 671B DeepSeek-V3 model by +2.4%') requires that the rubric-based reward be a valid proxy for output quality. The abstract asserts this ('rubrics serve as structured, model-interpretable criteria for automatic scoring of subjective outputs') but does not report any evidence that rubric scores correlate with human preferences or with held-out quality judgments. Given only ~5K training samples, RL can exploit rubric-specific surface features (reward hacking) and inflate benchmark scores if the benchmark evaluators share stylistic or structural criteria with the rubrics. Open-ended benchmarks in humanities often use LLM-as-judge or rubric-like scoring, so the +5.2% and +2.4% margins may partly reflect reward-evaluator overlap rather than genuine quality gains. The supplied full text is corrupted, so I could not verify whether Sections 4-5 include a human study; the abstract itself only promises 'key lessons' and 'limitations and future releases.' The paper deserves credit for open-sourcing the model and the 10,000-rubric system, but those artifacts do not by themselves validate the reward signal.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes extending Reinforcement Learning from Verifiable Rewards (RLVR) to open-ended tasks by replacing verifiable outcome checkers with rubric-based rewards. The authors report constructing, to their knowledge, the largest rubric reward system to date, with over 10,000 rubrics sourced from humans, LLMs, or hybrid human-LLM collaboration. They train an open-sourced Qwen-30B-A3B model and claim that with only 5K+ samples it improves by +5.2% on open-ended benchmarks (especially humanities), outperforms a 671B DeepSeek-V3 model by +2.4%, and provides fine-grained stylistic control that reduces 'AI-like' tone. The supplied full text is corrupted and largely unreadable, so the review is necessarily based on the abstract and on the fragments that can be recovered; no experimental details, tables, or derivations are verifiable from the submitted material.","tokens_in":21582,"tokens_out":2363,"duration_ms":26972,"significance":"If the claims hold, this is a meaningful contribution: it broadens RLVR to subjective, open-ended domains; it introduces a large-scale rubric resource and an open-sourced model; and it proposes a concrete mechanism for stylistic control via 'rubric anchors.' The open release of the model and the 10,000-rubric system is a tangible strength. However, the central quantitative claims (+5.2% and +2.4%) are asserted in the abstract without statistical support, baseline details, or an evaluation protocol, and the submitted manuscript text does not permit independent verification. The circularity risk between rubric-based training rewards and rubric-like benchmark evaluation is real and must be addressed with concrete evidence, such as correlation with human judgments. The paper is promising but not yet verifiable in its current form.","major_comments":[{"comment":"The submitted manuscript text is corrupted and unreadable beyond the abstract; equations, tables, and any Sections 4-5 content cannot be inspected. This blocks verification of the training framework, the construction of the 10,000+ rubrics, the experimental setup, and the reported gains. Please provide a clean, decodable version before a full review can be completed.","section":"Full text"},{"comment":"The central claims '+5.2% on open-ended benchmarks' and '+2.4% over 671B DeepSeek-V3' are presented without error bars, number of evaluation instances, number of seeds, or significance tests. Given that the training set is only 5K+ samples, it is important to report variance across seeds and to specify which benchmarks constitute 'open-ended' and 'humanities.'","section":"Abstract"},{"comment":"The load-bearing premise is that rubric-based automatic scoring reflects genuine output quality. The abstract does not provide evidence that rubric scores correlate with human preferences or with held-out quality judgments. Please report a rubric-vs-human correlation study on a held-out set, and include an independent human evaluation of the trained model's outputs to rule out the alternative explanation that the gain is driven by overlap between the training reward and the benchmark evaluator.","section":"Abstract (reward validity)"},{"comment":"With only 5K+ training samples and open-ended benchmarks, data leakage is a serious concern. The abstract does not state whether the training prompts, rubric instances, or their evaluator prompts overlap with the benchmark test sets. Please clarify the separation between training data, rubric construction data, and evaluation data, and report a contamination check.","section":"Abstract (data contamination)"}],"minor_comments":[{"comment":"The phrase '5K+ samples' is ambiguous: specify whether this refers to training prompts, generated rollouts, or rubric instances.","section":"Abstract"},{"comment":"The statement 'rubrics serve as structured, model-interpretable criteria' would benefit from a precise definition of rubric structure and how the score is computed from the rubric components.","section":"Abstract"},{"comment":"The claim of 'more human-like, expressive responses' should be supported with a concrete evaluation rubric or an inter-annotator agreement measure, rather than relying on the authors' stylistic description.","section":"Abstract"},{"comment":"Please provide a breakdown of the 10,000+ rubrics by source (human, LLM, hybrid) and report quality-control measures used during rubric construction.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The main issue is that the submitted file is unreadable, so I cannot verify the methods or results. The authors should resubmit a clean, decodable manuscript. The circularity concern about rubric reward and benchmark evaluation is substantive and should be treated as a required experimental addition, not a discussion point. The open-sourced artifacts are a positive signal, but the paper is not ready for acceptance until the verification gap is closed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Abstract-level take: the claim is genuinely interesting, but the supplied full text is unreadable encoding garbage, so no one should treat this as a verified result. The reader's UNVERDICTED is exactly right.\n\nWhat's actually new: extending RLVR to open-ended tasks with rubric-based rewards at 10,000+ scale, with human/LLM/hybrid construction, and using rubrics as stylistic anchors to reduce \"AI-like\" tone. Those are concrete, actionable ideas, and open-sourcing the model and rubric system is real credit. If the +5.2% open-ended gain over base and +2.4% over DeepSeek-V3 with 5K samples hold up, this is a meaningful training axis for subjective domains.\n\nSoft spots: the load-bearing premise is that rubric rewards track genuine quality. The abstract asserts this but shows no correlation with human preferences or held-out judgments. With only 5K samples, RL can exploit rubric-specific surface cues, and if the benchmark evaluators use rubric-like LLM-judge scoring, part of the gain is reward-evaluator overlap rather than quality. That is a real circularity risk, not a manufactured one. There are also no error bars, baseline controls, or leakage analysis visible in the abstract. The bigger problem is that I cannot inspect Sections 4-5 at all due to the corruption; the paper may well include human studies, reward-hacking checks, and ablations. I can't tell.\n\nThe open-sourced artifacts don't by themselves validate the reward signal. But they do make the work worth engaging: if the artifact and rubrics are released, a referee can test the reward-quality correlation independently. The idea is plausible enough, and the stakes are high enough, that this deserves peer review rather than a desk reject. I would ask reviewers to focus on three things: correlation of rubric rewards with human preference, overlap between training rubrics and evaluation metrics, and whether the 2.4% margin over DeepSeek-V3 survives different judges and seeds. My own verdict remains unverdictable until the full text is readable.","headline":"The abstract makes a significant, plausible claim about rubric-based RLVR for open-ended tasks, but the supplied full text is corrupted, so this is an abstract-level note: send it to review, but only after the full text and reward-validity evidence are actually inspectable.","tokens_in":21991,"tokens_out":2322,"would_cite":false,"duration_ms":26229,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Rubric-based rewards extend reinforcement learning from verifiable tasks to open-ended writing and reasoning, and a 30B-parameter model trained on about 5,000 samples beats a 671B-parameter model on open-ended benchmarks.","keywords":["RLVR","rubric-based rewards","open-ended tasks","large language models","reward design","style control","human-LLM rubric construction"],"falsifier":"Measure whether the model's gains on rubric scores coincide with gains in human preference judgments on held-out open-ended responses, and train a control model with shuffled rubric rewards; the central claim fails if benchmark gains appear without human preference gains or appear with shuffled rewards.","tokens_in":21037,"feed_emoji":"📋","tokens_out":9705,"duration_ms":85846,"temperature":0.7,"pith_summary":"Reinforcement learning from verifiable rewards (RLVR) has worked for tasks with objectively checkable answers, such as code that passes tests or math with a right answer, but it has not transferred to open-ended writing and reasoning. The paper's central claim is that replacing verifiable signals with rubric-based automatic scoring, built at a scale of more than 10,000 rubrics, extends the same RL machinery to subjective tasks. On only 5,000-plus training samples, the resulting Qwen-30B-A3B model improves by 5.2% on open-ended benchmarks, especially humanities, and outperforms the 671B DeepSeek-V3 model by 2.4% while preserving general and reasoning abilities. Rubrics also act as anchors that steer style, producing less 'AI-like' and more human, expressive responses. If the claim holds, verifiability is no longer the bottleneck that confines RL training of language models.","feed_headline":"Rubric rewards let a 30B model beat a 671B model on open-ended tasks","feed_subtitle":"This extends reward-driven training from code and math to subjective writing while preserving reasoning skills.","key_machinery":"The load-bearing machinery is the rubric-based reward: a rubric is a set of structured criteria that a model can interpret, each marking out what a good subjective response contains, and the criteria are converted into an automatic score used as the RL reward. The paper ties this reward to 'rubric anchors', rubrics that steer not only quality but also style, so that training reinforces human-like phrasing rather than only factual correctness. The scale of over 10,000 rubrics, from human, LLM, and hybrid sources, is what turns a single scoring rubric into a general reward signal usable across diverse open-ended tasks.","core_discovery":"The paper claims that rubrics, defined as structured, model-interpretable criteria for scoring subjective outputs, can serve as the reward in reinforcement learning for open-ended tasks, removing the need for a mechanically verifiable answer. It reports building the largest such rubric reward system to date, with over 10,000 rubrics sourced from humans, LLMs, or a hybrid of both, and a training framework that makes rubric-based RL stable. The resulting open-sourced model, Qwen-30B-A3B, trained on just over five thousand samples, gains 5.2% on open-ended benchmarks, especially humanities, outperforms the much larger DeepSeek-V3 671B model by 2.4%, and preserves general and reasoning performance. The paper further claims that rubrics can be used as 'anchors' for fine-grained stylistic control, reducing the generic AI-like tone and producing more human-like, expressive responses.","pith_inferences":["Inference: the most direct test of the mechanism is whether rubric-score gains transfer to human preference; if they do not, the reported benchmark gains may partly reflect overlap between rubric criteria and benchmark evaluation metrics.","Inference: the rubric-anchor idea suggests a general recipe: define any measurable stylistic or quality dimension as a rubric and optimize it, which could enable controllable RL for voice, tone, formatting, and safety.","Inference: rubric quality and coverage likely matter more than raw count, so comparing the same model trained on a curated subset versus a random subset of rubrics would isolate the value of scale.","Inference: hybrid human-LLM rubric construction points toward scalable RLVR where an LLM proposes criteria and humans validate a sample, rather than hand-writing every reward rule."],"forward_implications":["Open-ended domains such as essay writing, summarization, and creative reasoning become trainable by the same RLVR recipe that worked for code and math.","A 30B-parameter model can outperform a 671B-parameter model on open-ended benchmarks, suggesting rubric-guided RL is a data- and compute-efficient route to strong subjective-task performance.","Because rubrics are interpretable, users can inspect and alter the criteria, making reward design debuggable rather than a black-box reward model.","Stylistic control can be achieved through the reward itself, so output tone is a training objective rather than a post-hoc prompt adjustment.","General and reasoning abilities are preserved, indicating that rubric rewards do not trade away core skills for benchmark gains."],"supporting_citations":[],"fun_headline_variants":["Rubric anchors let a 30B model beat a 671B on open-ended","Rewards from rubrics: 30B model tops 671B on subjective tasks","Open-ended RL with rubrics gives small model big win","Rubric-based RL lifts 30B model past 671B on humanities","Anchoring RL with rubrics: small model, large edge"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's central claim rests on the assumption that automated scoring with a rubric actually measures the quality human readers care about, rather than just matching the rubric's wording or the benchmark's answer key.","fun_headline_variants_meta":{"raw":{"variants":["Rubric anchors let a 30B model beat a 671B on open-ended","Rewards from rubrics: 30B model tops 671B on subjective tasks","Open-ended RL with rubrics gives small model big win","Rubric-based RL lifts 30B model past 671B on humanities","Anchoring RL with rubrics: small model, large edge"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000666,"raw_usage":{"total_tokens":3069,"prompt_tokens":1001,"completion_tokens":2068,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":617,"completion_tokens_details":{"reasoning_tokens":1969}},"tokens_in":617,"tokens_out":2068,"duration_ms":13988,"temperature":1.0,"reasoning_tokens":1969,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:17:37.533080+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure whether the model's gains on rubric scores coincide with gains in human preference judgments on held-out open-ended responses, and train a control model with shuffled rubric rewards; the central claim fails if benchmark gains appear without human preference gains or appear with shuffled rewards.","supporting_citations":[],"review_version":2}