Pith. sign in

REVIEW 3 major objections 5 minor 69 references

Language models can self-improve on open-ended tasks at test time by co-evolving their own rubrics, response archives, and policy—without labels or external judges.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-30 18:31 UTC pith:MULZLS7H

load-bearing objection Solid open-ended TTRL methods paper: real gap, strengthened baselines, large ID gains; self-judge alignment is the standing caveat, not a hidden collapse. the 3 major comments →

arxiv 2607.26873 v1 pith:MULZLS7H submitted 2026-07-29 cs.CL

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning

classification cs.CL
keywords test-time reinforcement learningopen-ended generationself-evolving rubricsG-N-B archivesprobabilistic criterion scoringGRPOlabel-free adaptationout-of-distribution transfer
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Standard test-time reinforcement learning rewards answers that match a majority vote, which works when answers can be reduced to a shared canonical form but fails for open-ended writing, where many valid responses differ in wording, completeness, and safety. SERPO instead builds rewards from the model’s own outputs by keeping, for each prompt, ordered Good–Normal–Bad response archives, a query-specific rubric of atomic criteria that best separate those archives, and a shared actor updated with group-relative policy optimization. Criterion scores come from the probability the frozen self-judge assigns to Pass versus Fail verdict tokens, not hard binary labels, so borderline cases still yield graded advantages. On two model sizes, this closed loop lifts HealthBench and ResearchQA by as much as about twenty points over the base models, improves a six-benchmark macro-average by up to eight points, transfers to held-out medical and science tasks, and keeps improving when the same actor is switched from one benchmark to another.

Core claim

Under a fixed, label-free test-time budget with no external reward models or stronger judges, co-evolving maximally separated Good–Normal–Bad response archives, query-specific rubrics that retain only discriminating criteria, and actor parameters via probabilistic Pass/Fail rewards is enough to produce large in-domain gains on open-ended medical and research QA, plus out-of-distribution transfer and continued cross-benchmark improvement.

What carries the argument

The three-way SERPO loop: G-N-B archives store the most separated rollout triple per visit; rubric evolution keeps criteria with high response-score variance and G-N-B order agreement; probabilistic criterion scoring turns post-reasoning true/false token likelihoods into oriented, archive-calibrated rewards for GRPO policy updates that refresh the next rollouts.

Load-bearing premise

A frozen copy of the same deployed model, used only as rubric writer and Pass/Fail judge, must supply a stable enough quality signal that optimizing against it improves real open-ended quality rather than merely fitting the model’s own biases.

What would settle it

If, after SERPO adaptation, independent human or stronger-judge ratings on HealthBench and ResearchQA show no gain over the base model—or if gains reverse when the self-judge is replaced by a held-out human rubric—then the claim that self-evolved criteria yield genuine quality improvement would fail.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Open-ended domains without extractable answers become viable targets for label-free test-time RL, not only multiple-choice or symbolic tasks.
  • Evolved policies can transfer to related held-out benchmarks and keep rising when the adaptation stream switches domains.
  • Rubric evolution and policy evolution are complementary: fixed rubrics alone do not improve, and freezing the actor or archives erodes most of the gain.
  • Probabilistic verdict scoring matters; hard Pass/Fail collapses advantages and cuts in-domain performance.
  • Longer single-benchmark and sequential multi-benchmark runs can continue to extract signal beyond a standard 30-epoch budget.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If self-judge bias is the main risk, pairing SERPO with occasional cheap human vetoes on high-utility criteria could harden the loop without restoring full labeled RL.
  • The same G-N-B-plus-rubric machinery might apply to other graded open-ended settings (code review comments, legal memos, tutoring) where majority answer voting is undefined.
  • Continual cyclic benchmark streams with replay and rollback, as the authors sketch, would test whether test-time self-evolution can become a standing post-deployment habit rather than a one-shot adaptation.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes SERPO, a fixed-set, label-free test-time RL method for open-ended generation. It replaces answer voting with a closed loop that co-evolves query-local Good–Normal–Bad (G-N-B) response archives, query-specific rubrics retained for archive discrimination, and shared actor parameters updated by GRPO. Criterion rewards come from frozen initial-weight copies of the deployed model via post-reasoning Pass/Fail token probabilities (Eqs. 3–4, 8), with utility dm = vm am (Eq. 7) driving retention and weighting. On Qwen3-4B and Qwen3.5-9B, SERPO improves HealthBench and ResearchQA by up to +20.63 and +20.31 over Base, raises the six-benchmark macro-average by up to +8.06, and reports OOD transfer plus continued HealthBench→ResearchQA evolution (Table 1; Figs. 3–4). Baselines include strengthened open-ended voting (response vote; claim consensus with G=16), static and evolving rubric-only TTS, and a privileged external-judge + official-rubric reference.

Significance. Open-ended TTRL without extractable answers is a genuine gap: majority-vote TTRL does not transfer cleanly when responses lack a canonical form, and most rubric co-evolution work assumes broader post-training budgets. SERPO’s contribution is a concrete, fully specified closed loop under a strict information budget (fixed prompts, self-rollouts, frozen self-judge/generator only). Strengths include strengthened open-ended voting baselines, complementary rubric-only vs. full-policy controls, multi-benchmark ID/OOD design, length analyses that separate verbosity from quality, sequential cross-benchmark evolution, and unusually complete implementation detail (Algorithm 1, hyperparameter tables, prompt templates). If the self-constructed rewards track externally graded quality rather than shared self-preference, the result is a useful recipe for label-free adaptation on long-form tasks.

major comments (3)
  1. [§4 Eqs. 3–8; Table 1–2; §6] The central claim—that maximizing rewards from a frozen same-model judge/generator improves true open-ended quality under a label-free budget—rests on an alignment premise that is only indirectly tested. §4 fixes evaluator roles and builds archives, utilities, and GRPO rewards entirely from that self-judge (Eqs. 3–8); Table 2 shows that training the judge/generator or freezing G-N-B archives hurts external scores, and Table 1 shows large GPT-5.1/official-rubric gains plus OOD transfer. That is supportive but not decisive: there is no human or independent-family audit that retained criteria (especially HealthBench negative safety items) are correct rather than correlated self-biases, and no criterion-level agreement analysis between SERPO rewards and the hidden official rubrics on the same responses. For a safety-sensitive medical benchmark this is load-bearing. Please add (i) a criterion
  2. [§4 Eq. (7); Algorithm 1; §B.3] Main-text utility and the appendix algorithm disagree on order-agreement. Eq. (7) defines am = [(Cm − Dm)/Pm]+ with Pm described as the comparison count, while Algorithm 1 (and §B.3) use am = max{0, (Cm − Dm)/(Cm + Dm + Tm)} with explicit ties Tm and margin τ. These are not equivalent when ties are common, and am directly controls elimination, admission, and reward weights wrew_m. Please reconcile the definition, state which was used in all reported runs, and note whether results are sensitive to the tie handling.
  3. [Table 1; Table 6; RQ2] OOD and transfer claims are directionally positive but uneven, and the paper sometimes over-aggregates them. In Table 1, several SERPO OOD lifts are small (e.g., Qwen3-4B RaR-Science +0.83; MedQA +0.92) relative to evaluation SD in Table 6, while ID gains are large; the privileged external-judge reference is stronger in-domain but weaker on all eight OOD cells—an interesting ID–OOD reversal that deserves a clearer causal discussion (self-evolved criteria vs. official-rubric overfitting) rather than a blanket “supports OOD transfer.” Please report significance or confidence intervals for small OOD deltas and qualify the transfer claim by magnitude and benchmark type (MCQ extraction vs. rubric-graded open-ended).
minor comments (5)
  1. [Figure 1; Abstract; §1] Figure 1 and several early paragraphs have missing spaces after periods/commas (e.g., “donotextend”, “closedloop”, “Good–Normal–Bad” rendering). A full copy-edit pass is needed.
  2. [§5 Baselines; Appendix C.1] Claim-consensus is an important baseline; consider moving the full reward equation and failure/fallback behavior from the appendix into a short main-text subsection so readers can assess fairness of G=16 vs. SERPO’s G=8 without leaving the main paper.
  3. [Figure 3; RQ5] The long-horizon linear projection to epoch ~99 reaching the privileged ID reference (Figure 3) is descriptive only; label it explicitly as non-predictive extrapolation so it is not read as a forecast.
  4. [§4 Eqs. 5 and 8] Notation: r_arc vs. final GRPO r, and wm,t vs. wrew_m, are clear in §4 but easy to confuse on first read; a one-row “score used for archives vs. score used for GRPO” callout would help.
  5. [§1; §2] Related work cites concurrent LLM-as-a-Verifier and several 2026 rubric-evolution papers; ensure camera-ready citations are stable and that the “first to combine post-reasoning Boolean verdict probabilities with evolving query-specific rubrics” claim remains accurate against the final concurrent set.

Circularity Check

0 steps flagged

No derivation circularity: self-referential TTRL rewards are the stated method; headline gains are measured by hidden external graders and official rubrics never used in adaptation.

full rationale

SERPO is an empirical methods paper, not a first-principles derivation. Its load-bearing claim is that a closed loop—G-N-B archives ordered by self-scores (Eqs. 5–6), criteria retained for archive discrimination (Eq. 7), and GRPO rewards from frozen same-model Pass/Fail token likelihoods (Eqs. 3–4, 8)—improves open-ended quality under a fixed label-free budget. That loop is intentionally self-referential by the TTRL problem statement (§1, §3): rewards are built only from the test prompts and the model’s own rollouts, with frozen initial-weight copies as rubric generator and judge. This is not a hidden tautology in which a reported “prediction” equals a fitted input by construction. Reported ID/OOD gains (Table 1; abstract) use GPT-5.1 and dataset-official rubrics that remain hidden during adaptation; the privileged external-judge + official-rubric reference is explicitly marked non-label-free. Ablations (Table 2) and OOD/cross-benchmark transfer further separate the claim from pure self-consistency. No step reduces a claimed external result to a self-definition, a fitted parameter renamed as prediction, a load-bearing uniqueness theorem from overlapping authors, or a renamed known law. Concerns that the self-judge may share biases with reporting graders are validity/alignment risks, not circularity of the derivation chain.

Axiom & Free-Parameter Ledger

7 free parameters · 6 axioms · 3 invented entities

The central empirical claim rests on standard RL/LM tooling plus several domain assumptions about self-judging and open-ended quality, plus many hand-set hyperparameters that shape archives, elimination, and rewards. No new physical entities; invented constructs are algorithmic state (G-N-B archives, query-local rubric pools, discrimination utility).

free parameters (7)
  • Elimination fraction ζ = 0.25
    Fraction of low-utility criteria marked at-risk each refresh; directly controls rubric churn.
  • Rollout group size G (SERPO) = 8 (baselines use 16)
    Number of on-policy samples per prompt; affects separation triples and GRPO advantages.
  • Rubric refresh interval F / archive width W = F=3 visits, W=3
    How often criteria are proposed and how many G-N-B visits are retained; shapes evidence freshness.
  • Deletion patience (strikes) = 3
    Consecutive at-risk rounds before criterion deletion; stabilizes the pool.
  • Calibration minimum range δ / utility floor εu / tie margin τ = δ=0.05, εu=0.01, τ=0.05
    Hand-set thresholds for G-N-B score calibration, weight flooring, and pairwise order ties.
  • Claim-consensus κ and response-vote Jaccard threshold = κ=0.5, τ_resp=0.5
    Baseline pseudo-label thresholds chosen for open-ended voting comparisons.
  • GRPO/optimization knobs (lr, KL coeff, batch sizes, epochs) = lr=1e-6, KL=0.001, 30 epochs default
    Shared training recipe from Qwen3/GRPO practice; not fit per benchmark but load-bearing for reported gains.
axioms (6)
  • domain assumption Fixed-set transductive TTRL may use only unlabeled test prompts, self-sampled responses, and frozen evaluation-blind auxiliaries—no references, human feedback, or external RMs during adaptation.
    Defines the problem budget in §3; all method claims are relative to this information constraint.
  • domain assumption A frozen initial-weight copy of the actor is a sufficiently reliable rubric generator and criterion judge for constructing informative rewards.
    §4 keeps generator/judge fixed; ablations that train them degrade results, but external validity of self-judgment is assumed.
  • ad hoc to paper Criteria with high response-score variance and G-N-B order agreement (utility dm = vm am) are better reward features for true open-ended quality.
    Eq. (7) and recoverable elimination operationalize ‘useful criterion’; this is a design axiom, not derived from external theory.
  • domain assumption Softmax over Pass/Fail verdict-token logprobs yields a usable [0,1] criterion-satisfaction probability for GRPO.
    Eq. (3); standard LLM-as-judge probability interface, assumed calibrated enough after G-N-B orientation/calibration.
  • standard math GRPO group-relative advantages with KL regularization are an appropriate actor update given scalar rubric rewards.
    §3 cites Shao et al. 2024 / PPO-style updates; optimizer choice is standard, not proved optimal here.
  • domain assumption Official benchmark rubrics scored by GPT-5.1 (and accuracy extraction on MC tasks) are adequate external measures of open-ended quality for claiming improvement.
    §5 evaluation protocol; reporting graders are privileged and hidden during adaptation.
invented entities (3)
  • Query-local Good–Normal–Bad (G-N-B) response archives with max-separation triple selection Dt no independent evidence
    purpose: Provide ordered, contrastive evidence for rubric proposal, utility, and reward calibration without reference answers.
    Algorithmic state introduced in §4; not a physical entity. Independent evidence is only via ablations (freezing archives hurts).
  • Discrimination utility dm combining normalized archive variance and pairwise G-N-B order agreement no independent evidence
    purpose: Rank, weight, admit, and eliminate query-specific criteria.
    Paper-defined scalar; success is measured by downstream benchmark scores, not an external measurement of dm.
  • Recoverable elimination region ΔR with strike counters for rubric pool maintenance no independent evidence
    purpose: Stabilize evolving rubrics while removing persistently weak criteria.
    Engineering mechanism specific to SERPO’s loop; no external falsifiable handle beyond end-task metrics.

pith-pipeline@v1.2.0-daily-grok45 · 28128 in / 4346 out tokens · 125176 ms · 2026-07-30T18:31:06.002748+00:00 · methodology

0 comments
read the original abstract

Test-time reinforcement learning (TTRL) enables language models to self-evolve at inference time without labeled feedback. Existing methods rely on answer voting and therefore do not extend naturally to open-ended generation, where valid responses cannot be mapped to a shared canonical answer. Without external reward models or stronger judges, adaptation must instead construct reliable rewards from the model's own outputs. We introduce SERPO (Self-Evolving Rubric Policy Optimization), which replaces answer voting with a closed loop that co-evolves response evidence, query-specific rubrics, and policy parameters. Good-Normal-Bad (G-N-B) response evolution organizes maximally separated rollouts into ordered archives; rubric evolution retains criteria that discriminate these archives; probabilistic criterion scoring converts verdict-token likelihoods into reward signals; and policy evolution optimizes the actor with the resulting signals. New actor rollouts then refresh both the archives and rubrics, closing the three-way evolution loop. Across two model configurations, two in-domain benchmarks, and four OOD benchmarks, SERPO improves HealthBench and ResearchQA by up to 20.63 and 20.31 points over the corresponding base models, raises the six-benchmark macro-average by up to 8.06 points, and supports OOD transfer and continued cross-benchmark evolution.

Figures

Figures reproduced from arXiv: 2607.26873 by Hua Yang, Jianze Wang, Jinlong Chen, Kunwang Zheng, Qianglong Chen, Qilong Zhang, Ying Liu, Yu Cao.

Figure 1
Figure 1. Figure 1: Criterion-level evidence for open-ended TTRL. Claim consensus can retain frequent but incomplete advice, whereas SERPO forms G-N-B evidence using evolving query-specific criteria without reference answers. reasoning Boolean verdict probabilities with evolving, query￾specific rubrics for policy optimization; concurrent LLM￾as-a-Verifier independently studies a related fixed-criterion interface (Kwok et al. … view at source ↗
Figure 2
Figure 2. Figure 2: Overview of SERPO. Here q is the input query, while oi and ri are the i-th policy rollout and its scalar reward. Et, Rt, and θt denote the query-local G-N-B archive state, query-specific rubric state, and shared actor parameters at encounter t, respectively. Criterion utility combines response-score variance vm and G-N-B order agreement am; ∆R denotes the low-utility elimination region. Snowflakes mark fix… view at source ↗
Figure 3
Figure 3. Figure 3: Qwen3-4B evolution on HealthBench (HB). The [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figure 5
Figure 5. Figure 5: Response-length ablations on Qwen3-4B. Both panels share the same variant rows. Left: the three-benchmark mean [PITH_FULL_IMAGE:figures/full_fig_p015_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Response-length behavior of Qwen3-4B. Left: six-benchmark macro-average versus mean length over the five bench [PITH_FULL_IMAGE:figures/full_fig_p015_6.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

69 extracted references · 35 linked inside Pith

  1. [1]

    2025 , eprint=

    TTRL: Test-Time Reinforcement Learning , author=. 2025 , eprint=

  2. [2]

    2026 , eprint=

    Rubric-Guided Self-Distillation: Post-Training Without Rubric Verifiers , author=. 2026 , eprint=

  3. [3]

    Beyond Majority Voting: Self-Reflective Test-Time Reinforcement Learning for

    Wu, Sitong and Tan, Haoru and Zhang, Xichen and Xia, Bin and Zhang, Shaofeng and Qi, Xiaojuan and Yu, Bei and Jia, Jiaya , booktitle=. Beyond Majority Voting: Self-Reflective Test-Time Reinforcement Learning for. 2026 , note=

  4. [4]

    2025 , eprint=

    Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains , author=. 2025 , eprint=

  5. [6]

    Applied Sciences , volume=

    What Disease Does This Patient Have? A Large-Scale Open Domain Question Answering Dataset from Medical Exams , author=. Applied Sciences , volume=. doi:10.3390/app11146421 , year=

  6. [7]

    2025 , pages=

    Zhang, Ming and Shen, Yujiong and Li, Zelin and Sha, Huayu and Hu, Binze and Wang, Yuhui and Huang, Chenhao and Liu, Shichun and Tong, Jingqi and Jiang, Changhao and Chai, Mingxu and Xi, Zhiheng and Dou, Shihan and Gui, Tao and Zhang, Qi and Huang, Xuanjing , booktitle=. 2025 , pages=

  7. [8]

    and Zhang, Hao and Gonzalez, Joseph E

    Zheng, Lianmin and Chiang, Wei-Lin and Sheng, Ying and Zhuang, Siyuan and Wu, Zhanghao and Zhuang, Yonghao and Lin, Zi and Li, Zhuohan and Li, Dacheng and Xing, Eric P. and Zhang, Hao and Gonzalez, Joseph E. and Stoica, Ion , year=. Judging. 2306.05685 , archivePrefix=

  8. [9]

    2026 , eprint=

    Detecting and Mitigating the Correct-Answer Extinction Window in Test-Time Reinforcement Learning with Majority Voting , author=. 2026 , eprint=

  9. [10]

    2026 , eprint=

    When More Sampling Hurts: The Modal Ceiling and Correlation Ceiling of Test-Time Scaling , author=. 2026 , eprint=

  10. [15]

    2026 , eprint=

    Reinforcing Chain-of-Thought Reasoning with Self-Evolving Rubrics , author=. 2026 , eprint=

  11. [17]

    2026 , eprint=

    Rubric-based On-policy Distillation , author=. 2026 , eprint=

  12. [20]

    2025 , eprint=

    Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models , author=. 2025 , eprint=

  13. [21]

    and Wei, Jason and Hicks, Rebecca Soskin and Bowman, Preston and Qui

    Arora, Rahul K. and Wei, Jason and Hicks, Rebecca Soskin and Bowman, Preston and Qui. 2025 , eprint=

  14. [23]

    2026 , eprint=

    Recursive Self-Evolving Agents via Held-Out Selection , author=. 2026 , eprint=

  15. [31]

    2026 , eprint=

    Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text , author=. 2026 , eprint=

  16. [33]

    2023 , eprint=

    Self-Consistency Improves Chain of Thought Reasoning in Language Models , author=. 2023 , eprint=

  17. [34]

    2023 , eprint=

    Tree of Thoughts: Deliberate Problem Solving with Large Language Models , author=. 2023 , eprint=

  18. [35]

    2023 , eprint=

    Self-Refine: Iterative Refinement with Self-Feedback , author=. 2023 , eprint=

  19. [36]

    2023 , eprint=

    Reflexion: Language Agents with Verbal Reinforcement Learning , author=. 2023 , eprint=

  20. [38]

    2024 , eprint=

    Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models , author=. 2024 , eprint=

  21. [42]

    2017 , eprint=

    Proximal Policy Optimization Algorithms , author=. 2017 , eprint=

  22. [43]

    2505.09388 , archivePrefix=

    An Yang and Anfeng Li and Baosong Yang and Beichen Zhang and Binyuan Hui and Bo Zheng and Bowen Yu and Chang Gao and Chengen Huang and Chenxu Lv and Chujie Zheng and Dayiheng Liu and Fan Zhou and Fei Huang and Feng Hu and Hao Ge and Haoran Wei and Huan Lin and Jialong Tang and Jian Yang and Jianhong Tu and Jianwei Zhang and Jianxin Yang and Jiaxi Yang and...

  23. [44]

    2026 , doi=

    Transactions of the Association for Computational Linguistics , volume=. 2026 , doi=

  24. [45]

    Agarwal, S.; Zhang, Z.; Yuan, L.; Han, J.; and Peng, H. 2025. The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning. arXiv:2505.15134

  25. [46]

    K.; Wei, J.; Hicks, R

    Arora, R. K.; Wei, J.; Hicks, R. S.; Bowman, P.; Qui \ n onero-Candela, J.; Tsimpourlas, F.; Sharman, M.; Shah, M.; Vallone, A.; Beutel, A.; Heidecke, J.; and Singhal, K. 2025. HealthBench : Evaluating Large Language Models Towards Improved Human Health. arXiv:2505.08775

  26. [47]

    Y.; and Yearick, K

    Bay, Y. Y.; and Yearick, K. A. 2026. When More Sampling Hurts: The Modal Ceiling and Correlation Ceiling of Test-Time Scaling. arXiv:2606.28661

  27. [48]

    Chen, X.; Li, G.; Wang, Z.; Jin, B.; Qian, C.; Wang, Y.; Wang, H.; Zhang, Y.; Zhang, D.; Zhang, T.; Tong, H.; and Ji, H. 2025. RM-R1 : Reward Modeling as Reasoning. arXiv:2505.02387

  28. [49]

    Ding, H.; Huang, B.; Fang, Y.; Liao, W.; Li, Z.; Zhang, J.; Wu, Z.; Zhao, J.; and Wang, Y. 2026. EvoRubrics : Dynamic Rubrics as Rewards via Adversarial Co-Evolution for LLM Reinforcement Learning. arXiv:2606.23038

  29. [50]

    Fang, J.; Hong, Z.; Zheng, M.; Song, M.; Li, G.; Jiang, H.; Zhang, D.; Guo, H.; Wang, X.; and Chua, T.-S. 2026. Rubric-based On-policy Distillation. arXiv:2605.07396

  30. [51]

    Guan, X.; Hu, X.; Huang, S.; Wang, Z.; Zhang, B.; Li, Z.; Xie, P.; Liu, B.; and Cao, J. 2026. EvoRubric : Self-Evolving Rubric-Driven RL for Open-Ended Generation. arXiv:2605.29847

  31. [52]

    Gunjal, A.; Wang, A.; Lau, E.; Nath, V.; He, Y.; Liu, B.; and Hendryx, S. 2025. Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains. arXiv:2507.17746

  32. [53]

    Huang, C.; Chou, S.-Y.; Zhang, Z.; and Cardie, C. 2026 a . Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text. arXiv:2604.20051

  33. [54]

    Huang, C.; Liu, H.; Zheng, T.; Dai, R.; Huang, L.; Li, J.; Li, Z.; Wei, Z.; Meng, Y.; and Huang, J. 2026 b . G-Zero : Self-Play for Open-Ended Generation from Zero Data. arXiv:2605.09959

  34. [55]

    Jin, D.; Pan, E.; Oufattole, N.; Weng, W.-H.; Fang, H.; and Szolovits, P. 2021. What Disease Does This Patient Have? A Large-Scale Open Domain Question Answering Dataset from Medical Exams. Applied Sciences, 11(14): 6421

  35. [56]

    Y.; Shin, J.; Welleck, S.; Neubig, G.; Lee, M.; Lee, K.; and Seo, M

    Kim, S.; Suk, J.; Longpre, S.; Lin, B. Y.; Shin, J.; Welleck, S.; Neubig, G.; Lee, M.; Lee, K.; and Seo, M. 2024. Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models. arXiv:2405.01535

  36. [57]

    P.; Leang, J

    Kwan, W.-C.; Gema, A. P.; Leang, J. O. J.; and Minervini, P. 2026. SCOPE : Self-Play via Co-Evolving Policies for Open-Ended Tasks. arXiv:2605.31433

  37. [58]

    Kwok, J.; Li, S.; Atreya, P.; Liu, Y.; Jiang, Y.; Finn, C.; Pavone, M.; Stoica, I.; and Mirhoseini, A. 2026. LLM-as-a-Verifier : A General-Purpose Verification Framework. arXiv:2607.05391

  38. [59]

    Li, S.; Zhao, J.; Wei, M.; Ren, H.; Zhou, Y.; Yang, J.; Liu, S.; Zhang, K.; and Chen, W. 2026 a . RubricHub : A Comprehensive and Highly Discriminative Rubric Dataset via Automated Coarse-to-Fine Generation. arXiv:2601.08430

  39. [60]

    S.; Xin, R.; Xiao, T.; Wang, Y.; Shao, R.; Hao, Z.; Sclar, M.; Oh, S.; Brahman, F.; Koh, P

    Li, S. S.; Xin, R.; Xiao, T.; Wang, Y.; Shao, R.; Hao, Z.; Sclar, M.; Oh, S.; Brahman, F.; Koh, P. W.; and Tsvetkov, Y. 2026 b . EvoLM : Self-Evolving Language Models through Co-Evolved Discriminative Rubrics. arXiv:2605.03871

  40. [61]

    Yifei ; Chang, A.; Malaviya, C.; and Yatskar, M

    Li S. Yifei ; Chang, A.; Malaviya, C.; and Yatskar, M. 2026. ResearchQA : Evaluating Scholarly Question Answering at Scale Across 75 Fields with Survey-Mined Questions and Rubrics. Transactions of the Association for Computational Linguistics, 14: 1344--1368

  41. [62]

    Lin, H.; Kuai, Z.; Xue, E.; and Wang, L. 2026. Detecting and Mitigating the Correct-Answer Extinction Window in Test-Time Reinforcement Learning with Majority Voting. arXiv:2605.19444

  42. [63]

    Liu, M.; Shen, Y.; Xu, Z.; Cao, Y.; Cho, E.; Kumar, V.; Ghanadan, R.; and Huang, L. 2024. X-Eval : Generalizable Multi-aspect Text Evaluation via Augmented Instruction Tuning with Auxiliary Evaluation Aspects. arXiv:2311.08788

  43. [64]

    Liu, Y.; Iter, D.; Xu, Y.; Wang, S.; Xu, R.; and Zhu, C. 2023. G-Eval : NLG Evaluation using GPT-4 with Better Human Alignment. arXiv:2303.16634

  44. [65]

    Liu, Z.; Zhang, L.; Wang, X.; Xu, Z.; Zhan, S.; Shan, X.; Huang, W.; Dai, T.; Xia, S.-T.; Huo, C.; and Ding, L. 2026. ARBOR : Online Process Rewards via a Reusable Rubric Buffer for Search Agents. arXiv:2606.03239

  45. [66]

    P.; Hermann, K.; Welleck, S.; Yazdanbakhsh, A.; and Clark, P

    Madaan, A.; Tandon, N.; Gupta, P.; Hallinan, S.; Gao, L.; Wiegreffe, S.; Alon, U.; Dziri, N.; Prabhumoye, S.; Yang, Y.; Gupta, S.; Majumder, B. P.; Hermann, K.; Welleck, S.; Yazdanbakhsh, A.; and Clark, P. 2023. Self-Refine: Iterative Refinement with Self-Feedback. arXiv:2303.17651

  46. [67]

    W.; Iyyer, M.; Zettlemoyer, L.; and Hajishirzi, H

    Min, S.; Krishna, K.; Lyu, X.; Lewis, M.; Yih, W.-t.; Koh, P. W.; Iyyer, M.; Zettlemoyer, L.; and Hajishirzi, H. 2023. FActScore : Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation. arXiv:2305.14251

  47. [68]

    Nguyen, M.; Nguyen, Q.; and Vuong, P. 2026. Recursive Self-Evolving Agents via Held-Out Selection. arXiv:2606.28374

  48. [69]

    OpenAI . 2025. GPT-5.1 Instant and GPT-5.1 Thinking System Card Addendum. https://openai.com/index/gpt-5-system-card-addendum-gpt-5-1/. Accessed: 2026-07-29

  49. [70]

    Qwen Team . 2025. Qwen3-4B-Instruct-2507 Model Card. https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507. Accessed: 2026-07-29

  50. [71]

    Qwen Team . 2026 a . Qwen3.5-9B Model Card. https://huggingface.co/Qwen/Qwen3.5-9B. Accessed: 2026-07-29

  51. [72]

    Qwen Team . 2026 b . Qwen3.6-27B Model Card. https://huggingface.co/Qwen/Qwen3.6-27B. Accessed: 2026-07-29

  52. [73]

    L.; Stickland, A

    Rein, D.; Hou, B. L.; Stickland, A. C.; Petty, J.; Pang, R. Y.; Dirani, J.; Michael, J.; and Bowman, S. R. 2023. GPQA : A Graduate-Level Google-Proof Q&A Benchmark. arXiv:2311.12022

  53. [74]

    Rezaei, M.; Mahmoud, A.; Wang, Z.; Tyagi, U.; Gosai, A.; Dumitru, R.-G.; Sabharwal, A.; Liu, B.; and He, Y. 2026. Rubric-Guided Self-Distillation: Post-Training Without Rubric Verifiers. arXiv:2606.12507

  54. [75]

    Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017. Proximal Policy Optimization Algorithms. arXiv:1707.06347

  55. [76]

    Z.; Ivison, H.; Kishore, V.; Zhuo, J.; Zhao, X.; Park, M.; Finlayson, S

    Shao, R.; Asai, A.; Shen, S. Z.; Ivison, H.; Kishore, V.; Zhuo, J.; Zhao, X.; Park, M.; Finlayson, S. G.; Sontag, D.; Murray, T.; Min, S.; Dasigi, P.; Soldaini, L.; Brahman, F.; Yih, W.-t.; Wu, T.; Zettlemoyer, L.; Kim, Y.; Hajishirzi, H.; and Koh, P. W. 2025. DR Tulu : Reinforcement Learning with Evolving Rubrics for Deep Research. arXiv:2511.19399

  56. [77]

    K.; Wu, Y.; and Guo, D

    Shao, Z.; Wang, P.; Zhu, Q.; Xu, R.; Song, J.; Bi, X.; Zhang, H.; Zhang, M.; Li, Y. K.; Wu, Y.; and Guo, D. 2024. DeepSeekMath : Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv:2402.03300

  57. [78]

    Sheng, L.; Ma, W.; Hong, R.; Wang, X.; Zhang, A.; and Chua, T.-S. 2026. Reinforcing Chain-of-Thought Reasoning with Self-Evolving Rubrics. arXiv:2602.10885

  58. [79]

    Shinn, N.; Cassano, F.; Berman, E.; Gopinath, A.; Narasimhan, K.; and Yao, S. 2023. Reflexion: Language Agents with Verbal Reinforcement Learning. arXiv:2303.11366

  59. [80]

    Wang, B.; Su, W.; Tian, H.; Kong, H.; Yang, T.; Yao, T.; Pan, Q.; Wu, Y.; Ai, Q.; Zhang, M.; and Liu, Y. 2026. Co-Evolving LLM Evaluators and Policies via DynamicRubric . arXiv:2607.20083

  60. [81]

    Wang, X.; Wei, J.; Schuurmans, D.; Le, Q.; Chi, E.; Narang, S.; Chowdhery, A.; and Zhou, D. 2023. Self-Consistency Improves Chain of Thought Reasoning in Language Models. arXiv:2203.11171

  61. [82]

    Whitehouse, C.; Wang, T.; Yu, P.; Li, X.; Weston, J.; Kulikov, I.; and Saha, S. 2025. J1 : Incentivizing Thinking in LLM -as-a-Judge via Reinforcement Learning. arXiv:2505.10320

  62. [83]

    Wu, S.; Tan, H.; Zhang, X.; Xia, B.; Zhang, S.; Qi, X.; Yu, B.; and Jia, J. 2026. Beyond Majority Voting: Self-Reflective Test-Time Reinforcement Learning for LLM Reasoning. In Proceedings of the 43rd International Conference on Machine Learning, volume 306 of Proceedings of Machine Learning Research. To appear

  63. [84]

    Xie, W.; Zhao, H.; Liu, W.; Zhu, Y.; Chen, L.; Ye, M.; Chen, Z.; Xu, Y.; Dong, S.; Wang, Z.; Xu, X.; Shi, K.; Wu, R.; Zhang, X.; Shao, W.; Chang, B.; Duan, N.; and Wang, J. 2026. Step-wise Rubric Rewards for LLM Reasoning. arXiv:2605.17291

  64. [85]

    Yang, C.; Xiang, Z.; Tang, Y.; Teng, Z.; Huang, C.; Long, F.; Liu, Y.; and Su, J. 2026. TTCS : Test-Time Curriculum Synthesis for Self-Evolving. arXiv:2601.22628

  65. [86]

    L.; Cao, Y.; and Narasimhan, K

    Yao, S.; Yu, D.; Zhao, J.; Shafran, I.; Griffiths, T. L.; Cao, Y.; and Narasimhan, K. 2023. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. arXiv:2305.10601

  66. [87]

    Ye, S.; Kim, D.; Kim, S.; Hwang, H.; Kim, S.; Jo, Y.; Thorne, J.; Kim, J.; and Seo, M. 2024. FLASK : Fine-grained Language Model Evaluation based on Alignment Skill Sets. arXiv:2307.10928

  67. [88]

    Zhang, M.; Shen, Y.; Li, Z.; Sha, H.; Hu, B.; Wang, Y.; Huang, C.; Liu, S.; Tong, J.; Jiang, C.; Chai, M.; Xi, Z.; Dou, S.; Gui, T.; Zhang, Q.; and Huang, X. 2025 a . LLMEval-Med : A Real-world Clinical Benchmark for Medical LLM s with Physician Validation. In Findings of the Association for Computational Linguistics: EMNLP 2025, 4888--4914. Association f...

  68. [89]

    Zhang, Q.; Hu, C.; Upasani, S.; Ma, B.; Hong, F.; Kamanuru, V.; Rainton, J.; Wu, C.; Ji, M.; Li, H.; Thakker, U.; Zou, J.; and Olukotun, K. 2025 b . Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models. arXiv:2510.04618

  69. [90]

    Zuo, Y.; Zhang, K.; Sheng, L.; Qu, S.; Cui, G.; Zhu, X.; Li, H.; Zhang, Y.; Long, X.; Hua, E.; Qi, B.; Sun, Y.; Ma, Z.; Yuan, L.; Ding, N.; and Zhou, B. 2025. TTRL: Test-Time Reinforcement Learning. arXiv:2504.16084