REVIEW 3 major objections 5 minor 69 references
Language models can self-improve on open-ended tasks at test time by co-evolving their own rubrics, response archives, and policy—without labels or external judges.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-30 18:31 UTC pith:MULZLS7H
load-bearing objection Solid open-ended TTRL methods paper: real gap, strengthened baselines, large ID gains; self-judge alignment is the standing caveat, not a hidden collapse. the 3 major comments →
SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Under a fixed, label-free test-time budget with no external reward models or stronger judges, co-evolving maximally separated Good–Normal–Bad response archives, query-specific rubrics that retain only discriminating criteria, and actor parameters via probabilistic Pass/Fail rewards is enough to produce large in-domain gains on open-ended medical and research QA, plus out-of-distribution transfer and continued cross-benchmark improvement.
What carries the argument
The three-way SERPO loop: G-N-B archives store the most separated rollout triple per visit; rubric evolution keeps criteria with high response-score variance and G-N-B order agreement; probabilistic criterion scoring turns post-reasoning true/false token likelihoods into oriented, archive-calibrated rewards for GRPO policy updates that refresh the next rollouts.
Load-bearing premise
A frozen copy of the same deployed model, used only as rubric writer and Pass/Fail judge, must supply a stable enough quality signal that optimizing against it improves real open-ended quality rather than merely fitting the model’s own biases.
What would settle it
If, after SERPO adaptation, independent human or stronger-judge ratings on HealthBench and ResearchQA show no gain over the base model—or if gains reverse when the self-judge is replaced by a held-out human rubric—then the claim that self-evolved criteria yield genuine quality improvement would fail.
If this is right
- Open-ended domains without extractable answers become viable targets for label-free test-time RL, not only multiple-choice or symbolic tasks.
- Evolved policies can transfer to related held-out benchmarks and keep rising when the adaptation stream switches domains.
- Rubric evolution and policy evolution are complementary: fixed rubrics alone do not improve, and freezing the actor or archives erodes most of the gain.
- Probabilistic verdict scoring matters; hard Pass/Fail collapses advantages and cuts in-domain performance.
- Longer single-benchmark and sequential multi-benchmark runs can continue to extract signal beyond a standard 30-epoch budget.
Where Pith is reading between the lines
- If self-judge bias is the main risk, pairing SERPO with occasional cheap human vetoes on high-utility criteria could harden the loop without restoring full labeled RL.
- The same G-N-B-plus-rubric machinery might apply to other graded open-ended settings (code review comments, legal memos, tutoring) where majority answer voting is undefined.
- Continual cyclic benchmark streams with replay and rollback, as the authors sketch, would test whether test-time self-evolution can become a standing post-deployment habit rather than a one-shot adaptation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes SERPO, a fixed-set, label-free test-time RL method for open-ended generation. It replaces answer voting with a closed loop that co-evolves query-local Good–Normal–Bad (G-N-B) response archives, query-specific rubrics retained for archive discrimination, and shared actor parameters updated by GRPO. Criterion rewards come from frozen initial-weight copies of the deployed model via post-reasoning Pass/Fail token probabilities (Eqs. 3–4, 8), with utility dm = vm am (Eq. 7) driving retention and weighting. On Qwen3-4B and Qwen3.5-9B, SERPO improves HealthBench and ResearchQA by up to +20.63 and +20.31 over Base, raises the six-benchmark macro-average by up to +8.06, and reports OOD transfer plus continued HealthBench→ResearchQA evolution (Table 1; Figs. 3–4). Baselines include strengthened open-ended voting (response vote; claim consensus with G=16), static and evolving rubric-only TTS, and a privileged external-judge + official-rubric reference.
Significance. Open-ended TTRL without extractable answers is a genuine gap: majority-vote TTRL does not transfer cleanly when responses lack a canonical form, and most rubric co-evolution work assumes broader post-training budgets. SERPO’s contribution is a concrete, fully specified closed loop under a strict information budget (fixed prompts, self-rollouts, frozen self-judge/generator only). Strengths include strengthened open-ended voting baselines, complementary rubric-only vs. full-policy controls, multi-benchmark ID/OOD design, length analyses that separate verbosity from quality, sequential cross-benchmark evolution, and unusually complete implementation detail (Algorithm 1, hyperparameter tables, prompt templates). If the self-constructed rewards track externally graded quality rather than shared self-preference, the result is a useful recipe for label-free adaptation on long-form tasks.
major comments (3)
- [§4 Eqs. 3–8; Table 1–2; §6] The central claim—that maximizing rewards from a frozen same-model judge/generator improves true open-ended quality under a label-free budget—rests on an alignment premise that is only indirectly tested. §4 fixes evaluator roles and builds archives, utilities, and GRPO rewards entirely from that self-judge (Eqs. 3–8); Table 2 shows that training the judge/generator or freezing G-N-B archives hurts external scores, and Table 1 shows large GPT-5.1/official-rubric gains plus OOD transfer. That is supportive but not decisive: there is no human or independent-family audit that retained criteria (especially HealthBench negative safety items) are correct rather than correlated self-biases, and no criterion-level agreement analysis between SERPO rewards and the hidden official rubrics on the same responses. For a safety-sensitive medical benchmark this is load-bearing. Please add (i) a criterion
- [§4 Eq. (7); Algorithm 1; §B.3] Main-text utility and the appendix algorithm disagree on order-agreement. Eq. (7) defines am = [(Cm − Dm)/Pm]+ with Pm described as the comparison count, while Algorithm 1 (and §B.3) use am = max{0, (Cm − Dm)/(Cm + Dm + Tm)} with explicit ties Tm and margin τ. These are not equivalent when ties are common, and am directly controls elimination, admission, and reward weights wrew_m. Please reconcile the definition, state which was used in all reported runs, and note whether results are sensitive to the tie handling.
- [Table 1; Table 6; RQ2] OOD and transfer claims are directionally positive but uneven, and the paper sometimes over-aggregates them. In Table 1, several SERPO OOD lifts are small (e.g., Qwen3-4B RaR-Science +0.83; MedQA +0.92) relative to evaluation SD in Table 6, while ID gains are large; the privileged external-judge reference is stronger in-domain but weaker on all eight OOD cells—an interesting ID–OOD reversal that deserves a clearer causal discussion (self-evolved criteria vs. official-rubric overfitting) rather than a blanket “supports OOD transfer.” Please report significance or confidence intervals for small OOD deltas and qualify the transfer claim by magnitude and benchmark type (MCQ extraction vs. rubric-graded open-ended).
minor comments (5)
- [Figure 1; Abstract; §1] Figure 1 and several early paragraphs have missing spaces after periods/commas (e.g., “donotextend”, “closedloop”, “Good–Normal–Bad” rendering). A full copy-edit pass is needed.
- [§5 Baselines; Appendix C.1] Claim-consensus is an important baseline; consider moving the full reward equation and failure/fallback behavior from the appendix into a short main-text subsection so readers can assess fairness of G=16 vs. SERPO’s G=8 without leaving the main paper.
- [Figure 3; RQ5] The long-horizon linear projection to epoch ~99 reaching the privileged ID reference (Figure 3) is descriptive only; label it explicitly as non-predictive extrapolation so it is not read as a forecast.
- [§4 Eqs. 5 and 8] Notation: r_arc vs. final GRPO r, and wm,t vs. wrew_m, are clear in §4 but easy to confuse on first read; a one-row “score used for archives vs. score used for GRPO” callout would help.
- [§1; §2] Related work cites concurrent LLM-as-a-Verifier and several 2026 rubric-evolution papers; ensure camera-ready citations are stable and that the “first to combine post-reasoning Boolean verdict probabilities with evolving query-specific rubrics” claim remains accurate against the final concurrent set.
Circularity Check
No derivation circularity: self-referential TTRL rewards are the stated method; headline gains are measured by hidden external graders and official rubrics never used in adaptation.
full rationale
SERPO is an empirical methods paper, not a first-principles derivation. Its load-bearing claim is that a closed loop—G-N-B archives ordered by self-scores (Eqs. 5–6), criteria retained for archive discrimination (Eq. 7), and GRPO rewards from frozen same-model Pass/Fail token likelihoods (Eqs. 3–4, 8)—improves open-ended quality under a fixed label-free budget. That loop is intentionally self-referential by the TTRL problem statement (§1, §3): rewards are built only from the test prompts and the model’s own rollouts, with frozen initial-weight copies as rubric generator and judge. This is not a hidden tautology in which a reported “prediction” equals a fitted input by construction. Reported ID/OOD gains (Table 1; abstract) use GPT-5.1 and dataset-official rubrics that remain hidden during adaptation; the privileged external-judge + official-rubric reference is explicitly marked non-label-free. Ablations (Table 2) and OOD/cross-benchmark transfer further separate the claim from pure self-consistency. No step reduces a claimed external result to a self-definition, a fitted parameter renamed as prediction, a load-bearing uniqueness theorem from overlapping authors, or a renamed known law. Concerns that the self-judge may share biases with reporting graders are validity/alignment risks, not circularity of the derivation chain.
Axiom & Free-Parameter Ledger
free parameters (7)
- Elimination fraction ζ =
0.25
- Rollout group size G (SERPO) =
8 (baselines use 16)
- Rubric refresh interval F / archive width W =
F=3 visits, W=3
- Deletion patience (strikes) =
3
- Calibration minimum range δ / utility floor εu / tie margin τ =
δ=0.05, εu=0.01, τ=0.05
- Claim-consensus κ and response-vote Jaccard threshold =
κ=0.5, τ_resp=0.5
- GRPO/optimization knobs (lr, KL coeff, batch sizes, epochs) =
lr=1e-6, KL=0.001, 30 epochs default
axioms (6)
- domain assumption Fixed-set transductive TTRL may use only unlabeled test prompts, self-sampled responses, and frozen evaluation-blind auxiliaries—no references, human feedback, or external RMs during adaptation.
- domain assumption A frozen initial-weight copy of the actor is a sufficiently reliable rubric generator and criterion judge for constructing informative rewards.
- ad hoc to paper Criteria with high response-score variance and G-N-B order agreement (utility dm = vm am) are better reward features for true open-ended quality.
- domain assumption Softmax over Pass/Fail verdict-token logprobs yields a usable [0,1] criterion-satisfaction probability for GRPO.
- standard math GRPO group-relative advantages with KL regularization are an appropriate actor update given scalar rubric rewards.
- domain assumption Official benchmark rubrics scored by GPT-5.1 (and accuracy extraction on MC tasks) are adequate external measures of open-ended quality for claiming improvement.
invented entities (3)
-
Query-local Good–Normal–Bad (G-N-B) response archives with max-separation triple selection Dt
no independent evidence
-
Discrimination utility dm combining normalized archive variance and pairwise G-N-B order agreement
no independent evidence
-
Recoverable elimination region ΔR with strike counters for rubric pool maintenance
no independent evidence
read the original abstract
Test-time reinforcement learning (TTRL) enables language models to self-evolve at inference time without labeled feedback. Existing methods rely on answer voting and therefore do not extend naturally to open-ended generation, where valid responses cannot be mapped to a shared canonical answer. Without external reward models or stronger judges, adaptation must instead construct reliable rewards from the model's own outputs. We introduce SERPO (Self-Evolving Rubric Policy Optimization), which replaces answer voting with a closed loop that co-evolves response evidence, query-specific rubrics, and policy parameters. Good-Normal-Bad (G-N-B) response evolution organizes maximally separated rollouts into ordered archives; rubric evolution retains criteria that discriminate these archives; probabilistic criterion scoring converts verdict-token likelihoods into reward signals; and policy evolution optimizes the actor with the resulting signals. New actor rollouts then refresh both the archives and rubrics, closing the three-way evolution loop. Across two model configurations, two in-domain benchmarks, and four OOD benchmarks, SERPO improves HealthBench and ResearchQA by up to 20.63 and 20.31 points over the corresponding base models, raises the six-benchmark macro-average by up to 8.06 points, and supports OOD transfer and continued cross-benchmark evolution.
Figures
Reference graph
Works this paper leans on
-
[1]
2025 , eprint=
TTRL: Test-Time Reinforcement Learning , author=. 2025 , eprint=
2025
-
[2]
2026 , eprint=
Rubric-Guided Self-Distillation: Post-Training Without Rubric Verifiers , author=. 2026 , eprint=
2026
-
[3]
Beyond Majority Voting: Self-Reflective Test-Time Reinforcement Learning for
Wu, Sitong and Tan, Haoru and Zhang, Xichen and Xia, Bin and Zhang, Shaofeng and Qi, Xiaojuan and Yu, Bei and Jia, Jiaya , booktitle=. Beyond Majority Voting: Self-Reflective Test-Time Reinforcement Learning for. 2026 , note=
2026
-
[4]
2025 , eprint=
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains , author=. 2025 , eprint=
2025
-
[6]
What Disease Does This Patient Have? A Large-Scale Open Domain Question Answering Dataset from Medical Exams , author=. Applied Sciences , volume=. doi:10.3390/app11146421 , year=
-
[7]
2025 , pages=
Zhang, Ming and Shen, Yujiong and Li, Zelin and Sha, Huayu and Hu, Binze and Wang, Yuhui and Huang, Chenhao and Liu, Shichun and Tong, Jingqi and Jiang, Changhao and Chai, Mingxu and Xi, Zhiheng and Dou, Shihan and Gui, Tao and Zhang, Qi and Huang, Xuanjing , booktitle=. 2025 , pages=
2025
-
[8]
and Zhang, Hao and Gonzalez, Joseph E
Zheng, Lianmin and Chiang, Wei-Lin and Sheng, Ying and Zhuang, Siyuan and Wu, Zhanghao and Zhuang, Yonghao and Lin, Zi and Li, Zhuohan and Li, Dacheng and Xing, Eric P. and Zhang, Hao and Gonzalez, Joseph E. and Stoica, Ion , year=. Judging. 2306.05685 , archivePrefix=
-
[9]
2026 , eprint=
Detecting and Mitigating the Correct-Answer Extinction Window in Test-Time Reinforcement Learning with Majority Voting , author=. 2026 , eprint=
2026
-
[10]
2026 , eprint=
When More Sampling Hurts: The Modal Ceiling and Correlation Ceiling of Test-Time Scaling , author=. 2026 , eprint=
2026
-
[15]
2026 , eprint=
Reinforcing Chain-of-Thought Reasoning with Self-Evolving Rubrics , author=. 2026 , eprint=
2026
-
[17]
2026 , eprint=
Rubric-based On-policy Distillation , author=. 2026 , eprint=
2026
-
[20]
2025 , eprint=
Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models , author=. 2025 , eprint=
2025
-
[21]
and Wei, Jason and Hicks, Rebecca Soskin and Bowman, Preston and Qui
Arora, Rahul K. and Wei, Jason and Hicks, Rebecca Soskin and Bowman, Preston and Qui. 2025 , eprint=
2025
-
[23]
2026 , eprint=
Recursive Self-Evolving Agents via Held-Out Selection , author=. 2026 , eprint=
2026
-
[31]
2026 , eprint=
Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text , author=. 2026 , eprint=
2026
-
[33]
2023 , eprint=
Self-Consistency Improves Chain of Thought Reasoning in Language Models , author=. 2023 , eprint=
2023
-
[34]
2023 , eprint=
Tree of Thoughts: Deliberate Problem Solving with Large Language Models , author=. 2023 , eprint=
2023
-
[35]
2023 , eprint=
Self-Refine: Iterative Refinement with Self-Feedback , author=. 2023 , eprint=
2023
-
[36]
2023 , eprint=
Reflexion: Language Agents with Verbal Reinforcement Learning , author=. 2023 , eprint=
2023
-
[38]
2024 , eprint=
Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models , author=. 2024 , eprint=
2024
-
[42]
2017 , eprint=
Proximal Policy Optimization Algorithms , author=. 2017 , eprint=
2017
-
[43]
An Yang and Anfeng Li and Baosong Yang and Beichen Zhang and Binyuan Hui and Bo Zheng and Bowen Yu and Chang Gao and Chengen Huang and Chenxu Lv and Chujie Zheng and Dayiheng Liu and Fan Zhou and Fei Huang and Feng Hu and Hao Ge and Haoran Wei and Huan Lin and Jialong Tang and Jian Yang and Jianhong Tu and Jianwei Zhang and Jianxin Yang and Jiaxi Yang and...
-
[44]
2026 , doi=
Transactions of the Association for Computational Linguistics , volume=. 2026 , doi=
2026
-
[45]
Agarwal, S.; Zhang, Z.; Yuan, L.; Han, J.; and Peng, H. 2025. The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning. arXiv:2505.15134
Pith/arXiv arXiv 2025
-
[46]
Arora, R. K.; Wei, J.; Hicks, R. S.; Bowman, P.; Qui \ n onero-Candela, J.; Tsimpourlas, F.; Sharman, M.; Shah, M.; Vallone, A.; Beutel, A.; Heidecke, J.; and Singhal, K. 2025. HealthBench : Evaluating Large Language Models Towards Improved Human Health. arXiv:2505.08775
Pith/arXiv arXiv 2025
-
[47]
Bay, Y. Y.; and Yearick, K. A. 2026. When More Sampling Hurts: The Modal Ceiling and Correlation Ceiling of Test-Time Scaling. arXiv:2606.28661
Pith/arXiv arXiv 2026
-
[48]
Chen, X.; Li, G.; Wang, Z.; Jin, B.; Qian, C.; Wang, Y.; Wang, H.; Zhang, Y.; Zhang, D.; Zhang, T.; Tong, H.; and Ji, H. 2025. RM-R1 : Reward Modeling as Reasoning. arXiv:2505.02387
arXiv 2025
-
[49]
Ding, H.; Huang, B.; Fang, Y.; Liao, W.; Li, Z.; Zhang, J.; Wu, Z.; Zhao, J.; and Wang, Y. 2026. EvoRubrics : Dynamic Rubrics as Rewards via Adversarial Co-Evolution for LLM Reinforcement Learning. arXiv:2606.23038
Pith/arXiv arXiv 2026
-
[50]
Fang, J.; Hong, Z.; Zheng, M.; Song, M.; Li, G.; Jiang, H.; Zhang, D.; Guo, H.; Wang, X.; and Chua, T.-S. 2026. Rubric-based On-policy Distillation. arXiv:2605.07396
Pith/arXiv arXiv 2026
-
[51]
Guan, X.; Hu, X.; Huang, S.; Wang, Z.; Zhang, B.; Li, Z.; Xie, P.; Liu, B.; and Cao, J. 2026. EvoRubric : Self-Evolving Rubric-Driven RL for Open-Ended Generation. arXiv:2605.29847
Pith/arXiv arXiv 2026
-
[52]
Gunjal, A.; Wang, A.; Lau, E.; Nath, V.; He, Y.; Liu, B.; and Hendryx, S. 2025. Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains. arXiv:2507.17746
Pith/arXiv arXiv 2025
-
[53]
Huang, C.; Chou, S.-Y.; Zhang, Z.; and Cardie, C. 2026 a . Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text. arXiv:2604.20051
Pith/arXiv arXiv 2026
-
[54]
Huang, C.; Liu, H.; Zheng, T.; Dai, R.; Huang, L.; Li, J.; Li, Z.; Wei, Z.; Meng, Y.; and Huang, J. 2026 b . G-Zero : Self-Play for Open-Ended Generation from Zero Data. arXiv:2605.09959
Pith/arXiv arXiv 2026
-
[55]
Jin, D.; Pan, E.; Oufattole, N.; Weng, W.-H.; Fang, H.; and Szolovits, P. 2021. What Disease Does This Patient Have? A Large-Scale Open Domain Question Answering Dataset from Medical Exams. Applied Sciences, 11(14): 6421
2021
-
[56]
Y.; Shin, J.; Welleck, S.; Neubig, G.; Lee, M.; Lee, K.; and Seo, M
Kim, S.; Suk, J.; Longpre, S.; Lin, B. Y.; Shin, J.; Welleck, S.; Neubig, G.; Lee, M.; Lee, K.; and Seo, M. 2024. Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models. arXiv:2405.01535
Pith/arXiv arXiv 2024
-
[57]
Kwan, W.-C.; Gema, A. P.; Leang, J. O. J.; and Minervini, P. 2026. SCOPE : Self-Play via Co-Evolving Policies for Open-Ended Tasks. arXiv:2605.31433
Pith/arXiv arXiv 2026
-
[58]
Kwok, J.; Li, S.; Atreya, P.; Liu, Y.; Jiang, Y.; Finn, C.; Pavone, M.; Stoica, I.; and Mirhoseini, A. 2026. LLM-as-a-Verifier : A General-Purpose Verification Framework. arXiv:2607.05391
Pith/arXiv arXiv 2026
-
[59]
Li, S.; Zhao, J.; Wei, M.; Ren, H.; Zhou, Y.; Yang, J.; Liu, S.; Zhang, K.; and Chen, W. 2026 a . RubricHub : A Comprehensive and Highly Discriminative Rubric Dataset via Automated Coarse-to-Fine Generation. arXiv:2601.08430
arXiv 2026
-
[60]
S.; Xin, R.; Xiao, T.; Wang, Y.; Shao, R.; Hao, Z.; Sclar, M.; Oh, S.; Brahman, F.; Koh, P
Li, S. S.; Xin, R.; Xiao, T.; Wang, Y.; Shao, R.; Hao, Z.; Sclar, M.; Oh, S.; Brahman, F.; Koh, P. W.; and Tsvetkov, Y. 2026 b . EvoLM : Self-Evolving Language Models through Co-Evolved Discriminative Rubrics. arXiv:2605.03871
Pith/arXiv arXiv 2026
-
[61]
Yifei ; Chang, A.; Malaviya, C.; and Yatskar, M
Li S. Yifei ; Chang, A.; Malaviya, C.; and Yatskar, M. 2026. ResearchQA : Evaluating Scholarly Question Answering at Scale Across 75 Fields with Survey-Mined Questions and Rubrics. Transactions of the Association for Computational Linguistics, 14: 1344--1368
2026
-
[62]
Lin, H.; Kuai, Z.; Xue, E.; and Wang, L. 2026. Detecting and Mitigating the Correct-Answer Extinction Window in Test-Time Reinforcement Learning with Majority Voting. arXiv:2605.19444
Pith/arXiv arXiv 2026
-
[63]
Liu, M.; Shen, Y.; Xu, Z.; Cao, Y.; Cho, E.; Kumar, V.; Ghanadan, R.; and Huang, L. 2024. X-Eval : Generalizable Multi-aspect Text Evaluation via Augmented Instruction Tuning with Auxiliary Evaluation Aspects. arXiv:2311.08788
Pith/arXiv arXiv 2024
-
[64]
Liu, Y.; Iter, D.; Xu, Y.; Wang, S.; Xu, R.; and Zhu, C. 2023. G-Eval : NLG Evaluation using GPT-4 with Better Human Alignment. arXiv:2303.16634
Pith/arXiv arXiv 2023
-
[65]
Liu, Z.; Zhang, L.; Wang, X.; Xu, Z.; Zhan, S.; Shan, X.; Huang, W.; Dai, T.; Xia, S.-T.; Huo, C.; and Ding, L. 2026. ARBOR : Online Process Rewards via a Reusable Rubric Buffer for Search Agents. arXiv:2606.03239
Pith/arXiv arXiv 2026
-
[66]
P.; Hermann, K.; Welleck, S.; Yazdanbakhsh, A.; and Clark, P
Madaan, A.; Tandon, N.; Gupta, P.; Hallinan, S.; Gao, L.; Wiegreffe, S.; Alon, U.; Dziri, N.; Prabhumoye, S.; Yang, Y.; Gupta, S.; Majumder, B. P.; Hermann, K.; Welleck, S.; Yazdanbakhsh, A.; and Clark, P. 2023. Self-Refine: Iterative Refinement with Self-Feedback. arXiv:2303.17651
Pith/arXiv arXiv 2023
-
[67]
W.; Iyyer, M.; Zettlemoyer, L.; and Hajishirzi, H
Min, S.; Krishna, K.; Lyu, X.; Lewis, M.; Yih, W.-t.; Koh, P. W.; Iyyer, M.; Zettlemoyer, L.; and Hajishirzi, H. 2023. FActScore : Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation. arXiv:2305.14251
Pith/arXiv arXiv 2023
-
[68]
Nguyen, M.; Nguyen, Q.; and Vuong, P. 2026. Recursive Self-Evolving Agents via Held-Out Selection. arXiv:2606.28374
Pith/arXiv arXiv 2026
-
[69]
OpenAI . 2025. GPT-5.1 Instant and GPT-5.1 Thinking System Card Addendum. https://openai.com/index/gpt-5-system-card-addendum-gpt-5-1/. Accessed: 2026-07-29
2025
-
[70]
Qwen Team . 2025. Qwen3-4B-Instruct-2507 Model Card. https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507. Accessed: 2026-07-29
2025
-
[71]
Qwen Team . 2026 a . Qwen3.5-9B Model Card. https://huggingface.co/Qwen/Qwen3.5-9B. Accessed: 2026-07-29
2026
-
[72]
Qwen Team . 2026 b . Qwen3.6-27B Model Card. https://huggingface.co/Qwen/Qwen3.6-27B. Accessed: 2026-07-29
2026
-
[73]
Rein, D.; Hou, B. L.; Stickland, A. C.; Petty, J.; Pang, R. Y.; Dirani, J.; Michael, J.; and Bowman, S. R. 2023. GPQA : A Graduate-Level Google-Proof Q&A Benchmark. arXiv:2311.12022
Pith/arXiv arXiv 2023
-
[74]
Rezaei, M.; Mahmoud, A.; Wang, Z.; Tyagi, U.; Gosai, A.; Dumitru, R.-G.; Sabharwal, A.; Liu, B.; and He, Y. 2026. Rubric-Guided Self-Distillation: Post-Training Without Rubric Verifiers. arXiv:2606.12507
Pith/arXiv arXiv 2026
-
[75]
Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017. Proximal Policy Optimization Algorithms. arXiv:1707.06347
Pith/arXiv arXiv 2017
-
[76]
Z.; Ivison, H.; Kishore, V.; Zhuo, J.; Zhao, X.; Park, M.; Finlayson, S
Shao, R.; Asai, A.; Shen, S. Z.; Ivison, H.; Kishore, V.; Zhuo, J.; Zhao, X.; Park, M.; Finlayson, S. G.; Sontag, D.; Murray, T.; Min, S.; Dasigi, P.; Soldaini, L.; Brahman, F.; Yih, W.-t.; Wu, T.; Zettlemoyer, L.; Kim, Y.; Hajishirzi, H.; and Koh, P. W. 2025. DR Tulu : Reinforcement Learning with Evolving Rubrics for Deep Research. arXiv:2511.19399
Pith/arXiv arXiv 2025
-
[77]
Shao, Z.; Wang, P.; Zhu, Q.; Xu, R.; Song, J.; Bi, X.; Zhang, H.; Zhang, M.; Li, Y. K.; Wu, Y.; and Guo, D. 2024. DeepSeekMath : Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv:2402.03300
Pith/arXiv arXiv 2024
-
[78]
Sheng, L.; Ma, W.; Hong, R.; Wang, X.; Zhang, A.; and Chua, T.-S. 2026. Reinforcing Chain-of-Thought Reasoning with Self-Evolving Rubrics. arXiv:2602.10885
arXiv 2026
-
[79]
Shinn, N.; Cassano, F.; Berman, E.; Gopinath, A.; Narasimhan, K.; and Yao, S. 2023. Reflexion: Language Agents with Verbal Reinforcement Learning. arXiv:2303.11366
Pith/arXiv arXiv 2023
-
[80]
Wang, B.; Su, W.; Tian, H.; Kong, H.; Yang, T.; Yao, T.; Pan, Q.; Wu, Y.; Ai, Q.; Zhang, M.; and Liu, Y. 2026. Co-Evolving LLM Evaluators and Policies via DynamicRubric . arXiv:2607.20083
Pith/arXiv arXiv 2026
-
[81]
Wang, X.; Wei, J.; Schuurmans, D.; Le, Q.; Chi, E.; Narang, S.; Chowdhery, A.; and Zhou, D. 2023. Self-Consistency Improves Chain of Thought Reasoning in Language Models. arXiv:2203.11171
Pith/arXiv arXiv 2023
-
[82]
Whitehouse, C.; Wang, T.; Yu, P.; Li, X.; Weston, J.; Kulikov, I.; and Saha, S. 2025. J1 : Incentivizing Thinking in LLM -as-a-Judge via Reinforcement Learning. arXiv:2505.10320
arXiv 2025
-
[83]
Wu, S.; Tan, H.; Zhang, X.; Xia, B.; Zhang, S.; Qi, X.; Yu, B.; and Jia, J. 2026. Beyond Majority Voting: Self-Reflective Test-Time Reinforcement Learning for LLM Reasoning. In Proceedings of the 43rd International Conference on Machine Learning, volume 306 of Proceedings of Machine Learning Research. To appear
2026
-
[84]
Xie, W.; Zhao, H.; Liu, W.; Zhu, Y.; Chen, L.; Ye, M.; Chen, Z.; Xu, Y.; Dong, S.; Wang, Z.; Xu, X.; Shi, K.; Wu, R.; Zhang, X.; Shao, W.; Chang, B.; Duan, N.; and Wang, J. 2026. Step-wise Rubric Rewards for LLM Reasoning. arXiv:2605.17291
Pith/arXiv arXiv 2026
-
[85]
Yang, C.; Xiang, Z.; Tang, Y.; Teng, Z.; Huang, C.; Long, F.; Liu, Y.; and Su, J. 2026. TTCS : Test-Time Curriculum Synthesis for Self-Evolving. arXiv:2601.22628
arXiv 2026
-
[86]
L.; Cao, Y.; and Narasimhan, K
Yao, S.; Yu, D.; Zhao, J.; Shafran, I.; Griffiths, T. L.; Cao, Y.; and Narasimhan, K. 2023. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. arXiv:2305.10601
Pith/arXiv arXiv 2023
-
[87]
Ye, S.; Kim, D.; Kim, S.; Hwang, H.; Kim, S.; Jo, Y.; Thorne, J.; Kim, J.; and Seo, M. 2024. FLASK : Fine-grained Language Model Evaluation based on Alignment Skill Sets. arXiv:2307.10928
Pith/arXiv arXiv 2024
-
[88]
Zhang, M.; Shen, Y.; Li, Z.; Sha, H.; Hu, B.; Wang, Y.; Huang, C.; Liu, S.; Tong, J.; Jiang, C.; Chai, M.; Xi, Z.; Dou, S.; Gui, T.; Zhang, Q.; and Huang, X. 2025 a . LLMEval-Med : A Real-world Clinical Benchmark for Medical LLM s with Physician Validation. In Findings of the Association for Computational Linguistics: EMNLP 2025, 4888--4914. Association f...
2025
-
[89]
Zhang, Q.; Hu, C.; Upasani, S.; Ma, B.; Hong, F.; Kamanuru, V.; Rainton, J.; Wu, C.; Ji, M.; Li, H.; Thakker, U.; Zou, J.; and Olukotun, K. 2025 b . Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models. arXiv:2510.04618
Pith/arXiv arXiv 2025
-
[90]
Zuo, Y.; Zhang, K.; Sheng, L.; Qu, S.; Cui, G.; Zhu, X.; Li, H.; Zhang, Y.; Long, X.; Hua, E.; Qi, B.; Sun, Y.; Ma, Z.; Yuan, L.; Ding, N.; and Zhou, B. 2025. TTRL: Test-Time Reinforcement Learning. arXiv:2504.16084
Pith/arXiv arXiv 2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.