CoNL lets LLMs self-improve on non-verifiable tasks by rewarding critiques that produce better solutions in multi-agent conversations, jointly optimizing generation and judging without external feedback.
Meta-rewarding language models: Self-improving alignment with LLM-as-a-meta-judge
3 Pith papers cite this work, alongside 7 external citations. Polarity classification is still indexing.
years
2026 3verdicts
UNVERDICTED 3representative citing papers
Quality signals from LLM self-judgments and token entropy yield up to 18.6% gains over base model on saturated arithmetic tasks, outperforming SFT, but produce mixed or negative results on GSM8K depending on the signal used.
Proxy RL produces a staged proxy-internalization capability that emerges before and predicts reward hacking in coding environments.
citing papers explorer
-
Conversation for Non-verifiable Learning: Self-Evolving LLMs through Meta-Evaluation
CoNL lets LLMs self-improve on non-verifiable tasks by rewarding critiques that produce better solutions in multi-agent conversations, jointly optimizing generation and judging without external feedback.
-
Learning from Saturated Data: Signals Beyond Correctness for LLM Training
Quality signals from LLM self-judgments and token entropy yield up to 18.6% gains over base model on saturated arithmetic tasks, outperforming SFT, but produce mixed or negative results on GSM8K depending on the signal used.
-
Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization
Proxy RL produces a staged proxy-internalization capability that emerges before and predicts reward hacking in coding environments.