Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T12:01:42.290502Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 100 of 104 outbound references and 87 inbound Pith citation observations for arXiv:2401.10020.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T12:01:42.290502Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T00:42:59.909123Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
100 of 104 outbound references displayed
External citation measurements
9
pith, observed 2026-08-05T02:28:24.338817Z
Observation 31994b54-2ec2-418c-9940-0e5dbf79ce86 · outbound
Self-Rewarding Language Models Advances in Neural Information Processing Systems , volume=
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 725116c1-126f-4659-a96b-355d62b9efca · outbound
Self-Rewarding Language Models Think you have solved question answering?
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d5b18464-4550-4bde-aee1-8b86426bb5c2 · outbound
Self-Rewarding Language Models 2019 , journal =
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0af92e01-73e4-4365-a060-411220df14de · outbound
Self-Rewarding Language Models EMNLP , year=
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b1454ac6-0334-4701-9ff2-01acf8007344 · outbound
Self-Rewarding Language Models 9th International Conference on Learning Representations
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f1b61372-b3e4-48a4-87f8-14087334fbc9 · outbound
Self-Rewarding Language Models Thirty-Fourth AAAI Conference on Artificial Intelligence , year =
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation adfac2cf-8595-43c9-940a-ae239c7daced · outbound
Self-Rewarding Language Models ROUGE : A Package for Automatic Evaluation of Summaries
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2933b852-04f0-4fac-9f69-a952b05ad815 · outbound
Self-Rewarding Language Models 2023 , howpublished =
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 64f1b9d9-6415-40b2-9a78-b7c4d56d4a17 · outbound
Self-Rewarding Language Models Improving Neural Machine Translation Models with Monolingual Data
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c52f7301-bc29-450e-9902-5dda04c09567 · outbound
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4fa5bd23-995b-4df4-b2e0-cf9d79204e56 · outbound
Self-Rewarding Language Models QLoRA: Efficient Finetuning of Quantized LLMs
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 40ff08f2-550a-491f-98dd-61d8f18cd0ab · outbound
Self-Rewarding Language Models LongForm: Effective Instruction Tuning with Reverse Instructions
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7af784e0-5d8f-47ae-b031-6af3d7ad3cec · outbound
Self-Rewarding Language Models Advances in Neural Information Processing Systems , volume=
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9e057a06-22ec-426c-aa0f-14e7df7081ed · outbound
Self-Rewarding Language Models LIMA: Less Is More for Alignment
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a7fa6976-dfbf-4bc7-93c7-59c428293d37 · outbound
Self-Rewarding Language Models The False Promise of Imitating Proprietary LLMs
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 98fcc26b-8ee0-40ed-a112-783ac403c827 · outbound
Self-Rewarding Language Models WizardLM: Empowering large pre-trained language models to follow complex instructions
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation be8aa6ee-88d2-4e99-b978-321456987a9c · outbound
Self-Rewarding Language Models and Stoica, Ion and Xing, Eric P
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d4616484-7a1e-4372-9165-d2878530372c · outbound
Self-Rewarding Language Models Hashimoto , title =
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 19e8d7e5-02d2-4c8f-91b9-8ca1703ed691 · outbound
Self-Rewarding Language Models Instruction Tuning with GPT-4
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b9509c37-137a-4f8e-b156-24fccf42a9d2 · outbound
Self-Rewarding Language Models OPT-IML: Scaling Language Model Instruction Meta Learning through the Lens of Generalization
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation eafb547f-1599-40e1-9392-7deb52fd7233 · outbound
Self-Rewarding Language Models arXiv e-prints , pages=
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b9bed052-c2c1-4124-a23e-78a9478662e7 · outbound
Self-Rewarding Language Models Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 763f42ac-7f9f-4424-95d4-44f2f663cf0f · outbound
Self-Rewarding Language Models Finetuned Language Models Are Zero-Shot Learners
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6cd73ad8-b993-4883-87fd-704103e19a2b · outbound
Self-Rewarding Language Models Hashimoto , title =
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 647d196f-cc9c-42cd-969a-cc80e46d4829 · outbound
Self-Rewarding Language Models Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages=
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7ab2ef3d-d2a8-4202-8b2a-70a8251ae45d · outbound
Self-Rewarding Language Models Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a544a119-f541-422e-8442-33020c80b8f8 · outbound
Self-Rewarding Language Models 2023 , eprint=
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 06370a74-d99d-4f57-af55-7ad9b48d23a0 · outbound
Self-Rewarding Language Models 2024 , url=
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6b530596-6414-4a8c-af4a-04358233053c · outbound
Self-Rewarding Language Models Multitask Prompted Training Enables Zero-Shot Task Generalization
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0149cd98-8d07-4f7b-9395-33996d42f4a1 · outbound
Self-Rewarding Language Models Cross-Task Generalization via Natural Language Crowdsourcing Instructions
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c2dd0cec-a04b-46da-af65-812fed6f2564 · outbound
Self-Rewarding Language Models Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ee4707a8-990e-4426-b4aa-8027e2b0051c · outbound
Self-Rewarding Language Models Constitutional
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cd30226b-d5df-49ec-b9c9-d52ecdb4ddd9 · outbound
Self-Rewarding Language Models Self-critiquing models for assisting human evaluators
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 944ba9d4-8dff-42ff-bbf9-6744e12ce992 · outbound
Self-Rewarding Language Models Self-Refine: Iterative Refinement with Self-Feedback
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5d7793c5-945d-4384-81b3-b545bf8c11df · outbound
Self-Rewarding Language Models Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2e3b20d5-15ee-49e4-b670-2681968984bb · outbound
Self-Rewarding Language Models 2023 , month =
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation eb8e6915-5f59-40cf-a93f-5257cf003227 · outbound
Self-Rewarding Language Models Enhancing Chat Language Models by Scaling High-quality Instructional Conversations
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3b635495-6310-4e9d-8e82-9a7e81e62345 · outbound
Self-Rewarding Language Models The Curious Case of Neural Text Degeneration
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 55a05645-df9d-4464-a153-b3b80b675902 · outbound
Self-Rewarding Language Models arXiv e-prints , pages=
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a93c8510-cb55-44e1-8637-8e8266c155c7 · outbound
Self-Rewarding Language Models CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 72a95905-f221-497e-9e5c-f88af0db26c3 · outbound
Self-Rewarding Language Models How Far Can Camels Go? Exploring the State of Instruction Tuning on Open Resources
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation eb703f43-de70-429b-a021-c5c08bb8d600 · outbound
Self-Rewarding Language Models Proceedings of the AAAI conference on artificial intelligence , volume=
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 318c89fe-b485-45d8-ba2d-b0110e84726e · outbound
Self-Rewarding Language Models Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6312acd9-a68e-4fbf-9686-f9e530a27c64 · outbound
Self-Rewarding Language Models Measuring Massive Multitask Language Understanding
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0076f5d9-fd96-43dc-a896-db3d2a8ee47c · outbound
Self-Rewarding Language Models 2023 , url =
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 38d8faae-22d3-4009-b78d-500213db6a87 · outbound
Self-Rewarding Language Models Thirty-seventh Conference on Neural Information Processing Systems , year=
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 979de74b-3b4e-4f83-b644-3401c22f9297 · outbound
Self-Rewarding Language Models The Twelfth International Conference on Learning Representations , year=
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 267b1716-1260-4025-aa8d-b650af90d68e · outbound
Self-Rewarding Language Models Gonzalez and Ion Stoica , booktitle=
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 81db4109-a291-40aa-8c62-fd54d523033e · outbound
Self-Rewarding Language Models Unresolved cited work
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 113fe07b-8d74-443a-825f-c8d1106ad10a · outbound
Self-Rewarding Language Models Visualizing data using
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f3117206-b70b-4fb1-b066-26a5f495d2e6 · outbound
Self-Rewarding Language Models Unresolved cited work
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1a4f39e1-53b7-42e8-8ba2-8dbe710e112c · outbound
Self-Rewarding Language Models Advances in Neural Information Processing Systems , volume=
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 61554148-4f04-411a-b2b1-9c9c7c779555 · outbound
Self-Rewarding Language Models Unresolved cited work
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 67b346b2-bd1e-48e5-b27f-5e654ed2b8aa · outbound
Self-Rewarding Language Models 2023 , url=
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation af17cc94-be69-4df7-9715-dd52c6095034 · outbound
Self-Rewarding Language Models Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track , year=
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 60f6000e-afc4-4544-ae44-1ac1b3202b15 · outbound
Self-Rewarding Language Models LLaMA: Open and Efficient Foundation Language Models
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4dbcd9d2-7203-4380-825d-af05804473e1 · outbound
Self-Rewarding Language Models Proceedings of the 25th International Conference on Machine Learning , pages=
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6b0fc452-10b3-4156-97f4-06fe4158b913 · outbound
Self-Rewarding Language Models Machine learning , volume=
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1db4f041-1d64-4170-9258-3463d54c8fee · outbound
Self-Rewarding Language Models OpenAI blog , volume=
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2b56ee4e-a1e1-4214-8842-259d35c94304 · outbound
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7e47feb6-f96a-4fa4-acf1-15dd074e394f · outbound
Self-Rewarding Language Models The CRINGE loss: Learning what language not to model
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 37dafee1-c794-4784-8c85-d9ee68b506fc · outbound
Self-Rewarding Language Models Claude 2
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 279a7133-f1d4-4e09-8ff2-22512915be9f · outbound
Self-Rewarding Language Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3fce83ed-2ffd-4bba-b7a3-4605e7a517f3 · outbound
Self-Rewarding Language Models Constitutional AI: Harmlessness from AI Feedback
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3f0f706b-e2db-4877-b4d4-7115fe0dd19e · outbound
Self-Rewarding Language Models Benchmarking foundation models with language-model-as-an-examiner
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4c230694-83eb-4bec-a20b-dfb20451651b · outbound
Self-Rewarding Language Models Piqa: Reasoning about physical commonsense in natural language
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9c103a83-b4f7-4d84-8aea-f09699b7c539 · outbound
Self-Rewarding Language Models AlpaGasus : Training a better alpaca with fewer data
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7d47b342-7b40-4b9c-ade3-65f777a95b5c · outbound
Self-Rewarding Language Models Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 226a7d65-ffa5-4f6e-8a4a-42f636bdb605 · outbound
Self-Rewarding Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e46a2f2d-c1f5-45d5-816e-90e4bc1f6f57 · outbound
Self-Rewarding Language Models Training Verifiers to Solve Math Word Problems
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c595846b-de6b-4f54-897a-edc878437d05 · outbound
Self-Rewarding Language Models A unified architecture for natural language processing: Deep neural networks with multitask learning
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 650f1700-58a3-4bc1-8196-94d98ab8091e · outbound
Self-Rewarding Language Models AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 40c30f43-a60e-4da2-9bfb-2b749953fefb · outbound
Self-Rewarding Language Models The Devil Is in the Errors: Leveraging Large Language Models for Fine-grained Machine Translation Evaluation
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fac3085a-c512-4da3-bb82-1aeb948c1407 · outbound
Self-Rewarding Language Models Reinforced Self-Training (ReST) for Language Modeling
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f0c6dc93-5c2c-47b6-a2d7-f58b65dc8000 · outbound
Self-Rewarding Language Models Measuring massive multitask language understanding
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 255488b6-6486-47df-8b8e-fe4ecff2c815 · outbound
Self-Rewarding Language Models Unnatural Instructions: Tuning Language Models with (Almost) No Human Labor
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 027108b2-b368-4f88-889c-94b34120d6c1 · outbound
Self-Rewarding Language Models Prometheus: Inducing Fine-grained Evaluation Capability in Language Models
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5e117b42-fa35-4fe3-976c-c8c37299e9cc · outbound
Self-Rewarding Language Models OpenAssistant Conversations -- Democratizing Large Language Model Alignment
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 938c560b-6c88-41b8-ab7e-b24082894ade · outbound
Self-Rewarding Language Models Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1703f623-039c-4602-998e-9b96a625b743 · outbound
Self-Rewarding Language Models RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
Reference 103
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 76b75438-d0c9-4048-abf9-dcd6f057f7a3 · outbound
Self-Rewarding Language Models Self-alignment with instruction backtranslation
Reference 104
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 23e4108e-751c-4857-9eed-cc753150129c · outbound
Self-Rewarding Language Models Hashimoto
Reference 105
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1b866eb6-2d81-42e0-be46-b9554cd7985c · outbound
Self-Rewarding Language Models ROUGE : A package for automatic evaluation of summaries
Reference 106
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5142f443-50db-45a3-a6c3-f960cc0b8358 · outbound
Self-Rewarding Language Models Can a suit of armor conduct electricity? a new dataset for open book question answering
Reference 107
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5f8a03f5-922c-48a2-b2af-b1c2a79ac6ab · outbound
Self-Rewarding Language Models Training language models to follow instructions with human feedback
Reference 108
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6fc8b0e2-3af5-4882-b55f-9c25c10421cd · outbound
Self-Rewarding Language Models Automatically Correcting Large Language Models: Surveying the landscape of diverse self-correction strategies
Reference 109
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 597a14cc-3aa2-40bc-9927-a695fd3ebd26 · outbound
Self-Rewarding Language Models Language models are unsupervised multitask learners
Reference 110
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0bfc952c-61d4-4ffe-885e-17984c4e86b0 · outbound
Self-Rewarding Language Models Direct preference optimization: Your language model is secretly a reward model
Reference 111
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 13dbc281-ed12-4238-b341-4b8a4c0eb765 · outbound
Self-Rewarding Language Models Branch-Solve-Merge Improves Large Language Model Evaluation and Generation
Reference 112
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 088bd612-4aa5-4c4b-b45a-d8bc5861ae84 · outbound
Self-Rewarding Language Models SocialIQA: Commonsense Reasoning about Social Interactions
Reference 113
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 02e2bcbc-c53a-4f8f-b30f-f535d39c9082 · outbound
Self-Rewarding Language Models Proximal Policy Optimization Algorithms
Reference 114
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9a7d8f85-8216-48fd-9e80-24998d0c5188 · outbound
Self-Rewarding Language Models Learning to summarize with human feedback
Reference 115
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e7a2d0cb-808e-483c-8e33-de3d0441bf7b · outbound
Self-Rewarding Language Models Hashimoto
Reference 116
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f70294a5-f8e1-41c3-bdce-ad929aefd96a · outbound
Self-Rewarding Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 117
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dc17169d-6b90-404d-9479-58691c9bb72d · outbound
Self-Rewarding Language Models Visualizing data using t-SNE
Reference 118
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 542183c1-5c4d-405c-ae3d-19d037588e8f · outbound
Self-Rewarding Language Models Smith, Daniel Khashabi, and Hannaneh Hajishirzi
Reference 119
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation be79fbfa-d555-494a-8c4a-1fae5e1094ea · outbound
Self-Rewarding Language Models Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 120
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bdd04f55-6a75-4367-b275-0d6bb227348b · outbound
Self-Rewarding Language Models Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss
Reference 121
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 87ecf850-2d3c-4cb1-b334-e4ba3f24daf1 · outbound
Self-Rewarding Language Models RRHF : Rank responses to align language models with human feedback
Reference 122
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9d7f4113-d9d8-4e29-a26f-e2e1170e3243 · outbound
Self-Rewarding Language Models URL https:// doi.org/10.18653/v1/p19-1472
Reference 123
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0ddbc15f-8636-45db-8b9e-b53b88bcc368 · inbound
Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models Self-Rewarding Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b498cc2c-6bf4-4fc5-ad15-a9f91c662656 · inbound
KTO: Model Alignment as Prospect Theoretic Optimization Self-Rewarding Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 93501320-7e32-4ca5-885f-c7d18ed1bd81 · inbound
LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models Self-Rewarding Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a1388664-2fa5-4511-bca0-c965b7fa8165 · inbound
TextGrad: Automatic "Differentiation" via Text Self-Rewarding Language Models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 759d9ff3-ec21-4452-9a23-6c63f6773f81 · inbound
The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale Self-Rewarding Language Models
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d7d6c5fd-dc09-4622-953b-957c1f47e449 · inbound
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Self-Rewarding Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 43567dc9-29ee-4850-a561-374a6966c29b · inbound
Reference 199
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 985a03ac-cbb1-4116-aff7-6d47da53886b · inbound
LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Self-Rewarding Language Models
Reference 285
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 31843f9f-dd1b-4618-a10b-18ab76d89998 · inbound
Reinforcement Learning from Human Feedback Self-Rewarding Language Models
Reference 285
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 632c9860-33a5-4e8f-b341-5a12f7ac8a02 · inbound
A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence Self-Rewarding Language Models
Reference 195
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 596f078d-45fb-46bf-a7ad-bcc89012defb · inbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Self-Rewarding Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b58b8565-ac9d-418f-baf7-a7c69adfc8d5 · inbound
Self-Rewarding Vision-Language Model via Reasoning Decomposition Self-Rewarding Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6202109f-f96a-4bf0-a89b-444b0fda9ba0 · inbound
Improving Alignment in LVLMs with Debiased Self-Judgment Self-Rewarding Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f6d1485-5ea9-4e63-8b1f-def683ab6f0a · inbound
Igniting Creative Writing in Small Language Models: LLM-as-a-Judge versus Multi-Agent Refined Rewards Self-Rewarding Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46ff7d88-df70-4e82-ba8b-5a1480e5df3a · inbound
Improving Large Vision and Language Models by Learning from a Panel of Peers Self-Rewarding Language Models
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 457351ee-529c-446f-86c1-9bac5eb10b8b · inbound
Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training Self-Rewarding Language Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5f0889fe-3ffc-4605-8f3c-129dde1d2f9e · inbound
Enhancing Speech Large Language Models through Reinforced Behavior Alignment Self-Rewarding Language Models
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9072b3cc-72b5-4a2d-8b39-6bfc30a446b7 · inbound
On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization Self-Rewarding Language Models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0c669d5e-9341-4ae3-9449-1814176dc4a3 · inbound
Breaking the Self-Confirming Loop: Diagnosing and Mitigating Systemic Reward Bias in Self-Rewarding RL Self-Rewarding Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b66f49f-04bf-4706-aedb-d3cab0e19629 · inbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning Self-Rewarding Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23a6939e-eb50-4155-bcc8-347fa703c964 · inbound
CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning Self-Rewarding Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f426e85-277b-48ca-a7f0-571f93d4a738 · inbound
A Survey of Reinforcement Learning For Economics Self-Rewarding Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3533ebd3-39cb-403a-aeed-c2988b2bfc98 · inbound
Toward Epistemic Stability: Engineering Consistent Procedures for Industrial LLM Hallucination Reduction Self-Rewarding Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation eba0c7ab-2891-47b2-8a99-792297af4b3e · inbound
Visual-ERM: Reward Modeling for Visual Equivalence Self-Rewarding Language Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9fc9fe92-33da-447d-bf00-19bbdd65bba2 · inbound
AdaRubric: Task-Adaptive Rubrics for Reliable LLM Agent Evaluation and Reward Learning Self-Rewarding Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c7b12d90-1405-4881-b3fe-0620883b79b4 · inbound
Can LLMs Learn to Reason Robustly under Noisy Supervision? Self-Rewarding Language Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0b86fdbb-a4f8-40b7-b35f-7d04808a240c · inbound
Pioneer Agent: Continual Improvement of Small Language Models in Production Self-Rewarding Language Models
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b8982aa8-85e4-4c2b-aaa4-469763691a53 · inbound
Utilizing and Calibrating Hindsight Process Rewards via Reinforcement with Mutual Information Self-Evaluation Self-Rewarding Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d9b6c656-ea4e-4272-8718-995a3aac259d · inbound
Beyond Verifiable Rewards: Rubric-Based GRM for Reinforced Fine-Tuning SWE Agents Self-Rewarding Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d66bfeb8-2c50-445c-a393-a8a05754e4df · inbound
PoliLegalLM: A Technical Report on a Large Language Model for Political and Legal Affairs Self-Rewarding Language Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation edb53ea2-d14a-4f08-a049-cd13b4ddc6ea · inbound
Neural Garbage Collection: Learning to Forget while Learning to Reason Self-Rewarding Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fdc661dd-a1be-4a94-9a16-75e501a7e4e6 · inbound
IRIS: Interpolative R\'enyi Iterative Self-play for Large Language Model Fine-Tuning Self-Rewarding Language Models
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 541196b6-1eb7-4750-9d55-e6443b126665 · inbound
Skills-Coach: A Self-Evolving Skill Optimizer via Training-Free GRPO Self-Rewarding Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 26262ddf-924c-41f0-817c-81ffcecffe56 · inbound
Enhanced LLM Reasoning by Optimizing Reward Functions with Search-Driven Reinforcement Learning Self-Rewarding Language Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bb5d7e04-394e-4cf1-a2a8-843f79faa19a · inbound
Enhanced LLM Reasoning by Optimizing Reward Functions with Search-Driven Reinforcement Learning Self-Rewarding Language Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 72394f1c-9f91-4a31-872c-bd37d4cad5da · inbound
ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration Self-Rewarding Language Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2ade9202-b664-440e-8877-40ef044adc74 · inbound
StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction Self-Rewarding Language Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ee02b45c-13a6-472d-b9b2-2ebc2d966ee3 · inbound
SEIF: Self-Evolving Reinforcement Learning for Instruction Following Self-Rewarding Language Models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 74acb89b-9a59-48b2-bcba-d1196ba93e95 · inbound
MedThink: Enhancing Diagnostic Accuracy in Small Models via Teacher-Guided Reasoning Correction Self-Rewarding Language Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a8e753a1-485a-4b49-b492-234c6d25770d · inbound
Beyond Static Bias: Adaptive Multi-Fidelity Bandits with Improving Proxies Self-Rewarding Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1744a6ae-c52c-4356-b868-0588f0b8a435 · inbound
CoDistill-GRPO: A Co-Distillation Recipe for Efficient Group Relative Policy Optimization Self-Rewarding Language Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 33059efc-3cad-4536-8292-f914a96910dc · inbound
Primal Generation, Dual Judgment: Self-Training from Test-Time Scaling Self-Rewarding Language Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 88791b15-faf8-4e8d-ae65-0c2d42a27a52 · inbound
Primal Generation, Dual Judgment: Self-Training from Test-Time Scaling Self-Rewarding Language Models
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bd14f2d7-e531-4521-8c78-c8b52055df3b · inbound
From Generic Correlation to Input-Specific Credit in On-Policy Self Distillation Self-Rewarding Language Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 754e0a15-3f02-4279-b195-0293cbe81b07 · inbound
Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization Self-Rewarding Language Models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation af8ddb94-9ef2-4529-9439-fac1feeffc13 · inbound
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Self-Rewarding Language Models
Reference 140
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2fbb6b58-ab9a-49a9-b585-7807f0343689 · inbound
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Self-Rewarding Language Models
Reference 140
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 00f64250-d0ca-494e-841a-b32e878676b4 · inbound
FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale Self-Rewarding Language Models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3b80c03c-943c-4d49-a5b8-e2a91782a24b · inbound
Video-Zero: Self-Evolution Video Understanding Self-Rewarding Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a4fd26ba-f14e-45a1-9ef4-da2ad38dc904 · inbound
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d4bf3bf9-45f9-4391-a1f3-3593c8db321e · inbound
DuIVRS-2: An LLM-based Interactive Voice Response System for Large-scale POI Attribute Acquisition Self-Rewarding Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8d10e24a-c840-471a-8ef0-a4b4a6b7cef6 · inbound
SOLAR: A Self-Optimizing Open-Ended Autonomous Agent for Lifelong Learning and Continual Adaptation Self-Rewarding Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 66035c1a-aac5-4f27-bc4d-4f300070f43d · inbound
Reinforcing Human Behavior Simulation via Verbal Feedback Self-Rewarding Language Models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7d53a406-9368-45d7-9cbe-58e7d17e1046 · inbound
Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable Oversight Self-Rewarding Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4c04b105-014c-4767-84ed-167cb91426d6 · inbound
Ryze: Evidence-Enriched Data Synthesis from Biomedical Papers Self-Rewarding Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3f85ff3b-6de0-4164-9839-a75efce56570 · inbound
Deep Research as Rubric for Reinforcement Learning Self-Rewarding Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e3ae54f1-7685-4b9e-8614-c3c8a25126c1 · inbound
Trust Region On-Policy Distillation Self-Rewarding Language Models
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a2729d11-3347-4b24-91bd-84380aa0c271 · inbound
Exploiting Verification-Generation Gap: Test-Time Reinforcement Learning with Confidence-Conditioned Verification Self-Rewarding Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f325a507-3750-4a3c-ab35-355bd749c9d0 · inbound
Policy-Conditioned Counterfactual Credit for Verifiable Reinforcement Learning of Long-Horizon Language Agents Self-Rewarding Language Models
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c2ef25eb-40bc-4a71-8fea-c8e652265e88 · inbound
Provenance-Grounded Gating and Adaptive Recovery in Synthetic Post-Training Data Curation Self-Rewarding Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c8f387f7-9032-4d3a-9bc6-ead6f3f9782d · inbound
Ouroboros-Spatial: Closing the Data-Model Loop for Spatial Reasoning Self-Rewarding Language Models
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c47cb9a6-e728-43a6-bc43-322a6d365ff7 · inbound
Ouroboros-Spatial: Closing the Data-Model Loop for Spatial Reasoning Self-Rewarding Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2378224-a61f-43f3-af19-e24cc36c68e7 · inbound
Embodied-BenchClaw: An Autonomous Multi-Agent System for Embodied Spatial Intelligence Benchmark Construction Self-Rewarding Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6957cad2-ce98-4d45-9395-7b1907aa4b4e · inbound
Self-Improvement Can Self-Regress: The Rise-and-Collapse Failure Mode of LLM Self-Training Self-Rewarding Language Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1dc97d87-059a-4ada-8bdc-4668aa4cba7f · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Self-Rewarding Language Models
Reference 252
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f4494c5f-6319-4502-acd6-7ff65dac4369 · inbound
Grounded Scaling: Why Agentic AI Needs Deterministic Environments Self-Rewarding Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d9270c68-6889-43cc-9d2b-564542049922 · inbound
PRIDE: Privileged Information-enhanced Distillation for Empathetic Dialogue Generation Self-Rewarding Language Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e49e096e-fe76-41bf-b4d4-74b231a50bbf · inbound
Vision-driven Preference Synthesis for Mitigating Hallucinations in VLMs Self-Rewarding Language Models
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 98e891cc-9b8b-4972-9a39-c755dc36ee82 · inbound
PASTA: A Paraphrasing And Self-Training Approach for Knowledge Updating in LLMs Self-Rewarding Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3f49bf9b-6d01-49a2-aa20-a5051b328941 · inbound
Active-GRPO: Adaptive Imitation and Self-Improving Reasoning for Molecular Optimization Self-Rewarding Language Models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 733794e4-1aa9-4014-8757-65abe37e24ba · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Self-Rewarding Language Models
Reference 210
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d57da815-ead9-4632-b797-8533f778b58e · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Self-Rewarding Language Models
Reference 211
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 615534d4-4def-4d21-ad04-1e6b9a311a42 · inbound
Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation Self-Rewarding Language Models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c51813a2-1fc2-4974-8364-f77c13fce7b9 · inbound
Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation Self-Rewarding Language Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ec5e70a-4643-450b-835c-7f868db92e59 · inbound
Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation Self-Rewarding Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37b4c542-cd58-4b90-b3f2-5b6767c91cf2 · inbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Self-Rewarding Language Models
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 36f27840-916c-40e0-a746-c4ccd5e21f1b · inbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Self-Rewarding Language Models
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f3cf442-54df-48bf-9996-6e8ce059e79d · inbound
More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges Self-Rewarding Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 47f64fd4-1aa6-4d0e-89f7-1d4fef9652ac · inbound
Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops Self-Rewarding Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d69d0e54-baf0-44b8-8252-193322e536a8 · inbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Self-Rewarding Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52d4ddb4-5b6b-4cca-a130-6451205cf9a5 · inbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Self-Rewarding Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7c1d6fa-76c3-40ae-a3d0-e44b3ee16705 · inbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Self-Rewarding Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be6e5875-72f5-48ee-b4e1-3f8dd21fcfb8 · inbound
DeepBias: Adaptive In-depth Probing of Social Biases in LVLMs Self-Rewarding Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 026459f0-1812-4435-92d0-5f8ab495ed82 · inbound
RRPO: Reference-Relative Policy Optimization with Stratified Conditional Rollouts Self-Rewarding Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08c10460-733c-472e-98c8-325aedeb5194 · inbound
FormulaSPIN: Self-Play Fine-Tuning for Natural Language to Spreadsheet Formula Generation Self-Rewarding Language Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 547954f7-9c66-4dd1-af5b-2907541648ac · inbound
Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback Self-Rewarding Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b7c10b3-f874-4fe8-8bb0-b59dfb1efac2 · inbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Self-Rewarding Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.