Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T23:18:40.394477Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 4 inbound Pith citation observations for arXiv:2502.08922.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T23:18:40.394477Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:09:46.031769Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-13T17:08:01.250732Z
49 of 49 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation eb139baf-7580-4c82-8155-cc86f5db0d28 · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Meta-rewarding language models: Self-improving alignment with LLM -as-a-meta-judge
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e724814b-ff4b-49ce-bd8d-18587919ea51 · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022 a
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e892fca-77c4-4940-9743-4061a4d54fa9 · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88a07b6c-c796-4124-84bb-16bfdcd50f94 · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models E., Fort, S., Lanham, T., Telleen-Lawton, T., Conerly, T., Henighan, T., Hume, T., Bowman, S
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7b5a1aec-238c-4517-8551-cdacf8658b67 · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models On the Opportunities and Risks of Foundation Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8421daa7-8b6a-424e-8e6b-1658d8da8220 · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0305702-5474-4972-9d5e-3d4203c67e02 · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e509c3c-332b-4979-9a0d-763c8681ae98 · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models H., Baker, B., Gao, L., Aschenbrenner, L., Chen, Y., Ecoffet, A., Joglekar, M., Leike, J., et al
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2c615e10-6313-4219-b04a-95e09e6e31af · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Discovering latent knowledge in language models without supervision
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d46c78dd-3243-433c-9bfb-621697db17cf · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1163456b-9200-41ef-9695-08a2fd6e6d56 · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1249323-4353-49ba-bdd0-b7cbd6e58b82 · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c7fafe0c-56c6-451f-b9cd-cf2034c80ae6 · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Training verifiers to solve math word problems, 2021
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6f2cf57-9420-4492-867a-7d51ea2b01a1 · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models MetaRM: Shifted Distributions Alignment via Meta-Learning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a435c9e9-de7c-40ac-ac1c-dbf3d4c2c9f1 · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0faf2016-1197-42bf-9496-d41fc3b97446 · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models and Bengio, Y
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 81736b11-9a44-490a-b19a-5d4fe0d9ae56 · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models The Llama 3 Herd of Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4e7d5dd-168e-4055-9f3a-aca257409d27 · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models A Survey on LLM-as-a-Judge
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a82e82e3-79b1-4e5c-8908-5d982ded6360 · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Measuring massive multitask language understanding
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c61ef4f5-603d-4696-bfcf-d7a32ee4e1da · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Large Language Models Can Self-Improve
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89abcfe5-4980-4830-a70f-a54d6375978a · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa350cd6-20a9-4804-aa81-23c4452eae54 · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Mistral 7B
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4af70f34-2ea6-4aab-8e0f-5eda5533175c · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models A survey of reinforcement learning from human feedback, 2024
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 967174d5-25f9-43bf-b316-f897652d2b9c · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models o pf, A., Kilcher, Y., von R \
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c9581455-2476-41a8-b899-d9ccae36168b · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models RewardBench: Evaluating Reward Models for Language Modeling
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cfd3253-0bdc-4c0a-accf-e50797e06f7a · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 295294a0-e936-4b54-972d-5442e5f58595 · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models and Hutter, F
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e09f68dd-25f7-442b-9268-0b9981e5ea09 · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models BioMedGPT: Open Multimodal Generative Pre-trained Transformer for BioMedicine
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5a02edd-f8ea-4c74-9aca-1e8d39ffe7e1 · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Introducing ChatGPT
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 20395fd6-42f6-4f65-821c-fdbb441104fd · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Training language models to follow instructions with human feedback
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7de1d6b8-31cd-43dd-8c9e-2f83d7250299 · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Iterative Reasoning Preference Optimization
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70fe3b7b-f63f-46f8-ab25-0070b9bc4792 · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Disentangling length from quality in direct preference optimization
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cc8016c-02b9-41da-a29a-b5e2a5431e7a · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models D., Ermon, S., and Finn, C
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 903318f1-9c08-4bdf-8239-36e367371e71 · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Proximal Policy Optimization Algorithms
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17deff01-442c-461d-9e67-b0f3dc9307cc · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Loose lips sink ships: Mitigating length bias in reinforcement learning from human feedback
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7f7c52d3-b3d6-435b-9a49-90fde801f122 · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models A Long Way to Go: Investigating Length Correlations in RLHF
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25f07d7e-6a95-4b3b-ac2c-c14c9fdce18b · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Unresolved cited work
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 530e25d9-8427-48e0-b08e-4145d69cc392 · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Llama: Open and efficient foundation language models, 2023
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88a9ef1a-fdf6-43df-9a9a-48aa06193a4e · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Aligning Large Language Models with Human: A Survey
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7164f128-cdf4-44ef-91c6-d7204eaa7eed · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models CREAM: Consistency Regularized Self-Rewarding Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45ba3edd-7f76-45ac-9694-d4a798e03727 · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64166ff1-2b6d-4123-b172-19f7c6def6a6 · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Unsupervised data augmentation for consistency training
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 598702dc-9791-4f54-a428-300b670adccc · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b18fa00a-7ce2-4f66-b358-f64b253c7e37 · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Y., Cho, K., Li, X., Sukhbaatar, S., Xu, J., and Weston, J
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 99baab46-c283-4db2-bea8-20b2703cc36e · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Consistency regularization for cross-lingual fine-tuning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e33d1246-26e6-4c54-b606-ade210b624df · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models E., and Stoica, I
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 176818cb-b9d0-4c68-8670-6542311d105e · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cb36be4-555d-4d63-b195-7f4f991998b5 · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Lima: Less is more for alignment, 2023
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c8fedeb3-353b-4ef0-a942-1cd7d86361a2 · outbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models write newline
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a523748-4ad1-4f59-b1a4-e6c6a8b30bc4 · inbound
From System 1 to System 2: A Survey of Reasoning Large Language Models Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models
Reference 171
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 97b1f075-2d3f-4a02-bc76-f2266f53f3a7 · inbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18db4afd-80b6-4da9-b0f4-c8e5b09750d9 · inbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b43603d-a844-4c6f-9943-96b546c31597 · inbound
Can LLMs Learn to Reason Robustly under Noisy Supervision? Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.