Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:49:51.597998Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 0 inbound Pith citation observations for arXiv:2608.06310.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:49:51.597998Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
77 of 77 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b12cabb1-843c-474d-a18d-35c787b36ab9 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Advances in Neural Information Processing Systems , volume=
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 22df6d52-8ba6-4011-9982-cd97d2cc0ea1 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Proceedings of the AAAI Conference on Artificial Intelligence , volume=
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 874c2201-b413-4803-bccf-15fc82642af2 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Unified Reward Model for Multimodal Understanding and Generation
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 872aecdc-9cd9-4a4f-a1a1-b8c05bff6acf · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Advances in neural information processing systems , volume=
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f2875dae-92c1-47ac-88bf-3db3247f8185 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e834e349-dfe1-4c35-b6a7-699bc4d1b514 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction WorldPM: Scaling Human Preference Modeling
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1eae25c-98dc-46f4-bd31-67057d34c166 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5ac76aea-1516-4274-aa1a-c86127d9e040 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Reward is enough: Llms are in-context reinforcement learners , volume =
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d9c87546-9603-4be6-9e86-b1c393da434a · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 88ac1e2a-c41b-44a4-821e-b08628422215 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Scaling laws for neural language models , volume =
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3c3453f5-e0ce-45b3-8d12-65c133728908 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Large language models are not fair evaluators , year =
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a26e3229-0fc7-499a-9efc-b1b29c362500 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Theoretical and empirical evaluation of data reduction for exact Kemeny rank aggregation , volume =
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation abc954f5-9411-4fb6-88fb-fb668c94d847 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Advances in Neural Information Processing Systems , volume=
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a288283c-417e-43f7-b702-0afebe30db57 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Improved parameterized algorithms for the Kemeny aggregation problem , year =
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a3233059-6b6a-45b9-9c8d-52dbd3cd32d7 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Are we done with mmlu? , year =
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9dddf58f-edbd-4289-ae91-9ba7ba432359 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Length-controlled alpacaeval: A simple debiasing of automatic evaluators , year =
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3fd36120-3b59-402d-8147-bfbd4713ea10 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Let's Verify Step by Step , year =
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cfcb2e39-e491-480b-a1b7-85d637b189bc · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Gpqa: A graduate-level google-proof q&a benchmark , year =
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0fa929c0-9b64-4521-a979-12c017d6a46f · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction From crowdsourced data to high-quality benchmarks: Arena-hard and benchbuilder pipeline , volume =
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3fe14d06-9186-4c56-9830-380a238aacd7 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Wildbench: Benchmarking llms with challenging tasks from real users in the wild , volume =
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3e82869b-6700-4aab-a177-234bab8f057f · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Hashimoto , howpublished =
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 355a6e70-0d2d-43f0-89b1-a2fda7ff1101 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction SimPO: Simple Preference Optimization with a Reference-Free Reward , year =
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 19117320-1377-4f01-b117-5727f531c0b1 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction HelpSteer3-Preference: Open Human-Annotated Preference Data across Diverse Tasks and Languages , volume =
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9d034b94-468a-43be-adde-ba6bbeade5ea · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Qwen2 technical report , volume =
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 365a72b4-5c00-4d89-bb4f-1ce13ea0d5f8 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction The llama 3 herd of models , volume =
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e5370f44-3ffc-4801-8fbe-d55fc5530a3a · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Rrhf: Rank responses to align language models with human feedback without tears , year =
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 85fedc31-14ce-4255-b22b-5d720dbc7908 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction A computational study of the Kemeny rule for preference aggregation , year =
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 59d360b3-9b3d-456f-8839-3bdac42719d9 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Judgebench: A benchmark for evaluating llm-based judges , year =
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 76e7c7a9-8a8b-4d97-8ccb-726744604d13 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Rm-bench: Benchmarking reward models of language models with subtlety and style , year =
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fa601f3e-928d-4972-b376-ebb6eb3ad227 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Proximal policy optimization algorithms , year =
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ff36b12-62a6-450f-bea7-5ed06b613eb1 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Rank analysis of incomplete block designs: I
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 85e72548-907a-4bea-85fa-b08ebf5e0fd7 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Reinforcement Learning with Verifiable Rewards: GRPO's Effective Loss, Dynamics, and Success Amplification , year =
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8a42ff62-ad78-4892-9759-5b79eb8f1844 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Language models that think, chat better , year =
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ef34f03e-4bb5-40be-849c-abb9c80bb355 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Dissecting Long Reasoning Models: An Empirical Study , year =
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation defa5713-cebe-4c74-93d7-09fa59f6532f · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Reinforcement learning with verifiable rewards implicitly incentivizes correct reasoning in base llms , year =
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b4c28ff7-c18d-4f2b-ac58-2b744eb86fe8 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1b3741e8-8e3a-45a2-9816-bb41d5a2b51c · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Pre-Trained Policy Discriminators are General Reward Models , year =
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ea53c487-dc4c-4eeb-ba1a-bdbd04c02b01 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Dynamic Reward Adjustment in Multi-Reward Reinforcement Learning for Counselor Reflection Generation , year =
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 91696147-46b0-43a5-866b-90ec5bcbbe09 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Contrastive Preference Optimization: Pushing the Boundaries of
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e2cfe807-ba89-40b6-af9b-ae97d316bef8 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction From system 1 to system 2: A survey of reasoning large language models , year =
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 42037de2-a9df-499d-9cc1-42189ddef0e3 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b566abb0-1aa6-4453-b6c6-97ae51c03c6d · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Prior constraints-based reward model training for aligning large language models , year =
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation de7f27d8-4851-4122-94fc-73e36ae6b8bd · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Improving In-Context Learning via Sequentially Selection and Preference Alignment for Few-Shot Aspect-Based Sentiment Analysis , year =
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 119bb83d-96b8-4f36-a699-fe18ba729caa · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models , year =
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 38c72927-3206-40a4-98de-0a66f7aa5a28 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Manning and Stefano Ermon and Chelsea Finn , booktitle =
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5f6a0495-ddc0-49fa-8cfd-e859a86ded4f · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Discriminative Reranking for Neural Machine Translation , year =
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e4bcb13c-a52a-4763-809d-d513742ce78b · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Dapo: An open-source llm reinforcement learning system at scale , year =
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b28d81b-3420-4344-b60c-6f5b0b42fc86 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Deepseekmath: Pushing the limits of mathematical reasoning in open language models , year =
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 057e7e02-d2c4-4063-a2e4-9bb69bf4c57a · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Generative reward modeling via synthetic criteria preference learning , year =
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6efafa89-5f2b-4bb9-9ea4-e4d74a342fb2 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Unified multimodal chain-of-thought reward model through reinforcement fine-tuning , year =
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d03689ab-3acf-429a-b7ab-fc5322fbf02f · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Rm-r1: Reward modeling as reasoning , year =
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cce62ad7-7a71-4235-b790-53db65196ea3 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Reward reasoning model , year =
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2b9f8ed9-f566-4459-9c84-dc3dbcae46fc · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction GRAM-R ^2 : Self-Training Generative Foundation Reward Models for Reward Reasoning , year =
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fe4d15e3-bc62-408c-851f-2489c4f00019 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction GRAM: A Generative Foundation Reward Model for Reward Generalization , year =
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fc0e704b-c1bf-4185-b0f1-5273ea7b7dab · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Inference-time scaling for generalist reward modeling , year =
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5a571064-f449-4e59-bd29-3d741065ed9a · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Reward Model Ensembles Help Mitigate Overoptimization , year =
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8b55f41c-0c8f-4224-a776-77118909d455 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy , year =
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0c1242af-9be7-489e-b125-e1785b84d53a · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Rovrm: A robust visual reward model optimized via auxiliary textual preference data , year =
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b1d32292-7cd2-4a90-bbde-433e37e585c5 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Specialist or Generalist? Instruction Tuning for Specific
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ea9c6137-14cb-4ddb-8dad-99dcf3f0622d · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Unveiling the Generalization Power of Fine-Tuned Large Language Models , year =
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aef914da-2aac-4694-baa9-bfca318d3932 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning , year =
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 823d99c6-f758-4b11-a3bd-c88f0f608d21 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Chi and Quoc V
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf973903-1f16-46b4-ac31-9a8109a64c7e · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Scaling instruction-finetuned language models , year =
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 22f12c7c-1afd-4d0a-bd3a-4081930d89c9 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction ArXiv preprint , title =
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bc442f6e-970a-4c0d-ac49-e85517058b0d · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Generative verifiers: Reward modeling as next-token prediction , year =
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9680b05a-c1c1-4367-81c7-ec56d5854582 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Foundations of large language models , year =
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d3e6dc32-ce32-44c9-ae07-f5a2d468a4ab · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Step-level verifier-guided hybrid test-time scaling for large language models , year =
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f1499bb0-3728-49e7-ba81-b386ef8fedb2 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction s1: Simple test-time scaling , year =
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6962bda8-8b9a-482d-885d-ab4b684d4b79 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Ziegler and Ryan Lowe and Chelsea Voss and Alec Radford and Dario Amodei and Paul F
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2e27ccf8-bc49-44e0-8a17-e6ad72a664fb · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Christiano and Jan Leike and Tom B
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8475e862-f0f7-479e-9626-806ee73707bc · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Pku-saferlhf: Towards multi-level safety alignment for llms with human preference , year =
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c4d27562-bd7a-497e-9384-c35c7d3df281 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Training a helpful and harmless assistant with reinforcement learning from human feedback , year =
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a7abe3c-2669-44aa-a261-06f63631e3f5 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Hybrid alignment training for large language models , year =
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 42a3a541-e686-4a72-8b1d-e169d7cf4070 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Unresolved cited work
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4fe81ff3-0c90-4e39-99a0-995292c5cf9f · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Scaling Learning Algorithms Towards
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85bf7a5a-2af4-417c-ad49-306fdf7e79fd · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction and Osindero, Simon and Teh, Yee Whye , journal =
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88a5e574-b9ad-4d50-83de-e772ac4b4c88 · outbound
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction 2016 , publisher=
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.