Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:34:53.874909Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 15 inbound Pith citation observations for arXiv:2505.14674.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:34:53.874909Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:15:37.528799Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T06:49:38.201535Z
87 of 87 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2b5dc76e-1feb-4983-8214-fa8e19d1cf8a · outbound
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef46c4eb-d007-43dd-bcac-af4fe3fbe2f0 · outbound
Reward Reasoning Model Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10d6e181-a23c-4326-a1ea-b681e34819c8 · outbound
Reward Reasoning Model Atla Selene Mini: A General Purpose Evaluation Model
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 324dd137-dbb8-4b26-ad1d-bfb5723149e7 · outbound
Reward Reasoning Model The claude 3 model family: Opus, sonnet, haiku.Claude-3 Model Card, 1:1, 2024
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abb5fcfe-f89d-4712-9a0a-c907f47bcfec · outbound
Reward Reasoning Model Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 47b45be1-fd2f-4c65-b61d-eb409a76aef7 · outbound
Reward Reasoning Model Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6cb77be-ec9f-4514-991a-2a89a49521dc · outbound
Reward Reasoning Model Constitutional AI: Harmlessness from AI Feedback
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9e244ad-e186-419b-a6b6-3427b84bc06e · outbound
Reward Reasoning Model Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d2dc476-2b95-4d48-8f9e-b5fce7fb796a · outbound
Reward Reasoning Model Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2be30d7-3095-4200-b422-7a1795082215 · outbound
Reward Reasoning Model Sparks of artificial general intelligence: Early experiments with GPT-4, 2023
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f46c7d44-57de-479f-95fe-2c2c9e4dd482 · outbound
Reward Reasoning Model CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecaf28c3-59e1-4e5f-9512-4f3d20541476 · outbound
Reward Reasoning Model Judgelrm: Large reasoning models as a judge.arXiv preprint arXiv:2504.00050, 2025
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6981fcf-1d7b-42ff-b33b-3da231f42faf · outbound
Reward Reasoning Model Seal: Steerable reasoning calibration of large language models for free
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89dbad2b-12a3-4097-97aa-2fe0d4550f18 · outbound
Reward Reasoning Model Rm-r1: Reward modeling as reasoning.arXiv preprint arXiv:2505.02387, 2025
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4325b17c-a6d6-41a3-89e8-0f630abe60e5 · outbound
Reward Reasoning Model Christiano, Jan Leike, Tom B
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c864530-7a15-455c-a7f2-e29d36148472 · outbound
Reward Reasoning Model Training Verifiers to Solve Math Word Problems
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99b94b53-9660-4767-8b08-eb18dde2642a · outbound
Reward Reasoning Model Elo.The Rating of Chessplayers, Past and Present
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e5287c2-99f5-44e1-9322-1a6eb4ea874c · outbound
Reward Reasoning Model Gonzalez, and Ion Stoica
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f88e40d-533c-4b66-bd39-e43d86ce9f77 · outbound
Reward Reasoning Model On Designing Effective RL Reward at Training Time for LLM Reasoning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 033762ca-1f77-4f47-a0ad-dbf9d7e01395 · outbound
Reward Reasoning Model Scaling laws for reward model overoptimization
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca1e8f79-71e1-4408-afda-e93bfc3593fb · outbound
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9251b517-e7ab-4e26-8341-e109762f8faf · outbound
Reward Reasoning Model DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8aff6e2-39e4-47fe-9c9e-541080cb20fd · outbound
Reward Reasoning Model Mcranker: Generating diverse criteria on-the-fly to improve pointwise llm rankers
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b2953dc-7c4b-4ecd-a612-2d49d768aa75 · outbound
Reward Reasoning Model Skywork open reasoner series
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6099a2b-3bf3-4ae4-aec6-0aec91939039 · outbound
Reward Reasoning Model Measuring mathematical problem solving with the MATH dataset
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f12a8b94-3cf5-4ae0-851e-f9a5556d8a5d · outbound
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31bd0b84-c20e-4a88-b147-8b2dd8f52a19 · outbound
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 536d8b20-32fd-46c3-a843-3794f23bc031 · outbound
Reward Reasoning Model LLM-blender: Ensembling large language models with pairwise ranking and generative fusion
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd8e1793-3b22-41ac-acd8-1db43cb79403 · outbound
Reward Reasoning Model SWE-bench: Can language models resolve real-world github issues? InThe Twelfth International Conference on Learning Representations, 2024
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 825d3a07-6a2b-4c64-aec2-07b755b4daa0 · outbound
Reward Reasoning Model Scaling Laws for Neural Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8afb08ed-a533-41f1-a60d-44f323f602d8 · outbound
Reward Reasoning Model Maple: Multi-modal prompt learning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95386e55-d1bd-4352-9799-9053ebb635a1 · outbound
Reward Reasoning Model Prometheus 2: An open source language model specialized in evaluating other language models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce635670-f63c-447b-afd0-6a86ceaab1fe · outbound
Reward Reasoning Model Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ce3f9a5-3861-4ee6-b727-1c15d0e0512d · outbound
Reward Reasoning Model RewardBench: Evaluating Reward Models for Language Modeling
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d59bc306-acfb-42b5-9cef-d49b096d8c08 · outbound
Reward Reasoning Model Generative judge for evaluating alignment
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 97e93a02-a458-4cf8-bb3d-1fa1e9b98994 · outbound
Reward Reasoning Model From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0dc8f84-b9f5-42c8-8c22-50f1fcd63e53 · outbound
Reward Reasoning Model Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 17a41f16-cc95-4c35-847c-587a961c0c0d · outbound
Reward Reasoning Model Rouge: A package for automatic evaluation of summaries
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ef38318-1a13-4659-b948-b3841db37b33 · outbound
Reward Reasoning Model Let’s verify step by step
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10f0424d-d52f-4e3b-bd01-22350bb450f2 · outbound
Reward Reasoning Model PairJudge RM: Perform Best-of-N Sampling with Knockout Tournament
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1944545c-20ca-4cd3-a1b0-57e105fd9972 · outbound
Reward Reasoning Model Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f772cdb-2858-463c-9082-4001e982e021 · outbound
Reward Reasoning Model General-reasoner: Advancing llm reasoning across all domains
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fbbf5d0d-d666-4fea-beba-18ae91c32152 · outbound
Reward Reasoning Model Inference-time scaling for generalist reward modeling.arXiv preprint arXiv:2504.02495, 2025
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca842ffc-23a6-46a5-a14a-e0cee2affeb2 · outbound
Reward Reasoning Model Training language models to follow instructions with human feedback
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 04e333ca-9936-4e0a-b645-56707c172df2 · outbound
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e156917e-cdb5-489a-9e9f-32173197839d · outbound
Reward Reasoning Model Bleu: a method for automatic evaluation of machine translation
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e9deba9-8e08-489e-9407-3eeaba038f60 · outbound
Reward Reasoning Model Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d40c1b1c-1b3f-4874-8d8c-5e9cad63712b · outbound
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f077f48-f53e-4671-8c10-968b40c4aba1 · outbound
Reward Reasoning Model OffsetBias: Leveraging debiased data for tuning evaluators
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d207513-7fda-4cb8-bf09-5dedac3090c0 · outbound
Reward Reasoning Model Bradley Knox, Chelsea Finn, and Scott Niekum
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 93b2115a-a6af-4994-afc3-20a88c48f4dd · outbound
Reward Reasoning Model Direct preference optimization: Your language model is secretly a reward model
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b28ebea-f75d-48d4-8953-78e11cbea415 · outbound
Reward Reasoning Model DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5b06b04-de8e-479d-85e1-a53711d9db2d · outbound
Reward Reasoning Model The limits of automatic summarisation according to rouge
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a9d9763e-c9e5-4662-920b-503fffe91314 · outbound
Reward Reasoning Model Skywork critic model se- ries
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c60e5510-ccfa-4c84-8520-879502e2d277 · outbound
Reward Reasoning Model HybridFlow: A Flexible and Efficient RLHF Framework
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8540d3f5-3be1-4a25-9b67-af2314e4b64a · outbound
Reward Reasoning Model LLaMA: Open and Efficient Foundation Language Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfef8121-af75-407d-85ed-ec9b6453c90e · outbound
Reward Reasoning Model Scaling LLM test- time compute optimally can be more effective than scaling parameters for reasoning
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a812a115-e4e8-4789-8eca-a48adb5d8948 · outbound
Reward Reasoning Model Secrets of RLHF in Large Language Models Part II: Reward Modeling
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e4f17b6-25e3-441e-a397-57d7508cd05e · outbound
Reward Reasoning Model Foundational autoraters: Taming large language models for better automatic evaluation
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72abce9f-4137-4ad2-b582-5787a5b1cfe9 · outbound
Reward Reasoning Model Pandalm: An automatic evaluation benchmark for llm instruction tuning optimization
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d720deb4-efae-4f39-b636-0d2064001d37 · outbound
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b22891b-1cb1-4eb9-8f68-fa2c7b066dd8 · outbound
Reward Reasoning Model HelpSteer: Multi-attribute Helpfulness Dataset for SteerLM
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 341ee96e-1efe-4a1d-bfac-48453a6ff084 · outbound
Reward Reasoning Model Unresolved cited work
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 34bbcde9-445e-44e9-a73d-b4270e62c790 · outbound
Reward Reasoning Model Reinforcement learning for reasoning in large language models with one training example
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ba7f0d83-0f17-420d-866d-a61a987a3d3a · outbound
Reward Reasoning Model J1: Incentivizing thinking in llm-as-a-judge via reinforcement learning.arXiv preprint arXiv:2505.10320, 2025
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 514a7db1-078e-4cde-b405-c3bca95521a6 · outbound
Reward Reasoning Model Zhang, Makesh Narsimhan Sreedhar, and Oleksii Kuchaiev
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a79a40e6-edfa-40cc-b0df-2f7e4ae87bf4 · outbound
Reward Reasoning Model Chi, Quoc V Le, and Denny Zhou
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6245467e-b388-4761-9d3c-06b0d0775aed · outbound
Reward Reasoning Model Inference scaling laws: An empirical analysis of compute-optimal inference for LLM problem-solving
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3914fb9b-10c8-4674-9571-b66f9dc9e801 · outbound
Reward Reasoning Model Metametrics: Calibrating metrics for generation tasks using human preferences
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 84ed587f-2345-4098-8b28-592c52a1dbb8 · outbound
Reward Reasoning Model Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed3d3685-5069-4f6d-a549-19afa6aca6af · outbound
Reward Reasoning Model Learning LLM-as-a-judge for preference alignment
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 24c0ae7f-4b60-4b0e-9cc0-07b08648a264 · outbound
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85e6dfe3-c4c2-442f-bd24-53cb550d790f · outbound
Reward Reasoning Model Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de494628-796f-48eb-8e34-1871d3509850 · outbound
Reward Reasoning Model Self-rewarding language models
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6fdcaf4-4f08-469b-9515-792345129f24 · outbound
Reward Reasoning Model DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e64558c4-476e-4df2-888d-8be1b6db038b · outbound
Reward Reasoning Model Self-generated critiques boost reward modeling for language models
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3868ed14-945a-4658-936a-d99a0c28f3ae · outbound
Reward Reasoning Model Gonzalez, and Ion Stoica
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf5a87e0-1d19-497d-8395-e141e26afb38 · outbound
Reward Reasoning Model Generative verifiers: Reward modeling as next-token prediction
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a5b784e3-5bd9-4f6e-949f-953ce030b548 · outbound
Reward Reasoning Model The Lessons of Developing Process Reward Models in Mathematical Reasoning
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09dd05a2-446f-4202-8241-8fb9c98a7b3c · outbound
Reward Reasoning Model JudgeLM: Fine-tuned large language mod- els are scalable judges
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9df9a508-d550-4afe-bdfb-f965df1da3ef · outbound
Reward Reasoning Model Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e82255d5-fafd-40d0-857b-3536af06b0d7 · outbound
Reward Reasoning Model A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f22bab45-4e75-432a-8778-e7f7e5105b1a · outbound
Reward Reasoning Model Partially Adhered
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5065b370-aea8-49c4-9f0b-5c6e0df169ea · outbound
Reward Reasoning Model Useful but Incomplete
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 51c7af70-e049-4117-9136-927f002999ed · outbound
Reward Reasoning Model Not Detailed
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7be62865-2f07-45af-bfcd-045c54341bcd · outbound
Reward Reasoning Model doi: 10.18653/v1/2024.emnlp-main.248
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6be746c7-c633-497b-844a-2731a542c577 · outbound
Reward Reasoning Model \boxed{Assistant 1}
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation aab485b9-07c5-43db-a0cd-505f144055de · inbound
Tournament of Prompts: Evolving LLM Instructions Through Structured Debates and Elo Ratings Reward Reasoning Model
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04ec0a88-7b73-41a4-92c7-b46d6cf7ab9b · inbound
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains Reward Reasoning Model
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a2818eef-f5ce-4311-b868-782e9556d367 · inbound
GenSelect: A Generative Approach to Best-of-N Reward Reasoning Model
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf18a61e-8d83-4445-8171-7d29f8bc4780 · inbound
VRPRM: Process Reward Modeling via Visual Reasoning Reward Reasoning Model
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a5c71336-4ebc-41fc-83b2-a6491ddc19d0 · inbound
VRPRM: Process Reward Modeling via Visual Reasoning Reward Reasoning Model
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25a4676c-98de-4856-8bc4-f4660d2fd46d · inbound
Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs Reward Reasoning Model
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 589bb886-b043-4291-8a39-8c5ed8c35d8c · inbound
A Survey of Reinforcement Learning for Large Reasoning Models Reward Reasoning Model
Reference 173
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c3900aa1-9feb-460d-9ceb-9b19a2dc5c8d · inbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Reward Reasoning Model
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ee2d70b-a2a2-46cf-ad4e-9e71f89c5c87 · inbound
PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling Reward Reasoning Model
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c37af3d7-af75-4336-af67-24164bd603c7 · inbound
OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning Reward Reasoning Model
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a625218-dee7-4b6e-b395-3d97d857f5b3 · inbound
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e63bc315-3599-4d2c-9fd4-993a8f2154b0 · inbound
Reference 155
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ad125d65-8338-4ff8-b3a2-25cb045d814e · inbound
Biological Reasoning-Informed Regression for Interpretable Regulatory DNA Activity Prediction Reward Reasoning Model
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d329b7c8-0292-4432-947f-e065d7525ebf · inbound
Counsel: A Meta-Evaluation Dataset for Agentic Tasks Reward Reasoning Model
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 12a1c266-d55d-436a-b877-aadc29b49771 · inbound
TAPAS: Throughput-adaptive Perception for Autonomous Systems Reward Reasoning Model
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.