Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:09:38.960214Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2506.09096.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:09:38.960214Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
52 of 52 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4132a0e4-49f0-476a-91db-a8c37c4f63d7 · outbound
Intra-Trajectory Consistency for Reward Modeling Training language models to follow instructions with human feedback
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bec86bc4-369b-4086-b50c-f16b257b9303 · outbound
Intra-Trajectory Consistency for Reward Modeling Safe RLHF : Safe reinforcement learning from human feedback
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2fcd8772-3748-4012-983b-cf3ccc1b3c56 · outbound
Intra-Trajectory Consistency for Reward Modeling Model alignment as prospect theoretic optimization
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 99d4cb26-b296-413f-a0f2-d56a0e88cb02 · outbound
Intra-Trajectory Consistency for Reward Modeling A Survey of Direct Preference Optimization
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fde3e4a-7d27-4dd2-8899-cc5c84a54d8f · outbound
Intra-Trajectory Consistency for Reward Modeling Generative verifiers: Reward modeling as next-token prediction
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b9fe2fbb-cf39-4bd0-ab3a-daf565f93a65 · outbound
Intra-Trajectory Consistency for Reward Modeling Rewarding progress: Scaling automated process verifiers for LLM reasoning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6661a49b-0f37-47d9-b026-caa023ae1740 · outbound
Intra-Trajectory Consistency for Reward Modeling Scaling laws for reward model overoptimization
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 726cb910-5925-4f6c-b4a9-48241819474f · outbound
Intra-Trajectory Consistency for Reward Modeling Regularizing hidden states enables learning generalizable reward model for LLM s
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 53f1914c-cd13-4cac-aaaa-4e94bc21976d · outbound
Intra-Trajectory Consistency for Reward Modeling Reward model ensembles help mitigate overoptimization
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation edff0ebb-c86c-41a1-a74a-25c15ce9b1c3 · outbound
Intra-Trajectory Consistency for Reward Modeling Warm: On the benefits of weight averaged reward models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 32e99435-33a6-4fd5-a7d6-30f7ff031b90 · outbound
Intra-Trajectory Consistency for Reward Modeling The trickle-down impact of reward inconsistency on RLHF
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 51ecde41-b6f3-43e0-b887-94d27783c4df · outbound
Intra-Trajectory Consistency for Reward Modeling Rrm: Robust reward model training mitigates reward hacking
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bbc1ae30-c7e8-4a0c-860c-cac1e1c9f90a · outbound
Intra-Trajectory Consistency for Reward Modeling Length-controlled alpacaeval: A simple debiasing of automatic evaluators
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2b9af422-9ecf-4211-89ba-30c3ec0a3010 · outbound
Intra-Trajectory Consistency for Reward Modeling Odin: Disentangled reward mitigates hacking in rlhf
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fdc4bee4-9ae1-446c-aa16-55969f355381 · outbound
Intra-Trajectory Consistency for Reward Modeling Improving discriminative capability of reward models in rlhf using contrastive learning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 49f91bbc-1be1-4b00-9db4-52a9d972270c · outbound
Intra-Trajectory Consistency for Reward Modeling Rethinking reward modeling in preference-based large language model alignment
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9049e081-f6f4-4c1f-a5fa-e908588fa405 · outbound
Intra-Trajectory Consistency for Reward Modeling Let's verify step by step
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e927d83-6394-4273-ac20-844e64b50103 · outbound
Intra-Trajectory Consistency for Reward Modeling Math-shepherd: Verify and reinforce llms step-by-step without human annotations
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7a7ca000-670d-417c-8a6a-7a8db72ec35f · outbound
Intra-Trajectory Consistency for Reward Modeling Rest-mcts*: Llm self-training via process reward guided tree search
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0e4c4813-eb08-4f65-bbb2-cbf80ae1a160 · outbound
Intra-Trajectory Consistency for Reward Modeling Gemma: Open models based on gemini research and technology
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0b1e1919-0da9-432f-98da-33781cc30cd4 · outbound
Intra-Trajectory Consistency for Reward Modeling The llama 3 herd of models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 379d687f-f425-49c7-bf25-29e493da0680 · outbound
Intra-Trajectory Consistency for Reward Modeling Training a helpful and harmless assistant with reinforcement learning from human feedback
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e15458fe-de22-4138-89a3-13ee36986d9f · outbound
Intra-Trajectory Consistency for Reward Modeling Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2c87082c-53e2-4bcf-8941-90a39fabb102 · outbound
Intra-Trajectory Consistency for Reward Modeling Solving math word problems with process-and outcome-based feedback
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b6be155d-90e6-4dfb-8466-2f519790f09e · outbound
Intra-Trajectory Consistency for Reward Modeling Token-level direct preference optimization
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 835f56c4-ee2f-4d5e-b96b-004c8aae0fb5 · outbound
Intra-Trajectory Consistency for Reward Modeling Process reinforcement through implicit rewards
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 521f59ea-afed-41d1-be65-117b7ccc2e9c · outbound
Intra-Trajectory Consistency for Reward Modeling Gen ARM : Reward guided generation with autoregressive reward model for test-time alignment
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1be149c5-8847-42b9-a1aa-f1115b8585de · outbound
Intra-Trajectory Consistency for Reward Modeling Fine-grained human feedback gives better rewards for language model training
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 196c5c63-3519-4ae4-95de-a01b34f6c74a · outbound
Intra-Trajectory Consistency for Reward Modeling Safety alignment should be made more than just a few tokens deep
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6237dda3-1d2e-4d98-b747-9e20ee8a911f · outbound
Intra-Trajectory Consistency for Reward Modeling Process reward model with q-value rankings
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ce204ed8-5904-45fb-9405-ced160567cf9 · outbound
Intra-Trajectory Consistency for Reward Modeling Coda: Contrast-enhanced and diversity-promoting data augmentation for natural language understanding
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 98bdd41b-ec82-485a-943d-8e407496663e · outbound
Intra-Trajectory Consistency for Reward Modeling A survey on data augmentation for text classification
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 44606918-c76c-4a6c-bd26-703a45d2b9f7 · outbound
Intra-Trajectory Consistency for Reward Modeling Cubuk, Alex Kurakin, Han Zhang, and Colin Raffel
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8e096c90-a31e-446b-8466-1a2f3dbd2c6d · outbound
Intra-Trajectory Consistency for Reward Modeling Flexmatch: Boosting semi-supervised learning with curriculum pseudo labeling
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 62cc8540-6c22-42c8-bf54-bd30eaa62a12 · outbound
Intra-Trajectory Consistency for Reward Modeling Qwen2 technical report
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5b5989df-c9f7-487c-abac-0ae2dad00254 · outbound
Intra-Trajectory Consistency for Reward Modeling Measuring mathematical problem solving with the math dataset
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 73391db7-3f72-47da-8f0c-1a3638a177ca · outbound
Intra-Trajectory Consistency for Reward Modeling Free process rewards without process labels
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3bc09aff-6ed7-4505-8345-735472625b6b · outbound
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2697b11-fb0a-4340-8c89-b747571cf43d · outbound
Intra-Trajectory Consistency for Reward Modeling Direct preference optimization: Your language model is secretly a reward model
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 10235350-5bec-4bbc-a2e7-56b114fc2575 · outbound
Intra-Trajectory Consistency for Reward Modeling Rlhf workflow: From reward modeling to online rlhf
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6e5a6512-f8d7-4337-b74d-7e6e175ebe34 · outbound
Intra-Trajectory Consistency for Reward Modeling Rewardbench: Evaluating reward models for language modeling
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e05e6164-2d74-45cd-a9a0-df3fa4aefba7 · outbound
Intra-Trajectory Consistency for Reward Modeling Llama 2: Open foundation and fine-tuned chat models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 43d67609-7c95-4db5-b8a2-0c528715a0b9 · outbound
Intra-Trajectory Consistency for Reward Modeling Secrets of rlhf in large language models part ii: Reward modeling
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f81486a0-a0ec-4cec-83a8-fc12d2ec4624 · outbound
Intra-Trajectory Consistency for Reward Modeling Reward model ensembles help mitigate overoptimization
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3c2aebab-b8cb-4e42-b102-a5ccbcf6015e · outbound
Intra-Trajectory Consistency for Reward Modeling Alpacafarm: A simulation framework for methods that learn from human feedback
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 30b3becf-60e8-4b2f-b91a-08d6e6de7e84 · outbound
Intra-Trajectory Consistency for Reward Modeling Loose lips sink ships: Mitigating length bias in reinforcement learning from human feedback
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d8e979ef-ed7f-4cfb-ac10-6836fa9f4ea1 · outbound
Intra-Trajectory Consistency for Reward Modeling Pfohl, Deepak Ramachandran, Peter Shaw, and Jonathan Berant
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7d9606f7-8284-4b33-8833-7550696c0906 · outbound
Intra-Trajectory Consistency for Reward Modeling Proximal policy optimization algorithms
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 49fbd15f-3dca-4368-b712-494501cc534c · outbound
Intra-Trajectory Consistency for Reward Modeling Understanding the learning dynamics of alignment with human feedback
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f245d885-99b7-4bee-be2f-03ae603fcd91 · outbound
Intra-Trajectory Consistency for Reward Modeling Improve mathematical reasoning in language models by automated process supervision
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dab70296-bd7d-493d-b28f-2bcddad4ce2f · outbound
Intra-Trajectory Consistency for Reward Modeling Llamafactory: Unified efficient fine-tuning of 100+ language models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1327d8a5-e129-4bc2-a7e8-75d29996d2b4 · outbound
Intra-Trajectory Consistency for Reward Modeling Lora: Low-rank adaptation of large language models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.