Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T03:14:13.563948Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2607.26094.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T03:14:13.563948Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
20 of 20 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 51040481-5837-49d3-8358-f99c6d01eb97 · outbound
Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fb2c7a8-6ead-4e92-b780-9321a19dc83a · outbound
Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback UltraFeedback: Boosting Language Models with Scaled AI Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d10d2cd2-a4c2-4cb7-88a7-3f85aa373f0c · outbound
Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback RewardBench: Evaluating Reward Models for Language Modeling
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48a605e5-6313-47ef-8ce5-c44dc9c7a684 · outbound
Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0645fb08-2d27-4bd5-8b20-734940f2ba5c · outbound
Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback Let's Verify Step by Step
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60712595-d350-4103-9b40-62289167d629 · outbound
Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback Highly Efficient Self-Adaptive Reward Shaping for Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e68c4f83-74cb-4aa6-ae17-d92108e4637f · outbound
Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 994fa857-840e-4df8-8a72-8c376f28cd8e · outbound
Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback Proximal Policy Optimization Algorithms
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 463e2fcb-5601-4877-bee7-ed326d82e249 · outbound
Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee29d25c-1280-4512-ba25-21417f06d751 · outbound
Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback Solving math word problems with process- and outcome-based feedback
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36eec2bb-fe64-48d6-9715-b767de896496 · outbound
Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback Secrets of RLHF in Large Language Models Part II: Reward Modeling
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7d42c34-5f16-46a9-9cb9-82521729b748 · outbound
Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback Shepherd: A Critic for Language Model Generation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78f09d93-52dd-4add-9bea-8dfd62e6d254 · outbound
Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback Text2Reward: Reward Shaping with Language Models for Reinforcement Learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 547954f7-9c66-4dd1-af5b-2907541648ac · outbound
Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback Self-Rewarding Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afc01fc8-d4da-4fb7-83db-221c62e941f5 · outbound
Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback Instruction-Following Evaluation for Large Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03ec0e56-5877-4bdf-856e-e1f22df7607a · outbound
Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback Reward Shaping via Meta-Learning
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbeb53db-f0cd-4a26-8a74-66021b08a0fb · outbound
Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback LoRA: Low-Rank Adaptation of Large Language Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff93314f-7801-4b9c-9a7d-3773201a0dfc · outbound
Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e737dde-6869-45d3-9274-77ad0eb7785c · outbound
Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e3c978e-8427-4234-aef3-759f534a98e3 · outbound
Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.