Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2504.04950.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:34:54.497908Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T20:46:14.356541Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 6b797002-0f2c-4569-b19c-7217b5948dae · inbound
Seed1.5-VL Technical Report A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization
Reference 157
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 79be537c-ece5-4495-b294-07b595d9fd67 · inbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e894e89-aaa8-40d4-a55e-a065552fb817 · inbound
WebDancer: Towards Autonomous Information Seeking Agency A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 972cb59f-2d79-474c-be90-22b1351106c5 · inbound
PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96ff8a98-aaa1-4426-84c9-7d72a24e8711 · inbound
RewardDance: Reward Scaling in Visual Generation A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64bc43cc-a96f-4ce5-b2fb-14d01e7ae83d · inbound
Voting with the Graph: Stable RLAIF via Topological Consistency Maximization A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa743bd7-6033-40e5-8722-9cbe0250b7bd · inbound
PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f7d6ff37-2f61-492e-a2ec-ca2b96a8755e · inbound
Leveraging Verifier-Based Reinforcement Learning in Image Editing A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b758234d-2223-4943-955d-1e58f95674e2 · inbound
Leveraging Verifier-Based Reinforcement Learning in Image Editing A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 16782ab1-f999-4139-bbfe-0df663bba745 · inbound
Pairwise Preference Reward and Group-Based Diversity Enhancement for Superior Open-Ended Generation A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5104f631-dd80-4742-9611-63305d747dbf · inbound
Trust Region On-Policy Distillation A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization
Reference 197
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.