Pith. sign in

Paper Citation Record · LEDGER

A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2504.04950.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.04950 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:34:54.497908Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:46:14.356541Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6b797002-0f2c-4569-b19c-7217b5948dae · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization

Reference 157

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.751164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:eea083a1f4f526f28d0b2fbd947eedbff35e09bac4e5cccfb7f0b709d73f8a91

Observation 79be537c-ece5-4495-b294-07b595d9fd67 · inbound

Generative RLHF-V: Learning Principles from Multi-modal Human Preference cites this paper.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:54.497908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:54.497908Z digest=sha256:d19efd9200d721aa178bb6a6a6d99f5ff2454c46d6bb8bb06d909d82c2a00f40

Observation 8e894e89-aaa8-40d4-a55e-a065552fb817 · inbound

WebDancer: Towards Autonomous Information Seeking Agency cites this paper.

WebDancer: Towards Autonomous Information Seeking Agency A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:00.584149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:09:00.584149Z digest=sha256:771fa4d02dda80fe1d885ca47bbf6a9fc3f95ffd3b93a82ac7145780079f525d

Observation 972cb59f-2d79-474c-be90-22b1351106c5 · inbound

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier cites this paper.

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T04:39:37.841550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:39:37.841550Z digest=sha256:cf5510dee2e4b0a3c08af86cf742b6b459827900951bd3a14e8a225c74ec90a0

Observation 96ff8a98-aaa1-4426-84c9-7d72a24e8711 · inbound

RewardDance: Reward Scaling in Visual Generation cites this paper.

RewardDance: Reward Scaling in Visual Generation A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T20:09:01.240055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:09:01.240055Z digest=sha256:6a4bc3447d2981cc8fdaf933cfe14c9d7e9ef9ba7a92667afa3590a5df8ff7e0

Observation 64bc43cc-a96f-4ce5-b2fb-14d01e7ae83d · inbound

Voting with the Graph: Stable RLAIF via Topological Consistency Maximization cites this paper.

Voting with the Graph: Stable RLAIF via Topological Consistency Maximization A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T09:28:13.017326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:28:13.017326Z digest=sha256:9fe4bda3d28ea7a802d342ea4648bd44d0bc37845f19a8d32fa86fead526a061

Observation aa743bd7-6033-40e5-8722-9cbe0250b7bd · inbound

PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling cites this paper.

PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:20:49.249595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T03:15:57.744706Z digest=sha256:6a7933ac54e895838424ffeba80764d845e064e09777bcff9dea72cb94a0c0cf

Observation f7d6ff37-2f61-492e-a2ec-ca2b96a8755e · inbound

Leveraging Verifier-Based Reinforcement Learning in Image Editing cites this paper.

Leveraging Verifier-Based Reinforcement Learning in Image Editing A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:06:27.558499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T08:00:33.307429Z digest=sha256:7da3084a3174d5e595d082c531b524395bd65286350908503741bb077cfd06e5

Observation b758234d-2223-4943-955d-1e58f95674e2 · inbound

Leveraging Verifier-Based Reinforcement Learning in Image Editing cites this paper.

Leveraging Verifier-Based Reinforcement Learning in Image Editing A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-21T09:14:05.945999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T09:11:02.183133Z digest=sha256:d155ab4c394c72abbf711de51b585238e01e70b7a2aa07946b8f2ecb3b22a09f

Observation 16782ab1-f999-4139-bbfe-0df663bba745 · inbound

Pairwise Preference Reward and Group-Based Diversity Enhancement for Superior Open-Ended Generation cites this paper.

Pairwise Preference Reward and Group-Based Diversity Enhancement for Superior Open-Ended Generation A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T10:18:11.891575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T10:15:22.633882Z digest=sha256:2dd01123a4ea4f375acc9d5f9492f309275e80aa75638ddd4cdbc5798ad62a56

Observation 5104f631-dd80-4742-9611-63305d747dbf · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization

Reference 197

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:46:14.358178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:c6e2060d56223845814ec8cad23bc963d7e2a9dfd9824a57c3accd8b75416cca