Pith. sign in

Paper Citation Record · LEDGER

WorldPM: Scaling Human Preference Modeling

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2505.10527.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.10527 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:49:45.723481Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T00:07:27.940098Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c6887f61-aa82-4a10-9f0e-16a07b78309b · inbound

RewardDance: Reward Scaling in Visual Generation cites this paper.

RewardDance: Reward Scaling in Visual Generation WorldPM: Scaling Human Preference Modeling

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T20:09:00.504202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:09:00.504202Z digest=sha256:44414b5190b5c19948bd2746569b761f1423e3c80f7528f6a1d0b6797ca80d50

Observation f423c084-1aae-4047-82a9-ee19963d9dc0 · inbound

AI Can Learn Scientific Taste cites this paper.

AI Can Learn Scientific Taste WorldPM: Scaling Human Preference Modeling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T18:14:49.961366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:14:49.961366Z digest=sha256:be4fbde615eb7557e9acdfce3a2c3689e34e181485975ac1df3c00d8acd1b320

Observation c7f984f2-8052-45cb-abbe-747d8cfe1a63 · inbound

Beyond Overlap Metrics: Rewarding Reasoning and Preferences for Faithful Multi-Role Dialogue Summarization cites this paper.

Beyond Overlap Metrics: Rewarding Reasoning and Preferences for Faithful Multi-Role Dialogue Summarization WorldPM: Scaling Human Preference Modeling

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:51:46.328497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T06:47:34.623340Z digest=sha256:66e16b2d12e7a595267b44953d542d0dd83a3f7a3d82be0d3d6363c6c479b197

Observation d8ebd339-9560-40c2-9808-0fb0e0b52c20 · inbound

Leveraging Verifier-Based Reinforcement Learning in Image Editing cites this paper.

Leveraging Verifier-Based Reinforcement Learning in Image Editing WorldPM: Scaling Human Preference Modeling

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:06:27.446483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T08:00:33.307429Z digest=sha256:40c9d0d92c67390468104bbe95bc23082e1ec553ee168978430e6ba7cc10ecef

Observation b54261bb-8481-4f67-9bc9-04a6944e3572 · inbound

Leveraging Verifier-Based Reinforcement Learning in Image Editing cites this paper.

Leveraging Verifier-Based Reinforcement Learning in Image Editing WorldPM: Scaling Human Preference Modeling

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-21T09:14:05.932676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T09:11:02.183133Z digest=sha256:a2e7a2df96c377d456f7595512dc00fa58e1ddab56646fd8039432d693e225bb

Observation 2619ad0c-fc3c-45c9-8374-4ed6ab95b580 · inbound

RewardHarness: Self-Evolving Agentic Post-Training cites this paper.

RewardHarness: Self-Evolving Agentic Post-Training WorldPM: Scaling Human Preference Modeling

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:31:23.995255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:9f2960817a41261befaf2b466b74bb77cf9a2330a80d4eff8926d4b3e7b51c07

Observation 7808e38c-29ba-4e74-a93b-4aad4a8ebd20 · inbound

Z-Reward: Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions cites this paper.

Z-Reward: Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions WorldPM: Scaling Human Preference Modeling

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:07:27.941638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T17:30:57.001021Z digest=sha256:47b0588b7b49c423ed26207c18a7f38de9d7f344a88225dbd89ff38c078327f7

Observation 2d25d118-c7e0-45f1-8ef5-4658c4a9af16 · inbound

Z-Reward: Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions cites this paper.

Z-Reward: Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions WorldPM: Scaling Human Preference Modeling

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-15T10:53:37.186361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T10:53:37.186361Z digest=sha256:0dbd76ca70d76dfed273514ee462a0f4e8e2afb9404a612dd98a36bc6f71539d

Observation e834e349-dfe1-4c35-b6a7-699bc4d1b514 · inbound

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction cites this paper.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction WorldPM: Scaling Human Preference Modeling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:45.723481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:49:45.723481Z digest=sha256:4cf4dc5a4f8a2ee738a528a7adb429b1ae4422b4e57c0d61eac17b6c19aac197