Pith. sign in

Paper Citation Record · LEDGER

ReflectRM: Boosting Generative Reward Models via Self-Reflection within a Unified Judgment Framework

As of 4 August 2026, this Paper Citation Record lists 5 of 5 outbound references and 0 inbound Pith citation observations for arXiv:2604.07506.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.07506 v2

Coverage vector

measured 5 of 5 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T18:01:04.393062Z

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

5 of 5 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8100f572-e939-4952-a7cb-e28305aca3a1 · outbound

This paper cites Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains.

ReflectRM: Boosting Generative Reward Models via Self-Reflection within a Unified Judgment Framework Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:07:56.884714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:01:04.393062Z digest=sha256:257d4997e23a99a60edb0cae96e5e94700b30c5eff1a47276dda0eb11d7944ed

Observation 7a1cbc90-4fe2-4ab9-b500-0e68866dcaaa · outbound

This paper cites Generative Reward Models.

ReflectRM: Boosting Generative Reward Models via Self-Reflection within a Unified Judgment Framework Generative Reward Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:40:57.754795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:01:04.393062Z digest=sha256:ea4528879249d160265156f6ddb988495338714aa2b9e611665de9682616ee6c

Observation dfaab1f6-00f2-46c1-9490-45ea81b54b21 · outbound

This paper cites Seed1.5-thinking: Advancing superb reasoning models with reinforce- ment learning.

ReflectRM: Boosting Generative Reward Models via Self-Reflection within a Unified Judgment Framework Seed1.5-thinking: Advancing superb reasoning models with reinforce- ment learning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:40:57.715450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:01:04.393062Z digest=sha256:b393732a24ff49d829a8d1caec54d4ca784afd0663579e8c3abcbe027ddad31c

Observation a73a7e5e-3ddb-4fd7-9012-903b049b95d5 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

ReflectRM: Boosting Generative Reward Models via Self-Reflection within a Unified Judgment Framework Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:40:57.742179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:01:04.393062Z digest=sha256:3b614ffd4a9d7219ca189fd0e89d215f54c43e9788dca4eefdcd23deb6d23d8c

Observation d6d6597b-00ba-4141-861f-4ce5338722b3 · outbound

This paper cites Qwen3 Technical Report.

ReflectRM: Boosting Generative Reward Models via Self-Reflection within a Unified Judgment Framework Qwen3 Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:40:57.734395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:01:04.393062Z digest=sha256:6047b1f0a61920db7a134eed47316bd556a59d57e44786ec5e99a00c5bbe7aee

Pith citing papers

No inbound Pith citation observations are available.