Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T18:01:04.393062Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 5 of 5 outbound references and 0 inbound Pith citation observations for arXiv:2604.07506.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T18:01:04.393062Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
5 of 5 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8100f572-e939-4952-a7cb-e28305aca3a1 · outbound
ReflectRM: Boosting Generative Reward Models via Self-Reflection within a Unified Judgment Framework Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7a1cbc90-4fe2-4ab9-b500-0e68866dcaaa · outbound
ReflectRM: Boosting Generative Reward Models via Self-Reflection within a Unified Judgment Framework Generative Reward Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dfaab1f6-00f2-46c1-9490-45ea81b54b21 · outbound
ReflectRM: Boosting Generative Reward Models via Self-Reflection within a Unified Judgment Framework Seed1.5-thinking: Advancing superb reasoning models with reinforce- ment learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a73a7e5e-3ddb-4fd7-9012-903b049b95d5 · outbound
ReflectRM: Boosting Generative Reward Models via Self-Reflection within a Unified Judgment Framework Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d6d6597b-00ba-4141-861f-4ce5338722b3 · outbound
ReflectRM: Boosting Generative Reward Models via Self-Reflection within a Unified Judgment Framework Qwen3 Technical Report
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
No inbound Pith citation observations are available.