Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T18:12:55.726473Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2604.07484.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T18:12:55.726473Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
14 of 14 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a2c5f952-0625-401f-b6ee-d5d35e16f2ef · outbound
ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training Deep Think with Confidence
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 06d2ddd8-1325-45e9-89bb-dc941196c020 · outbound
ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 07d7b8a6-e0a2-4e6d-b4f6-98930e1cad61 · outbound
ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training Generative Reward Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5f890e66-3a40-4204-a50f-f731830dcea3 · outbound
ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 370d0ac6-6f41-4d36-bca1-4fd51f43450a · outbound
ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training Qwen3 Technical Report
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5ecaf77b-8d39-4576-b78f-eca6fc09c686 · outbound
ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c3c920da-d8a9-44c7-9d48-5b26fcb69339 · outbound
ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 54dd87a1-56bb-4821-a9cf-c777e9f2d5da · outbound
ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d4baa38a-00cf-4d09-ac61-44f62aa32209 · outbound
ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5ff1ec89-606a-46b0-a8fb-680d7a218b7c · outbound
ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training </Criterion> <Analysis> Response 1: Response 1 provides a broad overview of NHI systems and correctly explains the concepts of both top-down and bottom-up approaches in healthcare
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e0d1a515-5b34-46f1-b4ab-427baea06963 · outbound
ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0b054f66-50e0-4232-974a-5aacc5adf354 · outbound
ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8687569f-0162-4e09-a581-2367bd9fcc4a · outbound
ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1a0e2a63-45e1-4361-b2bd-dd5b142ed282 · outbound
ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training </Criterion> <Analysis> Response 1: Response 1 provides a comprehensive overview of the NHI system and explains both top-down and bottom-up approaches in the context of healthcare
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
No inbound Pith citation observations are available.