Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2509.20357.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-01T01:13:13.329619Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 98ccfcce-41d7-4632-acc8-ffe15a349367 · inbound
A Survey of Reinforcement Learning for Large Reasoning Models Language Models that Think, Chat Better
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e3769ba7-2e51-41fe-9a9d-db38190c9f93 · inbound
Learning to Pose Problems: Reasoning-Driven and Solver-Adaptive Data Synthesis Language Models that Think, Chat Better
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 67c5f086-2392-481f-9936-c73202b9170d · inbound
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Language Models that Think, Chat Better
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a589ca64-2c63-4b47-ae6e-637a78f008eb · inbound
SUPERNOVA: Eliciting General Reasoning in LLMs with Reinforcement Learning on Natural Instructions Language Models that Think, Chat Better
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 74af8204-2fa4-4a52-b14d-0f8b2908afa6 · inbound
CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts Language Models that Think, Chat Better
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation fd9b56e7-cd87-4440-bbe6-e2af9be40c77 · inbound
CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts Language Models that Think, Chat Better
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9bb049ef-a430-474c-8ed1-e0ac8aae0fa2 · inbound
Wait, am I Being Fair? Characterizing Deductive Stereotyping and Mitigating It with Fair-GCG Language Models that Think, Chat Better
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.