Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 43 inbound Pith citation observations for arXiv:2403.04642.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T15:11:08.137792Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-23T02:42:26.083324Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation e620d219-e4d2-4d99-b7cc-e5b38a53b4cf · inbound
Training Language Models to Self-Correct via Reinforcement Learning Teaching Large Language Models to Reason with Reinforcement Learning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5e742e69-9d5d-4e1b-ab04-905424c5eb95 · inbound
Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning Teaching Large Language Models to Reason with Reinforcement Learning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 81073105-9561-4640-9bdf-431d133e25bb · inbound
Training Large Language Models to Reason in a Continuous Latent Space Teaching Large Language Models to Reason with Reinforcement Learning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0d0b062e-1036-4c1f-b7fe-403edeb2d6d2 · inbound
Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Teaching Large Language Models to Reason with Reinforcement Learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 943e4545-6612-4952-a7c6-53c23682a60a · inbound
On the Emergence of Thinking in LLMs I: Searching for the Right Intuition Teaching Large Language Models to Reason with Reinforcement Learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c98a33f2-13cf-4e37-974c-51909915dc0d · inbound
Process Reward Models for LLM Agents: Practical Framework and Directions Teaching Large Language Models to Reason with Reinforcement Learning
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23390e2f-76d4-4c61-b3fb-2f047a4552c5 · inbound
Learning to Reason at the Frontier of Learnability Teaching Large Language Models to Reason with Reinforcement Learning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1a9aa297-dd75-445b-9d9f-0739788963a9 · inbound
Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space Teaching Large Language Models to Reason with Reinforcement Learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49fc1608-b320-4489-bbfa-ae4d409b03ac · inbound
HCRMP: A LLM-Hinted Contextual Reinforcement Learning Framework for Autonomous Driving Teaching Large Language Models to Reason with Reinforcement Learning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d19f54ef-db2a-4b5b-867e-f353c9648dd4 · inbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Teaching Large Language Models to Reason with Reinforcement Learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2120d7eb-4477-465d-8dea-891dc5dd9fb2 · inbound
Learning to Select In-Context Demonstration Preferred by Large Language Model Teaching Large Language Models to Reason with Reinforcement Learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b280998b-7aeb-4866-9259-74fd9d347658 · inbound
LoVeC: Reinforcement Learning for Better Verbalized Confidence in Long-Form Generations Teaching Large Language Models to Reason with Reinforcement Learning
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7e35f750-3716-42a7-88b1-4dbb4526b9d0 · inbound
SCOUT: Teaching Pre-trained Language Models to Enhance Reasoning via Flow Chain-of-Thought Teaching Large Language Models to Reason with Reinforcement Learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e135594d-1c46-4d77-8a61-9b2fe77ab4d0 · inbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Teaching Large Language Models to Reason with Reinforcement Learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a53a291c-66f3-4219-8f0f-dc8fd58cf89f · inbound
Truly Self-Improving Agents Require Intrinsic Metacognitive Learning Teaching Large Language Models to Reason with Reinforcement Learning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28fa47da-3602-46cb-bef4-1b7022a8e596 · inbound
RePO: Replay-Enhanced Policy Optimization Teaching Large Language Models to Reason with Reinforcement Learning
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c3bf2f5-bbaa-4cbe-aa0d-b4f53a75aaf8 · inbound
Intent Factored Generation: Unleashing the Diversity in Your Language Model Teaching Large Language Models to Reason with Reinforcement Learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31b4fc56-00b6-434c-9a07-5a53afa3bab2 · inbound
RAST: Reasoning Activation in LLMs via Small-model Transfer Teaching Large Language Models to Reason with Reinforcement Learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30cebf1b-9edf-4c93-a4c6-0d06a81f86d1 · inbound
Learning Efficient Robotic Garment Manipulation with Standardization Teaching Large Language Models to Reason with Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e441468-f270-49bd-8275-064d1d46a9b2 · inbound
Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models Teaching Large Language Models to Reason with Reinforcement Learning
Reference 139
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32d4999e-f949-4bcf-ba50-b84d03b796ca · inbound
When LLMs Copy to Think: Uncovering Copy-Guided Attacks in Reasoning LLMs Teaching Large Language Models to Reason with Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ee5efad-7c3b-49cb-b87e-062e32774c0e · inbound
Med-R$^3$: Enhancing Medical Retrieval-Augmented Reasoning of LLMs via Progressive Reinforcement Learning Teaching Large Language Models to Reason with Reinforcement Learning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cccdf65-b86f-4eaa-8584-d24c5534aadb · inbound
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Teaching Large Language Models to Reason with Reinforcement Learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d1ae46cd-7879-4233-b5ff-171a85413323 · inbound
Beyond Memorization: Extending Reasoning Depth with Recurrence, Memory and Test-Time Compute Scaling Teaching Large Language Models to Reason with Reinforcement Learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7c481f30-7042-4931-9867-16411b6c3d47 · inbound
TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Teaching Large Language Models to Reason with Reinforcement Learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86d251eb-e19e-42fc-aca2-e7b41d675a63 · inbound
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Teaching Large Language Models to Reason with Reinforcement Learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f7fc2164-4bbc-444f-8172-de9b9d233944 · inbound
Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning Teaching Large Language Models to Reason with Reinforcement Learning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f81f4d0a-b7c7-4490-9bbc-57456cdccdf8 · inbound
SeLaR: Selective Latent Reasoning in Large Language Models Teaching Large Language Models to Reason with Reinforcement Learning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3fa2893a-9d5a-40d5-a0d3-b56879453ce5 · inbound
LiFT: Does Instruction Fine-Tuning Improve In-Context Learning for Longitudinal Modelling by Large Language Models? Teaching Large Language Models to Reason with Reinforcement Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6a38efa3-6279-45c4-b446-85466f863fef · inbound
Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding Teaching Large Language Models to Reason with Reinforcement Learning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 94c2ee82-3972-41d2-9102-f160f9404b65 · inbound
Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning Teaching Large Language Models to Reason with Reinforcement Learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d02a8372-ccfc-427e-8cbc-d60d3256639e · inbound
Logic-Regularized Verifier Elicits Reasoning from LLMs Teaching Large Language Models to Reason with Reinforcement Learning
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dda6e555-6545-49a4-8474-436f8a5da3d9 · inbound
NoisyCoconut: Counterfactual Consensus via Latent Space Reasoning Teaching Large Language Models to Reason with Reinforcement Learning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d6d976ea-a5e7-40f6-8fe6-1d56bae82ff5 · inbound
Epistemic Uncertainty for Test-Time Discovery Teaching Large Language Models to Reason with Reinforcement Learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c81a34d1-d755-4ac4-914f-1d137bb1437d · inbound
When Reasoning Traces Become Performative: Step-Level Evidence that Chain-of-Thought Is an Imperfect Oversight Channel Teaching Large Language Models to Reason with Reinforcement Learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9a1e33e8-a197-49ed-ae07-e9a23d3c2147 · inbound
From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI Teaching Large Language Models to Reason with Reinforcement Learning
Reference 139
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40d461bd-51fb-4411-a91d-074baac150f9 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Teaching Large Language Models to Reason with Reinforcement Learning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15e17f8c-3d1d-4117-a752-b61ea7913795 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Teaching Large Language Models to Reason with Reinforcement Learning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58f57867-c9d6-4227-84b9-37c8ca272be2 · inbound
Structured Synthetic Reasoning Data for Arithmetic Fine-Tuning of Small Language Models Teaching Large Language Models to Reason with Reinforcement Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3178038-8e0d-482d-b8e8-4a088c93762d · inbound
LeAct: Learning to Reason from Expert Actions Teaching Large Language Models to Reason with Reinforcement Learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46f0c937-c1c1-42b7-b42e-7b9169cd431c · inbound
AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning Teaching Large Language Models to Reason with Reinforcement Learning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 639afc79-1014-4fc7-896c-13a2914bd6db · inbound
RepoReasoner: Evaluating Repository-Level Code Reasoning Ability of Long-Context Language Models Teaching Large Language Models to Reason with Reinforcement Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b9c5573-8ad2-4bc5-88c1-ea212ae184dc · inbound
Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details Teaching Large Language Models to Reason with Reinforcement Learning
Reference 193
Source-reported events for the cited work
Unavailable: canonical work link unavailable.