Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2503.23829.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T17:57:52.481344Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 7745fa08-d56a-4cc0-aa9c-f06996f5d17d · inbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0b6983d6-f37e-4f3d-acf4-257c3df171fc · inbound
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 500c22dd-a511-42f3-b0b0-879be3b0da6b · inbound
Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21974947-640f-41f2-956a-4a4e42720073 · inbound
DeepSearch: Overcome the Bottleneck of Reinforcement Learning with Verifiable Rewards via Monte Carlo Tree Search Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5678ad83-14b1-4b82-b552-9a4d16097830 · inbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6d6e0de-6812-4916-84ef-98aeff15b891 · inbound
Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 49959a33-24e4-4b0d-861a-15ade233c683 · inbound
DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 553adc0d-e76a-4293-a47a-83b6d66855ed · inbound
CURE-Med: Curriculum-Informed Reinforcement Learning for Multilingual Medical Reasoning Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 53947a8e-5330-48d8-88c6-f47d8edd3061 · inbound
Specificity-aware reinforcement learning for fine-grained open-world classification Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1570ade8-c9af-4dc6-bf9f-63b133f018d8 · inbound
Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 106
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3b7189d9-5807-4d9e-aa2f-658142edfee5 · inbound
LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a62ac229-1b51-47de-b77c-2b9ccaf8fd5b · inbound
Interpretable Electrophysiological Features of Resting-State EEG Capture Cortical Network Dynamics in Parkinsons Disease Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbb70e96-f794-4e69-a379-9ec69c09355e · inbound
Trust Your Memory: Verifiable Control of Smart Homes through Reinforcement Learning with Multi-dimensional Rewards Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 369ad831-a096-4da0-8abc-13cba5e99796 · inbound
V-tableR1: Process-Supervised Multimodal Table Reasoning with Critic-Guided Policy Optimization Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5803ac88-fbb8-4ee0-bba0-c31b7b85402d · inbound
Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoring Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b3786b5d-2427-4e79-8e10-970e17cfd6c1 · inbound
Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoring Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1dd83977-cef3-48d3-8789-8ceb9697ab7a · inbound
Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 105
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6e20c37c-f7c7-41eb-ab24-00109f9543f2 · inbound
M2A: Synergizing Mathematical and Agentic Reasoning in Large Language Models Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c2ca204c-6241-40e0-afc0-1f11f80191cb · inbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1f31b556-3816-4be0-bc71-8924d6b7bc72 · inbound
Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bfe80b35-bcae-4429-a91f-e2caf0a75a60 · inbound
CAST: Non-Privileged Clipped Asymmetric Self-Teaching with Advantage Flipping for GRPO Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f305aa37-c9eb-46a2-9aa9-e80b02291d7d · inbound
CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0569465f-718a-4fb8-91cd-7afc15e533d5 · inbound
GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 69cbd603-7104-40b4-853a-246a243078a3 · inbound
CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 55788ff5-3d72-4a1a-888c-f2014b89a2f6 · inbound
Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1bb171b7-5dd2-4e15-b5ba-78bbaeb0644f · inbound
PointVG-R: Internalizing Geometric Reasoning in MLLMs for Precise Pointing Localization via Visual Chain of Thought Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dcf008c4-a9b1-4f38-9e35-7bf8114e9a0c · inbound
BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 61f6e4fe-1eae-489a-950e-6014c1940867 · inbound
Benchmarking Frontier LLMs on Arabic Cultural and Sociolinguistic Knowledge: A Cross-Evaluation Framework with Human SME Ground Truth Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0c369fed-3427-4cc4-84b1-b2f2437d9bc2 · inbound
Text-Driven 3D Indoor Scene Synthesis in Non-Manhattan Environments Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.