Pith. sign in

Paper Citation Record · LEDGER

Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2503.23829.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.23829 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T17:57:52.481344Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7745fa08-d56a-4cc0-aa9c-f06996f5d17d · inbound

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks cites this paper.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:52:14.093197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:4eba7e0be6b1a4f1ca04486c3ab91864a161f79156922f40f7dfa44fd20c68ff

Observation 0b6983d6-f37e-4f3d-acf4-257c3df171fc · inbound

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains cites this paper.

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:07:56.851371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T06:07:56.678339Z digest=sha256:20e29ec51377b807b8c2d5f29e361295b9efb518591b416754ff43baa7658a2f

Observation 500c22dd-a511-42f3-b0b0-879be3b0da6b · inbound

Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs cites this paper.

Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T17:57:52.481344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:57:52.481344Z digest=sha256:81e2631e6a72dd59caa6a0815a883481c6935acfd7b9f7b9a69f8e225ba4d93d

Observation 21974947-640f-41f2-956a-4a4e42720073 · inbound

DeepSearch: Overcome the Bottleneck of Reinforcement Learning with Verifiable Rewards via Monte Carlo Tree Search cites this paper.

DeepSearch: Overcome the Bottleneck of Reinforcement Learning with Verifiable Rewards via Monte Carlo Tree Search Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:12:35.834905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T12:12:25.437344Z digest=sha256:a53929a2730cc388f8236f82cadae34d5e6f6d3887c0d1c0b7d14cbee5a5f413

Observation 5678ad83-14b1-4b82-b552-9a4d16097830 · inbound

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts cites this paper.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:48.485814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:48.485814Z digest=sha256:c47e828a65fe88960506d222f53a24609b20fa54865f1f6c64286c4c95417a7b

Observation f6d6e0de-6812-4916-84ef-98aeff15b891 · inbound

Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning cites this paper.

Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:21:04.360338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T07:20:01.505216Z digest=sha256:5a7b8a31c5b3467b486d39bc913c30095903616f4853215725f16d7bfa425aec

Observation 49959a33-24e4-4b0d-861a-15ade233c683 · inbound

DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training cites this paper.

DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:48:50.980608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T01:46:21.744857Z digest=sha256:3bdd2858c87e0bbb0e6bccc4f09f214db0fea19620db7d1c78bc0602b203af95

Observation 553adc0d-e76a-4293-a47a-83b6d66855ed · inbound

CURE-Med: Curriculum-Informed Reinforcement Learning for Multilingual Medical Reasoning cites this paper.

CURE-Med: Curriculum-Informed Reinforcement Learning for Multilingual Medical Reasoning Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:20:57.559827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T13:20:49.919833Z digest=sha256:9aeda7a38dd2eccda18ff293bea424c2313a70654a9f3487069dd4714bd310e9

Observation 53947a8e-5330-48d8-88c6-f47d8edd3061 · inbound

Specificity-aware reinforcement learning for fine-grained open-world classification cites this paper.

Specificity-aware reinforcement learning for fine-grained open-world classification Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:46:17.815819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T16:44:47.866126Z digest=sha256:072082e4401c3ab9d85f5b6272a230aa009eeac6c8ced00d96afaa3a5a64e244

Observation 1570ade8-c9af-4dc6-bf9f-63b133f018d8 · inbound

Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation cites this paper.

Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 106

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:51:40.644170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T21:51:10.972744Z digest=sha256:0e74d073ab31b7fc00bf0203e0f0f723b1af6bc753ae5a0b66815c7bc68cc8f1

Observation 3b7189d9-5807-4d9e-aa2f-658142edfee5 · inbound

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks cites this paper.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:f781a3c36a8a69fad8bf3f412f40fb2ed99e96c15cae851b709a705b1c80ebdf

Observation a62ac229-1b51-47de-b77c-2b9ccaf8fd5b · inbound

Interpretable Electrophysiological Features of Resting-State EEG Capture Cortical Network Dynamics in Parkinsons Disease cites this paper.

Interpretable Electrophysiological Features of Resting-State EEG Capture Cortical Network Dynamics in Parkinsons Disease Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T14:23:49.346727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:23:49.346727Z digest=sha256:b5e6a23cc2ab536d00cb0388d220415ae16ea45cfd8a6ca03b5fa0ba5d15140d

Observation cbb70e96-f794-4e69-a379-9ec69c09355e · inbound

Trust Your Memory: Verifiable Control of Smart Homes through Reinforcement Learning with Multi-dimensional Rewards cites this paper.

Trust Your Memory: Verifiable Control of Smart Homes through Reinforcement Learning with Multi-dimensional Rewards Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:06:01.172294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:49:00.343580Z digest=sha256:ef9005c5b42f4196d36d9d8abd5152005414a780694ae845c8501ca24257cddd

Observation 369ad831-a096-4da0-8abc-13cba5e99796 · inbound

V-tableR1: Process-Supervised Multimodal Table Reasoning with Critic-Guided Policy Optimization cites this paper.

V-tableR1: Process-Supervised Multimodal Table Reasoning with Critic-Guided Policy Optimization Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:01:04.148354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T23:48:32.613988Z digest=sha256:fc602bc61af40412d9520fd5a69465cb4bb119189320896b78004fbae86335af

Observation 5803ac88-fbb8-4ee0-bba0-c31b7b85402d · inbound

Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoring cites this paper.

Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoring Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T19:35:39.078929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T19:33:35.690030Z digest=sha256:37504de3aa3e37fe84ce9246100d50797da32e6cf6341c99db79691c95efe6a7

Observation b3786b5d-2427-4e79-8e10-970e17cfd6c1 · inbound

Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoring cites this paper.

Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoring Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T02:15:52.434975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T01:56:43.707252Z digest=sha256:f4fc52fb140ab32a954b5817dfccb3b9828418887bc8b8b065284cca866490f2

Observation 1dd83977-cef3-48d3-8789-8ceb9697ab7a · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:49.250698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:86f4376faeb8d22dd693d36fe07bc55be856a280ca6790be37ecc751bf6d1d66

Observation 6e20c37c-f7c7-41eb-ab24-00109f9543f2 · inbound

M2A: Synergizing Mathematical and Agentic Reasoning in Large Language Models cites this paper.

M2A: Synergizing Mathematical and Agentic Reasoning in Large Language Models Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:22.662442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:43:50.384980Z digest=sha256:f0fc9b24f4918ceb440f6fdda75d8f324247564ea8f9b6ebe16491d7e26e0316

Observation c2ca204c-6241-40e0-afc0-1f11f80191cb · inbound

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents cites this paper.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:38:05.525894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:14de9f1f5ba8d62842f0d7931d47332dd3127fc147b7ae088fc788a42bcbd151

Observation 1f31b556-3816-4be0-bc71-8924d6b7bc72 · inbound

Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination cites this paper.

Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:12:47.109855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T23:08:05.597810Z digest=sha256:6851576f84cb8727317ff980542da455f60147a7cf579cab811a4e1567f1ac33

Observation bfe80b35-bcae-4429-a91f-e2caf0a75a60 · inbound

CAST: Non-Privileged Clipped Asymmetric Self-Teaching with Advantage Flipping for GRPO cites this paper.

CAST: Non-Privileged Clipped Asymmetric Self-Teaching with Advantage Flipping for GRPO Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:26:01.064488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T22:25:04.067339Z digest=sha256:980c8661949e693b3425e456ef8970915eafe210112e1fe0712594c959efe81a

Observation f305aa37-c9eb-46a2-9aa9-e80b02291d7d · inbound

CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts cites this paper.

CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:02:33.947566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T19:01:14.340754Z digest=sha256:0f3b9794759edd81294a938b10fd5d1569c63eabf4be3af1276a9189652cd3c9

Observation 0569465f-718a-4fb8-91cd-7afc15e533d5 · inbound

GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling cites this paper.

GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T06:06:40.787545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T07:45:43.320339Z digest=sha256:2ceb2a2a13ba4de1ab2c63e3a6b7f5590a2b6c3c28f9d0ec545ca0075a7e2218

Observation 69cbd603-7104-40b4-853a-246a243078a3 · inbound

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning cites this paper.

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:17:29.019297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T17:21:38.543724Z digest=sha256:27d04989658931b846458840777fccc95f2e98921ea66d8fdcf0082aef011ef5

Observation 55788ff5-3d72-4a1a-888c-f2014b89a2f6 · inbound

Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views cites this paper.

Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:39:45.338242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T08:38:46.044079Z digest=sha256:c123ee7374ab114fe51e479533971a7ae3ea185bfad2b3d4c6511d653b2662b4

Observation 1bb171b7-5dd2-4e15-b5ba-78bbaeb0644f · inbound

PointVG-R: Internalizing Geometric Reasoning in MLLMs for Precise Pointing Localization via Visual Chain of Thought cites this paper.

PointVG-R: Internalizing Geometric Reasoning in MLLMs for Precise Pointing Localization via Visual Chain of Thought Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:19:57.794350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T00:46:17.094339Z digest=sha256:4bd53d7ad8a9d73138e8a5844ea869a667b1066def50e062ad8e035d834650bd

Observation dcf008c4-a9b1-4f38-9e35-7bf8114e9a0c · inbound

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards cites this paper.

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T12:44:40.146628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-30T10:07:39.554999Z digest=sha256:ce6770fc78ccc1729d7f7d167aa72d44585202f2feb4d13b00bc9dfc838cc63b

Observation 61f6e4fe-1eae-489a-950e-6014c1940867 · inbound

Benchmarking Frontier LLMs on Arabic Cultural and Sociolinguistic Knowledge: A Cross-Evaluation Framework with Human SME Ground Truth cites this paper.

Benchmarking Frontier LLMs on Arabic Cultural and Sociolinguistic Knowledge: A Cross-Evaluation Framework with Human SME Ground Truth Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:27:18.482991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T19:26:21.833145Z digest=sha256:2eba16d4646883b6657ed5fc8d44402e552155cb239ab7fdcdcbb34a00a6c1d5

Observation 0c369fed-3427-4cc4-84b1-b2f2437d9bc2 · inbound

Text-Driven 3D Indoor Scene Synthesis in Non-Manhattan Environments cites this paper.

Text-Driven 3D Indoor Scene Synthesis in Non-Manhattan Environments Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T13:28:18.294461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-07-03T13:21:06.380698Z digest=sha256:5a65d1a8e7a588be52205753621768eec03c5bb70ce23065b19c1224989b2553