Pith. sign in

Paper Citation Record · LEDGER

Teaching Large Language Models to Reason with Reinforcement Learning

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 55 inbound Pith citation observations for arXiv:2403.04642.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.04642 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 55 of 55 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:18:12.224876Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T02:42:26.083324Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e620d219-e4d2-4d99-b7cc-e5b38a53b4cf · inbound

Training Language Models to Self-Correct via Reinforcement Learning cites this paper.

Training Language Models to Self-Correct via Reinforcement Learning Teaching Large Language Models to Reason with Reinforcement Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:04:10.444859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-17T12:04:10.210508Z digest=sha256:dafa13d29834024f31f83bef279e348c32ddd546a75eb153bf10bf61e0c6f9f9

Observation 5e742e69-9d5d-4e1b-ab04-905424c5eb95 · inbound

Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning cites this paper.

Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning Teaching Large Language Models to Reason with Reinforcement Learning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:42:19.179140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T01:42:19.004468Z digest=sha256:21a79825f191aa6cc04501c2604d14df633b1a7f4ed470c128232c1c19e98588

Observation 0a875894-36eb-4c88-9ac9-491c91e8e7e4 · inbound

ROSE: A Reward-Oriented Data Selection Framework for LLM Task-Specific Instruction Tuning cites this paper.

ROSE: A Reward-Oriented Data Selection Framework for LLM Task-Specific Instruction Tuning Teaching Large Language Models to Reason with Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T05:13:41.429873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:13:41.429873Z digest=sha256:744fbc5262544ed3740466a0f0f41df27f7792dd85397c7fac24df7e765bc73e

Observation 71f6d78d-5621-4854-a1ab-cd6d43502736 · inbound

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models cites this paper.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Teaching Large Language Models to Reason with Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.615413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.615413Z digest=sha256:4202550b069d5f89a2c70766a1606482c58184600fc3610f7e4925efcd406d5c

Observation cf63e4ed-bb18-479c-aa5d-50872c9e5c05 · inbound

Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models cites this paper.

Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models Teaching Large Language Models to Reason with Reinforcement Learning

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T22:57:01.684212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:57:01.684212Z digest=sha256:b57cb938a569aa7f150ae5227dd54f3015edd1f71d118cb2373fddc847748fa2

Observation 81073105-9561-4640-9bdf-431d133e25bb · inbound

Training Large Language Models to Reason in a Continuous Latent Space cites this paper.

Training Large Language Models to Reason in a Continuous Latent Space Teaching Large Language Models to Reason with Reinforcement Learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:29:05.706544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T10:29:05.384381Z digest=sha256:bc3e1b10843592ad52f5523c8d5f78619ac547480fdb9fb160167e85d3141d18

Observation ba1be8f6-1e9c-46e5-9890-b523a464221c · inbound

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners cites this paper.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Teaching Large Language Models to Reason with Reinforcement Learning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:00.359505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:00.359505Z digest=sha256:5f54152af245fc215ec9febe7f139fa774a6b3c18381aac9d2deb002729478ed

Observation f8af6422-b7ed-4ed2-9ca0-443a8db8c5e7 · inbound

Aviary: training language agents on challenging scientific tasks cites this paper.

Aviary: training language agents on challenging scientific tasks Teaching Large Language Models to Reason with Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:33.546798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:33.546798Z digest=sha256:b9b578399f5bb40c16a4933795252468f2ed35fd3af27e980f86e6772d72d848

Observation d351d8bb-4310-4a0b-879d-69e4d1b00f30 · inbound

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos cites this paper.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Teaching Large Language Models to Reason with Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:19.915886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:19.915886Z digest=sha256:038ef7bb1b4146c192af46ba095b72018ba1ac8da3f1ae5a8a49112b10f0b8d8

Observation b3f98590-752b-423a-adb8-a0d733587b04 · inbound

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step cites this paper.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Teaching Large Language Models to Reason with Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.309279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.309279Z digest=sha256:7aa8c71fe85e135b3ff0255aa60dc661be03bd9e26d9e88e9a97478df32ac6ff

Observation 5844cb03-48df-4753-95b5-fc08356f6552 · inbound

Reinforcement Learning for Long-Horizon Interactive LLM Agents cites this paper.

Reinforcement Learning for Long-Horizon Interactive LLM Agents Teaching Large Language Models to Reason with Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T14:56:00.206362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:56:00.206362Z digest=sha256:41857858aa3dba7592ff1889addcfac42e159dd65a35ee428f757f107867b929

Observation 0d0b062e-1036-4c1f-b7fe-403edeb2d6d2 · inbound

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning cites this paper.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Teaching Large Language Models to Reason with Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.137792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.137792Z digest=sha256:5c744574d6cceb473139aa4e2b289689f850d5938ebe1ffd0eede3e23bef4f66

Observation 943e4545-6612-4952-a7c6-53c23682a60a · inbound

On the Emergence of Thinking in LLMs I: Searching for the Right Intuition cites this paper.

On the Emergence of Thinking in LLMs I: Searching for the Right Intuition Teaching Large Language Models to Reason with Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:53.357179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:25:53.357179Z digest=sha256:6a069f5fbb3bd756aeae0b2bb2ded09623a9dbe5e230572c73cc54eadcc2784a

Observation c98a33f2-13cf-4e37-974c-51909915dc0d · inbound

Process Reward Models for LLM Agents: Practical Framework and Directions cites this paper.

Process Reward Models for LLM Agents: Practical Framework and Directions Teaching Large Language Models to Reason with Reinforcement Learning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:38.006409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:38.006409Z digest=sha256:8dec9b6fe19ebb5647445182346221add8f9625009c2f60c7da6b1afb450b5bf

Observation 23390e2f-76d4-4c61-b3fb-2f047a4552c5 · inbound

Learning to Reason at the Frontier of Learnability cites this paper.

Learning to Reason at the Frontier of Learnability Teaching Large Language Models to Reason with Reinforcement Learning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.085841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:d745097bbdca7fa4ce5ea3a2253595d5a40d7733cca1fc34a72398e3ba9cc347

Observation 109e504e-7317-40f8-925c-907e65c62b00 · inbound

Aligning Constraint Generation with Design Intent in Parametric CAD cites this paper.

Aligning Constraint Generation with Design Intent in Parametric CAD Teaching Large Language Models to Reason with Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T12:18:12.224876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:18:12.224876Z digest=sha256:c5f0db0f6b4c645f046a04a0ce852ec7ef9a0163585b89ea0aa9d29dae7b235f

Observation 004b573a-7359-4ea7-a06b-d8ef1fa19cf8 · inbound

Everything You Wanted to Know About LLM-based Vulnerability Detection But Were Afraid to Ask cites this paper.

Everything You Wanted to Know About LLM-based Vulnerability Detection But Were Afraid to Ask Teaching Large Language Models to Reason with Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T12:11:39.537883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:11:39.537883Z digest=sha256:5b604e5871d3c8b9e27972fac3b3b24b69c5a57b5e90b3caf63589a131c20142

Observation a798774f-f250-4210-9b7a-9fbe4aafbeff · inbound

Multi-agent Embodied AI: Advances and Future Directions cites this paper.

Multi-agent Embodied AI: Advances and Future Directions Teaching Large Language Models to Reason with Reinforcement Learning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T23:16:15.524069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:16:15.524069Z digest=sha256:2ad50e10e549f052c30fddb248bd8a2765fdb2c80a7f49572e9b5e2eb35a72e3

Observation c3b46dae-82da-4cce-ab08-518e97aea87d · inbound

EfficientLLM: Efficiency in Large Language Models cites this paper.

EfficientLLM: Efficiency in Large Language Models Teaching Large Language Models to Reason with Reinforcement Learning

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:35.115307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:35.115307Z digest=sha256:d33ff3687234ba05cf1158f7c37dc7ccb2a6da6c281a75203ec4a315c6dfbf70

Observation 1a9aa297-dd75-445b-9d9f-0739788963a9 · inbound

Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space cites this paper.

Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space Teaching Large Language Models to Reason with Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:43.520186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:43.520186Z digest=sha256:6e1c2c6a94ded4941ca98cf4bd702d838582d850ab38afa8eb68eb461fc99271

Observation 49fc1608-b320-4489-bbfa-ae4d409b03ac · inbound

HCRMP: A LLM-Hinted Contextual Reinforcement Learning Framework for Autonomous Driving cites this paper.

HCRMP: A LLM-Hinted Contextual Reinforcement Learning Framework for Autonomous Driving Teaching Large Language Models to Reason with Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:13:43.610133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:13:43.610133Z digest=sha256:d15825ed25465487ee6bf416f7d7fa54b97bb908712d2721625affa181392b0f

Observation d19f54ef-db2a-4b5b-867e-f353c9648dd4 · inbound

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models cites this paper.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Teaching Large Language Models to Reason with Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:53.173407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:53.173407Z digest=sha256:26f25391c5a21e6224095ed4f85ac51317f6cfc8cd9f5dbe0e7eb64267828f9c

Observation 2120d7eb-4477-465d-8dea-891dc5dd9fb2 · inbound

Learning to Select In-Context Demonstration Preferred by Large Language Model cites this paper.

Learning to Select In-Context Demonstration Preferred by Large Language Model Teaching Large Language Models to Reason with Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:22.127164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:09:22.127164Z digest=sha256:ec402f64ba65f56bba9add0a1e8b58b08607ce7361296f411e3cca4b8536b441

Observation b280998b-7aeb-4866-9259-74fd9d347658 · inbound

LoVeC: Reinforcement Learning for Better Verbalized Confidence in Long-Form Generations cites this paper.

LoVeC: Reinforcement Learning for Better Verbalized Confidence in Long-Form Generations Teaching Large Language Models to Reason with Reinforcement Learning

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:47:17.959953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T12:43:51.019983Z digest=sha256:5c5284ae1ee73f40d47ab9920ef823d0d6b78b69092fd2a491982a7d7678cc74

Observation 7e35f750-3716-42a7-88b1-4dbb4526b9d0 · inbound

SCOUT: Teaching Pre-trained Language Models to Enhance Reasoning via Flow Chain-of-Thought cites this paper.

SCOUT: Teaching Pre-trained Language Models to Enhance Reasoning via Flow Chain-of-Thought Teaching Large Language Models to Reason with Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:59.544322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:59.544322Z digest=sha256:63536d7689c19bdc451fdbf35568814059e74b55ecff148dcd1de0a9f160f236

Observation e135594d-1c46-4d77-8a61-9b2fe77ab4d0 · inbound

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning cites this paper.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Teaching Large Language Models to Reason with Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:28.455709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:28.455709Z digest=sha256:138d0494b9488e7f0acf4bef6a9aef758e7d9e187672991f7b94c9737b11c57a

Observation a53a291c-66f3-4219-8f0f-dc8fd58cf89f · inbound

Truly Self-Improving Agents Require Intrinsic Metacognitive Learning cites this paper.

Truly Self-Improving Agents Require Intrinsic Metacognitive Learning Teaching Large Language Models to Reason with Reinforcement Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:19.442762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:28:19.442762Z digest=sha256:998133272e163cf806ea29b31ca9d2a32015feff1adc84d35129e5563155d298

Observation 28fa47da-3602-46cb-bef4-1b7022a8e596 · inbound

RePO: Replay-Enhanced Policy Optimization cites this paper.

RePO: Replay-Enhanced Policy Optimization Teaching Large Language Models to Reason with Reinforcement Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T04:53:56.348495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:53:56.348495Z digest=sha256:0f4798b9ee9fef803d0ba1b8bbdef3d5eb9c1cb2457f65feaba2d6f086e4a75a

Observation 1c3bf2f5-bbaa-4cbe-aa0d-b4f53a75aaf8 · inbound

Intent Factored Generation: Unleashing the Diversity in Your Language Model cites this paper.

Intent Factored Generation: Unleashing the Diversity in Your Language Model Teaching Large Language Models to Reason with Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:47:42.354458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:47:42.354458Z digest=sha256:7c88355ceaefb4e97d334c7846177a660651cd79377736d5b04146d6631b1d66

Observation 31b4fc56-00b6-434c-9a07-5a53afa3bab2 · inbound

RAST: Reasoning Activation in LLMs via Small-model Transfer cites this paper.

RAST: Reasoning Activation in LLMs via Small-model Transfer Teaching Large Language Models to Reason with Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:53.893423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:53.893423Z digest=sha256:e14bda31beff3763e26d2b1353e48084de8fc0f2af8e5eba8dc386718891697b

Observation 30cebf1b-9edf-4c93-a4c6-0d06a81f86d1 · inbound

Learning Efficient Robotic Garment Manipulation with Standardization cites this paper.

Learning Efficient Robotic Garment Manipulation with Standardization Teaching Large Language Models to Reason with Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:00.026680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:00.026680Z digest=sha256:4799a7f855bd4be2b180ef7749bfdbd1ea1ccf0d8c5817f5f7a6692f80f25547

Observation 1e441468-f270-49bd-8275-064d1d46a9b2 · inbound

Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models cites this paper.

Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models Teaching Large Language Models to Reason with Reinforcement Learning

Reference 139

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:04.082222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:04.082222Z digest=sha256:d72ea7f1d4ebb626de5158ece214605c42f235a63d7026e1a4b0811b66bfa02f

Observation 32d4999e-f949-4bcf-ba50-b84d03b796ca · inbound

When LLMs Copy to Think: Uncovering Copy-Guided Attacks in Reasoning LLMs cites this paper.

When LLMs Copy to Think: Uncovering Copy-Guided Attacks in Reasoning LLMs Teaching Large Language Models to Reason with Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:07:27.843871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:07:27.843871Z digest=sha256:30279b79eac760b62087c5835e38f9c11b33d7109653978463aa06dceb5aa56b

Observation 2ee5efad-7c3b-49cb-b87e-062e32774c0e · inbound

Med-R$^3$: Enhancing Medical Retrieval-Augmented Reasoning of LLMs via Progressive Reinforcement Learning cites this paper.

Med-R$^3$: Enhancing Medical Retrieval-Augmented Reasoning of LLMs via Progressive Reinforcement Learning Teaching Large Language Models to Reason with Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T10:44:26.838066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:44:26.838066Z digest=sha256:3d6b49bcb9442d288a9d8da9b7289ed98fb89f8e60b43a53671f9424e6d163f6

Observation 4cccdf65-b86f-4eaa-8584-d24c5534aadb · inbound

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization cites this paper.

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Teaching Large Language Models to Reason with Reinforcement Learning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-19T01:16:57.097300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T01:15:49.658412Z digest=sha256:86b4d3d509586fc34cbd539a7b1289c07c15942afccc3295e3e4a9c92d8a8e55

Observation d1ae46cd-7879-4233-b5ff-171a85413323 · inbound

Beyond Memorization: Extending Reasoning Depth with Recurrence, Memory and Test-Time Compute Scaling cites this paper.

Beyond Memorization: Extending Reasoning Depth with Recurrence, Memory and Test-Time Compute Scaling Teaching Large Language Models to Reason with Reinforcement Learning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:51:50.809262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T20:49:26.966293Z digest=sha256:c964ec3d640a3cf08d681f294dc3b800243d876e0c4234536cbd714ceb6a7bc0

Observation 7c481f30-7042-4931-9867-16411b6c3d47 · inbound

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance cites this paper.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Teaching Large Language Models to Reason with Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:08.022212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:08.022212Z digest=sha256:ca1e4273dbba2584b0b377a2455d5e3afa6144ed736c1fa34805c48e4a9f2acd

Observation 86d251eb-e19e-42fc-aca2-e7b41d675a63 · inbound

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment cites this paper.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Teaching Large Language Models to Reason with Reinforcement Learning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:20:54.606098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:10e9660df4523dda65e5b338dae86cfb651adfadd0a2b48027d35eefd263a5db

Observation f7fc2164-4bbc-444f-8172-de9b9d233944 · inbound

Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning cites this paper.

Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning Teaching Large Language Models to Reason with Reinforcement Learning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:13:47.929498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T01:11:26.411893Z digest=sha256:70d2fb75214e960bf785599caae410a1c10b560ba374dde9893b5974b17ed4c6

Observation f81f4d0a-b7c7-4490-9bbc-57456cdccdf8 · inbound

SeLaR: Selective Latent Reasoning in Large Language Models cites this paper.

SeLaR: Selective Latent Reasoning in Large Language Models Teaching Large Language Models to Reason with Reinforcement Learning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:35:49.556259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-10T18:27:36.132030Z digest=sha256:ce1b0212fc17835b34b6fd8babad376cec292b546a16a81eccf27d6e542cb0a1

Observation 3fa2893a-9d5a-40d5-a0d3-b56879453ce5 · inbound

LiFT: How to Enable In-Context Learning for Longitudinal Modelling cites this paper.

LiFT: How to Enable In-Context Learning for Longitudinal Modelling Teaching Large Language Models to Reason with Reinforcement Learning

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:18:22.899179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T00:13:44.570359Z digest=sha256:91c8c504bb6c595d836775ed8089c1e6caeaf4b981e421bcde5ab69d7b4c5fe2

Observation 6a38efa3-6279-45c4-b446-85466f863fef · inbound

Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding cites this paper.

Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding Teaching Large Language Models to Reason with Reinforcement Learning

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:16:06.034886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-09T23:05:05.251150Z digest=sha256:dbe569034a6309c40c225d36468614c3b019ee4684e8b35cd5ffba1a75131aa2

Observation 94c2ee82-3972-41d2-9102-f160f9404b65 · inbound

Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning cites this paper.

Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning Teaching Large Language Models to Reason with Reinforcement Learning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:26:27.797962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T06:13:09.898530Z digest=sha256:74cfde5d1346add2de620d7fe148df77b1a49f14a21a4ecfefcba2634b33951b

Observation d02a8372-ccfc-427e-8cbc-d60d3256639e · inbound

Logic-Regularized Verifier Elicits Reasoning from LLMs cites this paper.

Logic-Regularized Verifier Elicits Reasoning from LLMs Teaching Large Language Models to Reason with Reinforcement Learning

Reference 72

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:51:10.941397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-08T10:54:01.229934Z digest=sha256:0ea5fc68198acd9563d1be0496f9a80230685a4f2e7604ddc9bbf049bb601c0d

Observation dda6e555-6545-49a4-8474-436f8a5da3d9 · inbound

NoisyCoconut: Counterfactual Consensus via Latent Space Reasoning cites this paper.

NoisyCoconut: Counterfactual Consensus via Latent Space Reasoning Teaching Large Language Models to Reason with Reinforcement Learning

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:41:24.258184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-12T00:51:40.815981Z digest=sha256:2a013420895e37d43c6120e799b2cb512e05c812bf2c7da92c2d06d9e312bf17

Observation d6d976ea-a5e7-40f6-8fe6-1d56bae82ff5 · inbound

Epistemic Uncertainty for Test-Time Discovery cites this paper.

Epistemic Uncertainty for Test-Time Discovery Teaching Large Language Models to Reason with Reinforcement Learning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:57:06.061818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T01:52:41.192353Z digest=sha256:56fb9f2a6cd078c5899d630ca03b69456d57d94b884fd722c30c6726d0b0f67a

Observation c81a34d1-d755-4ac4-914f-1d137bb1437d · inbound

When Reasoning Traces Become Performative: Step-Level Evidence that Chain-of-Thought Is an Imperfect Oversight Channel cites this paper.

When Reasoning Traces Become Performative: Step-Level Evidence that Chain-of-Thought Is an Imperfect Oversight Channel Teaching Large Language Models to Reason with Reinforcement Learning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:32:24.335451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T06:30:12.558660Z digest=sha256:df77708295a65339150eed2379c49d1379b90caf07217fe6b3634a70a6c5b219

Observation 9a1e33e8-a197-49ed-ae07-e9a23d3c2147 · inbound

From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI cites this paper.

From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI Teaching Large Language Models to Reason with Reinforcement Learning

Reference 139

Resolution
unresolved
no resolver link, observed 2026-08-02T11:29:31.754741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:29:31.754741Z digest=sha256:6e5d695049290dbf9859b7c1745b6a2301109b3ebdb342cc1d4f629de5b34c87

Observation 40d461bd-51fb-4411-a91d-074baac150f9 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Teaching Large Language Models to Reason with Reinforcement Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:e4affb250b38073d73e792e8202e7b4ed7958e0955d8b337ae4190e49e98efa5

Observation 15e17f8c-3d1d-4117-a752-b61ea7913795 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Teaching Large Language Models to Reason with Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:36.004785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:36.004785Z digest=sha256:28f27de4adfdeed1b126525d1c7e7535ae1b159b9497063f26e4f1a2a40b2be6

Observation 58f57867-c9d6-4227-84b9-37c8ca272be2 · inbound

Structured Synthetic Reasoning Data for Arithmetic Fine-Tuning of Small Language Models cites this paper.

Structured Synthetic Reasoning Data for Arithmetic Fine-Tuning of Small Language Models Teaching Large Language Models to Reason with Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T13:26:28.373498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:26:28.373498Z digest=sha256:8de0391b5a93268d8ad77c72568332935202f645729653682f51e4e758ce552a

Observation d3178038-8e0d-482d-b8e8-4a088c93762d · inbound

LeAct: Learning to Reason from Expert Actions cites this paper.

LeAct: Learning to Reason from Expert Actions Teaching Large Language Models to Reason with Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T06:32:19.440448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:32:19.440448Z digest=sha256:8043904a66c9db6cc9b1714b0f489e3020e750bce19e65aa997e184fd089036d

Observation 46f0c937-c1c1-42b7-b42e-7b9169cd431c · inbound

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning cites this paper.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning Teaching Large Language Models to Reason with Reinforcement Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:56.790463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:56.790463Z digest=sha256:9008f42cd4efaef2b3907adda985c983aed358e365008bf6b20f522bcc32791e

Observation 639afc79-1014-4fc7-896c-13a2914bd6db · inbound

RepoReasoner: Evaluating Repository-Level Code Reasoning Ability of Long-Context Language Models cites this paper.

RepoReasoner: Evaluating Repository-Level Code Reasoning Ability of Long-Context Language Models Teaching Large Language Models to Reason with Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T00:57:34.320933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:57:34.320933Z digest=sha256:227e15f4b884755ae7f95cb789f7ea3fbddf564ae9b37260fdfd60ce04172c33

Observation 3b9c5573-8ad2-4bc5-88c1-ea212ae184dc · inbound

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details cites this paper.

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details Teaching Large Language Models to Reason with Reinforcement Learning

Reference 193

Resolution
unresolved
no resolver link, observed 2026-08-05T15:25:40.180493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:25:40.180493Z digest=sha256:cd3a0d08c0a88dcd49927460dbc4297933ee4db5bae0fcad576f6fcbc083a59e