Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

As of 11 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 86 inbound Pith citation observations for arXiv:2506.14245.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.14245 v2

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T12:44:27.769465Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 86 of 86 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:17:53.443925Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 7a17aa33-fcff-43ff-a872-6e30272278cc · outbound

This paper cites Chain-of-Thought Reasoning In The Wild Is Not Always Faithful.

Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-06-01T03:02:49.618587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T12:44:27.769465Z digest=sha256:57f660ab56b9272ef7674eef309f90558e32da3f2128d18c66e189e332b5606f

Observation 0adda03e-aaa3-4f17-9821-0de309ce15c0 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T12:44:27.787791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T12:44:27.769465Z digest=sha256:7abf65a8f76d0f20f9e0ef91b43f130c1e30d31de6d39d63392ecae648218661

Observation 795ebf42-5265-4b53-9011-6cfc92740d34 · outbound

This paper cites Qwen2.5 Technical Report.

Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs Qwen2.5 Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-13T12:44:27.790554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T12:44:27.769465Z digest=sha256:c8f2323b169b44e3f30e0dfb0f0f3226a287c554e7f1a25f249e968ebab8d355

Observation e2883e1e-efb1-48f4-90f8-43236ee1a0f5 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-13T12:44:27.792652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T12:44:27.769465Z digest=sha256:dfe16cb52c34e21cc6fdbeafeea9ded8fc8af0a836ea6202bd7f73ce7a17be6e

Observation 5f272f10-a7ba-463b-b59d-c1f1e7b57e98 · outbound

This paper cites Clas- sify them into the following categories (if applicable): - **Calculation Error**: Mistakes in arithmetic, algebraic manipulation, or numerical computation.

Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs Clas- sify them into the following categories (if applicable): - **Calculation Error**: Mistakes in arithmetic, algebraic manipulation, or numerical computation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:44:27.795018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T12:44:27.769465Z digest=sha256:b286b7cac8553574bab1a327e2c41a9a67c9c12360fb5130012523100f8bf434

Observation 1d25eb30-e5e1-40a6-b95c-2dc798e76616 · outbound

This paper cites unideal case.

Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs unideal case

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:44:27.796840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T12:44:27.769465Z digest=sha256:4c1d4c0f526e79da01ae15e264cb5aba045b7add41f36c05bf79b3e7e9080fde

Observation 87c3b3a5-2075-4791-9ddd-c9a9e968da19 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-13T12:44:27.799090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T12:44:27.769465Z digest=sha256:40836508f1562297d1e134296ee4e3781fe520b2d7085bae7b12afa52bca98f6

Observation 1b19361c-1b56-417e-a1eb-152206b607bd · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-05-13T12:44:27.801141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T12:44:27.769465Z digest=sha256:773ed78fda327809509ce657cbdc8c9707197ef37e3779f2a2775c070fc60f46

Observation e20cefe5-ceab-4a7e-85b3-e80ddb853411 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-13T12:44:27.803305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T12:44:27.769465Z digest=sha256:30c1bd1c8f2c04199f530adc69e3284c8457ba91bce95e27d4be592018a4e3db

Observation 09412fc5-719c-4de9-94db-2a383e178685 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-05-13T12:44:27.805028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T12:44:27.769465Z digest=sha256:3fc7577a8087bb8b340527505878e6d8da7b24b70113b18558ac3eeac77df678

Observation d526927e-4834-41c1-bfee-9b5eeee96687 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-13T12:44:27.807099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T12:44:27.769465Z digest=sha256:7ee36d9d3d3794cb912abb994f1c52cff0997577fbdc013f7d506d4316cbe489

Observation 85840522-6502-4e21-91a9-ff09888f1da6 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-13T12:44:27.808830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T12:44:27.769465Z digest=sha256:2590d043e96b457b662447d18175c7497220475ffaedf87e6e9f48380e9900cf

Observation c090d2ba-3d59-42ef-b97e-bfb1ddf7cd43 · outbound

This paper cites This distance can be written in the form m√n p , wherem,n, andpare positive integers,mandpare relatively prime, andn is not divisible by the square of any prime.

Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs This distance can be written in the form m√n p , wherem,n, andpare positive integers,mandpare relatively prime, andn is not divisible by the square of any prime

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:44:27.811046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T12:44:27.769465Z digest=sha256:9838520d5bd61bdc38168e9b901a8f81ee82afd1896ec1b5ab7dc8021c55dbcc

Observation d14f7b43-69e3-4837-9229-6bbc04583f5b · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-13T12:44:27.812573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T12:44:27.769465Z digest=sha256:7248e56c1681e58b43cb6ea50b9affe369e851866bc8c690aba9177194f578cb

Observation 33efeb71-c8c3-4231-b61d-9c1cd6355450 · outbound

This paper cites DeepSeek-R1-0528-Qwen3-8B verify: The area calculation for triangle MNE uses DE + EG as a base, which is not a valid base unless DE and EG are collinear.

Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs DeepSeek-R1-0528-Qwen3-8B verify: The area calculation for triangle MNE uses DE + EG as a base, which is not a valid base unless DE and EG are collinear

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:44:27.814556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T12:44:27.769465Z digest=sha256:e33349f5ba818e4cfa7a21fb988eb23387cc015b03bbb20a7bb595dfba8491b8

Observation a477e74d-705c-4c47-89b4-e430c0beb7a7 · outbound

This paper cites DeepSeek-R1-0528-Qwen3-8B verify: - **Omission / Incompleteness** - The so- lution does not provide a complete justification for why the point (1, -1) gives the maximum value.

Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs DeepSeek-R1-0528-Qwen3-8B verify: - **Omission / Incompleteness** - The so- lution does not provide a complete justification for why the point (1, -1) gives the maximum value

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:44:27.816322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T12:44:27.769465Z digest=sha256:d2bd222bbe5409b74ce7b0e3635146aab83a53ec2d2c1643982711f582887c4e

Pith citing papers

Observation c8c83820-e531-4d45-9f9c-500a770b25d5 · inbound

From Reasoning to Code: GRPO Optimization for Underrepresented Languages cites this paper.

From Reasoning to Code: GRPO Optimization for Underrepresented Languages Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:21.129156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:21.129156Z digest=sha256:216536a2b39239c0727885a1c6bc2222e0958e307e400aa39d1d652bc8f15eb5

Observation df4d5d63-c7c8-4137-ac54-04eaa3f322c9 · inbound

Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR cites this paper.

Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-21T23:24:26.306377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-21T23:20:45.685446Z digest=sha256:91254b7c6f3c135198609fac178ce6b9f560446abba5065f78794829831c68ea

Observation c8671764-0901-48f0-b025-98b5b4980efb · inbound

Cascaded Information Disclosure for Generalized Evaluation of Problem Solving Capabilities cites this paper.

Cascaded Information Disclosure for Generalized Evaluation of Problem Solving Capabilities Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T10:28:55.422601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:28:55.422601Z digest=sha256:57511648289382003946c20b45632088b04e294d514f88acf49901828513991b

Observation f0813485-2174-467c-af2d-4067382cdc97 · inbound

CLPO: Curriculum Learning meets Policy Optimization for LLM Reasoning cites this paper.

CLPO: Curriculum Learning meets Policy Optimization for LLM Reasoning Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T13:53:35.444747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:53:35.444747Z digest=sha256:44c4ee956cecaefb0ac420dd5d5c67c06214ccc483e879beaee13e9c52ff53d1

Observation 7b46c176-c529-496e-a393-add5908b0750 · inbound

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers cites this paper.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-25T07:40:28.693870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:117a4c336c7d9378c008c2917e3c3f58e2035d9fca1d7c81936772d59e48748e

Observation 79f1bcaf-9f0f-41a5-af74-317c703a9f33 · inbound

Dynamic Generation of Multi-LLM Agents Communication Topologies with Graph Diffusion Models cites this paper.

Dynamic Generation of Multi-LLM Agents Communication Topologies with Graph Diffusion Models Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-21T20:40:35.717644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-21T20:40:13.721163Z digest=sha256:32c6ac04d315638bc8be6ad5b3ed12fac5bb3879d0aae3eee4ee96e730402ab9

Observation 7704fabb-9215-492a-b1c0-e5dd933473f1 · inbound

Auditing Data Membership in Reinforcement Learning With Verifiable Rewards cites this paper.

Auditing Data Membership in Reinforcement Learning With Verifiable Rewards Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:40:17.708608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T21:37:54.702010Z digest=sha256:8d6c44912b7bbdb308e503b2de77d5e9743a4bb5022c27eead35a05932a506a0

Observation aa1dca30-acca-487f-b4b0-1086a8389b6c · inbound

No More Stale Feedback: Co-Evolving Critics for Open-World Agent Learning cites this paper.

No More Stale Feedback: Co-Evolving Critics for Open-World Agent Learning Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T16:03:04.310183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T16:01:48.789986Z digest=sha256:6d9a99a2b9c48f34762a6cdb7d2cf108e543e80ecebab6d9f5854d9ba5697958

Observation 0fca62d3-9fac-4363-924b-df5ea3a77cc7 · inbound

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training cites this paper.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:54.268990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:54.268990Z digest=sha256:278e5b4bd7f98053ebbdb967e11fade06291cdd84f8f38f70f0f6894853acefc

Observation 0011bb1d-6b65-4435-bdae-0363d142794b · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 108

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:44.643557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:44.643557Z digest=sha256:ff90c4d2d406bc49127a7abb1123e5f089947e6dc73d886f786f7942b6738c50

Observation ab7ef238-a2c3-494d-9dd6-7b4ac94ed9cb · inbound

VI-CuRL: Stabilizing Verifier-Independent RL Reasoning via Confidence-Guided Variance Reduction cites this paper.

VI-CuRL: Stabilizing Verifier-Independent RL Reasoning via Confidence-Guided Variance Reduction Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-25T07:26:41.719095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-25T07:26:35.767179Z digest=sha256:c6bb0d3e68f3dd1bfb008f87367428f6f837793a08a0e25587848749f3a8d28a

Observation 63ab4a7f-3bc9-4d31-878f-e8afd30fe1e5 · inbound

On the Emergence of Implicit Curriculum in RLVR Learning Dynamics cites this paper.

On the Emergence of Implicit Curriculum in RLVR Learning Dynamics Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T23:11:56.783061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T23:11:56.783061Z digest=sha256:3c1ce4042c1e40ccfb95dfe7ec2acb453bbe32c4b9b8d2c1a6a52090d6915d54

Observation 33ad160b-697c-4eb7-9458-1ef6d97d3929 · inbound

Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation cites this paper.

Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 124

Resolution
verified exact
local_arxiv, observed 2026-05-15T21:51:40.674171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T21:51:10.972744Z digest=sha256:dd1a88670cc59d4a9ba3a16b7944c83bacee83a4776679ab6071c904b9a177b4

Observation 8c8861f0-cdee-46ed-80e5-299a150835d9 · inbound

Vibe Coding XR: Accelerating AI + XR Prototyping with XR Blocks and Gemini cites this paper.

Vibe Coding XR: Accelerating AI + XR Prototyping with XR Blocks and Gemini Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-15T00:18:21.564249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T00:17:41.692833Z digest=sha256:ffe44b79e428f2e59e01886fb66bf1cea67cf8cfa685bad0e3282bb6575bcd53

Observation e968aa27-6968-4ba1-968a-1ef12cd44a4d · inbound

C2F-Thinker: Coarse-to-Fine Reasoning with Hint-Guided Reinforcement Learning for Multimodal Sentiment Analysis cites this paper.

C2F-Thinker: Coarse-to-Fine Reasoning with Hint-Guided Reinforcement Learning for Multimodal Sentiment Analysis Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-15T13:55:53.249343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T13:51:40.334057Z digest=sha256:086b5fdf7bc9163b326f23ea6166c38e6dca7dd4e7da1c6b888fe0a4dd2a26ec

Observation 77482315-5405-4744-b986-a54514bc38d3 · inbound

Interpretable Electrophysiological Features of Resting-State EEG Capture Cortical Network Dynamics in Parkinsons Disease cites this paper.

Interpretable Electrophysiological Features of Resting-State EEG Capture Cortical Network Dynamics in Parkinsons Disease Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T14:23:49.346727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:23:49.346727Z digest=sha256:31799dec239083e5152db1bf7ecb488034a2fcfe88b33d36cd1d2e3d8c472773

Observation 3a833d6e-6843-479b-a267-1a0efbee4e22 · inbound

Chart-RL: Policy Optimization Reinforcement Learning for Enhanced Visual Reasoning in Chart Question Answering with Vision Language Models cites this paper.

Chart-RL: Policy Optimization Reinforcement Learning for Enhanced Visual Reasoning in Chart Question Answering with Vision Language Models Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-13T20:03:12.149433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T20:02:44.124502Z digest=sha256:7850d011dbd648c61050b8705ec4f02ece4d3034087fcd95fe62a2b7d62ec658

Observation 8906811d-7d64-4e30-b030-cbe09a2f9ab1 · inbound

The Stepwise Informativeness Assumption: Why are Entropy Dynamics and Reasoning Correlated in LLMs? cites this paper.

The Stepwise Informativeness Assumption: Why are Entropy Dynamics and Reasoning Correlated in LLMs? Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T13:05:38.304562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T13:05:26.484571Z digest=sha256:81f183d0848f7cc222acd68ab7d3281bbf81a05919971fe2b910c3bc8e91edae

Observation 3348998f-c5a2-4a65-bca7-2bc632d57b7f · inbound

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems cites this paper.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:44:27.816864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:08ede5b84a654e35c70b3cbe612c9451ed5fafd47740fc1c60d01dc44a82718b

Observation a7993578-5777-4df4-9c32-c04bc96f62a0 · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:44:27.816864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:25d07d757195aedcf66ecf9bf4e1a114b5eab9a4d7d64ba6cc3dc0424f8cd210

Observation ff11fec3-bbdb-43d0-95ed-2af57c26a296 · inbound

AutoOR: Scalably Post-training LLMs to Autoformalize Operations Research Problems cites this paper.

AutoOR: Scalably Post-training LLMs to Autoformalize Operations Research Problems Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:44:27.816864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T07:02:02.992871Z digest=sha256:99ba9a541c410617023a01616f39d1fefd36bd8d08a54b2bb6b2dc721bb86886

Observation 6a234c22-3534-41f6-8b1c-5159bbca5794 · inbound

WebGen-R1: Incentivizing Large Language Models to Generate Functional and Aesthetic Websites with Reinforcement Learning cites this paper.

WebGen-R1: Incentivizing Large Language Models to Generate Functional and Aesthetic Websites with Reinforcement Learning Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:44:27.816864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T00:18:18.340391Z digest=sha256:0304f131bc4a96b0836a95c661ab1134b39139177000b40535590d495d449a64

Observation bcb35224-6a9e-4f1d-bb75-7b1284dc1b73 · inbound

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models cites this paper.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:44:27.816864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:f249add17204886e627b399279fcc8c3abfd1baaac6f0cbec925a4c1d8e425d2

Observation a16d9e53-f368-4c31-874f-d5e50ee45a38 · inbound

Discovering Agentic Safety Specifications from 1-Bit Danger Signals cites this paper.

Discovering Agentic Safety Specifications from 1-Bit Danger Signals Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:44:27.816864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T08:23:38.003940Z digest=sha256:bd7f061faf58bde98d117edb28e1dbe0b9b73869a691107c40d2f8c75bed8fe6

Observation 0415a4d8-153f-42b7-962d-863014de4441 · inbound

Optimizing ground state preparation protocols with autoresearch cites this paper.

Optimizing ground state preparation protocols with autoresearch Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:44:27.816864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-07T17:04:34.912728Z digest=sha256:72f9729be260965a5b25cc90dab0d4fb69f46c8831b422aae491c03efcaa0aac

Observation 9c23c0a8-f7a1-4099-8c52-e6e58b50e643 · inbound

Optimizing ground state preparation protocols with autoresearch cites this paper.

Optimizing ground state preparation protocols with autoresearch Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:44:27.816864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T02:09:12.325998Z digest=sha256:836d5597c8a505a97d659de87653ee15141b1b60b8f4955cd22f4dd125182eb3

Observation 0dddffff-f69d-4493-8f53-72248841d2be · inbound

Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake Detection cites this paper.

Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake Detection Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:44:27.816864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-09T14:04:52.065878Z digest=sha256:ff118f23d9aa3e715834f8eef56e7bdb15dc42840813e4313ad02c533e43f5d9

Observation 2dffd715-a828-4070-b5ec-2dfd112bdbaa · inbound

Reference-Sampled Boltzmann Projection for KL-Regularized RLVR: Target-Matched Weighted SFT, Finite One-Shot Gaps, and Policy Mirror Descent cites this paper.

Reference-Sampled Boltzmann Projection for KL-Regularized RLVR: Target-Matched Weighted SFT, Finite One-Shot Gaps, and Policy Mirror Descent Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:44:27.816864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T19:34:22.546508Z digest=sha256:a288926028f0002b2ef8e91e8028b146ef58bedc0a49028d4b4bafba83377434

Observation d8239cca-5738-471b-b1f7-e3ab3a817cbc · inbound

Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMs cites this paper.

Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMs Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 68

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T12:44:27.816864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T16:58:10.013475Z digest=sha256:cb3fb2126baf8a6444bb89f6ea6aa76795b853903254d0a2b512f452476e25c1

Observation 605fb9f0-173f-4166-ad33-2f33f6655427 · inbound

Adapt to Thrive! Adaptive Power-Mean Policy Optimization for Improved LLM Reasoning cites this paper.

Adapt to Thrive! Adaptive Power-Mean Policy Optimization for Improved LLM Reasoning Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T12:44:27.816864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T16:51:19.555272Z digest=sha256:545b09cdb3f46ac493a12e18e8bef6ce6aa256b6ec4fa530086406bd8c9d4f38

Observation 62d51b8f-6e7f-4e60-8fea-085f7967a7eb · inbound

Efficiently Aligning Language Models with Online Natural Language Feedback cites this paper.

Efficiently Aligning Language Models with Online Natural Language Feedback Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 1

Resolution
malformed identifier
arxiv_id, observed 2026-05-13T12:44:27.816864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T16:55:12.949494Z digest=sha256:96e71d25867c3c1f026b0b96a3a5c42203619b70f149ba2e00a77852c5426430

Observation c9eb816b-b0a9-4b62-894a-54e42f5ba987 · inbound

Efficiently Aligning Language Models with Online Natural Language Feedback cites this paper.

Efficiently Aligning Language Models with Online Natural Language Feedback Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 1

Resolution
malformed identifier
local_arxiv, observed 2026-06-30T23:45:07.593611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T23:44:39.010504Z digest=sha256:8b91dcede703d1aa5c651ec4230efbbbc0d66be703f9a9ef8772a7211fb29042

Observation 68663c42-881c-4f77-bc12-9d7b5a9779f4 · inbound

Pen-Strategist: A Reasoning Framework for Penetration Testing Strategy Formation and Analysis cites this paper.

Pen-Strategist: A Reasoning Framework for Penetration Testing Strategy Formation and Analysis Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:44:27.816864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T17:48:09.078064Z digest=sha256:42cfdf79780c6b482e9790bb1eb96a3103b1da5ea9428857444db077de518895

Observation ccdf6be5-6d72-4abf-a9a4-d721b68be54f · inbound

Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning cites this paper.

Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:44:27.816864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T06:13:09.898530Z digest=sha256:a1e8f281015dac8da27da00c764fd132e0eab7fd0bc817516904df9b7f9bd337

Observation f92258b0-2e12-41a5-9b7f-d0f4f6ad2edd · inbound

Gradient Extrapolation-Based Policy Optimization cites this paper.

Gradient Extrapolation-Based Policy Optimization Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:44:27.816864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-11T02:07:08.792030Z digest=sha256:5cd639cadf938f4f36c439c16018027cbe5036d938800d6bbfbbc979b43dfe9d

Observation 3e8e6f52-3640-4926-bd8d-e9c90a41bfa5 · inbound

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR cites this paper.

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:44:27.816864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T01:09:27.543625Z digest=sha256:88586cbb64039bff92832472d69dc678700eb8e4c19e1b7f3e458241aac3e9b3

Observation c6580ab5-e3ba-422e-ac3f-cde0e5d54dad · inbound

CauSim: Scaling Causal Reasoning with Increasingly Complex Causal Simulators cites this paper.

CauSim: Scaling Causal Reasoning with Increasingly Complex Causal Simulators Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:44:27.816864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T02:33:04.549851Z digest=sha256:fe7144855f13de081da3e575761520662f5b0b95fc1e0d3eadb76efb49d91b1f

Observation 1ddbe0bb-a75f-4756-9302-7b2c86852aad · inbound

expo: Exploration-prioritized policy optimization via adaptive kl regulation and gaussian curriculum sampling cites this paper.

expo: Exploration-prioritized policy optimization via adaptive kl regulation and gaussian curriculum sampling Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:44:27.816864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:20:30.301154Z digest=sha256:e7a2b3f02de8f7d09a15377249ffd08cb2e1270bd291cf51987f18a13d54f024

Observation 1ed35065-6247-423a-b2e5-64c5873eeff2 · inbound

fg-expo: Frontier-guided exploration-prioritized policy optimization via adaptive kl and gaussian curriculum cites this paper.

fg-expo: Frontier-guided exploration-prioritized policy optimization via adaptive kl and gaussian curriculum Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:44:27.816864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T03:06:51.437906Z digest=sha256:dfca4e931beed19a23d99e63002c9e06469323aed4af59824253b19f71a25c20

Observation ebc3dd4d-5909-4b02-b835-5cec98e3d1e0 · inbound

Learn to Think: Improving Multimodal Reasoning through Vision-Aware Self-Improvement Training cites this paper.

Learn to Think: Improving Multimodal Reasoning through Vision-Aware Self-Improvement Training Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:44:27.816864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T06:17:57.264809Z digest=sha256:2cb5c91507ccc5ca3afac5e3bdae4c0ffa644f64f304c142ad6d602b6b46b70d

Observation e14f5791-9aec-4dfa-bad1-72e99380fa95 · inbound

Holder Policy Optimisation cites this paper.

Holder Policy Optimisation Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T12:44:27.816864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-13T06:08:28.855671Z digest=sha256:b126bcae38b0f1487f43b976d2b4aeef4ab76aca8b75dd2e43a009233db301b6

Observation e61f42e1-8c5c-4827-8c7b-f6674608dfbe · inbound

Holder Policy Optimisation cites this paper.

Holder Policy Optimisation Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T10:01:23.109055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:e337ce40966926e388ce63e080d830836c1e13ce69fc32825b5a03d2b910a2da

Observation c3b606eb-4205-4ba6-9a09-d807c49b625e · inbound

Reasoning Can Be Restored by Correcting a Few Decision Tokens cites this paper.

Reasoning Can Be Restored by Correcting a Few Decision Tokens Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-19T20:57:46.844803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T20:56:48.771058Z digest=sha256:f52f71dcf37a0b3641926e27c3c4761a37f9017288b796cb6e642d4353d32491

Observation 6e9872b8-fd74-45a5-9b40-e0927155b269 · inbound

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment cites this paper.

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:13:16.104340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T12:11:23.775843Z digest=sha256:3e15296d670599a4f7a5434a9e0cfbd4fc1784d452a5f9e613fc06a80d724e64

Observation ea38ffc0-0870-4c71-a473-9bb34a44d38b · inbound

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment cites this paper.

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:21:21.445174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T09:19:39.848194Z digest=sha256:54fc4155894d15b009fd1faa36bcaf406039877b6f1f6443af14c4d845551545

Observation 385c5967-c3f6-471e-b7b9-801c40ae247b · inbound

TimeSRL: Generalizable Time-Series Behavioral Modeling via Semantic RL-Tuned LLMs -- A Case Study in Mental Health cites this paper.

TimeSRL: Generalizable Time-Series Behavioral Modeling via Semantic RL-Tuned LLMs -- A Case Study in Mental Health Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-05-21T06:13:59.541184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-21T06:09:44.172188Z digest=sha256:1fb5cf87fa1fda7ea30e87992fe57a7d30383f52b9c35a6d8b9a71e568178ace

Observation 501dfa38-bae4-4d24-adb8-4753e01e0e59 · inbound

RL with Learnable Textual Feedback: A Bilevel Approach cites this paper.

RL with Learnable Textual Feedback: A Bilevel Approach Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-06-30T14:44:45.389511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T14:36:31.047856Z digest=sha256:c02112818b24bb57f38d2cc2413024d88cece2f1f4f00054ff6b6df30e517dc8

Observation 42e53109-e362-4219-a074-f2e066a92577 · inbound

Detecting Unfaithful Chain-of-Thought via Circuit-Guided Internal-External Discrepancy cites this paper.

Detecting Unfaithful Chain-of-Thought via Circuit-Guided Internal-External Discrepancy Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-06-29T21:53:59.515085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T21:47:17.894881Z digest=sha256:09a42ffdffb10819864d686b3bed5b57bf0dcd19901c619edfb5f62ea766fb14

Observation aa53460f-d11c-4698-a351-66cd561626f2 · inbound

Single-Rollout Hidden-State Dynamics for Training-Free RLVR Data Selection cites this paper.

Single-Rollout Hidden-State Dynamics for Training-Free RLVR Data Selection Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-06-29T14:13:30.091451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T14:08:40.968105Z digest=sha256:3dc0897e093078875285ab1e20d06ca3c038d91095f6a63392a88f8991fd4955

Observation be6dab5a-7f68-4565-b122-f32f9ef196e7 · inbound

Learning Design Skills as Memory Policies for Agentic Photonic Inverse Design cites this paper.

Learning Design Skills as Memory Policies for Agentic Photonic Inverse Design Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T08:03:13.854769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T08:02:27.359309Z digest=sha256:71dd2ba69be9df19b1acce139bd33b8830b7d0f29b7afa856a060e61fc8c6a09

Observation 0c0dcf0e-1c86-490f-b523-8b830b02350f · inbound

Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination cites this paper.

Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-28T23:12:47.130224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T23:08:05.597810Z digest=sha256:a1b1ec6e8a93e73020621fe5d2d804de0b52842ac72d6e01c29297b25b2a8be0

Observation d61e6317-f63a-4d56-9aaf-91f9163c9f3f · inbound

DRIFT: Decoupled Rollouts and Importance-Weighted Fine-Tuning for Efficient Multi-Turn Optimization cites this paper.

DRIFT: Decoupled Rollouts and Importance-Weighted Fine-Tuning for Efficient Multi-Turn Optimization Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-06-28T23:22:47.446580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T23:16:49.358792Z digest=sha256:a90f6dbd7976b50d019615b39fbd6b2e5d28025829b96a618e5e672c438830b5

Observation 2da60c93-eec6-4ef7-8097-6a86456ffc7b · inbound

ARCA: Adapter-Residual Credit Assignment When Token Signals Degenerate cites this paper.

ARCA: Adapter-Residual Credit Assignment When Token Signals Degenerate Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:16:00.826631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-28T22:54:53.415067Z digest=sha256:4c8fd481a7800a5ea7d30fb16506472affbe44561eedc7a83c31fae26fcdedb0

Observation 3f073902-f3a6-4dff-82be-b1ab3471e0e4 · inbound

An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models cites this paper.

An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-01T21:36:14.954767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T16:45:33.046568Z digest=sha256:09eed6202cb4df626cfabb06bd1cb6a19d2a84c979b7e37c2791a5cfe5996266

Observation e4db9203-f400-4198-ba53-7fb855936f39 · inbound

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching cites this paper.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:36:27.134981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:a7e86a9de007fe6025a94e9ea646e81338deb014d09dfcdd4d905efaea9a6b46

Observation 875eacab-6215-4ed2-9e69-9d0407c66210 · inbound

RUBAS: Rubric-Based Reinforcement Learning for Agent Safety cites this paper.

RUBAS: Rubric-Based Reinforcement Learning for Agent Safety Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T02:16:26.545699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T11:07:47.814115Z digest=sha256:fdd1b480c4da1b14c0894cb387c0d1ffa3ad81bc66d31c80c1730b614fe3c673

Observation 8bb9c31c-a77d-4bd1-81a7-27ba35b847db · inbound

Smart Picks in the Dark: Towards Efficient RLVR for Reasoning via Tracing Metacognitive Pivots cites this paper.

Smart Picks in the Dark: Towards Efficient RLVR for Reasoning via Tracing Metacognitive Pivots Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T07:26:46.249936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T06:55:09.927034Z digest=sha256:711dbc0110c00ef7d47073454c0378cc7b019e8639c0fef3365f6ecb6b4f1468

Observation ad98ace7-89ff-444b-ba3b-9b0823f72811 · inbound

GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling cites this paper.

GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T06:06:40.860348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-28T07:45:43.320339Z digest=sha256:f4633e7928b999a41bb78061a8f1ea8f2987c93f16b682a129159907004268dc

Observation 195d58b3-f851-4598-8de3-4768f6b75741 · inbound

On Advantage Estimates for Max@K Policy Gradients cites this paper.

On Advantage Estimates for Max@K Policy Gradients Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.454152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:1dccea9e9786ad63bf04ec0bbcba8d39ae620f70a2e382b3567dc528ca5106e3

Observation 74730984-bc4d-4a21-8f11-f8778707e8c8 · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 271

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T01:37:30.535790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:6abf6f85c568fa184f893f1e619b9e8462cd2f8d525e33e2b41de075473681b9

Observation 2fa2167b-67a9-4e62-8a36-4ad5e412a7b5 · inbound

From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI cites this paper.

From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 120

Resolution
unresolved
no resolver link, observed 2026-08-02T11:29:29.541010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:29:29.541010Z digest=sha256:4a5df5c88e28e5bcb0e891a60f8b8cc1215c1430f23edc6e2905ab395bf9fea3

Observation 9b71737c-66b1-497a-b21c-82bb7dd11b1b · inbound

See First, Answer Later: Visual Evidence Pre-Alignment via Sufficiency-Driven RL cites this paper.

See First, Answer Later: Visual Evidence Pre-Alignment via Sufficiency-Driven RL Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 38

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T19:28:52.863616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T01:44:24.586195Z digest=sha256:21213879c4973860326b5c48262fc89d2f0a604e67473a0f96eb1160317befa7

Observation 2b53a7e4-1e48-4f8e-aabf-22a7877b44b3 · inbound

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning cites this paper.

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T20:38:55.869950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T01:13:11.483599Z digest=sha256:4c1319308ce3ea188f8196278135d8da3759c4635fe12ab0847841ef3abcb464

Observation a0e97a9f-e319-40a3-9b8d-541ea2f15184 · inbound

Learning from Your Own Mistakes: Constructing Learnable Micro-Reflective Trajectories for Self-Distillation cites this paper.

Learning from Your Own Mistakes: Constructing Learnable Micro-Reflective Trajectories for Self-Distillation Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:09:14.091061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T21:30:30.336251Z digest=sha256:ba59adcc0b917da9b74767cc6ec9e322eec62881073cad0610bdc3d2dd905280

Observation 4e67d117-1fec-4793-9a39-d6b064b6e480 · inbound

Seeing Before Reasoning: Decoupling Perception and Reasoning for Shortcut-Resilient Multimodal On-Policy Self-Distillation cites this paper.

Seeing Before Reasoning: Decoupling Perception and Reasoning for Shortcut-Resilient Multimodal On-Policy Self-Distillation Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:29:15.777562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T21:12:26.064012Z digest=sha256:769dc7758ab9ca55cb41b03d01ff1db93189539c9c9fcc61d7f9b21f5fae5a50

Observation 4592c602-423d-4e8b-ac5d-c0dcdd8719ef · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 224

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:09:40.646504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:c766d133c3fa2fb29470a85b86b9c00bc5fb8b9b046a0619cf80938ab0887526

Observation 5d64ef3b-3bae-482f-8f5a-2ea59e43cdd5 · inbound

ReNIO: Reweighting Negative Trajectory Importance for LLM On-Policy Distillation cites this paper.

ReNIO: Reweighting Negative Trajectory Importance for LLM On-Policy Distillation Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:09:45.128946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-26T09:05:36.937295Z digest=sha256:29f711f3ffeebbffe4eefb9e404253bc91fb9b4f60231979cec365250b89e955

Observation cf8f9fac-b21d-42c9-9ea3-1245d14be7bc · inbound

Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views cites this paper.

Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 50

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T10:39:45.365913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T08:38:46.044079Z digest=sha256:10a06b078ac443b9124e3003ad622476888b2ef628130aab1e4fb8f573bcb85b

Observation 4bcd50ff-c2a1-4cbd-a5f9-ab7fe76b7ae7 · inbound

PointVG-R: Internalizing Geometric Reasoning in MLLMs for Precise Pointing Localization via Visual Chain of Thought cites this paper.

PointVG-R: Internalizing Geometric Reasoning in MLLMs for Precise Pointing Localization via Visual Chain of Thought Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 96

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:19:57.755265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T00:46:17.094339Z digest=sha256:335d2a705eefac01f007360d7521d43d9182d41b7cff0c0684af2eff78c6bbe3

Observation 6395ed83-137f-43e6-ba83-af5a45b904f4 · inbound

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards cites this paper.

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T12:44:40.152604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-30T10:07:39.554999Z digest=sha256:6f720006cdafb176f98b7527273ea2fb385569f42d5fc594450bc1889f5089ec

Observation a9dd725b-7137-4538-b79b-8bc1416704ef · inbound

Evo-PI: Aligning Medical Reasoning via Evolving Principle-Guided Supervision cites this paper.

Evo-PI: Aligning Medical Reasoning via Evolving Principle-Guided Supervision Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T10:25:41.162175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-01T05:35:37.578632Z digest=sha256:f1713f69c0c9c7bdf65826ef89e3672bbc1446a4e9c97db6166011ff05120ee4

Observation 160329a2-0799-4988-9f09-815b43d1694f · inbound

HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better cites this paper.

HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-11T11:58:14.396969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T11:58:14.396969Z digest=sha256:240a4143bdb26e38046017512bcc1e608f247f5ed61b70e61f8f9bcb7f4999f0

Observation 15d9a061-9375-45f8-9d03-7f72e2f173fb · inbound

From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier cites this paper.

From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 249

Resolution
verified exact
local_arxiv, observed 2026-07-10T18:17:33.858859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-10T18:16:31.176239Z digest=sha256:c1703f130672054da6e590ad7a714f0501ac4095cf98b3fb91d2ed9239f1d842

Observation a3c556d4-8710-45e5-8b1c-d7bf522a8c8a · inbound

OpenProver: Agentic and Interactive Theorem Proving with Lean 4 cites this paper.

OpenProver: Agentic and Interactive Theorem Proving with Lean 4 Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-13T04:38:11.219373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T04:38:11.219373Z digest=sha256:069c4e921bab75eb379d95290cc8f70ae187d4166a93ee02ae270ed96ba85c0e

Observation 2eb0d7ae-ad35-43ee-96db-33d83135d706 · inbound

The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy cites this paper.

The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 263

Resolution
unresolved
no resolver link, observed 2026-07-14T06:30:16.612345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:30:16.612345Z digest=sha256:4dfe6548c0342ca7ea85a08c4140326032018c5b69985d6069d21aa60b2fbf46

Observation 9ea837e8-8479-47c6-b016-841b48f2eb7a · inbound

Think Through a Bottleneck: Hourglass Reasoning for Rigorous Induction cites this paper.

Think Through a Bottleneck: Hourglass Reasoning for Rigorous Induction Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-14T03:48:31.314623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T03:48:31.314623Z digest=sha256:87556e51c1e37f7a1ab8928182d24a5b040df7bb9586f869bace56a10b37fe3e

Observation 971d7c0f-ea51-4964-9eb0-fd15aeb544a9 · inbound

Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL cites this paper.

Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-02T14:50:16.901566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:50:16.901566Z digest=sha256:43ee9938abe4b7a3acb8f01d0c9e6183f34fed58e10f34f6c4936a874b527694

Observation 0f0e6076-2c99-483a-921b-6606026c0a5e · inbound

PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization cites this paper.

PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T14:39:41.442301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:39:41.442301Z digest=sha256:7732e3d6a700e3b9255dc70d8575caeb7b238124830501e09af36db5db376abb

Observation 1c88179b-b75e-4a29-a01d-1511f31d0b37 · inbound

TraversRL: Traversable Pedestrian Pathway Generation With Reinforcement Learning cites this paper.

TraversRL: Traversable Pedestrian Pathway Generation With Reinforcement Learning Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T17:54:29.191419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:54:29.191419Z digest=sha256:94259c987ce51257d0f4ce3b7274eec60dd17692968171b8b6369a09af4ad0b6

Observation e7eeed5c-9805-4e0c-ba32-16f4e5399aac · inbound

SLPO: Scaling Latent Reasoning via a Surrogate Policy cites this paper.

SLPO: Scaling Latent Reasoning via a Surrogate Policy Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T12:06:50.831758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:06:50.831758Z digest=sha256:ee8a20fea0d83b970f523c51b5ea66291db21d3f8d1455930ffbf684a8d2ca14

Observation 2ccb2f31-6291-4be4-81d2-34523d9259e8 · inbound

How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift cites this paper.

How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T07:44:11.904578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:44:11.904578Z digest=sha256:d09857090b988e71762827342b625a07f256114f2415b9aeccea36817053ed52

Observation 2f4202ad-7ff8-481b-9067-c516c9cb5121 · inbound

Bridging Compute- and Data-Optimal Pretraining cites this paper.

Bridging Compute- and Data-Optimal Pretraining Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-01T03:02:07.807759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:02:07.807759Z digest=sha256:fdcd280d6e963227227537d2c9e5a3092c305e3e8cfed8cd6a0c0421afbf6fcc

Observation e44d51b0-8ae4-4ad0-9230-592978058a1f · inbound

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning cites this paper.

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-31T23:39:13.429089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:39:13.429089Z digest=sha256:02620db0b323dfe005a8cf6399a8a0d7f36ac4015c71843800ea863ec73e2a96

Observation 7df2fb86-6cf9-48a0-be5c-c82780243a82 · inbound

LEEPS: Latent-Guided Explore-Exploit Prompt Sampling for Efficient RLVR in Large Language Models cites this paper.

LEEPS: Latent-Guided Explore-Exploit Prompt Sampling for Efficient RLVR in Large Language Models Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-31T18:34:16.348704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T18:34:16.348704Z digest=sha256:d3cdca3b5d6f70fdf88034cb7569f7000788fe18fc19aeee3fe27ce8df2f7cbb

Observation 57cb7ae4-e6b9-4b1b-bdc4-76ccba8c7c9b · inbound

Uncertainty-Aware Simulation-Based Inference for Operations Research with Large Language Models cites this paper.

Uncertainty-Aware Simulation-Based Inference for Operations Research with Large Language Models Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T02:03:06.482231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T02:03:06.482231Z digest=sha256:14123a3428985f9ddcec25c3368c64d206d68cc0f6be809d21b357ef3411580d

Observation d7cabc12-5fd3-4ad5-9e9c-693f2d03cd75 · inbound

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs cites this paper.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.443925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.443925Z digest=sha256:61bfded05cf75978e2c620422d394fcb49601c77fa1a0242f2f3c96eb7331067