Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T12:44:27.769465Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 86 inbound Pith citation observations for arXiv:2506.14245.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T12:44:27.769465Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T17:17:53.443925Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
16 of 16 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
Observation 7a17aa33-fcff-43ff-a872-6e30272278cc · outbound
Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0adda03e-aaa3-4f17-9821-0de309ce15c0 · outbound
Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 795ebf42-5265-4b53-9011-6cfc92740d34 · outbound
Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs Qwen2.5 Technical Report
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e2883e1e-efb1-48f4-90f8-43236ee1a0f5 · outbound
Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5f272f10-a7ba-463b-b59d-c1f1e7b57e98 · outbound
Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs Clas- sify them into the following categories (if applicable): - **Calculation Error**: Mistakes in arithmetic, algebraic manipulation, or numerical computation
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1d25eb30-e5e1-40a6-b95c-2dc798e76616 · outbound
Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs unideal case
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 87c3b3a5-2075-4791-9ddd-c9a9e968da19 · outbound
Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1b19361c-1b56-417e-a1eb-152206b607bd · outbound
Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e20cefe5-ceab-4a7e-85b3-e80ddb853411 · outbound
Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 09412fc5-719c-4de9-94db-2a383e178685 · outbound
Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d526927e-4834-41c1-bfee-9b5eeee96687 · outbound
Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 85840522-6502-4e21-91a9-ff09888f1da6 · outbound
Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c090d2ba-3d59-42ef-b97e-bfb1ddf7cd43 · outbound
Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs This distance can be written in the form m√n p , wherem,n, andpare positive integers,mandpare relatively prime, andn is not divisible by the square of any prime
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d14f7b43-69e3-4837-9229-6bbc04583f5b · outbound
Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 33efeb71-c8c3-4231-b61d-9c1cd6355450 · outbound
Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs DeepSeek-R1-0528-Qwen3-8B verify: The area calculation for triangle MNE uses DE + EG as a base, which is not a valid base unless DE and EG are collinear
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a477e74d-705c-4c47-89b4-e430c0beb7a7 · outbound
Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs DeepSeek-R1-0528-Qwen3-8B verify: - **Omission / Incompleteness** - The so- lution does not provide a complete justification for why the point (1, -1) gives the maximum value
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c8c83820-e531-4d45-9f9c-500a770b25d5 · inbound
From Reasoning to Code: GRPO Optimization for Underrepresented Languages Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df4d5d63-c7c8-4137-ac54-04eaa3f322c9 · inbound
Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c8671764-0901-48f0-b025-98b5b4980efb · inbound
Cascaded Information Disclosure for Generalized Evaluation of Problem Solving Capabilities Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0813485-2174-467c-af2d-4067382cdc97 · inbound
CLPO: Curriculum Learning meets Policy Optimization for LLM Reasoning Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b46c176-c529-496e-a393-add5908b0750 · inbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 79f1bcaf-9f0f-41a5-af74-317c703a9f33 · inbound
Dynamic Generation of Multi-LLM Agents Communication Topologies with Graph Diffusion Models Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7704fabb-9215-492a-b1c0-e5dd933473f1 · inbound
Auditing Data Membership in Reinforcement Learning With Verifiable Rewards Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation aa1dca30-acca-487f-b4b0-1086a8389b6c · inbound
No More Stale Feedback: Co-Evolving Critics for Open-World Agent Learning Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0fca62d3-9fac-4363-924b-df5ea3a77cc7 · inbound
StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0011bb1d-6b65-4435-bdae-0363d142794b · inbound
Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 108
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab7ef238-a2c3-494d-9dd6-7b4ac94ed9cb · inbound
VI-CuRL: Stabilizing Verifier-Independent RL Reasoning via Confidence-Guided Variance Reduction Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 63ab4a7f-3bc9-4d31-878f-e8afd30fe1e5 · inbound
On the Emergence of Implicit Curriculum in RLVR Learning Dynamics Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33ad160b-697c-4eb7-9458-1ef6d97d3929 · inbound
Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 124
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8c8861f0-cdee-46ed-80e5-299a150835d9 · inbound
Vibe Coding XR: Accelerating AI + XR Prototyping with XR Blocks and Gemini Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e968aa27-6968-4ba1-968a-1ef12cd44a4d · inbound
C2F-Thinker: Coarse-to-Fine Reasoning with Hint-Guided Reinforcement Learning for Multimodal Sentiment Analysis Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 77482315-5405-4744-b986-a54514bc38d3 · inbound
Interpretable Electrophysiological Features of Resting-State EEG Capture Cortical Network Dynamics in Parkinsons Disease Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a833d6e-6843-479b-a267-1a0efbee4e22 · inbound
Chart-RL: Policy Optimization Reinforcement Learning for Enhanced Visual Reasoning in Chart Question Answering with Vision Language Models Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8906811d-7d64-4e30-b030-cbe09a2f9ab1 · inbound
The Stepwise Informativeness Assumption: Why are Entropy Dynamics and Reasoning Correlated in LLMs? Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3348998f-c5a2-4a65-bca7-2bc632d57b7f · inbound
Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a7993578-5777-4df4-9c32-c04bc96f62a0 · inbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ff11fec3-bbdb-43d0-95ed-2af57c26a296 · inbound
AutoOR: Scalably Post-training LLMs to Autoformalize Operations Research Problems Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6a234c22-3534-41f6-8b1c-5159bbca5794 · inbound
WebGen-R1: Incentivizing Large Language Models to Generate Functional and Aesthetic Websites with Reinforcement Learning Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bcb35224-6a9e-4f1d-bb75-7b1284dc1b73 · inbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a16d9e53-f368-4c31-874f-d5e50ee45a38 · inbound
Discovering Agentic Safety Specifications from 1-Bit Danger Signals Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0415a4d8-153f-42b7-962d-863014de4441 · inbound
Optimizing ground state preparation protocols with autoresearch Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9c23c0a8-f7a1-4099-8c52-e6e58b50e643 · inbound
Optimizing ground state preparation protocols with autoresearch Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0dddffff-f69d-4493-8f53-72248841d2be · inbound
Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake Detection Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2dffd715-a828-4070-b5ec-2dfd112bdbaa · inbound
Reference-Sampled Boltzmann Projection for KL-Regularized RLVR: Target-Matched Weighted SFT, Finite One-Shot Gaps, and Policy Mirror Descent Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d8239cca-5738-471b-b1f7-e3ab3a817cbc · inbound
Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMs Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 605fb9f0-173f-4166-ad33-2f33f6655427 · inbound
Adapt to Thrive! Adaptive Power-Mean Policy Optimization for Improved LLM Reasoning Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 62d51b8f-6e7f-4e60-8fea-085f7967a7eb · inbound
Efficiently Aligning Language Models with Online Natural Language Feedback Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c9eb816b-b0a9-4b62-894a-54e42f5ba987 · inbound
Efficiently Aligning Language Models with Online Natural Language Feedback Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 68663c42-881c-4f77-bc12-9d7b5a9779f4 · inbound
Pen-Strategist: A Reasoning Framework for Penetration Testing Strategy Formation and Analysis Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ccdf6be5-6d72-4abf-a9a4-d721b68be54f · inbound
Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f92258b0-2e12-41a5-9b7f-d0f4f6ad2edd · inbound
Gradient Extrapolation-Based Policy Optimization Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3e8e6f52-3640-4926-bd8d-e9c90a41bfa5 · inbound
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c6580ab5-e3ba-422e-ac3f-cde0e5d54dad · inbound
CauSim: Scaling Causal Reasoning with Increasingly Complex Causal Simulators Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1ddbe0bb-a75f-4756-9302-7b2c86852aad · inbound
expo: Exploration-prioritized policy optimization via adaptive kl regulation and gaussian curriculum sampling Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1ed35065-6247-423a-b2e5-64c5873eeff2 · inbound
fg-expo: Frontier-guided exploration-prioritized policy optimization via adaptive kl and gaussian curriculum Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ebc3dd4d-5909-4b02-b835-5cec98e3d1e0 · inbound
Learn to Think: Improving Multimodal Reasoning through Vision-Aware Self-Improvement Training Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e14f5791-9aec-4dfa-bad1-72e99380fa95 · inbound
Holder Policy Optimisation Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e61f42e1-8c5c-4827-8c7b-f6674608dfbe · inbound
Holder Policy Optimisation Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c3b606eb-4205-4ba6-9a09-d807c49b625e · inbound
Reasoning Can Be Restored by Correcting a Few Decision Tokens Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6e9872b8-fd74-45a5-9b40-e0927155b269 · inbound
AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ea38ffc0-0870-4c71-a473-9bb34a44d38b · inbound
AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 385c5967-c3f6-471e-b7b9-801c40ae247b · inbound
TimeSRL: Generalizable Time-Series Behavioral Modeling via Semantic RL-Tuned LLMs -- A Case Study in Mental Health Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 501dfa38-bae4-4d24-adb8-4753e01e0e59 · inbound
RL with Learnable Textual Feedback: A Bilevel Approach Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 42e53109-e362-4219-a074-f2e066a92577 · inbound
Detecting Unfaithful Chain-of-Thought via Circuit-Guided Internal-External Discrepancy Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation aa53460f-d11c-4698-a351-66cd561626f2 · inbound
Single-Rollout Hidden-State Dynamics for Training-Free RLVR Data Selection Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation be6dab5a-7f68-4565-b122-f32f9ef196e7 · inbound
Learning Design Skills as Memory Policies for Agentic Photonic Inverse Design Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0c0dcf0e-1c86-490f-b523-8b830b02350f · inbound
Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d61e6317-f63a-4d56-9aaf-91f9163c9f3f · inbound
DRIFT: Decoupled Rollouts and Importance-Weighted Fine-Tuning for Efficient Multi-Turn Optimization Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2da60c93-eec6-4ef7-8097-6a86456ffc7b · inbound
ARCA: Adapter-Residual Credit Assignment When Token Signals Degenerate Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3f073902-f3a6-4dff-82be-b1ab3471e0e4 · inbound
An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e4db9203-f400-4198-ba53-7fb855936f39 · inbound
Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 875eacab-6215-4ed2-9e69-9d0407c66210 · inbound
RUBAS: Rubric-Based Reinforcement Learning for Agent Safety Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8bb9c31c-a77d-4bd1-81a7-27ba35b847db · inbound
Smart Picks in the Dark: Towards Efficient RLVR for Reasoning via Tracing Metacognitive Pivots Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ad98ace7-89ff-444b-ba3b-9b0823f72811 · inbound
GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 195d58b3-f851-4598-8de3-4768f6b75741 · inbound
On Advantage Estimates for Max@K Policy Gradients Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 74730984-bc4d-4a21-8f11-f8778707e8c8 · inbound
Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 271
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2fa2167b-67a9-4e62-8a36-4ad5e412a7b5 · inbound
From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 120
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b71737c-66b1-497a-b21c-82bb7dd11b1b · inbound
See First, Answer Later: Visual Evidence Pre-Alignment via Sufficiency-Driven RL Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2b53a7e4-1e48-4f8e-aabf-22a7877b44b3 · inbound
From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a0e97a9f-e319-40a3-9b8d-541ea2f15184 · inbound
Learning from Your Own Mistakes: Constructing Learnable Micro-Reflective Trajectories for Self-Distillation Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4e67d117-1fec-4793-9a39-d6b064b6e480 · inbound
Seeing Before Reasoning: Decoupling Perception and Reasoning for Shortcut-Resilient Multimodal On-Policy Self-Distillation Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4592c602-423d-4e8b-ac5d-c0dcdd8719ef · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 224
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5d64ef3b-3bae-482f-8f5a-2ea59e43cdd5 · inbound
ReNIO: Reweighting Negative Trajectory Importance for LLM On-Policy Distillation Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cf8f9fac-b21d-42c9-9ea3-1245d14be7bc · inbound
Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4bcd50ff-c2a1-4cbd-a5f9-ab7fe76b7ae7 · inbound
PointVG-R: Internalizing Geometric Reasoning in MLLMs for Precise Pointing Localization via Visual Chain of Thought Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6395ed83-137f-43e6-ba83-af5a45b904f4 · inbound
BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a9dd725b-7137-4538-b79b-8bc1416704ef · inbound
Evo-PI: Aligning Medical Reasoning via Evolving Principle-Guided Supervision Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 160329a2-0799-4988-9f09-815b43d1694f · inbound
HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15d9a061-9375-45f8-9d03-7f72e2f173fb · inbound
From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 249
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a3c556d4-8710-45e5-8b1c-d7bf522a8c8a · inbound
OpenProver: Agentic and Interactive Theorem Proving with Lean 4 Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2eb0d7ae-ad35-43ee-96db-33d83135d706 · inbound
The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 263
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ea837e8-8479-47c6-b016-841b48f2eb7a · inbound
Think Through a Bottleneck: Hourglass Reasoning for Rigorous Induction Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 971d7c0f-ea51-4964-9eb0-fd15aeb544a9 · inbound
Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 103
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f0e6076-2c99-483a-921b-6606026c0a5e · inbound
PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c88179b-b75e-4a29-a01d-1511f31d0b37 · inbound
TraversRL: Traversable Pedestrian Pathway Generation With Reinforcement Learning Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7eeed5c-9805-4e0c-ba32-16f4e5399aac · inbound
SLPO: Scaling Latent Reasoning via a Surrogate Policy Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ccb2f31-6291-4be4-81d2-34523d9259e8 · inbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f4202ad-7ff8-481b-9067-c516c9cb5121 · inbound
Bridging Compute- and Data-Optimal Pretraining Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 105
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e44d51b0-8ae4-4ad0-9230-592978058a1f · inbound
Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7df2fb86-6cf9-48a0-be5c-c82780243a82 · inbound
LEEPS: Latent-Guided Explore-Exploit Prompt Sampling for Efficient RLVR in Large Language Models Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57cb7ae4-e6b9-4b1b-bdc4-76ccba8c7c9b · inbound
Uncertainty-Aware Simulation-Based Inference for Operations Research with Large Language Models Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7cabc12-5fd3-4ad5-9e9c-693f2d03cd75 · inbound
LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.