Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-27T00:48:56.892634Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2606.17735.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-27T00:48:56.892634Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
37 of 37 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bb33b206-c858-4eb3-88ed-e13ab3645cda · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d5faa1f9-4e02-4910-85cd-3fda3b150e26 · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs A Survey of Reinforcement Learning for Large Reasoning Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ca03c586-95e5-467b-a080-52bc8ff08e3b · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c52240c1-8467-4371-9340-c9db88b9fb3a · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 60250b10-1048-4cf3-a447-b61dc719b8d5 · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Large language model reasoning failures.arXiv preprint arXiv:2602.06176
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 10196b18-3880-4df7-a95e-e48a0fa945f5 · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs arXiv preprint arXiv:2602.01288 , year=
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ce81ecab-1249-4d57-ba80-818dd8c6825b · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Reasoning Can Be Restored by Correcting a Few Decision Tokens
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4bf21333-ba45-47d0-8aa3-d524641dcc6a · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs YuLan: An Open-source Large Language Model
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3e9b128a-5262-44cf-a893-034adde9778d · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs org/abs/2603.16790
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5eb652c5-a99a-4dd6-8538-e5b53434f88c · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Iquest-coder-v1 technical report.arXiv preprint arXiv:2603.16733, 2026b
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ea62d952-3925-4f51-9dc8-d2fe78c4813d · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f13d3bd3-8636-49bf-9f23-ca2825a91b1d · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Reward under attack: Analyzing the robustness and hackability of process reward models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9aac31a8-6bc2-4a4a-a4b0-ac0388280c0b · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fca526ff-4e66-4cca-8a5d-6cc9e535664d · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs More test-time compute can hurt: Overestimation bias in llm beam search.arXiv preprint arXiv:2603.15377,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4dd49caf-9621-40ee-ad3f-5fe8b710e869 · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs From Curiosity to Caution: Mitigating Reward Hacking for Best-of-N with Pessimism
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b1a72380-ad28-4828-af6e-72db290a8683 · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Junbo Li, Peng Zhou, Rui Meng, Meet P
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 47791efb-a617-49b5-bc38-87f1fdeec26e · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Int: Self-proposed interventions enable credit assignment in llm reasoning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c7fe5dad-8229-4c64-84c9-62bd1b7988c1 · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation aa43bd65-83b7-400c-83d7-429dca86b6c5 · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs ArborKV: Structure-Aware KV Cache Management for Scaling Tree-based LLM Reasoning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3315d004-c551-43b2-9254-19e8a79e0799 · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Qwen3 Technical Report
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b09bb816-eef6-443a-bba6-cee66ed2240f · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Closed-loop transformers: Autoregressive modeling as iterative latent equilibrium.arXiv preprint arXiv:2511.21882,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2369deb1-4c14-4f6b-8905-f0fb3c79a0bd · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Mmrpt: Multimodal reinforcement pre-training via masked vision-dependent reasoning.arXiv preprint arXiv:2512.07203, 2025c
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 70e36fa2-6a44-4fcb-8375-76f2025b988d · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs arXiv preprint arXiv:2603.09065 , year=
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 82e3a9ec-f72b-42da-b747-44ef9f62ab08 · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4e5c98fd-1684-4b7a-b208-b492eb431894 · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Early Decisions Matter: Proximity Bias and Initial Trajectory Shaping in Non-Autoregressive Diffusion Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8ddeb62a-f527-43c7-b8ce-79a8a1b90ca1 · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Corefine: Confidence-guided self-refinement for adaptive test-time compute.arXiv preprint arXiv:2602.08948, 2026a
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fa684804-bab1-498f-a8d2-6b28a0d03fa7 · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Understanding reasoning in thinking language models via steering vectors.arXiv preprint arXiv:2506.18167
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a84dbcc5-af56-4be8-99a1-7aa055112616 · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs PILOT: Planning via Internalized Latent Optimization Trajectories for Large Language Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 00686bbf-8d21-4b3c-9757-0d3be16d8793 · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs HiMAC: Hierarchical Macro-Micro Learning for Long-Horizon LLM Agents
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fc57e249-3b1a-4fa1-828e-a015484f138d · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Probabilistic soundness guarantees in llm reasoning chains
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89b0f538-d471-446b-ba6b-a47cfcb5424d · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Prism: Pushing the frontier of deep think via process reward model-guided inference.arXiv preprint arXiv:2603.02479,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 65c61e8d-c18c-4ffd-94d5-6f94025a9aec · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs v 1: Unifying generation and self-verification for parallel reasoners.arXiv preprint arXiv:2603.04304, 2026a
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 365dd4e4-44fb-4a6a-980e-69e8631d0689 · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 10591c9e-6fd0-420f-a861-07e92f28ec32 · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Group Sequence Policy Optimization
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8124e925-d29c-4936-a827-8f80984bd642 · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Disco: Re- inforcing large reasoning models with discriminative con- strained optimization.arXiv preprint arXiv:2505.12366
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 24d57e2e-a177-4eea-86b4-6178b2760c23 · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs The proposed method does not involve human subjects, private user data, or personally identifiable information
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c5c2945-2225-4fb8-a332-9faebfa4a26b · outbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
No inbound Pith citation observations are available.