Pith. sign in

Paper Citation Record · LEDGER

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs

As of 6 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2606.17735.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.17735 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T00:48:56.892634Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact32
  • verified fuzzy0
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bb33b206-c858-4eb3-88ed-e13ab3645cda · outbound

This paper cites Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:08:58.849123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:a83cbb97c230e9fb9f414af6e4a0d936f570f64c000864e607660fa5adbe839d

Observation d5faa1f9-4e02-4910-85cd-3fda3b150e26 · outbound

This paper cites A Survey of Reinforcement Learning for Large Reasoning Models.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs A Survey of Reinforcement Learning for Large Reasoning Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:08:58.861020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:4e74cd0fe3ecc641a1514164d5d91b02442ca3179efbf9707805a40d0dc00ff8

Observation ca03c586-95e5-467b-a080-52bc8ff08e3b · outbound

This paper cites Reinforcement Learning for Long-Horizon Interactive LLM Agents.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:18:58.371419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:3057f3ec4e72c6a7b71e0fad1c5063a9a2308b38e7c401e06a6dc8c1c88c5f05

Observation c52240c1-8467-4371-9340-c9db88b9fb3a · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:18:58.398809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:9be0086e60c6dbe8c685f3a120593dedc124c89661d5a323f1df18fed09db1b2

Observation 60250b10-1048-4cf3-a447-b61dc719b8d5 · outbound

This paper cites Large language model reasoning failures.arXiv preprint arXiv:2602.06176.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Large language model reasoning failures.arXiv preprint arXiv:2602.06176

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T21:18:58.374020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:e2009e23a0af430acb43074924d96075388332b278539851730fa4ae1f4ca914

Observation 10196b18-3880-4df7-a95e-e48a0fa945f5 · outbound

This paper cites arXiv preprint arXiv:2602.01288 , year=.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs arXiv preprint arXiv:2602.01288 , year=

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:18:58.392439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:30eabeb775899a717b797207826a6d0195f194dc98edb0f4b2f807194ceed18f

Observation ce81ecab-1249-4d57-ba80-818dd8c6825b · outbound

This paper cites Reasoning Can Be Restored by Correcting a Few Decision Tokens.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Reasoning Can Be Restored by Correcting a Few Decision Tokens

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:18:58.340911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:c5ad06b64e0ee61e8b4b354f81313093239768b2ee8a582ed288928ec150ec85

Observation 4bf21333-ba45-47d0-8aa3-d524641dcc6a · outbound

This paper cites YuLan: An Open-source Large Language Model.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs YuLan: An Open-source Large Language Model

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:18:58.386862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:eb9ecbeaeb6e64b05faecec6e8601acac6cfb6121962587b1392c611b88561c5

Observation 3e9b128a-5262-44cf-a893-034adde9778d · outbound

This paper cites org/abs/2603.16790.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs org/abs/2603.16790

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T21:18:58.383757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:247729bd7a72556e7d35d15f49a6a599ab2a249bdb55f97b5b399eca4f369c5d

Observation 5eb652c5-a99a-4dd6-8538-e5b53434f88c · outbound

This paper cites Iquest-coder-v1 technical report.arXiv preprint arXiv:2603.16733, 2026b.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Iquest-coder-v1 technical report.arXiv preprint arXiv:2603.16733, 2026b

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:18:58.384774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:609262d726273e19e3004099934f749ab375454577045d165adcbc2bde2b03d2

Observation ea62d952-3925-4f51-9dc8-d2fe78c4813d · outbound

This paper cites A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:18:58.377454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:649ad321a9ef59f90529f506a8b8debad7b3643b4c34dfcdd8ea3850a628a8dd

Observation f13d3bd3-8636-49bf-9f23-ca2825a91b1d · outbound

This paper cites Reward under attack: Analyzing the robustness and hackability of process reward models.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Reward under attack: Analyzing the robustness and hackability of process reward models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:18:58.359171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:42348dc75b45b7979e4460fd28b6494350baec9c99a025f8f3496d618366cd35

Observation 9aac31a8-6bc2-4a4a-a4b0-ac0388280c0b · outbound

This paper cites Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:18:58.372138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:104ad5f9fdcee5b5ef326f206ae186ffc6bce42678fc0b034aa9371c3f0fe6f6

Observation fca526ff-4e66-4cca-8a5d-6cc9e535664d · outbound

This paper cites More test-time compute can hurt: Overestimation bias in llm beam search.arXiv preprint arXiv:2603.15377,.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs More test-time compute can hurt: Overestimation bias in llm beam search.arXiv preprint arXiv:2603.15377,

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:18:58.402092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:819681e1c871b9b44b82f18094a1956be4743b50effe8a2a5b47246a0196694b

Observation 4dd49caf-9621-40ee-ad3f-5fe8b710e869 · outbound

This paper cites From Curiosity to Caution: Mitigating Reward Hacking for Best-of-N with Pessimism.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs From Curiosity to Caution: Mitigating Reward Hacking for Best-of-N with Pessimism

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:18:58.404824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:0974158a229b9ce843f1fe79f1cabf89e7b5b2d3de28a893b8371739131195cb

Observation b1a72380-ad28-4828-af6e-72db290a8683 · outbound

This paper cites Junbo Li, Peng Zhou, Rui Meng, Meet P.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Junbo Li, Peng Zhou, Rui Meng, Meet P

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T21:18:58.333883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:9f69eed0f6824a37019e129417188f40255611e61cf073ab180d1396dec7288c

Observation 47791efb-a617-49b5-bc38-87f1fdeec26e · outbound

This paper cites Int: Self-proposed interventions enable credit assignment in llm reasoning.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Int: Self-proposed interventions enable credit assignment in llm reasoning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:18:58.387573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:5f10356d87421a186450c6cce682d2a6a76e748c8ee58b9ba994782de9317e99

Observation c7fe5dad-8229-4c64-84c9-62bd1b7988c1 · outbound

This paper cites Adaptive KV Cache Reuse for Fast Long-Context LLM Serving.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Adaptive KV Cache Reuse for Fast Long-Context LLM Serving

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:18:58.407427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:3a65107ec01116c9b73ccaaf2197974aeb6c39244cb09418a4c3200c6b0dbfd1

Observation aa43bd65-83b7-400c-83d7-429dca86b6c5 · outbound

This paper cites ArborKV: Structure-Aware KV Cache Management for Scaling Tree-based LLM Reasoning.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs ArborKV: Structure-Aware KV Cache Management for Scaling Tree-based LLM Reasoning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:08:58.857588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:f41170414ff0fc0546b739408146cdf13b44f0e259d64bea9b793e339257f2f6

Observation 3315d004-c551-43b2-9254-19e8a79e0799 · outbound

This paper cites Qwen3 Technical Report.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Qwen3 Technical Report

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:18:58.389583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:710056d7c15827f674698a11d87ea191669b2e3b2f054950cfb32cf362ae3e84

Observation b09bb816-eef6-443a-bba6-cee66ed2240f · outbound

This paper cites Closed-loop transformers: Autoregressive modeling as iterative latent equilibrium.arXiv preprint arXiv:2511.21882,.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Closed-loop transformers: Autoregressive modeling as iterative latent equilibrium.arXiv preprint arXiv:2511.21882,

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:18:58.379863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:37625a134a597aeb42f7c704abaa53403d9714326878603d083f132dd4108720

Observation 2369deb1-4c14-4f6b-8905-f0fb3c79a0bd · outbound

This paper cites Mmrpt: Multimodal reinforcement pre-training via masked vision-dependent reasoning.arXiv preprint arXiv:2512.07203, 2025c.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Mmrpt: Multimodal reinforcement pre-training via masked vision-dependent reasoning.arXiv preprint arXiv:2512.07203, 2025c

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:18:58.380481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:31a589ec648082d9c9426c044d8d94b7a4c0fce7f30f9b69df65a30786dcf04f

Observation 70e36fa2-6a44-4fcb-8375-76f2025b988d · outbound

This paper cites arXiv preprint arXiv:2603.09065 , year=.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs arXiv preprint arXiv:2603.09065 , year=

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:18:58.356277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:ab3f5700fc45fae012e63debe108265ba032da692e39626fe043fb0793fc51c5

Observation 82e3a9ec-f72b-42da-b747-44ef9f62ab08 · outbound

This paper cites AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:18:58.374761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:45c679d8242660c09555a0b3ee9a12144a27e13e9eeca6952d917096cd83f114

Observation 4e5c98fd-1684-4b7a-b208-b492eb431894 · outbound

This paper cites Early Decisions Matter: Proximity Bias and Initial Trajectory Shaping in Non-Autoregressive Diffusion Language Models.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Early Decisions Matter: Proximity Bias and Initial Trajectory Shaping in Non-Autoregressive Diffusion Language Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:18:58.395252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:94fe9eb17f1e8aa2580f6f0846955f8dd4008802607af942a44b0ecd4b2a2f94

Observation 8ddeb62a-f527-43c7-b8ce-79a8a1b90ca1 · outbound

This paper cites Corefine: Confidence-guided self-refinement for adaptive test-time compute.arXiv preprint arXiv:2602.08948, 2026a.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Corefine: Confidence-guided self-refinement for adaptive test-time compute.arXiv preprint arXiv:2602.08948, 2026a

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:18:58.392569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:9c20696eb763cf1d27ac5046a2ada60d831db68fd70a8c187b9b9993060e48a2

Observation fa684804-bab1-498f-a8d2-6b28a0d03fa7 · outbound

This paper cites Understanding reasoning in thinking language models via steering vectors.arXiv preprint arXiv:2506.18167.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Understanding reasoning in thinking language models via steering vectors.arXiv preprint arXiv:2506.18167

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:18:58.353528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:3779860d085e2b8351e15c72cca713fc985fce84fdc2b254276e1faa81ac76a0

Observation a84dbcc5-af56-4be8-99a1-7aa055112616 · outbound

This paper cites PILOT: Planning via Internalized Latent Optimization Trajectories for Large Language Models.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs PILOT: Planning via Internalized Latent Optimization Trajectories for Large Language Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:18:58.382207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:53cbe9f289da1262ce949e2b3b8e5650b95879218a9cef129ace08d472e28806

Observation 00686bbf-8d21-4b3c-9757-0d3be16d8793 · outbound

This paper cites HiMAC: Hierarchical Macro-Micro Learning for Long-Horizon LLM Agents.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs HiMAC: Hierarchical Macro-Micro Learning for Long-Horizon LLM Agents

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:18:58.338366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:58555d3355a2a77446405b26aa7628fd7afbc5c1b424badb92b9db29f60c12b6

Observation fc57e249-3b1a-4fa1-828e-a015484f138d · outbound

This paper cites Probabilistic soundness guarantees in llm reasoning chains.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Probabilistic soundness guarantees in llm reasoning chains

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-27T00:48:56.892634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:d3af3bcd6929359c8a89be7429e18ed8c45baaf8f2aa6a065de58e6a191267db

Observation 89b0f538-d471-446b-ba6b-a47cfcb5424d · outbound

This paper cites Prism: Pushing the frontier of deep think via process reward model-guided inference.arXiv preprint arXiv:2603.02479,.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Prism: Pushing the frontier of deep think via process reward model-guided inference.arXiv preprint arXiv:2603.02479,

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:18:58.359524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:024affa4feb864c5bc6981d7a21324b2ce8a20c8c15c7aa30970b03a4a4eaef0

Observation 65c61e8d-c18c-4ffd-94d5-6f94025a9aec · outbound

This paper cites v 1: Unifying generation and self-verification for parallel reasoners.arXiv preprint arXiv:2603.04304, 2026a.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs v 1: Unifying generation and self-verification for parallel reasoners.arXiv preprint arXiv:2603.04304, 2026a

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:18:58.366082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:9fd7abdd516637e8d0ee4de9647ecd85618844bf513bc33dcb9d3f894611890f

Observation 365dd4e4-44fb-4a6a-980e-69e8631d0689 · outbound

This paper cites DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:18:58.389882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:7370e20db960dddeec35ce8a0ea625e01fb50e688c6846caaae00e731e391def

Observation 10591c9e-6fd0-420f-a861-07e92f28ec32 · outbound

This paper cites Group Sequence Policy Optimization.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Group Sequence Policy Optimization

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:08:58.863995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:1629da9f73a983ab243b0c922fd3abaf0a861d3daea31bb4c30928b9f2d71224

Observation 8124e925-d29c-4936-a827-8f80984bd642 · outbound

This paper cites Disco: Re- inforcing large reasoning models with discriminative con- strained optimization.arXiv preprint arXiv:2505.12366.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Disco: Re- inforcing large reasoning models with discriminative con- strained optimization.arXiv preprint arXiv:2505.12366

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:08:58.868789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:70923812577aca9b22a0ac7da87b412b32f5a6cf750a6081d93ae69fdfd1e6c7

Observation 24d57e2e-a177-4eea-86b4-6178b2760c23 · outbound

This paper cites The proposed method does not involve human subjects, private user data, or personally identifiable information.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs The proposed method does not involve human subjects, private user data, or personally identifiable information

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-27T00:48:56.892634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:6b0a803285321bc6a056fb48c7bc3257fe25b77c53b9e1683399b75736b630c9

Observation 9c5c2945-2225-4fb8-a332-9faebfa4a26b · outbound

This paper cites an unresolved cited work.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Unresolved cited work

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:08:58.871510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:58bd1cb23ea3f0cde71334445fae84f6532fb0d2045d948e9bcf7872c7cfe792

Pith citing papers

No inbound Pith citation observations are available.