Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T18:53:02.820543Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 11 inbound Pith citation observations for arXiv:2509.09675.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T18:53:02.820543Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T10:36:23.165454Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T16:39:57.677897Z
34 of 34 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7c8d4574-d36e-4a66-9513-d4dc796d50d7 · outbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models A Survey of Exploration Methods in Reinforcement Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7086f6d9-3732-48fc-84fc-8e693773f390 · outbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models std ´␣ ϕJpwpkq n,h ˇˇ1ďkďK (¯ . Elliptical (“count-based
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5aa8ec28-2077-4742-851a-d95c0b788c38 · outbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models Qwen3-4B-Base-GRPO
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e18ab364-7475-43b4-9f0e-aa8b67c2eb79 · outbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6747d781-40bb-453b-a586-a9ba2a630f5f · outbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models Breach in the Shield: Unveiling the Vulnerabilities of Large Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7a85f9e-f9a8-4575-8593-467edb25e281 · outbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edf8c7ba-09fc-4dea-9987-691406fe9548 · outbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82224b45-c9a0-41ba-b27b-66d47ad1791c · outbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6d24b4f-2202-4ed8-8b6c-22d35406aa26 · outbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models Continuous control with deep reinforcement learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90686f68-f452-49b0-a015-954971bc96dc · outbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3fe07f1-8719-4bc0-9cad-4ef046aac618 · outbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models Proximal Policy Optimization Algorithms
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d2781dd-027f-4a65-a4c0-b329ef2326e0 · outbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85d9f907-6cf2-486c-9306-6d340ed9c89a · outbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models On entropy control in llm-rl algorithms.arXiv preprint arXiv:2509.03493,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f4aa9d9-897b-4962-9273-2271b2e68222 · outbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models Thermometer: Towards Universal Calibration for Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 555c6c82-7c77-4038-a2f9-43400cb59813 · outbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models Solving math word problems with process- and outcome-based feedback
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f22aacd1-a65f-4627-9d38-08d962735cec · outbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models LiteSearch: Efficacious Tree Search for LLM
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation beefff04-48de-4cf9-b381-03c75111e970 · outbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models Towards Self-Improvement of LLMs via MCTS: Leveraging Stepwise Knowledge with Curriculum Preference Learning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c00dd1be-6d63-497a-aaf2-83d01e19e1ec · outbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models Qwen3 Technical Report
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14604fba-6511-4875-b9b9-6e734e0d09cb · outbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 785be2f1-8e1f-4b8f-86da-ff8b432162eb · outbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f693e32-c32b-4e35-8cc7-c3c537d19128 · outbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models One Token to Fool LLM-as-a-Judge
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 832c0674-f33c-4253-8d81-84f819a15d64 · outbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models Learning to Reason via Mixture-of-Thought for Logical Reasoning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 398b7170-af66-43e5-b8c6-f39c12d228aa · outbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training
Reference 1997
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b36915be-0427-4923-9bf1-971eaa09ee47 · outbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 2010
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9bacd8f-e76d-40d2-82f6-9ca30d510465 · outbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models Navigate the unknown: Enhancing llm reasoning with intrinsic motivation guided exploration.arXiv preprint arXiv:2505.17621,
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c617351-2487-4a29-9d57-4246570be9c3 · outbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12a00344-c489-4e71-83da-94b86ecbb9da · outbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models Measuring Mathematical Problem Solving With the MATH Dataset
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d9c5d8f-b005-415a-a257-49ea048cacfc · outbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models Training Verifiers to Solve Math Word Problems
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fe00bae-31bc-4acd-938e-f62307b1658f · outbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models Supervising the search process produces reliable and generalizable information-seeking agents
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0ebcc62-be89-4ba5-b101-e6d25f04a06e · outbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models Online Preference Alignment for Language Models via Count-based Exploration
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cdec345-3fa4-4bd6-ba62-80192386feb1 · outbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models Calibrating Large Language Models Using Their Generations Only
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8f1606b-f184-41d0-b092-a942c03e5d23 · outbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models Deep Think with Confidence
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fce78d2-28a7-4bd5-b754-c5491bdec338 · outbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a58071c-5895-4ae0-ab48-1c5826335408 · outbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models Exploration by Random Network Distillation
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09f464ea-0ad7-4475-8c3d-19cd3b1ec900 · inbound
StatEval: A Comprehensive Benchmark for Large Language Models in Statistics CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93ca5813-c903-41f5-8c9b-4dd664148796 · inbound
Calibration-Aware Policy Optimization for Reasoning LLMs CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c50ad9dd-2bad-4881-a5e8-43369204a2e7 · inbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 45a82209-4970-43e8-9141-2cf073ebb7ec · inbound
Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 368f6c5e-2403-4604-9882-89a8c77393e9 · inbound
Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 236cf7f9-3abe-43d0-8427-5a59be816a68 · inbound
Reinforcing Multimodal Reasoning Against Visual Degradation CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9d32ec0f-a3ba-43e5-bbd1-5de9823bc9c3 · inbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3039b200-07d3-4b2c-a9be-e5fcab8a40e4 · inbound
Epistemic Uncertainty for Test-Time Discovery CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1358ee36-5a8f-4a5b-a9b9-66951295ece0 · inbound
SALT: When More Rollouts Don't Help in Group-Based Policy Optimization and How to Make Them Matter CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 92d15d31-edab-49ac-955e-6afb9bba302f · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 59aabf2b-ef39-4e89-967b-6d107da89b62 · inbound
ExTra: Exploratory Trajectory Optimization for Language Model Reinforcement Learning CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.