Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T22:57:45.104367Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 8 inbound Pith citation observations for arXiv:2509.06923.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T22:57:45.104367Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T06:47:16.811511Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T20:48:56.162598Z
47 of 47 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 93596ae4-2186-456b-8029-e254c01e1a43 · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f01f223-f52c-4166-aade-c1a4a12ac517 · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Item Response Theory -- A Statistical Framework for Educational and Psychological Measurement , August 2021
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f72a7458-a5a3-4c0d-b06e-ada7f85cd201 · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Le, Sergey Levine, and Yi Ma
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0c404b07-eea7-406f-806f-656d7a4b564f · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Think you have Solved Question Answering ? Try ARC , the AI2 Reasoning Challenge , March 2018
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 42a70a16-2e3f-4666-839f-92788dbf398b · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Training Verifiers to Solve Math Word Problems , November 2021
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d9d605fe-c2cf-4b35-89ee-81d908e2c15c · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a698d10c-3860-45c8-aeaa-acfabdc15edf · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 99a581eb-bb46-4b04-b1fd-195bbee553d2 · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Improving RL Exploration for LLM Reasoning through Retrospective Replay , July 2025
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b5f9c3cd-c6c7-4d1f-bda7-b4f83407fd66 · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding SRFT : A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning , June 2025
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ee9c28cb-9cb1-48b0-92d5-219e85d86a43 · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2f0548de-d50d-44aa-9e93-ff7fe17aa989 · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Navigate the Unknown : Enhancing LLM Reasoning with Intrinsic Motivation Guided Exploration , July 2025
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 64a1f6c6-343e-4a1f-b2df-68829cac0f7b · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding OlympiadBench : A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fc34b23-ffb7-43b1-9b4c-c1d52cc4a21c · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding DeepMath-103K : A Large-Scale , Challenging , Decontaminated , and Verifiable Mathematical Dataset for Advancing Reasoning , May 2025
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 128cd7d8-352d-48c6-b544-6c6f867fa5d3 · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Measuring Mathematical Problem Solving With the MATH Dataset
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e6ed8466-8181-49ef-94a1-76896f5398ea · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Mathruler
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 83710495-76b5-4dd1-a7ee-0d251dee9197 · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Boosting MLLM Reasoning with Text-Debiased Hint-GRPO , June 2025 a
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 258492ce-cfe1-47b0-9d46-7284a74843c5 · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Ponti, and Ivan Titov
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1523fefe-0cc4-4194-b2bd-eafbbf98d6b1 · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding o pf, Yannic Kilcher, Dimitri Von R \
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2a5ccdd0-5724-4b9e-b8ed-9c913285f29f · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Solving Quantitative Reasoning Problems with Language Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c8abb6ee-5f90-4724-b2d2-bcf58a1be024 · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Numinamath
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6586b87-5bf7-476e-a3d2-6ecf533b3494 · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding UFT : Unifying Supervised and Reinforcement Fine-Tuning , May 2025 a
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 52566960-7312-407a-bcaa-08293655759c · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Understanding R1-Zero-Like Training : A Critical Perspective , March 2025 b
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d4f5f2dc-0568-401a-84ba-a8e30d5b49e6 · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Lmfit: Non-linear least-squares minimization and curve-fitting for python, July 2025
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 351e1c8d-a033-4fbb-8e11-123c0533b1c4 · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 27e2390e-3676-4df5-b62c-41388186a7f7 · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Qwen2.5 Technical Report , January 2025
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 440ee493-f604-4abd-8b83-3b4186e47f61 · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0b1990f7-f032-4d9a-bf1d-04074a768a88 · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding LLMs are Greedy Agents : Effects of RL Fine-tuning on Decision-Making Abilities , April 2025
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation da49e5ef-5b6b-4446-81c8-402d62e139e1 · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8a1355ea-9f4b-4af1-bb47-3b9076efcf10 · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding HybridFlow : A Flexible and Efficient RLHF Framework
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a9b3187-dec7-4f36-9d57-9b7164f1c263 · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Uncertainty and influence aware reward model refinement for reinforcement learning from human feedback
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c4afeeb9-bb6e-4f67-bcc8-36fa5a480d1e · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a042cc48-e491-4367-a15f-ba10ddfc2889 · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Reinforcement Learning for Reasoning in Large Language Models with One Training Example , May 2025
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7d642c7a-26ee-4614-a887-56f5d4f13efc · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding MMLU-Pro : A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f3a9728c-70db-4145-b27d-c442f6cf7f71 · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Thought- Augmented Policy Optimization : Bridging External Guidance and Internal Capabilities , May 2025
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation dc81abd4-de63-48b4-ad50-b15edc21dfc4 · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Learning to Reason under Off-Policy Guidance , May 2025
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7de74d27-420a-46d5-accb-76d188db3c9a · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding DAPO : An Open-Source LLM Reinforcement Learning System at Scale , May 2025
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation fbd018f8-824d-4551-a4b1-a7e4d18bc2e7 · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model ?, May 2025
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a1ee149a-4753-4272-9eb4-ef54db05394b · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding SimpleRL-Zoo : Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild , August 2025
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 325aeb21-5b34-4a23-8e05-3647e2110441 · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding StepHint : Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason , July 2025 a
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d7169412-fd5a-4bec-adf4-814bad7d7b24 · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bef1975-8d57-4f09-9fb7-75dd75b0e87f · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding On-policy rl meets off-policy experts: Harmonizing supervised fine-tuning and reinforcement learning via dynamic weighting
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50486305-c0d1-4b4e-965b-454f40074b50 · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Echo Chamber : RL Post-training Amplifies Behaviors Learned in Pretraining , August 2025
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 22592da0-e03e-4ceb-9c7a-c9e9b6d2cb2e · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Group Sequence Policy Optimization , July 2025
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5fec02c9-cf42-4118-a9b8-5cef6f0466b6 · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding write newline
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3edaec59-f97c-47cd-878a-02b920373b33 · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding @esa (Ref
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5219136e-5216-4351-8e00-d35529b61fd5 · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Unresolved cited work
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aace86a0-1040-45b1-9f85-196a0abd9acb · outbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Unresolved cited work
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bae0d82-242d-44fe-b444-564415b7a378 · inbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f658452d-372c-446c-8f25-3536e859720a · inbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3eb694fb-4925-4ce6-9c72-fe6ffc122af1 · inbound
Rollout Pass-Rate Control: Steering Binary-Reward RL Toward Its Most Informative Regime Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5555d0e5-89ba-4846-9d58-27c7ba9e3fed · inbound
Rollout Pass-Rate Control: Steering Binary-Reward RL Toward Its Most Informative Regime Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 29675988-4aa4-4541-a9a8-fc710c32bfec · inbound
Rollout Pass-Rate Control: Steering Binary-Reward RL Toward Its Most Informative Regime Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1741a719-b144-4594-bf07-22b160e3657b · inbound
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5c8df385-a953-4309-8779-b0c3f383755d · inbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26c06b32-36c7-43d3-ba09-8a252be4e03a · inbound
AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.