Pith. sign in

Paper Citation Record · LEDGER

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding

As of 22 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 8 inbound Pith citation observations for arXiv:2509.06923.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.06923 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T22:57:45.104367Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T06:47:16.811511Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:48:56.162598Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 93596ae4-2186-456b-8029-e254c01e1a43 · outbound

This paper cites SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T22:57:44.986207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:57:44.986207Z digest=sha256:316e0c744d4ac3ae4e67f3f53765fa859c11861964b17ae2853c79e0b8bd16dd

Observation 9f01f223-f52c-4166-aade-c1a4a12ac517 · outbound

This paper cites Item Response Theory -- A Statistical Framework for Educational and Psychological Measurement , August 2021.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Item Response Theory -- A Statistical Framework for Educational and Psychological Measurement , August 2021

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:57:45.472117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:44.989509Z digest=sha256:c5c674a22ca3eec9b1dd6867f57304de92cfe55bbbc5a36271eb1a111c0a091b

Observation f72a7458-a5a3-4c0d-b06e-ada7f85cd201 · outbound

This paper cites Le, Sergey Levine, and Yi Ma.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Le, Sergey Levine, and Yi Ma

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:57:45.464284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:44.992098Z digest=sha256:97dc116a6eedb9960fe2b561209c8e45b8124506a25304261bffff36639a2e25

Observation 0c404b07-eea7-406f-806f-656d7a4b564f · outbound

This paper cites Think you have Solved Question Answering ? Try ARC , the AI2 Reasoning Challenge , March 2018.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Think you have Solved Question Answering ? Try ARC , the AI2 Reasoning Challenge , March 2018

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:57:45.456985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:44.994980Z digest=sha256:41a9b63b29798d195bd196597f7deb0286526a0e7712ec4b8f31ac1b018b3a9f

Observation 42a70a16-2e3f-4666-839f-92788dbf398b · outbound

This paper cites Training Verifiers to Solve Math Word Problems , November 2021.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Training Verifiers to Solve Math Word Problems , November 2021

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:57:45.449215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:44.997663Z digest=sha256:f5576832902e547df81d7aa41a6a1c2ad057ed627e7aaa1e6bcfbff277ef9593

Observation d9d605fe-c2cf-4b35-89ee-81d908e2c15c · outbound

This paper cites an unresolved cited work.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-04T22:57:45.441467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:45.000383Z digest=sha256:1a3721575b435a02c8a397d69a2c717c62431f3313efa826d191ef28b14dad10

Observation a698d10c-3860-45c8-aeaa-acfabdc15edf · outbound

This paper cites Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:57:45.433454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:45.003588Z digest=sha256:88b668f8df450a817f66c3547f4a808e595466cbfcdbcb80016c706e3975d5c8

Observation 99a581eb-bb46-4b04-b1fd-195bbee553d2 · outbound

This paper cites Improving RL Exploration for LLM Reasoning through Retrospective Replay , July 2025.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Improving RL Exploration for LLM Reasoning through Retrospective Replay , July 2025

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:57:45.425536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:45.006413Z digest=sha256:0ae1b95a3bd5227f8bdc6442b6c380a976b2ba1b84c41bad4c98a7ea7c179186

Observation b5f9c3cd-c6c7-4d1f-bda7-b4f83407fd66 · outbound

This paper cites SRFT : A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning , June 2025.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding SRFT : A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning , June 2025

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:57:45.417758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:45.008794Z digest=sha256:6c423c04f5011f5b7e1c63fe65526dbb75fff7df6e0bd75dedbec1f22c55ced2

Observation ee9c28cb-9cb1-48b0-92d5-219e85d86a43 · outbound

This paper cites an unresolved cited work.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-04T22:57:45.409813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:45.011241Z digest=sha256:273a83b4b1d7100f51ea4cb5fa73f57a5abcd1e7c8a84373b8de4e158dc36c24

Observation 2f0548de-d50d-44aa-9e93-ff7fe17aa989 · outbound

This paper cites Navigate the Unknown : Enhancing LLM Reasoning with Intrinsic Motivation Guided Exploration , July 2025.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Navigate the Unknown : Enhancing LLM Reasoning with Intrinsic Motivation Guided Exploration , July 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:57:45.402185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:45.013676Z digest=sha256:c384e0fc7ce3747720da293635d976d0e14e1d7c37b3f7a3b8032e045a7125a6

Observation 64a1f6c6-343e-4a1f-b2df-68829cac0f7b · outbound

This paper cites OlympiadBench : A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding OlympiadBench : A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T22:57:45.016293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:57:45.016293Z digest=sha256:45769a5dc0732eeaec18b8ad9ec735a89f4a8e2ba3eda813cad50a67a844d217

Observation 0fc34b23-ffb7-43b1-9b4c-c1d52cc4a21c · outbound

This paper cites DeepMath-103K : A Large-Scale , Challenging , Decontaminated , and Verifiable Mathematical Dataset for Advancing Reasoning , May 2025.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding DeepMath-103K : A Large-Scale , Challenging , Decontaminated , and Verifiable Mathematical Dataset for Advancing Reasoning , May 2025

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:57:45.389184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:45.018722Z digest=sha256:8cebfd78454966b890eb5f0e5fe0db3fceafe20a8f95b25b4c941d3cd6427198

Observation 128cd7d8-352d-48c6-b544-6c6f867fa5d3 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Measuring Mathematical Problem Solving With the MATH Dataset

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:57:45.381414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:45.021109Z digest=sha256:3a2845f58309076a5009859e14d774dc5f625bca5e6b7d6db4e1b2568945c5e9

Observation e6ed8466-8181-49ef-94a1-76896f5398ea · outbound

This paper cites Mathruler.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Mathruler

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:57:45.373069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:45.023497Z digest=sha256:827b734929279fe18e89dde3acb92b4f1f88081bd7541f248a45318940558db5

Observation 83710495-76b5-4dd1-a7ee-0d251dee9197 · outbound

This paper cites Boosting MLLM Reasoning with Text-Debiased Hint-GRPO , June 2025 a.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Boosting MLLM Reasoning with Text-Debiased Hint-GRPO , June 2025 a

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:57:45.364754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:45.026025Z digest=sha256:0adcee240a76057bf7183d7507c809a20e70667c40c59b01b4318441a6eabb37

Observation 258492ce-cfe1-47b0-9d46-7284a74843c5 · outbound

This paper cites Ponti, and Ivan Titov.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Ponti, and Ivan Titov

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:57:45.356909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:45.028454Z digest=sha256:b944923323ddeccc759ccf5e17ea9c0e1a77bbdbb2a78eeb14a4d63ac5fcbaf6

Observation 1523fefe-0cc4-4194-b2bd-eafbbf98d6b1 · outbound

This paper cites o pf, Yannic Kilcher, Dimitri Von R \.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding o pf, Yannic Kilcher, Dimitri Von R \

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:57:45.349070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:45.030895Z digest=sha256:87b1b33a0f51b206793ab77480f768e64b430727d2105019ce1d1001f5a83650

Observation 2a5ccdd0-5724-4b9e-b8ed-9c913285f29f · outbound

This paper cites Solving Quantitative Reasoning Problems with Language Models.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Solving Quantitative Reasoning Problems with Language Models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:57:45.341399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:45.033305Z digest=sha256:ae79de29fd4404b0861f9cb7b17090bcf55f3e5ae0955dd1b4ccf68030f2bb13

Observation c8abb6ee-5f90-4724-b2d2-bcf58a1be024 · outbound

This paper cites Numinamath.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Numinamath

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T22:57:45.035812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:57:45.035812Z digest=sha256:99e1b1e5b5d91456849acb2aa211b1f500d46886d04ff735a95a900ba060a757

Observation b6586b87-5bf7-476e-a3d2-6ecf533b3494 · outbound

This paper cites UFT : Unifying Supervised and Reinforcement Fine-Tuning , May 2025 a.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding UFT : Unifying Supervised and Reinforcement Fine-Tuning , May 2025 a

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:57:45.327724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:45.038216Z digest=sha256:241197fbc4123232cb1c032ca84803e58afd81d9f0df3a20747ce7dc553dd6ac

Observation 52566960-7312-407a-bcaa-08293655759c · outbound

This paper cites Understanding R1-Zero-Like Training : A Critical Perspective , March 2025 b.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Understanding R1-Zero-Like Training : A Critical Perspective , March 2025 b

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:57:45.320373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:45.040623Z digest=sha256:79ed15bea785eb16af99cb40acca677f1a26b8550cb65fd611dc7aa2aede175d

Observation d4f5f2dc-0568-401a-84ba-a8e30d5b49e6 · outbound

This paper cites Lmfit: Non-linear least-squares minimization and curve-fitting for python, July 2025.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Lmfit: Non-linear least-squares minimization and curve-fitting for python, July 2025

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T22:57:45.042861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:57:45.042861Z digest=sha256:a633b7840d863f879837ffa66ac9c35d21d389e88b780c3e65115dbb4d3ef0f3

Observation 351e1c8d-a033-4fbb-8e11-123c0533b1c4 · outbound

This paper cites an unresolved cited work.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-04T22:57:45.313024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:45.045203Z digest=sha256:3f8cf4679291b1aa61f3bad230a7451545c8020415f11fc763eeb568ec6273f3

Observation 27e2390e-3676-4df5-b62c-41388186a7f7 · outbound

This paper cites Qwen2.5 Technical Report , January 2025.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Qwen2.5 Technical Report , January 2025

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:57:45.305385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:45.047861Z digest=sha256:e5424238674f24245270b2508db16687db824b5e8bcfde37d2d9948479b6ae0c

Observation 440ee493-f604-4abd-8b83-3b4186e47f61 · outbound

This paper cites an unresolved cited work.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-04T22:57:45.297896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:45.050279Z digest=sha256:56599dace090f31d673d8c5b564b332176543689904701fa61df810b0a483306

Observation 0b1990f7-f032-4d9a-bf1d-04074a768a88 · outbound

This paper cites LLMs are Greedy Agents : Effects of RL Fine-tuning on Decision-Making Abilities , April 2025.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding LLMs are Greedy Agents : Effects of RL Fine-tuning on Decision-Making Abilities , April 2025

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:57:45.290551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:45.052662Z digest=sha256:fd196f49a5978969d3dd15aed48aa0566bc469eda2a6814847fa341832b86ce6

Observation da49e5ef-5b6b-4446-81c8-402d62e139e1 · outbound

This paper cites an unresolved cited work.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-04T22:57:45.283399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:45.055037Z digest=sha256:ac624ee2c7647d7edaebbed75039f6cdbbaffe026c7e4130b29f32bc9bc47fb5

Observation 8a1355ea-9f4b-4af1-bb47-3b9076efcf10 · outbound

This paper cites HybridFlow : A Flexible and Efficient RLHF Framework.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding HybridFlow : A Flexible and Efficient RLHF Framework

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T22:57:45.057545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:57:45.057545Z digest=sha256:de62b9886ef517e8cdfae9d9b8790db093f8c5757d91a29199b33ef5ade9233c

Observation 5a9b3187-dec7-4f36-9d57-9b7164f1c263 · outbound

This paper cites Uncertainty and influence aware reward model refinement for reinforcement learning from human feedback.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Uncertainty and influence aware reward model refinement for reinforcement learning from human feedback

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:57:45.276390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:45.060059Z digest=sha256:e6d9fee1a60781165672b5042b7bf609fe402c724567244eaa8abbf59dc2deaa

Observation c4afeeb9-bb6e-4f67-bcc8-36fa5a480d1e · outbound

This paper cites an unresolved cited work.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-04T22:57:45.268789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:45.062279Z digest=sha256:7f1a1beff172657f6eaec761bfee2e83dc71ac1e2082e385665cfabf853618b3

Observation a042cc48-e491-4367-a15f-ba10ddfc2889 · outbound

This paper cites Reinforcement Learning for Reasoning in Large Language Models with One Training Example , May 2025.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Reinforcement Learning for Reasoning in Large Language Models with One Training Example , May 2025

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:57:45.261294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:45.064634Z digest=sha256:ebdc08e28643a31b32886e9fce06c737dc55514655bc9745ac3453db3a51ac0c

Observation 7d642c7a-26ee-4614-a887-56f5d4f13efc · outbound

This paper cites MMLU-Pro : A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding MMLU-Pro : A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:57:45.253931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:45.067201Z digest=sha256:694ca193d46e1bee59ec02b03d48ebd04c0ec24a3acef9825e47bb6bc6820ae9

Observation f3a9728c-70db-4145-b27d-c442f6cf7f71 · outbound

This paper cites Thought- Augmented Policy Optimization : Bridging External Guidance and Internal Capabilities , May 2025.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Thought- Augmented Policy Optimization : Bridging External Guidance and Internal Capabilities , May 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:57:45.246019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:45.069522Z digest=sha256:1b2e2dedcde1aa2a5a0bdd83d9e30591b3b35de55e01cadd4a10b90653c8f5ef

Observation dc81abd4-de63-48b4-ad50-b15edc21dfc4 · outbound

This paper cites Learning to Reason under Off-Policy Guidance , May 2025.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Learning to Reason under Off-Policy Guidance , May 2025

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:57:45.238451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:45.071834Z digest=sha256:990979a3a2c0c53ca745388c19069aaf9caccacc88ee3ff042661a067efc861b

Observation 7de74d27-420a-46d5-accb-76d188db3c9a · outbound

This paper cites DAPO : An Open-Source LLM Reinforcement Learning System at Scale , May 2025.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding DAPO : An Open-Source LLM Reinforcement Learning System at Scale , May 2025

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:57:45.230201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:45.074426Z digest=sha256:191bbbafaa13325073462162c9d38df1fd6e278cef09f92d636dc52a6e168cd5

Observation fbd018f8-824d-4551-a4b1-a7e4d18bc2e7 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model ?, May 2025.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model ?, May 2025

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:57:45.221829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:45.076905Z digest=sha256:677934e7b741cba2338d22725c0c823e6049370974c3569dfe93ddc59d35ae31

Observation a1ee149a-4753-4272-9eb4-ef54db05394b · outbound

This paper cites SimpleRL-Zoo : Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild , August 2025.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding SimpleRL-Zoo : Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild , August 2025

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:57:45.213621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:45.079396Z digest=sha256:1c6fd9c02f2e70957d60171d11794643593729c2f7e5ccc52bf725f392094d97

Observation 325aeb21-5b34-4a23-8e05-3647e2110441 · outbound

This paper cites StepHint : Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason , July 2025 a.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding StepHint : Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason , July 2025 a

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:57:45.205387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:45.081843Z digest=sha256:721fc222bbb31c495d6d77f99fb5d38dfc7b2241a3da8199cc8b0fbaf699e3c2

Observation d7169412-fd5a-4bec-adf4-814bad7d7b24 · outbound

This paper cites A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T22:57:45.084185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:57:45.084185Z digest=sha256:0ef8eae4b1cca99c96a173435175a04529b11544be0690ed7a4e4a2d4e5e12d0

Observation 3bef1975-8d57-4f09-9fb7-75dd75b0e87f · outbound

This paper cites On-policy rl meets off-policy experts: Harmonizing supervised fine-tuning and reinforcement learning via dynamic weighting.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding On-policy rl meets off-policy experts: Harmonizing supervised fine-tuning and reinforcement learning via dynamic weighting

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T22:57:45.088987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:57:45.088987Z digest=sha256:75f15f6e726f5389bc7ebfc70884eec4990d611c98cf8d2f0e16d48f04dc0cbf

Observation 50486305-c0d1-4b4e-965b-454f40074b50 · outbound

This paper cites Echo Chamber : RL Post-training Amplifies Behaviors Learned in Pretraining , August 2025.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Echo Chamber : RL Post-training Amplifies Behaviors Learned in Pretraining , August 2025

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:57:45.197229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:45.091369Z digest=sha256:946f838153b4de3ab9e65adfc933525a213a25f986814fca2bef15beefe6b1ce

Observation 22592da0-e03e-4ceb-9c7a-c9e9b6d2cb2e · outbound

This paper cites Group Sequence Policy Optimization , July 2025.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Group Sequence Policy Optimization , July 2025

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:57:45.188053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-04T22:57:45.093647Z digest=sha256:d2cd2cee8ff1ae695604f43d6c8b5e3447356b62f26670fc6b25a8ce7598bad7

Observation 5fec02c9-cf42-4118-a9b8-5cef6f0466b6 · outbound

This paper cites write newline.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding write newline

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T22:57:45.096137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:57:45.096137Z digest=sha256:ed0552a4ebee381f7acae4ea9f8a3e1d5ea20fac6a70ebe599c1b1ff43e2c64a

Observation 3edaec59-f97c-47cd-878a-02b920373b33 · outbound

This paper cites @esa (Ref.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding @esa (Ref

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T22:57:45.099078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:57:45.099078Z digest=sha256:3a5359921a3e5a37eb9c030512122e610b36ef6ab92df505670eae91b6cd7027

Observation 5219136e-5216-4351-8e00-d35529b61fd5 · outbound

This paper cites an unresolved cited work.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T22:57:45.101843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:57:45.101843Z digest=sha256:d5ad2b1cee6c6981e050d400fdd7c0bc9d88fb77df41692834fdd457aa6725ee

Observation aace86a0-1040-45b1-9f85-196a0abd9acb · outbound

This paper cites an unresolved cited work.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T22:57:45.104367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:57:45.104367Z digest=sha256:7f6d84f19b996907f8b15dfa9aa9c7e91bcb60b05ef75ca714b73380db52f24c

Pith citing papers

Observation 6bae0d82-242d-44fe-b444-564415b7a378 · inbound

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards cites this paper.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T19:38:20.902182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:38:20.902182Z digest=sha256:e18f044c5b456fb94f7d51d61ca42ae568b4b876bbc0f02e3d57534eb631566b

Observation f658452d-372c-446c-8f25-3536e859720a · inbound

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards cites this paper.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.811511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.811511Z digest=sha256:65e9d8f8efac3be15b2b33692eea9cb7e4c30f6b8a5b678b285261100f03863e

Observation 3eb694fb-4925-4ce6-9c72-fe6ffc122af1 · inbound

Rollout Pass-Rate Control: Steering Binary-Reward RL Toward Its Most Informative Regime cites this paper.

Rollout Pass-Rate Control: Steering Binary-Reward RL Toward Its Most Informative Regime Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:55:42.890276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T17:58:38.234197Z digest=sha256:835fb362146e291477a742f36767899e4f3cd87b2f0ff094681b6fc2bb92c7f5

Observation 5555d0e5-89ba-4846-9d58-27c7ba9e3fed · inbound

Rollout Pass-Rate Control: Steering Binary-Reward RL Toward Its Most Informative Regime cites this paper.

Rollout Pass-Rate Control: Steering Binary-Reward RL Toward Its Most Informative Regime Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:54.622226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T02:09:39.645781Z digest=sha256:0ca6e64c77c3069df97b6e46eda015dc64b34c18e493d8581770d1daf0834f58

Observation 29675988-4aa4-4541-a9a8-fc710c32bfec · inbound

Rollout Pass-Rate Control: Steering Binary-Reward RL Toward Its Most Informative Regime cites this paper.

Rollout Pass-Rate Control: Steering Binary-Reward RL Toward Its Most Informative Regime Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:32:41.462913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T17:32:34.585808Z digest=sha256:7c0c5fee3ee60a63e73c14f3bcc17e009176787671e8deea666eabb93a7de949

Observation 1741a719-b144-4594-bf07-22b160e3657b · inbound

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients cites this paper.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:56.164138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T01:08:52.981296Z digest=sha256:5f8dc6580e4bbc30f82ef5971b44980dd05c84417ab0cbec086d6ec458c96369

Observation 5c8df385-a953-4309-8779-b0c3f383755d · inbound

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information cites this paper.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:33.943411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:33.943411Z digest=sha256:6fd3a32458a869e03dbb31c9d3aa2c83c34a327fe0a5debc08b6c0d5137ecebc

Observation 26c06b32-36c7-43d3-ba09-8a252be4e03a · inbound

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning cites this paper.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:54.765756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:54.765756Z digest=sha256:a1b0a1c4f7c46d7f6e4360eeb7216764ecda476985842f6943cde07ed1a8623b