Pith. sign in

Paper Citation Record · LEDGER

Thinking-Based Non-Thinking: Solving the Reward Hacking Problem in Training Hybrid Reasoning Models via Reinforcement Learning

As of 9 August 2026, this Paper Citation Record lists 6 of 6 outbound references and 0 inbound Pith citation observations for arXiv:2601.04805.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.04805 v2

Coverage vector

measured 6 of 6 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T11:57:56.105261Z

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

6 of 6 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c3668722-a034-4959-8801-38af745e7961 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Thinking-Based Non-Thinking: Solving the Reward Hacking Problem in Training Hybrid Reasoning Models via Reinforcement Learning Understanding R1-Zero-Like Training: A Critical Perspective

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T11:57:55.717976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:57:55.717976Z digest=sha256:865fbf62330bf63606a9a6f6f5cb18aa8d89044f17d30ccaa3437419ee3879f0

Observation ca149062-f3c4-4f18-b192-c05c9a81c0c0 · outbound

This paper cites Feedback Loops With Language Models Drive In-Context Reward Hacking.

Thinking-Based Non-Thinking: Solving the Reward Hacking Problem in Training Hybrid Reasoning Models via Reinforcement Learning Feedback Loops With Language Models Drive In-Context Reward Hacking

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T11:57:55.899447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:57:55.899447Z digest=sha256:73b29a9cda2504c86f1cc1560b032761f41cb7dd01ac90daa74e4ef3f2d98d7e

Observation f52f6b69-0b3a-4462-a845-512b93b01a08 · outbound

This paper cites Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models.

Thinking-Based Non-Thinking: Solving the Reward Hacking Problem in Training Hybrid Reasoning Models via Reinforcement Learning Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T11:57:56.105261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:57:56.105261Z digest=sha256:40812bc8e6c8c51ce303d9adb868eae3face3c4557fc24aa10a0a509065fb3ad

Observation 92821b1a-0390-4fcf-8b03-debe79311163 · outbound

This paper cites Let's Verify Step by Step.

Thinking-Based Non-Thinking: Solving the Reward Hacking Problem in Training Hybrid Reasoning Models via Reinforcement Learning Let's Verify Step by Step

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T11:57:55.505392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:57:55.505392Z digest=sha256:0d209befa18f21a36f92d0f4377acf2da0a4bb2777f0a93216aa6d4930191f68

Observation ac5ae07d-7a07-4554-bcb2-15a2561fdcd8 · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

Thinking-Based Non-Thinking: Solving the Reward Hacking Problem in Training Hybrid Reasoning Models via Reinforcement Learning Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T11:57:55.221525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:57:55.221525Z digest=sha256:e29b019fba413c7bb87ebf008350faef4712961bf64bb8325b47deadb18c5e4d

Observation 395467f8-3c9a-47b1-adba-a8327d70034c · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

Thinking-Based Non-Thinking: Solving the Reward Hacking Problem in Training Hybrid Reasoning Models via Reinforcement Learning ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T11:57:55.346251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:57:55.346251Z digest=sha256:b49358d51eee129f794dbe1d57d9cb14dce595190e9261b12a04f9ad58fa8d72

Pith citing papers

No inbound Pith citation observations are available.