Pith. sign in

Paper Citation Record · LEDGER

Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2412.14135.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.14135 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T23:33:52.937067Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 647261e7-faf9-46f4-8233-cd650ecf5e9e · inbound

HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs cites this paper.

HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:36:50.283750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T12:36:50.060335Z digest=sha256:32173c8df04d0c6ff63d6c715da848252a1e00bc5da17652dba9a30022b8b494

Observation ff67168c-269a-401b-ae22-67cffee03c28 · inbound

Search-o1: Agentic Search-Enhanced Large Reasoning Models cites this paper.

Search-o1: Agentic Search-Enhanced Large Reasoning Models Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:36:27.640538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T17:36:27.515468Z digest=sha256:bdd713721cf12a0036201616e07ef2afe609535ba8313ea96de089a3cf764ad4

Observation 454dedb2-e3f9-4e64-9a26-2b809b0efec5 · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:36:24.628273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:aec8b05a03b3ddf5517d5e50e691b9edff12921cb92b4e7df34a429bd2edc9b9

Observation e9b1f88b-5d58-4037-921a-64db6211d813 · inbound

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks cites this paper.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:52:14.111450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:759786233daf15fbce607368b727563c1dfbcbcb6bd8bb2d2cd5bfb1f24a9d6b

Observation 3372b021-128e-463a-9bf3-b4508077d489 · inbound

TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework cites this paper.

TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:52.937067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:33:52.937067Z digest=sha256:c402247c8b49abfc4a5be14568a66396ea12ff640e8b7c986c1101299693c532

Observation a1b73dc4-bc02-4a79-bf8c-24778fe13642 · inbound

Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMs cites this paper.

Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMs Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective

Reference 266

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:45:59.486005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T16:58:10.013475Z digest=sha256:cd4fbfebfb6610820f45605943baea373d7a12ab312249a05675ab765823b8bd

Observation 9ffcf2c6-bd94-470f-b32f-550aa083dcb3 · inbound

Adapt to Thrive! Adaptive Power-Mean Policy Optimization for Improved LLM Reasoning cites this paper.

Adapt to Thrive! Adaptive Power-Mean Policy Optimization for Improved LLM Reasoning Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective

Reference 251

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:01:00.754150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T16:51:19.555272Z digest=sha256:4ec50405f138432e7df0026d73c07c4f5d7d3d640347479ce7ba0ac50d4c0c1c

Observation 58f25752-65fa-4ac8-bbaa-43dd3fcbd1fa · inbound

CAT: Confidence-Adaptive Thinking for Efficient Reasoning of Large Reasoning Models cites this paper.

CAT: Confidence-Adaptive Thinking for Efficient Reasoning of Large Reasoning Models Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:26:58.101193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-02T13:22:27.566432Z digest=sha256:5907692768a38a94106b31620e3bfbc6252f75d539783becbcd6181cf03669c2