Pith. sign in

Paper Citation Record · LEDGER

Towards Self-Improvement of LLMs via MCTS: Leveraging Stepwise Knowledge with Curriculum Preference Learning

As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 2 inbound Pith citation observations for arXiv:2410.06508.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.06508 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 2 of 2 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T18:53:02.770333Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:56:29.951632Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation beefff04-48de-4cf9-b381-03c75111e970 · inbound

CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models cites this paper.

CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models Towards Self-Improvement of LLMs via MCTS: Leveraging Stepwise Knowledge with Curriculum Preference Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T18:53:02.770333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:53:02.770333Z digest=sha256:7e1e656412d5a6521938a28f0b3bd96669a02e7c8b98eb1978a2a9cc55b65468

Observation a2352f16-e9ff-4379-9aaf-4d06f4f24489 · inbound

Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling cites this paper.

Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling Towards Self-Improvement of LLMs via MCTS: Leveraging Stepwise Knowledge with Curriculum Preference Learning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:56:29.953706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T10:25:10.559953Z digest=sha256:c93f93c5cc37c61f1680d9bee0d30f02a3b53c59224129237874b1997cd73120