Pith. sign in

Paper Citation Record · LEDGER

Self-Improving Large Language Models via Progressive Experience Evolution

As of 8 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2608.02139.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.02139 v2

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:14:53.840638Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a86376f1-91ec-4e3f-a42a-ab7681d51475 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Self-Improving Large Language Models via Progressive Experience Evolution Training Verifiers to Solve Math Word Problems

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:51.628042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:51.628042Z digest=sha256:01b32836dd14da3c20abf8b29d98670244364f0e0e1c08efdb2fc941962a8736

Observation 8ae55e69-9a49-4504-b1b4-43e0f52c452d · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Self-Improving Large Language Models via Progressive Experience Evolution Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:52.132451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:52.132451Z digest=sha256:a4bfd87a9ec8b740839632885b2033bee6c04895ab876761c1394cdc010b9136

Observation 7f9e970c-930a-4876-a105-f50ebe8d9dfc · outbound

This paper cites Let's Verify Step by Step.

Self-Improving Large Language Models via Progressive Experience Evolution Let's Verify Step by Step

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:52.433047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:52.433047Z digest=sha256:9d8033ba672af98e16e03d5eaf5e1a69a0b65753b542bafe4af7670d83dcb22a

Observation fd8b029b-c614-420d-963a-c6901beb1782 · outbound

This paper cites ReFT: Reasoning with Reinforced Fine-Tuning.

Self-Improving Large Language Models via Progressive Experience Evolution ReFT: Reasoning with Reinforced Fine-Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:52.493528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:52.493528Z digest=sha256:115e76a29bbb62bae68ef62c96ec7294f571200976a31cc185cd6db63919ecfc

Observation 8ab77ad3-d340-48c2-8ca2-09ad79301ced · outbound

This paper cites Mexico City, Mexico: Association for Compu- tational Linguistics.

Self-Improving Large Language Models via Progressive Experience Evolution Mexico City, Mexico: Association for Compu- tational Linguistics

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:14:54.527323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:14:52.584832Z digest=sha256:04a8c282a53b290147b95d595006d5eccaad9a298bbb4a8f11e7746a28b15d90

Observation 17bae05d-6cd2-4496-a84d-ee2afd62df82 · outbound

This paper cites Privileged Information Distillation for Language Models.

Self-Improving Large Language Models via Progressive Experience Evolution Privileged Information Distillation for Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:52.667860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:52.667860Z digest=sha256:948b211ea90a39d31d0509959a07164f17ce393c7d8de05cfc660fca2674329e

Observation f931f582-dfb6-4887-89ad-1bd468f449cd · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Self-Improving Large Language Models via Progressive Experience Evolution DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:52.737140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:52.737140Z digest=sha256:3f080f61ad1d8d12645e143792f4a6bd62dfd99c9461364b8c873436166a6018

Observation bf57b854-9cf4-4a50-9db3-79684a2b7804 · outbound

This paper cites Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning.

Self-Improving Large Language Models via Progressive Experience Evolution Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:14:54.057384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:14:52.852963Z digest=sha256:040e49807f2f5746af592f52a26e8484ca0d6afdeb1a7cbaa941fecb40d9014d

Observation cb45d8de-3697-4783-9d55-96d7e867809c · outbound

This paper cites Learning to summarize from human feedback.

Self-Improving Large Language Models via Progressive Experience Evolution Learning to summarize from human feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:52.945625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:52.945625Z digest=sha256:65669276acfbf90967d372e0dd850f934ab93a37d95b6e3948c55741de9fcabb

Observation 6fc808de-7b5f-48e0-8ca0-15ed7ed9a7d0 · outbound

This paper cites an unresolved cited work.

Self-Improving Large Language Models via Progressive Experience Evolution Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:14:54.399597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:14:53.053392Z digest=sha256:07919fc7144fd2131340fd2f223fb37f784ea146c27347a0b0f7846155954c6d

Observation 7ee965ff-f280-4811-a330-c9a58e959cee · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Self-Improving Large Language Models via Progressive Experience Evolution Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:53.167075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:53.167075Z digest=sha256:70115783382e132aa4a07fe57b8db12f335924c448c23f51e71c68aa02c3cd96

Observation 917f942a-4e44-4f62-9862-d96d96fa09b2 · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

Self-Improving Large Language Models via Progressive Experience Evolution Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:53.270921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:53.270921Z digest=sha256:4415e6bc01e8ed007360d069e03dbbb566a4eaaed5a59f5ea4658656c4dfa516

Observation ae2fa355-1d36-42e6-b2c9-856a172610a1 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Self-Improving Large Language Models via Progressive Experience Evolution Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:53.385502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:53.385502Z digest=sha256:cb5608e7f50f12a737b375b7c2b18ecd86b23a312b3118a8b3705586aa297563

Observation 90a4fbce-1e03-4d14-8fca-69c32febfd6f · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

Self-Improving Large Language Models via Progressive Experience Evolution RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:53.476750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:53.476750Z digest=sha256:01c04a2f087cee273b96c99478a25ed9096122e1a98279b0413928a34b4c93d5

Observation 552e9929-b0b7-4a2f-a4b3-f196aa85df3c · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Self-Improving Large Language Models via Progressive Experience Evolution ReAct: Synergizing Reasoning and Acting in Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:53.569449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:53.569449Z digest=sha256:449c07c0a741f3d64fda31fccdae6c62011b59f4ab154fdaa07cb321f2a1dcae

Observation 69d5ac64-ecc7-4360-ad19-e41cc7ad756c · outbound

This paper cites Expel:Llmagentsareexperientiallearners.

Self-Improving Large Language Models via Progressive Experience Evolution Expel:Llmagentsareexperientiallearners

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:14:54.242192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:14:53.661199Z digest=sha256:d45cec373244543af6d4991e0aa44fa24ed616190e537eeb3f6c49af4ae1e9f4

Observation 6884dd0e-b56d-48ac-b1f8-b0a7828cbbea · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

Self-Improving Large Language Models via Progressive Experience Evolution Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:53.742184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:53.742184Z digest=sha256:540eb802c042cf735ddbe13aa9906ae48be8f5c5b1d9f94e574d9db8a4d2fd7e

Observation f922c2e8-bb6f-4c45-8aeb-fe438a28dee5 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Self-Improving Large Language Models via Progressive Experience Evolution Fine-Tuning Language Models from Human Preferences

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:53.840638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:53.840638Z digest=sha256:379726e5962b968595eda63c44c420a962db0992a3c728b0447ff7fda465d37a

Observation 3d7fd242-6134-4c8c-87ed-4cc0e74efd5a · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Self-Improving Large Language Models via Progressive Experience Evolution Distilling the Knowledge in a Neural Network

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:52.018406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:52.018406Z digest=sha256:7fc1218390c33eca2a806731850eb827be2a158c9c6741e72872dc917e10d0e6

Observation 936ce180-9f81-4839-9535-6323bcc33898 · outbound

This paper cites Scaling Laws for Neural Language Models.

Self-Improving Large Language Models via Progressive Experience Evolution Scaling Laws for Neural Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:52.247773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:52.247773Z digest=sha256:017f519c36be1b6db9e38a9d8f62b8517d4e2ad35249b3668b85132a74533e70

Observation 4058ba69-5813-41c8-83de-99ec992faf60 · outbound

This paper cites CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning.

Self-Improving Large Language Models via Progressive Experience Evolution CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:52.339631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:52.339631Z digest=sha256:8bddc5fe845db31df552217ce053af33aa3e4761b2473e8f25db2f98e75f2ca3

Observation 1261451e-4e58-4460-b123-6dab728803a0 · outbound

This paper cites GPT-4 Technical Report.

Self-Improving Large Language Models via Progressive Experience Evolution GPT-4 Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:51.435355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:51.435355Z digest=sha256:2b4e8a39f7549ff560b2d4e1d1fbbb282f5d57250a14ecc487a8c78658aae1bd

Observation 562d20f0-28e4-49f3-a75b-25b9f7dc0ce7 · outbound

This paper cites On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes.

Self-Improving Large Language Models via Progressive Experience Evolution On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:51.525171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:51.525171Z digest=sha256:a8e5242b89c5d8db272f0cb7d142ccc2d2fac2f3a647fa3e9a425dbe02d107be

Observation e3dc3a32-dfd4-493b-91d6-fa61925135e4 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Self-Improving Large Language Models via Progressive Experience Evolution DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:51.898510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:51.898510Z digest=sha256:24fb63a4a5a73dac3b08f1ea54502221edfe8241309b286642bd9c880f375999

Observation 4c0bdd01-606b-4211-8a2b-71f6b2bf37e1 · outbound

This paper cites MiniLLM: On-Policy Distillation of Large Language Models.

Self-Improving Large Language Models via Progressive Experience Evolution MiniLLM: On-Policy Distillation of Large Language Models

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:51.747842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:51.747842Z digest=sha256:5d6af8e94411c5581b3fdc82a0681dca76eabf451c2ab7475870d895cd80a998

Pith citing papers

No inbound Pith citation observations are available.