Pith. sign in

Paper Citation Record · LEDGER

ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2312.10003.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.10003 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:18:48.417057Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ebcb3ca3-be8a-4d86-805c-d3123f16e5e8 · inbound

METEOR: Evolutionary Journey of Large Language Models from Guidance to Self-Growth cites this paper.

METEOR: Evolutionary Journey of Large Language Models from Guidance to Self-Growth ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T18:22:41.319035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:22:41.319035Z digest=sha256:3e79193e722737bdc12d92638babd935ea16428d09120e02bf2de075e4f0f50a

Observation a140117e-8ccd-4d50-820a-96af5bac56c3 · inbound

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision cites this paper.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.225484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.225484Z digest=sha256:b6835cdea01771249bd925cd6785655229775da7dab0e5674d564aacd1f75ac1

Observation 77782fd0-038e-4755-90dc-7407c55db867 · inbound

Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments cites this paper.

Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T18:56:44.399417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:56:44.399417Z digest=sha256:d721c83172d6761279f35d11b6fbdcd22cb4668c6cdd16c8bf1e7e92a199964f

Observation 98cd6fc3-0cdb-4b1e-9952-546dc8930ff9 · inbound

Exploring Expert Failures Improves LLM Agent Tuning cites this paper.

Exploring Expert Failures Improves LLM Agent Tuning ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T12:18:48.417057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:18:48.417057Z digest=sha256:0c9a36c7b6a1294f0d670c2dfd868acff559ad5f59fb0e35b438673321937577

Observation 0d5d98b6-93e2-4658-8d54-4003accc3b97 · inbound

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review cites this paper.

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:57:38.252446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T02:57:37.873567Z digest=sha256:6177cd2daae07af065c07087b5948e5b66cd625519f1c35f892fe3714147daba

Observation 4e2a6e8b-cb11-462c-b840-ba4da87be4b4 · inbound

WebEvolver: Enhancing Web Agent Self-Improvement with Coevolving World Model cites this paper.

WebEvolver: Enhancing Web Agent Self-Improvement with Coevolving World Model ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-16T11:09:37.258867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:09:37.258867Z digest=sha256:c2167ea4ddebf285137d038d2b759f94cdc286f6ddb714ee0ea3d54a93294599

Observation e146ce08-8442-47ea-baba-77c1d7e78c41 · inbound

Truly Self-Improving Agents Require Intrinsic Metacognitive Learning cites this paper.

Truly Self-Improving Agents Require Intrinsic Metacognitive Learning ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:16.864213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:28:16.864213Z digest=sha256:b4b253a1fb4c7f9ba2a67d579fcadb696bc5a082ee30fcb64755d6c94edb3ef0

Observation 40e8a6d1-dc91-4263-8cc7-410eb534c813 · inbound

UProp: Investigating the Uncertainty Propagation of LLMs in Multi-Step Agentic Decision-Making cites this paper.

UProp: Investigating the Uncertainty Propagation of LLMs in Multi-Step Agentic Decision-Making ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T19:17:00.501720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:17:00.501720Z digest=sha256:9d1e979c7f33d40f25a79e16a78283928745a272442fd56cb97725ba7ba31b31

Observation 7bf92f9b-25fa-45a2-949f-e52db5e4775e · inbound

AI Reasoning for Wireless Communications and Networking: A Survey and Perspectives cites this paper.

AI Reasoning for Wireless Communications and Networking: A Survey and Perspectives ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T19:34:28.337466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:34:28.337466Z digest=sha256:868b1217104025570f64591b27ea39af8535069b6023c195381eb373c76412b2

Observation af7de01d-ebad-4268-b54c-87434c6f3bcd · inbound

Estimating the Empowerment of Language Model Agents cites this paper.

Estimating the Empowerment of Language Model Agents ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T14:53:14.599607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:53:14.599607Z digest=sha256:ba624fa7704c51ac80e08f8e78e38d63443193b89e029f113a68aef9654456d3

Observation 70ba945b-1d86-4212-81f1-419d5feffc55 · inbound

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory cites this paper.

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 234

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:13:15.679804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-14T23:13:15.016486Z digest=sha256:0a4757b696dc49f1aad039dfb16931818c0ba75f588a6f8f5a14ad6974ceee93

Observation cb05160e-8cee-452a-8503-e66f5b5c1c96 · inbound

Agentic Reasoning for Large Language Models cites this paper.

Agentic Reasoning for Large Language Models ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:14:25.988543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T15:14:25.558878Z digest=sha256:422f7ebcd81f83f3ba8c3e56f9ae76a391cb42dfb86f546ee5586da9946844f3

Observation c3be1648-17fc-4441-ab3d-2bf62885c8f5 · inbound

Toward Efficient Agents: Memory, Tool learning, and Planning cites this paper.

Toward Efficient Agents: Memory, Tool learning, and Planning ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T09:21:32.684964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:21:32.684964Z digest=sha256:f4bbdcbed8aa0f1e3c4cfadac45f1ab06c7beab1a3004fe78e797706f06a5616

Observation 037ae162-fc22-492f-bde3-b91cfe59372d · inbound

Quantifying Trust: Financial Risk Management for Trustworthy AI Agents cites this paper.

Quantifying Trust: Financial Risk Management for Trustworthy AI Agents ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:18:01.300645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T17:16:17.464937Z digest=sha256:e3c04f7bee702e87d720adb4d50d679a90f293613a051f3ba1eb11618b7c9992

Observation 7258b9e7-2cdf-4f52-9335-e577337d4d82 · inbound

DuIVRS-2: An LLM-based Interactive Voice Response System for Large-scale POI Attribute Acquisition cites this paper.

DuIVRS-2: An LLM-based Interactive Voice Response System for Large-scale POI Attribute Acquisition ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:33:12.922872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-20T10:30:48.957568Z digest=sha256:1bc04a8641a7bb4431cb687b715f24a2661d734fca13645bb271e48d9b5c4697

Observation 3352d076-4318-4d6b-b64a-b506e727b548 · inbound

EvoGens: A Population-Based Heuristic Search Framework for Scientific Idea Generation cites this paper.

EvoGens: A Population-Based Heuristic Search Framework for Scientific Idea Generation ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:32:44.146779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T22:29:44.078911Z digest=sha256:be2c5ef2191a875ab3eb7278a5def13adf0ee834d3c3762a0f4e05a7408418ef

Observation 4c09d938-fcb1-48b2-865a-423e2a18bd28 · inbound

SlimSearcher: Training Efficiency-Aware Web Agents via Adaptive Reward Gating cites this paper.

SlimSearcher: Training Efficiency-Aware Web Agents via Adaptive Reward Gating ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:17:08.710917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-27T22:54:44.329613Z digest=sha256:eef09057260a8cd5c7ef1dec038a60e4530384500479c22d829ea5302af044f8

Observation d379452a-44fc-4c5a-bf92-c4ce6d53b990 · inbound

Graph2Idea:Retrieval-Augmented Scientific Idea Generation with Graph-Structured Contexts cites this paper.

Graph2Idea:Retrieval-Augmented Scientific Idea Generation with Graph-Structured Contexts ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:17:31.037482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T16:39:36.945808Z digest=sha256:ec520eabef1f5b9a867803782aed214f317261f72d42c27d11ba601304b6eada

Observation 079675ac-308f-40a4-a1e8-8918767a667b · inbound

Self-Improvement Can Self-Regress: The Rise-and-Collapse Failure Mode of LLM Self-Training cites this paper.

Self-Improvement Can Self-Regress: The Rise-and-Collapse Failure Mode of LLM Self-Training ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:29:16.726094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T21:07:28.660119Z digest=sha256:d2ae543943bc775dca56773bef195921b68ab4dd4fe55ae431ad4aa012de2752

Observation b5f02c49-29fc-4e89-8f0d-210781c0d1e7 · inbound

Stop Hand-Holding Your Coding Agent: Engineering the Loops that Replace Step-by-Step Prompting cites this paper.

Stop Hand-Holding Your Coding Agent: Engineering the Loops that Replace Step-by-Step Prompting ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:47:22.277629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-02T20:40:11.059493Z digest=sha256:118d93660d13d9ddf24e3291a798c5642b847f79128623365864cd78963a9194

Observation 8ceacf65-bebc-4793-b2c4-7fc4ebf54b23 · inbound

LLMs for Agentic Home Energy Management cites this paper.

LLMs for Agentic Home Energy Management ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-11T17:08:27.234032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:08:27.234032Z digest=sha256:b1e17ff247723078338566e29dc145865bf0badf3f3391f7ed1b009dff317372

Observation ee669dc4-56fa-4ab4-9f3a-51b5978ba771 · inbound

LLMs for Agentic Home Energy Management cites this paper.

LLMs for Agentic Home Energy Management ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:33.501969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:40:33.501969Z digest=sha256:733c46136f317d081b5372275ce302a1bbd6457fae3d1741d4d33d585b2c83b0

Observation 7ffba785-912c-490c-9786-776a4b314105 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 78

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:499b2e56d1c101538ffa63ebedb7f941aa14e012f92d3375d1a3657c62916330

Observation 6f322826-12d6-45ff-ba02-ecde46fa59be · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:40.326974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:40.326974Z digest=sha256:43a14ace78bd7fb257193b60e84219804abe395acec145b9a536fc9114b1271d

Observation a194edc7-af73-4ddd-8eb2-ca258a572d92 · inbound

Engineering Trustworthy Agentic AI for Critical Systems cites this paper.

Engineering Trustworthy Agentic AI for Critical Systems ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:28.624570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:28.624570Z digest=sha256:a38ca6d01708d2fa1e0bb7ac55f8a45a54772bb000cd20983fa706ba5d3db3bb

Observation 63ef0a9b-0eb2-4062-acbb-75db8b113e97 · inbound

Hybrid LLM-Augmented Reinforcement Learning Agents for Complex Sequential Decision Tasks cites this paper.

Hybrid LLM-Augmented Reinforcement Learning Agents for Complex Sequential Decision Tasks ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T17:52:44.777206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:52:44.777206Z digest=sha256:2d99746e3891d7f5560839f987f7c7a2c3081e7cd68fd7930e90ad7a72ba5f40