Pith. sign in

Paper Citation Record · LEDGER

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models

As of 7 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 1 inbound Pith citation observation for arXiv:2607.01480.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.01480 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-03T20:12:06.882343Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-09T03:36:57.168246Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T03:45:55.520373Z

Reference resolution

66 of 66 outbound references displayed

  • verified exact35
  • verified fuzzy17
  • unresolved8
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fc4d51fe-6fb8-4581-a029-b5b8d121d135 · outbound

This paper cites On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T02:30:40.993998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:0762267b15352e75ed5e8e5c54ae4b737b85ed9ebe60b3cc9106163f390505df

Observation fff6b2cd-6e04-4605-bd8e-9581ca7b4f5c · outbound

This paper cites X-KD: General Experiential Knowledge Distillation for Large Language Models.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models X-KD: General Experiential Knowledge Distillation for Large Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:56.042521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:4ee221da68305bf72ea5d3f50b50985728101e0ce79d3e42b61dd3c226bbcca5

Observation 1d4b53f4-ab9d-411b-a908-49fa86d3301d · outbound

This paper cites Mem$^2$Evolve: Towards Self-Evolving Agents via Co-Evolutionary Capability Expansion and Experience Distillation.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Mem$^2$Evolve: Towards Self-Evolving Agents via Co-Evolutionary Capability Expansion and Experience Distillation

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.051955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:09c4427ae623c726e7ceaf034538069aa1130e706a54794fe51225a8e2d94f4e

Observation 87811c71-02d1-4b6d-879d-d3d75b05d6ad · outbound

This paper cites 2024 , url =.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models 2024 , url =

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:18:56.063464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:46f8a44c9d9a6c2c44e5859f62c7566b5e08a916217c6afce9350d3af61457e4

Observation 7ba9b806-7857-402c-bed6-1918f5806784 · outbound

This paper cites MiniLLM: Knowledge Distillation of Large Language Models.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models MiniLLM: Knowledge Distillation of Large Language Models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T02:30:40.978661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:805b7bc283c3d1abd4391d9e64bfa659a70b709b94ddcfdcfcfe9ee20102cd33

Observation 0b1753ab-02ec-4f94-bc55-9673cd8230a0 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:55.998872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:3700e0bc23839505bef754645e688c54711072d6dad97474b4db573d27f1553c

Observation 5680b1aa-4d1a-47fb-b781-767415228ee5 · outbound

This paper cites Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.022571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:28e975c08e995f03ba96d597b12d58aadc1af5700a1b7525e4d7325822c58ca5

Observation 089bd3e1-5eb0-4757-b04d-dca08dd5b761 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Distilling the Knowledge in a Neural Network

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T02:30:41.003031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:3d4c4679ff3520c2c53b9f37224fdf993873b30a2607bbb788036b2d80ba5f44

Observation a11baa40-fae8-46b3-9765-17eafbc1ebf7 · outbound

This paper cites R-Zero: Self-Evolving Reasoning LLM from Zero Data.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.024959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:2f09620bed1df69f1ae674b531a11a9a53ff463b6d251304d5d57a40e17bafa9

Observation bce8b0a5-814f-4bde-8d2b-34066addb28e · outbound

This paper cites Reinforcement Learning via Self-Distillation.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Reinforcement Learning via Self-Distillation

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.034169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:22dd4e6e2a3216db114c06f1904b99599986014095129e360c18ae50042a1326

Observation 4a69fee8-75c3-474f-ab0d-7a521c93678f · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:55.992008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:439c4e043ce88722e887296a97293ccad9e66e48205a6f3077f4e3a223a0f222

Observation e6bbf87d-3bf1-408b-8cb8-8556994f040e · outbound

This paper cites Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.102957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:3cfaf764be13996bbca590122075f1aac5afd3317eb4a6a09f57bf320c3e3c5a

Observation d062eae2-1b50-4831-a1c4-1b7e6cd55aa2 · outbound

This paper cites uttler, Mike Lewis, Wen-tau Yih, Tim Rockt.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models uttler, Mike Lewis, Wen-tau Yih, Tim Rockt

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T02:20:38.106111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:9c5be6a83227aeb0fea56b3a39b0b2e0e92e73dc6ef2bbeae5ba0cd2ceee26f2

Observation 81047703-ce6d-4161-bb41-388488b2b128 · outbound

This paper cites Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:55.971108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:5c7074f418b24f18e84d08fd49fac9cf97b833feb2aaff84e557e0d6bd5f4fd4

Observation 670ff9ac-a164-4ffe-a885-5d34d06be270 · outbound

This paper cites Thinking Machines Lab: Con- nectionism.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Thinking Machines Lab: Con- nectionism

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T02:20:38.108167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:ce0b5edf4afdba04cfa3848431f43181b2d3658bc213c213b3e868a1c1899b92

Observation 9fdb8892-ecc3-47de-8e4b-f93fcc0840ba · outbound

This paper cites SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.006533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:4a72bab8507704e0cc437b413a5fbfefa4eff6971913e01ab20fb80cf6d3e1ea

Observation e7853188-04b8-42fc-9048-bc80633a380f · outbound

This paper cites Self-Refine: Iterative Refinement with Self-Feedback.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Self-Refine: Iterative Refinement with Self-Feedback

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T02:20:38.110479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:836c0043f893784aeb9a24c97b5cf3d23944e763279ef5ab41119536fa31ac25

Observation 6900ad8c-6747-4390-a81f-46f3692f5e2e · outbound

This paper cites Olmo 3.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Olmo 3

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.017511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:69854441f1f1155193f0b566dbd33a79d5ee9cd6ad554197bc8cb94aeee63b52

Observation 4817e270-b6ec-4d05-8370-dae361075871 · outbound

This paper cites ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.072054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:d4129456752e0a6c80142f8d5295e4fa552c1ccab0f6a4534ddeba95d841ad67

Observation 674b87a2-8d5f-4abb-89f4-9ba64fda4437 · outbound

This paper cites MemGPT: Towards LLMs as Operating Systems.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models MemGPT: Towards LLMs as Operating Systems

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.038402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:025b2218033f8407225018d807d2fb7fec5bfcffcd2ed4a463251cc4775e364d

Observation 62124e6c-a77b-4658-85d2-67fb4a3746c3 · outbound

This paper cites Generative Agents: Interactive Simulacra of Human Behavior.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Generative Agents: Interactive Simulacra of Human Behavior

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T02:20:38.067274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:a1623ecf390af01a486cde26eebea0ce1a6ed5334671b689252f1a91b4099287

Observation bf280084-6390-44c1-ba84-3d8efc1888a4 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T02:20:38.102087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:6927a21b81c7eeaf4409e326a3c76cad6ffe537640117927951fc58e56b1e413

Observation fb573098-16fa-4032-8706-3fb825bfb784 · outbound

This paper cites A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T02:20:38.103740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:3459f9849e6a3c8326c6c36247bca3c3eeda50f62c5750fa82c28b729ef1ad02

Observation 29e02a43-0d98-434e-a366-a303f6903300 · outbound

This paper cites CRISP: Compressed Reasoning via Iterative Self-Policy Distillation.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models CRISP: Compressed Reasoning via Iterative Self-Policy Distillation

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.076998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:33aa7de6d49d6843d73381ce18f2fdd4512629b94242674fc7cc8e549fabb362

Observation 40cde8c1-6444-4e98-ba9c-c25ddaa38b10 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Proximal Policy Optimization Algorithms

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.033961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:1af0e72686bd816b13486241862c93bee25d46460ccf1e2a50886b7b7bef517c

Observation e90a683d-a983-48bc-acc3-1cf3784e9dd7 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.002504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:8fa057653662447639a70e08b1921daa50f47b9ac8c89f8f3003c09a3e78f1da

Observation bf1e3ccc-7beb-4018-89e8-53dc9d500b73 · outbound

This paper cites Self-Distillation Enables Continual Learning.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Self-Distillation Enables Continual Learning

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.026859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:ae604a64552a21b12d01197c8df10c62a6ccb56c2dcb207ff45bd98d9a207b7e

Observation 56208162-0e12-435e-9886-4b3fc49b2bce · outbound

This paper cites arXiv preprint arXiv:2602.13949 , year=.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models arXiv preprint arXiv:2602.13949 , year=

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:18:56.031192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:1eab4059a92cc6ea507bcfd551dc09192cdd1534ef168b2c16630b44b40875a8

Observation a096af70-762f-4649-a1e8-ae31dee55dc7 · outbound

This paper cites Reflexion: Language Agents with Verbal Reinforcement Learning.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Reflexion: Language Agents with Verbal Reinforcement Learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T02:20:38.112836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:9bf4ef564d34a9a46e0d02488f0901f3f4ee6f5e9f7ba084010e4e6bd8a91997

Observation 3fece100-cff2-4eaf-91bf-0e0843dace55 · outbound

This paper cites A Survey of On-Policy Distillation for Large Language Models.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models A Survey of On-Policy Distillation for Large Language Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.097509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:4a9e524274bd71ae81a212c57876ef030fa06e961c56830f9a1644fa05980db1

Observation db5a2bff-2fdb-409b-b8b4-4d2825affed9 · outbound

This paper cites arXiv preprint arXiv:2512.16962 , year=.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models arXiv preprint arXiv:2512.16962 , year=

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:18:56.047583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:dcd5de47cc7f741a4fb77a0b1b35766eaf56564eacb31074b6bb18a8cb127cf6

Observation 89180346-94e3-4128-8574-9f93e2463a5f · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.052252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:59e3f0ea29b3dd189d0534e050adbe3ef93ee6b996a5cff111284eabbce30a68

Observation 92cfd3ce-54e4-479f-9536-bc24da8582a6 · outbound

This paper cites Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:55.981242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:f64e1d902a5cedc9647be7e7c8d17898b03148a6d22e5395181b1531b3f3981d

Observation 1da99ab1-a6c3-4100-83d2-702a6600178c · outbound

This paper cites Skillorchestra: Learning to route agents via skill transfer.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Skillorchestra: Learning to route agents via skill transfer

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:55.975149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:bd142aaf9c5dcc5f6579c556ba10723e0e7d672226c4166058cdb423cfd52bb2

Observation 4d9917ee-8029-4bc9-b57f-4cead8e44db8 · outbound

This paper cites Mem-{\alpha}: Learning Memory Construction via Reinforcement Learning.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Mem-{\alpha}: Learning Memory Construction via Reinforcement Learning

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.027410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:8ae866a8e5dd26e61b1b95be5ec420b37cc224e1b53238743a70ab244b2316b8

Observation e9b1c90b-af14-4dc4-8d59-7e9a9cd62ed7 · outbound

This paper cites Q-learning.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Q-learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T02:20:38.114987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:c4fbcf66ccd59751b072aaf4024bf25958ac7f7f998ec0653a32a7cc55bc7e2e

Observation 8aae3176-e03b-4f4a-896c-46b15aa18b42 · outbound

This paper cites EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:55.991514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:a005db47a86c7b6731fc77c6fc49466214bff41a788642f10bc4de033690345a

Observation 1faae88e-5ffe-406e-95e5-451a4bf81e3f · outbound

This paper cites TokMem: One-token procedural memory for large language models.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models TokMem: One-token procedural memory for large language models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:56.011962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:0e55086bea37c7d8eaf854e7f15dba983b21033aedc5b53beee56e5cf8e85cac

Observation 723905de-d7a4-44d6-a1f5-e7a4a56d650f · outbound

This paper cites SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:55.962509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:0f67f3d40354c5eb6dd0c2e9b3a1f76f6356dd04f12e587c90928052dad11466

Observation 81af51f9-74f0-4e9d-8337-3e4c2448dda2 · outbound

This paper cites Smith, and Hannaneh Hajishirzi.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Smith, and Hannaneh Hajishirzi

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:18:56.092025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:56bed2d20a42ca4232918fcf3382882cd02bd74e3b34a25b6879c55126e5dd70

Observation 58d2b13a-a53e-4ab4-a6c1-b82e0b9f3eb0 · outbound

This paper cites A-mem: Agentic memory for llm agents.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models A-mem: Agentic memory for llm agents

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T02:30:40.992228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:edaab55e6ec899337366071376f90d705b2061e438d6685197eb950cfbce0e3b

Observation c1825435-f391-414d-8378-e3eb128426be · outbound

This paper cites Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.110948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:c19803b55285ba62b5dfd24acada2b58dac14e41b3beb4dd9679b3182b9e7530

Observation 7e6c3d50-1c1b-4e75-bfc6-4b4ab203f981 · outbound

This paper cites Qwen3 Technical Report.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Qwen3 Technical Report

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.055985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:2c3bbad6ef329ba31d96afc4dae6cf7b59af6e16751ff77fe7164e601cc3f27f

Observation 55154ca4-4ad4-4327-9cd8-dce7ce0a44ed · outbound

This paper cites Self-Distilled RLVR.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Self-Distilled RLVR

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.016907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:9b44a4a488c2776dabdc0ec7355fd40435429f100473f11ff6109cd1f8f3ab8a

Observation a71b7b0a-151c-4b64-953d-91a58e966412 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models ReAct: Synergizing Reasoning and Acting in Language Models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T02:30:40.983835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:4ee1091a3f11e836345f9ffc27c5b37ed25113432985f70321d83de3ec20e6d5

Observation 94f8087b-693c-4ce0-9cd8-3c3a9105ae97 · outbound

This paper cites Online Experiential Learning for Language Models.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Online Experiential Learning for Language Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.047601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:db58c4ff477b4be8f74bc3dfe9a09b960467f7061e9893ff4077f5c2eb67fe79

Observation 18e3537b-5c2c-40ae-80ac-bd6ca651053a · outbound

This paper cites On-Policy Context Distillation for Language Models.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models On-Policy Context Distillation for Language Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.020870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:50e99d65c7c8b83311f30f38991c87c2271dd361361a2efd7edc43425d83f3b4

Observation d5bb9787-45e6-4721-88e4-957545d6e9a7 · outbound

This paper cites Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model Agents.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model Agents

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:55.999248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:402ce1be8837b9dfc7e164d87d396a345b2b4c4a0665c42b1dd61315b080a95a

Observation 83c2c239-5d14-401b-a70c-95c1ad02da25 · outbound

This paper cites Appagent: Multimodal agents as smartphone users.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Appagent: Multimodal agents as smartphone users

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:18:55.951347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:461e980d389dd771eefc37b0295a457004d4a000d2207408f69bb08cc684309f

Observation f7b69478-4de6-403f-9bbc-fbe3c568c477 · outbound

This paper cites Embarrassingly Simple Self-Distillation Improves Code Generation.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Embarrassingly Simple Self-Distillation Improves Code Generation

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:55.984794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:34276c33a31bd87b478ecfe611b764b78b418712442da0bab8df4abfdd379389

Observation b14c456a-aeb6-4ff6-8935-18dc82d63618 · outbound

This paper cites Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.067076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:30da90788a6c4ef8829a20dd670073f6cf4a2155b718fa3a16f402cff5ce8bf5

Observation 06e1150e-1063-41cf-8c7b-612a00060226 · outbound

This paper cites MemFly: On-the-Fly Memory Optimization via Information Bottleneck.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models MemFly: On-the-Fly Memory Optimization via Information Bottleneck

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:07:05.092025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:9e0e5e8a9b4800d9b10519f363ca71a22fa34a6f263039ba11b96eb5da435162

Observation d2d34747-665b-488b-a40c-6b988b835b79 · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.019411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:82a4170e259b1e4cb429cf8c0e7e335d6b3aeccd044bb2163dec7583b01e2e42

Observation bfdcb1ff-8825-482b-8d19-b400e993c334 · outbound

This paper cites Memorybank: Enhancing large language models with long-term memory.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Memorybank: Enhancing large language models with long-term memory

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T02:30:40.976558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:4ec809458ab4367023ed6f2c559d4ea790c3e2184c23790e58271de6e2c7ff48

Observation acaae7d4-6702-4255-a879-3460198171d9 · outbound

This paper cites Memento: Fine-tuning LLM Agents without Fine-tuning LLMs.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Memento: Fine-tuning LLM Agents without Fine-tuning LLMs

Reference 55

Resolution
malformed identifier
arxiv_id, observed 2026-07-03T20:18:56.014458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:8136fb6cab72bf5e510418fa566d6dcc42ecc36eef8cbb0a90bb27d48f9f067f

Observation d0a3161e-ba21-41c9-ba8d-7fa83ab0c402 · outbound

This paper cites an unresolved cited work.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-07-05T02:30:40.980383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:cde294ff38d12618bf273bc3f63c69f99468f75f6bf13a6f812c0674a7454bb9

Observation 146ed652-00b0-4022-ac36-1c578f473177 · outbound

This paper cites strategies.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models strategies

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T02:30:40.985478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:30fb21840e09427240fef141b9bce7a3b568f8393677c367003868445eaea7c4

Observation f8500994-61ce-43ff-b86d-2d3dcf817c82 · outbound

This paper cites an unresolved cited work.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-07-05T02:30:41.001001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:0792740c4854b21858dde57408576960823567bafe1323838165a10b66c0c3eb

Observation 44e5f288-da0c-4c10-ac1c-c2149c6b69f5 · outbound

This paper cites an unresolved cited work.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-07-05T02:30:41.006527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:716ee23caf373ee1fbfdef78445af6cf99d7f0e3928e24a3309ad31087bfa360

Observation c568cb7f-2553-475c-b470-619019826c99 · outbound

This paper cites an unresolved cited work.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-07-05T02:30:40.995573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:7e8c7dada8006df6469fe0e5554dbdfa08f702b9f0d80dc92fb24de3c1f32390

Observation 73a5f3c2-fca4-4cbe-8657-32b71e6e757a · outbound

This paper cites an unresolved cited work.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-07-05T02:30:40.998728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:c4040fc08123d45ef5985292a44a8f684539d4325783b99e5b0941cc4da29071

Observation 7ecde5b2-94bd-4ad0-b551-3ea4dae42b3e · outbound

This paper cites behaviors.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models behaviors

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T02:30:40.982118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:57ef4382045a744921d4492d361bde1d269faa591be2d301687fd7c6c6e1270e

Observation d981465f-f665-460c-bfa3-1b007ba4bb12 · outbound

This paper cites an unresolved cited work.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-07-05T02:30:40.987268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:a078bc053e4c23173e9f1cc9ee9b14a037d7a7d6e25f9370cd661b444666816d

Observation 4c11d6a0-6440-4bcb-ad9c-f14fccb86cc7 · outbound

This paper cites an unresolved cited work.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-07-05T02:30:40.990581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:0203f44cc4f33aa6c53b33b51fc0c9896dde1f1db9cb1859d2283e1141afdca0

Observation 566235dd-86cc-4878-8e32-c40995a704a1 · outbound

This paper cites an unresolved cited work.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-07-05T02:30:41.008095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:1aa9883ee2f04ace9b3a51ea8d6d3fb2717e051b3ea5660e5ae97a90e6e5f0a1

Observation 22875d9b-6d43-4d26-8fc9-36cb067fce14 · outbound

This paper cites actions": [ {.

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models actions": [ {

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T02:30:41.004833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:12:06.882343Z digest=sha256:2dde538346916dcb391d524527b87bd8bff02dce866b2a5254470d448b5c492a

Pith citing papers

Observation 08602f70-20ab-430b-9ad5-903980551299 · inbound

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops cites this paper.

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops Procedural Memory Distillation: Online Reflection for Self-Improving Language Models

Reference 113

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:45:55.521846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T03:36:57.168246Z digest=sha256:08dd2d262a5bf8e133031f306adfcf7fba9ffcdce4f7b9107f2a7e06d3027da9