Pith. sign in

Paper Citation Record · LEDGER

Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents

As of 5 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 1 inbound Pith citation observation for arXiv:2605.30159.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.30159 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T07:31:00.010342Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T10:07:16.700499Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-02T03:26:28.755836Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact17
  • verified fuzzy0
  • unresolved2
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b585e253-a62f-420b-9a00-216649f204a1 · outbound

This paper cites Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory.

Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-29T07:33:13.440771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:31:00.010342Z digest=sha256:9e88f4e3179ba0ddabd8b06a3ed7dda4589eaa54e86a5e177ec4add97ee87590

Observation bebaa282-e6ff-49e3-91c6-b9c8ddb06350 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents Process Reinforcement through Implicit Rewards

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-29T07:33:13.443443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:31:00.010342Z digest=sha256:f16a8cfaf5c602f412f67866259a4f619b10a9ce017b7cbd2b15fe936339410b

Observation 8452a397-d936-41a6-b52c-1f1547935804 · outbound

This paper cites DeepSeek-V3 Technical Report.

Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents DeepSeek-V3 Technical Report

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T07:33:13.451103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:31:00.010342Z digest=sha256:3c2340f8da5eac8e8b5523d80f5e9fb30c52a4d57231ec79639296c1844f1b49

Observation 04094eff-3cb8-4982-b4b2-1901071e3489 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-29T07:33:13.427264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:31:00.010342Z digest=sha256:fe585b458d6ac10bb192480e46fac01f8c5abeec177501a29c3a84766b95b7f5

Observation 359d6454-51cb-451a-89d6-ad8e082fec26 · outbound

This paper cites Memory for autonomous llm agents: Mechanisms, evaluation, and emerging frontiers.

Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents Memory for autonomous llm agents: Mechanisms, evaluation, and emerging frontiers

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:33:13.446261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:31:00.010342Z digest=sha256:3e017c0a0cf8bbc210200254f6ab4de4a0d14d0be0ac7c37c7d2d5fed43f9d22

Observation 98e01cf1-c8ee-41b9-944c-fc76ca98086e · outbound

This paper cites The Llama 3 Herd of Models.

Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents The Llama 3 Herd of Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-06-29T07:33:13.418528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:31:00.010342Z digest=sha256:6f19fe6218b65ecf5fe74f79e3104e9ea1e6f4b919c152ce0ebf2e086c140c2d

Observation 4df8dfc9-3156-404a-90ff-4d5ce39d6888 · outbound

This paper cites Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective.

Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-29T07:33:13.417177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:31:00.010342Z digest=sha256:d0ab165a0a4626f8f2db7ce7c8313c8ab3004830113172c39ebf4142a30d6b9a

Observation 4888566a-4858-4236-a375-faf78d8a3704 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-06-29T07:33:13.419713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:31:00.010342Z digest=sha256:6dad6b3b17891f66183537e08d8de5b89f1a94b22b6dcde1cfa8744c9ca272ba

Observation 78574955-d4b7-4858-8373-56432812f4fe · outbound

This paper cites Language Models (Mostly) Know What They Know.

Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents Language Models (Mostly) Know What They Know

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-06-29T07:33:13.421934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:31:00.010342Z digest=sha256:792eb99a62f94fc4240eea3777b5ba6b7d961d58d271865b8403116d49906748

Observation 632117b4-0a02-4cb0-bda5-82a73380ecf6 · outbound

This paper cites MemOS: An Operating System for Memory-Augmented Generation (MAG) in Large Language Models.

Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents MemOS: An Operating System for Memory-Augmented Generation (MAG) in Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:33:13.424789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:31:00.010342Z digest=sha256:3b7b0941ade0f0733c2e6db75d2553a3b9feb0c01ee487a36385970deaced56f

Observation 35e0d7ff-494d-4f4f-9086-a51e9ba35b79 · outbound

This paper cites CoRR , volume =.

Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents CoRR , volume =

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T07:33:13.449553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:31:00.010342Z digest=sha256:e2e9be67e3550103a477cbcf7265b71c447d0a6cadfa108e874f646bf7b7a218

Observation 651aa359-ad3a-482a-80b2-e287fb96455c · outbound

This paper cites MemGPT: Towards LLMs as Operating Systems.

Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents MemGPT: Towards LLMs as Operating Systems

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T07:33:13.413069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:31:00.010342Z digest=sha256:f7c800b7d4a8e6699afca069d9bbb4c3ed2c1fe1450245f8a2be9e313dd00f1b

Observation 94f834d4-9b01-44a7-8798-b16d8535b2f7 · outbound

This paper cites Sufficientstatisticsintheoptimumcontrolofstochasticsystems.

Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents Sufficientstatisticsintheoptimumcontrolofstochasticsystems

Reference 13

Resolution
verified exact
doi, observed 2026-06-29T07:33:12.803942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:31:00.010342Z digest=sha256:e60e1c0494d37b3a8ed332907f85f623a41d1f1d55cd583a59dc2edc4eda5213

Observation a017bad6-bd36-4321-8257-45c8144e4699 · outbound

This paper cites When to mem- orize and when to stop: Gated recurrent memory for long-context reasoning.arXiv preprint arXiv:2602.10560.

Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents When to mem- orize and when to stop: Gated recurrent memory for long-context reasoning.arXiv preprint arXiv:2602.10560

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:33:13.410692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:31:00.010342Z digest=sha256:34741d7bd4028fba94877c61dc5004376e8bf83ae11d979a51943a432ab572c4

Observation 95236416-bdb3-4e50-8797-557039790449 · outbound

This paper cites An Exact Solution Approach for Portfolio Op- timization Problems Under Stochastic and Integer Constraints.

Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents An Exact Solution Approach for Portfolio Op- timization Problems Under Stochastic and Integer Constraints

Reference 15

Resolution
malformed identifier
doi_truncated, observed 2026-06-29T07:33:12.804635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:31:00.010342Z digest=sha256:3f5692787c267725eda16e4e8f603ed60e15722e0fa59f49a628ee9544d8cc75

Observation aa27fdd6-a2aa-4ece-8a39-4bc4fe442977 · outbound

This paper cites Mem-{\alpha}: Learning Memory Construction via Reinforcement Learning.

Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents Mem-{\alpha}: Learning Memory Construction via Reinforcement Learning

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T07:33:13.448579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:31:00.010342Z digest=sha256:4d5eea04fb5b3356b6c7d36e85f3e92beed843068349e3babcce6d0e96b4445d

Observation e17ebc96-e040-4f10-a35e-e316a78efe2e · outbound

This paper cites A-MEM: Agentic Memory for LLM Agents.

Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents A-MEM: Agentic Memory for LLM Agents

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-06-29T07:33:13.453403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:31:00.010342Z digest=sha256:648c2d32ec985e9de2c55681a12b2d55ea09337c7c515d179a5a27f4175accc9

Observation a23a96b3-fd48-4893-9fcc-49dd1682cf3e · outbound

This paper cites MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent.

Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-29T07:33:13.455884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:31:00.010342Z digest=sha256:19a27f7376b4b288c2b5b9fab1d1645e18f739ded658737bdb1d0fefd75e1838

Observation 72c9f816-fa26-4953-b09a-8c18f05d28f8 · outbound

This paper cites Curing Miracle Steps in LLM Mathematical Reasoning with Rubric Rewards.

Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents Curing Miracle Steps in LLM Mathematical Reasoning with Rubric Rewards

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-06-29T07:33:13.441071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:31:00.010342Z digest=sha256:51a4e631629dd99c3f30a7b7e13e0f2d4fa1ae34fdfa463f6c89882e2a2985c6

Observation 7e2bdf7f-a62d-40e0-bbfe-2779a55f5ccd · outbound

This paper cites A Survey of Reinforcement Learning for Large Reasoning Models.

Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents A Survey of Reinforcement Learning for Large Reasoning Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-06-29T07:33:13.446098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:31:00.010342Z digest=sha256:9a23dd5dab8ca4b4c54aa7f2399792ff46f3f00f33f70af9791eb30923793281

Observation ae60c598-1e6c-45c4-9fae-cdfddec9e110 · outbound

This paper cites Deepresearcher: Scaling deep research via reinforcement learning in real-world environments.

Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents Deepresearcher: Scaling deep research via reinforcement learning in real-world environments

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-29T07:31:00.010342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:31:00.010342Z digest=sha256:265ea87d6ab958735c757a7295cb0639e11dbcab23c24049cdafd885c436ad71

Observation fe5963fd-333f-4947-943d-ca58fb2be704 · outbound

This paper cites MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents.

Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-06-29T07:33:13.435675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:31:00.010342Z digest=sha256:f24717b3a7fffb535bdd5498bef3039ca326b60de5ed3b353aac774e47d5cecf

Observation 458f09ff-6840-40d5-9099-64b0f79ced31 · outbound

This paper cites The key point is architectural: once the history is compressed into memory, downstream reasoning and action selection can only access the information preserved inm t.

Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents The key point is architectural: once the history is compressed into memory, downstream reasoning and action selection can only access the information preserved inm t

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-29T07:31:00.010342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:31:00.010342Z digest=sha256:fd6a43ca0fa9a0312d46055193c5c095126a50bb5b27c74527e17aaf7e519db2

Observation 1e938751-8bf7-4b8c-9f8a-aa1ad74fe6a6 · outbound

This paper cites Based on current memory, what is the answer to the question?.

Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents Based on current memory, what is the answer to the question?

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:33:13.433128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:31:00.010342Z digest=sha256:c4cc763b116e05a0c32736f4fe37dad9737e6eecc8cdfaf4a12ec82b148597f5

Pith citing papers

Observation 07c7eee8-2297-4021-9873-67667021732c · inbound

ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning cites this paper.

ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-02T03:26:28.757256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T10:07:16.700499Z digest=sha256:524dd2d5b9ad67148353b3283295f0be67ff14af4096c97ed9705ce5d40808b0