Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T07:31:00.010342Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 1 inbound Pith citation observation for arXiv:2605.30159.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T07:31:00.010342Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-28T10:07:16.700499Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-02T03:26:28.755836Z
24 of 24 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b585e253-a62f-420b-9a00-216649f204a1 · outbound
Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bebaa282-e6ff-49e3-91c6-b9c8ddb06350 · outbound
Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents Process Reinforcement through Implicit Rewards
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8452a397-d936-41a6-b52c-1f1547935804 · outbound
Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents DeepSeek-V3 Technical Report
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 04094eff-3cb8-4982-b4b2-1901071e3489 · outbound
Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 359d6454-51cb-451a-89d6-ad8e082fec26 · outbound
Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents Memory for autonomous llm agents: Mechanisms, evaluation, and emerging frontiers
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 98e01cf1-c8ee-41b9-944c-fc76ca98086e · outbound
Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents The Llama 3 Herd of Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4df8dfc9-3156-404a-90ff-4d5ce39d6888 · outbound
Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4888566a-4858-4236-a375-faf78d8a3704 · outbound
Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 78574955-d4b7-4858-8373-56432812f4fe · outbound
Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents Language Models (Mostly) Know What They Know
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 632117b4-0a02-4cb0-bda5-82a73380ecf6 · outbound
Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents MemOS: An Operating System for Memory-Augmented Generation (MAG) in Large Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 35e0d7ff-494d-4f4f-9086-a51e9ba35b79 · outbound
Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents CoRR , volume =
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 651aa359-ad3a-482a-80b2-e287fb96455c · outbound
Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents MemGPT: Towards LLMs as Operating Systems
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 94f834d4-9b01-44a7-8798-b16d8535b2f7 · outbound
Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents Sufficientstatisticsintheoptimumcontrolofstochasticsystems
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a017bad6-bd36-4321-8257-45c8144e4699 · outbound
Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents When to mem- orize and when to stop: Gated recurrent memory for long-context reasoning.arXiv preprint arXiv:2602.10560
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 95236416-bdb3-4e50-8797-557039790449 · outbound
Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents An Exact Solution Approach for Portfolio Op- timization Problems Under Stochastic and Integer Constraints
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation aa27fdd6-a2aa-4ece-8a39-4bc4fe442977 · outbound
Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents Mem-{\alpha}: Learning Memory Construction via Reinforcement Learning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e17ebc96-e040-4f10-a35e-e316a78efe2e · outbound
Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents A-MEM: Agentic Memory for LLM Agents
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a23a96b3-fd48-4893-9fcc-49dd1682cf3e · outbound
Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 72c9f816-fa26-4953-b09a-8c18f05d28f8 · outbound
Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents Curing Miracle Steps in LLM Mathematical Reasoning with Rubric Rewards
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7e2bdf7f-a62d-40e0-bbfe-2779a55f5ccd · outbound
Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents A Survey of Reinforcement Learning for Large Reasoning Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ae60c598-1e6c-45c4-9fae-cdfddec9e110 · outbound
Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents Deepresearcher: Scaling deep research via reinforcement learning in real-world environments
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe5963fd-333f-4947-943d-ca58fb2be704 · outbound
Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 458f09ff-6840-40d5-9099-64b0f79ced31 · outbound
Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents The key point is architectural: once the history is compressed into memory, downstream reasoning and action selection can only access the information preserved inm t
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e938751-8bf7-4b8c-9f8a-aa1ad74fe6a6 · outbound
Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents Based on current memory, what is the answer to the question?
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 07c7eee8-2297-4021-9873-67667021732c · inbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.