Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2503.09501.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T10:01:25.490489Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-09T22:56:37.763786Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 31a01da1-b436-419f-9ef4-4a8789cbfc67 · inbound
Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning
Reference 169
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 03613c47-f757-4d02-85a3-20eb907510f6 · inbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1839bc5c-f1d9-4f5c-a6b8-67beba00fd52 · inbound
PASK: Toward Intent-Aware Proactive Agents with Long-Term Memory ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 803fa2c0-1856-4033-93e7-2e4a992ee665 · inbound
Weak-Link Optimization for Multi-Agent Reasoning and Collaboration ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation be9f6412-e539-4728-9fed-975d56a6b911 · inbound
HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0d912134-9d6c-4f83-a529-03f72eb64e6b · inbound
AIPO: Learning to Reason from Active Interaction ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 30be449b-bb44-4dcf-98b4-f100817906e0 · inbound
AIPO: Learning to Reason from Active Interaction ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 645471cc-29c6-4711-8751-8270dc713942 · inbound
Implicit Hierarchical GRPO: Decoupling Tool Invocation from Execution for Tool-Integrated Mathematical Reasoning ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b4ed6ee3-62ee-4e03-a436-b0852d366c7c · inbound
Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 600d072f-bf08-44dd-a9b8-39a11b16c190 · inbound
Design First, Code Later: Aesthetically Pleasing Template-Free Slides Generation ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd45f8b0-9842-42cc-8c33-61a942f2e337 · inbound
Where Do CoT Training Gains Land in LLM based Agents? ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6d755d8b-ccb0-41b4-a7e8-8d2bb96f16c4 · inbound
Mathematical methods of reinforcement learning ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning
Reference 117
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 27b1aa33-e165-4331-be6f-5a30bce0a4e6 · inbound
MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf496511-afe6-448c-aa03-daec8daf50ec · inbound
Training Language Models to Cooperate with Inference-Time Controllers ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.