Pith. sign in

Paper Citation Record · LEDGER

ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2503.09501.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.09501 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T10:01:25.490489Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T22:56:37.763786Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 31a01da1-b436-419f-9ef4-4a8789cbfc67 · inbound

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems cites this paper.

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning

Reference 169

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:42:10.681342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T21:39:49.832151Z digest=sha256:ba30c8fa739ba8ae4435ac7d4c363d89ae6ae837050c30aaafaefa91e144b6ac

Observation 03613c47-f757-4d02-85a3-20eb907510f6 · inbound

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning cites this paper.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:25.490489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:25.490489Z digest=sha256:5b7eac2d27c3b02c9bba75e896aa4cd67bdbfacfd861c5a9a0a4bd81a4acd827

Observation 1839bc5c-f1d9-4f5c-a6b8-67beba00fd52 · inbound

PASK: Toward Intent-Aware Proactive Agents with Long-Term Memory cites this paper.

PASK: Toward Intent-Aware Proactive Agents with Long-Term Memory ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:30:59.641074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:05:40.242925Z digest=sha256:93c2cf6e9ad7542a59bac04b502a906c67d253d01f1e3a68e979d0799db00f92

Observation 803fa2c0-1856-4033-93e7-2e4a992ee665 · inbound

Weak-Link Optimization for Multi-Agent Reasoning and Collaboration cites this paper.

Weak-Link Optimization for Multi-Agent Reasoning and Collaboration ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:58:13.232273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T08:53:03.420926Z digest=sha256:2625a83817e7e2198df673f2a4aacb141986809605a407c472f69acaede5086d

Observation be9f6412-e539-4728-9fed-975d56a6b911 · inbound

HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs cites this paper.

HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:54:48.595081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:51:37.506096Z digest=sha256:c7862c37cfb5f3c7d84a7f9d2995e8e58a30da57d25c922eac4b87d604e1e0f1

Observation 0d912134-9d6c-4f83-a529-03f72eb64e6b · inbound

AIPO: Learning to Reason from Active Interaction cites this paper.

AIPO: Learning to Reason from Active Interaction ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:06:31.567058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T01:17:28.124867Z digest=sha256:c36d8d153fa5233d4628f47fa75796a2ac5d88b5a2e26754726b7dce65ea7b12

Observation 30be449b-bb44-4dcf-98b4-f100817906e0 · inbound

AIPO: Learning to Reason from Active Interaction cites this paper.

AIPO: Learning to Reason from Active Interaction ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-19T18:07:42.344993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T18:07:27.492419Z digest=sha256:3c9192c02c4a4b26ac2ad5f0f38d40be1915760e1ae14bdbbb057cebadab34f9

Observation 645471cc-29c6-4711-8751-8270dc713942 · inbound

Implicit Hierarchical GRPO: Decoupling Tool Invocation from Execution for Tool-Integrated Mathematical Reasoning cites this paper.

Implicit Hierarchical GRPO: Decoupling Tool Invocation from Execution for Tool-Integrated Mathematical Reasoning ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T11:23:14.035547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T11:21:39.492072Z digest=sha256:f79bd9a0b15dea58e97889320a11373db93e5d59f53a3128ed14664ba7642848

Observation b4ed6ee3-62ee-4e03-a436-b0852d366c7c · inbound

Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents cites this paper.

Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:21:21.800398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T09:18:15.753115Z digest=sha256:b1f6dde9cbb3b1a49aff6b34f12dffcfbc330b5cdf603da4cd671cf2da1c8dbf

Observation 600d072f-bf08-44dd-a9b8-39a11b16c190 · inbound

Design First, Code Later: Aesthetically Pleasing Template-Free Slides Generation cites this paper.

Design First, Code Later: Aesthetically Pleasing Template-Free Slides Generation ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T08:52:44.531560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T08:52:44.531560Z digest=sha256:d073de8b9fdc8ae99d4d01ce7e716ede049fe8b40ea31cf579fbab11f35edbc0

Observation fd45f8b0-9842-42cc-8c33-61a942f2e337 · inbound

Where Do CoT Training Gains Land in LLM based Agents? cites this paper.

Where Do CoT Training Gains Land in LLM based Agents? ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:49:51.510271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-26T04:55:14.293452Z digest=sha256:f270a548b63dd5c664f1b9ddda8def521c16cc2b3cfa60eea66f96cc16b321f3

Observation 6d755d8b-ccb0-41b4-a7e8-8d2bb96f16c4 · inbound

Mathematical methods of reinforcement learning cites this paper.

Mathematical methods of reinforcement learning ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning

Reference 117

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:56:37.765076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:da2c4b8aa0fac92571e8b92297cc5d444b62670647903a5fa39a040d1050efbf

Observation 27b1aa33-e165-4331-be6f-5a30bce0a4e6 · inbound

MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning cites this paper.

MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-14T03:32:02.298042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:32:02.298042Z digest=sha256:70c0f6df426dd5d5663bf82f0a484cffd8303240273d5410eba64d96a6bae836

Observation cf496511-afe6-448c-aa03-daec8daf50ec · inbound

Training Language Models to Cooperate with Inference-Time Controllers cites this paper.

Training Language Models to Cooperate with Inference-Time Controllers ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-30T13:09:01.454452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T13:09:01.454452Z digest=sha256:bff6919c60314aafeef3fea8ef9f69fa254e52e8ebd0fbdff25af8cb21188e01