Pith. sign in

Paper Citation Record · LEDGER

MARFT: Multi-Agent Reinforcement Fine-Tuning

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2504.16129.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.16129 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T07:56:46.663818Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T13:49:51.542821Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 36f56b3d-8acd-4ef2-91d1-c7668399995e · inbound

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems cites this paper.

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 168

Resolution
verified exact
arxiv_id, observed 2026-06-02T02:03:31.668542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T21:39:49.832151Z digest=sha256:9e153c5a8e1f05f0f44c7aa2085d9163c3c7173879ee608029933dc5f9366dc7

Observation f15f0f2e-b42b-43d9-8268-dc97d30a64e2 · inbound

A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems cites this paper.

A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-06-02T02:03:31.668542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T23:21:42.029285Z digest=sha256:7a81469a4fd0f526e93831dfce72edd797516e73603fcfec434bef477af723d1

Observation a4c744a8-7fc4-4fc7-a50f-d1ac819ee512 · inbound

MASPRM: Multi-Agent System Process Reward Model cites this paper.

MASPRM: Multi-Agent System Process Reward Model MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T07:56:46.663818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:56:46.663818Z digest=sha256:8a39bfd8d1aab62d2570ef1046554c09815ed166fbca53426cb8cff400de58e9

Observation 5b74506c-3120-45dc-b88d-5ae8aa0f055b · inbound

Memory in the Age of AI Agents cites this paper.

Memory in the Age of AI Agents MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-06-02T02:03:31.668542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T18:18:19.911342Z digest=sha256:9fbe8a03db43533c26685d772e936c3eb2762dc4e004d0dbc25b40b0a1eae031

Observation 3c7a5d6a-ba58-4d98-b73b-599239c6c62f · inbound

Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic cites this paper.

Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-02T02:03:31.668542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T09:47:47.051969Z digest=sha256:4197f5e0f683372b632bf161c2238df8a59b48c6eee3b388cd30b4601566903d

Observation f9d5ae1a-8d7a-4685-a6ea-d898d0abbce9 · inbound

Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic cites this paper.

Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:39.487517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:39.487517Z digest=sha256:8e02bc6658c9af8187a8dfaa45d560faed7f80b78a1a84e749a416a55d341775

Observation d52a709f-a445-4e56-8c54-613947da1d0a · inbound

Learning from Self-Debate: Preparing Reasoning Models for Multi-Agent Debate cites this paper.

Learning from Self-Debate: Preparing Reasoning Models for Multi-Agent Debate MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-02T02:03:31.668542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T14:29:15.751497Z digest=sha256:b7c313cebaf6f5bdac8c03b1e0c9a0e67fb68742dfbefa5fbb951c473a606028

Observation d4118c98-6069-45c6-8a3f-0565b18f7ca4 · inbound

Joint Optimization of Multi-agent Memory System cites this paper.

Joint Optimization of Multi-agent Memory System MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-02T02:03:31.668542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T12:12:35.056095Z digest=sha256:e2348f059f21ba247f00b8e373909602160bafbc66d5633dda9a928e7be3002e

Observation 38277402-9449-402c-98e9-4dd9706248c1 · inbound

Reinforcement Learning for LLM-based Multi-Agent Systems through Orchestration Traces cites this paper.

Reinforcement Learning for LLM-based Multi-Agent Systems through Orchestration Traces MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-02T02:03:31.668542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T18:44:27.685266Z digest=sha256:94109047e5cb68cb28f99768eed854115cb28af1b12c34117c45c63c9ea7cc6c

Observation 5290e0a4-4aef-4a39-967a-aba1123935de · inbound

Tree-based Credit Assignment for Multi-Agent Memory System cites this paper.

Tree-based Credit Assignment for Multi-Agent Memory System MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-02T02:03:31.668542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T15:40:56.527868Z digest=sha256:ad1fcf8dfebb344e7244fba08912b7cbf3b74fc04173267d16178932a3efce68

Observation d95171eb-2196-449b-a94b-089862af93c0 · inbound

AIPO: Learning to Reason from Active Interaction cites this paper.

AIPO: Learning to Reason from Active Interaction MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-02T02:03:31.668542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T01:17:28.124867Z digest=sha256:ff2ff0a41267d4236b3426a4eda8335353368d31920bc0cd49e6360a71266a40

Observation 3a1659cf-20a2-4c84-ab5b-74896ef270df · inbound

AIPO: Learning to Reason from Active Interaction cites this paper.

AIPO: Learning to Reason from Active Interaction MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-02T02:03:31.668542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T18:07:27.492419Z digest=sha256:4c5b2113d2bd4ecc963d1a3d32f6968f5c4b4233b2cc98b7e1f6ec1701404304

Observation 5913973e-4df3-422b-9779-3c5c0fabb6be · inbound

Reinforced Collaboration in Multi-Agent Flow Networks cites this paper.

Reinforced Collaboration in Multi-Agent Flow Networks MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-02T02:03:31.668542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T20:42:48.438057Z digest=sha256:065dc847b02d8826fa251d8c70f11a063abd171cba1a61a7393bc35e49772107

Observation 3f0151e0-bd36-4962-acfb-4667cca22f31 · inbound

Position: Agentic AI System Is a Foreseeable Pathway to AGI cites this paper.

Position: Agentic AI System Is a Foreseeable Pathway to AGI MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-06-02T02:03:31.668542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-14T20:10:36.101426Z digest=sha256:f2fe5d457e5b528e541ab070ae73f0e2b9aebfdc8b3c30425605c7f5ce283aaa

Observation f771318b-14e8-45eb-841f-7b4481dbacac · inbound

Traj-Evolve: A Self-Evolving Multi-Agent System for Patient Trajectory Modeling in Lung Cancer Early Detection cites this paper.

Traj-Evolve: A Self-Evolving Multi-Agent System for Patient Trajectory Modeling in Lung Cancer Early Detection MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 48

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:26:22.822170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T14:20:06.381334Z digest=sha256:0a293df0be51e1a1bb8f10476c4a2eeeddff580347203a9ed1d6007d6e654646

Observation 112dcb44-05b2-4c69-8901-48c21d5a6403 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 113

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:09:40.650675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:a9a279622d1230592747b4822b0e76aeebc33d554f0d93ab3da5ea9494b9f8ef

Observation fec67690-d95e-4a18-9243-724f69b41a88 · inbound

Where Do CoT Training Gains Land in LLM based Agents? cites this paper.

Where Do CoT Training Gains Land in LLM based Agents? MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-04T13:49:51.544355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-26T04:55:14.293452Z digest=sha256:5271c9d639991e6ef353e84c1b824437155eba45d924eda6e95a02bb6dbea435