Pith. sign in

Paper Citation Record · LEDGER

MARFT: Multi-Agent Reinforcement Fine-Tuning

As of 23 August 2026, this Paper Citation Record lists 100 of 118 outbound references and 24 inbound Pith citation observations for arXiv:2504.16129.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.16129 v5

Coverage vector

measured 100 of 118 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:42:31.475666Z

measured 124 of 124 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:28:56.795142Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T13:49:51.542821Z

Reference resolution

100 of 118 outbound references displayed

  • verified exact6
  • verified fuzzy7
  • unresolved86
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9eb51eaa-6401-4b10-a06b-35eb7f9d0be8 · outbound

This paper cites Springer New York, New York, NY, 2008.

MARFT: Multi-Agent Reinforcement Fine-Tuning Springer New York, New York, NY, 2008

Reference 1

Resolution
verified exact
doi, observed 2026-08-16T11:42:31.618499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-16T11:42:31.153282Z digest=sha256:5a801fe33ca35bd0747a9da8bbbf481fc3a2c3a196f082a9c9254b5518949f25

Observation cf42ef0a-fe0e-4fe4-a7c9-a7207bbca069 · outbound

This paper cites URL https://arxiv.org/abs/2407.21075.

MARFT: Multi-Agent Reinforcement Fine-Tuning URL https://arxiv.org/abs/2407.21075

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.157278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.157278Z digest=sha256:83280a931cc93df75a52c378ad8489f6c20d44253a26284fbe201ea62b5b5ac5

Observation aeb3c174-d91c-44ba-92be-bb5d8092154d · outbound

This paper cites Introducing the model context protocol.

MARFT: Multi-Agent Reinforcement Fine-Tuning Introducing the model context protocol

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.160564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.160564Z digest=sha256:83bc461526b3c9c45ce495ec6870eb275c3ca8003b080e76fad4f808ae4e0cd0

Observation bf72fa2d-d3c1-4bd0-acef-ff85a52afa0f · outbound

This paper cites A comprehensive survey of multiagent reinforcement learning.

MARFT: Multi-Agent Reinforcement Fine-Tuning A comprehensive survey of multiagent reinforcement learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.163852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.163852Z digest=sha256:29f104358f4e758e904f262e7e515206557842a19ec0d166a11b91648fd5d9cd

Observation 45951589-a087-4bcc-a148-cab6eebc559e · outbound

This paper cites Owl: Optimized workforce learning for general multi-agent assistance in real-world task automation.

MARFT: Multi-Agent Reinforcement Fine-Tuning Owl: Optimized workforce learning for general multi-agent assistance in real-world task automation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.167173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.167173Z digest=sha256:47d7d0083c048869fda31f26ffb6705303a2394bebb28905147d7e34b3e40f27

Observation 312e0411-4f12-4c75-8e17-f5336fce08a5 · outbound

This paper cites On the utility of learning about humans for human-ai coordination.

MARFT: Multi-Agent Reinforcement Fine-Tuning On the utility of learning about humans for human-ai coordination

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.170448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.170448Z digest=sha256:6d9c6eff3c82b49fca0c8db961062e72d404da794efb08d8f5f7e3be8bb8fa83

Observation 07c696e2-ad0c-493e-bfcb-74ab506ae3df · outbound

This paper cites Grounding large language models in interactive environments with online reinforcement learning.

MARFT: Multi-Agent Reinforcement Fine-Tuning Grounding large language models in interactive environments with online reinforcement learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.174184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.174184Z digest=sha256:e9ffe789410a1524289c9f44db6123365bf629a19f6070674aa4c20f69b53d8f

Observation 577a50e7-6282-4b9f-91bb-502b5d646618 · outbound

This paper cites Why Do Multi-Agent LLM Systems Fail?.

MARFT: Multi-Agent Reinforcement Fine-Tuning Why Do Multi-Agent LLM Systems Fail?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.177076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.177076Z digest=sha256:86015c781f8c10674ff3953c95ca1dae43ac008b2771c5dc802128aa64b2ed70

Observation 0c542a79-84f9-4874-9031-76dfda89e004 · outbound

This paper cites Communication-efficient actor-critic methods for homogeneous markov games.

MARFT: Multi-Agent Reinforcement Fine-Tuning Communication-efficient actor-critic methods for homogeneous markov games

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.180382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.180382Z digest=sha256:0b3534796795886c5737e8cb72d5f36e4097f98705bdb98b617bc63d82cda06e

Observation 66559784-6238-484e-98b1-08360b7ece00 · outbound

This paper cites Octopus: On-device language model for function calling of software APIs.

MARFT: Multi-Agent Reinforcement Fine-Tuning Octopus: On-device language model for function calling of software APIs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.183464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.183464Z digest=sha256:c9469403613220f1795ae17317e1e815ec1ea26772560a2a5ae2e4ae5304fcaa

Observation 9cf33de4-bcd9-4672-ba52-678857627928 · outbound

This paper cites an unresolved cited work.

MARFT: Multi-Agent Reinforcement Fine-Tuning Unresolved cited work

Reference 11

Resolution
verified exact
doi, observed 2026-08-16T11:42:31.608308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-16T11:42:31.186767Z digest=sha256:287c0c2787ce3560b9b6405f50aa8d49bf22996a97d43451717ff46c40353853

Observation cba3c9cb-5505-4d76-b00a-0d29e0b695dd · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

MARFT: Multi-Agent Reinforcement Fine-Tuning Training Verifiers to Solve Math Word Problems

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.189868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.189868Z digest=sha256:33fb0ccb598de4de1c4a8407f80281a13f0b314a9bfd1885a460a5302307cbad

Observation 6f9537ee-898b-42ea-b158-e0839a9f2c9f · outbound

This paper cites AD-AutoGPT: An Autonomous GPT for Alzheimer's Disease Infodemiology.

MARFT: Multi-Agent Reinforcement Fine-Tuning AD-AutoGPT: An Autonomous GPT for Alzheimer's Disease Infodemiology

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.193301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.193301Z digest=sha256:b62ec8d60a6b3b9e6239cdf7fbd4a3e6bd199793b789532197039cf2beb42eb0

Observation 904c9c91-663e-4b44-a6b4-5ac6a6cd9fc7 · outbound

This paper cites Independent policy gradient methods for competitive reinforcement learning.

MARFT: Multi-Agent Reinforcement Fine-Tuning Independent policy gradient methods for competitive reinforcement learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.196620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.196620Z digest=sha256:6be5f01a0453b5d624469eb0678e8b6e43135e503e4f02cc588d91dc548ab274

Observation 5420bff4-53cb-4801-81fa-1c2f4fea28af · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

MARFT: Multi-Agent Reinforcement Fine-Tuning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.200098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.200098Z digest=sha256:1e33daf09788c7b393142d1ffd2f0bed044118d7e7857c5d7fe633190b786aa2

Observation e0b33f79-d612-448c-8da5-1ec5681b3c5e · outbound

This paper cites Pilarski, and Richard S.

MARFT: Multi-Agent Reinforcement Fine-Tuning Pilarski, and Richard S

Reference 16

Resolution
metadata mismatch
raw_fallback, observed 2026-08-16T11:42:32.522835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-16T11:42:31.203739Z digest=sha256:38bb3782154a94e6f62e42ca2be8c8abc39e67831bb92c0d89b3bc1db4656376

Observation cd43666d-c93b-48f1-94d7-bd83db5f9518 · outbound

This paper cites Interactive debugging and steering of multi-agent ai systems.

MARFT: Multi-Agent Reinforcement Fine-Tuning Interactive debugging and steering of multi-agent ai systems

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.206758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.206758Z digest=sha256:4de604c1f3ff3fb70b0ff72d0b1cb61b33690d59e475e327b12f86318a86c9a7

Observation b2614f40-55cf-45e7-ac1a-b8a168a2a712 · outbound

This paper cites TinyAgent: Function Calling at the Edge.

MARFT: Multi-Agent Reinforcement Fine-Tuning TinyAgent: Function Calling at the Edge

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.209789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.209789Z digest=sha256:96d80fbfef548137521155e9135bcf42f37ab36c292b5026a0ea173459038813

Observation b7ce1f31-b946-472a-ba34-a710a67a8dc5 · outbound

This paper cites Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution.

MARFT: Multi-Agent Reinforcement Fine-Tuning Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.213132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.213132Z digest=sha256:e475ae0830a7588dd36b71a3412ef269d7c12db12ff77cfdb204890cf34c8448

Observation e3188bd8-a64f-4335-b449-081648881a23 · outbound

This paper cites Counterfactual multi-agent policy gradients.

MARFT: Multi-Agent Reinforcement Fine-Tuning Counterfactual multi-agent policy gradients

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.216461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.216461Z digest=sha256:2dda32f081602751daf97aa4660bda50331162a76e1405962c25cb60d4c18ca8

Observation 2fede6fb-7a62-486c-a991-836c9ecbd096 · outbound

This paper cites ANP - Agent Network Protocol.

MARFT: Multi-Agent Reinforcement Fine-Tuning ANP - Agent Network Protocol

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.219523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.219523Z digest=sha256:3b1692289a3a81c164e5628d5f617c8da79a323a84936c5c005cf2cef806f404

Observation 70a6b63b-ba86-4339-8e35-3828bf2b45be · outbound

This paper cites Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use.

MARFT: Multi-Agent Reinforcement Fine-Tuning Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.222958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.222958Z digest=sha256:b25e100b2b09460d3be90dc1fddbe9427cfe6ac2d3e3702cd0738a059bf69c38

Observation a5c1dfb9-e03f-43b0-a3a8-345d3cfcacc2 · outbound

This paper cites Towards an AI co-scientist.

MARFT: Multi-Agent Reinforcement Fine-Tuning Towards an AI co-scientist

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.226476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.226476Z digest=sha256:f8ff04649f3248ee582d013ac2f9b2cea9a411451c6b31dfcec226fa4a5125d6

Observation aab2846d-3a9c-49c1-92ae-0f80103dca37 · outbound

This paper cites Variance reduction techniques for gradient estimates in reinforcement learning.

MARFT: Multi-Agent Reinforcement Fine-Tuning Variance reduction techniques for gradient estimates in reinforcement learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.229700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.229700Z digest=sha256:6a8998f50cbc38bcf9f3e21e02d925e173975156f34f515ae6022556c897ac4e

Observation cbf625a3-4efb-4d3b-91af-96cb2efb4149 · outbound

This paper cites Reinforcement learning with deep energy-based policies.

MARFT: Multi-Agent Reinforcement Fine-Tuning Reinforcement learning with deep energy-based policies

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.232779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.232779Z digest=sha256:fa0b941feb7d28cb0a50e3b025326d391dcf1dbad3b59e80349884e30fcc0594

Observation 28ad4c4a-f641-4155-a519-27a45a4bc102 · outbound

This paper cites Emergence of Locomotion Behaviours in Rich Environments.

MARFT: Multi-Agent Reinforcement Fine-Tuning Emergence of Locomotion Behaviours in Rich Environments

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.235834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.235834Z digest=sha256:5a9d9cf9010a0d106945b39b91a039c6fed4f03384f4d7c4f383146cd028c7a5

Observation b239c855-55f4-4a2c-b414-e59222095fa9 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

MARFT: Multi-Agent Reinforcement Fine-Tuning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.239163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.239163Z digest=sha256:a5eb5e0ed2f4c669de8c5f99677f2bd6715df78338039e2ef363cb445eaa8b81

Observation ddd5af33-fe5b-475c-9bd5-b876fa125048 · outbound

This paper cites Meta GPT : Meta programming for a multi-agent collaborative framework.

MARFT: Multi-Agent Reinforcement Fine-Tuning Meta GPT : Meta programming for a multi-agent collaborative framework

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.242092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.242092Z digest=sha256:f89ffbc0cc0142773ea6544e4adc7857b4e4b9bfdad18f2730c42d5854c0a175

Observation bce5ffdc-368f-40f9-b5cc-f202f8c6b546 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

MARFT: Multi-Agent Reinforcement Fine-Tuning REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.245094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.245094Z digest=sha256:2a73353e9f717ff7becda77101046e0f87934e874bd459212f81b3b13ca4ed0d

Observation a7e9e30b-fc99-4ce2-89f2-bc2a4dd02fa8 · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

MARFT: Multi-Agent Reinforcement Fine-Tuning OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.248307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.248307Z digest=sha256:abc37d3343e58703d8b7ddb919f9398db9e15c85b0aa98c416a1af9d6efa3a9c

Observation 84b1caae-229c-4e9c-bb34-14ebd3c5eb52 · outbound

This paper cites AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation.

MARFT: Multi-Agent Reinforcement Fine-Tuning AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.251836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.251836Z digest=sha256:baeb3c9759ca1ac2b93b05941cccae478cd0533bd6da6d7bf17ad72964297811

Observation 06230c10-326e-4d0e-b490-1ba0a9f651fa · outbound

This paper cites Qwen2.5-Coder Technical Report.

MARFT: Multi-Agent Reinforcement Fine-Tuning Qwen2.5-Coder Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.255201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.255201Z digest=sha256:525dde6304551465028dfd6166bc13a3b7b496384d1e73b229efd57bd326f33f

Observation 4f38a103-9f70-47f2-8cbd-a79961ef8dbe · outbound

This paper cites Deep reinforcement learning for swarm systems.

MARFT: Multi-Agent Reinforcement Fine-Tuning Deep reinforcement learning for swarm systems

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.258403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.258403Z digest=sha256:d0eda5b5cb86f33efad1dbb48c742973f071a5e43f746be1bebff447b5445411

Observation 99477a6a-5bf2-4736-bf47-2283c6002fbe · outbound

This paper cites From LLMs to LLM-based Agents for Software Engineering: A Survey of Current, Challenges and Future.

MARFT: Multi-Agent Reinforcement Fine-Tuning From LLMs to LLM-based Agents for Software Engineering: A Survey of Current, Challenges and Future

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.261523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.261523Z digest=sha256:188ae34c989132e3bc45caf6fb461e7715a280e34a44b03e6d586d870cefb0b2

Observation 7fb57dc9-2e40-4db8-b948-6aabb2fc9d2c · outbound

This paper cites Joshi, Hanna Moazam, Heather Miller, Matei Zaharia, and Christopher Potts.

MARFT: Multi-Agent Reinforcement Fine-Tuning Joshi, Hanna Moazam, Heather Miller, Matei Zaharia, and Christopher Potts

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.264758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.264758Z digest=sha256:d15cb26e3aa54394468aaa28c8f7768d60023ca290841e8f04b1a888d7cfd02f

Observation 0c77f10d-abb5-4e00-ab3d-dff731c2990d · outbound

This paper cites an unresolved cited work.

MARFT: Multi-Agent Reinforcement Fine-Tuning Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.268014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.268014Z digest=sha256:05b2bcd95772b1f44fcca407a099d2899fad03c893601c2a39b53d5544dad325

Observation bf3dcbf6-bff3-4654-b789-9f1afa0a456e · outbound

This paper cites Multi-agent reinforcement learning for traffic signal control.

MARFT: Multi-Agent Reinforcement Fine-Tuning Multi-agent reinforcement learning for traffic signal control

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.271026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.271026Z digest=sha256:d91302007246ff8d74ec98c9c751189bc93d1b79584008a60b702084bac5ff04

Observation b8f90959-4df6-41df-b33a-34e3c08413e7 · outbound

This paper cites Actor-critic algorithms.

MARFT: Multi-Agent Reinforcement Fine-Tuning Actor-critic algorithms

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.273919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.273919Z digest=sha256:82116eec61d70f37a01f62179bf0761b43ee51996fd00f423f2d2962c0380677

Observation 106d41aa-b274-4ea1-a282-d0518ae88c12 · outbound

This paper cites Buy 4 REINFORCE samples, get a baseline for free!, 2019.

MARFT: Multi-Agent Reinforcement Fine-Tuning Buy 4 REINFORCE samples, get a baseline for free!, 2019

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.277306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.277306Z digest=sha256:25c8e156eee1103f0a3992e06f46fcdbd94709e257a1544c957b2a965d93491b

Observation 9e78c747-d96d-44cb-aa7f-74dbe1086b73 · outbound

This paper cites Settling the variance of multi-agent policy gradients.

MARFT: Multi-Agent Reinforcement Fine-Tuning Settling the variance of multi-agent policy gradients

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.280253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.280253Z digest=sha256:31925a34717cd5d69216d45d7b9c6ee1abf5392c3a4413ca818a1cd9a63abddb

Observation e195a704-7623-436b-a813-b4ed759a2c00 · outbound

This paper cites Trust region policy optimisation in multi-agent reinforcement learning.

MARFT: Multi-Agent Reinforcement Fine-Tuning Trust region policy optimisation in multi-agent reinforcement learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.283446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.283446Z digest=sha256:7cba8145e6c0b64b484a6d746bd09e482ed3ef7c14434da5acec1a3c2a0ae932

Observation b3728b33-b2d8-407a-979b-f449c7e58788 · outbound

This paper cites Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.

MARFT: Multi-Agent Reinforcement Fine-Tuning Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.286207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.286207Z digest=sha256:e3f3e8c6d66f11400c5b34cd69ed11a2494e0f817d6bad8387d45b01a506a289

Observation 2fdfa293-6285-44c5-bc15-baf26a262a11 · outbound

This paper cites Multi-agent actor-critic for mixed cooperative-competitive environments.

MARFT: Multi-Agent Reinforcement Fine-Tuning Multi-agent actor-critic for mixed cooperative-competitive environments

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.289275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.289275Z digest=sha256:411756872565196a0be7a2cba360de4e83606abf9e2537cbcb2747ff21562eba

Observation dedd5286-1cc7-4070-b1ca-b34d47eb14cd · outbound

This paper cites ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities.

MARFT: Multi-Agent Reinforcement Fine-Tuning ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.292282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.292282Z digest=sha256:a239ddf69be134b09541abc2867b168d44b2e2af79251b944db429b6ef2a56da

Observation a1930bc1-f3c3-4bc4-8148-4e877240b7ec · outbound

This paper cites Coevolving with the other you: Fine-tuning llm with sequential cooperative multi-agent reinforcement learning.

MARFT: Multi-Agent Reinforcement Fine-Tuning Coevolving with the other you: Fine-tuning llm with sequential cooperative multi-agent reinforcement learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.295863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.295863Z digest=sha256:593f745068c91bfb8985835cbbe61afc4093aa87063f1c791d4d247a69a05c1d

Observation 5cae160b-b902-4473-8189-1e6ab5444370 · outbound

This paper cites Roco: Dialectic multi-robot collaboration with large language models, 2023.

MARFT: Multi-Agent Reinforcement Fine-Tuning Roco: Dialectic multi-robot collaboration with large language models, 2023

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.298864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.298864Z digest=sha256:c0cfec332e2b36232b43e430b0a35b1de62a8d4a6b8912cf2a896ab5da7c19f2

Observation aeb4b120-4dff-4b94-aec4-dbc876e81ca5 · outbound

This paper cites Laurent, and Nadine Le Fort-Piat.

MARFT: Multi-Agent Reinforcement Fine-Tuning Laurent, and Nadine Le Fort-Piat

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.302085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.302085Z digest=sha256:b9d07107246fd27afde9645fb70f695d9e1adf9349e738a32cf36ab90fe60549

Observation 7591dcbf-c723-4311-b81b-5f86c713f77d · outbound

This paper cites Coordinated multi-robot exploration under communication constraints using decentralized markov decision processes.

MARFT: Multi-Agent Reinforcement Fine-Tuning Coordinated multi-robot exploration under communication constraints using decentralized markov decision processes

Reference 48

Resolution
verified exact
doi, observed 2026-08-16T11:42:31.587327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-16T11:42:31.305215Z digest=sha256:89e90732b8bf1459b24b942c27d503a14b9f475bfdc1666d6328c5364651af04

Observation 5d2593b8-0dd2-44d3-95f6-9c5cc34ba5d6 · outbound

This paper cites GAIA : a benchmark for general AI assistants.

MARFT: Multi-Agent Reinforcement Fine-Tuning GAIA : a benchmark for general AI assistants

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.308241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.308241Z digest=sha256:43a96248e4ce62488fc1b1b249227939f6a621ff99302605c00a87dd685ee50a

Observation 437a705b-88c4-41b2-b813-1d2dc5b3214e · outbound

This paper cites Steps toward artificial intelligence.

MARFT: Multi-Agent Reinforcement Fine-Tuning Steps toward artificial intelligence

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.311461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.311461Z digest=sha256:cb3f6a00fadf43c7f6b83de26b3af9e8b7f1249909bd17e86f9fa6366bab9328

Observation 0e1fb39d-a66a-40f7-8d96-19b58d0e0c04 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

MARFT: Multi-Agent Reinforcement Fine-Tuning Playing Atari with Deep Reinforcement Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.314775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.314775Z digest=sha256:c9a26f8025e3ed6764c12b1a8b5119c85652677f88bb7022acd11e04da84f6ea

Observation 336b26d2-1126-4858-b8e6-8fb3f19bcce3 · outbound

This paper cites Asynchronous methods for deep reinforcement learning.

MARFT: Multi-Agent Reinforcement Fine-Tuning Asynchronous methods for deep reinforcement learning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.318241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.318241Z digest=sha256:2c286124dd8446462a56ac6af95b686bbab4d57721defde8eb470ae021607470

Observation 28b9f397-7cca-4723-a715-cb4a1cef0559 · outbound

This paper cites Ppo improvement in different environments.

MARFT: Multi-Agent Reinforcement Fine-Tuning Ppo improvement in different environments

Reference 53

Resolution
verified exact
raw_fallback, observed 2026-08-16T11:42:32.163762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-16T11:42:31.321695Z digest=sha256:52e10c8eb0d768b22eb60005871fc28709b5b92a200b70f4d8820f41d2b6e9af

Observation ce14e7d8-f984-4da2-bd76-27e96c9df2f7 · outbound

This paper cites Richard Yu.

MARFT: Multi-Agent Reinforcement Fine-Tuning Richard Yu

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.324657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.324657Z digest=sha256:cf7e757191c49e73aa3b3758d655415ed4f9a155a5c18e9674edcfd77a8a573f

Observation de7cc274-a03d-4c3d-a610-5716d64fdf0a · outbound

This paper cites Oliehoek and Christopher Amato.

MARFT: Multi-Agent Reinforcement Fine-Tuning Oliehoek and Christopher Amato

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.330409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.330409Z digest=sha256:041b58ba6b401da3673d34ff136e20c351dbfb672fd2c70b2a7b2e2f5bf62282

Observation f1addb96-2a43-4f80-9922-8a470753e8aa · outbound

This paper cites Learning to reason with LLMs , sep 2024 a.

MARFT: Multi-Agent Reinforcement Fine-Tuning Learning to reason with LLMs , sep 2024 a

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.333687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.333687Z digest=sha256:007d798a5adf331a3bff41bc15d392ba09bb2ff600edbe444c550d45f9093537

Observation 58b2b400-5f50-4f79-8f6e-0f95bc292948 · outbound

This paper cites Openai's reinforcement fine-tuning research program, dec 2024 b.

MARFT: Multi-Agent Reinforcement Fine-Tuning Openai's reinforcement fine-tuning research program, dec 2024 b

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.336819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.336819Z digest=sha256:52ec8299762ba95ebfe1aad50624ba5b28b0012e4441863254d29bb6117d3506

Observation 2d675879-d0cf-499d-9c55-3d3a4995d6ba · outbound

This paper cites Training language models to follow instructions with human feedback.

MARFT: Multi-Agent Reinforcement Fine-Tuning Training language models to follow instructions with human feedback

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.339882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.339882Z digest=sha256:bb24414f15d53abb5a5a661477d047d57872889accde8817734d00f60b8c049b

Observation d2292abc-50ce-4013-8620-a735503e9408 · outbound

This paper cites MAPoRL: Multi-Agent Post-Co-Training for Collaborative Large Language Models with Reinforcement Learning.

MARFT: Multi-Agent Reinforcement Fine-Tuning MAPoRL: Multi-Agent Post-Co-Training for Collaborative Large Language Models with Reinforcement Learning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.343032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.343032Z digest=sha256:7b514596b3da7166ba7d697df182f9275079ef0884be2ffd78653e8b57fc5d45

Observation 12b03d0f-c3b0-43e3-a890-00f70113ab0d · outbound

This paper cites Gorilla: Large Language Model Connected with Massive APIs.

MARFT: Multi-Agent Reinforcement Fine-Tuning Gorilla: Large Language Model Connected with Massive APIs

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.346302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.346302Z digest=sha256:a68052e03db2d35bfc0859c96ff518bfb25de95a7129c9d56bf7da0ed6a599dd

Observation 3a9b14ca-6ca2-40be-b437-dbf6e5fa87ca · outbound

This paper cites Codeforces.

MARFT: Multi-Agent Reinforcement Fine-Tuning Codeforces

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.349806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.349806Z digest=sha256:b9424dbd4927b478848d06b0869860e60b4a1332c4e8ec86677b337fe34526ba

Observation f3b191a7-7944-4fef-a23e-8c9e5c33586b · outbound

This paper cites Virtualhome: Simulating household activities via programs.

MARFT: Multi-Agent Reinforcement Fine-Tuning Virtualhome: Simulating household activities via programs

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.352967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.352967Z digest=sha256:b6fd294eef75296b0f277acd7c27f0e25fbc108214d2824b477a9f78ff39376c

Observation f1b428c4-94c6-46f2-8986-9b0166c890e2 · outbound

This paper cites C hat D ev: Communicative agents for software development.

MARFT: Multi-Agent Reinforcement Fine-Tuning C hat D ev: Communicative agents for software development

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.356348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.356348Z digest=sha256:12a3eb43fc84210d65af0cefa442b7c564704bee8377fce2097f3b1c35e1c9a0

Observation b5503f34-32f3-4b32-878f-8ad8f95259d3 · outbound

This paper cites Qwq-32b: Unveiling the power of reinforcement learning in reasoning, March 2025.

MARFT: Multi-Agent Reinforcement Fine-Tuning Qwq-32b: Unveiling the power of reinforcement learning in reasoning, March 2025

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.359527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.359527Z digest=sha256:f58ee46b9cb63ffbd6a34acacd175492d251682cac2eec3ad33e6cdf1b7dd337

Observation 009e447b-a2a8-424d-bac2-12a3096920f0 · outbound

This paper cites Sensitivity of ${^{44}}$Ti and ${^{56}}$Ni production in CCSN shock-driven nucleosynthesis to reaction rates.

MARFT: Multi-Agent Reinforcement Fine-Tuning Sensitivity of ${^{44}}$Ti and ${^{56}}$Ni production in CCSN shock-driven nucleosynthesis to reaction rates

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:42:32.022059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-16T11:42:31.362578Z digest=sha256:b1bcd4eb358abd7483043278358755465eb1640f568b072ab51f0cf906394d6c

Observation 8094e94a-62b2-4df2-a4fe-0c12701b3bd3 · outbound

This paper cites Toolformer: language models can teach themselves to use tools.

MARFT: Multi-Agent Reinforcement Fine-Tuning Toolformer: language models can teach themselves to use tools

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.365967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.365967Z digest=sha256:8707a23d3d21e97a6509859b02fd2f77c7bcfab12ebdee410905fbcb748580c5

Observation 6697cf3f-da24-42c7-8728-d241bdc41596 · outbound

This paper cites Trust region policy optimization.

MARFT: Multi-Agent Reinforcement Fine-Tuning Trust region policy optimization

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.368973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.368973Z digest=sha256:ca8d8f8aa94d9340e2bb9b9a7a2c31f1d20ff144e4bf40c9a3ea631d5045fd93

Observation c72428e9-ad27-45df-b7b7-a199fa46206c · outbound

This paper cites Proximal Policy Optimization Algorithms.

MARFT: Multi-Agent Reinforcement Fine-Tuning Proximal Policy Optimization Algorithms

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.372268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.372268Z digest=sha256:cd50f217a6291172de2634f2eca0bfcd49a3b3c1aa19170d0b27d0be2faa8bbe

Observation 513b33d0-0bca-41c5-a0a6-42780e8273a5 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

MARFT: Multi-Agent Reinforcement Fine-Tuning High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.375174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.375174Z digest=sha256:774b3ba3715f51c4df09fbdf5d186cd95d4cc34f5b2737fd12f61360a4931f00

Observation 0320e9f5-14ba-4e54-9054-4f5589b84895 · outbound

This paper cites Safe, Multi-Agent, Reinforcement Learning for Autonomous Driving.

MARFT: Multi-Agent Reinforcement Fine-Tuning Safe, Multi-Agent, Reinforcement Learning for Autonomous Driving

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.378425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.378425Z digest=sha256:dbf0b6856313727c33c670084c529bbdb4c413d113d99b61b36d5cc2fd9b8005

Observation 82de1eca-c568-46cc-8a41-02b38eaff3bd · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

MARFT: Multi-Agent Reinforcement Fine-Tuning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.381905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.381905Z digest=sha256:819f937cf635f763b8faab6c3cddb601ff982fa2955aeb3bedfcecb0b1d3daf0

Observation 9ddc8129-b31e-40fc-9bc7-a1b0f3e88dfb · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

MARFT: Multi-Agent Reinforcement Fine-Tuning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.385183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.385183Z digest=sha256:ce6149a69bc3b5df5d1f472c46b0e22676f9f6f628df3fb2daf777a277cc7612

Observation 0986bcd4-b67b-4544-ade3-a3857defb731 · outbound

This paper cites Multiagent Systems: Algorithmic, Game-Theoretic, and Logical Foundations.

MARFT: Multi-Agent Reinforcement Fine-Tuning Multiagent Systems: Algorithmic, Game-Theoretic, and Logical Foundations

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.388778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.388778Z digest=sha256:06e8e0c00a1477629e1df75c882bd2923a033e7dc4453ffda998996366da1bab

Observation 8d0d90d0-8ac4-4e87-90f2-c9bc45b2f409 · outbound

This paper cites An empirical study on google research football multi-agent scenarios.

MARFT: Multi-Agent Reinforcement Fine-Tuning An empirical study on google research football multi-agent scenarios

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.391930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.391930Z digest=sha256:420d983ec2746358c1ede61e9918250597cea075bf89611c40a30dca9388be09

Observation 0e431dec-b1ef-4729-a7dd-b006147b42d0 · outbound

This paper cites Multiagent systems: A survey from a machine learning perspective.

MARFT: Multi-Agent Reinforcement Fine-Tuning Multiagent systems: A survey from a machine learning perspective

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.395316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.395316Z digest=sha256:7a74bfd2eadc20bfe99d93ceb31bf0b51d2d601b03df72feb01d3e22c190127e

Observation ff1b245d-d016-49e8-bd93-b17eb0b4dc47 · outbound

This paper cites Multiagent Finetuning: Self Improvement with Diverse Reasoning Chains.

MARFT: Multi-Agent Reinforcement Fine-Tuning Multiagent Finetuning: Self Improvement with Diverse Reasoning Chains

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.398412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.398412Z digest=sha256:72fb445b9bae273328461d1e6c5d7ba933b04fdfe1bef1484b7672592df0a71d

Observation 6643af35-db93-41bc-9c17-4594389a4ba9 · outbound

This paper cites Learning multiagent communication with backpropagation.

MARFT: Multi-Agent Reinforcement Fine-Tuning Learning multiagent communication with backpropagation

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.401844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.401844Z digest=sha256:873092dc29246acfb27f5b84b4ae56addd380ce1320fa4c2ffb53967f9c0dd39

Observation 01a9a781-19cb-459f-adb8-a8e9a4bdff1a · outbound

This paper cites Announcing the agent2agent protocol (a2a).

MARFT: Multi-Agent Reinforcement Fine-Tuning Announcing the agent2agent protocol (a2a)

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.405142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.405142Z digest=sha256:07a0194a553d010defc44a70bd42cdd168fb42b8f82e05a64674907cc884d8ff

Observation 3b004d92-20ac-45ab-b693-ab7a195df80c · outbound

This paper cites Dyna, an integrated architecture for learning, planning, and reacting.

MARFT: Multi-Agent Reinforcement Fine-Tuning Dyna, an integrated architecture for learning, planning, and reacting

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.408398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.408398Z digest=sha256:d0f555f30be305d10655e611746eacee14a12af852979d08830340845c78ef58

Observation dc8a11ab-5c68-4e6b-8f1f-37ac56d635bc · outbound

This paper cites Sutton and Andrew G.

MARFT: Multi-Agent Reinforcement Fine-Tuning Sutton and Andrew G

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.411534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.411534Z digest=sha256:69a4469d4d2d51513df8807bb4ccad82cdbcfdedd1cc792e77f98bb8aac030d6

Observation 938c959b-3c16-4d2d-a380-cc6c9705db27 · outbound

This paper cites Policy gradient methods for reinforcement learning with function approximation.

MARFT: Multi-Agent Reinforcement Fine-Tuning Policy gradient methods for reinforcement learning with function approximation

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.414449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.414449Z digest=sha256:4012ccd4e4258471d6d4de857ce681b40bc51746950405ccdf63e9ddb0f47fa8

Observation 52e2fb41-fcae-4ba3-8964-9384554f4d7d · outbound

This paper cites Temporal credit assignment in reinforcement learning.

MARFT: Multi-Agent Reinforcement Fine-Tuning Temporal credit assignment in reinforcement learning

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:42:32.742824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-16T11:42:31.417478Z digest=sha256:8346326a0e64cb52bdc4c7126641865888300432f116722564b30dca21a6c160

Observation e603ba33-1f2f-4cfd-93e9-ae996f1df2e1 · outbound

This paper cites Multi-agent reinforcement learning: independent versus cooperative agents.

MARFT: Multi-Agent Reinforcement Fine-Tuning Multi-agent reinforcement learning: independent versus cooperative agents

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:42:32.733310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-16T11:42:31.420883Z digest=sha256:08ce585dae20bff9ab84fbb8550efcc2fcf880828076e1bfe474f518e53e6696

Observation 7d4a8e78-f5a1-482f-8315-5d44b4939e2e · outbound

This paper cites True knowledge comes from practice: Aligning large language models with embodied environments via reinforcement learning.

MARFT: Multi-Agent Reinforcement Fine-Tuning True knowledge comes from practice: Aligning large language models with embodied environments via reinforcement learning

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:42:32.723674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-16T11:42:31.424024Z digest=sha256:ffb9b37fb2872fbc22cf26b5df43c77553a8831d308c2f952d61161a5804d1af

Observation ab0eec97-f14b-4070-8c17-c9e0a97e3576 · outbound

This paper cites Magis: Llm-based multi-agent framework for github issue resolution.

MARFT: Multi-Agent Reinforcement Fine-Tuning Magis: Llm-based multi-agent framework for github issue resolution

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:42:32.713798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-16T11:42:31.427024Z digest=sha256:60bc83568f426ab80e97cb45a4b274a7a98ecd27e82875abd9dac0dfb272cb52

Observation 996e0e0e-fdaf-4b55-ac50-b731aac57806 · outbound

This paper cites Mujoco: A physics engine for model-based control.

MARFT: Multi-Agent Reinforcement Fine-Tuning Mujoco: A physics engine for model-based control

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.430133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.430133Z digest=sha256:1eebd7ef8c919bf93a5f9f33e7214b2afc3289acc73051fffd428eb848c2cc00

Observation 3a04697a-d7e3-4f31-b89a-530667564d16 · outbound

This paper cites Multiagent learning: Basics, challenges, and prospects.

MARFT: Multi-Agent Reinforcement Fine-Tuning Multiagent learning: Basics, challenges, and prospects

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:42:32.703500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-16T11:42:31.433534Z digest=sha256:f88c22b90912e295f3dd03a68333ff69c20bf2a78531920b75fb0cfc32d77f67

Observation 51d8aef0-f0b1-40aa-a26b-843ed84c2e9e · outbound

This paper cites Reinforcement Learning and Markov Decision Processes, pp.\ 3--42.

MARFT: Multi-Agent Reinforcement Fine-Tuning Reinforcement Learning and Markov Decision Processes, pp.\ 3--42

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.436615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.436615Z digest=sha256:256587281f827bda081af0ec6fe07b1843d52455c0e32712eb12cabe17b9c856

Observation 3b761b9d-ba48-45d4-9211-6ce3200d4b98 · outbound

This paper cites Networked Agents in the Dark: Team Value Learning under Partial Observability.

MARFT: Multi-Agent Reinforcement Fine-Tuning Networked Agents in the Dark: Team Value Learning under Partial Observability

Reference 89

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:42:31.893996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-16T11:42:31.439907Z digest=sha256:664c66b1b9c7c2e5c324cc7e0cad426cc60fccd0dd3933af4a6ed7c925b0a5a8

Observation ae4875e3-81d5-4413-8ce1-2313e219f597 · outbound

This paper cites ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning.

MARFT: Multi-Agent Reinforcement Fine-Tuning ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.443115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.443115Z digest=sha256:2e1aa951c68c40b89acf5d930443108ff186dfb023df4b6f36359ce4d278b0ba

Observation ed4908c4-f17e-415f-b2cf-4e67c6c70550 · outbound

This paper cites OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models.

MARFT: Multi-Agent Reinforcement Fine-Tuning OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.446573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.446573Z digest=sha256:a938d32e72f0a63a3a93b23f5be496cf6855a058f44a238cef3704a8f745f468

Observation 267ecab9-1e6f-4f33-80e3-5e1436897793 · outbound

This paper cites Order matters: Agent-by-agent policy optimization.

MARFT: Multi-Agent Reinforcement Fine-Tuning Order matters: Agent-by-agent policy optimization

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.449916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.449916Z digest=sha256:afbed0549fe4097c25ef1ba6168d1e6823d6f99bfa6b880e49f30f4e3097c50c

Observation 967ed46c-f685-42df-a9c7-72582359b6db · outbound

This paper cites an unresolved cited work.

MARFT: Multi-Agent Reinforcement Fine-Tuning Unresolved cited work

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.453087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.453087Z digest=sha256:a162c76f38bfbd08af2b257a169a96c41a81b5295b79e5d05225556380ae950f

Observation 98767025-9dac-4a2a-9d2b-1fc78eb5dca6 · outbound

This paper cites CMATH: Can Your Language Model Pass Chinese Elementary School Math Test?.

MARFT: Multi-Agent Reinforcement Fine-Tuning CMATH: Can Your Language Model Pass Chinese Elementary School Math Test?

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.456270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.456270Z digest=sha256:bdcd805d04c806ab1d02d8aac7bdb4e7c38974f1f1225805bf2a6ff1556a5149

Observation e3f2efb9-0b21-4043-a730-ccd206397fcc · outbound

This paper cites Multiagent systems: a modern approach to distributed artificial intelligence.

MARFT: Multi-Agent Reinforcement Fine-Tuning Multiagent systems: a modern approach to distributed artificial intelligence

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:42:32.688412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-16T11:42:31.459790Z digest=sha256:a55927b5f466327893ee52e12df5f315dab762b8d8a148e581caa9e219a14fd7

Observation 766d4c40-d984-4175-8ce8-45bdddc97a5e · outbound

This paper cites Multi-agent reinforcement learning is a sequence modeling problem.

MARFT: Multi-Agent Reinforcement Fine-Tuning Multi-agent reinforcement learning is a sequence modeling problem

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:42:32.678258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-16T11:42:31.462892Z digest=sha256:0c0fb300a65edd0abb9e7faa4557ae63a6126b0a6a2bec263745fbed471145fd

Observation 896bea58-9d53-4cca-a2d7-d69207883b66 · outbound

This paper cites Entropy-Regularized Token-Level Policy Optimization for Language Agent Reinforcement.

MARFT: Multi-Agent Reinforcement Fine-Tuning Entropy-Regularized Token-Level Policy Optimization for Language Agent Reinforcement

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.466006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.466006Z digest=sha256:a3eb32dbf8f89e1a2b9d143681bda0fce245f9276e7450c5fd37b32acf6d9a03

Observation 769bf109-d734-4115-9536-0e8a72edd4af · outbound

This paper cites Reinforcing Language Agents via Policy Optimization with Action Decomposition.

MARFT: Multi-Agent Reinforcement Fine-Tuning Reinforcing Language Agents via Policy Optimization with Action Decomposition

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.469238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.469238Z digest=sha256:2f03a00861176dbcef63c9b093ef872cdf8259b253b562302c5814ca340064c8

Observation b9e2ef32-faeb-48ca-ae59-f0f9ddea3c87 · outbound

This paper cites Williams.

MARFT: Multi-Agent Reinforcement Fine-Tuning Williams

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.472645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.472645Z digest=sha256:d286ebdec828bafa0c3bdfbf2f684543e93fc667e82cbfae7f2fe1e3c69bf761

Observation 5694d241-be68-4389-b33b-8079b83ce287 · outbound

This paper cites AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation.

MARFT: Multi-Agent Reinforcement Fine-Tuning AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.475666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.475666Z digest=sha256:bd527db3e78b1fa0af8767f7363496f17f336461611e2b9204d4d883e1d0f9d4

Pith citing papers

Observation 36f56b3d-8acd-4ef2-91d1-c7668399995e · inbound

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems cites this paper.

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 168

Resolution
verified exact
arxiv_id, observed 2026-06-02T02:03:31.668542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T21:39:49.832151Z digest=sha256:e15f83cc5b6ce0b253c72946cd8f3704379bfa5fa90eca6f551e77e96e3ac003

Observation e61494a9-9d23-4fd6-a885-542d46411b31 · inbound

Reasoning Like an Economist: Post-Training on Economic Problems Induces Strategic Generalization in LLMs cites this paper.

Reasoning Like an Economist: Post-Training on Economic Problems Induces Strategic Generalization in LLMs MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:06:25.439085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:06:25.439085Z digest=sha256:928db4f008784a60b9e447c08c3d1a7e586a46b49baa6fc4509d73b195de5391

Observation 11df9874-f606-47d4-b944-f2b6b3ae4a03 · inbound

Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models cites this paper.

Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:58.017209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:41:58.017209Z digest=sha256:dbcdbedd172ea35d6c1b2f14598ef1fa838e0f2cbbadea484b5a434efd07975f

Observation a90916ee-2ddb-465e-bff1-f8c54a8c7fda · inbound

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs cites this paper.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:58.937766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:58.937766Z digest=sha256:2de44b45c2fb0ec35622b1648266dd3964579acb30020c92aa2fd091a9e02103

Observation f15f0f2e-b42b-43d9-8268-dc97d30a64e2 · inbound

A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems cites this paper.

A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-06-02T02:03:31.668542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T23:21:42.029285Z digest=sha256:547ab3b094ce5c5848e0b1a6700490e315049b03439f11ae9b9ca6b187be9024

Observation 7284a6bd-a9f8-495d-a1aa-f45307bd4211 · inbound

SWIRL: A Staged Workflow for Interleaved Reinforcement Learning in Mobile GUI Control cites this paper.

SWIRL: A Staged Workflow for Interleaved Reinforcement Learning in Mobile GUI Control MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:53.497423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:18:53.497423Z digest=sha256:13d9720c5085cf0283e3b22df8e337647d27903cd95748400b72ef60bb164021

Observation 42e39dd7-55d2-400f-aab7-650c5e6df32d · inbound

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing cites this paper.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.103829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.103829Z digest=sha256:321844b48cadb85bb20d9ff4644ebdb64b182ac757b6bb1f17f6a6be5331e3da

Observation a4c744a8-7fc4-4fc7-a50f-d1ac819ee512 · inbound

MASPRM: Multi-Agent System Process Reward Model cites this paper.

MASPRM: Multi-Agent System Process Reward Model MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T07:56:46.663818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:56:46.663818Z digest=sha256:1d049e51af683823fe09509d69716778d5db5ff405860a78fd41be4a6a1943a1

Observation 5b74506c-3120-45dc-b88d-5ae8aa0f055b · inbound

Memory in the Age of AI Agents cites this paper.

Memory in the Age of AI Agents MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-06-02T02:03:31.668542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-11T18:18:19.911342Z digest=sha256:d79c566c22e350945bbc06d0c06bc587274ffbdab2700afd3c8e9dd4d5c7660d

Observation 3c7a5d6a-ba58-4d98-b73b-599239c6c62f · inbound

Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic cites this paper.

Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-02T02:03:31.668542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T09:47:47.051969Z digest=sha256:6f0232c17bbedac6a55a0db44994f6a7384a619cbc23f517e64d9adf4f75c527

Observation f9d5ae1a-8d7a-4685-a6ea-d898d0abbce9 · inbound

Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic cites this paper.

Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:39.487517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:39.487517Z digest=sha256:b269d5d726fc8f14f3d31c14c9a74fa7f830a398cbcbf341d98b3b72a19361fb

Observation d52a709f-a445-4e56-8c54-613947da1d0a · inbound

Learning from Self-Debate: Preparing Reasoning Models for Multi-Agent Debate cites this paper.

Learning from Self-Debate: Preparing Reasoning Models for Multi-Agent Debate MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-02T02:03:31.668542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T14:29:15.751497Z digest=sha256:befe9bb54e82a8a085eb8fc7852502349c0ccd922cf96e40cc31969b1cb13633

Observation d4118c98-6069-45c6-8a3f-0565b18f7ca4 · inbound

Joint Optimization of Multi-agent Memory System cites this paper.

Joint Optimization of Multi-agent Memory System MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-02T02:03:31.668542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T12:12:35.056095Z digest=sha256:d53a110f7a6851f357485b31da8ac9ff88003b9d3cbd0e3714e3c683a95a8e0b

Observation 38277402-9449-402c-98e9-4dd9706248c1 · inbound

Reinforcement Learning for LLM-based Multi-Agent Systems through Orchestration Traces cites this paper.

Reinforcement Learning for LLM-based Multi-Agent Systems through Orchestration Traces MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-02T02:03:31.668542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T18:44:27.685266Z digest=sha256:2e2d1dfa3f5cacae7452058f921952e5e1eef75b5544b7f8129402bc2f2e4115

Observation 5290e0a4-4aef-4a39-967a-aba1123935de · inbound

Tree-based Credit Assignment for Multi-Agent Memory System cites this paper.

Tree-based Credit Assignment for Multi-Agent Memory System MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-02T02:03:31.668542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T15:40:56.527868Z digest=sha256:3420f007aa5f937e5c44a0716902ec8a4abfb77151a87ef93276a681a9d7078a

Observation d95171eb-2196-449b-a94b-089862af93c0 · inbound

AIPO: Learning to Reason from Active Interaction cites this paper.

AIPO: Learning to Reason from Active Interaction MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-02T02:03:31.668542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T01:17:28.124867Z digest=sha256:2a139a06e83e0c02e1734bbea6faff5ecbe602fc04a974e09245d7c4eac62c6a

Observation 3a1659cf-20a2-4c84-ab5b-74896ef270df · inbound

AIPO: Learning to Reason from Active Interaction cites this paper.

AIPO: Learning to Reason from Active Interaction MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-02T02:03:31.668542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T18:07:27.492419Z digest=sha256:905c99fa56ff83878210d4103c5359d64fbb716b053783a1a7dfd90d75e97a88

Observation 5913973e-4df3-422b-9779-3c5c0fabb6be · inbound

Reinforced Collaboration in Multi-Agent Flow Networks cites this paper.

Reinforced Collaboration in Multi-Agent Flow Networks MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-02T02:03:31.668542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-14T20:42:48.438057Z digest=sha256:13e98006c67d6ef6384d2370cd79b228c3d60cdb9d114f0c39a51c59b2e5fd3c

Observation 3f0151e0-bd36-4962-acfb-4667cca22f31 · inbound

Position: Agentic AI System Is a Foreseeable Pathway to AGI cites this paper.

Position: Agentic AI System Is a Foreseeable Pathway to AGI MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-06-02T02:03:31.668542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-14T20:10:36.101426Z digest=sha256:f55fb7942778ff8ad881931abf85559c313b0de9cd40aefceee37788efff0e01

Observation f771318b-14e8-45eb-841f-7b4481dbacac · inbound

Traj-Evolve: A Self-Evolving Multi-Agent System for Patient Trajectory Modeling in Lung Cancer Early Detection cites this paper.

Traj-Evolve: A Self-Evolving Multi-Agent System for Patient Trajectory Modeling in Lung Cancer Early Detection MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 48

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:26:22.822170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T14:20:06.381334Z digest=sha256:188f0669f79dcb79ef489c330e1fe8a109c05d42e87739f1a8246473ecae820f

Observation 112dcb44-05b2-4c69-8901-48c21d5a6403 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 113

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:09:40.650675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:2981b73032204808daf653c7fea4eb41cb2ceb08e57148de8f2b89a5a8f8ded6

Observation fec67690-d95e-4a18-9243-724f69b41a88 · inbound

Where Do CoT Training Gains Land in LLM based Agents? cites this paper.

Where Do CoT Training Gains Land in LLM based Agents? MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-04T13:49:51.544355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-26T04:55:14.293452Z digest=sha256:2bdaec8d1d2816a322e58cdb7c1069059d18b57f3a41fcf265767ca007b580af

Observation 92e17d1a-d46b-402e-90f5-755126482633 · inbound

MoRSE: Task-Oriented Multi-Agent System with Mixture of Role-Subtask Experts cites this paper.

MoRSE: Task-Oriented Multi-Agent System with Mixture of Role-Subtask Experts MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T14:31:57.712831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:31:57.712831Z digest=sha256:3cd47fc3a082a3021baee4a2417dd23517041ba1ae4f980905664943fb1c616e

Observation 8452a54b-71b8-4916-96a9-74b4ff346e3e · inbound

ExRole: From Team Trajectories to Executable Roles in Multi-Agent Language Models cites this paper.

ExRole: From Team Trajectories to Executable Roles in Multi-Agent Language Models MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-16T00:28:56.795142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:28:56.795142Z digest=sha256:c1ea41f052e72df76b594fa73d223db4148d9d0fc3f249cb648b8717f309d34e