Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:42:31.475666Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 100 of 118 outbound references and 24 inbound Pith citation observations for arXiv:2504.16129.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:42:31.475666Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:28:56.795142Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T13:49:51.542821Z
100 of 118 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9eb51eaa-6401-4b10-a06b-35eb7f9d0be8 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Springer New York, New York, NY, 2008
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation cf42ef0a-fe0e-4fe4-a7c9-a7207bbca069 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning URL https://arxiv.org/abs/2407.21075
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aeb3c174-d91c-44ba-92be-bb5d8092154d · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Introducing the model context protocol
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf72fa2d-d3c1-4bd0-acef-ff85a52afa0f · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning A comprehensive survey of multiagent reinforcement learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45951589-a087-4bcc-a148-cab6eebc559e · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Owl: Optimized workforce learning for general multi-agent assistance in real-world task automation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 312e0411-4f12-4c75-8e17-f5336fce08a5 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning On the utility of learning about humans for human-ai coordination
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07c696e2-ad0c-493e-bfcb-74ab506ae3df · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Grounding large language models in interactive environments with online reinforcement learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 577a50e7-6282-4b9f-91bb-502b5d646618 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Why Do Multi-Agent LLM Systems Fail?
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c542a79-84f9-4874-9031-76dfda89e004 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Communication-efficient actor-critic methods for homogeneous markov games
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66559784-6238-484e-98b1-08360b7ece00 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Octopus: On-device language model for function calling of software APIs
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cf33de4-bcd9-4672-ba52-678857627928 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation cba3c9cb-5505-4d76-b00a-0d29e0b695dd · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Training Verifiers to Solve Math Word Problems
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f9537ee-898b-42ea-b158-e0839a9f2c9f · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning AD-AutoGPT: An Autonomous GPT for Alzheimer's Disease Infodemiology
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 904c9c91-663e-4b44-a6b4-5ac6a6cd9fc7 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Independent policy gradient methods for competitive reinforcement learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5420bff4-53cb-4801-81fa-1c2f4fea28af · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0b33f79-d612-448c-8da5-1ec5681b3c5e · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Pilarski, and Richard S
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation cd43666d-c93b-48f1-94d7-bd83db5f9518 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Interactive debugging and steering of multi-agent ai systems
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2614f40-55cf-45e7-ac1a-b8a168a2a712 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning TinyAgent: Function Calling at the Edge
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7ce1f31-b946-472a-ba34-a710a67a8dc5 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3188bd8-a64f-4335-b449-081648881a23 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Counterfactual multi-agent policy gradients
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fede6fb-7a62-486c-a991-836c9ecbd096 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning ANP - Agent Network Protocol
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70a6b63b-ba86-4339-8e35-3828bf2b45be · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5c1dfb9-e03f-43b0-a3a8-345d3cfcacc2 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Towards an AI co-scientist
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aab2846d-3a9c-49c1-92ae-0f80103dca37 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Variance reduction techniques for gradient estimates in reinforcement learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbf625a3-4efb-4d3b-91af-96cb2efb4149 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Reinforcement learning with deep energy-based policies
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28ad4c4a-f641-4155-a519-27a45a4bc102 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Emergence of Locomotion Behaviours in Rich Environments
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b239c855-55f4-4a2c-b414-e59222095fa9 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Measuring Mathematical Problem Solving With the MATH Dataset
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddd5af33-fe5b-475c-9bd5-b876fa125048 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Meta GPT : Meta programming for a multi-agent collaborative framework
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bce5ffdc-368f-40f9-b5cc-f202f8c6b546 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7e9e30b-fc99-4ce2-89f2-bc2a4dd02fa8 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84b1caae-229c-4e9c-bb34-14ebd3c5eb52 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06230c10-326e-4d0e-b490-1ba0a9f651fa · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Qwen2.5-Coder Technical Report
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f38a103-9f70-47f2-8cbd-a79961ef8dbe · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Deep reinforcement learning for swarm systems
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99477a6a-5bf2-4736-bf47-2283c6002fbe · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning From LLMs to LLM-based Agents for Software Engineering: A Survey of Current, Challenges and Future
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fb57dc9-2e40-4db8-b948-6aabb2fc9d2c · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Joshi, Hanna Moazam, Heather Miller, Matei Zaharia, and Christopher Potts
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c77f10d-abb5-4e00-ab3d-dff731c2990d · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf3dcbf6-bff3-4654-b789-9f1afa0a456e · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Multi-agent reinforcement learning for traffic signal control
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8f90959-4df6-41df-b33a-34e3c08413e7 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Actor-critic algorithms
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 106d41aa-b274-4ea1-a282-d0518ae88c12 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Buy 4 REINFORCE samples, get a baseline for free!, 2019
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e78c747-d96d-44cb-aa7f-74dbe1086b73 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Settling the variance of multi-agent policy gradients
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e195a704-7623-436b-a813-b4ed759a2c00 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Trust region policy optimisation in multi-agent reinforcement learning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3728b33-b2d8-407a-979b-f449c7e58788 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fdfa293-6285-44c5-bc15-baf26a262a11 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Multi-agent actor-critic for mixed cooperative-competitive environments
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dedd5286-1cc7-4070-b1ca-b34d47eb14cd · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1930bc1-f3c3-4bc4-8148-4e877240b7ec · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Coevolving with the other you: Fine-tuning llm with sequential cooperative multi-agent reinforcement learning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cae160b-b902-4473-8189-1e6ab5444370 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Roco: Dialectic multi-robot collaboration with large language models, 2023
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aeb4b120-4dff-4b94-aec4-dbc876e81ca5 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Laurent, and Nadine Le Fort-Piat
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7591dcbf-c723-4311-b81b-5f86c713f77d · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Coordinated multi-robot exploration under communication constraints using decentralized markov decision processes
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5d2593b8-0dd2-44d3-95f6-9c5cc34ba5d6 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning GAIA : a benchmark for general AI assistants
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 437a705b-88c4-41b2-b813-1d2dc5b3214e · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Steps toward artificial intelligence
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e1fb39d-a66a-40f7-8d96-19b58d0e0c04 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Playing Atari with Deep Reinforcement Learning
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 336b26d2-1126-4858-b8e6-8fb3f19bcce3 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Asynchronous methods for deep reinforcement learning
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28b9f397-7cca-4723-a715-cb4a1cef0559 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Ppo improvement in different environments
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ce14e7d8-f984-4da2-bd76-27e96c9df2f7 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Richard Yu
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de7cc274-a03d-4c3d-a610-5716d64fdf0a · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Oliehoek and Christopher Amato
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1addb96-2a43-4f80-9922-8a470753e8aa · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Learning to reason with LLMs , sep 2024 a
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58b2b400-5f50-4f79-8f6e-0f95bc292948 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Openai's reinforcement fine-tuning research program, dec 2024 b
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d675879-d0cf-499d-9c55-3d3a4995d6ba · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Training language models to follow instructions with human feedback
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2292abc-50ce-4013-8620-a735503e9408 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning MAPoRL: Multi-Agent Post-Co-Training for Collaborative Large Language Models with Reinforcement Learning
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12b03d0f-c3b0-43e3-a890-00f70113ab0d · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Gorilla: Large Language Model Connected with Massive APIs
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a9b14ca-6ca2-40be-b437-dbf6e5fa87ca · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Codeforces
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3b191a7-7944-4fef-a23e-8c9e5c33586b · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Virtualhome: Simulating household activities via programs
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1b428c4-94c6-46f2-8986-9b0166c890e2 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning C hat D ev: Communicative agents for software development
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5503f34-32f3-4b32-878f-8ad8f95259d3 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Qwq-32b: Unveiling the power of reinforcement learning in reasoning, March 2025
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 009e447b-a2a8-424d-bac2-12a3096920f0 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Sensitivity of ${^{44}}$Ti and ${^{56}}$Ni production in CCSN shock-driven nucleosynthesis to reaction rates
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8094e94a-62b2-4df2-a4fe-0c12701b3bd3 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Toolformer: language models can teach themselves to use tools
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6697cf3f-da24-42c7-8728-d241bdc41596 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Trust region policy optimization
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c72428e9-ad27-45df-b7b7-a199fa46206c · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Proximal Policy Optimization Algorithms
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 513b33d0-0bca-41c5-a0a6-42780e8273a5 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0320e9f5-14ba-4e54-9054-4f5589b84895 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Safe, Multi-Agent, Reinforcement Learning for Autonomous Driving
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82de1eca-c568-46cc-8a41-02b38eaff3bd · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ddc8129-b31e-40fc-9bc7-a1b0f3e88dfb · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning HybridFlow: A Flexible and Efficient RLHF Framework
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0986bcd4-b67b-4544-ade3-a3857defb731 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Multiagent Systems: Algorithmic, Game-Theoretic, and Logical Foundations
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d0d90d0-8ac4-4e87-90f2-c9bc45b2f409 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning An empirical study on google research football multi-agent scenarios
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e431dec-b1ef-4729-a7dd-b006147b42d0 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Multiagent systems: A survey from a machine learning perspective
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff1b245d-d016-49e8-bd93-b17eb0b4dc47 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Multiagent Finetuning: Self Improvement with Diverse Reasoning Chains
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6643af35-db93-41bc-9c17-4594389a4ba9 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Learning multiagent communication with backpropagation
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01a9a781-19cb-459f-adb8-a8e9a4bdff1a · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Announcing the agent2agent protocol (a2a)
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b004d92-20ac-45ab-b693-ab7a195df80c · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Dyna, an integrated architecture for learning, planning, and reacting
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc8a11ab-5c68-4e6b-8f1f-37ac56d635bc · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Sutton and Andrew G
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 938c959b-3c16-4d2d-a380-cc6c9705db27 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Policy gradient methods for reinforcement learning with function approximation
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52e2fb41-fcae-4ba3-8964-9384554f4d7d · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Temporal credit assignment in reinforcement learning
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e603ba33-1f2f-4cfd-93e9-ae996f1df2e1 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Multi-agent reinforcement learning: independent versus cooperative agents
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7d4a8e78-f5a1-482f-8315-5d44b4939e2e · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning True knowledge comes from practice: Aligning large language models with embodied environments via reinforcement learning
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ab0eec97-f14b-4070-8c17-c9e0a97e3576 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Magis: Llm-based multi-agent framework for github issue resolution
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 996e0e0e-fdaf-4b55-ac50-b731aac57806 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Mujoco: A physics engine for model-based control
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a04697a-d7e3-4f31-b89a-530667564d16 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Multiagent learning: Basics, challenges, and prospects
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 51d8aef0-f0b1-40aa-a26b-843ed84c2e9e · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Reinforcement Learning and Markov Decision Processes, pp.\ 3--42
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b761b9d-ba48-45d4-9211-6ce3200d4b98 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Networked Agents in the Dark: Team Value Learning under Partial Observability
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ae4875e3-81d5-4413-8ce1-2313e219f597 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed4908c4-f17e-415f-b2cf-4e67c6c70550 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 267ecab9-1e6f-4f33-80e3-5e1436897793 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Order matters: Agent-by-agent policy optimization
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 967ed46c-f685-42df-a9c7-72582359b6db · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Unresolved cited work
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98767025-9dac-4a2a-9d2b-1fc78eb5dca6 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning CMATH: Can Your Language Model Pass Chinese Elementary School Math Test?
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3f2efb9-0b21-4043-a730-ccd206397fcc · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Multiagent systems: a modern approach to distributed artificial intelligence
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 766d4c40-d984-4175-8ce8-45bdddc97a5e · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Multi-agent reinforcement learning is a sequence modeling problem
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 896bea58-9d53-4cca-a2d7-d69207883b66 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Entropy-Regularized Token-Level Policy Optimization for Language Agent Reinforcement
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 769bf109-d734-4115-9536-0e8a72edd4af · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning Reinforcing Language Agents via Policy Optimization with Action Decomposition
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9e2ef32-faeb-48ca-ae59-f0f9ddea3c87 · outbound
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5694d241-be68-4389-b33b-8079b83ce287 · outbound
MARFT: Multi-Agent Reinforcement Fine-Tuning AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36f56b3d-8acd-4ef2-91d1-c7668399995e · inbound
Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems MARFT: Multi-Agent Reinforcement Fine-Tuning
Reference 168
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e61494a9-9d23-4fd6-a885-542d46411b31 · inbound
Reasoning Like an Economist: Post-Training on Economic Problems Induces Strategic Generalization in LLMs MARFT: Multi-Agent Reinforcement Fine-Tuning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11df9874-f606-47d4-b944-f2b6b3ae4a03 · inbound
Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models MARFT: Multi-Agent Reinforcement Fine-Tuning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a90916ee-2ddb-465e-bff1-f8c54a8c7fda · inbound
How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs MARFT: Multi-Agent Reinforcement Fine-Tuning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f15f0f2e-b42b-43d9-8268-dc97d30a64e2 · inbound
A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems MARFT: Multi-Agent Reinforcement Fine-Tuning
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7284a6bd-a9f8-495d-a1aa-f45307bd4211 · inbound
SWIRL: A Staged Workflow for Interleaved Reinforcement Learning in Mobile GUI Control MARFT: Multi-Agent Reinforcement Fine-Tuning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42e39dd7-55d2-400f-aab7-650c5e6df32d · inbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing MARFT: Multi-Agent Reinforcement Fine-Tuning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4c744a8-7fc4-4fc7-a50f-d1ac819ee512 · inbound
MASPRM: Multi-Agent System Process Reward Model MARFT: Multi-Agent Reinforcement Fine-Tuning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b74506c-3120-45dc-b88d-5ae8aa0f055b · inbound
Memory in the Age of AI Agents MARFT: Multi-Agent Reinforcement Fine-Tuning
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3c7a5d6a-ba58-4d98-b73b-599239c6c62f · inbound
Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic MARFT: Multi-Agent Reinforcement Fine-Tuning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f9d5ae1a-8d7a-4685-a6ea-d898d0abbce9 · inbound
Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic MARFT: Multi-Agent Reinforcement Fine-Tuning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d52a709f-a445-4e56-8c54-613947da1d0a · inbound
Learning from Self-Debate: Preparing Reasoning Models for Multi-Agent Debate MARFT: Multi-Agent Reinforcement Fine-Tuning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d4118c98-6069-45c6-8a3f-0565b18f7ca4 · inbound
Joint Optimization of Multi-agent Memory System MARFT: Multi-Agent Reinforcement Fine-Tuning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 38277402-9449-402c-98e9-4dd9706248c1 · inbound
Reinforcement Learning for LLM-based Multi-Agent Systems through Orchestration Traces MARFT: Multi-Agent Reinforcement Fine-Tuning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5290e0a4-4aef-4a39-967a-aba1123935de · inbound
Tree-based Credit Assignment for Multi-Agent Memory System MARFT: Multi-Agent Reinforcement Fine-Tuning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d95171eb-2196-449b-a94b-089862af93c0 · inbound
AIPO: Learning to Reason from Active Interaction MARFT: Multi-Agent Reinforcement Fine-Tuning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3a1659cf-20a2-4c84-ab5b-74896ef270df · inbound
AIPO: Learning to Reason from Active Interaction MARFT: Multi-Agent Reinforcement Fine-Tuning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5913973e-4df3-422b-9779-3c5c0fabb6be · inbound
Reinforced Collaboration in Multi-Agent Flow Networks MARFT: Multi-Agent Reinforcement Fine-Tuning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3f0151e0-bd36-4962-acfb-4667cca22f31 · inbound
Position: Agentic AI System Is a Foreseeable Pathway to AGI MARFT: Multi-Agent Reinforcement Fine-Tuning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f771318b-14e8-45eb-841f-7b4481dbacac · inbound
Traj-Evolve: A Self-Evolving Multi-Agent System for Patient Trajectory Modeling in Lung Cancer Early Detection MARFT: Multi-Agent Reinforcement Fine-Tuning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 112dcb44-05b2-4c69-8901-48c21d5a6403 · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning MARFT: Multi-Agent Reinforcement Fine-Tuning
Reference 113
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fec67690-d95e-4a18-9243-724f69b41a88 · inbound
Where Do CoT Training Gains Land in LLM based Agents? MARFT: Multi-Agent Reinforcement Fine-Tuning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 92e17d1a-d46b-402e-90f5-755126482633 · inbound
MoRSE: Task-Oriented Multi-Agent System with Mixture of Role-Subtask Experts MARFT: Multi-Agent Reinforcement Fine-Tuning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8452a54b-71b8-4916-96a9-74b4ff346e3e · inbound
ExRole: From Team Trajectories to Executable Roles in Multi-Agent Language Models MARFT: Multi-Agent Reinforcement Fine-Tuning
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.