Pith. sign in

Paper Citation Record · LEDGER

Speaking the Language of Teamwork: LLM-Guided Credit Assignment in Multi-Agent Reinforcement Learning

As of 9 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 2 inbound Pith citation observations for arXiv:2502.03723.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.03723 v2

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T01:03:57.025841Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T07:56:46.797195Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

15 of 15 outbound references displayed

  • verified exact2
  • verified fuzzy2
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 85dd0847-b73e-4d6f-b06c-61836ec1bfe7 · outbound

This paper cites GPT-4 Technical Report.

Speaking the Language of Teamwork: LLM-Guided Credit Assignment in Multi-Agent Reinforcement Learning GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T01:03:56.956730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T01:03:56.956730Z digest=sha256:260aa65797edbe4badfe681c0848252ede2421915a8c5bc38f63d28da428fd61

Observation da73d674-217f-498e-b824-a679cfa9a230 · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

Speaking the Language of Teamwork: LLM-Guided Credit Assignment in Multi-Agent Reinforcement Learning RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T01:03:56.972646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T01:03:56.972646Z digest=sha256:32b85103530c6d687f016a5ce7509c941e94f24ba364ea8c175fb17ba5a341a6

Observation 67a4b0f7-931d-4a38-82ba-3573b732027a · outbound

This paper cites PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training.

Speaking the Language of Teamwork: LLM-Guided Credit Assignment in Multi-Agent Reinforcement Learning PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T01:03:56.977920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T01:03:56.977920Z digest=sha256:ff22ae3e8a4efbd4fa12200698bb7da000ca92294a3a4490b3f48748ea6da904

Observation f4c0a592-b42c-4738-a2c3-0f7e97a1021d · outbound

This paper cites Value-Decomposition Multi-Agent Actor-Critics.

Speaking the Language of Teamwork: LLM-Guided Credit Assignment in Multi-Agent Reinforcement Learning Value-Decomposition Multi-Agent Actor-Critics

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-09T01:03:57.167804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T01:03:56.995039Z digest=sha256:3a52ccd25ecfcd2bbb995898e273a2513d07059ca7b7ff27a1605fa547ecbcca

Observation 07a7d512-4dd4-49e0-a511-8fea0fc40018 · outbound

This paper cites A Minimaximalist Approach to Reinforcement Learning from Human Feedback.

Speaking the Language of Teamwork: LLM-Guided Credit Assignment in Multi-Agent Reinforcement Learning A Minimaximalist Approach to Reinforcement Learning from Human Feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T01:03:57.000001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T01:03:57.000001Z digest=sha256:b018368bb2616e480a3b2de5fdfc5cc57042d7a001f281ad5b9b7a510f5fde0d

Observation a4bd5be9-8ec5-452e-9614-099159e616da · outbound

This paper cites QPLEX: Duplex Dueling Multi-Agent Q-Learning.

Speaking the Language of Teamwork: LLM-Guided Credit Assignment in Multi-Agent Reinforcement Learning QPLEX: Duplex Dueling Multi-Agent Q-Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T01:03:57.010803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T01:03:57.010803Z digest=sha256:ea897fb3fcdf9f91a84a19de1d4738e6ee3428a890cd49374ca63fafd1b77721

Observation f1e0ea21-0d5d-40ff-9d48-64a1ab8391f9 · outbound

This paper cites The Surprising Effectiveness of PPO in Cooperative, Multi-Agent Games.

Speaking the Language of Teamwork: LLM-Guided Credit Assignment in Multi-Agent Reinforcement Learning The Surprising Effectiveness of PPO in Cooperative, Multi-Agent Games

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T01:03:57.021190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T01:03:57.021190Z digest=sha256:506a7f2399ab902c2f7ee2d6fd601ac07c4408301128cbfe494c059589816b8d

Observation 1a12b5d7-adb3-437f-803b-e196361904d1 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Speaking the Language of Teamwork: LLM-Guided Credit Assignment in Multi-Agent Reinforcement Learning Fine-Tuning Language Models from Human Preferences

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T01:03:57.025841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T01:03:57.025841Z digest=sha256:f590faaaf7d89de60eaef2f2643b9e1ab386132aff9f8c4c23c6386509b379b6

Observation c015bae3-3cdc-4418-9e2b-b7d71171aa35 · outbound

This paper cites Safe, Multi-Agent, Reinforcement Learning for Autonomous Driving.

Speaking the Language of Teamwork: LLM-Guided Credit Assignment in Multi-Agent Reinforcement Learning Safe, Multi-Agent, Reinforcement Learning for Autonomous Driving

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-09T01:03:56.990709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T01:03:56.990709Z digest=sha256:69ba002db4cbeb9b5db87011c4b933eea950ac26ff6ade21c543a439fee0daad

Observation 71e1e9c4-e878-4489-a434-6e91cfa7c096 · outbound

This paper cites M., Stepputtis, S., Campbell, J., and Sycara, K.

Speaking the Language of Teamwork: LLM-Guided Credit Assignment in Multi-Agent Reinforcement Learning M., Stepputtis, S., Campbell, J., and Sycara, K

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T01:03:57.276158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T01:03:56.982350Z digest=sha256:f3e262a8676e8e5464f35eefc9b62a45ed9159b8c27e83d3a92494c892c214a1

Observation fd87cffd-a3a8-4b0c-9528-235a4f84bdcb · outbound

This paper cites an unresolved cited work.

Speaking the Language of Teamwork: LLM-Guided Credit Assignment in Multi-Agent Reinforcement Learning Unresolved cited work

Reference 2020

Resolution
unresolved
raw_fallback, observed 2026-08-09T01:03:57.260440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T01:03:56.986632Z digest=sha256:dc815afd147d9fc702bd20ee613398268a168ea7d9fabc09cc5a5decdd8e413d

Observation 081a3b8a-f58e-4007-8964-81aa60c348bd · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Speaking the Language of Teamwork: LLM-Guided Credit Assignment in Multi-Agent Reinforcement Learning LLaMA: Open and Efficient Foundation Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-09T01:03:57.005707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T01:03:57.005707Z digest=sha256:f5e052e866fbaaa82461009a7955129f97fb4da59ff59ba6778e6c29f35c3569

Observation b66be455-9538-49cb-a6d0-4c0848d20552 · outbound

This paper cites Multi-Agent Reinforcement Learning is a Sequence Modeling Problem.

Speaking the Language of Teamwork: LLM-Guided Credit Assignment in Multi-Agent Reinforcement Learning Multi-Agent Reinforcement Learning is a Sequence Modeling Problem

Reference 2022

Resolution
verified exact
local_arxiv, observed 2026-08-09T01:03:57.100578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T01:03:57.016285Z digest=sha256:480a9480dd04eef1d376835d45fc6c21b2823095d94560d3fb0825249973b927

Observation 8ea196f9-74d8-4fd3-b484-b654409ff70e · outbound

This paper cites A variational approach to mutual information-based coordination for multi-agent reinforcement learning.

Speaking the Language of Teamwork: LLM-Guided Credit Assignment in Multi-Agent Reinforcement Learning A variational approach to mutual information-based coordination for multi-agent reinforcement learning

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T01:03:57.290368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T01:03:56.967507Z digest=sha256:34770f3d8e42ce690783aa58e76fb6730526acc3212661b46bd07e128dd7a466

Observation cff30953-901f-495c-928c-8f3660cacebc · outbound

This paper cites Assigning Credit with Partial Reward Decoupling in Multi-Agent Proximal Policy Optimization.

Speaking the Language of Teamwork: LLM-Guided Credit Assignment in Multi-Agent Reinforcement Learning Assigning Credit with Partial Reward Decoupling in Multi-Agent Proximal Policy Optimization

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-09T01:03:56.962583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T01:03:56.962583Z digest=sha256:5e828c407f759197ff815f393517ba19913dc70c50d9dd48815198f14c64c1ab

Pith citing papers

Observation 61540494-045a-46b4-b6e9-baa27c71ad69 · inbound

MASPRM: Multi-Agent System Process Reward Model cites this paper.

MASPRM: Multi-Agent System Process Reward Model Speaking the Language of Teamwork: LLM-Guided Credit Assignment in Multi-Agent Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T07:56:46.797195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:56:46.797195Z digest=sha256:00c0ac68caeaeb15dccc6d44b46ba4a673fd6bb5ae56cde1174459db7e6045a0

Observation be47df70-0bf2-4a7c-ad41-141f96d55027 · inbound

MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation cites this paper.

MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation Speaking the Language of Teamwork: LLM-Guided Credit Assignment in Multi-Agent Reinforcement Learning

Reference 81

Resolution
unresolved
no resolver link, observed 2026-07-31T21:55:17.425466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T21:55:17.425466Z digest=sha256:3f1f675ccc843ab80cccb33eef4865a3c03f8011dab86542446195ae7a20d485