Pith. sign in

Paper Citation Record · LEDGER

Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise

As of 21 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 1 inbound Pith citation observation for arXiv:2507.18867.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.18867 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:11:51.843712Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T22:11:35.277901Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact1
  • verified fuzzy26
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 86df1eb7-31ff-499f-8a05-7582d87f2d42 · outbound

This paper cites An overview of recent progress in the study of distributed multi-agent coordination,.

Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise An overview of recent progress in the study of distributed multi-agent coordination,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:11:52.346192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:11:51.721038Z digest=sha256:e20b6c9066c019be7c78110ebcf0658f61641bd9381cbc4196f0497a40d98645

Observation 7aae10c2-ec9c-40f6-a210-59ccbd00f5a4 · outbound

This paper cites Coordinated multi-agent reinforcement learning in networked distributed pomdps,.

Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise Coordinated multi-agent reinforcement learning in networked distributed pomdps,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:11:52.332714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:11:51.725330Z digest=sha256:ef0cb1e29950acbd0c557d543106f27530e4e457869bd484b4f89194bedecc41

Observation 928bbe21-dea1-4800-92a2-d5c09fa0eed6 · outbound

This paper cites Interpretation of neural networks is fragile,.

Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise Interpretation of neural networks is fragile,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:11:52.319777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:11:51.729192Z digest=sha256:18e098ba580d399c8f1754a22757d3e3c377769f884eeee5f289d48ec8ed6a82

Observation f02dc4b7-aa54-4ba2-83bc-20b554f92d04 · outbound

This paper cites Guided Deep Reinforcement Learning for Swarm Systems.

Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise Guided Deep Reinforcement Learning for Swarm Systems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T18:11:51.732896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:11:51.732896Z digest=sha256:fe5993af27816c1fa6523090a956582fbcd7dbf4e1f6b5935a4fe380ee555e50

Observation 0ae5c23c-fd1b-49c2-8e86-f864dec45f2a · outbound

This paper cites Q-value path decomposition for deep multiagent reinforce- ment learning,.

Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise Q-value path decomposition for deep multiagent reinforce- ment learning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:11:52.291297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:11:51.737474Z digest=sha256:44723a43471b0690a94818bbac9fd82fb165aececc134e1051d7c0910e0eacb1

Observation f6ec4a4e-66db-4745-9627-d47e0086627b · outbound

This paper cites RODE: Learning roles to decompose multi-agent tasks,.

Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise RODE: Learning roles to decompose multi-agent tasks,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:11:52.275210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:11:51.741371Z digest=sha256:576b05da56ce064538542c9eb33c507e812b0feec63251eb348b921ac4e8d9df

Observation dc522ff2-f8d9-4f2d-8920-66be7dec5564 · outbound

This paper cites ROMA: Multi-agent reinforcement learning with emergent roles,.

Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise ROMA: Multi-agent reinforcement learning with emergent roles,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:11:52.251242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:11:51.745927Z digest=sha256:a0978cb2464339384bf5af5ea142fed6c7cfad72ab80fcfdcb4cb7c296018d31

Observation 17163ee2-9a0a-4b59-8fa9-6e68f2bfd633 · outbound

This paper cites Value-decomposition networks for cooperative multi-agent learning based on team reward,.

Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise Value-decomposition networks for cooperative multi-agent learning based on team reward,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:11:52.238985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:11:51.749662Z digest=sha256:68e11958e688a890ce232c3119db322d9f5ad87d1d264d3c9afdf9f7d8b7ea87

Observation 72450c64-c9f6-4d6b-a816-b14b5614cc75 · outbound

This paper cites QMIX: Monotonic value function factorisation for deep multi-agent reinforcement learning,.

Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise QMIX: Monotonic value function factorisation for deep multi-agent reinforcement learning,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:11:52.225317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:11:51.753469Z digest=sha256:5d29522ee1bf08dbfb45b0e05c3bfec28da480fb993f26ee07ca373010766b5a

Observation 7644ba24-bc25-4d19-b215-65910de58041 · outbound

This paper cites QPLEX: Duplex dueling multi-agent Q-learning,.

Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise QPLEX: Duplex dueling multi-agent Q-learning,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:11:52.211438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:11:51.757145Z digest=sha256:e733c51b8d70a43b7ccee36abfdfc9da8860bb18b46188295b332d4e31187498

Observation dd050ab3-74cd-4686-a4dd-6b607000b8a6 · outbound

This paper cites Mixrts: Toward interpretable multi-agent reinforcement learning via mixing recurrent soft decision trees,.

Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise Mixrts: Toward interpretable multi-agent reinforcement learning via mixing recurrent soft decision trees,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T18:11:51.760852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:11:51.760852Z digest=sha256:5616dbb08ea158c8625b83f92d099313c61bdb0f9c18d359faea36b7b945fc51

Observation 7a55ea35-d8c7-4c25-b52b-79801420ac5c · outbound

This paper cites NA 2Q: Neural attention additive model for interpretable multi-agent q-learning,.

Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise NA 2Q: Neural attention additive model for interpretable multi-agent q-learning,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:11:52.190260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:11:51.765266Z digest=sha256:b4f3133eab3160160dacf39b18e8361c466b2437840b35e518fa745efc9c5e7e

Observation f4847e28-195d-45e3-b790-be2c8c1e4fec · outbound

This paper cites A comprehensive survey of multiagent reinforcement learning,.

Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise A comprehensive survey of multiagent reinforcement learning,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:11:52.176156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:11:51.769428Z digest=sha256:ec2d649248be0ba167200989b3e76bcddbbcfe4ed782926af63bea00c0e4c92b

Observation cd4549f7-d7b7-467e-847c-81185da4cabc · outbound

This paper cites Dynamic agent-based reward shaping for multi-agent systems,.

Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise Dynamic agent-based reward shaping for multi-agent systems,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:11:52.157085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:11:51.773617Z digest=sha256:924f8cc07fbb82c5616be31444d7c72da7b7b8c4c50cd22f10bb1092594726ea

Observation 135661cd-6a0e-4ecf-9350-bc2dede56e62 · outbound

This paper cites Deep Multiagent Reinforcement Learning: Challenges and Directions.

Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise Deep Multiagent Reinforcement Learning: Challenges and Directions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T18:11:51.777431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:11:51.777431Z digest=sha256:6c59cd7e709bd8c509db5a084bd49ad586ea2f8d930c106fe28e02db00090428

Observation c2afbab4-841b-45e2-a6c2-6f345f056307 · outbound

This paper cites Cooperative exploration for multi-agent deep reinforcement learning,.

Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise Cooperative exploration for multi-agent deep reinforcement learning,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:11:52.139282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:11:51.781506Z digest=sha256:61b03a82a5ee3c53f0ab464f80d54368c0d7b21bd03cf11b08707c702c765f42

Observation df13e225-1866-4ce0-8d4a-c80628d35866 · outbound

This paper cites Maven: Multi-agent variational exploration,.

Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise Maven: Multi-agent variational exploration,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:11:52.124487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:11:51.785349Z digest=sha256:323084e1b4750c1db4294e46a07f3388ce1aae12de4a39ec762ab38c07c47864

Observation cb4634f4-e8f4-4aa7-9e1b-188afe0bbfb8 · outbound

This paper cites Liir: Learning individual intrinsic reward in multi-agent reinforcement learning,.

Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise Liir: Learning individual intrinsic reward in multi-agent reinforcement learning,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:11:52.112144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:11:51.788773Z digest=sha256:87b248f47f2392e515423cb49675c13448976126569699c6823fa1345a9dbf17

Observation e152f0ad-e6b0-4619-be9d-de4d721080e5 · outbound

This paper cites MASER: Multi-agent reinforcement learning with subgoals generated from experience replay buffer,.

Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise MASER: Multi-agent reinforcement learning with subgoals generated from experience replay buffer,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:11:52.099099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:11:51.792385Z digest=sha256:b9690eee0774b466839798b1cea51b9d0a344966adc2438dd11f6473b8d2eb7a

Observation 4c71277c-3f95-4217-ad70-7494054548b5 · outbound

This paper cites Individual reward assisted multi-agent reinforcement learning,.

Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise Individual reward assisted multi-agent reinforcement learning,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:11:52.086940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:11:51.796299Z digest=sha256:91150b108918b5da3b0f2147b716f64593917c0eb7e10e50351778877fc56b46

Observation 0bd24b21-7f41-4da1-84a3-ca66dac89248 · outbound

This paper cites Haven: hierarchical cooperative multi-agent reinforcement learning with dual coordination mechanism,.

Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise Haven: hierarchical cooperative multi-agent reinforcement learning with dual coordination mechanism,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:11:52.074201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:11:51.799555Z digest=sha256:f9685bc22077b0c7d81734d130e9f4d259765d82ee52ad1888248cc93deb1791

Observation 6b51b7fa-82ed-4b49-a18c-f78853a97a1d · outbound

This paper cites Hierarchical Reinforcement Learning in StarCraft II with Human Expertise in Subgoals Selection.

Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise Hierarchical Reinforcement Learning in StarCraft II with Human Expertise in Subgoals Selection

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-15T18:11:51.913702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:11:51.803074Z digest=sha256:b3d97a083347976df68eb13650380520282d1de23bae9c85417d119d4fd7c17d

Observation b0495970-2ab0-467c-9552-02fae64ee742 · outbound

This paper cites Dl2: training and querying neural networks with logic,.

Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise Dl2: training and querying neural networks with logic,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:11:52.060717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:11:51.806607Z digest=sha256:a1f9c6f88178fc1de08dbc309a214253555e93221408f8cfb9f7a3f3a4debddf

Observation 47143ef4-d7de-4bbb-96df-ecc229f0b9ad · outbound

This paper cites Rule-based reinforce- ment learning for efficient robot navigation with space reduction,.

Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise Rule-based reinforce- ment learning for efficient robot navigation with space reduction,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:11:52.047291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:11:51.811090Z digest=sha256:118eed3154fa455953bac27512a66356b8b8ba1acf1a5048d9d1dd6f2354abe4

Observation ff1811f7-7cea-4374-9ce0-e181e1928c31 · outbound

This paper cites Extracting decision tree from trained deep reinforcement learning in traffic signal control,.

Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise Extracting decision tree from trained deep reinforcement learning in traffic signal control,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:11:52.030720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:11:51.814905Z digest=sha256:13f9a21812dad0ff78bf49864c0f17c794a86502686050d12c576a3b9be69e24

Observation 3139a28d-2881-4a15-adec-4878867ed2ae · outbound

This paper cites KoGuN: Accelerating Deep Reinforcement Learning via Integrating Human Suboptimal Knowledge.

Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise KoGuN: Accelerating Deep Reinforcement Learning via Integrating Human Suboptimal Knowledge

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T18:11:51.819017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:11:51.819017Z digest=sha256:66d7c8c063cb5f7fa9d583dd7f8927536f6857e6e9f43dfe6a8e6a1903da7f67

Observation 81a19c4a-408d-451b-b086-272ac28135e1 · outbound

This paper cites QTRAN: Learning to factorize with transformation for cooperative multi-agent reinforcement learning,.

Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise QTRAN: Learning to factorize with transformation for cooperative multi-agent reinforcement learning,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:11:52.005740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:11:51.824232Z digest=sha256:a05538810196e1d270a57806c0cc6ed00ec96e05763306d3f1fbd7deca49426b

Observation c6b5f62d-b164-4fc5-949d-95b81035ffcf · outbound

This paper cites Exploration with unreliable intrinsic reward in multi-agent reinforcement learning,.

Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise Exploration with unreliable intrinsic reward in multi-agent reinforcement learning,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:11:51.990828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:11:51.828539Z digest=sha256:de938a4af0c96f5f0bebc3baee98caa1b0f978a2e809c783cd9bb5c05307d6b1

Observation 4e73d029-9e32-4d44-a71f-1a45254f122b · outbound

This paper cites Discovering generalizable multi-agent coordination skills from multi-task offline data,.

Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise Discovering generalizable multi-agent coordination skills from multi-task offline data,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:11:51.977204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:11:51.833112Z digest=sha256:ec24ee7fdc211f329eb6b46e450baf3cf567ec3fc768b505e29b0fb18d3dfaa3

Observation b16e8539-3a27-4d54-afe7-78d763cbaf48 · outbound

This paper cites Shared experience actor- critic for multi-agent reinforcement learning,.

Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise Shared experience actor- critic for multi-agent reinforcement learning,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:11:51.963741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:11:51.837940Z digest=sha256:c2a5a0aad723962994edbe3a597f188abb989d1d76febc5f21ee1904e3c66d71

Observation 2f20e365-ddce-4432-a5f0-8a1708048f2e · outbound

This paper cites The StarCraft Multi-Agent Challenge,.

Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise The StarCraft Multi-Agent Challenge,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:11:51.951454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:11:51.843712Z digest=sha256:e226b30d2b7cb78b77ef904b5892fdb58fb4019f65617878d33c6899565ae505

Pith citing papers

Observation 4595a430-4bdd-4eb1-a04b-a8175adb6fd6 · inbound

Robust Instruction Compliance in Cooperative Multi-Agent Reinforcement Learning cites this paper.

Robust Instruction Compliance in Cooperative Multi-Agent Reinforcement Learning Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:15:06.017440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-30T22:11:35.277901Z digest=sha256:9e0b041cec811e19b9af81a2ee717348f68588f81b87242111940ed609f95a89