Pith. sign in

Paper Citation Record · LEDGER

Multi-Agent Reinforcement Learning via Agent-Specific Preference

As of 23 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2608.08604.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08604 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:36:17.857891Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact2
  • verified fuzzy34
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3af094d7-4ce1-4aea-844c-a39dc5b985b4 · outbound

This paper cites Distributed multiagent coordinated learning for autonomous driving in highways based on dynamic coordination graphs,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Distributed multiagent coordinated learning for autonomous driving in highways based on dynamic coordination graphs,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:36:21.271161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:16.550285Z digest=sha256:c06700649de69032f3ee05b53849ee3d42e2947ecd7c60a0d4d2d63eaa54d4ec

Observation 8cd10382-4c85-45db-8a97-b414a458fa7d · outbound

This paper cites Beyond static populations: Efficient delay-constrained scheduling for dynamic users via deep reinforcement learning,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Beyond static populations: Efficient delay-constrained scheduling for dynamic users via deep reinforcement learning,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:36:21.162074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:16.573209Z digest=sha256:2383d126fdacf5ac8219a95e73820fe068404ce6418b135bdfab8aa4cf36a070

Observation 84b729e2-6981-4689-8fcc-f606dcad4ab1 · outbound

This paper cites Dota 2 with Large Scale Deep Reinforcement Learning.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Dota 2 with Large Scale Deep Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-14T04:36:16.608292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:36:16.608292Z digest=sha256:357f84e0cdce71872fdd5cca461c2543bc9dad921fbe7032845ee770be9f1742

Observation 672e6592-83e5-4223-8171-43728f0721d6 · outbound

This paper cites The surprising effectiveness of ppo in cooperative multi-agent games,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference The surprising effectiveness of ppo in cooperative multi-agent games,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T04:36:16.649417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:36:16.649417Z digest=sha256:bbcd8748e6761f7d0ee440893c4512347c7374df5e2ce84f2ce6fa7635863485

Observation d91d45e6-93c3-4e59-b7f8-1be87f64aef5 · outbound

This paper cites Value-decomposition networks for cooperative multi-agent learning based on team reward,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Value-decomposition networks for cooperative multi-agent learning based on team reward,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:36:21.014746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:16.688288Z digest=sha256:3a1c068b5ba43ebf39bf4bcd4a30a301e268de9452dbf3820470793b1b7ab521

Observation 19f60790-e190-4f43-9d4b-3a7c9c4c957b · outbound

This paper cites Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:36:20.978488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:16.714747Z digest=sha256:247d0dffc19ecf2a90e64b9c909a02201baf087e6b2ef3bc4a4e705ed4afafba

Observation 2519376a-93f8-4e40-8f62-148aa8849795 · outbound

This paper cites Beyond shallow behavior: Task-efficient value-based multi-task offline marl via skill discovery,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Beyond shallow behavior: Task-efficient value-based multi-task offline marl via skill discovery,

Reference 7

Resolution
verified exact
raw_fallback, observed 2026-08-14T04:36:18.350792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:16.753954Z digest=sha256:5c263a629e08effb179f06deef2e0f4b41ff7e8f927df3d7c9575052b1407746

Observation c71f5dc0-2a61-4241-a0a7-cb1c02f58d6c · outbound

This paper cites Reward learning from human preferences and demonstrations in atari,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Reward learning from human preferences and demonstrations in atari,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T04:36:16.770050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:36:16.770050Z digest=sha256:c8f5f108272a6618b02f10f20e58c66040f5ef300d3e0d0aac83ee00b1ada631

Observation 286c746c-db69-420f-8113-3c9c16227677 · outbound

This paper cites Pebble: Feedback-efficient interactive reinforcement learning via relabeling experience and unsuper- vised pre-training,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Pebble: Feedback-efficient interactive reinforcement learning via relabeling experience and unsuper- vised pre-training,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T04:36:16.808147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:36:16.808147Z digest=sha256:33195f893171aa5d063d71ef8e39f76ff2b6b120d05d3ef43bbfb2a076ff2d64

Observation 89ac83be-13c5-4498-903b-d5f98857b25a · outbound

This paper cites S-EPOA: Overcoming the Indistinguishability of Segments with Skill-Driven Preference-Based Reinforcement Learning,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference S-EPOA: Overcoming the Indistinguishability of Segments with Skill-Driven Preference-Based Reinforcement Learning,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:36:20.840813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:16.825120Z digest=sha256:9153b069b8bce14f196ac4fedf2f2d1f71155dc9d78173670beb3db22c6d7f1d

Observation 7f329cea-5943-40af-b280-b5c3902bb3e5 · outbound

This paper cites CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambigu- ous Queries,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambigu- ous Queries,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:36:20.754749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:16.840199Z digest=sha256:c179786e6793f6b4f1eb7be07e30d3c114e7b7ede76ffebbf0ac790ebdc71054

Observation 92602b96-256a-4d30-8a97-7bf1908acc8c · outbound

This paper cites STAIR: Addressing stage misalignment through temporal-aligned preference reinforcement learning,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference STAIR: Addressing stage misalignment through temporal-aligned preference reinforcement learning,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:36:20.644829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:16.845103Z digest=sha256:af7533ee38b9d5ec4dae8af627b8376ab9fde5d56dce802605a4fab8d88c058b

Observation a6b65122-d2cf-410d-843f-4ce4352deed5 · outbound

This paper cites MAGPIE: Utilizing Agent-Specific Preferences for Multi-Agent Reinforcement Learning,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference MAGPIE: Utilizing Agent-Specific Preferences for Multi-Agent Reinforcement Learning,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:36:20.567668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:16.860741Z digest=sha256:0a0b1cef9c278d5d09b4a90fd6fe5b980f9e671d2b934c856f445ae9798bfe1b

Observation 78293d41-b356-4e8c-aff9-f1a302f60518 · outbound

This paper cites Senior: Efficient query selection and preference-guided exploration in preference-based rein- forcement learning,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Senior: Efficient query selection and preference-guided exploration in preference-based rein- forcement learning,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:36:20.510295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:16.875042Z digest=sha256:64ae8fe01bf904f9ac61e131fc5142e593cb3ff8fd40e09d08de2912203695c8

Observation 8fcfbad5-977a-4eab-8d80-074a0bdbe692 · outbound

This paper cites Training language models to follow instructions with human feedback,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Training language models to follow instructions with human feedback,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:36:20.461605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:16.894848Z digest=sha256:41dba5138d35d4c081c3d9eef0d267f9c53ce575ef62ac963f0b3d50c53b8bad

Observation 1baf82a0-27b7-419c-b71e-6bbbb7ac8342 · outbound

This paper cites Preference-based Multi-Objective Reinforcement Learning,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Preference-based Multi-Objective Reinforcement Learning,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:36:20.445367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:16.911743Z digest=sha256:d3524dd73b42b4e0f66d39c04925dbd129e016555cc99b5ea34074dfc0895c79

Observation 32adde6a-ccb5-49e5-b8cb-e57f8486e132 · outbound

This paper cites COLLIE: Guiding Skill Discovery in Semantically Coherent Latent Space,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference COLLIE: Guiding Skill Discovery in Semantically Coherent Latent Space,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:36:20.415619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:16.928871Z digest=sha256:f733fe588a92f88d186235dbbfbceba4c3b6daee11ddfe22368a457347b78e5e

Observation 50e544e2-8d1b-488a-adec-069260b18d2b · outbound

This paper cites Offline multi-agent preference-based reinforcement learning with agent-aware direct preference optimization,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Offline multi-agent preference-based reinforcement learning with agent-aware direct preference optimization,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:36:20.356116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:16.955588Z digest=sha256:85e7ff5b9ca41dc71368e63c6fb48e930fad3a1b69d21e0765a03c0e2d601fa3

Observation aa0a22fc-1975-4cdf-a19f-d1ac2c8c3a10 · outbound

This paper cites O-MAPL: Offline multi- agent preference learning,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference O-MAPL: Offline multi- agent preference learning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:36:20.309771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:16.975795Z digest=sha256:d786277ca3308e509aafeeb43a50dc37a99081a0f8e7dfac11d67776a91a1770

Observation 4c097685-29d3-4959-9f1a-dd45e116419b · outbound

This paper cites Decoding global preferences: Tem- poral and cooperative dependency modeling in multi-agent preference- based reinforcement learning,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Decoding global preferences: Tem- poral and cooperative dependency modeling in multi-agent preference- based reinforcement learning,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:36:20.224815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:16.986437Z digest=sha256:cfce1d416feeaa119076225c274d5d7dfbc3550b75acb406664515400f776591

Observation c12da3cf-cc42-460e-b52c-039311c79143 · outbound

This paper cites DPM: Dual preferences-based multi-agent reinforcement learning,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference DPM: Dual preferences-based multi-agent reinforcement learning,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:36:20.103435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:16.993601Z digest=sha256:2657dfaf08ccdd1c033c894b21b9958ab192e78ed882bd7479470472ea338cc2

Observation cd7c8ac9-fb25-4c05-b6a4-5ed466ed21be · outbound

This paper cites Multi-agent reinforcement learning from human feedback: Data coverage and algorithmic techniques,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Multi-agent reinforcement learning from human feedback: Data coverage and algorithmic techniques,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:36:20.066353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:17.021089Z digest=sha256:7a771d2f1bdcdb8a465fe4b228a0168b250169fb6e255f1c17ffe4aab1aef95e

Observation 9a875b21-3ef4-4c71-90d5-1a40993bf5db · outbound

This paper cites Mastering the game of go without human knowledge,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Mastering the game of go without human knowledge,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:36:20.049859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:17.041463Z digest=sha256:d28efa61eb5e2b1d854f97126d830f306d6256a07437452e55dce800458aa71a

Observation 69448ec0-32ba-477d-8031-c9ceaa81f02d · outbound

This paper cites Deep reinforcement learning for the control of robotic manipulation: a focussed mini-review,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Deep reinforcement learning for the control of robotic manipulation: a focussed mini-review,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-14T04:36:17.074751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:36:17.074751Z digest=sha256:bddcf691f7bed17bd038fa63b607af2eea6e30b4d1432ba89f61e7c2540e2eb5

Observation e8ea2054-58d0-40cf-999c-0fd7eaea0ac8 · outbound

This paper cites Integrating Mechanism and Data: Reinforcement Learning Based on Multi-fidelity Model for Data Center Cooling Control,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Integrating Mechanism and Data: Reinforcement Learning Based on Multi-fidelity Model for Data Center Cooling Control,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:36:19.986627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:17.114746Z digest=sha256:4c04e9a4621bd21e1a76e8d34cb8a9be144ce0f121c46e5105ad267382aef0a1

Observation 63e9da64-a45a-4f14-b412-a2ec9c8a03e2 · outbound

This paper cites Large-scale Data Center Cooling Control via Sample-efficient Reinforcement Learning,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Large-scale Data Center Cooling Control via Sample-efficient Reinforcement Learning,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:36:19.825850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:17.130324Z digest=sha256:df69367bf7a02a69488516b1932ea7b5c7227fde29a809dbd34b2ebc347132b3

Observation 455e0aea-83a9-4856-a961-bda2efcd8ea3 · outbound

This paper cites E-mapp: Efficient Multi- Agent Reinforcement Learning with Parallel Program Guidance,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference E-mapp: Efficient Multi- Agent Reinforcement Learning with Parallel Program Guidance,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:36:19.716791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:17.157426Z digest=sha256:4bd5e8caecfff7d344a196064a7a8b32f87358df8b2a3287aa003e92f63ddb9c

Observation 5e5d5239-89e1-4d51-838d-79d943d84d02 · outbound

This paper cites From solo to symphony: Orchestrating multi-agent collaboration with single-agent demos,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference From solo to symphony: Orchestrating multi-agent collaboration with single-agent demos,

Reference 28

Resolution
verified exact
raw_fallback, observed 2026-08-14T04:36:18.265953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:17.179907Z digest=sha256:cbb905009352b0142eb83ccdf9cab364f2cc75e3970c901716c985ab01bae5eb

Observation a48d4323-7fc4-4eb6-ab98-175952a2b5f7 · outbound

This paper cites GlobeDiff: State Diffusion Process for Partial Ob- servability in Multi-Agent Systems,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference GlobeDiff: State Diffusion Process for Partial Ob- servability in Multi-Agent Systems,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-14T04:36:17.215298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:36:17.215298Z digest=sha256:50473304dc7bad7e8526acc5790a0b855fc63c675f644b125c4b0a48aa75ce44

Observation b6dbcd36-759c-4bc3-a562-1acb4a884737 · outbound

This paper cites Multi-agent reinforcement learning for resources allocation optimization: a survey,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Multi-agent reinforcement learning for resources allocation optimization: a survey,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:36:19.660553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:17.231443Z digest=sha256:595f5c8f3b1b9bb3aab0cf57c76463367f647211b8def4686b9b169d204e2ecb

Observation b9601f21-9653-406d-b5ac-2eed6a9c5d6e · outbound

This paper cites A review of cooperative multiagent deep reinforcement learning,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference A review of cooperative multiagent deep reinforcement learning,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:36:19.581804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:17.250163Z digest=sha256:8a19847b0c23a03d09e14cd213525d4a40a05d95c258497fef394c24b5d90745

Observation 73380aa3-8e5a-4436-a6e5-465071410ecd · outbound

This paper cites Reinforcement learning with sparse rewards using guidance from offline demonstration,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Reinforcement learning with sparse rewards using guidance from offline demonstration,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:36:19.543357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:17.274746Z digest=sha256:c6428da130d3d89915af7a70195e8a233806c39d2499c4a31cfdc189fbe38075

Observation 87a477fc-ae7c-42c3-9c30-63f5fe8fe6cc · outbound

This paper cites Reward function design in reinforcement learning,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Reward function design in reinforcement learning,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-14T04:36:17.294934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:36:17.294934Z digest=sha256:4715946fec69052bcf5eac0d8dbe0fb4b6ce62d61b048674925f627fcb692032

Observation 9e6e6999-c57b-4404-bbb5-07e52cf796d0 · outbound

This paper cites Liir: Learning indi- vidual intrinsic reward in multi-agent reinforcement learning,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Liir: Learning indi- vidual intrinsic reward in multi-agent reinforcement learning,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:36:19.444850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:17.309179Z digest=sha256:d2c182bf7a38f97c505a206de85fc97d8f42872f2925e11030d5f2cc28970bd0

Observation 9b8c8573-55ce-4d36-89f8-84e7c55fcf1d · outbound

This paper cites Quality assessment of 3d human animation: Subjective and objective evaluation,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Quality assessment of 3d human animation: Subjective and objective evaluation,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:36:19.382619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:17.345848Z digest=sha256:9555af9cd9cd7a67bbce3ac6da8ac063aa3039de564ac79fbd25551536f1ec5c

Observation 8f839c82-5b23-4326-b4d7-017f20882504 · outbound

This paper cites Deep reinforcement learning from human preferences,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Deep reinforcement learning from human preferences,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-14T04:36:17.353907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:36:17.353907Z digest=sha256:cc3344709fe40597cc100ffadc8e13d41d8f629a7f9cfd6c06e7eecab4c5f7e5

Observation 0cdd9fec-c79b-4520-832c-01fec04eaadf · outbound

This paper cites Preference-Based Multi-Objective Reinforcement Learning with Explicit Reward Modeling,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Preference-Based Multi-Objective Reinforcement Learning with Explicit Reward Modeling,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:36:19.287309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:17.361993Z digest=sha256:0e91484ebb1feaed5086cca118eea370e5c635b7b6afa38d7395cc10dfa12521

Observation dbb60e09-a73f-4594-9b6d-4749e80a317b · outbound

This paper cites Human Implicit Preference-Based Policy Fine-tuning for Multi-Agent Reinforcement Learning in USV Swarm.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Human Implicit Preference-Based Policy Fine-tuning for Multi-Agent Reinforcement Learning in USV Swarm

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-14T04:36:17.374053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:36:17.374053Z digest=sha256:5abbc509da6645ac4aea7b296b35294be7e3d01bf5abe4260ddd4b090a49fefc

Observation bcbf0b58-b3b3-4722-aeab-a4db5eebd672 · outbound

This paper cites A bayesian approach for policy learning from trajectory preference queries,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference A bayesian approach for policy learning from trajectory preference queries,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-14T04:36:17.394925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:36:17.394925Z digest=sha256:49965c0cba3ad1a2898a79e7eef6d01512e1da574f8b3b2c26bbd2bd9b3c7551

Observation ee756fbc-34e2-4a60-8b66-8b7f10dff330 · outbound

This paper cites Convergence of q-learning: A simple proof,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Convergence of q-learning: A simple proof,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-14T04:36:17.441628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:36:17.441628Z digest=sha256:0abd68cce695a9bd70b2d0f61477df737e1e18d1adc7721b96382ff4869a369e

Observation 26e99042-100d-4a8c-863e-52586cd6aa13 · outbound

This paper cites Decentralized multi-agent reinforcement learning: An off-policy method,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Decentralized multi-agent reinforcement learning: An off-policy method,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:36:19.098389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:17.474913Z digest=sha256:fedaaa42702cdcb30c81e4e66eac9161efdd0cdaadf82b54d09e19294a354c6c

Observation ca4f6545-5dc4-4165-8e14-aba295910844 · outbound

This paper cites An ocba-based method for efficient sample collection in reinforcement learning,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference An ocba-based method for efficient sample collection in reinforcement learning,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-14T04:36:17.492763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:36:17.492763Z digest=sha256:a15b554f46a3731885173eb6bf87712dd64e7d84c63054b06c64e75974a81f4c

Observation 0d47f4ff-4521-408e-a44b-cafb28994727 · outbound

This paper cites Equilibrium points in n-person games,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Equilibrium points in n-person games,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-14T04:36:17.526288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:36:17.526288Z digest=sha256:9da79c569b0b6eab6d9a6faf3a59eec413a71b8d2be782393bc4ccd01d237624

Observation de9c269f-a0ca-40cd-a4f1-846e3a380b07 · outbound

This paper cites Surf: Semi- supervised reward learning with data augmentation for feedback-efficient preference-based reinforcement learning,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Surf: Semi- supervised reward learning with data augmentation for feedback-efficient preference-based reinforcement learning,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-14T04:36:17.560521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:36:17.560521Z digest=sha256:bba46c50499911acebd76a4ecd410449ffc13d5cc7adbd7f5a1ba81d9dffdbc7

Observation 465e0962-6210-4547-b08c-a93257f75ba5 · outbound

This paper cites Rank analysis of incomplete block designs: I. the method of paired comparisons,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Rank analysis of incomplete block designs: I. the method of paired comparisons,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-14T04:36:17.595278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:36:17.595278Z digest=sha256:3cf8bc1ea649d01f3abdd8626e8d1586a8ce1164dcc8c2ffbf71c46d5bc039a6

Observation 98fb9612-b991-46bc-a7e9-db0ca9c575f7 · outbound

This paper cites Rademacher and gaussian complex- ities: Risk bounds and structural results,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Rademacher and gaussian complex- ities: Risk bounds and structural results,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:36:18.840876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:17.634871Z digest=sha256:2ee450efb06a19ffa783b39f1fef6a6d661afed05b004e1653c0cd5933b8c742

Observation c783cef6-1bb3-42c1-acbf-b3149e517db9 · outbound

This paper cites Emergence of Grounded Compositional Language in Multi-Agent Populations.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Emergence of Grounded Compositional Language in Multi-Agent Populations

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-14T04:36:17.671891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:36:17.671891Z digest=sha256:fedbecfe52d142fa78fe653fc25751d7d6ddc59d0f7a8647da5f3a60ab0ca8d7

Observation 62480f3e-fcb9-47fb-8794-effd0ae7a10a · outbound

This paper cites Multi- agent actor-critic for mixed cooperative-competitive environments,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Multi- agent actor-critic for mixed cooperative-competitive environments,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:36:18.785759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:17.694784Z digest=sha256:4b1a95916b1da9327277f31ac1814a566bdb43284d64377cb82f09286a51403e

Observation f13f5b2b-d5dd-45f6-96c3-8c0723c3a478 · outbound

This paper cites Balancing and scheduling of surface mount technology lines,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Balancing and scheduling of surface mount technology lines,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:36:18.695957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:17.722888Z digest=sha256:897856128624f371f0edff99328aa8c9c0d67899cc63d9956e003279595d60e1

Observation f36bf8b1-6198-4369-bdb4-3d275b6fd0e1 · outbound

This paper cites Flow shop scheduling prob- lems with assembly operations: a review and new trends,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Flow shop scheduling prob- lems with assembly operations: a review and new trends,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:36:18.611135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:17.741371Z digest=sha256:018e40402048c097139a5ae06350615ceb884a8908a5f024dda0109e1536cc14

Observation 562ce88a-aa78-40ea-9145-6753305ccc0a · outbound

This paper cites Modeling semiconductor testing job schedul- ing and dynamic testing machine configuration,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Modeling semiconductor testing job schedul- ing and dynamic testing machine configuration,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:36:18.529247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:17.764799Z digest=sha256:87c24cf28f5529ae01cd4b603e136376de2a3dd68b404963b029f44e9e2cb13e

Observation 3e5ec873-3062-46b0-ab6f-8975a7d77a97 · outbound

This paper cites Reducing Overestimation Bias in Multi-Agent Domains Using Double Centralized Critics.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Reducing Overestimation Bias in Multi-Agent Domains Using Double Centralized Critics

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-14T04:36:17.780718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:36:17.780718Z digest=sha256:45f3a161dc525e71c660d1c9010381df2465a17d858c675d0dc3e7b60bf73dc6

Observation 0c94c674-9dba-48e0-abb1-a0e0a538cd8b · outbound

This paper cites The o.d.e. method for convergence of stochastic approximation and reinforcement learning,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference The o.d.e. method for convergence of stochastic approximation and reinforcement learning,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:36:18.448334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:36:17.811075Z digest=sha256:21ba35ccbe021e94ebda81e4500e395106c955d1e60c00e78475b54a191f6868

Observation 107f356d-9fc7-46d1-872b-29368bfd01c9 · outbound

This paper cites an unresolved cited work.

Multi-Agent Reinforcement Learning via Agent-Specific Preference Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-14T04:36:17.829554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:36:17.829554Z digest=sha256:a760267c16eaa9f579092902456693599312cbb4afc30d8a788f96af58c7e819

Observation d9f12a40-c978-4c32-b782-58ebad53e6c0 · outbound

This paper cites A stochastic approximation method,.

Multi-Agent Reinforcement Learning via Agent-Specific Preference A stochastic approximation method,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-14T04:36:17.857891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:36:17.857891Z digest=sha256:391f15b55905c85af51e05c816e10df622526c3dcad13905c4bf71078202cd32

Pith citing papers

No inbound Pith citation observations are available.