Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T04:36:17.857891Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2608.08604.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T04:36:17.857891Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
55 of 55 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3af094d7-4ce1-4aea-844c-a39dc5b985b4 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Distributed multiagent coordinated learning for autonomous driving in highways based on dynamic coordination graphs,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8cd10382-4c85-45db-8a97-b414a458fa7d · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Beyond static populations: Efficient delay-constrained scheduling for dynamic users via deep reinforcement learning,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 84b729e2-6981-4689-8fcc-f606dcad4ab1 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Dota 2 with Large Scale Deep Reinforcement Learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 672e6592-83e5-4223-8171-43728f0721d6 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference The surprising effectiveness of ppo in cooperative multi-agent games,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d91d45e6-93c3-4e59-b7f8-1be87f64aef5 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Value-decomposition networks for cooperative multi-agent learning based on team reward,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 19f60790-e190-4f43-9d4b-3a7c9c4c957b · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2519376a-93f8-4e40-8f62-148aa8849795 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Beyond shallow behavior: Task-efficient value-based multi-task offline marl via skill discovery,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c71f5dc0-2a61-4241-a0a7-cb1c02f58d6c · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Reward learning from human preferences and demonstrations in atari,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 286c746c-db69-420f-8113-3c9c16227677 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Pebble: Feedback-efficient interactive reinforcement learning via relabeling experience and unsuper- vised pre-training,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89ac83be-13c5-4498-903b-d5f98857b25a · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference S-EPOA: Overcoming the Indistinguishability of Segments with Skill-Driven Preference-Based Reinforcement Learning,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7f329cea-5943-40af-b280-b5c3902bb3e5 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambigu- ous Queries,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 92602b96-256a-4d30-8a97-7bf1908acc8c · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference STAIR: Addressing stage misalignment through temporal-aligned preference reinforcement learning,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a6b65122-d2cf-410d-843f-4ce4352deed5 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference MAGPIE: Utilizing Agent-Specific Preferences for Multi-Agent Reinforcement Learning,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 78293d41-b356-4e8c-aff9-f1a302f60518 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Senior: Efficient query selection and preference-guided exploration in preference-based rein- forcement learning,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8fcfbad5-977a-4eab-8d80-074a0bdbe692 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Training language models to follow instructions with human feedback,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1baf82a0-27b7-419c-b71e-6bbbb7ac8342 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Preference-based Multi-Objective Reinforcement Learning,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 32adde6a-ccb5-49e5-b8cb-e57f8486e132 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference COLLIE: Guiding Skill Discovery in Semantically Coherent Latent Space,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 50e544e2-8d1b-488a-adec-069260b18d2b · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Offline multi-agent preference-based reinforcement learning with agent-aware direct preference optimization,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation aa0a22fc-1975-4cdf-a19f-d1ac2c8c3a10 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference O-MAPL: Offline multi- agent preference learning,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4c097685-29d3-4959-9f1a-dd45e116419b · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Decoding global preferences: Tem- poral and cooperative dependency modeling in multi-agent preference- based reinforcement learning,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c12da3cf-cc42-460e-b52c-039311c79143 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference DPM: Dual preferences-based multi-agent reinforcement learning,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation cd7c8ac9-fb25-4c05-b6a4-5ed466ed21be · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Multi-agent reinforcement learning from human feedback: Data coverage and algorithmic techniques,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9a875b21-3ef4-4c71-90d5-1a40993bf5db · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Mastering the game of go without human knowledge,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 69448ec0-32ba-477d-8031-c9ceaa81f02d · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Deep reinforcement learning for the control of robotic manipulation: a focussed mini-review,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8ea2054-58d0-40cf-999c-0fd7eaea0ac8 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Integrating Mechanism and Data: Reinforcement Learning Based on Multi-fidelity Model for Data Center Cooling Control,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 63e9da64-a45a-4f14-b412-a2ec9c8a03e2 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Large-scale Data Center Cooling Control via Sample-efficient Reinforcement Learning,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 455e0aea-83a9-4856-a961-bda2efcd8ea3 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference E-mapp: Efficient Multi- Agent Reinforcement Learning with Parallel Program Guidance,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5e5d5239-89e1-4d51-838d-79d943d84d02 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference From solo to symphony: Orchestrating multi-agent collaboration with single-agent demos,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a48d4323-7fc4-4eb6-ab98-175952a2b5f7 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference GlobeDiff: State Diffusion Process for Partial Ob- servability in Multi-Agent Systems,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6dbcd36-759c-4bc3-a562-1acb4a884737 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Multi-agent reinforcement learning for resources allocation optimization: a survey,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b9601f21-9653-406d-b5ac-2eed6a9c5d6e · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference A review of cooperative multiagent deep reinforcement learning,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 73380aa3-8e5a-4436-a6e5-465071410ecd · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Reinforcement learning with sparse rewards using guidance from offline demonstration,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 87a477fc-ae7c-42c3-9c30-63f5fe8fe6cc · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Reward function design in reinforcement learning,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e6e6999-c57b-4404-bbb5-07e52cf796d0 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Liir: Learning indi- vidual intrinsic reward in multi-agent reinforcement learning,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9b8c8573-55ce-4d36-89f8-84e7c55fcf1d · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Quality assessment of 3d human animation: Subjective and objective evaluation,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8f839c82-5b23-4326-b4d7-017f20882504 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Deep reinforcement learning from human preferences,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cdd9fec-c79b-4520-832c-01fec04eaadf · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Preference-Based Multi-Objective Reinforcement Learning with Explicit Reward Modeling,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation dbb60e09-a73f-4594-9b6d-4749e80a317b · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Human Implicit Preference-Based Policy Fine-tuning for Multi-Agent Reinforcement Learning in USV Swarm
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcbf0b58-b3b3-4722-aeab-a4db5eebd672 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference A bayesian approach for policy learning from trajectory preference queries,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee756fbc-34e2-4a60-8b66-8b7f10dff330 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Convergence of q-learning: A simple proof,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26e99042-100d-4a8c-863e-52586cd6aa13 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Decentralized multi-agent reinforcement learning: An off-policy method,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ca4f6545-5dc4-4165-8e14-aba295910844 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference An ocba-based method for efficient sample collection in reinforcement learning,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d47f4ff-4521-408e-a44b-cafb28994727 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Equilibrium points in n-person games,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de9c269f-a0ca-40cd-a4f1-846e3a380b07 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Surf: Semi- supervised reward learning with data augmentation for feedback-efficient preference-based reinforcement learning,
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 465e0962-6210-4547-b08c-a93257f75ba5 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Rank analysis of incomplete block designs: I. the method of paired comparisons,
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98fb9612-b991-46bc-a7e9-db0ca9c575f7 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Rademacher and gaussian complex- ities: Risk bounds and structural results,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c783cef6-1bb3-42c1-acbf-b3149e517db9 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Emergence of Grounded Compositional Language in Multi-Agent Populations
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62480f3e-fcb9-47fb-8794-effd0ae7a10a · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Multi- agent actor-critic for mixed cooperative-competitive environments,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f13f5b2b-d5dd-45f6-96c3-8c0723c3a478 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Balancing and scheduling of surface mount technology lines,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f36bf8b1-6198-4369-bdb4-3d275b6fd0e1 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Flow shop scheduling prob- lems with assembly operations: a review and new trends,
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 562ce88a-aa78-40ea-9145-6753305ccc0a · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Modeling semiconductor testing job schedul- ing and dynamic testing machine configuration,
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3e5ec873-3062-46b0-ab6f-8975a7d77a97 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Reducing Overestimation Bias in Multi-Agent Domains Using Double Centralized Critics
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c94c674-9dba-48e0-abb1-a0e0a538cd8b · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference The o.d.e. method for convergence of stochastic approximation and reinforcement learning,
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 107f356d-9fc7-46d1-872b-29368bfd01c9 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference Unresolved cited work
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9f12a40-c978-4c32-b782-58ebad53e6c0 · outbound
Multi-Agent Reinforcement Learning via Agent-Specific Preference A stochastic approximation method,
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.