Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T14:53:31.878138Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2606.21085.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T14:53:31.878138Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5e56c8ae-9b97-47ef-a257-e3ab9e932fe3 · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4da55aa2-6ad8-4802-b7e0-98228f68d925 · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition D4RL: Datasets for Deep Data-Driven Reinforcement Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cb33f670-bd87-434c-9cdb-ef1ee4dad3ec · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition Accelerating online reinforcement learning with offline datasets,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a15bbbce-6ff8-4103-8f59-f4e651534e2f · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition Offline reinforcement learning with implicit q-learning,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a7b4799-8310-4170-9c84-13c087906096 · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6791f76b-ff93-446f-b8af-404759e7aa66 · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition A minimalist approach to offline reinforcement learning,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c84a435-1ae6-4e59-8284-b3a2895a5301 · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition Policy expansion for bridging offline-to- online reinforcement learning,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04a7f2e6-32f9-49cf-8bd8-b433c0661fb8 · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition Bayesian design principles for offline-to-online reinforcement learning,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fdf59f5-5d5f-4c84-ae03-3b28cc7a7343 · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition FACMAC: Factored Multi-Agent Centralised Policy Gradients
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3a793b8a-4899-4a34-9253-3efceb7105b2 · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition Opride: Efficient offline preference-based reinforcement learning via in-dataset exploration,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3454fff4-12e0-48ac-a4e9-41f5ff6020eb · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition Batch policy learning under constraints,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18438a75-e91c-4842-8308-1f619c6451a9 · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition Flow to control: Offline reinforcement learning with lossless primitive discovery,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 648d7c2e-9240-434f-b9df-8ba1dc4e955e · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition Off-policy deep reinforcement learning without exploration,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6993e16f-fcb4-4349-9f1d-5890129fb102 · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition Stabilizing off-policy q-learning via bootstrapping error reduction,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c801630-ff42-4b77-aeda-774471b68a88 · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition Behavior regularized offline reinforcement learning,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87c45629-542b-458e-8fdc-8758e736f135 · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition Critic regularized regression,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03d9989d-ce4d-4349-812c-7da7d6a898ab · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition Advantage-weighted regression: Simple and scalable offline reinforcement learning,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9695e938-52a9-43c4-a9db-dde80651fe69 · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition A survey on offline multi-agent reinforcement learning,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba2c7537-050c-4a56-92a0-ba15e1e27f03 · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition Offline multi-agent reinforcement learning: Guidelines and benchmarks,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2acc15f1-db4f-47b1-8ccc-6aee027fe0ca · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition Offline multi-agent reinforcement learning with knowledge distillation,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 762331be-8f84-425d-af83-eccf69c2c4ba · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition Offline decentralized multi-agent reinforcement learning,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 426bc625-36cd-42cc-96b5-4446e1dae85f · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition Globediff: State diffusion process for partial observability in multi-agent systems,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9466d7ba-4cd8-44a4-ab55-7a09b5094875 · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition Believe what you see: Implicit constraint approach for offline multi-agent reinforcement learning,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 487d87dd-ba1a-4789-9a00-575b8473d661 · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition Plan better amid conservatism: Offline multi-agent reinforcement learning with actor rectification,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a4cfc7e-10bb-4c23-b975-d17992854b68 · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition Counterfactual conservative q learning for offline multi-agent reinforcement learning,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68e56995-11cc-410d-aaeb-817cbc992a6b · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition Offline multi-agent rein- forcement learning with implicit global-to-local value regularization,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d46a85c3-6b36-40cf-8976-2b4de62608dc · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a11b9d84-2657-476f-a518-92bdf153ea89 · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition Qplex: Duplex dueling multi-agent q-learning,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1eb8af10-ba6d-463f-a3f2-bfcbee1efdb8 · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition Cal-ql: Calibrated offline reinforcement learning,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 868ee271-a8e6-41f7-a3dc-027c8fb10325 · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition Proto: Iterative policy regularized offline-to-online reinforcement learning,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a982434-94a4-4bd5-b4be-d5fc93266f52 · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition Efficient online reinforcement learning with offline data,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53038486-179f-4854-bc5c-385ff3cdb7d0 · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition Trust region policy optimisation in multi-agent reinforcement learning,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0061c53-725e-4e4b-9428-c7d416096b13 · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition First, we incorporate the standardized public offline datasets released by the OMIGA benchmark [26]
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ed56350-bd92-47a8-97ab-c33aafd14583 · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e75f9f9-1697-48c1-b93a-ec860bb1549a · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition Unresolved cited work
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14a89a6a-86fd-4cdd-9998-99d699ab03af · outbound
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition To establish a rigor- ous multi-agent comparison, we extend these algorithms to their multi-agent counterparts, denoted as PEX-MA, AW AC-MA, RLPD-MA, and PROTO-MA
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.