Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2109.11251.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:23:07.654898Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
83
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 1cff9fb9-eb41-4802-a372-4046bfc8210e · inbound
Light Aircraft Game : Basic Implementation and training results analysis Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9f7b077-a641-4f1e-b536-eba274abbe39 · inbound
Improving monotonic optimization in heterogeneous multi-agent reinforcement learning with optimal marginal deterministic policy gradient Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecd55247-3864-45aa-92d3-e5006119891a · inbound
Centralized Adaptive Sampling for Reliable Co-Training of Independent Multi-Agent Policies Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 924df6fc-5af8-4b3f-8a73-96966cc68a29 · inbound
Multi-Agent Trust Region Policy Optimisation: A Joint Constraint Approach Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce1cbd52-a50d-43df-9d3c-f212f995948f · inbound
TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation acc98b05-d732-4449-9d73-9c2f0350a4cc · inbound
TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d29723a-b446-4a5e-9b8e-cae77fc18570 · inbound
Topology-Driven Anti-Entanglement Control for Soft Robots Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b778fc2f-accf-435a-8f96-be38f0240d6c · inbound
Robust Instruction Compliance in Cooperative Multi-Agent Reinforcement Learning Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning
Reference 122
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 401883ea-8851-4bfa-9b3c-1bfcc7509581 · inbound
Self-Supervised On-Policy Reinforcement Learning via Contrastive Proximal Policy Optimisation Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1a6b8c58-8480-4861-9f48-66da0251baad · inbound
TeamTR: Trust-Region Fine-Tuning for Multi-Agent LLM Coordination Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 113eb189-443b-40bf-9fcc-d7de7a1804dc · inbound
Phi-Actor-Critic: Steering General-Sum Games to Pareto-Efficient Correlated Equilibria Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bbbcca0b-11cc-4c9c-93b3-d697a1a39799 · inbound
Reference-Free Heterogeneous Multi-Agent Reinforcement Learning for Grid-Friendly Tie-Line Power Shaping in Industrial Microgrids Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 235d54c4-9e76-462f-9e21-80b4af10140d · inbound
Action-Factored Multi-Agent Reinforcement Learning for Scalable Quantum Device Tuning Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d46f04a-1a2d-4c1b-87ba-8bdd023b8cd3 · inbound
Heterogeneous Multi-Agent Reinforcement Learning for Radio Resource Management under Coupled Finite-Horizon Constraints Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.