Pith. sign in

Paper Citation Record · LEDGER

Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2109.11251.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2109.11251 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:23:07.654898Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

83
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1cff9fb9-eb41-4802-a372-4046bfc8210e · inbound

Light Aircraft Game : Basic Implementation and training results analysis cites this paper.

Light Aircraft Game : Basic Implementation and training results analysis Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:07.654898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:07.654898Z digest=sha256:dbde5e4a4856a63fdae36663317c2a2827c572e534d8e0f3f7fb33cc62690838

Observation b9f7b077-a641-4f1e-b536-eba274abbe39 · inbound

Improving monotonic optimization in heterogeneous multi-agent reinforcement learning with optimal marginal deterministic policy gradient cites this paper.

Improving monotonic optimization in heterogeneous multi-agent reinforcement learning with optimal marginal deterministic policy gradient Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:37.242154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:37.242154Z digest=sha256:67fa610064ac1bf003b40914f5c58f88938f7b3a052159b4b719913778f968cc

Observation ecd55247-3864-45aa-92d3-e5006119891a · inbound

Centralized Adaptive Sampling for Reliable Co-Training of Independent Multi-Agent Policies cites this paper.

Centralized Adaptive Sampling for Reliable Co-Training of Independent Multi-Agent Policies Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T01:06:57.409927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T01:05:17.086399Z digest=sha256:66a65c5e0eae7a6c7deba6ed717efd50bb6eacb13abad150de27a831f8ed61d2

Observation 924df6fc-5af8-4b3f-8a73-96966cc68a29 · inbound

Multi-Agent Trust Region Policy Optimisation: A Joint Constraint Approach cites this paper.

Multi-Agent Trust Region Policy Optimisation: A Joint Constraint Approach Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:02.434225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:31:02.434225Z digest=sha256:357688c4c8399f9c6baeb3e15e1180bf999aeb36fcdf3a45a1c64d1e8672d291

Observation ce1cbd52-a50d-43df-9d3c-f212f995948f · inbound

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning cites this paper.

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:45:29.793674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T07:41:48.634351Z digest=sha256:ff0343288a640b80196ebc61f38480dea424e6136d303b8230a2be00e05084bc

Observation acc98b05-d732-4449-9d73-9c2f0350a4cc · inbound

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning cites this paper.

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T05:37:38.440274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:37:38.440274Z digest=sha256:55991b9b9526d7a9a1b7ae957964439b3bc8d4494fd4f8c9b02b575ef7021a24

Observation 1d29723a-b446-4a5e-9b8e-cae77fc18570 · inbound

Topology-Driven Anti-Entanglement Control for Soft Robots cites this paper.

Topology-Driven Anti-Entanglement Control for Soft Robots Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:41:28.058511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T19:28:39.040470Z digest=sha256:689dbdf3d2af6734f0873c6f67da0389e99bce77e56267bddd3a57b09d7e3d29

Observation b778fc2f-accf-435a-8f96-be38f0240d6c · inbound

Robust Instruction Compliance in Cooperative Multi-Agent Reinforcement Learning cites this paper.

Robust Instruction Compliance in Cooperative Multi-Agent Reinforcement Learning Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 122

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:15:05.903774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-30T22:11:35.277901Z digest=sha256:4b8d63c0487bb0e6083cf156bc52db3785512217ef2a56fc6e36a03cfe448577

Observation 401883ea-8851-4bfa-9b3c-1bfcc7509581 · inbound

Self-Supervised On-Policy Reinforcement Learning via Contrastive Proximal Policy Optimisation cites this paper.

Self-Supervised On-Policy Reinforcement Learning via Contrastive Proximal Policy Optimisation Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:12:54.902012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T20:12:07.920183Z digest=sha256:1b4efe7045f56b1831e93eedd8b3881732e111b4d53834669dbec1e97db69443

Observation 1a6b8c58-8480-4861-9f48-66da0251baad · inbound

TeamTR: Trust-Region Fine-Tuning for Multi-Agent LLM Coordination cites this paper.

TeamTR: Trust-Region Fine-Tuning for Multi-Agent LLM Coordination Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T18:02:42.323932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T18:01:06.649723Z digest=sha256:5230add8006f11f73c22b560d34766b7cb6f071343cac3d43c225f222dd6bbf2

Observation 113eb189-443b-40bf-9fcc-d7de7a1804dc · inbound

Phi-Actor-Critic: Steering General-Sum Games to Pareto-Efficient Correlated Equilibria cites this paper.

Phi-Actor-Critic: Steering General-Sum Games to Pareto-Efficient Correlated Equilibria Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T08:47:50.453357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T10:42:17.852996Z digest=sha256:30056809b0cef17880572a211b36b1eed6f518ef41e2d46fbe7119289b1b87db

Observation bbbcca0b-11cc-4c9c-93b3-d697a1a39799 · inbound

Reference-Free Heterogeneous Multi-Agent Reinforcement Learning for Grid-Friendly Tie-Line Power Shaping in Industrial Microgrids cites this paper.

Reference-Free Heterogeneous Multi-Agent Reinforcement Learning for Grid-Friendly Tie-Line Power Shaping in Industrial Microgrids Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:50:11.104720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-25T19:36:51.509710Z digest=sha256:534fa58b6343c043d0e41a1cf22ff69c68be23fcb9d978d2fdd2f508da095d88

Observation 235d54c4-9e76-462f-9e21-80b4af10140d · inbound

Action-Factored Multi-Agent Reinforcement Learning for Scalable Quantum Device Tuning cites this paper.

Action-Factored Multi-Agent Reinforcement Learning for Scalable Quantum Device Tuning Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-13T03:08:01.590659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:08:01.590659Z digest=sha256:4eca52da32d356ebe6508e28ffad7f9baf2ca6d3f22dec0959c3e15d7b05ce11

Observation 8d46f04a-1a2d-4c1b-87ba-8bdd023b8cd3 · inbound

Heterogeneous Multi-Agent Reinforcement Learning for Radio Resource Management under Coupled Finite-Horizon Constraints cites this paper.

Heterogeneous Multi-Agent Reinforcement Learning for Radio Resource Management under Coupled Finite-Horizon Constraints Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T21:47:09.010614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:47:09.010614Z digest=sha256:6c0cc5bffa1e916885443ae820ff32df88e086203d5a8f38f0d37186f72eed9e