Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:21:12.730681Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2506.16995.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:21:12.730681Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
31 of 31 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ca57c22c-5d3a-46b6-bd4f-ab6757cc2ec8 · outbound
Policy Improvement with Style-Specific Demonstrations Dota 2 with Large Scale Deep Reinforcement Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02195ccc-98cc-4a8f-9f14-c9a707d2b8ce · outbound
Policy Improvement with Style-Specific Demonstrations Superhuman ai for multiplayer poker
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d4c915ad-951e-491c-ae5c-82dd43426633 · outbound
Policy Improvement with Style-Specific Demonstrations Nvidia redefines game ai with ace autonomous game characters, 2025
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5115219d-98dc-4a69-9215-f9a14cf93b5a · outbound
Policy Improvement with Style-Specific Demonstrations IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ef97c96-828f-430e-a05e-cec82b9bedd9 · outbound
Policy Improvement with Style-Specific Demonstrations Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8950b764-6685-4552-b0d6-b2502472311b · outbound
Policy Improvement with Style-Specific Demonstrations Deep Q-learning from Demonstrations
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cbbe87a-8ed4-49a0-8cd9-5b0668970259 · outbound
Policy Improvement with Style-Specific Demonstrations Generative Adversarial Imitation Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5de823e1-9b42-46d7-a430-401aa717c47a · outbound
Policy Improvement with Style-Specific Demonstrations Approximately optimal approximate reinforcement learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 126ecc27-9a9e-4d8d-a420-07c3e5a03ae1 · outbound
Policy Improvement with Style-Specific Demonstrations Policy optimization with demonstrations
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7c2a69dd-b3d0-4887-ba38-886e15e34db2 · outbound
Policy Improvement with Style-Specific Demonstrations Conservative Q-Learning for Offline Reinforcement Learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3a91af9-3366-4d28-97e0-31004a5c2463 · outbound
Policy Improvement with Style-Specific Demonstrations Method for constructing artificial intelligence player with abstractions to markov decision processes in multiplayer game of mahjong
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 78678a86-7a7c-4c24-8e98-06a005ec0a9d · outbound
Policy Improvement with Style-Specific Demonstrations A unified game-theoretic approach to multiagent reinforcement learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 087fd091-e809-47ce-9b97-90610e49ba26 · outbound
Policy Improvement with Style-Specific Demonstrations Suphx: Mastering Mahjong with Deep Reinforcement Learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 057031ea-6220-4578-97f8-66f1dd18b5a3 · outbound
Policy Improvement with Style-Specific Demonstrations Official international mahjong: A new playground for ai research
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d79b351d-50d2-44cd-9adc-6de5f26056b5 · outbound
Policy Improvement with Style-Specific Demonstrations Building a computer mahjong player based on monte carlo simulation and opponent models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 557e7deb-6015-456f-a317-5f0ac27be47e · outbound
Policy Improvement with Style-Specific Demonstrations Overcoming Exploration in Reinforcement Learning with Demonstrations
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 405fb5cf-c2c2-462c-b2d0-aecb0765c448 · outbound
Policy Improvement with Style-Specific Demonstrations Handbook on mahjong competition rules, 2016
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e96eacf4-8b1d-43e7-b269-6481250f5f57 · outbound
Policy Improvement with Style-Specific Demonstrations Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 330d45d5-fc97-4d21-a156-e61ef9e9d8b9 · outbound
Policy Improvement with Style-Specific Demonstrations High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc122948-32cc-43f1-beac-ac578771c7de · outbound
Policy Improvement with Style-Specific Demonstrations Jordan, and Pieter Abbeel
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59357414-e1ef-41d6-b9dd-8afa77bcaad9 · outbound
Policy Improvement with Style-Specific Demonstrations Proximal Policy Optimization Algorithms
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efa39ccf-c4bd-4677-b3bb-215f03bc8770 · outbound
Policy Improvement with Style-Specific Demonstrations Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcd816df-9cbd-4f93-b20c-ac64ed283c8a · outbound
Policy Improvement with Style-Specific Demonstrations Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c7b3f68-6930-4007-bed4-bf885d102ae6 · outbound
Policy Improvement with Style-Specific Demonstrations Czarnecki, Micha \"e l Mathieu, Andrew Dudzik, Junyoung Chung, David H
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d1f718d7-9ee3-4d05-baca-b05b3d4cc908 · outbound
Policy Improvement with Style-Specific Demonstrations Perfectdou: Dominating doudizhu with perfect information distillation, 2024
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 173442c0-ff24-4094-8167-6d2242f30bb4 · outbound
Policy Improvement with Style-Specific Demonstrations Towards Playing Full MOBA Games with Deep Reinforcement Learning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 302e7d29-979d-412d-a948-1fdda8081560 · outbound
Policy Improvement with Style-Specific Demonstrations Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 07e4e0b7-95b3-48af-a75a-5d9e1405fd81 · outbound
Policy Improvement with Style-Specific Demonstrations Botzone: an online multi-agent competitive platform for ai education
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4c7b16fc-609f-494f-a652-43c6948cbd90 · outbound
Policy Improvement with Style-Specific Demonstrations Learning Sparse Rewarded Tasks from Sub-Optimal Demonstrations
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c2d288bc-147d-4240-a106-970f8d88c848 · outbound
Policy Improvement with Style-Specific Demonstrations Behavior proximal policy optimization, 2023
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation affdf193-c327-4a13-a16a-86b1a70ef7f4 · outbound
Policy Improvement with Style-Specific Demonstrations write newline
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.