Pith. sign in

Paper Citation Record · LEDGER

Policy Improvement with Style-Specific Demonstrations

As of 23 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2506.16995.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.16995 v4

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:21:12.730681Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact2
  • verified fuzzy13
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ca57c22c-5d3a-46b6-bd4f-ab6757cc2ec8 · outbound

This paper cites Dota 2 with Large Scale Deep Reinforcement Learning.

Policy Improvement with Style-Specific Demonstrations Dota 2 with Large Scale Deep Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T19:21:12.593800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:21:12.593800Z digest=sha256:c96ae765a9d3b8f7c9d9bd7efa615d0841b5417ac0f2c3c04215b752147f099a

Observation 02195ccc-98cc-4a8f-9f14-c9a707d2b8ce · outbound

This paper cites Superhuman ai for multiplayer poker.

Policy Improvement with Style-Specific Demonstrations Superhuman ai for multiplayer poker

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:21:13.198763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T19:21:12.599247Z digest=sha256:e649bed3d8698a1cbe47ce4e95082453fbb2e494c9335767f466ae95a0eb1bd9

Observation d4c915ad-951e-491c-ae5c-82dd43426633 · outbound

This paper cites Nvidia redefines game ai with ace autonomous game characters, 2025.

Policy Improvement with Style-Specific Demonstrations Nvidia redefines game ai with ace autonomous game characters, 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:21:13.184902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T19:21:12.603625Z digest=sha256:03fa49be45f21f49debaa4d7442c6971092b82adab1e7a52eeb69c76b2a21ce0

Observation 5115219d-98dc-4a69-9215-f9a14cf93b5a · outbound

This paper cites IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures.

Policy Improvement with Style-Specific Demonstrations IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T19:21:12.608395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:21:12.608395Z digest=sha256:47fe7094212438720144011028ca67b5471f4ef7fbad0d2e4cafb3acce98e68a

Observation 2ef97c96-828f-430e-a05e-cec82b9bedd9 · outbound

This paper cites Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio.

Policy Improvement with Style-Specific Demonstrations Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T19:21:12.613322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:21:12.613322Z digest=sha256:d7bddc6aded872f318a82165999b93c3560471ab06dff8f65f7c633b7c839696

Observation 8950b764-6685-4552-b0d6-b2502472311b · outbound

This paper cites Deep Q-learning from Demonstrations.

Policy Improvement with Style-Specific Demonstrations Deep Q-learning from Demonstrations

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T19:21:12.617446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:21:12.617446Z digest=sha256:75652cfa5d1b15dbe0a02daa6621dd00d952d2c608167d078f7ae8a5c41c0ecd

Observation 8cbbe87a-8ed4-49a0-8cd9-5b0668970259 · outbound

This paper cites Generative Adversarial Imitation Learning.

Policy Improvement with Style-Specific Demonstrations Generative Adversarial Imitation Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T19:21:12.622735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:21:12.622735Z digest=sha256:01a45e19dd717fbd747fa6c32cfa502567466568b4338d725796c355e36ab80e

Observation 5de823e1-9b42-46d7-a430-401aa717c47a · outbound

This paper cites Approximately optimal approximate reinforcement learning.

Policy Improvement with Style-Specific Demonstrations Approximately optimal approximate reinforcement learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:21:13.160736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T19:21:12.627338Z digest=sha256:437e65cbafbb597544e212c4d4d337423df51c88ab8be6417389b6a945fb981b

Observation 126ecc27-9a9e-4d8d-a420-07c3e5a03ae1 · outbound

This paper cites Policy optimization with demonstrations.

Policy Improvement with Style-Specific Demonstrations Policy optimization with demonstrations

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:21:13.145999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T19:21:12.631657Z digest=sha256:71e8c19eca4473fcc6fd2defc5e33f7dd54a22851f7e4a284cffc63b3f634a19

Observation 7c2a69dd-b3d0-4887-ba38-886e15e34db2 · outbound

This paper cites Conservative Q-Learning for Offline Reinforcement Learning.

Policy Improvement with Style-Specific Demonstrations Conservative Q-Learning for Offline Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T19:21:12.635900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:21:12.635900Z digest=sha256:b59f73212c1698bc7cdf36347da04a3408683bf0a50d616ae996f71c43101ba8

Observation c3a91af9-3366-4d28-97e0-31004a5c2463 · outbound

This paper cites Method for constructing artificial intelligence player with abstractions to markov decision processes in multiplayer game of mahjong.

Policy Improvement with Style-Specific Demonstrations Method for constructing artificial intelligence player with abstractions to markov decision processes in multiplayer game of mahjong

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:21:13.130920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T19:21:12.640727Z digest=sha256:1b1382c5803338d0dbd4ace69419ac968eb0c98ba31bf036e1004d9c03696991

Observation 78678a86-7a7c-4c24-8e98-06a005ec0a9d · outbound

This paper cites A unified game-theoretic approach to multiagent reinforcement learning.

Policy Improvement with Style-Specific Demonstrations A unified game-theoretic approach to multiagent reinforcement learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:21:13.115449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T19:21:12.645328Z digest=sha256:082d446f999254a6f22f8946f983988f296a643152d61f40249adec0bea91641

Observation 087fd091-e809-47ce-9b97-90610e49ba26 · outbound

This paper cites Suphx: Mastering Mahjong with Deep Reinforcement Learning.

Policy Improvement with Style-Specific Demonstrations Suphx: Mastering Mahjong with Deep Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T19:21:12.649716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:21:12.649716Z digest=sha256:fd0f8a2b33f997aa8fe8a6109866a5c85491dbf31c6922f92acb8154f31f0dab

Observation 057031ea-6220-4578-97f8-66f1dd18b5a3 · outbound

This paper cites Official international mahjong: A new playground for ai research.

Policy Improvement with Style-Specific Demonstrations Official international mahjong: A new playground for ai research

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:21:13.101287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T19:21:12.654526Z digest=sha256:7c956bf8aaf1caa25a5076acbcf059a632a0f58f57f7fa68353cdd4594e05e61

Observation d79b351d-50d2-44cd-9adc-6de5f26056b5 · outbound

This paper cites Building a computer mahjong player based on monte carlo simulation and opponent models.

Policy Improvement with Style-Specific Demonstrations Building a computer mahjong player based on monte carlo simulation and opponent models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:21:13.086791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T19:21:12.658813Z digest=sha256:8baa16e4297627772f9715e36478dcd05587d6b64dabe5f2e40290f94d05369e

Observation 557e7deb-6015-456f-a317-5f0ac27be47e · outbound

This paper cites Overcoming Exploration in Reinforcement Learning with Demonstrations.

Policy Improvement with Style-Specific Demonstrations Overcoming Exploration in Reinforcement Learning with Demonstrations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T19:21:12.663108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:21:12.663108Z digest=sha256:3402ce6b12fbbe4a5c89bcdd7490ebaa263d196535ef925fe57f5f0a189eeee3

Observation 405fb5cf-c2c2-462c-b2d0-aecb0765c448 · outbound

This paper cites Handbook on mahjong competition rules, 2016.

Policy Improvement with Style-Specific Demonstrations Handbook on mahjong competition rules, 2016

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:21:13.071936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T19:21:12.667661Z digest=sha256:58d0a95779e73b038a1b9029bd6e872b8c983759681ade1fc1bc9ce80d5134ab

Observation e96eacf4-8b1d-43e7-b269-6481250f5f57 · outbound

This paper cites an unresolved cited work.

Policy Improvement with Style-Specific Demonstrations Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:21:13.057541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T19:21:12.671994Z digest=sha256:5d50619e62daad676c292d303677ccb46a5d2b34eebb180d9584b3254669d06f

Observation 330d45d5-fc97-4d21-a156-e61ef9e9d8b9 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Policy Improvement with Style-Specific Demonstrations High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T19:21:12.676241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:21:12.676241Z digest=sha256:72990725f6a8fd286f19909ecd1fc78e93f1179300cc3239702a72b368ece73a

Observation fc122948-32cc-43f1-beac-ac578771c7de · outbound

This paper cites Jordan, and Pieter Abbeel.

Policy Improvement with Style-Specific Demonstrations Jordan, and Pieter Abbeel

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:21:12.681331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:21:12.681331Z digest=sha256:f66b57b9525f0583eddf2b237025794dd26e185e42bc1cd3b0a05d60859028bd

Observation 59357414-e1ef-41d6-b9dd-8afa77bcaad9 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Policy Improvement with Style-Specific Demonstrations Proximal Policy Optimization Algorithms

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T19:21:12.685791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:21:12.685791Z digest=sha256:8e68194a40703ff0dfbf5ef7ba213b365ce9823ffa08527a41871c1d3afb4d2a

Observation efa39ccf-c4bd-4677-b3bb-215f03bc8770 · outbound

This paper cites Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm.

Policy Improvement with Style-Specific Demonstrations Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T19:21:12.690475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:21:12.690475Z digest=sha256:c25207093f2a81d2f0335ad41554edfd9ea7cce7be139758c8c9df46c11e1465

Observation dcd816df-9cbd-4f93-b20c-ac64ed283c8a · outbound

This paper cites Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards.

Policy Improvement with Style-Specific Demonstrations Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T19:21:12.694943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:21:12.694943Z digest=sha256:59944ad6a7e4e466d56aaa7dd13770c49c41410f0265e98ae49255b98759a2a4

Observation 3c7b3f68-6930-4007-bed4-bf885d102ae6 · outbound

This paper cites Czarnecki, Micha \"e l Mathieu, Andrew Dudzik, Junyoung Chung, David H.

Policy Improvement with Style-Specific Demonstrations Czarnecki, Micha \"e l Mathieu, Andrew Dudzik, Junyoung Chung, David H

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:21:13.032461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T19:21:12.699471Z digest=sha256:c6529a10b6eb471a395fbdc2557d8867c408fff8c15205dcf223ac5954e83027

Observation d1f718d7-9ee3-4d05-baca-b05b3d4cc908 · outbound

This paper cites Perfectdou: Dominating doudizhu with perfect information distillation, 2024.

Policy Improvement with Style-Specific Demonstrations Perfectdou: Dominating doudizhu with perfect information distillation, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:21:13.016806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T19:21:12.703785Z digest=sha256:ddd11e43df8cba3f45c94640476f3280c08775700c3371f690b391fc5f8380cd

Observation 173442c0-ff24-4094-8167-6d2242f30bb4 · outbound

This paper cites Towards Playing Full MOBA Games with Deep Reinforcement Learning.

Policy Improvement with Style-Specific Demonstrations Towards Playing Full MOBA Games with Deep Reinforcement Learning

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:21:12.795283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T19:21:12.708064Z digest=sha256:7f06bbc8777e7690cf9c6f7016198ce570d6d72420cde051c428d5f0674a2e4b

Observation 302e7d29-979d-412d-a948-1fdda8081560 · outbound

This paper cites an unresolved cited work.

Policy Improvement with Style-Specific Demonstrations Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:21:13.002262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T19:21:12.712541Z digest=sha256:289aa228f9c96f53c94e0da07601f03da022eb7cd226c14f079059b3a3c49f57

Observation 07e4e0b7-95b3-48af-a75a-5d9e1405fd81 · outbound

This paper cites Botzone: an online multi-agent competitive platform for ai education.

Policy Improvement with Style-Specific Demonstrations Botzone: an online multi-agent competitive platform for ai education

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:21:12.987785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T19:21:12.717186Z digest=sha256:a60792c464d9e9b27e0ca5c6256619984cd29bb446181749b9338c6cf4611911

Observation 4c7b16fc-609f-494f-a652-43c6948cbd90 · outbound

This paper cites Learning Sparse Rewarded Tasks from Sub-Optimal Demonstrations.

Policy Improvement with Style-Specific Demonstrations Learning Sparse Rewarded Tasks from Sub-Optimal Demonstrations

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:21:12.774626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T19:21:12.721639Z digest=sha256:511b8fea32ad70490fe342963423dd5c1cf28b866f6c113717ff2731b097c58f

Observation c2d288bc-147d-4240-a106-970f8d88c848 · outbound

This paper cites Behavior proximal policy optimization, 2023.

Policy Improvement with Style-Specific Demonstrations Behavior proximal policy optimization, 2023

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:21:12.973565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T19:21:12.726434Z digest=sha256:3c910db153939860d16b5e79a291bcc89cfcded13d5947e8db1876a2abc66d40

Observation affdf193-c327-4a13-a16a-86b1a70ef7f4 · outbound

This paper cites write newline.

Policy Improvement with Style-Specific Demonstrations write newline

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T19:21:12.730681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:21:12.730681Z digest=sha256:c86ee41a0609d8b4cb30c0cbb6f766ce175db1bd1cf8599aac61e1c74adf8ae7

Pith citing papers

No inbound Pith citation observations are available.