Pith. sign in

Paper Citation Record · LEDGER

Policy Improvement with Style-Specific Demonstrations

As of 18 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2506.16995.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.16995 v4

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:21:12.730681Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact2
  • verified fuzzy13
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ca57c22c-5d3a-46b6-bd4f-ab6757cc2ec8 · outbound

This paper cites Dota 2 with Large Scale Deep Reinforcement Learning.

Policy Improvement with Style-Specific Demonstrations Dota 2 with Large Scale Deep Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T19:21:12.593800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:21:12.593800Z digest=sha256:a8e88b8a6223a34b394a28f8319143e4d247cf490f871c2684d41a9ded6de701

Observation 02195ccc-98cc-4a8f-9f14-c9a707d2b8ce · outbound

This paper cites Superhuman ai for multiplayer poker.

Policy Improvement with Style-Specific Demonstrations Superhuman ai for multiplayer poker

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:21:13.198763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:21:12.599247Z digest=sha256:cc6bffa5169118563a87bdce1b51eccc7411ea987e1b909e37f8f7148342b078

Observation d4c915ad-951e-491c-ae5c-82dd43426633 · outbound

This paper cites Nvidia redefines game ai with ace autonomous game characters, 2025.

Policy Improvement with Style-Specific Demonstrations Nvidia redefines game ai with ace autonomous game characters, 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:21:13.184902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:21:12.603625Z digest=sha256:25c13a1ba7ee2ee8cf0ef3154e49fe977e09072aa161f51d8c85cb73b294ec89

Observation 5115219d-98dc-4a69-9215-f9a14cf93b5a · outbound

This paper cites IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures.

Policy Improvement with Style-Specific Demonstrations IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T19:21:12.608395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:21:12.608395Z digest=sha256:686d6f29dad81c8c7909f207dc9bdfd76d904ecee9afc379580b00d6a42800c2

Observation 2ef97c96-828f-430e-a05e-cec82b9bedd9 · outbound

This paper cites Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio.

Policy Improvement with Style-Specific Demonstrations Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T19:21:12.613322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:21:12.613322Z digest=sha256:4d89c4d96d1672c7731a6bc001eeb04e24dcf0bac6b04480b6d9a9864a7cc367

Observation 8950b764-6685-4552-b0d6-b2502472311b · outbound

This paper cites Deep Q-learning from Demonstrations.

Policy Improvement with Style-Specific Demonstrations Deep Q-learning from Demonstrations

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T19:21:12.617446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:21:12.617446Z digest=sha256:404e3878c12668ef362c73a2b11223cd7826de3b467c445f2ee3988fb40965cc

Observation 8cbbe87a-8ed4-49a0-8cd9-5b0668970259 · outbound

This paper cites Generative Adversarial Imitation Learning.

Policy Improvement with Style-Specific Demonstrations Generative Adversarial Imitation Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T19:21:12.622735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:21:12.622735Z digest=sha256:5a6b97604021748ac3505e84623c47234910f7395a95e2ca9c705cb67702ae7b

Observation 5de823e1-9b42-46d7-a430-401aa717c47a · outbound

This paper cites Approximately optimal approximate reinforcement learning.

Policy Improvement with Style-Specific Demonstrations Approximately optimal approximate reinforcement learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:21:13.160736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:21:12.627338Z digest=sha256:928ad32477d2dc53fdb472451d4679e8d43d63d4b92ff1a54335add60c87b645

Observation 126ecc27-9a9e-4d8d-a420-07c3e5a03ae1 · outbound

This paper cites Policy optimization with demonstrations.

Policy Improvement with Style-Specific Demonstrations Policy optimization with demonstrations

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:21:13.145999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:21:12.631657Z digest=sha256:2b92b3e92c931b68ed70ff94587d908e1ff3cd0945e3a7ac46d563b23c5fca6c

Observation 7c2a69dd-b3d0-4887-ba38-886e15e34db2 · outbound

This paper cites Conservative Q-Learning for Offline Reinforcement Learning.

Policy Improvement with Style-Specific Demonstrations Conservative Q-Learning for Offline Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T19:21:12.635900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:21:12.635900Z digest=sha256:5d3decc67e809b87e08d97f6defd49adeb3fb761090baa9f5855b124913f543f

Observation c3a91af9-3366-4d28-97e0-31004a5c2463 · outbound

This paper cites Method for constructing artificial intelligence player with abstractions to markov decision processes in multiplayer game of mahjong.

Policy Improvement with Style-Specific Demonstrations Method for constructing artificial intelligence player with abstractions to markov decision processes in multiplayer game of mahjong

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:21:13.130920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:21:12.640727Z digest=sha256:114d682d64d18d68183148fa2e6e48030a2d57526cd15dfe3c7b71f2d8b09bce

Observation 78678a86-7a7c-4c24-8e98-06a005ec0a9d · outbound

This paper cites A unified game-theoretic approach to multiagent reinforcement learning.

Policy Improvement with Style-Specific Demonstrations A unified game-theoretic approach to multiagent reinforcement learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:21:13.115449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:21:12.645328Z digest=sha256:59a738a089e9223d166a33f1a61805d58e6a7e3828e1f45dedbdcfd47601c9ec

Observation 087fd091-e809-47ce-9b97-90610e49ba26 · outbound

This paper cites Suphx: Mastering Mahjong with Deep Reinforcement Learning.

Policy Improvement with Style-Specific Demonstrations Suphx: Mastering Mahjong with Deep Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T19:21:12.649716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:21:12.649716Z digest=sha256:e45c0953ea55b27e229f9bbb15cf8026f976c54bbbc5c139dc64aaa14828c713

Observation 057031ea-6220-4578-97f8-66f1dd18b5a3 · outbound

This paper cites Official international mahjong: A new playground for ai research.

Policy Improvement with Style-Specific Demonstrations Official international mahjong: A new playground for ai research

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:21:13.101287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:21:12.654526Z digest=sha256:d29d9bd315c394a4dea578bfeb9ea9b72d0b4993a02f7a646dfb4157000cd612

Observation d79b351d-50d2-44cd-9adc-6de5f26056b5 · outbound

This paper cites Building a computer mahjong player based on monte carlo simulation and opponent models.

Policy Improvement with Style-Specific Demonstrations Building a computer mahjong player based on monte carlo simulation and opponent models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:21:13.086791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:21:12.658813Z digest=sha256:176df6c5641153059340fac91f256ed685a406f87e3f8df42905148256d0e1aa

Observation 557e7deb-6015-456f-a317-5f0ac27be47e · outbound

This paper cites Overcoming Exploration in Reinforcement Learning with Demonstrations.

Policy Improvement with Style-Specific Demonstrations Overcoming Exploration in Reinforcement Learning with Demonstrations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T19:21:12.663108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:21:12.663108Z digest=sha256:5024336bee85a33a3daba1d35587fa43b8260173585511acd01f8645d5b7ece3

Observation 405fb5cf-c2c2-462c-b2d0-aecb0765c448 · outbound

This paper cites Handbook on mahjong competition rules, 2016.

Policy Improvement with Style-Specific Demonstrations Handbook on mahjong competition rules, 2016

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:21:13.071936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:21:12.667661Z digest=sha256:ac3c37f18efb2cd489cd8e6e89abc185db176c1bed232e8c9a33b40fc7e2cd7b

Observation e96eacf4-8b1d-43e7-b269-6481250f5f57 · outbound

This paper cites an unresolved cited work.

Policy Improvement with Style-Specific Demonstrations Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:21:13.057541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:21:12.671994Z digest=sha256:b27930e541cb436ceba171b7c2ace442ec0d986c38927bdb5b5f514fb9ad1d12

Observation 330d45d5-fc97-4d21-a156-e61ef9e9d8b9 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Policy Improvement with Style-Specific Demonstrations High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T19:21:12.676241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:21:12.676241Z digest=sha256:e80e19faa0bcc9df2990a783ceb327cae13f91377428936c65c4483f03274962

Observation fc122948-32cc-43f1-beac-ac578771c7de · outbound

This paper cites Jordan, and Pieter Abbeel.

Policy Improvement with Style-Specific Demonstrations Jordan, and Pieter Abbeel

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:21:12.681331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:21:12.681331Z digest=sha256:09a2bd0f607906821f868f7d7d77a4c6290b9abb0419d2cea8f93fef761cfcc5

Observation 59357414-e1ef-41d6-b9dd-8afa77bcaad9 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Policy Improvement with Style-Specific Demonstrations Proximal Policy Optimization Algorithms

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T19:21:12.685791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:21:12.685791Z digest=sha256:70c9e6ef721cba15f384fd401d106b1db35496b3fb1abe30ab6908b58705c896

Observation efa39ccf-c4bd-4677-b3bb-215f03bc8770 · outbound

This paper cites Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm.

Policy Improvement with Style-Specific Demonstrations Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T19:21:12.690475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:21:12.690475Z digest=sha256:963df1ea8a860e159e0b25f5c81c90cb780b734d1a9f5c32830e886991081fd5

Observation dcd816df-9cbd-4f93-b20c-ac64ed283c8a · outbound

This paper cites Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards.

Policy Improvement with Style-Specific Demonstrations Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T19:21:12.694943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:21:12.694943Z digest=sha256:afdf20c2be32c27f555b8ae0ee2174b19084081b4880060a73c6f7da3d6e546d

Observation 3c7b3f68-6930-4007-bed4-bf885d102ae6 · outbound

This paper cites Czarnecki, Micha \"e l Mathieu, Andrew Dudzik, Junyoung Chung, David H.

Policy Improvement with Style-Specific Demonstrations Czarnecki, Micha \"e l Mathieu, Andrew Dudzik, Junyoung Chung, David H

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:21:13.032461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:21:12.699471Z digest=sha256:031d54ba1bf9af2287d8b000a5cf242825e206460f9743b183eea341c19f1499

Observation d1f718d7-9ee3-4d05-baca-b05b3d4cc908 · outbound

This paper cites Perfectdou: Dominating doudizhu with perfect information distillation, 2024.

Policy Improvement with Style-Specific Demonstrations Perfectdou: Dominating doudizhu with perfect information distillation, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:21:13.016806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:21:12.703785Z digest=sha256:28336e26b7c7be61b20bb73631504d7c57c55050b766570a0d60b5bf9cdd481f

Observation 173442c0-ff24-4094-8167-6d2242f30bb4 · outbound

This paper cites Towards Playing Full MOBA Games with Deep Reinforcement Learning.

Policy Improvement with Style-Specific Demonstrations Towards Playing Full MOBA Games with Deep Reinforcement Learning

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:21:12.795283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:21:12.708064Z digest=sha256:1474aadb68f93249fcc74f10316941c972562de3c408ce863a9c8d2c480c5fc9

Observation 302e7d29-979d-412d-a948-1fdda8081560 · outbound

This paper cites an unresolved cited work.

Policy Improvement with Style-Specific Demonstrations Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:21:13.002262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:21:12.712541Z digest=sha256:db388545920742527b017f31a489b1d8e9e50e4d8a322fb89ab90963657f646f

Observation 07e4e0b7-95b3-48af-a75a-5d9e1405fd81 · outbound

This paper cites Botzone: an online multi-agent competitive platform for ai education.

Policy Improvement with Style-Specific Demonstrations Botzone: an online multi-agent competitive platform for ai education

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:21:12.987785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:21:12.717186Z digest=sha256:8ed6dd697df4437344b67fb0f512e27e44735c47afe733f08ddd9da1d4a893c5

Observation 4c7b16fc-609f-494f-a652-43c6948cbd90 · outbound

This paper cites Learning Sparse Rewarded Tasks from Sub-Optimal Demonstrations.

Policy Improvement with Style-Specific Demonstrations Learning Sparse Rewarded Tasks from Sub-Optimal Demonstrations

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:21:12.774626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:21:12.721639Z digest=sha256:31661454ae259d757a8013fc19e76483287cb510742da0e2a60edd7d96ec9aad

Observation c2d288bc-147d-4240-a106-970f8d88c848 · outbound

This paper cites Behavior proximal policy optimization, 2023.

Policy Improvement with Style-Specific Demonstrations Behavior proximal policy optimization, 2023

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:21:12.973565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:21:12.726434Z digest=sha256:65ab492aa4c4120493d32382110f0881acd1116fc58c67920406c4aec0cd5f37

Observation affdf193-c327-4a13-a16a-86b1a70ef7f4 · outbound

This paper cites write newline.

Policy Improvement with Style-Specific Demonstrations write newline

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T19:21:12.730681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:21:12.730681Z digest=sha256:65015cb87b8844eac309cfccb6d0e235657d9cd035c56e5eb52180213657af59

Pith citing papers

No inbound Pith citation observations are available.