Pith. sign in

Paper Citation Record · LEDGER

Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 100 inbound Pith citation observations for arXiv:1712.01815.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1712.01815 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 100 of 146 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:17:01.988107Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1084
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0eaca5ad-6510-4b06-932c-d60eb3d17d03 · inbound

AI safety via debate cites this paper.

AI safety via debate Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:22:01.318969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-13T21:22:01.289451Z digest=sha256:ca047496cea8fe27c61cf2a2c1d24e728121934a72c8f0b57a64d6f9c544ef3e

Observation 7ee7e5c5-693f-44f8-8edc-2d9c24b21975 · inbound

Near-optimal Bayesian Solution For Unknown Discrete Markov Decision Process cites this paper.

Near-optimal Bayesian Solution For Unknown Discrete Markov Decision Process Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-25T20:11:11.531556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-25T20:08:03.014744Z digest=sha256:17ab71c3d9c12ff869c587a3f85d1d28c3d4d07c90950c09efe5852d4fbcece2

Observation fc83e076-d3cd-4d7e-9e6d-0e9a58688e09 · inbound

Inductive general game playing cites this paper.

Inductive general game playing Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-05-25T17:51:06.211247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-25T17:47:10.838972Z digest=sha256:0bc6fa309cf9b523ba52a900cae7256993b4c573b3c5391dda86bb9ba0897e36

Observation f0a6c56a-c69e-47ea-9569-3676dba578e4 · inbound

On Multi-Agent Learning in Team Sports Games cites this paper.

On Multi-Agent Learning in Team Sports Games Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-25T15:56:01.105162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-25T15:51:28.289206Z digest=sha256:fe7cbf7854bd2957fa954b3a5c6b7ebbdbe58efee267709925929a221f1ec4e7

Observation 5d96b3c8-3011-408c-8fc9-7a61619caf22 · inbound

Growing Action Spaces cites this paper.

Growing Action Spaces Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-25T13:25:52.372334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-25T13:24:59.682609Z digest=sha256:10fdfd8532e7e67d4722ad04370a26069973b0e5a2f3152ba5f4d094795194a8

Observation a640cb50-f99a-495d-8840-4f16cc9e329e · inbound

General Board Game Playing for Education and Research in Generic AI Game Learning cites this paper.

General Board Game Playing for Education and Research in Generic AI Game Learning Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-24T23:25:03.970949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-24T23:20:39.142086Z digest=sha256:e23dcf0454569785250972a49d354c42526dc9c0491e77f9f7c8fefa8a52142b

Observation 0abe9487-8397-4217-af11-baa2fe004dee · inbound

Solving Rubik's Cube with a Robot Hand cites this paper.

Solving Rubik's Cube with a Robot Hand Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 102

Resolution
verified exact
local_arxiv, observed 2026-05-15T09:38:28.895897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T09:38:28.621842Z digest=sha256:cc4d5b37036445e02be3475e8b90b1f8b27484323d3dfa18290417c83a250b88

Observation 0c02283a-22cb-4f6a-933e-a72bbff4bd6e · inbound

On the Measure of Intelligence cites this paper.

On the Measure of Intelligence Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:05:34.171922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T13:05:34.106110Z digest=sha256:abe2b78a1660f400c39f79950f8fafe3e150541c3b12605a08f1f2d0cad7bbba

Observation 8381a50b-b274-4684-9ff2-d1a1133fc4d7 · inbound

Generative Language Modeling for Automated Theorem Proving cites this paper.

Generative Language Modeling for Automated Theorem Proving Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-23T05:18:10.793298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T05:18:10.620262Z digest=sha256:e609f6176da52a893c49b21b9356a699b9df9c31b61614f79538149c8ee578d6

Observation 73302cc3-b4e0-40d2-86f1-59127324da1f · inbound

Language Models (Mostly) Know What They Know cites this paper.

Language Models (Mostly) Know What They Know Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 208

Resolution
verified exact
arxiv_id, observed 2026-05-10T15:42:47.476702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T15:42:47.274448Z digest=sha256:dad7cfdc85022ec7cd4532c7b8ae19efb89d1edc9dcb43e121805e0b4d66defd

Observation 5b094429-2807-4abd-bfe9-6fa891c7c3ef · inbound

Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling cites this paper.

Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 171

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:45:17.827643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T17:45:17.540282Z digest=sha256:3049ed8b1419482fd563ed6be30628bd652749d49938e09d925aa3b93abb1d57

Observation 5d1d5c01-630d-491d-9367-0b0ca6d33ad0 · inbound

Reasoning with Language Model is Planning with World Model cites this paper.

Reasoning with Language Model is Planning with World Model Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 130

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T01:49:28.964959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-17T01:49:28.796581Z digest=sha256:b03855b97ecaa87fb0288a0ba8da860f91c152a647831ebf2ff23aa74e61cad5

Observation bc59fdd5-f186-4d79-8d15-01b35dd647b3 · inbound

Reinforcement Learning with Foundation Priors: Let the Embodied Agent Efficiently Learn on Its Own cites this paper.

Reinforcement Learning with Foundation Priors: Let the Embodied Agent Efficiently Learn on Its Own Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-24T06:44:02.659097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-24T06:40:00.328012Z digest=sha256:95f88ed619883e9a4315de8a83519d01505b880beecfaf67219aa6bd713dfce7

Observation 82bf8c01-107d-4a33-bae1-912a5bf63842 · inbound

Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models cites this paper.

Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T23:26:04.520860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T23:26:04.445225Z digest=sha256:f9d74807a03bc83f3e412dda9624ce049a3bb2876f5f4ec3f46ebf9cb1b13e47

Observation 54a06b1a-4101-4aec-89aa-8058b0f629eb · inbound

Learning Interactive Real-World Simulators cites this paper.

Learning Interactive Real-World Simulators Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 56

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T02:15:18.500659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-16T02:15:18.265190Z digest=sha256:6ef476b87e1128b45a5ce8bf1a389eb9680c26205e982647e9ba505a68de1ac8

Observation 93501dce-6bb5-4f92-a496-7ee3a7ad24a9 · inbound

MTSpark: Enabling Multi-Task Learning with Spiking Neural Networks for Generalist Agents cites this paper.

MTSpark: Enabling Multi-Task Learning with Spiking Neural Networks for Generalist Agents Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T21:17:01.988107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:17:01.988107Z digest=sha256:2e3d99a0075032bb8c90dbcc5aad7bca9534b75f47701c89105292a16de841d3

Observation 02838376-9eab-42fe-a73f-7553b9b6589b · inbound

Exponential Speedups by Rerooting Levin Tree Search cites this paper.

Exponential Speedups by Rerooting Levin Tree Search Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T21:08:18.413650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:08:18.413650Z digest=sha256:ee8c20edb835d49a7a2d47012125922c837f63986b3f4cfbde27cb99fe79b131

Observation d6e5d3b2-1623-4529-a6c0-c9278cffbdb6 · inbound

Augmenting the action space with conventions to improve multi-agent cooperation in Hanabi cites this paper.

Augmenting the action space with conventions to improve multi-agent cooperation in Hanabi Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T19:51:19.986569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:51:19.986569Z digest=sha256:49c578d1c3bac2e8b99e01366bde900ed5c2164ce73b195aab9279a88b9e0fcf

Observation bb8e9ed8-20a7-4907-a60d-8687c5fc7c36 · inbound

Monte Carlo Tree Search based Space Transfer for Black-box Optimization cites this paper.

Monte Carlo Tree Search based Space Transfer for Black-box Optimization Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:12.875718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:09:12.875718Z digest=sha256:4ca24d05ff2c30df68665b3ebe106d3287e83dded22c89744830587f3a68c851

Observation 7dcbdc2d-a295-438d-ad05-1702f5c610a1 · inbound

Seed-CTS: Unleashing the Power of Tree Search for Superior Performance in Competitive Coding Tasks cites this paper.

Seed-CTS: Unleashing the Power of Tree Search for Superior Performance in Competitive Coding Tasks Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T14:02:49.125617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:02:49.125617Z digest=sha256:4a324bd17404de655f14deda6fa2d1855f564da6369b7ef3bb73386524033eb1

Observation 33e22423-7f97-4549-aba4-c0eb5d12a761 · inbound

RAG-Star: Enhancing Deliberative Reasoning with Retrieval Augmented Verification and Refinement cites this paper.

RAG-Star: Enhancing Deliberative Reasoning with Retrieval Augmented Verification and Refinement Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T13:42:49.630897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:42:49.630897Z digest=sha256:7b7a484068e643954711145ee5a6a23b9ab8ad984add94bc1e883691af7ce371

Observation 5ca98572-cd7b-402e-832b-2556fcb4081e · inbound

Ensembling Large Language Models with Process Reward-Guided Tree Search for Better Complex Reasoning cites this paper.

Ensembling Large Language Models with Process Reward-Guided Tree Search for Better Complex Reasoning Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T11:08:56.533394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:08:56.533394Z digest=sha256:1138589595a490cb2ab80ee9516cef63fa59d576a370648c75fb20a006441908

Observation fe3f1d92-ba3d-40d6-be20-3812083b914e · inbound

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning cites this paper.

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:05.125316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:05.125316Z digest=sha256:c6adc711517cce215844f9bfb7da52f092434bd948829dee72903d63c52dd346

Observation 6a231f84-78a8-49c5-bc27-805f2653b871 · inbound

Automating the Search for Artificial Life with Foundation Models cites this paper.

Automating the Search for Artificial Life with Foundation Models Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T05:11:36.105634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:11:36.105634Z digest=sha256:ebef76298bb1dea4f47b54b926443245049da9f6c9586ff16814726c03dcc82e

Observation 6ba3a3f3-7c77-478a-9912-b58697dd6e95 · inbound

NS-Gym: Open-Source Simulation Environments and Benchmarks for Non-Stationary Markov Decision Processes cites this paper.

NS-Gym: Open-Source Simulation Environments and Benchmarks for Non-Stationary Markov Decision Processes Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T19:52:42.047606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:52:42.047606Z digest=sha256:159e213d301a2f59a77252fc443e17aa161bea5e11ca53636f8219be2abb47c3

Observation 99a84f7f-9ae3-499d-913b-b7166dcb7dd7 · inbound

A Survey on Multi-Turn Interaction Capabilities of Large Language Models cites this paper.

A Survey on Multi-Turn Interaction Capabilities of Large Language Models Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-10T19:32:44.230387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:32:44.230387Z digest=sha256:2b9489000a95ad6fa0d5ef6505e2e403c8f7d126ad2d01ce4727ebd14b5d4d19

Observation 92a1232c-ee55-4e65-b270-8433a89c431c · inbound

AirRAG: Autonomous Strategic Planning and Reasoning Steer Retrieval Augmented Generation cites this paper.

AirRAG: Autonomous Strategic Planning and Reasoning Steer Retrieval Augmented Generation Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T19:27:46.327698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:27:46.327698Z digest=sha256:5835d7e8d5cd78678e9222f6ed1f3247935cc728a83d2b93921724a3d8b5c9fd

Observation c01dde67-806f-4de8-81c3-a6d3d5c47842 · inbound

Beyond the Sum: Unlocking AI Agents Potential Through Market Forces cites this paper.

Beyond the Sum: Unlocking AI Agents Potential Through Market Forces Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:30.315475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:04:30.315475Z digest=sha256:0cce07dbb16c5e03d6a3985832670074a67e546899a3cf67a6b64f64612eb2c3

Observation eeb9d69a-03ce-4f16-8622-951c2ef8f387 · inbound

Revisiting Rogers' Paradox in the Context of Human-AI Interaction cites this paper.

Revisiting Rogers' Paradox in the Context of Human-AI Interaction Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T19:53:42.711581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:53:42.711581Z digest=sha256:45784c560a479e3cc1cd61739404889946f0a37ec55e09d9bcd2a761a472899b

Observation f6ad1393-60cf-4f1d-bc29-dc0df14e6709 · inbound

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation cites this paper.

HEPPO-GAE: Hardware-Efficient Proximal Policy Optimization with Generalized Advantage Estimation Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-10T16:56:50.980566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:56:50.980566Z digest=sha256:f7bcd692c9a45812ee7fbf8842839fe41f50a4a9f9999f35f581761bdde948aa

Observation 7d8b1132-0942-45fb-bbf0-f08977538e45 · inbound

What if Eye...? Computationally Recreating Vision Evolution cites this paper.

What if Eye...? Computationally Recreating Vision Evolution Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T14:48:24.308054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:48:24.308054Z digest=sha256:b967293352e6ec2951a449a3aa9969ce93a13d86551adce4c973974ebf0535db

Observation e6bc3594-ea50-41aa-b4b8-b1c5686cd307 · inbound

Lipschitz Lifelong Monte Carlo Tree Search for Mastering Non-Stationary Tasks cites this paper.

Lipschitz Lifelong Monte Carlo Tree Search for Mastering Non-Stationary Tasks Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-09T18:23:05.665783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:23:05.665783Z digest=sha256:cf02d5e25ff4b807b5ea95699afb16ccf0d95907ceb3d9d7a5445f3445ce29a4

Observation 8f82252d-cc4f-4fde-8f0b-458d0bec97b8 · inbound

A Variational Inequality Approach to Independent Learning in Static Mean-Field Games cites this paper.

A Variational Inequality Approach to Independent Learning in Static Mean-Field Games Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-09T17:29:42.716190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:29:42.716190Z digest=sha256:61d0c029a7cb032d98e21efabd56acb04b3ba42fa21ab1ded6317ea5a7cf4ff1

Observation 71f9a99c-4ded-44b2-bf3f-f543e487623b · inbound

Develop AI Agents for System Engineering in Factorio cites this paper.

Develop AI Agents for System Engineering in Factorio Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T15:09:20.573488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:09:20.573488Z digest=sha256:3e2d9f873034937597e58b0b0921a633315c1fb3a0405a3862919ea2e999b2d6

Observation 42f2bac0-71e6-4c31-847b-d3c222489a9a · inbound

CH-MARL: Constrained Hierarchical Multiagent Reinforcement Learning for Sustainable Maritime Logistics cites this paper.

CH-MARL: Constrained Hierarchical Multiagent Reinforcement Learning for Sustainable Maritime Logistics Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T13:35:05.249257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:35:05.249257Z digest=sha256:3b1719252bac0d25be4cbe767592af0e776a626659533cd82abe0291fe289a78

Observation b2531ecd-f9c3-43ff-b4dd-243733311256 · inbound

Synthesis of Model Predictive Control and Reinforcement Learning: Survey and Classification cites this paper.

Synthesis of Model Predictive Control and Reinforcement Learning: Survey and Classification Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T13:17:17.462579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:17:17.462579Z digest=sha256:d94ccd8dc304ccd578707f15706109461043a3a5d080268c3203bb1771dea8bb

Observation b4b80ebf-5a1e-400e-8a2b-dc6c33c4488a · inbound

LIMO: Less is More for Reasoning cites this paper.

LIMO: Less is More for Reasoning Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 89

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T02:11:37.593984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-17T02:11:36.932541Z digest=sha256:7b44bbeeeaf21ad1beabbe1660493a422651a5e1b2f652dd78f401c554629979

Observation e3f38b74-f831-47e1-a2a6-91f7c19c4fca · inbound

Beyond Interpolation: Extrapolative Reasoning with Reinforcement Learning and Graph Neural Networks cites this paper.

Beyond Interpolation: Extrapolative Reasoning with Reinforcement Learning and Graph Neural Networks Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T00:33:49.251920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T00:33:49.251920Z digest=sha256:f9b66846553bedb442dfe84df18738093ae21de06f8dbad2dd86e7d22208c5c4

Observation 13770f35-3892-417f-b92e-e4aff51cea6e · inbound

Holistically Guided Monte Carlo Tree Search for Intricate Information Seeking cites this paper.

Holistically Guided Monte Carlo Tree Search for Intricate Information Seeking Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T21:40:26.896146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:40:26.896146Z digest=sha256:15e6c33974ccc2a76d21e7adfd4689e9a1d6775181e7b6b57784e64bf81c468e

Observation a36ca27a-c381-46e4-a85b-a0960fc6b2fa · inbound

Probabilistic Artificial Intelligence cites this paper.

Probabilistic Artificial Intelligence Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T20:51:10.107804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:51:10.107804Z digest=sha256:b91ae8a8efe2cc57c4d768830c178a9d3c0715e00cad16c5e10e5046f10fb70b

Observation 2ac05b3c-09a7-4cc2-8855-cce582f6580b · inbound

LLMs Can Teach Themselves to Better Predict the Future cites this paper.

LLMs Can Teach Themselves to Better Predict the Future Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T20:20:35.334579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:20:35.334579Z digest=sha256:57d4880723b899770bfec9641e9889c92819f893bc59c03c67e1f821ea8cbb9d

Observation 8ffa25c5-17da-454a-801f-0ae3455dcae6 · inbound

Two-Player Zero-Sum Differential Games with One-Sided Information cites this paper.

Two-Player Zero-Sum Differential Games with One-Sided Information Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T19:53:25.780992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T19:53:25.780992Z digest=sha256:f39afd574c4024cab3e845fef4ff0269fb7fd78ed0eb130e2c938c997a1ecd16

Observation 51fffe79-a280-413d-b851-57e036a9be1b · inbound

Policy Guided Tree Search for Enhanced LLM Reasoning cites this paper.

Policy Guided Tree Search for Enhanced LLM Reasoning Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:31.870961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:20:31.870961Z digest=sha256:b4ef707fcf80ca6af9eeefeab44aa91ccad4eefea62fcd3629a5005bad8e9ad0

Observation 058bfed7-c6c2-4d82-ac31-7778c2ce6739 · inbound

We Can't Understand AI Using our Existing Vocabulary cites this paper.

We Can't Understand AI Using our Existing Vocabulary Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:46.058042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:17:46.058042Z digest=sha256:e3b62cb0f9efe2e57581ea4a2088335099339709bf099254adb9659c23b43cf5

Observation 1d75d4da-4c03-46e7-a36e-48077c96761a · inbound

$\texttt{SEM-CTRL}$: Semantically Controlled Decoding cites this paper.

$\texttt{SEM-CTRL}$: Semantically Controlled Decoding Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T01:32:21.677175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:080c6819dbedc3936fc42c74fc3f51703346b812d397fb2d3d254b29b1127a9f

Observation bec547c2-e15e-4994-9ade-b935f7dcc8b1 · inbound

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding cites this paper.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:40:06.528100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:16f07e5a1ea4a4a456aa838b6e6f314a3faade8c6ec695323abb855bb75744d8

Observation 9d420f29-727e-4b67-8ded-c93a1f5a20f8 · inbound

Scalable Multi-Task Learning through Spiking Neural Networks with Adaptive Task-Switching Policy for Intelligent Autonomous Agents cites this paper.

Scalable Multi-Task Learning through Spiking Neural Networks with Adaptive Task-Switching Policy for Intelligent Autonomous Agents Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-22T19:47:01.203155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T19:45:07.172201Z digest=sha256:c7865dbe3eea33ba555704405024fd55120a03094c24c12a2f2704df98370cdf

Observation 143ce6d6-de68-44ed-84b8-5b90c3a4a49a · inbound

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models cites this paper.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:23.541553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:23.541553Z digest=sha256:fcce14684bc4df193ce5ed40497ab4cb38551df56d826e28902e21b9256025bc

Observation fb50a7b3-b697-4853-8960-64934c712f57 · inbound

Large Language Models for Planning: A Comprehensive and Systematic Survey cites this paper.

Large Language Models for Planning: A Comprehensive and Systematic Survey Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 221

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:05.275950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:05.275950Z digest=sha256:8cef56d32fa2a862cf9739f6610adfc07b48832f856d3c03b4d4df50e6d2133a

Observation ce2edbaa-5f35-45f9-8e1e-649dd05b8711 · inbound

A Framework for Adversarial Analysis of Decision Support Systems Prior to Deployment cites this paper.

A Framework for Adversarial Analysis of Decision Support Systems Prior to Deployment Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:42.027095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:42.027095Z digest=sha256:3f1ee8e9a7d7d0b607ad26621ebedefc4d6593cce1f185c3942bd6f1a015aacf

Observation f9703d2b-f199-4dd6-8544-44b9b91e86e5 · inbound

Decomposing Elements of Problem Solving: What "Math" Does RL Teach? cites this paper.

Decomposing Elements of Problem Solving: What "Math" Does RL Teach? Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:28.844213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:05:28.844213Z digest=sha256:8b07b9489be01af44c4ec2a7bd14106ea25af801f6fdba566c79b07ecd859b83

Observation bd313a9f-5794-4dbb-9f22-8b1be12cc3de · inbound

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training cites this paper.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:28.836184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:28.836184Z digest=sha256:65be147f93383c2585e7a1a90d290d3b6d921f9f83eba379346437cfe8e74900

Observation 9c343496-eb31-43df-98bc-4424873161c5 · inbound

Bregman Centroid Guided Cross-Entropy Method cites this paper.

Bregman Centroid Guided Cross-Entropy Method Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:02.041298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:02.041298Z digest=sha256:b3d6ec285672fcb52ca55b8f3a95e4aee362b63bce3daae92885e190daa4236f

Observation 419ef1f1-a177-49e1-8a10-5c862632ae18 · inbound

Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening cites this paper.

Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:02.529238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:31:02.529238Z digest=sha256:f1cd9a8d24631a73ffa46257cf11e96ebb1dd468c45b11450a464e0cb2d9fdf3

Observation eb642055-069f-49a9-a58c-65798539348c · inbound

LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning cites this paper.

LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:35.772152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:36:35.772152Z digest=sha256:f5e7efa8aa8b019ddcdc315212ec1a15b2ed24e0cea9a62d2e9e33f9f944beb5

Observation 96b5d1a6-42ac-479d-9522-b113d999f338 · inbound

Learning to Plan via Supervised Contrastive Learning and Strategic Interpolation: A Chess Case Study cites this paper.

Learning to Plan via Supervised Contrastive Learning and Strategic Interpolation: A Chess Case Study Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:37:34.688532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:37:34.688532Z digest=sha256:6f631cd2d9f700f34b0d6c6d8bf5ed255fe68d4c8647c0e35545c49a345b52d9

Observation 1c712f5b-949f-4280-a6ee-3b068b91deb3 · inbound

Boosting LLM Reasoning via Spontaneous Self-Correction cites this paper.

Boosting LLM Reasoning via Spontaneous Self-Correction Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 2008

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:30.719001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:30.719001Z digest=sha256:c0492545ae9b45b06f672d6729216b9a847e5cdd2b9d4d03a733e27897918cfc

Observation c8f2eee2-1bb1-44aa-bead-6ad9de2b6459 · inbound

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search cites this paper.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:24.323361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:24.323361Z digest=sha256:6a7f56bbfd8fc25a39c663bb8f746e2e608789025f6946e43c47b7efa566cd25

Observation 99368f62-42ff-4941-b9c3-b2d8ea26e4b2 · inbound

Data-Driven Policy Mapping for Safe RL-based Energy Management Systems cites this paper.

Data-Driven Policy Mapping for Safe RL-based Energy Management Systems Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:43.148148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:43.148148Z digest=sha256:7095b550d3d13a0184c8127c23795807ae6ac4873a8b522530b0e94943a051d6

Observation dd2ea97c-aa52-4054-999b-1fa2e8176f5e · inbound

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind cites this paper.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.387074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.387074Z digest=sha256:c34f50b0ba36295fc30be19c6cc24a71163730bd6d9a86239db4cac1773b941f

Observation 3d9dc32f-db41-44a7-8c75-c83962b8fee6 · inbound

A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search cites this paper.

A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:07:39.615289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:07:39.615289Z digest=sha256:764fa51d4c58a24d96a17c8415d2591ae7dc28f053f7223f415aaf2fb7c88ee9

Observation d5e27596-3d6c-4b64-9aa1-59f9f7e7627a · inbound

Gym4ReaL: A Suite for Benchmarking Real-World Reinforcement Learning cites this paper.

Gym4ReaL: A Suite for Benchmarking Real-World Reinforcement Learning Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T21:26:50.406662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:26:50.406662Z digest=sha256:cf89be454377a617398601dfc36feb839af02cc793b36548ca5b4217e9acae3f

Observation 9e9e4522-fbce-4ae5-ac6d-afa829209472 · inbound

Can Large Language Models Develop Strategic Reasoning? Post-training Insights from Learning Chess cites this paper.

Can Large Language Models Develop Strategic Reasoning? Post-training Insights from Learning Chess Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:13:15.204837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:13:15.204837Z digest=sha256:4868baacd6e82ececd0a0c506d4fb2510698d4392438d0438af50829c6e5bac5

Observation deed1063-266c-4cc0-a404-41faaf89ca6f · inbound

Reinforcement Learning for Automated Cybersecurity Penetration Testing cites this paper.

Reinforcement Learning for Automated Cybersecurity Penetration Testing Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:31:51.321684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:31:51.321684Z digest=sha256:6e20a383ad3a3862a41eac12ea78078a25a9b1e592faf4eeb203e2889b6b98b6

Observation edcbd904-06a2-410c-809f-67344d3ac2e1 · inbound

Partial Label Learning for Automated Theorem Proving cites this paper.

Partial Label Learning for Automated Theorem Proving Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:18:25.796266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:18:25.796266Z digest=sha256:147fd9c64c6bbeb2f75ee9f390bd5ab5b7148b022c095102513db2fc6f661db6

Observation f7ffe987-480b-4136-8cc5-1641a34a07b1 · inbound

Critique of World Model cites this paper.

Critique of World Model Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T19:42:27.951210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:42:27.951210Z digest=sha256:155dafa4e052602c6bce9eb76794e725385bf93c29f2272a88ad1c51d36c9209

Observation 22e431cc-f1da-48bd-92ef-a9c77a23ed9a · inbound

Artificial Generals Intelligence: Mastering Generals.io with Reinforcement Learning cites this paper.

Artificial Generals Intelligence: Mastering Generals.io with Reinforcement Learning Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:58:36.167097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:58:36.167097Z digest=sha256:cda6bc1423a4920122c7265d81266a1e16b4fad5f8671fe841d0f527e81eb5cb

Observation c36e4d87-d47b-44e8-8317-feb967e0fa88 · inbound

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities cites this paper.

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T16:34:25.194886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:34:25.194886Z digest=sha256:6ef8d6b86c81ed25317b19a96f5b1423c7da9b16f70e05f5d45f8379fdd0efa4

Observation d31759d9-4456-4e96-82a5-6807f12a6c34 · inbound

What Does it Mean for a Neural Network to Learn a "World Model"? cites this paper.

What Does it Mean for a Neural Network to Learn a "World Model"? Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:46.361511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:46.361511Z digest=sha256:ede3a4a0fa0a4ee36ee8abf4db66bcbf5a652e2d3616978c3c0c2fb5792d2e2e

Observation 774b86a0-9dd7-4312-9234-e4225be897b5 · inbound

General Agentic Planning Through Simulative Reasoning with World Models cites this paper.

General Agentic Planning Through Simulative Reasoning with World Models Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-22T12:41:33.535741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T12:37:10.956233Z digest=sha256:099f7bae72b62eb7cefc28c7c48c0ac8d526a9efe20d9174b8a8a8198c5e61c4

Observation f6ecdc91-0157-44c1-a375-e8a38252c7e3 · inbound

Tail-Risk-Safe Monte Carlo Tree Search under PAC-Level Guarantees cites this paper.

Tail-Risk-Safe Monte Carlo Tree Search under PAC-Level Guarantees Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T23:22:29.446275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:22:29.446275Z digest=sha256:6b372259dead5530e6518144bce2333f2e63ff5f1a4a5d0b177a2a6977dfc23e

Observation 562758e3-95f4-4845-be47-56ffaeeb86ab · inbound

Hybrid Physics-Machine Learning Models for Quantitative Electron Diffraction Refinements cites this paper.

Hybrid Physics-Machine Learning Models for Quantitative Electron Diffraction Refinements Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:27.677046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:06:27.677046Z digest=sha256:da648d320bf1c66c702e6c07b6213bb6baf9487f19714a7abcec7bb50d8209d3

Observation 504cf053-eea8-41b3-8015-1c90d66a0ba9 · inbound

The Fair Game: Auditing & Debiasing AI Algorithms Over Time cites this paper.

The Fair Game: Auditing & Debiasing AI Algorithms Over Time Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 117

Resolution
unresolved
no resolver link, observed 2026-08-05T22:45:20.775630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:45:20.775630Z digest=sha256:1f5838eeaa8b1403f2809c9d338b83f9088cd19b3d036d329e546b578c72ef1c

Observation 783b6d4e-df78-4688-b664-bf61f0902eba · inbound

Evolutionary Optimization of Deep Learning Agents for Sparrow Mahjong cites this paper.

Evolutionary Optimization of Deep Learning Agents for Sparrow Mahjong Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T22:06:19.698245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:06:19.698245Z digest=sha256:c04ec8b9f77371eac157f1f015fbc98e6b30481d74bf949b285b5c2454d3eaaa

Observation 14e5daec-6c9c-4667-8f5b-90fb8fe6d078 · inbound

Edge General Intelligence Through World Models and Agentic AI: Fundamentals, Solutions, and Challenges cites this paper.

Edge General Intelligence Through World Models and Agentic AI: Fundamentals, Solutions, and Challenges Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-05T21:02:04.453521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:02:04.453521Z digest=sha256:a36b7535067c646e61ab25f5239d1e24e566e0c729172b1aa2dfd9e12b54ae13

Observation cde19674-1179-4713-ad61-86d99b27d526 · inbound

Mirage or Method? How Model-Task Alignment Induces Divergent RL Conclusions cites this paper.

Mirage or Method? How Model-Task Alignment Induces Divergent RL Conclusions Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T14:35:51.350639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:35:51.350639Z digest=sha256:6d867c575352142bb9779c5e6b3d5c89bb95a7ab54b1b1a97cbd3116824fcffb

Observation b65bf5b5-13a2-49c3-8c8f-d92500571d16 · inbound

Scalable Option Learning in High-Throughput Environments cites this paper.

Scalable Option Learning in High-Throughput Environments Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 58

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T20:06:49.728038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-18T20:04:58.064472Z digest=sha256:cfea5f8a0a6df75bdf5d5381799ef10d4186b5aa96240b5a0dd83285f9550748

Observation e3d5166e-e192-4b50-9afd-3bbb249ff6f5 · inbound

TransZero: Parallel Tree Expansion in MuZero using Transformer Networks cites this paper.

TransZero: Parallel Tree Expansion in MuZero using Transformer Networks Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T16:58:12.987074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:58:12.987074Z digest=sha256:aba5ddb6deed1e9a20e84a975d83579fc169ec881b06d9fc2a24cfb28928e372

Observation 525fa1de-ba25-4a75-9403-e531e4b1fd88 · inbound

Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs cites this paper.

Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T17:57:52.477057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:57:52.477057Z digest=sha256:c02c87d6ae275323152d3f98a4610544561690f3751e0368f18d6420bc592109

Observation 72b104ca-c66f-4952-b4b4-a77ff7925207 · inbound

Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning cites this paper.

Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-21T21:34:22.300517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-21T21:33:40.229376Z digest=sha256:dba48422a9c9106218f9965bc0ff0b290834b9aa5ac7ca04c8de6142e09c8ef1

Observation 353cde8c-65a4-435e-b4be-b5577287cd77 · inbound

People use fast and flat simulation to reason about new games cites this paper.

People use fast and flat simulation to reason about new games Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T10:14:17.110418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:14:17.110418Z digest=sha256:853794be54d3e46b5080db94bd6abc28c17dc0654ea3fae45a32605bc70af189

Observation 7e257f57-fbcd-40c5-839b-f9aeb7887c89 · inbound

Optimal control of the future via prospective learning with control cites this paper.

Optimal control of the future via prospective learning with control Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:05:25.326088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T23:04:54.956095Z digest=sha256:1c1c7d2e37058b87ab377c237e8fb178100556bb533678e9977968490a4af8bf

Observation 7410afe3-f7c1-4cd2-b1c9-81ad21f0280e · inbound

Olmo 3 cites this paper.

Olmo 3 Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T21:31:16.815834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:619831be92cc5d9a0a4a8cb8327889b16525a507bbc477155a927381a1831d3f

Observation 1a2b69ba-8b4f-4e69-b893-e00b80a8a1c1 · inbound

Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making cites this paper.

Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:11:16.802600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T21:10:38.484582Z digest=sha256:e1d4bedd4435a8ba81cf7cc4c399390f3bfe768c962ccfce36572c4a09f3d44c

Observation 9cc4fbe1-2280-4386-8fd8-d51d67c38b86 · inbound

Toward Training Superintelligent Software Agents through Self-Play SWE-RL cites this paper.

Toward Training Superintelligent Software Agents through Self-Play SWE-RL Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-21T16:10:20.288129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-21T16:07:48.570995Z digest=sha256:56fc53705ceb4e6e78cba20ffdff7207a4b8e8b07b10a8a509a7e12d98e53295

Observation 718bc116-399b-4cb8-8ebc-c01cdc475e14 · inbound

Toward Training Superintelligent Software Agents through Self-Play SWE-RL cites this paper.

Toward Training Superintelligent Software Agents through Self-Play SWE-RL Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-03T15:02:13.565152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:02:13.565152Z digest=sha256:bf8027e1e8ba705de494f143b6fc4c863e0dd0f2e45df3f8006c1b3bd9772e02

Observation 2e6e7293-f947-406f-81aa-c91cd60daeb6 · inbound

Safety Alignment of LMs via Non-cooperative Games cites this paper.

Safety Alignment of LMs via Non-cooperative Games Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:24.762257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:24.762257Z digest=sha256:3f0ab6e8ffab0f753bed3d7ea3b15f95f719a25cd9392d73e03077d6778689f2

Observation c278cb97-5e22-497d-89f4-23db7ec1a477 · inbound

Multi-agent DRL-based Lane Change Decision Model for Cooperative Platooning in Mixed Traffic cites this paper.

Multi-agent DRL-based Lane Change Decision Model for Cooperative Platooning in Mixed Traffic Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T10:04:42.730729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:04:42.730729Z digest=sha256:4334eb34343afcca98915f226df5fa73ddcb63327428aec3bdc6b1cc94db4c46

Observation 1fe87430-de22-42e4-abfb-0a787823317f · inbound

Output-Space Search: Targeting LLM Generations in a Frozen Encoder-Defined Output Space cites this paper.

Output-Space Search: Targeting LLM Generations in a Frozen Encoder-Defined Output Space Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-03T07:08:33.035932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:08:33.035932Z digest=sha256:9fc07793bac2eaa88b4cfbce244fc2dfa693e2c8b951d7c0c195a1c36374503e

Observation 32557d67-3451-44be-8156-77395016f9c0 · inbound

CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning cites this paper.

CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T05:14:20.763891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:14:20.763891Z digest=sha256:3171a3717f08e5314739ceb6654678262261646e72527086fbd6c1684ec9d4db

Observation 3e56cbe2-06ef-401e-ac4c-1ef5b24d998a · inbound

Self-Supervised Bootstrapping of Action-Predictive Embodied Reasoning cites this paper.

Self-Supervised Bootstrapping of Action-Predictive Embodied Reasoning Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:04:11.930844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-21T14:03:48.795572Z digest=sha256:510b18a8acf88d3458a35b6ecb4e5468dac0c3b276cc864d390b658f5fec114f

Observation 349ee6f0-86de-41a1-8875-d019a4a94633 · inbound

Multi-agent imitation learning with function approximation: Linear Markov games and beyond cites this paper.

Multi-agent imitation learning with function approximation: Linear Markov games and beyond Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-02T20:42:07.940055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:42:07.940055Z digest=sha256:718f8d43de0e344e63a6dcfb4283e6d1ae989c4c2412716cba7dc5760cde3a23

Observation 06ebd88d-ee9c-40f6-a8db-7f3cc2982c46 · inbound

Quantum entanglement provides a competitive advantage in adversarial games cites this paper.

Quantum entanglement provides a competitive advantage in adversarial games Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:17.042117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:17.042117Z digest=sha256:3181e5857d6b7152f0be430e865a6c7c6d0997a503d55b80374958809ab3db41

Observation f58c790d-9f7d-4d1b-acad-e30429f8a847 · inbound

Computer Architecture's AlphaZero Moment: Automated Discovery in an Encircled World cites this paper.

Computer Architecture's AlphaZero Moment: Automated Discovery in an Encircled World Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:51:22.944814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T02:19:33.633470Z digest=sha256:44d35ad1c2f3051f283ffe10f7336b727c7ee7c57304f2a25e8fa5d2590d3f7f

Observation 0c3e1e3d-00f1-477e-bd1b-c30f092d37ea · inbound

Probabilistic Language Tries: A Unified Framework for Compression, Decision Policies, and Execution Reuse cites this paper.

Probabilistic Language Tries: A Unified Framework for Compression, Decision Policies, and Execution Reuse Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:22:58.916568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-14T21:21:47.095517Z digest=sha256:2d90e9a708777d256dd90f3b0d54c02c694ec02b068935524e6c76b1f96e8606

Observation 7e16cd7d-a8a1-487e-91e0-5dd59846ac29 · inbound

Advantage-Guided Diffusion for Model-Based Reinforcement Learning cites this paper.

Advantage-Guided Diffusion for Model-Based Reinforcement Learning Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:01:00.636827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T17:21:03.813720Z digest=sha256:30c522aa59eeb67028abc19c3412c0c3a1ce5a4ef9655bec4ed0359e9a80bf98

Observation 0b6162e5-857b-4a59-a128-e27154cf945a · inbound

AdverMCTS: Combating Pseudo-Correctness in Code Generation via Adversarial Monte Carlo Tree Search cites this paper.

AdverMCTS: Combating Pseudo-Correctness in Code Generation via Adversarial Monte Carlo Tree Search Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:40:57.429344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T16:35:16.056397Z digest=sha256:d2257b69b63b46fd4967dac19d71f5938660bcf91b90f9f74db68fa455115ee6

Observation 2324b110-0069-4883-ba49-ca862400a268 · inbound

AlphaCNOT: Learning CNOT Minimization with Model-Based Planning cites this paper.

AlphaCNOT: Learning CNOT Minimization with Model-Based Planning Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:10:25.759222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:10:05.119477Z digest=sha256:9f14c4e27b9cc556860d012f1a2c90c2960c097cc2abecec685567a341aad435

Observation 5d0a9622-d16a-48a0-bd6e-9fc7119fd27b · inbound

PAWN: Piece Value Analysis with Neural Networks cites this paper.

PAWN: Piece Value Analysis with Neural Networks Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T10:55:04.113495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T10:50:44.608469Z digest=sha256:518e510c120abe6b955d287727be2f7fc8699ed213bd6a3d8f0582817dcfe6d7

Observation 632febf0-5ba1-4a6b-ac89-72ce804b4ea6 · inbound

Causal inference for social network formation cites this paper.

Causal inference for social network formation Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:11:03.925868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T04:08:05.770405Z digest=sha256:7a7a92c80fd44b29aaa1fe00a6384355ea8f93d8e32e876394090a065a38913c