Pith. sign in

Paper Citation Record · LEDGER

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models

As of 5 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 2 inbound Pith citation observations for arXiv:2508.21365.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.21365 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:23:27.591640Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:34:01.106317Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact6
  • verified fuzzy6
  • unresolved27
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 3462f19f-6220-49a1-a503-049fa24058ce · outbound

This paper cites Cause and Effect: Can Large Language Models Truly Understand Causality?.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Cause and Effect: Can Large Language Models Truly Understand Causality?

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:23:28.247417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-08-05T14:23:23.843562Z digest=sha256:56bc5e68e9f9c70ddd1bb21245ddab245ba347db03daed8117c5eb0e5722d0ce

Observation 6039e9bf-ca8c-441f-97b9-581a2cd2d618 · outbound

This paper cites an unresolved cited work.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:23.929458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:23.929458Z digest=sha256:a66bd016d47edbf31a92d78062115e2f77cc16664fdda771ccf029609dd6ac47

Observation 9b2c6e68-646a-43be-92f7-971cf3f8ff64 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:24.020818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:24.020818Z digest=sha256:af0bf867d6d4f4d2b8602477aa86a24b81480711d15b02bcc06a655d55336193

Observation 65eeae8c-36e2-48ce-8939-8ceb51a810c7 · outbound

This paper cites MapEval: A Map-Based Evaluation of Geo-Spatial Reasoning in Foundation Models.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models MapEval: A Map-Based Evaluation of Geo-Spatial Reasoning in Foundation Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:24.072935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:24.072935Z digest=sha256:8d75ba0b98e205055718adcaa4cd6a5a550945852439a13f9b00e1fec0bf32b1

Observation 7a428987-6583-4222-9603-fb9898fd35bc · outbound

This paper cites Bayeschess: A computer chess program based on bayesian networks.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Bayeschess: A computer chess program based on bayesian networks

Reference 5

Resolution
verified exact
doi, observed 2026-08-05T14:23:28.080495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-08-05T14:23:24.156489Z digest=sha256:95e5a92c5dc1921382522f639dd884047a58a9e312b5d6cd1b5b262c5d8390ae

Observation d4b144ae-8637-4860-883c-01ca553fd71a · outbound

This paper cites Font and Tobias Mahlmann.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Font and Tobias Mahlmann

Reference 6

Resolution
metadata mismatch
raw_fallback, observed 2026-08-05T14:23:29.926244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-08-05T14:23:24.207525Z digest=sha256:f29a60e20d9e4951e850d04dd6848d9a4c0e6adfcc642f0de5547c9d64917205

Observation ce16c2fa-817c-4f36-9a6e-3bc437a6a1b5 · outbound

This paper cites Enabling self-improving agents to learn at test time with human-in-the-loop guidance.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Enabling self-improving agents to learn at test time with human-in-the-loop guidance

Reference 7

Resolution
verified exact
raw_fallback, observed 2026-08-05T14:23:29.665817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-08-05T14:23:24.271531Z digest=sha256:a4e98bab091369d82f97bec5f4faabe5dc1c72419ba489d0509ecb53e0063904

Observation 7bb921e6-e04c-48b4-83d7-f04eef122fb0 · outbound

This paper cites Measuring massive multitask language understanding.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Measuring massive multitask language understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:24.347757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:24.347757Z digest=sha256:180a348b38beeefe1d5384b051538bb47fc052d697a25f8a4f979d977882e99c

Observation 712a2b16-a84f-48c1-95db-a3c78fcec02a · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:24.434666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:24.434666Z digest=sha256:0cf7e07661f83f948fb079df396d37f447950f32a298496ec33a0ddf433b2e5c

Observation 4d97786b-a523-4a9f-a8cb-623b7b38872f · outbound

This paper cites A Survey on Large Language Model-Based Game Agents.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models A Survey on Large Language Model-Based Game Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:24.473717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:24.473717Z digest=sha256:cd8db0d38a19b2747661907b43eaba83a2e011c278d0480f6a5534cdcebc0085

Observation 37e15f8e-9e20-42f3-b515-e38236208f3f · outbound

This paper cites PokeLLMon: A Human-Parity Agent for Pokemon Battles with Large Language Models.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models PokeLLMon: A Human-Parity Agent for Pokemon Battles with Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:24.541597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:24.541597Z digest=sha256:93f5988fc91a653adeb7b51faa165160dfb27159bab9031a2d10d2453f669a3b

Observation b120537f-e788-4299-b64b-8bc3f74f543e · outbound

This paper cites C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:30.920211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-08-05T14:23:24.602432Z digest=sha256:14d76ab001c44ae4b192f649542982219032cf57b677eece038d824cdf1a8d6c

Observation 0a47fa39-e61c-4bd8-90a1-6ee078b5988d · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:24.667850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:24.667850Z digest=sha256:70bbd982a857f4d515c1c242993f563ad45f3e929de2127f0a1e04e2757d8780

Observation b8d24579-c921-478d-a8ca-613a937f29ae · outbound

This paper cites Mind the GAP: Glimpse-based Active Perception improves generalization and sample efficiency of visual reasoning.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Mind the GAP: Glimpse-based Active Perception improves generalization and sample efficiency of visual reasoning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:23:29.386562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-08-05T14:23:24.726570Z digest=sha256:13ffbcee63b0167b3fa0887a525ae3a170f9ba6f90ae81af3edd1b90a40cd6d2

Observation a7266d34-80bb-40ec-a956-b4b069ea27d6 · outbound

This paper cites School chinese benchmark, 2018.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models School chinese benchmark, 2018

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:30.727816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-08-05T14:23:24.798984Z digest=sha256:85d1ee8f5d7805f9506adbaa8218bb7c6a78bffc70ef1872cfc88b1260ee0182

Observation 4d0d153b-c670-4065-85c6-183d48962969 · outbound

This paper cites CodeI/O: Condensing Reasoning Patterns via Code Input-Output Prediction.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models CodeI/O: Condensing Reasoning Patterns via Code Input-Output Prediction

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:24.911042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:24.911042Z digest=sha256:0149feee3c2ed3817e044754ddeac7b52838ab08a373a751ef421493ea817c92

Observation d0e6f647-06ac-422b-be93-43d35f0c9b2e · outbound

This paper cites Simpo: Simple preference optimization with a reference-free reward.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Simpo: Simple preference optimization with a reference-free reward

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:24.975898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:24.975898Z digest=sha256:b5ede67c20765f1fcbb4d0b1818ab2b16eed8d6ead9d2083d964ea9e93760f09

Observation 8fe91552-3333-4fac-b23d-823e54708096 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Playing Atari with Deep Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:25.071210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:25.071210Z digest=sha256:dd8f835e111b8ea94308ba150341b547b4f2c5faeefb1d9c00ad6dabf235e02e

Observation 2ad2ca66-1586-4690-931c-d94db8641ab5 · outbound

This paper cites Creating pro-level AI for a real-time fighting game using deep reinforcement learning.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Creating pro-level AI for a real-time fighting game using deep reinforcement learning

Reference 19

Resolution
metadata mismatch
raw_fallback, observed 2026-08-05T14:23:29.160079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-08-05T14:23:25.134829Z digest=sha256:7852c4314881b4a358f75bef5f0b84db1898b9202f82c82234dd96db01156726

Observation 30ed4bd9-ea79-4473-95b0-6d637aab7769 · outbound

This paper cites Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, et al.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, et al

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:30.589098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-08-05T14:23:25.311280Z digest=sha256:e160d7bc1333015b9a03e898cbcb4fd04eb954efd35b1f70a8c74a593338967b

Observation cf1e4c3f-0971-4329-965f-2e2f5a9a7df6 · outbound

This paper cites Manning, Stefano Ermon, and Chelsea Finn.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Manning, Stefano Ermon, and Chelsea Finn

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:25.379171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:25.379171Z digest=sha256:bcf72588b1c2f2c1be22c8fe6be548d7f8f674dff3640f841d3c202934fc622d

Observation 95e7b3ed-3fcd-47db-89ce-a6f1e284b427 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Proximal Policy Optimization Algorithms

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:25.463390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:25.463390Z digest=sha256:ee42b183246e0dcda48d6852f80eb9a6513939b6ca135ac518ce06e6b814f4eb

Observation 6ad11b99-f8f3-4a63-b322-a6221a23776a · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:25.587387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:25.587387Z digest=sha256:4f92d416562e786b7471a3a91bf036aeb6f2415bf3f8f9ecfe8e29d1c1e9d56a

Observation 5678c3e7-8d6f-494f-b41c-ad52a791efca · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:25.667622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:25.667622Z digest=sha256:9f77cc5eb06c593e0d052dd5a766c67ef976453be217d571503971aa7f3572aa

Observation c3e1d71b-d402-466c-a341-5baf90e30292 · outbound

This paper cites Mastering the game of go with deep neural networks and tree search.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Mastering the game of go with deep neural networks and tree search

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:25.840989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:25.840989Z digest=sha256:e2d91e4046b80e60401cf27fc6b29c952da64fde1e9404d14604110d93a99430

Observation a01d5683-b46c-40b4-9454-fdbea02541dc · outbound

This paper cites Bayes' Bluff: Opponent Modelling in Poker.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Bayes' Bluff: Opponent Modelling in Poker

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:26.017512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:26.017512Z digest=sha256:9297f3ac434a29f8b1d02510287c28c697e7d3935ca21315304db12460f2fa8c

Observation 2eb77aa2-f1af-4a39-8bb1-bc87abb3f955 · outbound

This paper cites Brown, Adam Santoro, Aditya Gupta, et al.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Brown, Adam Santoro, Aditya Gupta, et al

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:30.402090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-08-05T14:23:26.131722Z digest=sha256:637c278c1a2ad9e2d7b9323027bb968ad8f01e2279591df4fc6423840947c754

Observation 8c0b0d61-6407-4193-b951-8fee35869ffe · outbound

This paper cites Meta-Reasoner: Dynamic Guidance for Optimized Inference-time Reasoning in Large Language Models.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Meta-Reasoner: Dynamic Guidance for Optimized Inference-time Reasoning in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:26.207323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:26.207323Z digest=sha256:c88629f6582bd56541f8d0fdb233fc3732b57d24e224355ec3d27c2d8d9310f2

Observation ad54ce16-f332-43d9-82dd-ddde84dc1c39 · outbound

This paper cites Le, Ed H.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Le, Ed H

Reference 29

Resolution
malformed identifier
no resolver link, observed 2026-08-05T14:23:26.320509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:26.320509Z digest=sha256:851ee4ef35b83311dab145ff217070bc51abf6efdcb3679702b43492d8ad7478

Observation 7cb96572-77c7-4a03-bca6-b78143d72d34 · outbound

This paper cites CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent Evaluation.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent Evaluation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:26.402477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:26.402477Z digest=sha256:bb127a61520794644adde699c8f8577513b0ab9661710da2a4f2faaabc0df337

Observation e1d6f15b-32c1-4415-95b7-cce9431820eb · outbound

This paper cites StarCraft II: A New Challenge for Reinforcement Learning.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models StarCraft II: A New Challenge for Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:26.486476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:26.486476Z digest=sha256:47513c0d8ba0e78b0094378faf941a626f6a3266be9341006e0ef454386b5054

Observation 6ec25b2f-cfd9-4b6c-a1f6-3fa2eeb994c4 · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:26.587398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:26.587398Z digest=sha256:be39977181b7780701aedd4e78dc54b928c476e6dd51de882b21d07f9d1cf7ac

Observation 6ef93714-881f-4b35-80d7-bf99c4f62fd2 · outbound

This paper cites Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:26.680923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:26.680923Z digest=sha256:9d0b9a81c41bfee8e296fe13804f21d749ca097598c5916ebfe7ff8b9fc8aa6b

Observation 1f538ca5-df5c-4728-8f1e-c26e3ae4beff · outbound

This paper cites Mind's Eye of LLMs: Visualization-of-Thought Elicits Spatial Reasoning in Large Language Models.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Mind's Eye of LLMs: Visualization-of-Thought Elicits Spatial Reasoning in Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:26.747722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:26.747722Z digest=sha256:d4e33ff035cab82c98404e3a4c9b8bfe63ac2c1a74ec1372c4620ef47c3e4f77

Observation e9e3842a-54be-4d53-beb0-b66881e83778 · outbound

This paper cites Agents Play Thousands of 3D Video Games.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Agents Play Thousands of 3D Video Games

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:26.797595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:26.797595Z digest=sha256:3d248c0f952117b0f5518342ceecd5e5312a7f91022b6b48f0ce3c0d8ca19733

Observation 5dbef2c2-1c83-443d-93f3-8b5e6f3c6e76 · outbound

This paper cites Policy-to-language: Train llms to explain decisions with flow-matching generated rewards.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Policy-to-language: Train llms to explain decisions with flow-matching generated rewards

Reference 36

Resolution
verified exact
raw_fallback, observed 2026-08-05T14:23:28.595131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-08-05T14:23:26.982264Z digest=sha256:4f988510eab15dd8b48acd0420789c8ee082629ab36cd47e670ac730b7824ad4

Observation e6705ad7-3bbe-4ea4-a259-aaef503742f2 · outbound

This paper cites Mastering complex control in moba games with deep reinforcement learning.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Mastering complex control in moba games with deep reinforcement learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:30.251013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-08-05T14:23:27.199232Z digest=sha256:af6c145d23bead2145b80b8555df2188b55de1712d50dda0de8ad4f4feae5940

Observation 34de5102-2fc4-47ed-a428-202845f6c721 · outbound

This paper cites C har P oet: A C hinese classical poetry generation system based on token-free LLM.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models C har P oet: A C hinese classical poetry generation system based on token-free LLM

Reference 38

Resolution
verified exact
doi, observed 2026-08-05T14:23:27.806325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-08-05T14:23:27.276998Z digest=sha256:0360d364c44f19e88fa16f0d06a04a72869fc4d9cc91e127ec0c30ae19aa67da

Observation 5a7d02ff-d198-40fd-88b0-c455de45c09f · outbound

This paper cites Training interactive agent in large fps game map with rule-enhanced reinforcement learning.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Training interactive agent in large fps game map with rule-enhanced reinforcement learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:30.097095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-08-05T14:23:27.367053Z digest=sha256:f064979ae9a9bd8814e25e6d98c6169c027a673d20b7615597ccf4ce55692f7f

Observation 0aeb653d-cc59-49f5-ac7c-a1e24532cf66 · outbound

This paper cites Ape210K: A Large-Scale and Template-Rich Dataset of Math Word Problems.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Ape210K: A Large-Scale and Template-Rich Dataset of Math Word Problems

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:27.444486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:27.444486Z digest=sha256:c263dcade0cfe3fa40cb116915d4df4b50c944dc36d1b9b2e2be9edf547b8a13

Observation 4007ec55-0812-434b-9b84-7b42004e4305 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Instruction-Following Evaluation for Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:27.530464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:27.530464Z digest=sha256:a02f5786159c490170a028fc9f28f5cb9f74cbac35b727bd0a953a8530a77ff8

Observation 9f49b6f6-5ae9-4788-8807-a56fc5849d5e · outbound

This paper cites PokerBench: Training Large Language Models to become Professional Poker Players.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models PokerBench: Training Large Language Models to become Professional Poker Players

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:27.591640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:27.591640Z digest=sha256:d4f5fb0b7f92b26f0ea1a4b00ce2e98f7f19a74d8d3836b5a8a22b5dd59925ee

Pith citing papers

Observation 682ba848-7a4b-4e77-a5d3-b9c69669f7e6 · inbound

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application cites this paper.

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models

Reference 288

Resolution
verified exact
arxiv_id, observed 2026-06-27T09:50:48.309200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T09:46:30.702256Z digest=sha256:d815aba946935a18b3aa235cbc7240e834114746e07a8344de1da3b0e77f4119

Observation f1e7d76b-4a07-474e-8ef8-18cc3018c509 · inbound

SyncPlan: Long-Horizon LLM Coordination with Explicit Synchronization and Adaptive Correction cites this paper.

SyncPlan: Long-Horizon LLM Coordination with Explicit Synchronization and Adaptive Correction Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:01.106317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:34:01.106317Z digest=sha256:49f2d952f6bf95e04dea4d4fae6a29b2286b16721cc743d5442f768dd1e40f6f