Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T14:23:27.591640Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 2 inbound Pith citation observations for arXiv:2508.21365.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T14:23:27.591640Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T23:34:01.106317Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
42 of 42 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 3462f19f-6220-49a1-a503-049fa24058ce · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Cause and Effect: Can Large Language Models Truly Understand Causality?
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6039e9bf-ca8c-441f-97b9-581a2cd2d618 · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b2c6e68-646a-43be-92f7-971cf3f8ff64 · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65eeae8c-36e2-48ce-8939-8ceb51a810c7 · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models MapEval: A Map-Based Evaluation of Geo-Spatial Reasoning in Foundation Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a428987-6583-4222-9603-fb9898fd35bc · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Bayeschess: A computer chess program based on bayesian networks
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d4b144ae-8637-4860-883c-01ca553fd71a · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Font and Tobias Mahlmann
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ce16c2fa-817c-4f36-9a6e-3bc437a6a1b5 · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Enabling self-improving agents to learn at test time with human-in-the-loop guidance
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7bb921e6-e04c-48b4-83d7-f04eef122fb0 · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Measuring massive multitask language understanding
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 712a2b16-a84f-48c1-95db-a3c78fcec02a · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d97786b-a523-4a9f-a8cb-623b7b38872f · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models A Survey on Large Language Model-Based Game Agents
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37e15f8e-9e20-42f3-b515-e38236208f3f · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models PokeLLMon: A Human-Parity Agent for Pokemon Battles with Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b120537f-e788-4299-b64b-8bc3f74f543e · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0a47fa39-e61c-4bd8-90a1-6ee078b5988d · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8d24579-c921-478d-a8ca-613a937f29ae · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Mind the GAP: Glimpse-based Active Perception improves generalization and sample efficiency of visual reasoning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a7266d34-80bb-40ec-a956-b4b069ea27d6 · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models School chinese benchmark, 2018
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4d0d153b-c670-4065-85c6-183d48962969 · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models CodeI/O: Condensing Reasoning Patterns via Code Input-Output Prediction
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0e6f647-06ac-422b-be93-43d35f0c9b2e · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Simpo: Simple preference optimization with a reference-free reward
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fe91552-3333-4fac-b23d-823e54708096 · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Playing Atari with Deep Reinforcement Learning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ad2ca66-1586-4690-931c-d94db8641ab5 · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Creating pro-level AI for a real-time fighting game using deep reinforcement learning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 30ed4bd9-ea79-4473-95b0-6d637aab7769 · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, et al
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cf1e4c3f-0971-4329-965f-2e2f5a9a7df6 · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Manning, Stefano Ermon, and Chelsea Finn
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95e7b3ed-3fcd-47db-89ce-a6f1e284b427 · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Proximal Policy Optimization Algorithms
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ad11b99-f8f3-4a63-b322-a6221a23776a · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5678c3e7-8d6f-494f-b41c-ad52a791efca · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3e1d71b-d402-466c-a341-5baf90e30292 · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Mastering the game of go with deep neural networks and tree search
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a01d5683-b46c-40b4-9454-fdbea02541dc · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Bayes' Bluff: Opponent Modelling in Poker
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2eb77aa2-f1af-4a39-8bb1-bc87abb3f955 · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Brown, Adam Santoro, Aditya Gupta, et al
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8c0b0d61-6407-4193-b951-8fee35869ffe · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Meta-Reasoner: Dynamic Guidance for Optimized Inference-time Reasoning in Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad54ce16-f332-43d9-82dd-ddde84dc1c39 · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Le, Ed H
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cb96572-77c7-4a03-bca6-b78143d72d34 · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent Evaluation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1d6f15b-32c1-4415-95b7-cce9431820eb · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models StarCraft II: A New Challenge for Reinforcement Learning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ec25b2f-cfd9-4b6c-a1f6-3fa2eeb994c4 · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Voyager: An Open-Ended Embodied Agent with Large Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ef93714-881f-4b35-80d7-bf99c4f62fd2 · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f538ca5-df5c-4728-8f1e-c26e3ae4beff · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Mind's Eye of LLMs: Visualization-of-Thought Elicits Spatial Reasoning in Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9e3842a-54be-4d53-beb0-b66881e83778 · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Agents Play Thousands of 3D Video Games
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dbef2c2-1c83-443d-93f3-8b5e6f3c6e76 · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Policy-to-language: Train llms to explain decisions with flow-matching generated rewards
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e6705ad7-3bbe-4ea4-a259-aaef503742f2 · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Mastering complex control in moba games with deep reinforcement learning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 34de5102-2fc4-47ed-a428-202845f6c721 · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models C har P oet: A C hinese classical poetry generation system based on token-free LLM
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5a7d02ff-d198-40fd-88b0-c455de45c09f · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Training interactive agent in large fps game map with rule-enhanced reinforcement learning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0aeb653d-cc59-49f5-ac7c-a1e24532cf66 · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Ape210K: A Large-Scale and Template-Rich Dataset of Math Word Problems
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4007ec55-0812-434b-9b84-7b42004e4305 · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Instruction-Following Evaluation for Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f49b6f6-5ae9-4788-8807-a56fc5849d5e · outbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models PokerBench: Training Large Language Models to become Professional Poker Players
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 682ba848-7a4b-4e77-a5d3-b9c69669f7e6 · inbound
Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models
Reference 288
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f1e7d76b-4a07-474e-8ef8-18cc3018c509 · inbound
SyncPlan: Long-Horizon LLM Coordination with Explicit Synchronization and Adaptive Correction Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.