Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T04:36:02.619219Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2509.06053.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T04:36:02.619219Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
16 of 16 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 76069711-87f4-4052-8c95-f8b4ad411aac · outbound
PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Reinforcement learning in robotics: A survey.The International Journal of Robotics Research, 32(11):1238–1274, 2013
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8adbcd51-4595-448a-a98f-249794217e54 · outbound
PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Sim-to-real robot learning from pixels with progressive nets
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a4d57b60-a009-40e3-86e1-db4827171bc6 · outbound
PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Google research football: A novel reinforcement learning environment
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 60035789-2acb-46d4-939c-dc77ac2ab32a · outbound
PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training A comprehensive review of multi-agent reinforcement learning in video games.IEEE Transactions on Games, 2025
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7aa75364-9c6c-42ee-89bc-0c6262147ced · outbound
PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Robustness and sample complexity of model-based marl for general-sum markov games.Dynamic Games and Applications, 13(1):56–88, 2023
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d96436a9-e1dc-4079-80a9-0716bb5a3cd9 · outbound
PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Multi-agent reinforcement learning for autonomous driving: A survey.CoRR, 2024
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cfcdfffe-0dfd-4d28-8cba-ca5f3d0de4ed · outbound
PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Deep reinforcement learning for autonomous driving: A survey.IEEE transactions on intelligent transportation systems, 23(6):4909–4926, 2021
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9825a211-cfb7-454e-98c8-5c2db3c02ff8 · outbound
PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Equilibrium selection for multi-agent reinforcement learning: A unified framework.arXiv preprint arXiv:2406.08844, 2024
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9987f2c2-1689-4a8e-8760-18d0312108a2 · outbound
PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Emergent reciprocity and team formation from randomized uncertain social preferences.Advances in neural information processing systems, 33:15786–15799, 2020
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 32ae72e4-5c42-4870-9f13-c7450bd67df1 · outbound
PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Taxai: A dynamic economic simulator and benchmark for multi-agent reinforcement learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation eb2d90d7-842f-419e-8e66-3481a3a2bc57 · outbound
PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training A new approach to solving smac task: Generating decision tree code from large language models.arXiv e-prints, pages arXiv–2410, 2024
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fee254cb-faf1-426e-88f5-9317392408f4 · outbound
PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training ADRD: LLM-Driven Autonomous Driving Based on Rule-based Decision Systems
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9929797c-8d60-4e4f-b4bb-40cab986b4bd · outbound
PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Language models speed up local search for finding programmatic policies.Transactions on Machine Learning Research, 20(X), 2024
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 55979992-ad1f-4405-abe3-e55b1f3cd0d8 · outbound
PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Synthesizing programmatic reinforcement learning policies with large language model guided search
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8af0bf21-4f54-4731-88bc-7bfe7d8d2ad2 · outbound
PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 261bd7e9-b332-4bd3-b498-d0d70c55ef7a · outbound
PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training React: Synergizing reasoning and acting in language models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
No inbound Pith citation observations are available.