Pith. sign in

Paper Citation Record · LEDGER

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training

As of 10 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2509.06053.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.06053 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T04:36:02.619219Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact2
  • verified fuzzy11
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 76069711-87f4-4052-8c95-f8b4ad411aac · outbound

This paper cites Reinforcement learning in robotics: A survey.The International Journal of Robotics Research, 32(11):1238–1274, 2013.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Reinforcement learning in robotics: A survey.The International Journal of Robotics Research, 32(11):1238–1274, 2013

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T04:36:01.718482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:36:01.718482Z digest=sha256:fb765ae3b2827ac76524b7282d059ee115797f10d61f1c4ef0a5da69d186b8f1

Observation 8adbcd51-4595-448a-a98f-249794217e54 · outbound

This paper cites Sim-to-real robot learning from pixels with progressive nets.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Sim-to-real robot learning from pixels with progressive nets

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:36:03.881420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:36:01.757271Z digest=sha256:180f20e13ef5d40778e55279e67eef854fcb42cd0b2396b7c190182f38f67b24

Observation a4d57b60-a009-40e3-86e1-db4827171bc6 · outbound

This paper cites Google research football: A novel reinforcement learning environment.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Google research football: A novel reinforcement learning environment

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:36:03.873689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:36:01.838723Z digest=sha256:7651ef9b5a9b58ff1df1d83e56cc37b40849d544a4c19641d0cf54a6c3b36e8d

Observation 60035789-2acb-46d4-939c-dc77ac2ab32a · outbound

This paper cites A comprehensive review of multi-agent reinforcement learning in video games.IEEE Transactions on Games, 2025.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training A comprehensive review of multi-agent reinforcement learning in video games.IEEE Transactions on Games, 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:36:03.866070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:36:01.932070Z digest=sha256:58d910338714d441a0f837ede3b07d36b0278e7e0e5d775c910f5de3742c7286

Observation 7aa75364-9c6c-42ee-89bc-0c6262147ced · outbound

This paper cites Robustness and sample complexity of model-based marl for general-sum markov games.Dynamic Games and Applications, 13(1):56–88, 2023.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Robustness and sample complexity of model-based marl for general-sum markov games.Dynamic Games and Applications, 13(1):56–88, 2023

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:36:03.858739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:36:01.986521Z digest=sha256:5d3aaf7f98367bcc964fe754cac103a9ecfb01472fd2ab5f92c1cceb646bff7d

Observation d96436a9-e1dc-4079-80a9-0716bb5a3cd9 · outbound

This paper cites Multi-agent reinforcement learning for autonomous driving: A survey.CoRR, 2024.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Multi-agent reinforcement learning for autonomous driving: A survey.CoRR, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:36:03.850909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:36:02.046727Z digest=sha256:1d54766063d53b3c99f8650cadfecedeb384107eae3d814105c1df8a1361f01d

Observation cfcdfffe-0dfd-4d28-8cba-ca5f3d0de4ed · outbound

This paper cites Deep reinforcement learning for autonomous driving: A survey.IEEE transactions on intelligent transportation systems, 23(6):4909–4926, 2021.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Deep reinforcement learning for autonomous driving: A survey.IEEE transactions on intelligent transportation systems, 23(6):4909–4926, 2021

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T04:36:02.077841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:36:02.077841Z digest=sha256:2d67350ebc1ee6826f3ebb4d3a1528372522dd6230802d482360299fdf496c41

Observation 9825a211-cfb7-454e-98c8-5c2db3c02ff8 · outbound

This paper cites Equilibrium selection for multi-agent reinforcement learning: A unified framework.arXiv preprint arXiv:2406.08844, 2024.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Equilibrium selection for multi-agent reinforcement learning: A unified framework.arXiv preprint arXiv:2406.08844, 2024

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-08-05T04:36:03.079984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:36:02.132762Z digest=sha256:5d835708865cbabe9f9828568f5f64ce7a7a5c488612676939d8846b52459483

Observation 9987f2c2-1689-4a8e-8760-18d0312108a2 · outbound

This paper cites Emergent reciprocity and team formation from randomized uncertain social preferences.Advances in neural information processing systems, 33:15786–15799, 2020.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Emergent reciprocity and team formation from randomized uncertain social preferences.Advances in neural information processing systems, 33:15786–15799, 2020

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:36:03.837935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:36:02.184973Z digest=sha256:16c06209b10851bd26971f78d483b1cbc3f13a7879a5813d964b522667ec8cb1

Observation 32ae72e4-5c42-4870-9f13-c7450bd67df1 · outbound

This paper cites Taxai: A dynamic economic simulator and benchmark for multi-agent reinforcement learning.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Taxai: A dynamic economic simulator and benchmark for multi-agent reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:36:03.772790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:36:02.269117Z digest=sha256:27fb4f5baca52c09144da4fbae35c1fc184e157df7360659f2c679b6d2a3d38a

Observation eb2d90d7-842f-419e-8e66-3481a3a2bc57 · outbound

This paper cites A new approach to solving smac task: Generating decision tree code from large language models.arXiv e-prints, pages arXiv–2410, 2024.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training A new approach to solving smac task: Generating decision tree code from large language models.arXiv e-prints, pages arXiv–2410, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:36:03.676990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:36:02.321874Z digest=sha256:337ebfdd2f208e55e090a4dd1828108ea2242e8b8a009a908f9aea6173841c8a

Observation fee254cb-faf1-426e-88f5-9317392408f4 · outbound

This paper cites ADRD: LLM-Driven Autonomous Driving Based on Rule-based Decision Systems.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training ADRD: LLM-Driven Autonomous Driving Based on Rule-based Decision Systems

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-05T04:36:02.769162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:36:02.379064Z digest=sha256:50dd0bd3cb8058c82f1a1df055c7ea95cd223bb0635381461e335864bb43db35

Observation 9929797c-8d60-4e4f-b4bb-40cab986b4bd · outbound

This paper cites Language models speed up local search for finding programmatic policies.Transactions on Machine Learning Research, 20(X), 2024.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Language models speed up local search for finding programmatic policies.Transactions on Machine Learning Research, 20(X), 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:36:03.500218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:36:02.461787Z digest=sha256:73b84fa0c4476a50cd2fe9086f1dccde15d2d6fc30fd79ceb5f9d23aaa167d25

Observation 55979992-ad1f-4405-abe3-e55b1f3cd0d8 · outbound

This paper cites Synthesizing programmatic reinforcement learning policies with large language model guided search.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Synthesizing programmatic reinforcement learning policies with large language model guided search

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:36:03.328505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:36:02.507355Z digest=sha256:c968fd10249997cb3622f2f65199f9cd7528df4e9d236c892dfdb0d0b508eb92

Observation 8af0bf21-4f54-4731-88bc-7bfe7d8d2ad2 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T04:36:02.563015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:36:02.563015Z digest=sha256:8f1faf9244b5d149e921437b2c01342f4e51963b8248fa84513f3f4bbd251c07

Observation 261bd7e9-b332-4bd3-b498-d0d70c55ef7a · outbound

This paper cites React: Synergizing reasoning and acting in language models.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training React: Synergizing reasoning and acting in language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:36:03.210132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T04:36:02.619219Z digest=sha256:f57e60bd3fe3a692849c592b4182c81d10ed041f1769beb22c544208e8aadf4e

Pith citing papers

No inbound Pith citation observations are available.