Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T20:30:30.167398Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 1 inbound Pith citation observation for arXiv:2508.10428.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T20:30:30.167398Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-26T21:02:26.081166Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T00:39:17.301001Z
32 of 32 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 402af738-96bf-48ea-a77c-3c528f4d0aa5 · outbound
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks , " * write output.state after.block = add.period write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcfd4c6f-7986-4566-bfbd-09aa7febcc19 · outbound
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks write newline
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ab399f5-2654-49db-b347-e4bfaeaa3917 · outbound
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks GPT-4 Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b936859-190d-4248-bc2a-732f483af256 · outbound
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks K.; and Johnson, B
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d7755b31-5bda-4475-bf0b-d9d04ef19bcf · outbound
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks DeepSeek-V3 Technical Report
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1b8d12e-1e9a-43ae-abf7-463dafdf8255 · outbound
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3304b1e-6bc5-4952-b5d3-00963259cf41 · outbound
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks L.; Yao, S.; Chen, Y.; Shen, P.; Yu, H.; Zhang, H.; Zhang, X.; Dong, Y.; et al
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 36a70c36-a638-4ce1-a441-ead97bbd793d · outbound
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks LLM-PySC2: Starcraft II learning environment for Large Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 505180f6-239f-4f55-a4c2-653ea13f0b53 · outbound
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks AvalonBench: Evaluating LLMs Playing the Game of Avalon
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f376d00-af2a-400e-bd75-d07db199621e · outbound
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks AgentBench: Evaluating LLMs as Agents
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fde970f7-9348-40e1-9558-6fcb47d8a8ff · outbound
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f87afcd7-dcb0-4cce-b4f7-3ef77efe648e · outbound
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b2f2ef39-5c7f-4313-a5a7-9521dd196d9d · outbound
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7dc2bd4a-cc3e-44b4-ab76-a4ef4532e340 · outbound
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks S.; Farquhar, G.; Foerster, J.; and Whiteson, S
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d1c389c-55fc-44b2-a709-f66e4b0f7af7 · outbound
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks TurtleBench: A Visual Programming Benchmark in Turtle Geometry
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fb317f2-1174-470a-a596-89e4fb743e01 · outbound
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks The StarCraft Multi-Agent Challenge
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f48fbc19-1b65-4b94-acd6-9684496eeb38 · outbound
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Exploring and Improving the Spatial Reasoning Abilities of Large Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 92afc32c-3678-4e37-b21d-cb174bd43a38 · outbound
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7390fba7-1b8b-42a2-843a-1b9e285ac15b · outbound
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5d87f50-03aa-4644-a248-302c63543b03 · outbound
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Qwen3 Technical Report
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52da062c-f74c-4ed1-8a55-a8b24269c144 · outbound
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks M.; Mathieu, M.; Dudzik, A.; Chung, J.; Choi, D
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f63a0b5f-976e-4d20-be07-1e6fc91df16b · outbound
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Voyager: An Open-Ended Embodied Agent with Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3884b65-701e-440e-ad64-110431ccac56 · outbound
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78f0f4cc-a38c-47d8-95eb-2a376f230b70 · outbound
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Avalon's Game of Thoughts: Battle Against Deception through Recursive Contemplation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a81d78a3-00e7-4120-bc75-c523599c43e7 · outbound
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks V.; Zhou, D.; et al
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ec19a5a-b486-460d-bf68-d71f42a55245 · outbound
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17bad6cb-08d8-4fa7-ae26-f9c03a337197 · outbound
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation decb996e-fb87-4d6e-bca9-4778c7a89e27 · outbound
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35b59c4f-6776-4b7f-b325-341e092b3b3b · outbound
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7053f5d5-dea0-40bc-9503-8717edc46fbd · outbound
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Advancing LLM Reasoning Generalists with Preference Trees
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a41886f-e875-406b-9aac-2dd5df8f76f6 · outbound
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Y.; Ju, J.; Nguyen, A
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 86275c19-63dd-4d84-b96b-ff9cbb9321a8 · outbound
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks WebArena: A Realistic Web Environment for Building Autonomous Agents
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b0044aa-d816-46ff-82d3-be3a138f2e0f · inbound
RTSGameBench: An RTS Benchmark for Strategic Reasoning by Vision-Language Models SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.