Pith. sign in

Paper Citation Record · LEDGER

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks

As of 9 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 1 inbound Pith citation observation for arXiv:2508.10428.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.10428 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:30:30.167398Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T21:02:26.081166Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T00:39:17.301001Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact2
  • verified fuzzy4
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 402af738-96bf-48ea-a77c-3c528f4d0aa5 · outbound

This paper cites , " * write output.state after.block = add.period write newline.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:28.708904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:28.708904Z digest=sha256:171209b10485b100ad1d9a24b71794e845ce8e8292aa8f11a99a2434a4ba3a1b

Observation fcfd4c6f-7986-4566-bfbd-09aa7febcc19 · outbound

This paper cites write newline.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:28.792837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:28.792837Z digest=sha256:65a41187107c21e6f114dbcc8442195219d4084faac5f4659e361cd4807b4fb7

Observation 3ab399f5-2654-49db-b347-e4bfaeaa3917 · outbound

This paper cites GPT-4 Technical Report.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:28.941922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:28.941922Z digest=sha256:b3ae7aa63eaae3e5f299903ac52955b751343d7041239c316d0c1c105ee56c4d

Observation 6b936859-190d-4248-bc2a-732f483af256 · outbound

This paper cites K.; and Johnson, B.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks K.; and Johnson, B

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:30:30.442231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:30:29.110846Z digest=sha256:a63d7249c2952f44689da7680e2f9399f78b632f452e749520fd7e16c48b46e4

Observation d7755b31-5bda-4475-bf0b-d9d04ef19bcf · outbound

This paper cites DeepSeek-V3 Technical Report.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks DeepSeek-V3 Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:29.216038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:29.216038Z digest=sha256:792a7c19e25546df4cf8f72e57fd2318fb4ed555f6d0d7fc53c182a1a169d388

Observation a1b8d12e-1e9a-43ae-abf7-463dafdf8255 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:29.352392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:29.352392Z digest=sha256:5919d9b0bdce4f2e3fbd98e386f86abb364de43810422c66265c8c486f5fe987

Observation a3304b1e-6bc5-4952-b5d3-00963259cf41 · outbound

This paper cites L.; Yao, S.; Chen, Y.; Shen, P.; Yu, H.; Zhang, H.; Zhang, X.; Dong, Y.; et al.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks L.; Yao, S.; Chen, Y.; Shen, P.; Yu, H.; Zhang, H.; Zhang, X.; Dong, Y.; et al

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:30:30.433855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:30:29.424757Z digest=sha256:91fd74534090e3de852411ea9fbb3ef21c56e42112a3c3adb9059dcc1b1a0711

Observation 36a70c36-a638-4ce1-a441-ead97bbd793d · outbound

This paper cites LLM-PySC2: Starcraft II learning environment for Large Language Models.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks LLM-PySC2: Starcraft II learning environment for Large Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-05T20:30:30.321803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:30:29.516705Z digest=sha256:77759c76ffb9119b1b4290c8534d2a7a42acd41fe8fe0d7fa161a56cb06a774d

Observation 505180f6-239f-4f55-a4c2-653ea13f0b53 · outbound

This paper cites AvalonBench: Evaluating LLMs Playing the Game of Avalon.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:29.614571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:29.614571Z digest=sha256:ea444fe35ca974768d9d093725f7afb41d78fff4d95489b345708fcb81342d0f

Observation 8f376d00-af2a-400e-bd75-d07db199621e · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks AgentBench: Evaluating LLMs as Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:29.700721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:29.700721Z digest=sha256:377d46d4d374e907e76031b86b75c3c524a03f31536a33ec1b7071e44d362f82

Observation fde970f7-9348-40e1-9558-6fcb47d8a8ff · outbound

This paper cites an unresolved cited work.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:30:30.425166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:30:29.807997Z digest=sha256:4ee05254310d0b1700ae46d2d90c0c688994cd3186ddc72e72f440ad6b40a58b

Observation f87afcd7-dcb0-4cce-b4f7-3ef77efe648e · outbound

This paper cites an unresolved cited work.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:30:30.415893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:30:29.895387Z digest=sha256:9a57160d1fcdc4589ca825dcd10e63a7b7ad8a2345fc61d0643d00a2989cfd14

Observation b2f2ef39-5c7f-4313-a5a7-9521dd196d9d · outbound

This paper cites an unresolved cited work.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:30:30.407634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:30:29.990030Z digest=sha256:d0a02015ff85d47dfe76fa691cf03d6ea46d8542527b6207d38a941502ec2878

Observation 7dc2bd4a-cc3e-44b4-ab76-a4ef4532e340 · outbound

This paper cites S.; Farquhar, G.; Foerster, J.; and Whiteson, S.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks S.; Farquhar, G.; Foerster, J.; and Whiteson, S

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:30.090360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:30.090360Z digest=sha256:3c89cf61f4c792ba24d524bc9898983991ab1cc29630b5cba12c4d508a8d1a6c

Observation 0d1c389c-55fc-44b2-a709-f66e4b0f7af7 · outbound

This paper cites TurtleBench: A Visual Programming Benchmark in Turtle Geometry.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks TurtleBench: A Visual Programming Benchmark in Turtle Geometry

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:30.112631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:30.112631Z digest=sha256:b8046079634a828b4cb21f888ddc22fadc6d78d57e20bb3a25ee41f46f4e81f5

Observation 0fb317f2-1174-470a-a596-89e4fb743e01 · outbound

This paper cites The StarCraft Multi-Agent Challenge.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks The StarCraft Multi-Agent Challenge

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:30.115627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:30.115627Z digest=sha256:f1928d1e2dd5d955e3e6200f334dd94604795779da308b80491a15e7934cb121

Observation f48fbc19-1b65-4b94-acd6-9684496eeb38 · outbound

This paper cites Exploring and Improving the Spatial Reasoning Abilities of Large Language Models.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Exploring and Improving the Spatial Reasoning Abilities of Large Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-05T20:30:30.273375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:30:30.118412Z digest=sha256:77ea3a3450920161a49cd0721dc646e13b9241c403814505d7d0a4ee61d2e517

Observation 92afc32c-3678-4e37-b21d-cb174bd43a38 · outbound

This paper cites Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:30.121068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:30.121068Z digest=sha256:327da727406949e9a2ec385d71490b7fa9404f807b466c6e18fe3ca6dc62fe03

Observation 7390fba7-1b8b-42a2-843a-1b9e285ac15b · outbound

This paper cites an unresolved cited work.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:30.123990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:30.123990Z digest=sha256:1b4bc058bee688bf6eda44a75fc8b13ee99ed9d2c58384c28dae41b24febfc6e

Observation e5d87f50-03aa-4644-a248-302c63543b03 · outbound

This paper cites Qwen3 Technical Report.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Qwen3 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:30.127668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:30.127668Z digest=sha256:613aa19d69c73a84067a09ce2c43ba227643891ecbfd55759a01af33fa5da704

Observation 52da062c-f74c-4ed1-8a55-a8b24269c144 · outbound

This paper cites M.; Mathieu, M.; Dudzik, A.; Chung, J.; Choi, D.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks M.; Mathieu, M.; Dudzik, A.; Chung, J.; Choi, D

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:30:30.388120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:30:30.130712Z digest=sha256:04f9bd23c3098396636eb03a95944d749fd8250976e81c9828c536042ade4214

Observation f63a0b5f-976e-4d20-be07-1e6fc91df16b · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:30.133129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:30.133129Z digest=sha256:5e69b2696ddccd64ef848b59a5e40a8692872494a24c32f4822aad0773beb289

Observation b3884b65-701e-440e-ad64-110431ccac56 · outbound

This paper cites Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:30.136060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:30.136060Z digest=sha256:1c764caa6d0d41c7725c42c2eb601e73db38b322160fc14e92d81af06b9d251b

Observation 78f0f4cc-a38c-47d8-95eb-2a376f230b70 · outbound

This paper cites Avalon's Game of Thoughts: Battle Against Deception through Recursive Contemplation.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Avalon's Game of Thoughts: Battle Against Deception through Recursive Contemplation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:30.139169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:30.139169Z digest=sha256:371570f4e0607d2146fd82e7b61c276a219989ce47b70636d76e390995f5a624

Observation a81d78a3-00e7-4120-bc75-c523599c43e7 · outbound

This paper cites V.; Zhou, D.; et al.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks V.; Zhou, D.; et al

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:30.146973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:30.146973Z digest=sha256:eddacbf6ac87b7d8811b980fe4c8c96470c7d030105e3414d81927bbc90e8cde

Observation 7ec19a5a-b486-460d-bf68-d71f42a55245 · outbound

This paper cites Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:30.149845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:30.149845Z digest=sha256:0d1ff1e6015dd6da034a8215ff65892a858e216f7ed758d96783da80196e297d

Observation 17bad6cb-08d8-4fa7-ae26-f9c03a337197 · outbound

This paper cites an unresolved cited work.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:30.152711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:30.152711Z digest=sha256:6ffee3a5933dd28cdcec51362dd7aa282c795444df0b52a7e333be03c933a3f1

Observation decb996e-fb87-4d6e-bca9-4778c7a89e27 · outbound

This paper cites an unresolved cited work.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:30.156489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:30.156489Z digest=sha256:1be02486c04618fc29285f66a568413e2a442075845dbb3187125c60b824fc3f

Observation 35b59c4f-6776-4b7f-b325-341e092b3b3b · outbound

This paper cites an unresolved cited work.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:30.159198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:30.159198Z digest=sha256:7312fd400668b059559643f80683cad0303bd45c827056c7ea25519b0d7862fc

Observation 7053f5d5-dea0-40bc-9503-8717edc46fbd · outbound

This paper cites Advancing LLM Reasoning Generalists with Preference Trees.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Advancing LLM Reasoning Generalists with Preference Trees

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:30.162000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:30.162000Z digest=sha256:bf054fec7e2825c0054fc3586e8ab86491bb5eedf687d96afdd0be9b2e7f1d2c

Observation 6a41886f-e875-406b-9aac-2dd5df8f76f6 · outbound

This paper cites Y.; Ju, J.; Nguyen, A.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Y.; Ju, J.; Nguyen, A

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:30:30.357438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:30:30.164783Z digest=sha256:fc881876f8e468a9f69ecac72ff0511d93e62dc15f1fa0d4496ab90403c2a7cd

Observation 86275c19-63dd-4d84-b96b-ff9cbb9321a8 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:30.167398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:30.167398Z digest=sha256:916829e2c59af7e03f5675eccf0ef6efb92359c114a1be6a50578d75b338a6f4

Pith citing papers

Observation 0b0044aa-d816-46ff-82d3-be3a138f2e0f · inbound

RTSGameBench: An RTS Benchmark for Strategic Reasoning by Vision-Language Models cites this paper.

RTSGameBench: An RTS Benchmark for Strategic Reasoning by Vision-Language Models SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:17.302344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T21:02:26.081166Z digest=sha256:8907c4158ff0d75602c9bd82d7cb8d3efab0bbd01a686dde997ee6517924ef2d