Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:34:40.588180Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2506.10264.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:34:40.588180Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:46:04.569599Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T04:46:06.449061Z
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5ef329c2-faa8-4c3c-91f6-9bdce7c3567b · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models A Survey of Large Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a0ed04e-5fa6-42f1-97a3-aafaf91785e6 · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fcf919a-7c90-4337-a613-36fc49f7c809 · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models DeepSeek-V3 Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5ffafba-0b18-46cb-9c61-3384db503749 · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models From System 1 to System 2: A Survey of Reasoning Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fc1952e-d8ff-4904-a2a9-b9b9d374da4e · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1a9a322-c273-4045-bfc8-e53b9d002f2b · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d739b67a-0cd2-485d-a83c-f8f98f421293 · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Large Language Models for Mathematical Reasoning: Progresses and Challenges
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e98680d2-4cf7-4ec3-b9b9-f7237ac17f62 · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Symbol-llm: Leverage language models for symbolic system in visual human activity reasoning,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a1320bd8-410b-40cf-a8b0-d6f97d3b1b38 · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Large language models as com- monsense knowledge for large-scale task planning,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 92f67507-c599-4a74-9854-40264fa754f0 · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models LLM as a Mastermind: A Survey of Strategic Reasoning with Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db5f9711-92a8-4130-83e3-98e13121d246 · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c92b787-ddad-496b-83e6-a9517e2cd855 · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models AvalonBench: Evaluating LLMs Playing the Game of Avalon
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c76d2130-52e1-4adf-bd7e-30f47eee7cde · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Large language models play starcraft ii: Benchmarks and a chain of summarization approach,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 548dc6b5-71ea-4142-bde3-ed7c49c19f2c · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Hierarchical Expert Prompt for Large-Language-Model: An Approach Defeat Elite AI in TextStarCraft II for the First Time
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d37a19b2-dfde-4166-84ee-8fc49b616ab4 · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cef86c4-0035-4502-a347-cb4f3d6ce03d · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models MAgIC: Investigation of Large Language Model Powered Multi-Agent in Cognition, Adaptability, Rationality and Collaboration
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74aa5be8-8b80-47cf-a647-7f0c6a55d467 · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models SmartPlay: A Benchmark for LLMs as Intelligent Agents
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a00ec7a-1e20-44a4-8a41-c6985d370f85 · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models MinePlanner: A Benchmark for Long-Horizon Planning in Large Minecraft Worlds
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5cc67543-6e4f-4d34-ba79-1051858155a5 · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Benchmarking Agentic Workflow Generation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f60b63af-7eef-4021-a1fa-69ac4cacbf93 · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Opendeception: Bench- marking and investigating ai deceptive behaviors via open-ended interaction simulation,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9eef73d2-8dc6-4004-bb3c-9b50647cc91e · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Intelligent decision making technol- ogy and challenge of wargame,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b0313ab1-a818-47fb-a7c0-fcbc6dfdbde4 · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b80d093f-01fe-4f2f-a1b2-9759cf70a94b · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models A five- dimensional framework for authentic assessment,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 95650192-16e4-4a14-86b5-250e7802eeca · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Appropriate criteria: Key to effective rubrics,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6915e39c-f832-4f67-9b9e-2119d88fea25 · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f1778ec8-9ea0-43c4-aad4-b6c2a46a3eb5 · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models A review of rubric use in higher education,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2d8f7a0b-7ce3-4ead-8e43-9d0e4c299ef8 · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Wiggins,Educative Assessment: Designing Assessments to Inform and Improve Student Performance
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0bd4715e-b19e-4b47-9f0b-8a57cf5c177b · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models A revision of bloom’s taxonomy: An overview,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 880d3f6e-76e3-425b-a96f-5da3de1616c7 · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3f2c76a7-d340-4289-b2d8-643704ab78a3 · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Improving instruction and assessment via bloom’s taxonomy and descriptive rubrics,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccde419f-35c4-4a70-a439-1e6ccf439aa2 · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Llm-rubric: A multidimensional, calibrated approach to auto- mated evaluation of natural language texts,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 29726c48-a90a-492d-985b-6a95f506daa0 · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Gptscore: Evaluate as you desire,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 978657f9-f8ad-441e-87cd-bd2e3649bc8d · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models G-eval: Nlg evaluation using gpt-4 with better human alignment,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fa1583f1-75a5-41f0-b858-5fee8b9242b1 · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Large language models are not fair evaluators,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d32b7a67-3b61-43d7-9c60-5a1138982931 · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Realtoxicityprompts: Evaluating neural toxic degeneration in language models,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7615c8bc-a834-4911-9077-33f9fefb4a88 · outbound
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Why we need new evaluation metrics for nlg,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7f9b05cd-6ed6-4e5e-b6d7-4f31ed705399 · inbound
Multi-level Value Alignment in Agentic AI Systems: Survey and Perspectives WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models
Reference 148
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.