Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-22T05:37:44.982932Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 11 of 11 outbound references and 0 inbound Pith citation observations for arXiv:2605.22238.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-22T05:37:44.982932Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
11 of 11 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6b3c7bfb-814c-4cef-86ca-bd27026cc3e8 · outbound
Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play Science , volume =
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b352fd9b-bde4-48db-9097-77c2079dc5c5 · outbound
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 708a3513-9f52-4af9-83c9-a4c96515b264 · outbound
Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play Human-level play in the game of diplomacy by combining language models with strategic reasoning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation aaa34980-a9de-4766-81ec-46c655554ffe · outbound
Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2f52693b-87b6-43c8-ba42-e929196dd843 · outbound
Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play Strategic Reasoning with Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1fd2599b-1768-4a2f-878f-c60e6f7610e4 · outbound
Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play Measuring Massive Multitask Language Understanding
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4248d872-252a-4963-bc25-47c423a5bb63 · outbound
Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play Holistic Evaluation of Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9c3e1682-da6d-4e13-831d-89dd2c7aa389 · outbound
Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play AgentBench: Evaluating LLMs as Agents
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a385a6bf-19c5-4b7f-99fd-95d7e89096fc · outbound
Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play GAIA: a benchmark for General AI Assistants
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 850250d3-cf0a-41a9-ae37-e18e2fbeca8c · outbound
Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 19e3369b-35d0-40d3-872e-3d0c9fd0433b · outbound
Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
No inbound Pith citation observations are available.