Pith. sign in

Paper Citation Record · LEDGER

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play

As of 8 August 2026, this Paper Citation Record lists 11 of 11 outbound references and 0 inbound Pith citation observations for arXiv:2605.22238.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.22238 v1

Coverage vector

measured 11 of 11 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-22T05:37:44.982932Z

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

11 of 11 outbound references displayed

  • verified exact9
  • verified fuzzy2
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6b3c7bfb-814c-4cef-86ca-bd27026cc3e8 · outbound

This paper cites Science , volume =.

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play Science , volume =

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T05:56:09.458656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-22T05:37:44.982932Z digest=sha256:dc45cbc488ef3ed9d5c65c7c12483af1e395d301085d5e159e3e041e4b78915d

Observation b352fd9b-bde4-48db-9097-77c2079dc5c5 · outbound

This paper cites 2024 , url =.

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play 2024 , url =

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T05:56:09.462337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-22T05:37:44.982932Z digest=sha256:20a363db7af6ea8d822120b9ce19c2cb08e3f8d94841cbcb2decbfe86107b4cd

Observation 708a3513-9f52-4af9-83c9-a4c96515b264 · outbound

This paper cites Human-level play in the game of diplomacy by combining language models with strategic reasoning.

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play Human-level play in the game of diplomacy by combining language models with strategic reasoning

Reference 10

Resolution
verified exact
doi, observed 2026-05-22T05:41:07.034526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-22T05:37:44.982932Z digest=sha256:bda1893f68f518df802d6ec2981734097805e5973accbf8c80c2cc2348ffe7dc

Observation aaa34980-a9de-4766-81ec-46c655554ffe · outbound

This paper cites GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents.

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:41:08.490213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-22T05:37:44.982932Z digest=sha256:82fc8264e4bf5747d62ae59ea56d89a920e79a07c9b2a060815208efa31cfa6b

Observation 2f52693b-87b6-43c8-ba42-e929196dd843 · outbound

This paper cites Strategic Reasoning with Language Models.

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play Strategic Reasoning with Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:41:08.486282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-22T05:37:44.982932Z digest=sha256:1f7cdef2b493ac5dc23bbc0ea9c125512e4e53f0e5a9bf626fe6a19273a3e418

Observation 1fd2599b-1768-4a2f-878f-c60e6f7610e4 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play Measuring Massive Multitask Language Understanding

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:41:08.482224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-22T05:37:44.982932Z digest=sha256:d58dbbeec3f948193e1326f0fb2d3ce2ffa5d492adcc818b8ec5ffd8f1c30798

Observation 4248d872-252a-4963-bc25-47c423a5bb63 · outbound

This paper cites Holistic Evaluation of Language Models.

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play Holistic Evaluation of Language Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:41:08.473414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-22T05:37:44.982932Z digest=sha256:a3d90a51620fa687820c5a857410c12a7be30e8ca45e3c7a6d33a9c572930c42

Observation 9c3e1682-da6d-4e13-831d-89dd2c7aa389 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play AgentBench: Evaluating LLMs as Agents

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:41:08.469034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-22T05:37:44.982932Z digest=sha256:2ea7ecddb22a52faebdeac3868050f04f891a46c512c5640fdd07591d727060d

Observation a385a6bf-19c5-4b7f-99fd-95d7e89096fc · outbound

This paper cites GAIA: a benchmark for General AI Assistants.

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play GAIA: a benchmark for General AI Assistants

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:41:08.463536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-22T05:37:44.982932Z digest=sha256:48b0caf66229bef363db70f132e85546454dfac1562147847d988f5779e0ff65

Observation 850250d3-cf0a-41a9-ae37-e18e2fbeca8c · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:41:08.478010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-22T05:37:44.982932Z digest=sha256:e7488f74d3d26218de6dd1e269b5e392b6503863ee4f80136739e5a601ad1d0a

Observation 19e3369b-35d0-40d3-872e-3d0c9fd0433b · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:41:08.494410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-22T05:37:44.982932Z digest=sha256:72701e19892dd909f863130416d6c6bdf8c8a0ca4a0ad0d7b1acfe9c5252c1ef

Pith citing papers

No inbound Pith citation observations are available.