Pith. sign in

Paper Citation Record · LEDGER

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play

As of 9 August 2026, this Paper Citation Record lists 11 of 11 outbound references and 0 inbound Pith citation observations for arXiv:2605.22238.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.22238 v1

Coverage vector

measured 11 of 11 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-22T05:37:44.982932Z

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

11 of 11 outbound references displayed

  • verified exact9
  • verified fuzzy2
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6b3c7bfb-814c-4cef-86ca-bd27026cc3e8 · outbound

This paper cites Science , volume =.

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play Science , volume =

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T05:56:09.458656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-22T05:37:44.982932Z digest=sha256:6ebb23d950e7d078e51f614ca539fffb803eac609ad5ee9f95473623a7fbee94

Observation b352fd9b-bde4-48db-9097-77c2079dc5c5 · outbound

This paper cites 2024 , url =.

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play 2024 , url =

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T05:56:09.462337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-22T05:37:44.982932Z digest=sha256:ba652ce3450a0c05d46d1f4371854ce18c87e5fce9bf9d05328059d52f9539b1

Observation 708a3513-9f52-4af9-83c9-a4c96515b264 · outbound

This paper cites Human-level play in the game of diplomacy by combining language models with strategic reasoning.

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play Human-level play in the game of diplomacy by combining language models with strategic reasoning

Reference 10

Resolution
verified exact
doi, observed 2026-05-22T05:41:07.034526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-22T05:37:44.982932Z digest=sha256:89a8fbbba847774beb30637d454213817fd0c7de0ec7025916197057069d9f7e

Observation aaa34980-a9de-4766-81ec-46c655554ffe · outbound

This paper cites GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents.

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:41:08.490213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-22T05:37:44.982932Z digest=sha256:f8f904243ebeac2de09540a564ba96751ffdf9030dd8acc712dd1cc0805b3aee

Observation 2f52693b-87b6-43c8-ba42-e929196dd843 · outbound

This paper cites Strategic Reasoning with Language Models.

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play Strategic Reasoning with Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:41:08.486282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-22T05:37:44.982932Z digest=sha256:11991431597784129d5111faa31a721bd6b4fd48f8e77e1961bb42639ef6b1d9

Observation 1fd2599b-1768-4a2f-878f-c60e6f7610e4 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play Measuring Massive Multitask Language Understanding

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:41:08.482224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-22T05:37:44.982932Z digest=sha256:3dce93935e117403afafc890ed67fc0e9fb09bc118ec4e4857c1f5f78bd47a76

Observation 4248d872-252a-4963-bc25-47c423a5bb63 · outbound

This paper cites Holistic Evaluation of Language Models.

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play Holistic Evaluation of Language Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:41:08.473414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-22T05:37:44.982932Z digest=sha256:4723ec1e4273fceb5e735f7a9328bf0f5b8b7d830a321ab4130a5e728e758023

Observation 9c3e1682-da6d-4e13-831d-89dd2c7aa389 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play AgentBench: Evaluating LLMs as Agents

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:41:08.469034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-22T05:37:44.982932Z digest=sha256:579f469122638d082ccf27b4532588d02f0f1d79a41b25ea7094144b6cbba2a9

Observation a385a6bf-19c5-4b7f-99fd-95d7e89096fc · outbound

This paper cites GAIA: a benchmark for General AI Assistants.

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play GAIA: a benchmark for General AI Assistants

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:41:08.463536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-22T05:37:44.982932Z digest=sha256:be2d7c3e288634dc77d771cb6ef1f4320d83d180c670ad503c76c84c61d3c017

Observation 850250d3-cf0a-41a9-ae37-e18e2fbeca8c · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:41:08.478010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-22T05:37:44.982932Z digest=sha256:2d85c0f075cd63298eeeb34f626119d1aeec728aeada1d9509d29c81eec42327

Observation 19e3369b-35d0-40d3-872e-3d0c9fd0433b · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:41:08.494410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-22T05:37:44.982932Z digest=sha256:c353abf6e1b45afb7c42f5b4a670641600f89817a71b90f1b0ac9594995e164d

Pith citing papers

No inbound Pith citation observations are available.