Pith. sign in

Paper Citation Record · LEDGER

Reasoning Capabilities of Large Language Models on Dynamic Tasks

As of 19 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 1 inbound Pith citation observation for arXiv:2505.10543.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.10543 v2

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:12:01.571261Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T21:36:44.033596Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T14:25:46.619265Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 775d0281-04ef-4006-b35f-d0d7237facb2 · outbound

This paper cites Language mod- els are few-shot learners,.

Reasoning Capabilities of Large Language Models on Dynamic Tasks Language mod- els are few-shot learners,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.466773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.466773Z digest=sha256:3935b79b25492bbd518c77056aa3f77110a838c3b57da462da8226e79975ff00

Observation ca51aa6e-73d8-4458-87c9-cb654602f008 · outbound

This paper cites Reflex- ion: Language agents with verbal reinforcement learning,.

Reasoning Capabilities of Large Language Models on Dynamic Tasks Reflex- ion: Language agents with verbal reinforcement learning,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.471784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.471784Z digest=sha256:cb9c4032aa50b97ec5d9945c248863a9279f1db9b7a9664c8b70f1ac41f785bc

Observation 5d6b2e2a-80bf-4586-a364-39ea9594b2ec · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Reasoning Capabilities of Large Language Models on Dynamic Tasks ReAct: Synergizing Reasoning and Acting in Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.477653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.477653Z digest=sha256:852be0278ba232a06b8ff9103c50784711ea1fa32c9faacc9212f72f0cfb8780

Observation 0ff6ec71-be91-42a3-8f17-8315825a5e20 · outbound

This paper cites The Rise and Potential of Large Language Model Based Agents: A Survey.

Reasoning Capabilities of Large Language Models on Dynamic Tasks The Rise and Potential of Large Language Model Based Agents: A Survey

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.483633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.483633Z digest=sha256:3de1ecd38bcc3d934270c211ffa7f02e9228b86385e0bdce767166b6fc1b59a2

Observation 397d00e7-3ba5-429e-a182-d106554c23df · outbound

This paper cites Text-based games as a challenging benchmark for large language models,.

Reasoning Capabilities of Large Language Models on Dynamic Tasks Text-based games as a challenging benchmark for large language models,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:12:01.979481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:12:01.488755Z digest=sha256:d9bd8e518f04166639baa37534cf81a275055cc6253404be4a4de7e3524910d2

Observation 1f7aa33f-e1db-4eb1-b571-debb7e5e8640 · outbound

This paper cites Training language models to follow instructions with human feedback,.

Reasoning Capabilities of Large Language Models on Dynamic Tasks Training language models to follow instructions with human feedback,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.493247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.493247Z digest=sha256:13232013e8deca8f65cf8b3ac87c50f3a87ca90bc1e8623818aec92bd3f4e776

Observation ba4a0d14-a25e-4426-9429-0b67fc784714 · outbound

This paper cites Prompt programming for large language models: Beyond the few-shot paradigm,.

Reasoning Capabilities of Large Language Models on Dynamic Tasks Prompt programming for large language models: Beyond the few-shot paradigm,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.498450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.498450Z digest=sha256:ed5ad95d1e728e03e28f702266693724ab269bc5f22d3204345ac6d42d332af4

Observation 564526f6-4b44-4d2e-9eb5-c1230ed93105 · outbound

This paper cites SmartPlay: A Benchmark for LLMs as Intelligent Agents.

Reasoning Capabilities of Large Language Models on Dynamic Tasks SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.502630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.502630Z digest=sha256:6bedeca91a08effc538376d6292d033b80aecc607dcc20c1962ab2d6aa93139f

Observation ac619244-a0bb-40b6-a592-7810617ec795 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

Reasoning Capabilities of Large Language Models on Dynamic Tasks Chain-of-thought prompting elicits reasoning in large language models,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.507211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.507211Z digest=sha256:5e1bf1a74bd36d691ee7035897407590007cee8edb58ba870927e422cea5aca2

Observation 50feaa8f-be13-44e6-a508-f20848d88d13 · outbound

This paper cites Self-refine: Iter- ative refinement with self-feedback,.

Reasoning Capabilities of Large Language Models on Dynamic Tasks Self-refine: Iter- ative refinement with self-feedback,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.511339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.511339Z digest=sha256:5eef3f03de1f9cd3904222455c766fb8622f10f60b22ccb38818f74f26500393

Observation 0259a033-897e-4c0e-8a1c-2a357cedce90 · outbound

This paper cites Agent-Pro: Learning to Evolve via Policy-Level Reflection and Optimization.

Reasoning Capabilities of Large Language Models on Dynamic Tasks Agent-Pro: Learning to Evolve via Policy-Level Reflection and Optimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.515887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.515887Z digest=sha256:cbc2fbec7d1c07ee685df65d29e58aadcf72a5f6e27cd2cc1997031d0822c98a

Observation b6cafd80-2a29-4413-8728-be1593027a34 · outbound

This paper cites AutoPlan: Automatic Planning of Interactive Decision-Making Tasks With Large Language Models.

Reasoning Capabilities of Large Language Models on Dynamic Tasks AutoPlan: Automatic Planning of Interactive Decision-Making Tasks With Large Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:12:01.733559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:12:01.520534Z digest=sha256:60f8a0a3ed1c18073ca15833fa3df3dc8417e12dd26de26877b3efbf8b73f5e4

Observation 377af352-bb2d-4d06-9771-d332ad7d3e72 · outbound

This paper cites Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents.

Reasoning Capabilities of Large Language Models on Dynamic Tasks Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.524931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.524931Z digest=sha256:e60b2e270f4304e893e603325f1e2427af151cba3e187ad6041c17d20077e710

Observation 4b535f04-ce16-415b-bc38-e7c19f0013d8 · outbound

This paper cites Mental Modeling of Reinforcement Learning Agents by Language Models.

Reasoning Capabilities of Large Language Models on Dynamic Tasks Mental Modeling of Reinforcement Learning Agents by Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.529561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.529561Z digest=sha256:d4b0741f992f654a97909b6ea9a296e32c3e107c8c09590f13ee9846de032638

Observation 62adbb95-fe10-4f37-be0c-6b79d85bdc0a · outbound

This paper cites EvoPrompt: Connecting LLMs with Evolutionary Algorithms Yields Powerful Prompt Optimizers.

Reasoning Capabilities of Large Language Models on Dynamic Tasks EvoPrompt: Connecting LLMs with Evolutionary Algorithms Yields Powerful Prompt Optimizers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.534625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.534625Z digest=sha256:b3a98161449050b7af1b7a4d454649f2b9d0daa8c0ab3a32717198d8b0cb94e7

Observation db7d2981-7220-4eb8-b4f8-d329ee8e128b · outbound

This paper cites Llamea: A large language model evolutionary algorithm for automatically generating metaheuristics,.

Reasoning Capabilities of Large Language Models on Dynamic Tasks Llamea: A large language model evolutionary algorithm for automatically generating metaheuristics,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:12:01.925131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:12:01.539176Z digest=sha256:7e85b9ba622ef018d77da97b3db87b86d53ba48bd5b779e3def2ae1dc827f34f

Observation 9d576973-9a70-4cc9-9ba4-d9da422eac3b · outbound

This paper cites Focused transformer: Contrastive training for context scaling,.

Reasoning Capabilities of Large Language Models on Dynamic Tasks Focused transformer: Contrastive training for context scaling,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:12:01.909695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:12:01.543682Z digest=sha256:90252f2ffe2fd746754bc591a0ff5705a36bea1e4fa0e1f3df13401b81390b31

Observation e1a413d7-8a89-420c-928b-dfd617713924 · outbound

This paper cites Lost in the Middle: How Language Models Use Long Contexts.

Reasoning Capabilities of Large Language Models on Dynamic Tasks Lost in the Middle: How Language Models Use Long Contexts

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.547721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.547721Z digest=sha256:0811486032bb6d296f296e6f46273e164a4494da26c8b45947a4e009182dba18

Observation 528de17c-6008-4c33-b393-b4b9f20b0797 · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

Reasoning Capabilities of Large Language Models on Dynamic Tasks Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.552813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.552813Z digest=sha256:ea098f4bf422d1ad06556e0794c4e68a6614a89e26006ccb0892d736123b6d5d

Observation 937adfb6-0f8f-40dc-aed3-cc8084c7ca7d · outbound

This paper cites Chain of thought- lessness? an analysis of cot in planning,.

Reasoning Capabilities of Large Language Models on Dynamic Tasks Chain of thought- lessness? an analysis of cot in planning,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:12:01.819066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:12:01.557485Z digest=sha256:a2eddba62e78556d3d0e6314ff84f61250ed2ce742ee67fa984786411fdbae95

Observation 6b30c080-4519-49ea-ae4d-2a4feb7e57d4 · outbound

This paper cites Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?.

Reasoning Capabilities of Large Language Models on Dynamic Tasks Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.562040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.562040Z digest=sha256:2a7a8c6d03c4a7537b912d888b8ba97937ac78039d09a1bab86d49d98eaa0995

Observation a1c75472-7ada-44f0-b4fb-ed9f59678d52 · outbound

This paper cites HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering.

Reasoning Capabilities of Large Language Models on Dynamic Tasks HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.566767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.566767Z digest=sha256:5b91aa3191775483debecba9aac677c930672128a6b061a5e9dcec912e573d06

Observation f0664dfa-eb55-42ae-9c61-6d5a51f64d30 · outbound

This paper cites BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games.

Reasoning Capabilities of Large Language Models on Dynamic Tasks BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.571261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.571261Z digest=sha256:ef51ffa7efc210a94262a2ccea21d8f5897ef73a9dc96db7ef265b87c2c55640

Pith citing papers

Observation 82031c4c-49ec-4ff8-99a9-b8e562076058 · inbound

Why We Need World Models for AGI: Where LLMs Fail and How World Models May Outperform cites this paper.

Why We Need World Models for AGI: Where LLMs Fail and How World Models May Outperform Reasoning Capabilities of Large Language Models on Dynamic Tasks

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:25:46.621005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T21:36:44.033596Z digest=sha256:d04e2234a933474461e9c295cf23ac7d60261f407b1e226d8f51a3700f0a98fb