Pith. sign in

Paper Citation Record · LEDGER

How Consistent Are LLM Agents? Measuring Behavioral Reproducibility in Multi-Step Tool-Calling Pipelines

As of 15 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 3 inbound Pith citation observations for arXiv:2605.28840.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.28840 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-04T21:40:04.594380Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:20:34.128103Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-11T20:08:14.591497Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact5
  • verified fuzzy9
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2820413a-b0d2-471e-b7ea-cae3e27d243a · outbound

This paper cites AI Agents That Matter.

How Consistent Are LLM Agents? Measuring Behavioral Reproducibility in Multi-Step Tool-Calling Pipelines AI Agents That Matter

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-04T21:40:08.959105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-07-04T21:40:04.594380Z digest=sha256:6d2b887a762fd1678f77935cda01181eeed02071d8536d5714af94d24a560e92

Observation 24bd9db3-4bf6-4676-8d65-159c7b04e6ae · outbound

This paper cites Simple and Scalable Predictive Uncertainty Estimation Using Deep Ensembles.

How Consistent Are LLM Agents? Measuring Behavioral Reproducibility in Multi-Step Tool-Calling Pipelines Simple and Scalable Predictive Uncertainty Estimation Using Deep Ensembles

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T21:40:08.971067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-07-04T21:40:04.594380Z digest=sha256:e5c33a7aaa37da3bcd3056cb6700b7da7abe92d581aa369e296dc3d507fab20c

Observation c4923d33-4936-449a-ba5f-50e2e39040cf · outbound

This paper cites API-Bank : A Comprehensive Benchmark for Tool-Augmented LLMs.

How Consistent Are LLM Agents? Measuring Behavioral Reproducibility in Multi-Step Tool-Calling Pipelines API-Bank : A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T21:40:08.980586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-07-04T21:40:04.594380Z digest=sha256:6efcd58dc4f9e1f445b409695145e5fdcc9f0cd737aca41edfc77654dac92ae2

Observation 29f4bf33-5b75-4cae-a1b0-dccee256f807 · outbound

This paper cites AgentBench : Evaluating LLMs as Agents.

How Consistent Are LLM Agents? Measuring Behavioral Reproducibility in Multi-Step Tool-Calling Pipelines AgentBench : Evaluating LLMs as Agents

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T21:40:08.978307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-07-04T21:40:04.594380Z digest=sha256:7056cdbb7b923c9b7a7e716242f046409c50a2e856e6ce6239d06d007ffd6072

Observation 8826c55f-c90b-4d34-afb6-e4ed7f29e5c4 · outbound

This paper cites Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity.

How Consistent Are LLM Agents? Measuring Behavioral Reproducibility in Multi-Step Tool-Calling Pipelines Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T21:40:08.985460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-07-04T21:40:04.594380Z digest=sha256:54c815081144d058a6691e0fe06c563a6f1f83e004ea2f19a564460c8e7874ca

Observation 58ea8c18-cbd7-4fe5-8569-fcf3b5f5f337 · outbound

This paper cites When Agents Disagree With Themselves: Behavioral Consistency as an Uncertainty Signal for LLM Agents.

How Consistent Are LLM Agents? Measuring Behavioral Reproducibility in Multi-Step Tool-Calling Pipelines When Agents Disagree With Themselves: Behavioral Consistency as an Uncertainty Signal for LLM Agents

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-16T01:21:45.045605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-07-04T21:40:04.594380Z digest=sha256:e4f0173223125dc2015d1c80807f31e661d7ca3481282026c24eca2ce2f9885b

Observation 5cc656bf-e0b3-437f-8d11-29abf3cd8675 · outbound

This paper cites Gorilla: Large Language Model Connected with Massive APIs.

How Consistent Are LLM Agents? Measuring Behavioral Reproducibility in Multi-Step Tool-Calling Pipelines Gorilla: Large Language Model Connected with Massive APIs

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T21:40:08.951739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-07-04T21:40:04.594380Z digest=sha256:9d109c3dab0b205e566d8db32eb94b828c17ad72da849e5ccbe1317be93a004d

Observation a0b67295-ebb5-4149-a60b-5683717d0fd4 · outbound

This paper cites True Few-Shot Learning with Language Models.

How Consistent Are LLM Agents? Measuring Behavioral Reproducibility in Multi-Step Tool-Calling Pipelines True Few-Shot Learning with Language Models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T21:40:08.975849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-07-04T21:40:04.594380Z digest=sha256:b341db2432056830ecfd253a1034da3a9c30eee45fcd0fe77cad0ae3e50b4ca1

Observation e3e80ddd-7a48-40a1-8a4f-eff821a61795 · outbound

This paper cites ToolLLM : Facilitating Large Language Models to Master 16000+ Real-World APIs.

How Consistent Are LLM Agents? Measuring Behavioral Reproducibility in Multi-Step Tool-Calling Pipelines ToolLLM : Facilitating Large Language Models to Master 16000+ Real-World APIs

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T21:40:08.988145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-07-04T21:40:04.594380Z digest=sha256:37fe925d9cdd323a3c70e1d29e3ffc72617b6e94da0a68d73ef82e46424ef13e

Observation 68a76fa4-6aeb-4cb1-9f99-57a638fb2337 · outbound

This paper cites Self-Reflection in LLM Agents: Effects on Problem-Solving Performance.

How Consistent Are LLM Agents? Measuring Behavioral Reproducibility in Multi-Step Tool-Calling Pipelines Self-Reflection in LLM Agents: Effects on Problem-Solving Performance

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-04T21:40:08.954429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-07-04T21:40:04.594380Z digest=sha256:740165cf00a0312a3322ddebfcbe6adc67cce752029ae4c93796fa4910085eca

Observation 7c2c9479-1b95-4a89-be6f-3fc238a75c18 · outbound

This paper cites Toolformer : Language Models Can Teach Themselves to Use Tools.

How Consistent Are LLM Agents? Measuring Behavioral Reproducibility in Multi-Step Tool-Calling Pipelines Toolformer : Language Models Can Teach Themselves to Use Tools

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T21:40:08.982920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-07-04T21:40:04.594380Z digest=sha256:a80160cffe875e7e010cb374d65a3c5651b24b5aff3d629d8301e01fea9d155b

Observation b750e918-9a7b-430d-a5cb-c5eef6908d31 · outbound

This paper cites Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting.

How Consistent Are LLM Agents? Measuring Behavioral Reproducibility in Multi-Step Tool-Calling Pipelines Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-04T21:40:08.966556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-07-04T21:40:04.594380Z digest=sha256:9694b4af3bce2defc67ac382b994350c5321773ddd8e9f13b72be57b890d13ca

Observation 50b41652-8c61-4d32-8590-7a1624e5e315 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

How Consistent Are LLM Agents? Measuring Behavioral Reproducibility in Multi-Step Tool-Calling Pipelines Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T21:40:08.983094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-07-04T21:40:04.594380Z digest=sha256:d8c2a38d9004d32b0a5ac9377793b159b66038c5e470b0b16e06b512041eb7fc

Observation 8da8965d-35cb-4628-a274-be211553ffdf · outbound

This paper cites Ethical and social risks of harm from Language Models.

How Consistent Are LLM Agents? Measuring Behavioral Reproducibility in Multi-Step Tool-Calling Pipelines Ethical and social risks of harm from Language Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-04T21:40:08.955390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-07-04T21:40:04.594380Z digest=sha256:5c902fd07f2507cff3dc4c210eea93f09ebed4a65098c489e235a572c8a395f4

Observation d33d085a-1893-442d-946c-b2933f12688c · outbound

This paper cites ReAct : Synergizing Reasoning and Acting in Language Models.

How Consistent Are LLM Agents? Measuring Behavioral Reproducibility in Multi-Step Tool-Calling Pipelines ReAct : Synergizing Reasoning and Acting in Language Models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T21:40:08.980637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-07-04T21:40:04.594380Z digest=sha256:180d20ca6ed474e6935946907d4f713f56b2777cb0c300c0fef113f026376a7e

Pith citing papers

Observation 1d6a2fe6-0b4f-4d8c-b9ae-2b04ab3f8272 · inbound

Determinants and Limits of LLM Security-Tool Orchestration: A Study with HexStrike-AI cites this paper.

Determinants and Limits of LLM Security-Tool Orchestration: A Study with HexStrike-AI How Consistent Are LLM Agents? Measuring Behavioral Reproducibility in Multi-Step Tool-Calling Pipelines

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-12T06:28:52.906906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:28:52.906906Z digest=sha256:9bba4998f3817ede7b1486c3150bde52deb617b4a83caa4449d878b4e60c62bf

Observation 8291d6d5-5a89-4c07-b573-664cdbabcc8f · inbound

OpenCodeReview: Determinism over Non-Determinism for Cost-Effective Agent-Based Code Review cites this paper.

OpenCodeReview: Determinism over Non-Determinism for Cost-Effective Agent-Based Code Review How Consistent Are LLM Agents? Measuring Behavioral Reproducibility in Multi-Step Tool-Calling Pipelines

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-11T20:08:14.595449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:08:14.363485Z digest=sha256:21b88696a429973ca09987e11827fd3dc4f748b86dae6fb8090688aea6871880

Observation 2cae528b-4854-46f2-851b-fdaec5cea5f4 · inbound

OpenCodeReview: Determinism over Non-Determinism for Cost-Effective Agent-Based Code Review cites this paper.

OpenCodeReview: Determinism over Non-Determinism for Cost-Effective Agent-Based Code Review How Consistent Are LLM Agents? Measuring Behavioral Reproducibility in Multi-Step Tool-Calling Pipelines

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-14T04:20:34.128103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:20:34.128103Z digest=sha256:836dd167d3546d17d2b5ea9853ca15bbbf52dcfc282a3eeefc71930eddc2b32b