Pith. sign in

Paper Citation Record · LEDGER

T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2312.14033.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.14033 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:17:46.338385Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:18:43.767493Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 44993b3d-a8d2-4b37-848d-8145c10d175b · inbound

InternLM2 Technical Report cites this paper.

InternLM2 Technical Report T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step

Reference 193

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:44:38.317801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T11:44:38.066501Z digest=sha256:d948e00dd570e1e500f47c950cbf874937d474607cc2634c3c1d68a2ac5704c9

Observation f8ecf101-dcae-4aee-8ff4-887f44a8337a · inbound

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations cites this paper.

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:46.338385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:46.338385Z digest=sha256:0c2905a43dada742761b6bd671681e4a803c7e1d462f2bd53b037b5207ca02b1

Observation 6512ffbd-b7dd-481e-88e9-49d88c10eb49 · inbound

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues cites this paper.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:38.371492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:38.371492Z digest=sha256:bcb8e6e76a1c9cc5e62376374b2217ef10f7fa52cdfcda72626b71268d0b3df4

Observation df1edc00-06ac-4a00-ba81-7c9cc172fbae · inbound

Teaching a Language Model to Speak the Language of Tools cites this paper.

Teaching a Language Model to Speak the Language of Tools T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:47.725080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:48:47.725080Z digest=sha256:db7ff4a4708d52b1587eda0979c32ebde3afc0a2d09f6185c0764869b58255ce

Observation 5e8428c8-01d7-4f9a-880f-66f929af1677 · inbound

A Survey of Context Engineering for Large Language Models cites this paper.

A Survey of Context Engineering for Large Language Models T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step

Reference 161

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:58:45.431204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T20:58:45.060041Z digest=sha256:c950fce28018b64a4c48b88b372c4aa16a3d93c2e9dd44a92fd595b7a85ab2fa

Observation 9c1b31bb-935b-4b46-8506-b73c01c8e05b · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:23:15.745837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:21140e3943aaeaef7dc5255a11b50bfb2c8e3699db76e3a36373741660668bcf

Observation af3d1478-45bd-4d65-b7e9-7e7899f8da05 · inbound

Evaluation and Benchmarking of LLM Agents: A Survey cites this paper.

Evaluation and Benchmarking of LLM Agents: A Survey T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.526623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.526623Z digest=sha256:062dcc16a4aeed9c32fec4d8ffe69dfbb0f20d0c37abfa97e74dc6c496dc81ea

Observation 0068ad15-fda3-4d1b-a81c-c670d82594d5 · inbound

Don't Start What You Can't Finish: A Counterfactual Audit of Support-State Triage in LLM Agents cites this paper.

Don't Start What You Can't Finish: A Counterfactual Audit of Support-State Triage in LLM Agents T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:57:15.534323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T07:55:07.930578Z digest=sha256:900fed7e6e67947f56f0aa629e02c62e39c2dbaf9cbc3bc6dca4c4ab66526980

Observation 144effff-7c90-45c6-bbfd-91d7e78da2a5 · inbound

Consistency as a Testable Property: Statistical Methods to Evaluate AI Agent Reliability cites this paper.

Consistency as a Testable Property: Statistical Methods to Evaluate AI Agent Reliability T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:41:21.751660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:41:15.286881Z digest=sha256:b72a1c326fe09137138c62215c000968417c48a129a2e7302b34714a1fc8cab5

Observation 69380883-da7f-4587-9e82-7b4c290fbf28 · inbound

Holistic Evaluation and Failure Diagnosis of AI Agents cites this paper.

Holistic Evaluation and Failure Diagnosis of AI Agents T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:19:43.250065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T03:17:28.794622Z digest=sha256:9d95820b44f80d2653e6a48701b543d7ab5a64169245efaa0dc3e9d28a162bd1

Observation a56ed124-4fff-42de-b142-8e02ced8ff85 · inbound

Synthesize and Reward -- Reinforcement Learning for Multi-Step Tool Use in Live Environments cites this paper.

Synthesize and Reward -- Reinforcement Learning for Multi-Step Tool Use in Live Environments T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:36:29.700075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T09:54:00.111238Z digest=sha256:2d86cc6c554973510b8352ad794256485703b4f0515f726ffa3b82ae4263b002

Observation 326e0d88-bee7-4345-99f9-7c3999f2fdd0 · inbound

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation cites this paper.

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step

Reference 227

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:18:43.768892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T04:19:26.332718Z digest=sha256:40c58d7d95113f8219d227b267536e97d604eaa6b532c4edf858e1d9df37d8cd

Observation 7b596bd0-182b-46a4-af88-d51e1343a432 · inbound

Qwen-Audio-VAE Technical Report cites this paper.

Qwen-Audio-VAE Technical Report T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step

Reference 234

Resolution
unresolved
no resolver link, observed 2026-07-14T03:31:19.309532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:31:19.309532Z digest=sha256:761f4c0e150eb964c5105a7f57536a38fb73e648c01ce15de77f98134c4ada34

Observation ee8046a5-274a-4651-92ce-9272532307be · inbound

WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing cites this paper.

WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T01:35:09.624996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:35:09.624996Z digest=sha256:2f9b248c9d97f211d20161cde54eb65de00e20d1895cce14907627408978c233