Pith. sign in

Paper Citation Record · LEDGER

MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2309.10691.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.10691 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:07:08.786395Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

16
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation dd0d9f5f-27ce-4fbb-81b7-6f27b5f70e55 · inbound

ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production cites this paper.

ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-22T15:44:58.119606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T15:42:05.266854Z digest=sha256:04b078e2d431391186cc5a9cc8a4c36a2cf63d13fdfaf5956865844f6de6b77c

Observation dbc2158c-70d8-4b05-a080-05b7abb71823 · inbound

ChemGraph: An Agentic Framework for Computational Chemistry Workflows cites this paper.

ChemGraph: An Agentic Framework for Computational Chemistry Workflows MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:08.786395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:08.786395Z digest=sha256:b90b5c17e17cd4f5ef392e629df7d0e3a9f23a30ed7f1809fcdb99da590ffe8a

Observation fe819c0f-4e05-4e74-b792-821941c5a143 · inbound

A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy cites this paper.

A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T04:52:00.986893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:52:00.986893Z digest=sha256:b114044ec095adae2b2caf944d4d948b747b6752e7aa1a73e53c35aae3190659

Observation be19d469-391a-47f3-8e42-b4a525f8d81e · inbound

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey cites this paper.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 190

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.896513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.896513Z digest=sha256:03104692b3a5cb0f06d72a1359f12438230add153e2f09294d8ceb41457a2cee

Observation 27374b34-1684-49b9-9c3e-101ca3db1365 · inbound

Automating Financial Statement Audits with Large Language Models cites this paper.

Automating Financial Statement Audits with Large Language Models MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:15.549331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:53:15.549331Z digest=sha256:ccf0ff7d4dc806408ebf869087abe7a3e2b8b1ac0c1cf596b23104a88d8310e7

Observation 1b717073-e197-45fc-ad3f-e3a5e158c86c · inbound

DrafterBench: Benchmarking Large Language Models for Tasks Automation in Civil Engineering cites this paper.

DrafterBench: Benchmarking Large Language Models for Tasks Automation in Civil Engineering MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T17:11:33.977624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:11:33.977624Z digest=sha256:c94f18142ad365ef562fabb48a2a1ea2944d17a1b49153339bf07555217ac53f

Observation 7e0cb956-51fa-47cc-be48-1410d9573aba · inbound

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning cites this paper.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:04.918820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:04.918820Z digest=sha256:77fe024b432f253d676f44c84c09bda2280d4406a01b25c85f8d25c04c14a133

Observation 2dce9817-e5db-4adb-a203-2e56ebe70506 · inbound

An Empirical Study of Testing Practices in Open Source AI Agent Frameworks and Agentic Applications cites this paper.

An Empirical Study of Testing Practices in Open Source AI Agent Frameworks and Agentic Applications MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:12:39.816366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T14:12:08.776876Z digest=sha256:74c3a6c3257481a0f965e9fda8cbd6e9912505956a87221c224e073fa66c5dfb

Observation 54ca4c07-5857-4c29-adaa-0985a18cbca1 · inbound

Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live cites this paper.

Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-18T02:00:39.797394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T01:58:23.234348Z digest=sha256:5cd0164d3c6af1b31fda15c3b40200956da926e97ab0ba73795ce8ae718c1a69

Observation f4751383-d045-463c-be3b-321ec00feebb · inbound

Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live cites this paper.

Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T00:19:21.129190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:19:21.129190Z digest=sha256:07e17a73f54567bedfd00c306a06f02701aec35c123dd5cc99691be6ae90ff5e

Observation 0164b675-c62d-424f-afc9-63de5f2c3494 · inbound

Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language Models cites this paper.

Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language Models MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T11:44:07.135108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:44:07.135108Z digest=sha256:583d2815d8ed7177b06c850893efa0eb1c64e688e0a706ec747231fc7295ec49

Observation a29eb7da-c3fa-4131-9675-0118d26be51c · inbound

LLM-Based Agentic Systems for Software Engineering: Challenges and Opportunities cites this paper.

LLM-Based Agentic Systems for Software Engineering: Challenges and Opportunities MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T10:30:41.526303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:30:41.526303Z digest=sha256:2b846d48dd1da10f36c42edd3acf216ee9280f3ccb39d5a935560a85e4e3a589

Observation d32cf886-8fce-4c05-b174-d866c7b8c931 · inbound

Qualixar OS: A Universal Operating System for AI Agent Orchestration cites this paper.

Qualixar OS: A Universal Operating System for AI Agent Orchestration MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:00:55.622769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:44:25.389675Z digest=sha256:0a35f15e7b4277f8e67e94f39577648f52a0fdef11f67a61588cb34178345ea5

Observation f05026b1-8666-411f-8889-4c4da5e839d3 · inbound

Evaluating Temporal Consistency in Multi-Turn Language Models cites this paper.

Evaluating Temporal Consistency in Multi-Turn Language Models MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:36:12.942416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T11:37:10.330674Z digest=sha256:6bccaffd0d160dde27a4ca182836fabd90361627a6340f713bd9172c45724a39

Observation fc8facb9-335a-4d59-a160-5a2a4d7691a5 · inbound

AgentFloor: How Far Up the tool use Ladder Can Small Open-Weight Models Go? cites this paper.

AgentFloor: How Far Up the tool use Ladder Can Small Open-Weight Models Go? MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:21:10.587173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-09T20:09:11.566825Z digest=sha256:549d85bdaf3a6c2214ecab4e9c135a8baba097155e033f12201a80681e256f22

Observation 29dad227-b50f-4304-88af-6fe780563325 · inbound

To Call or Not to Call: A Framework to Assess and Optimize LLM Tool Calling cites this paper.

To Call or Not to Call: A Framework to Assess and Optimize LLM Tool Calling MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:36:10.316387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-09T19:32:57.054584Z digest=sha256:17f8602cc65c18891ac000ba895f4d28d49822530cdd13814026ecce02c765b1

Observation 3ac429c7-71f4-4082-8f33-8e3c45480718 · inbound

A Language for Describing Agentic LLM Contexts cites this paper.

A Language for Describing Agentic LLM Contexts MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:21:09.182239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T17:18:16.256307Z digest=sha256:506ed4f9c0c901ad0b55e166551fbb11f6f083ab542063a2e53c9fd5efd9096c

Observation 077c04f4-2078-4930-8337-d5cb9dfc6c8d · inbound

Consistency as a Testable Property: Statistical Methods to Evaluate AI Agent Reliability cites this paper.

Consistency as a Testable Property: Statistical Methods to Evaluate AI Agent Reliability MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:41:21.833327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:41:15.286881Z digest=sha256:31cee42ea986563a9b7f682a889bed4dad618b109126503fa6edb7fa02457e76

Observation 310ebe67-b272-46ba-a4de-a4da00752f82 · inbound

The Scaling Laws of Skills in LLM Agent Systems cites this paper.

The Scaling Laws of Skills in LLM Agent Systems MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-20T18:13:37.568810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T18:10:08.737710Z digest=sha256:faa754eb9dad15e1712fd0073f2eac86fe3366e342ff84bdbb8bd6f8db65ae43

Observation 23514e88-d818-4ff8-81da-582867fa3c7a · inbound

Interactive Evaluation Requires a Design Science cites this paper.

Interactive Evaluation Requires a Design Science MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:58:14.000188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T10:55:08.135630Z digest=sha256:0a9b94586a15100a9517f59408d98567520baf16b857d46cbb18bd89712b579f

Observation 90999811-e646-45b4-a6bc-ed75fca53284 · inbound

SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science cites this paper.

SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:38:12.096119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T10:36:09.724234Z digest=sha256:c02c22016a8bb7bdedd1fe6e16fdefe0e6226a10a5a52637f46d88303ca5d9e1

Observation 63860cd3-a857-4659-a905-39e5599c0da3 · inbound

SeDT: Sentence-Transformer Decision-Transformer Conditioning for Multi-Turn Conversation Reliability cites this paper.

SeDT: Sentence-Transformer Decision-Transformer Conditioning for Multi-Turn Conversation Reliability MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T18:53:51.659118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T18:45:40.697579Z digest=sha256:4553277a185da817a369ead347789ca94ff138ff6a2834c7b7b217465405815d

Observation 6cb1c69f-d869-4e53-bca4-624b15c3ff28 · inbound

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns cites this paper.

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:19:23.870272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T19:47:59.090249Z digest=sha256:b3f4356194d8d318c836e31fbddde8b8fc78e44feaa9a73a94c171050edd356d

Observation 3fa3b650-be81-43d9-b8f2-967f5167901d · inbound

What Drives Interactive Improvement from Feedback? cites this paper.

What Drives Interactive Improvement from Feedback? MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T12:05:43.465485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T02:36:20.649687Z digest=sha256:58295ad79a29d33a28b0bff8aaf8e6d7e4069f208209933b7b8a37f2b997fc72

Observation 7761b0a8-c03b-43d6-9ee7-f924f4de7fff · inbound

DataClawEval: A Benchmark for Data Engineering Agents in Real Industrial Harness cites this paper.

DataClawEval: A Benchmark for Data Engineering Agents in Real Industrial Harness MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-31T19:54:08.546745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T19:54:08.546745Z digest=sha256:f1c691fc43d9b0fae4297a6cca8962e1b3f31f4bdd31b02600f2948e45097b56