Pith. sign in

Paper Citation Record · LEDGER

MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2309.10691.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.10691 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:07:08.786395Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

16
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation dd0d9f5f-27ce-4fbb-81b7-6f27b5f70e55 · inbound

ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production cites this paper.

ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-22T15:44:58.119606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T15:42:05.266854Z digest=sha256:5af16fb8c14cf35f469ee144ce257d0e8d79adc0a1cf927a1e5004442038a601

Observation dbc2158c-70d8-4b05-a080-05b7abb71823 · inbound

ChemGraph: An Agentic Framework for Computational Chemistry Workflows cites this paper.

ChemGraph: An Agentic Framework for Computational Chemistry Workflows MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:08.786395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:08.786395Z digest=sha256:4c77166a0f814803744d1721d623936dce43f31978e0dfa6b98d84d49086fd0b

Observation fe819c0f-4e05-4e74-b792-821941c5a143 · inbound

A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy cites this paper.

A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T04:52:00.986893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:52:00.986893Z digest=sha256:28b1f3f577e52b7534e8d468b8a0564d53268e6b80895262780acb28b182582e

Observation be19d469-391a-47f3-8e42-b4a525f8d81e · inbound

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey cites this paper.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 190

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.896513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.896513Z digest=sha256:c912764269f5e196aa145401c2c5e9dded0dd234f945ec781250348b52c7ffd5

Observation 27374b34-1684-49b9-9c3e-101ca3db1365 · inbound

Automating Financial Statement Audits with Large Language Models cites this paper.

Automating Financial Statement Audits with Large Language Models MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:15.549331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:53:15.549331Z digest=sha256:919a56a4d579f6ffafccb7357183e9248439c429f5e12d43a2849dae5a397218

Observation 1b717073-e197-45fc-ad3f-e3a5e158c86c · inbound

DrafterBench: Benchmarking Large Language Models for Tasks Automation in Civil Engineering cites this paper.

DrafterBench: Benchmarking Large Language Models for Tasks Automation in Civil Engineering MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T17:11:33.977624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:11:33.977624Z digest=sha256:7ee547704eea1406c426ba88b37da950d36a5f4e51f21f18748bb2c53de1c074

Observation 7e0cb956-51fa-47cc-be48-1410d9573aba · inbound

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning cites this paper.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:04.918820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:04.918820Z digest=sha256:250e752de7e0a4a0cd764e2377fd49792c51af7bf9ce011e296bec090561dc39

Observation 2dce9817-e5db-4adb-a203-2e56ebe70506 · inbound

An Empirical Study of Testing Practices in Open Source AI Agent Frameworks and Agentic Applications cites this paper.

An Empirical Study of Testing Practices in Open Source AI Agent Frameworks and Agentic Applications MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:12:39.816366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T14:12:08.776876Z digest=sha256:cd19a27a88f5ffeb41e690121733b5792d8a2dfc4139357479b71de30aa76e7b

Observation 54ca4c07-5857-4c29-adaa-0985a18cbca1 · inbound

Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live cites this paper.

Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-18T02:00:39.797394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T01:58:23.234348Z digest=sha256:271d786e6f747848b465237899d50a20ec84c73e3f9d050e40fce06e4991c75a

Observation f4751383-d045-463c-be3b-321ec00feebb · inbound

Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live cites this paper.

Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T00:19:21.129190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:19:21.129190Z digest=sha256:eb83c67e1ac83c41e4fe28f6723b58d8553f8a8466cf41211a8e81498ef5319b

Observation 0164b675-c62d-424f-afc9-63de5f2c3494 · inbound

Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language Models cites this paper.

Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language Models MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T11:44:07.135108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:44:07.135108Z digest=sha256:8b600185b28abaf294b286ad7fac2fb5a034c8a64dedaa9f6d3d549fc8eb43b9

Observation a29eb7da-c3fa-4131-9675-0118d26be51c · inbound

LLM-Based Agentic Systems for Software Engineering: Challenges and Opportunities cites this paper.

LLM-Based Agentic Systems for Software Engineering: Challenges and Opportunities MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T10:30:41.526303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:30:41.526303Z digest=sha256:12b786f514ad36ba0066dd787f608fbdf06159a59e914e3d0fa7456296eb42a0

Observation d32cf886-8fce-4c05-b174-d866c7b8c931 · inbound

Qualixar OS: A Universal Operating System for AI Agent Orchestration cites this paper.

Qualixar OS: A Universal Operating System for AI Agent Orchestration MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:00:55.622769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:44:25.389675Z digest=sha256:f3040d221ebdca727804e6c81ea44bbb4037b9c38c06dee8d5f72281037ea0a4

Observation f05026b1-8666-411f-8889-4c4da5e839d3 · inbound

Evaluating Temporal Consistency in Multi-Turn Language Models cites this paper.

Evaluating Temporal Consistency in Multi-Turn Language Models MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:36:12.942416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T11:37:10.330674Z digest=sha256:9257e44a18f79589a6e6fad3a9caf481bf2b41937c6575ded486e81c5895ec3e

Observation fc8facb9-335a-4d59-a160-5a2a4d7691a5 · inbound

AgentFloor: How Far Up the tool use Ladder Can Small Open-Weight Models Go? cites this paper.

AgentFloor: How Far Up the tool use Ladder Can Small Open-Weight Models Go? MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:21:10.587173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T20:09:11.566825Z digest=sha256:d6c879607874c9f9027d223b6e4751fc5c711e0d4702355aefeedc7aec8ac1d8

Observation 29dad227-b50f-4304-88af-6fe780563325 · inbound

To Call or Not to Call: A Framework to Assess and Optimize LLM Tool Calling cites this paper.

To Call or Not to Call: A Framework to Assess and Optimize LLM Tool Calling MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:36:10.316387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:32:57.054584Z digest=sha256:f2fcb20ed7b9d92023cb40c949d20000a577c1907b32bdee9d655a89cf4cf154

Observation 3ac429c7-71f4-4082-8f33-8e3c45480718 · inbound

A Language for Describing Agentic LLM Contexts cites this paper.

A Language for Describing Agentic LLM Contexts MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:21:09.182239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T17:18:16.256307Z digest=sha256:067a11cdd70e73cadc3142e31432fadb60a6612e59260049299855f7ea58b56d

Observation 077c04f4-2078-4930-8337-d5cb9dfc6c8d · inbound

Consistency as a Testable Property: Statistical Methods to Evaluate AI Agent Reliability cites this paper.

Consistency as a Testable Property: Statistical Methods to Evaluate AI Agent Reliability MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:41:21.833327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:41:15.286881Z digest=sha256:ddaac93c71c135f245108fdc3bcc773b534ef6febe7a97e3cc4de76474a604a3

Observation 310ebe67-b272-46ba-a4de-a4da00752f82 · inbound

The Scaling Laws of Skills in LLM Agent Systems cites this paper.

The Scaling Laws of Skills in LLM Agent Systems MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-20T18:13:37.568810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T18:10:08.737710Z digest=sha256:9cc043f47b6977fcbb0641ce94357d52dc901487d001a9e5f06df47c8fe89ec2

Observation 23514e88-d818-4ff8-81da-582867fa3c7a · inbound

Interactive Evaluation Requires a Design Science cites this paper.

Interactive Evaluation Requires a Design Science MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:58:14.000188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T10:55:08.135630Z digest=sha256:f662ab4e4f5dab904d858f96d44dc12513a88803f0dbb2b01ca51d43c5a53251

Observation 90999811-e646-45b4-a6bc-ed75fca53284 · inbound

SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science cites this paper.

SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:38:12.096119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T10:36:09.724234Z digest=sha256:038c290d4618d67314a85dde10c3f32750596f467d7c4c75d86590791f29c749

Observation 63860cd3-a857-4659-a905-39e5599c0da3 · inbound

SeDT: Sentence-Transformer Decision-Transformer Conditioning for Multi-Turn Conversation Reliability cites this paper.

SeDT: Sentence-Transformer Decision-Transformer Conditioning for Multi-Turn Conversation Reliability MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T18:53:51.659118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T18:45:40.697579Z digest=sha256:1f61479d2af79a0fd72c1b3c35358cefd9ffb46a1b6328dd7e2b0699f9bb3453

Observation 6cb1c69f-d869-4e53-bca4-624b15c3ff28 · inbound

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns cites this paper.

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:19:23.870272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:47:59.090249Z digest=sha256:69132980bd9aa4fffa0811cedadc62b40755e9dcdda76ed17fd13e0f09131362

Observation 3fa3b650-be81-43d9-b8f2-967f5167901d · inbound

What Drives Interactive Improvement from Feedback? cites this paper.

What Drives Interactive Improvement from Feedback? MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T12:05:43.465485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T02:36:20.649687Z digest=sha256:2f13f12f70ea894443c2d88fa15f271be89d4f0e3265b2b12828779d08959dab

Observation 7761b0a8-c03b-43d6-9ee7-f924f4de7fff · inbound

DataClawEval: A Benchmark for Data Engineering Agents in Real Industrial Harness cites this paper.

DataClawEval: A Benchmark for Data Engineering Agents in Real Industrial Harness MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-31T19:54:08.546745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T19:54:08.546745Z digest=sha256:1e4c9254143fb8c22f4fa075f21de7a9e57ea15193896ba067a7e2c829d2761d