Pith. sign in

Paper Citation Record · LEDGER

MAVEN: Improving Generalization in Agentic Tool Calling

As of 7 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 1 inbound Pith citation observation for arXiv:2605.30738.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.30738 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T22:42:11.459778Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T09:42:05.329537Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact11
  • verified fuzzy0
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 158a24d5-33b9-4218-b507-6547329b43dd · outbound

This paper cites gpt-oss-120b & gpt-oss-20b Model Card.

MAVEN: Improving Generalization in Agentic Tool Calling gpt-oss-120b & gpt-oss-20b Model Card

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T22:42:46.143973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T22:42:11.459778Z digest=sha256:7d40e422b52bf45f3ae4cc8da0d3ed2f6ed71b6d1e9990e4c84f489f5eea2643

Observation ec0e19cc-d42f-4d79-ad99-4bff5dda2520 · outbound

This paper cites HealthBench: Evaluating Large Language Models Towards Improved Human Health.

MAVEN: Improving Generalization in Agentic Tool Calling HealthBench: Evaluating Large Language Models Towards Improved Human Health

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T22:42:46.104453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T22:42:11.459778Z digest=sha256:023771eb6322d296fe9bab17057404fb33e8b2e774b3b8c5bddbe10ef04594ef

Observation 36137cd2-46f3-4581-959a-cbf91d22d2db · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

MAVEN: Improving Generalization in Agentic Tool Calling $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-28T22:42:46.123403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T22:42:11.459778Z digest=sha256:fa1c90c21d4b22e91e7b0cedc08712dcde77c6d66d437a851285b9103aab3b6b

Observation b03ccce6-64ce-4c26-b971-2d567fae845e · outbound

This paper cites Acebench: Who wins the match point in tool usage? arXiv preprint arXiv:2501.12851, 2025 a.

MAVEN: Improving Generalization in Agentic Tool Calling Acebench: Who wins the match point in tool usage? arXiv preprint arXiv:2501.12851, 2025 a

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:42:46.107975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T22:42:11.459778Z digest=sha256:f49a7a8987b28419cfea3494ab74599d30934aa3a9b0908f1318bdeb7908997a

Observation 2abfce88-501d-4451-86b0-31dbfc20acb5 · outbound

This paper cites Towards general agentic intelligence via environment scaling.arXiv preprint arXiv:2509.13311.

MAVEN: Improving Generalization in Agentic Tool Calling Towards general agentic intelligence via environment scaling.arXiv preprint arXiv:2509.13311

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:42:46.129089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T22:42:11.459778Z digest=sha256:a30009f493cbb4987421cc76959a375a78aaebc85f01b76143c4ffae7a5b3227

Observation cacb1980-62da-4e63-b519-bb6e4d88023a · outbound

This paper cites On Robustness and Reliability of Benchmark-Based Evaluation of LLMs.

MAVEN: Improving Generalization in Agentic Tool Calling On Robustness and Reliability of Benchmark-Based Evaluation of LLMs

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:42:46.135066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T22:42:11.459778Z digest=sha256:251ae18ec2687bb443567b3ca62df65276b0983d6a6b230107cfae14d212c9d5

Observation ce3c2b12-c59c-4d68-a973-4e6afe0b330d · outbound

This paper cites Exploring Code Analysis: Zero-Shot Insights on Syntax and Semantics with LLMs.

MAVEN: Improving Generalization in Agentic Tool Calling Exploring Code Analysis: Zero-Shot Insights on Syntax and Semantics with LLMs

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-28T22:42:46.141231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T22:42:11.459778Z digest=sha256:b77c033f0295f411494df8c9cbe50a5053d9fb09610d9bd633030055f81ac8ec

Observation 8a497ae1-a56e-4b3c-a908-df4ec41bc80c · outbound

This paper cites A Survey on Large Language Model Benchmarks.

MAVEN: Improving Generalization in Agentic Tool Calling A Survey on Large Language Model Benchmarks

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:42:46.138266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T22:42:11.459778Z digest=sha256:d02ebefa05b7e2481b85b308527514694fb35c8b7db43a109e06c6e5019a827f

Observation 4cdd3aa4-380b-452a-8cc8-13d1571aec9e · outbound

This paper cites Ac- cessed: 2025-10-06.

MAVEN: Improving Generalization in Agentic Tool Calling Ac- cessed: 2025-10-06

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-28T22:42:11.459778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:42:11.459778Z digest=sha256:8fa7e53cd71a4c393ea4522d7f4d0a290b451bbdc50d4c40907e748b2a25d780

Observation 14dcb97f-1f27-4e5a-ad39-ba703115f706 · outbound

This paper cites ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs.

MAVEN: Improving Generalization in Agentic Tool Calling ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-28T22:42:46.114300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T22:42:11.459778Z digest=sha256:17c037943753415f7a4f68ade0a2fa969b7009cf87485b1b099c4d7bf485d132

Observation 6939568b-f256-4022-9292-a412fca46502 · outbound

This paper cites and Tavor, A.

MAVEN: Improving Generalization in Agentic Tool Calling and Tavor, A

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-28T22:42:11.459778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:42:11.459778Z digest=sha256:8d04cd7efb40a55804b93a789e8cfef94b5c92009db30e3567d044f228b7a157

Observation 20bb1bdf-8214-4d9c-89ed-5968b9f0ad95 · outbound

This paper cites OpenAI GPT-5 System Card.

MAVEN: Improving Generalization in Agentic Tool Calling OpenAI GPT-5 System Card

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-06-28T22:42:46.126021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T22:42:11.459778Z digest=sha256:ac603e000053edca8752b26a160119f5142b21466f035512c1f5fdb6e8f91db9

Observation 7db4f82b-7ab8-4fb7-8e5b-9edf78dc28ad · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

MAVEN: Improving Generalization in Agentic Tool Calling Kimi K2: Open Agentic Intelligence

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-06-28T22:42:46.146810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T22:42:11.459778Z digest=sha256:31736b28c8b904006b11cd89adcc455971c4441f0b078c6d917f3ef2ccfde36d

Observation 52b1eff2-1d14-4594-998d-684ede8daa10 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

MAVEN: Improving Generalization in Agentic Tool Calling ReAct: Synergizing Reasoning and Acting in Language Models

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T22:42:46.131868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T22:42:11.459778Z digest=sha256:511160f22c0e10df4be05a2957a744d746c053de156f6967a339d3ff1931c755

Observation 48cd2a89-b5e4-41b8-b8b2-86bdd3ccf020 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

MAVEN: Improving Generalization in Agentic Tool Calling $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-06-28T22:42:46.120670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T22:42:11.459778Z digest=sha256:278e1effa3be61d6b600094474669d372ef8b7e0ba42b41365c6b19ac26179b7

Observation 826a85f9-54a9-40f1-8c8c-f40160d5e64b · outbound

This paper cites GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models.

MAVEN: Improving Generalization in Agentic Tool Calling GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-06-28T22:42:46.149739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T22:42:11.459778Z digest=sha256:ff8d3bfedee108eb0591ec404bcc5e4b92f9740ead2674f5faf8379a533a6945

Pith citing papers

Observation 2ed10347-f040-4e56-9800-24b1fe5ca8d7 · inbound

Benchmarking the Benchmarks: A Validity Audit of Tool-Calling Evaluation cites this paper.

Benchmarking the Benchmarks: A Validity Audit of Tool-Calling Evaluation MAVEN: Improving Generalization in Agentic Tool Calling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T09:42:05.329537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:42:05.329537Z digest=sha256:3f333bb4e5cdb2768805473a5728836799b19cdf2115b010d79ea1dba5c008e5