Pith. sign in

Paper Citation Record · LEDGER

Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2412.15606.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.15606 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:03:30.207465Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T06:36:52.552598Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 32f1ef0a-fe84-4df6-b804-d17cca32af85 · inbound

Allen: Rethinking MAS Design through Step-Level Policy Autonomy cites this paper.

Allen: Rethinking MAS Design through Step-Level Policy Autonomy Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T20:03:30.207465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:03:30.207465Z digest=sha256:97ca3f9cbb6a41b7ba10b9999f41a32dc5f79d612e7d0a64062cfe45ff35462a

Observation b9354345-b786-4627-a0ad-8a670e27034e · inbound

Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents cites this paper.

Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:06:17.811876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T11:02:55.529271Z digest=sha256:d5acd22eb26e92e0ada4fb6f96e0fb401194136f10f79e55996c9272316b05b3

Observation f2ccbbe1-25d1-455d-b2f5-9b54e8317461 · inbound

Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents cites this paper.

Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T12:44:12.396331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:44:12.396331Z digest=sha256:35768038b35b6e378b2602ed141b7d03f36c24b295b52ac70f2dcaa0d68b67ad

Observation 0e073128-d548-4e93-85b2-679c8043e069 · inbound

A Multimodal Foundation Model of Spatial Transcriptomics and Histology for Biological Discovery and Clinical Prediction cites this paper.

A Multimodal Foundation Model of Spatial Transcriptomics and Histology for Biological Discovery and Clinical Prediction Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-13T12:44:39.273341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:44:39.273341Z digest=sha256:d40d9c5b6a27140a62d44946ae0a10f28b9b2186b1e2dcaba3e13cb8ce2522d4

Observation 07a379d2-316c-441e-99c0-48a429f3d6e6 · inbound

Single-agent vs. Multi-agents for Automated Video Analysis of On-Screen Collaborative Learning Behaviors cites this paper.

Single-agent vs. Multi-agents for Automated Video Analysis of On-Screen Collaborative Learning Behaviors Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:33:02.306118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T17:32:06.256142Z digest=sha256:8b59cf5322725567b1b08a0ff6d3f4a304b5c6379c7e4c0146e5b182dc92874c

Observation bce18e27-d410-4a35-b9eb-1e7f402d7d0c · inbound

ReGRPO: Reflection-Augmented Policy Optimization for Tool-Using Agents cites this paper.

ReGRPO: Reflection-Augmented Policy Optimization for Tool-Using Agents Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:35:41.873614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T05:23:21.162004Z digest=sha256:9b2b4f35c11daa6a8d64785952643d6d5d584b4d50aa52bfe50575c880794822

Observation 614e779c-392d-4a80-a56f-5936bfe66faf · inbound

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing cites this paper.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.553796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:b2eb38ef9122d3c1b22d6b1afd78c638e19503c83c0259a20e77652098e7aff9

Observation 8d3a32d2-69e5-40f1-94b5-dd933246abd9 · inbound

ToolAnchor: Anchoring Counterfactual Context to Boost Agentic Tool-use Capability cites this paper.

ToolAnchor: Anchoring Counterfactual Context to Boost Agentic Tool-use Capability Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T06:36:48.974843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:36:48.974843Z digest=sha256:abb2f866d2bf2752db64b2143d2c96cf3cc39d6dae247baf4271851bfdb4582b