Pith. sign in

Paper Citation Record · LEDGER

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

As of 8 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 10 inbound Pith citation observations for arXiv:2505.20289.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20289 v2

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:59:40.120459Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:30:03.053598Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.147388Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bdd3af60-36f6-4ff6-9905-7332016ec311 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.679510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.679510Z digest=sha256:6f17dbbc85e05249954aaffd62f1034a54c1c05d601c42fb6fbd7e7a1cece505

Observation d08f3d56-2dfc-49e1-80c0-075079462928 · outbound

This paper cites Language models are few-shot learners.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Language models are few-shot learners

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:43.052275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:59:35.772704Z digest=sha256:e4e4e7cd50378f5b51bc35f7c3debd04aecada55827fa82d6696869787cdcda6

Observation ab323fc2-493c-4d18-b486-9882846af7be · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection LLaMA: Open and Efficient Foundation Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.839303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.839303Z digest=sha256:79a424997a8f32c0d5e866ee2ab255110b9eff46e0f257bc6bcc43788b516b3c

Observation cdba0ffa-9c85-4a20-a6ba-90a550228e0e · outbound

This paper cites GPT-4 Technical Report.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection GPT-4 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.888788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.888788Z digest=sha256:7c9db2947ac1c83c7f7f535c1e0633e0a9f98a70a77a1a1c63b5a6ed07443616

Observation 688941ca-0a04-47d9-b69b-2e2700a5f615 · outbound

This paper cites Visual instruction tuning.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Visual instruction tuning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.927742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.927742Z digest=sha256:7aede84de4e0973dfcbc6d667a7a3f51195c4424e62f836c263bb3cfa61d28d1

Observation afde5757-ba96-4b33-b939-a1c33aab25b7 · outbound

This paper cites Qwen2.5 Technical Report.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Qwen2.5 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:36.008311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:36.008311Z digest=sha256:d84292a596d7781f10662934daad9f446b5ce5299f70f7789d15682cbd5c2aca

Observation e01f5360-5b6f-4a5b-b824-68d7ebcb87cc · outbound

This paper cites GPT-4o System Card.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection GPT-4o System Card

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:36.102797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:36.102797Z digest=sha256:64f2c7f1db417de108f57838790f8d51bc74b0096f58f151020d0a552ff8e5ef

Observation 3f8024ca-e64d-44f5-8ce5-d7d481dfac4d · outbound

This paper cites Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:36.173060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:36.173060Z digest=sha256:874a9e57323796202ed012b9d4ad0a9743e1172f887f465a88d72b00c9a4fce0

Observation 8de9d6e9-170e-4728-98a1-1e4079043ccc · outbound

This paper cites ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:36.266705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:36.266705Z digest=sha256:d294765b2790ad240e998ba91b7e8707b2f30179b943aaac2f03c1c524a1d77b

Observation a73f13ec-d452-470f-9684-e0d8ba91a76f · outbound

This paper cites Pal: Program-aided language models.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Pal: Program-aided language models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:36.332759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:36.332759Z digest=sha256:1a80d774b022d6a243b455fe7f231f7ed2796d4639cf32bd29a9f41a7ed0cc7c

Observation 5e26eaad-20dc-4f94-a5a8-e12345f4ce40 · outbound

This paper cites Gorilla: Large language model connected with massive apis.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Gorilla: Large language model connected with massive apis

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:42.818668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:59:36.414302Z digest=sha256:afda64e910d2985af6b78323ec0ccc3b5f5d50cece406b0b7f4244a52514a460

Observation 2fa008a5-8011-47d7-abef-8f08ead072a1 · outbound

This paper cites Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:36.478765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:36.478765Z digest=sha256:5fbee02779c6d6f9d95b6b1ff77a15669e8a17e9eea954b0efb98517121c9a6b

Observation 8875f002-9e09-4fbc-8494-38984220fb19 · outbound

This paper cites Visual programming: Compositional visual reasoning without training.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Visual programming: Compositional visual reasoning without training

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:36.556198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:36.556198Z digest=sha256:c6b125618b9e9c4dccd075116b51e6a9561b3f8ce5361151b50d2cacd8d28573

Observation 9de7870e-40fa-46a8-8b5c-b1aa33c00c86 · outbound

This paper cites Vipergpt: Visual inference via python execution for reasoning.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Vipergpt: Visual inference via python execution for reasoning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:42.614098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:59:36.625081Z digest=sha256:97e3543c2d72a988ccb693de963093c6ef74813fc1396dfdf5e9376903f5b5ff

Observation bff801ff-b139-4cb6-929f-eeb09cfb4686 · outbound

This paper cites Toolformer: Language models can teach themselves to use tools.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Toolformer: Language models can teach themselves to use tools

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:42.461624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:59:36.697970Z digest=sha256:c1ccb9c769591a9efedf6abfaeffff3a86a4d02165604922c9c463434e89fcaf

Observation 6da95cc0-cc6c-4f67-96b8-8da502b1191f · outbound

This paper cites Llava-plus: Learning to use tools for creating multimodal agents.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Llava-plus: Learning to use tools for creating multimodal agents

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:36.748773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:36.748773Z digest=sha256:f95146994574e45829342b72ae1fdd9ccbd80ac7d17db6a0fc35a91cbe43d809

Observation 98792aca-f97f-4bc4-baaa-cb14ae54f20e · outbound

This paper cites OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:36.843545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:36.843545Z digest=sha256:0b5b85d055499d2b850f376a0612435c020155b7e23ae83af9650a12675aa263

Observation 4dd80ee7-9102-4d5c-b7ef-a280e31fca07 · outbound

This paper cites Reinforcement learning: An introduction, volume 1.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Reinforcement learning: An introduction, volume 1

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:36.900888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:36.900888Z digest=sha256:e0d010d230962256ba423a7df42d8448ccdcb2161beb3e3baf84da735274ee01

Observation 72fc4b0c-cf1c-40de-a9e1-66826631bbe0 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:36.982009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:36.982009Z digest=sha256:058c80f55718c4963f6cc6c34213f599200574eb80ee74bc286547779738fcc2

Observation 14f40af1-03f3-4515-ba06-7b0a00e206f9 · outbound

This paper cites Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:37.062514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:37.062514Z digest=sha256:bf5f1ea6783981ba5fd8059e4e804bdea199862ea769ff4987e1426cc6fc4f6e

Observation 28f6d4d3-57c7-46a2-bf52-6987a05e02fe · outbound

This paper cites Internet-Augmented Dialogue Generation.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Internet-Augmented Dialogue Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:37.132107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:37.132107Z digest=sha256:f517fc04d95a2187932ade99a76311294e28dc162bb88b23427fc4d264f625d3

Observation 4aefe993-3f2b-4298-84de-a1e815114d44 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Training Verifiers to Solve Math Word Problems

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:37.207212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:37.207212Z digest=sha256:9e2e34c2d1f96e2669627e9d855b93a403021d135a361864eca6edb3b04ba780

Observation 410dfcc5-d158-4ba6-8771-3b41ddc00ba4 · outbound

This paper cites Learning to reason with llms.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Learning to reason with llms

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:42.275098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:59:37.248131Z digest=sha256:96c9c88433b7d756258558b63c334d6ece7bd97619a6018ca64b14921266d23c

Observation ff8373f4-ace7-4f6b-a2cf-ffe6e4e5a260 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:37.341748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:37.341748Z digest=sha256:2009eee0b8422e968b7d13abf8680d68da36fb1578c663f0c9b9580b35bd9f35

Observation 100c0509-79ce-4388-bf35-80bd2e7076ea · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Chain-of-thought prompting elicits reasoning in large language models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:37.420694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:37.420694Z digest=sha256:6f9a6da3665503f253c83239c5ceb6b58733e8b7c5598556dd1bbbba6a84a318

Observation 4c15750b-4aaf-4190-9fda-8f7d5351c0ba · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:37.501506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:37.501506Z digest=sha256:f3ada9a39c2e8471a3a144e8ed37a5b8c2b23429e4aa9bf0594e17d10b01127c

Observation d494e456-33be-4295-947b-67488fb35dc9 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:37.624601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:37.624601Z digest=sha256:6e92fd2939b9145b999b36b951a9d041da797f7d751e1ded549ab420fe488fde

Observation 7414b3f1-a4ed-4a47-96b6-565e97e5c9cb · outbound

This paper cites ReTool: Reinforcement Learning for Strategic Tool Use in LLMs.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:37.694952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:37.694952Z digest=sha256:52f8c6bdc0d35f891b4d2dc41f0d4f8ca1e5568336c15da926b155abb48eeddb

Observation 301746a9-4f66-44b0-b061-cd7eb6fe2622 · outbound

This paper cites Blink: Multimodal large language models can see but not perceive.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Blink: Multimodal large language models can see but not perceive

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:37.771338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:37.771338Z digest=sha256:f7ba5da0af71fb0296578805b3a07fbe3e19c2490f093aa22303446f44bae13b

Observation c5905332-3d3c-4336-9b61-542c0e738dad · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:37.840955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:37.840955Z digest=sha256:79fefe25df473a5e2b044137587b54863363a1277ecca4d97a420c190f3a329e

Observation 79c0902b-3de5-453b-b578-d6bd921dc99a · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:37.936932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:37.936932Z digest=sha256:3ae9b4117b76246b44f01b435b35b9c52462e410037c678557113f7d1bca181f

Observation ed414bfd-688b-4641-afb4-c2a87b98951c · outbound

This paper cites Decoupled Weight Decay Regularization.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Decoupled Weight Decay Regularization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:38.012150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:38.012150Z digest=sha256:47729ad6943a05052112e9efe3455ff5cfdaeff156737b31673f89133a64b948

Observation 13414287-fbf4-4ebc-be64-4472e5db32cc · outbound

This paper cites UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:38.116421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:38.116421Z digest=sha256:a04f37ea2ba48989a8b043009e01dce1a72883cef92379074c3d02cd15e3ebbd

Observation ad3c211d-e2e9-42da-81f4-cc249b8695fe · outbound

This paper cites DePlot: One-shot visual language reasoning by plot-to-table translation.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection DePlot: One-shot visual language reasoning by plot-to-table translation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:38.207103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:38.207103Z digest=sha256:01f60ed950063a8f6f3a4de3ba4191573119a1ec8b8edcb90b662a1346239393

Observation 82d21851-fc21-4c4b-a93d-a59872686a91 · outbound

This paper cites ChartMoE: Mixture of Diversely Aligned Expert Connector for Chart Understanding.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection ChartMoE: Mixture of Diversely Aligned Expert Connector for Chart Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:38.348879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:38.348879Z digest=sha256:0251d9329220ad2464a5b506abf6a02f374b957a6dac4ed865656bf3c249c59f

Observation 3e2e839f-1f4f-446f-8e05-6288e7069970 · outbound

This paper cites The opencv library.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection The opencv library

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:42.104637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:59:38.471275Z digest=sha256:3c7f617cde240c4a7184baf283a6817ae99ae096d537628fec8d99ce2fea001b

Observation 09391139-cc9f-4177-b139-9a2919c9105d · outbound

This paper cites Context-aware chart element detection.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Context-aware chart element detection

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:41.869761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:59:38.589707Z digest=sha256:ef326ceb8c7357e9a25099c717b37b6194573f17b951587a04f5828def651cca

Observation 9a651a89-6b09-4ff8-b9b2-645b107eff7b · outbound

This paper cites Chartocr: Data extraction from charts images via a deep hybrid framework.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Chartocr: Data extraction from charts images via a deep hybrid framework

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:41.686448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:59:38.709688Z digest=sha256:84a131a92db19e1339b33e5faf5953f0099a6cfe28ff58603c8ef049e6894551

Observation 51f9c9fe-6506-45db-9fe9-db05c6fa83c3 · outbound

This paper cites ChartAssisstant: A Universal Chart Multimodal Language Model via Chart-to-Table Pre-training and Multitask Instruction Tuning.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection ChartAssisstant: A Universal Chart Multimodal Language Model via Chart-to-Table Pre-training and Multitask Instruction Tuning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:38.807547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:38.807547Z digest=sha256:b87e241b01715f46ba18774da2d0d97b62be89c3c787ce98c016f888f99480cd

Observation 250dc581-5b92-4aa4-9bf2-6ca6872fe5f6 · outbound

This paper cites ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:38.929113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:38.929113Z digest=sha256:33c705058367f7994c62360810e474cab79318284987ff06eba9492e0b499497

Observation 1d2389d6-7575-4bad-bc9a-1c287df02f90 · outbound

This paper cites Qwen2.5-VL Technical Report.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Qwen2.5-VL Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.039095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.039095Z digest=sha256:2e49f64a40af8bcbac20e1467cf068c678ddff995083ecb1d047f582f7dd3767

Observation a9d18b27-1551-49af-bdd2-6886ad4bac68 · outbound

This paper cites Diagram formalization enhanced multi-modal geometry problem solver.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Diagram formalization enhanced multi-modal geometry problem solver

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:41.489164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:59:39.114084Z digest=sha256:29c020f068ebc9f15d4850dc786e4f0268172c1a662df36ef630944ff9fe2b9f

Observation aa07cd9b-8274-4328-8380-29586848cb9b · outbound

This paper cites G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.192083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.192083Z digest=sha256:8f456a4bc0cbd5521ecc4908063434b61afd72db2fca62bfc19d7c58d1e792a2

Observation 5c99f6ec-2a88-4c96-82ec-425ab61a8f4b · outbound

This paper cites MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.265256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.265256Z digest=sha256:452c71e1294a1b64a82222cd1d81e0cbb4ca55e4145a956e28e05f2396758c19

Observation 10cac71c-70f2-4a17-9fb6-b6f207e1c2a6 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.319182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.319182Z digest=sha256:100c3f13ceafa548c644bd7819ff5897d017c574f24a76fb29671a57bde8580c

Observation 5ed9a8b4-64cd-4945-92a2-09d1c88075be · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku, 2024.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection The claude 3 model family: Opus, sonnet, haiku, 2024

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.365404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.365404Z digest=sha256:c0616fddd9aae2cbd2076a4a29e2d66cd33dce08a00320dd29a56ad8425be807

Observation dcc3e15b-a8b2-4a86-8408-0a0b0d4808b2 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection PaliGemma: A versatile 3B VLM for transfer

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.424086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.424086Z digest=sha256:f071af61656873b5693632d2aac7e3f087e6d75453376429bfd5b211420320b3

Observation 3dfde684-e0ac-4615-8c21-75abe6408fa0 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.471308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.471308Z digest=sha256:d23bd6bb9d5194d7d415dbbbb2d7980954f12f9bba6c45682fc7fd64fe52e423

Observation 9a6983b9-d7e6-45de-9416-b3084d58b5e0 · outbound

This paper cites Internvl2: Better than the best—expanding performance boundaries of open- source multimodal models with the progressive scaling strategy, July 2024.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Internvl2: Better than the best—expanding performance boundaries of open- source multimodal models with the progressive scaling strategy, July 2024

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:41.333416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:59:39.533698Z digest=sha256:0bc6a8c8963e9b2985d331f1c4a2175518f2c80197b6c6820b35321ab354fbd9

Observation 1f86a0a9-77b2-42c9-9c11-5294c7f4e863 · outbound

This paper cites The Llama 3 Herd of Models.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection The Llama 3 Herd of Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.588544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.588544Z digest=sha256:f3db0b6f0a76e8989e59a6718c704413d42cd540c7bace43e48ec354dc5cf639

Observation 6ea81de3-cb5c-44bb-b0c2-ccfb6e8f4fbf · outbound

This paper cites Cambrian-1: A fully open, vision- centric exploration of multimodal LLMs.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Cambrian-1: A fully open, vision- centric exploration of multimodal LLMs

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:41.165486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:59:39.642125Z digest=sha256:e339d1b5ddd86bddd6eadfe76f98af268f08c601cf41614db8080ecc9bebddcc

Observation 1b80ee2c-5315-4a66-8ea4-acccf288e58d · outbound

This paper cites Pixtral 12B.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Pixtral 12B

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.719031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.719031Z digest=sha256:a7a082254fc8989ee4f2d8df681305c278d30617762aa2bcf002f42b3794e727

Observation f25f61f5-941e-4ba4-9855-539b3bfd3da8 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.764847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.764847Z digest=sha256:f87bf5542eb75c0b3ac9d129bd59b06527dc597044a3b8e0ad9773e1c45b990a

Observation 3006adf5-42af-4746-b61d-4e572fd8820a · outbound

This paper cites Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.820900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.820900Z digest=sha256:4cbaf2b1f0cc82f193964cdd8b459cead35ddd4c231158b95da34c50b9154116

Observation 6cec24d0-f256-4c17-9cd9-8c44588cd93e · outbound

This paper cites RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.871961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.871961Z digest=sha256:aec208f1a584dd193f3805b20aa8994bb088c6daa3e983f280d2305df74e1aad

Observation a2e7b4b2-2738-4452-a79d-b384cda0cb89 · outbound

This paper cites Improved baselines with visual instruction 15 tuning.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Improved baselines with visual instruction 15 tuning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:40.966064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:59:39.941166Z digest=sha256:ca40bc3c870c8223617ae4a2b146a55895e5b4fd25ae73cb764afbd19bcb3dd4

Observation e1a78340-085f-4938-a3ba-adeab8c7d730 · outbound

This paper cites xGen-MM (BLIP-3): A family of open large multimodal models.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection xGen-MM (BLIP-3): A family of open large multimodal models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:40.006031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:40.006031Z digest=sha256:3bba1e7e2bab7eb05d5298b9da641b371834fdb81632468cf1c9faf60f075a10

Observation 5361dfab-5fa2-4314-a908-cde932edbe19 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection LLaVA-OneVision: Easy Visual Task Transfer

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:40.054356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:40.054356Z digest=sha256:4b7b8e80120f9f3a3eae5ab7942081010b6affcf04ae4181eaa09ad7a1c2261d

Observation 248a11fb-53c1-4729-a873-46e668f615f3 · outbound

This paper cites Vision language models are blind.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Vision language models are blind

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:40.721505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:59:40.120459Z digest=sha256:06eb3d6405a314779f2d31c90b3b0917f2321b22936ecb968770e364678ddd8e

Pith citing papers

Observation aa875bb9-c0b3-41b7-8c59-6dbbba4db11f · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

Reference 252

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.261893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:4cbda5970c3108440d81520265df1b4528995aff535e148616eab9eac2b9b9d0

Observation cb67b269-a8a5-41a4-974f-2f01fb059d25 · inbound

Latent Visual Reasoning cites this paper.

Latent Visual Reasoning VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T18:41:30.400550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T18:41:30.307521Z digest=sha256:31b64e2d75b4af913a510e6b8fdb39e079ae47f7552f4925d5c256ee7fb0d6e7

Observation 9b5f1228-cac0-48fb-9481-8f569a5422a9 · inbound

OpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial Agents cites this paper.

OpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial Agents VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T22:10:49.569781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:10:49.569781Z digest=sha256:144326548717070206e361ac5ae833a5da600245ebeb65d34438a412f275f5ab

Observation 36ae82ce-49a7-4938-abca-ade047ceada9 · inbound

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model cites this paper.

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:17:29.187051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T07:14:48.918959Z digest=sha256:1ca728e5d5de8fd50e45547e34fafceabb74131fc86ce79a50a532d3954852e3

Observation e149e1ea-637f-4b6b-b131-7f05cef92eae · inbound

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model cites this paper.

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:48:00.623176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T21:47:50.595481Z digest=sha256:cfcc30d0723b0490e20a08893bcebc0681cdb030dba0ebba636acf2701198228

Observation 11fce3b3-03fa-4695-a0bb-17051f699a0a · inbound

REVERSE: Reinforcing Evidence Verification and Search for Agentic Image geo-localization cites this paper.

REVERSE: Reinforcing Evidence Verification and Search for Agentic Image geo-localization VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:50.702685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T18:19:10.346738Z digest=sha256:2d78db4d4833832b91647772441c34628285701f10b8cc44b2922d1eb3fbaeba

Observation de041ff8-84d2-441e-9567-5a8772b38d6c · inbound

DeepLatent: Think with Images via Parallel Latent Visual Reasoning cites this paper.

DeepLatent: Think with Images via Parallel Latent Visual Reasoning VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T20:22:37.695015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T18:44:39.545911Z digest=sha256:a76cea4d503105ee34b612586144cd4f6a4608b8cf6b8e0f558553687afee36b

Observation f01cae5e-0002-47e1-b14a-147576225dbe · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

Reference 214

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.149238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:b31405b63db636d63e3198047e572a739cc04fbc52b9803d0ac08ef9cabb7b35

Observation 12c40016-1136-4111-a9a1-a24b220f201d · inbound

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing cites this paper.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:21.527737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:21.527737Z digest=sha256:bfb317d2e42d79509edf54be993faeac66a65eef04244611f5b187d199e5202c

Observation 2ecdc0a2-7efb-4598-a93b-4def83e45495 · inbound

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use cites this paper.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.053598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.053598Z digest=sha256:b7e7140d1a5f19597f492ca5be918cf6625c97eac571add22ee7738d4af4dd91