Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:59:40.120459Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 10 inbound Pith citation observations for arXiv:2505.20289.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:59:40.120459Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T00:30:03.053598Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T15:09:55.147388Z
59 of 59 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bdd3af60-36f6-4ff6-9905-7332016ec311 · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d08f3d56-2dfc-49e1-80c0-075079462928 · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Language models are few-shot learners
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ab323fc2-493c-4d18-b486-9882846af7be · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection LLaMA: Open and Efficient Foundation Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdba0ffa-9c85-4a20-a6ba-90a550228e0e · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection GPT-4 Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 688941ca-0a04-47d9-b69b-2e2700a5f615 · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Visual instruction tuning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afde5757-ba96-4b33-b939-a1c33aab25b7 · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Qwen2.5 Technical Report
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e01f5360-5b6f-4a5b-b824-68d7ebcb87cc · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection GPT-4o System Card
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f8024ca-e64d-44f5-8ce5-d7d481dfac4d · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8de9d6e9-170e-4728-98a1-1e4079043ccc · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a73f13ec-d452-470f-9684-e0d8ba91a76f · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Pal: Program-aided language models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e26eaad-20dc-4f94-a5a8-e12345f4ce40 · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Gorilla: Large language model connected with massive apis
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2fa008a5-8011-47d7-abef-8f08ead072a1 · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8875f002-9e09-4fbc-8494-38984220fb19 · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Visual programming: Compositional visual reasoning without training
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9de7870e-40fa-46a8-8b5c-b1aa33c00c86 · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Vipergpt: Visual inference via python execution for reasoning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bff801ff-b139-4cb6-929f-eeb09cfb4686 · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Toolformer: Language models can teach themselves to use tools
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6da95cc0-cc6c-4f67-96b8-8da502b1191f · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Llava-plus: Learning to use tools for creating multimodal agents
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98792aca-f97f-4bc4-baaa-cb14ae54f20e · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4dd80ee7-9102-4d5c-b7ef-a280e31fca07 · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Reinforcement learning: An introduction, volume 1
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72fc4b0c-cf1c-40de-a9e1-66826631bbe0 · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14f40af1-03f3-4515-ba06-7b0a00e206f9 · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28f6d4d3-57c7-46a2-bf52-6987a05e02fe · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Internet-Augmented Dialogue Generation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4aefe993-3f2b-4298-84de-a1e815114d44 · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Training Verifiers to Solve Math Word Problems
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 410dfcc5-d158-4ba6-8771-3b41ddc00ba4 · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Learning to reason with llms
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ff8373f4-ace7-4f6b-a2cf-ffe6e4e5a260 · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 100c0509-79ce-4388-bf35-80bd2e7076ea · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Chain-of-thought prompting elicits reasoning in large language models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c15750b-4aaf-4190-9fda-8f7d5351c0ba · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Visual-RFT: Visual Reinforcement Fine-Tuning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d494e456-33be-4295-947b-67488fb35dc9 · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7414b3f1-a4ed-4a47-96b6-565e97e5c9cb · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 301746a9-4f66-44b0-b061-cd7eb6fe2622 · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Blink: Multimodal large language models can see but not perceive
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5905332-3d3c-4336-9b61-542c0e738dad · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79c0902b-3de5-453b-b578-d6bd921dc99a · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed414bfd-688b-4641-afb4-c2a87b98951c · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Decoupled Weight Decay Regularization
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13414287-fbf4-4ebc-be64-4472e5db32cc · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad3c211d-e2e9-42da-81f4-cc249b8695fe · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection DePlot: One-shot visual language reasoning by plot-to-table translation
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82d21851-fc21-4c4b-a93d-a59872686a91 · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection ChartMoE: Mixture of Diversely Aligned Expert Connector for Chart Understanding
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e2e839f-1f4f-446f-8e05-6288e7069970 · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection The opencv library
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 09391139-cc9f-4177-b139-9a2919c9105d · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Context-aware chart element detection
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9a651a89-6b09-4ff8-b9b2-645b107eff7b · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Chartocr: Data extraction from charts images via a deep hybrid framework
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 51f9c9fe-6506-45db-9fe9-db05c6fa83c3 · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection ChartAssisstant: A Universal Chart Multimodal Language Model via Chart-to-Table Pre-training and Multitask Instruction Tuning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 250dc581-5b92-4aa4-9bf2-6ca6872fe5f6 · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d2389d6-7575-4bad-bc9a-1c287df02f90 · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Qwen2.5-VL Technical Report
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9d18b27-1551-49af-bdd2-6886ad4bac68 · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Diagram formalization enhanced multi-modal geometry problem solver
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation aa07cd9b-8274-4328-8380-29586848cb9b · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c99f6ec-2a88-4c96-82ec-425ab61a8f4b · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10cac71c-70f2-4a17-9fb6-b6f207e1c2a6 · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ed9a8b4-64cd-4945-92a2-09d1c88075be · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection The claude 3 model family: Opus, sonnet, haiku, 2024
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcc3e15b-a8b2-4a86-8408-0a0b0d4808b2 · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection PaliGemma: A versatile 3B VLM for transfer
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dfde684-e0ac-4615-8c21-75abe6408fa0 · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a6983b9-d7e6-45de-9416-b3084d58b5e0 · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Internvl2: Better than the best—expanding performance boundaries of open- source multimodal models with the progressive scaling strategy, July 2024
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1f86a0a9-77b2-42c9-9c11-5294c7f4e863 · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection The Llama 3 Herd of Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ea81de3-cb5c-44bb-b0c2-ccfb6e8f4fbf · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Cambrian-1: A fully open, vision- centric exploration of multimodal LLMs
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1b80ee2c-5315-4a66-8ea4-acccf288e58d · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Pixtral 12B
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f25f61f5-941e-4ba4-9855-539b3bfd3da8 · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3006adf5-42af-4746-b61d-4e572fd8820a · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cec24d0-f256-4c17-9cd9-8c44588cd93e · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2e7b4b2-2738-4452-a79d-b384cda0cb89 · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Improved baselines with visual instruction 15 tuning
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e1a78340-085f-4938-a3ba-adeab8c7d730 · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection xGen-MM (BLIP-3): A family of open large multimodal models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5361dfab-5fa2-4314-a908-cde932edbe19 · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection LLaVA-OneVision: Easy Visual Task Transfer
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 248a11fb-53c1-4729-a873-46e668f615f3 · outbound
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Vision language models are blind
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation aa875bb9-c0b3-41b7-8c59-6dbbba4db11f · inbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection
Reference 252
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cb67b269-a8a5-41a4-974f-2f01fb059d25 · inbound
Latent Visual Reasoning VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9b5f1228-cac0-48fb-9481-8f569a5422a9 · inbound
OpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial Agents VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36ae82ce-49a7-4938-abca-ade047ceada9 · inbound
Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e149e1ea-637f-4b6b-b131-7f05cef92eae · inbound
Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 11fce3b3-03fa-4695-a0bb-17051f699a0a · inbound
REVERSE: Reinforcing Evidence Verification and Search for Agentic Image geo-localization VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation de041ff8-84d2-441e-9567-5a8772b38d6c · inbound
DeepLatent: Think with Images via Parallel Latent Visual Reasoning VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f01cae5e-0002-47e1-b14a-147576225dbe · inbound
From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection
Reference 214
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 12c40016-1136-4111-a9a1-a24b220f201d · inbound
Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ecdc0a2-7efb-4598-a93b-4def83e45495 · inbound
In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.