Pith. sign in

Paper Citation Record · LEDGER

LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2311.05437.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.05437 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T19:25:33.751527Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T10:35:41.926606Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b96c18cb-aa7d-4802-81af-1c9b31884c8a · inbound

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices cites this paper.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:35:38.185520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:823793e2036bf35b00c297daf5aae72b10f6968b946f886827d13d98581edd66

Observation 4401e72f-91e8-4b34-a444-fad6f2db5712 · inbound

A Survey on Hallucination in Large Vision-Language Models cites this paper.

A Survey on Hallucination in Large Vision-Language Models LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:10:10.338212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T22:10:10.186950Z digest=sha256:7e32a84f02045a1dde37139e6073c89ae6e2559e71feae559b9adc2c39b1d961

Observation ffb75772-aa49-4876-8637-f0066da0672e · inbound

Large Language Models: A Survey cites this paper.

Large Language Models: A Survey LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 217

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:22:55.992348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T15:22:54.023279Z digest=sha256:f1d593909e9fdf7dd417c1480261e7f3f15e14a1ed1aa58dc235619a58397a8b

Observation 1ffe80b4-f047-46b4-9bbf-997312599f30 · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 77

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.536840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:f64f7869edb40a8d45f96026c9a7ff19e8456190a228118c51f033cd848d54b2

Observation 66ac61e0-3ecd-462d-97f2-86c17d6da27b · inbound

Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model cites this paper.

Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:57:26.964174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T05:57:26.887069Z digest=sha256:7635b49285d4f6d696a519bb41cebba2c0cd97758e0ff8a6f3af6f01083da1c9

Observation fd30c3a7-9fb9-4fe2-b9b9-cc64a6e87e46 · inbound

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction cites this paper.

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 136

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:15:47.062053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T19:15:21.695801Z digest=sha256:00cb54f39eae29d729be2764d32ed0afca7b482aefc6290029781bbdc23b91c9

Observation c09a4593-00e2-4dac-b352-4eecbd225704 · inbound

Image Embedding Sampling Method for Diverse Captioning cites this paper.

Image Embedding Sampling Method for Diverse Captioning LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T19:25:33.751527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:25:33.751527Z digest=sha256:88e219e6939e233dce7f16e1ef7df8e3f98270f6f8b3bd20f947e8d5dedf019e

Observation 42de5efa-0cb5-431a-83ba-d78fd3927afc · inbound

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models cites this paper.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:21:39.732552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:78f22b72501c4ddb791ee616ba033f917352f985986f1f940fb530c46a53fd60

Observation 5d56ce10-92a8-4789-9fc0-94d03b860da1 · inbound

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought cites this paper.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:23.566141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:23.566141Z digest=sha256:774213324786208773591f1e0069b71724a4ef311b762debb251ef0cd2bf7640

Observation b0e4dc90-d51d-4690-9e97-909398100931 · inbound

ZeroVO: Visual Odometry with Minimal Assumptions cites this paper.

ZeroVO: Visual Odometry with Minimal Assumptions LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:59.109392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:59.109392Z digest=sha256:8cde877c9b5f379c359aeac3571c4bc3b386f1483762e7820a345bd49b91fbc7

Observation 4720f529-b09c-480a-a074-accdfeef8031 · inbound

A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications cites this paper.

A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 152

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:20.777371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:20.777371Z digest=sha256:d7f62d485cd165580f74f3df8bd347022406b35d87d3fece9f7e6a01f020a006

Observation b6798a7d-cf83-4917-bc88-4894cf994b78 · inbound

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance cites this paper.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.756176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.756176Z digest=sha256:bb8b98d4e1714a17cc2139dff819b36252ce39b5dcde29f8822156bd4e9b2847

Observation 13d2176b-db56-4ac4-856e-40753e4f798d · inbound

Augmented Vision-Language Models: A Systematic Review cites this paper.

Augmented Vision-Language Models: A Systematic Review LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:37.251420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:37.251420Z digest=sha256:2c4a5504139d93595fc85e2c7bcb8891d2be7b48c540c230fc6ee46e76d7453d

Observation 3665d3b9-926c-4d2e-9dea-54a19b7b8306 · inbound

Reinforced Visual Perception with Tools cites this paper.

Reinforced Visual Perception with Tools LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:04.989084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:27:04.989084Z digest=sha256:5b17611dd0027dfd2d457b8f12cea7bb9ac3cb1be7e3ade249ff7443887e77e9

Observation ae82f0c8-2299-4db5-8847-af159577028b · inbound

Mitigating Coordinate Prediction Bias from Positional Encoding Failures cites this paper.

Mitigating Coordinate Prediction Bias from Positional Encoding Failures LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T05:12:23.519700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:11:30.734633Z digest=sha256:08185aab5a6f6671c518ed61233c539234fbb589f7d668507cbc0fb52fd9f641

Observation df93be91-3232-4a8c-b30b-18c5ae5848ed · inbound

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning cites this paper.

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T20:57:30.428570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:57:30.428570Z digest=sha256:df0e576d38af263f72a44f895443e40d7db12a6b3d14345251c0c853c2a4a5af

Observation 8089de90-037a-42c6-8409-83dbeba083cf · inbound

E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes cites this paper.

E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T09:40:35.631188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:40:35.631188Z digest=sha256:f0960e6e441c154531510ab869bd6bac4402cdddc3bf443b0bae673641ff856d

Observation a3d56488-2b30-43f2-b2fe-443ab946a08e · inbound

Less Detail, Better Answers: Degradation-Driven Prompting for VQA cites this paper.

Less Detail, Better Answers: Degradation-Driven Prompting for VQA LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:05:47.973219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T20:17:01.867903Z digest=sha256:4a84aee17e75bba97ebfa78f7bfe305c2bae1e89cb4254c13161fc63b4747d56

Observation ac93bcb1-9b2b-44dc-acb9-751745bcae77 · inbound

Waking Up Blind: Cold-Start Optimization of Supervision-Free Agentic Trajectories for Grounded Visual Perception cites this paper.

Waking Up Blind: Cold-Start Optimization of Supervision-Free Agentic Trajectories for Grounded Visual Perception LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:01.721126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T05:33:33.589721Z digest=sha256:dc72ca0841b022a72dd13b1fedcc46ae63fc8c27f838b18236c212afd0bd00f2

Observation b24b2b45-f5a1-4dce-8f76-ad08bd6af5a0 · inbound

Second-Order Multi-Level Variance Correction for Modality Competition in Multimodal Models cites this paper.

Second-Order Multi-Level Variance Correction for Modality Competition in Multimodal Models LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:03:39.571289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T19:02:02.266052Z digest=sha256:818a29b88d098506aff61b80aec5c391d624a7c4222a772656e4235d7f6ca89e

Observation 45a2a72a-1e30-471f-b9a5-d41b606ec131 · inbound

IndusAgent: Reinforcing Open-Vocabulary Industrial Anomaly Detection with Agentic Tools cites this paper.

IndusAgent: Reinforcing Open-Vocabulary Industrial Anomaly Detection with Agentic Tools LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:33:58.501835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T05:33:30.670201Z digest=sha256:03b958fdbadff4e8c80ced59a5dfb0b64823f8e6a87ff0a453645a18b6d16040

Observation d590c960-c4d2-457b-942d-8adab7f652f6 · inbound

ReGRPO: Reflection-Augmented Policy Optimization for Tool-Using Agents cites this paper.

ReGRPO: Reflection-Augmented Policy Optimization for Tool-Using Agents LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:35:41.928986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T05:23:21.162004Z digest=sha256:c6d10b9c245f206c18950b1b9fe47f7d97cfbd06d9c693db5f3532331e71e277

Observation 61b9a200-a680-4ee8-b4db-2cd88dac3504 · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 216

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:11048ef522a08ec89f3647b3a3898979cb49c48c88537b5be9f2f8bf3641fc58