Pith. sign in

Paper Citation Record · LEDGER

To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 53 inbound Pith citation observations for arXiv:2311.07574.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.07574 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 53 of 53 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:01:39.976358Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T17:07:25.768424Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fb08796f-a04b-43cd-99dd-3da7f1d1f638 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:42.628229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:ccc3a3a632fa67b05ed35771149200f786e62a3f293fedfbd137bbac67483027

Observation 246e4615-ae3c-4a41-b907-cb9847acbc77 · inbound

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices cites this paper.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 121

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:35:38.152397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:f458cda49415b56b9cfa54d9c5725dec21ba6e8f50d67149cc5bb5267a76e334

Observation dc43a9a1-d6cd-4c80-bb6b-37ce19e8e0e5 · inbound

InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model cites this paper.

InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:30:27.737221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T05:30:27.667126Z digest=sha256:3d3a01c8d72642ef4959859eddab54942864b060690c9d53dac689d7bcd9a0c2

Observation 3ced92d1-13d4-4817-80a1-a4ac61a6905e · inbound

MobileVLM V2: Faster and Stronger Baseline for Vision Language Model cites this paper.

MobileVLM V2: Faster and Stronger Baseline for Vision Language Model To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:27:52.101123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T15:27:51.839171Z digest=sha256:3e7fe82bde8ed043c7c517cc91de7e5d16ebc61a08088752fddea8831467e473

Observation 846e3349-49b1-4036-be1c-66763a84d450 · inbound

DeepSeek-VL: Towards Real-World Vision-Language Understanding cites this paper.

DeepSeek-VL: Towards Real-World Vision-Language Understanding To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:58:54.717931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T17:58:54.177359Z digest=sha256:34e4ccced1ed129596df5dee451e50014c814ca92cafc46af1fd587f465067e4

Observation 41079ee0-4245-4c6d-95d0-d2ae23429ab5 · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 113

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.300585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:506229059222fd3cb7b2f414fb6de62459aa17e63d60e6f478dca1be25b03552

Observation 91f73498-8512-43d5-b6a2-cc4d57037c79 · inbound

Are We on the Right Way for Evaluating Large Vision-Language Models? cites this paper.

Are We on the Right Way for Evaluating Large Vision-Language Models? To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:41:44.390189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T19:41:44.263663Z digest=sha256:68461ebf937e6e336fe43cb6fd705cb26b65c2f8a2c198d4d35a35f2c9903b67

Observation e4c6ebed-d918-48a8-a469-81cab83b8464 · inbound

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation cites this paper.

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:48:36.219428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T22:48:36.010306Z digest=sha256:127a98b2132c30b2f2b7a3b7b8c6251db5bd2d32b501bfd92acd1a2c976abb7b

Observation bdff8938-636e-4351-b0f2-108942cf6752 · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:59.228957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:2b05bb63931a3f22a8047bc400ada3b3e4ab63c00fe11bd8a71f680f1f873aa2

Observation 9be80528-7c79-4df7-a9ac-42d7cf12cd4e · inbound

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs cites this paper.

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 132

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:05:03.766949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T00:05:03.547664Z digest=sha256:f58d16a5d194bb5a16256cbb0b7bc0a98dd67eb1cb379d95cb6a578b48bcd393

Observation ec5e6333-e0bf-4e44-8a52-61b818a22b4c · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 150

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.754970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:f31a9006dc8985ddc58753eb19970e1be0b1dbc8ceeab193cd4a46dc0174bc61

Observation 3b444aa8-0c49-43fd-906d-460b82451ffc · inbound

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models cites this paper.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 110

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.693057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:010e99c06903a3e8d07aeed1d93a9ff2076246059ff3904ed100df47cf09bf79

Observation cba2aec2-877a-45a6-8f82-dbfb08cdcf5b · inbound

MIRe: Enhancing Multimodal Queries Representation via Fusion-Free Modality Interaction for Multimodal Retrieval cites this paper.

MIRe: Enhancing Multimodal Queries Representation via Fusion-Free Modality Interaction for Multimodal Retrieval To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T21:48:54.052069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:48:54.052069Z digest=sha256:3010ef8c2cd5b9cb750aef8744ff33334395df76ff2d40621a09bac8ba06e8a3

Observation 37c6eede-1eec-4593-8dac-5d4cdcd87a11 · inbound

BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices cites this paper.

BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 124

Resolution
unresolved
no resolver link, observed 2026-08-12T19:33:00.999591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:33:00.999591Z digest=sha256:4ae95fbd1bfe3493de54d0c473adfb36321b03fc2b1000468fb7d50d52c4c4c9

Observation 2353e9f8-d74d-4bf5-8e8e-4b59da94f0f5 · inbound

Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification cites this paper.

Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:00.346640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:00.346640Z digest=sha256:1be30951543ecf787bb59a30102f6dec84ae752399a6c563c30844c7b95d8131

Observation 6fb38b55-2166-4303-89fe-b62a5283848a · inbound

Agri-LLaVA: Knowledge-Infused Large Multimodal Assistant on Agricultural Pests and Diseases cites this paper.

Agri-LLaVA: Knowledge-Infused Large Multimodal Assistant on Agricultural Pests and Diseases To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T23:50:37.089010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:50:37.089010Z digest=sha256:a05eec58251aeeef767006ae47f6702663b9b02c7b9a3040cb62a253d747a79a

Observation f63519f4-1867-4e76-b350-9d23eca62241 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 245

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.101801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:e6ef25fe3506b3313bc00ea57f43b7752004fa3da0f1f60985acbe5e5c67aca0

Observation c14f66df-556f-4525-b716-5a3e9cc4b655 · inbound

POINTS1.5: Building a Vision-Language Model towards Real World Applications cites this paper.

POINTS1.5: Building a Vision-Language Model towards Real World Applications To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:36.335982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:36.335982Z digest=sha256:18c5c1037b1fc988a98749fcaab8d845d585ae1e7b409a5afd0fbeb1969e6816

Observation dbd852c3-0825-4781-914e-37e6b6a16722 · inbound

V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding cites this paper.

V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 120

Resolution
unresolved
no resolver link, observed 2026-08-11T16:58:03.454779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:58:03.454779Z digest=sha256:7f98e7cae21cb8a491864dfc514b9559d5cbf1fefe832fb0939f7701979c2283

Observation dbefc10c-0c29-47b3-a559-8f547d0cbcfa · inbound

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning cites this paper.

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 87

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T07:51:13.416130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-17T07:51:12.953777Z digest=sha256:6373135fce46d023e2040b3e33a1af9be2f563e1d2ae050b6abeb0fa2175b92c

Observation 35b58c34-35a2-4dd1-8441-376a04fccdd3 · inbound

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction cites this paper.

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:08:19.689938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T21:08:19.570050Z digest=sha256:b928f6c07d130c012df2e8e898550e3e5ee4816d5d559efd9c8a4372d97dc796

Observation 4151c0ed-3e43-4811-8928-2af738694700 · inbound

Visual Large Language Models for Generalized and Specialized Applications cites this paper.

Visual Large Language Models for Generalized and Specialized Applications To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 258

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.902114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.902114Z digest=sha256:466cd9af765ef8bcf5105fd7794d5fe4985d73b6496baeff1fcb8489449269ad

Observation a81bae1f-746a-4107-90b0-ebf99c150d8f · inbound

Pix2Cap-COCO: Advancing Visual Comprehension via Pixel-Level Captioning cites this paper.

Pix2Cap-COCO: Advancing Visual Comprehension via Pixel-Level Captioning To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:55.535190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:55.535190Z digest=sha256:7c0c422b420fd7e642951f7e41c500fbc6f14724a99f512edbb2fa669e3d5c4f

Observation 70518c8b-093a-44e4-9bac-08f2e0062082 · inbound

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models cites this paper.

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 126

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:34.507504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:34.507504Z digest=sha256:30a8331c96386460ce5084ac635b81225b968972ef7c6a6d3e0371517a442fcb

Observation e5d65e6f-ce52-4940-894f-d073327309c7 · inbound

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark cites this paper.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T11:26:29.210443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:26:29.210443Z digest=sha256:8ec48756c1620fc4bba97594174b4efd675cd5a4dc19a652014d66ea5a14acb9

Observation c6280041-f4b1-4ef6-ae4a-877331f0c2b9 · inbound

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models cites this paper.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.910395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.910395Z digest=sha256:b1d987eeb5fb6742796424bab234ac074951126169db0f74474c5fb636d0679d

Observation d87f0295-6e5f-4975-a7dc-91dd263cd06e · inbound

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding cites this paper.

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:52:01.931880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-22T19:49:00.961388Z digest=sha256:fb89f834433a0be813d5b90253db62d7e03d1871c165193e123046070039aa9f

Observation 8b9c1474-9b5a-4f77-9bfe-1f1c04b2e7e5 · inbound

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward cites this paper.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.976358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.976358Z digest=sha256:f874ec4f8413c1ac85ba07e5b98d1fcfc0737a16776c3b6f92a396908093e8e7

Observation 04a819bd-1fc3-4b4d-8bc9-74b786916413 · inbound

Critique Before Thinking: Mitigating Hallucination through Rationale-Augmented Instruction Tuning cites this paper.

Critique Before Thinking: Mitigating Hallucination through Rationale-Augmented Instruction Tuning To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T22:27:46.807706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:27:46.807706Z digest=sha256:9d186b37e0a2c2d84084bdfe0887c3787385916b1309dc060c8c04310be8e7ad

Observation e9fbeeaa-147d-48ab-8af4-d1dc59b836ec · inbound

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning cites this paper.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.670495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.670495Z digest=sha256:41b1eba6eebb1b4fa167ec4f53dcca9a40abdf1095733af831868be86be2458e

Observation 380b9a90-ab99-4a3d-8c4d-52846bee1ac0 · inbound

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion cites this paper.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:52.789583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:52.789583Z digest=sha256:a1f78bd4c789782a7b908787a3e332801d56cfa1df924b72e4a3a2c06058368f

Observation 24283774-8d39-46e8-b0b9-60d9db83d88c · inbound

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities cites this paper.

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 136

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:07.447847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:07.447847Z digest=sha256:f02ce4b4a01d3ecb5b8b16cd823c7de0b0082129b2f526f16834a86ded280755

Observation 00b24c17-26cc-4e54-9912-f809d9449f77 · inbound

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models cites this paper.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:28.020503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:28.020503Z digest=sha256:3eb2ed64709516b9a4c5ae7736c723f5558fd008cc0a9336078775968704574e

Observation 1df6994e-e983-4736-9c96-d1dc5913f244 · inbound

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought cites this paper.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:26.521528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:26.521528Z digest=sha256:f1b52e01e12b094c8a00d5c4f598cff17d73e13bd5e4afdcf59f323dfc0544fe

Observation aa307276-50eb-488a-b3c3-61cad5c90d37 · inbound

EvoMoE: Expert Evolution in Mixture of Experts for Multimodal Large Language Models cites this paper.

EvoMoE: Expert Evolution in Mixture of Experts for Multimodal Large Language Models To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:33.198713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:33.198713Z digest=sha256:5e9d2bbb329db1f2fd7aaa2176c69301f8b9a268abbfe7abaac713e3eeddd3c0

Observation 773e5efe-48b2-4f08-824e-8193d0d72480 · inbound

Fast or Slow? Integrating Fast Intuition and Deliberate Thinking for Enhancing Visual Question Answering cites this paper.

Fast or Slow? Integrating Fast Intuition and Deliberate Thinking for Enhancing Visual Question Answering To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:10.642601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:10.642601Z digest=sha256:74bf020f4ae5dfb00937928831d9e108307075458ca0b0daee9c1c5a2f89c619

Observation 464049c1-2cff-40ee-a310-a8cab0779050 · inbound

SMAR: Soft Modality-Aware Routing Strategy for MoE-based Multimodal Large Language Models Preserving Language Capabilities cites this paper.

SMAR: Soft Modality-Aware Routing Strategy for MoE-based Multimodal Large Language Models Preserving Language Capabilities To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:18.730814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:18.730814Z digest=sha256:19dba8d4c897d3688b8c5eac74d6c6f09c9019aeb3d13fc189b6a39965683d7c

Observation 7d658a2d-493f-4c3d-8b4e-c7ef1a5e2a14 · inbound

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation cites this paper.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:08.755303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:08.755303Z digest=sha256:a2930e2cc7479f68db79ddc051e7392c97355849c62194583970e11b561cd008

Observation 0e37eb4b-4166-45b0-b6df-e75df5fd21a9 · inbound

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation cites this paper.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:53.783559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:53.783559Z digest=sha256:0c86e659d12f797a70459680fe90ce7f24c96605f920ea64e4c3849bed25709d

Observation b970ecbc-9311-4a1b-bd3c-f83feadd74b4 · inbound

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models cites this paper.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:05.174737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:05.174737Z digest=sha256:c2f4b38317ec1d314d4260ae39be50663b3fd69821cc8eb604ed2a2ef55d3186

Observation 95153f93-0b08-4be7-a40f-5c9f3ef25d0c · inbound

See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs cites this paper.

See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T12:12:23.689342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:12:23.689342Z digest=sha256:a233902556db5d5375340bf3ee3bb3eb22cf6ad24a355361a1fd67938a8a2468

Observation 28cd8b76-9dd8-4a98-94ff-bdf1db514fcc · inbound

Metadata Management for AI-Augmented Data Workflows cites this paper.

Metadata Management for AI-Augmented Data Workflows To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:18.637019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:18.637019Z digest=sha256:b9a131e37aa5aecdc632eeb6af54c50639032294815d0cf2172b751a3382cb6d

Observation a8d5f0d5-a4b0-4106-a68a-5824764c9cde · inbound

UItron: Foundational GUI Agent with Advanced Perception and Planning cites this paper.

UItron: Foundational GUI Agent with Advanced Perception and Planning To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T14:03:38.489435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:03:38.489435Z digest=sha256:6df910a33765f83ba6a7fd88018c263f0e1be04c63130757627983f19ae86db1

Observation 97f328e9-a648-4b06-ba86-243fe4c8799c · inbound

Improving Large Vision and Language Models by Learning from a Panel of Peers cites this paper.

Improving Large Vision and Language Models by Learning from a Panel of Peers To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.502470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.502470Z digest=sha256:002116e899bc83effb2be77976c64d34156cead626d158ab49dd4bda2bf69bae

Observation 260f681b-8edd-4671-8ead-82b65f1cc844 · inbound

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs cites this paper.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:57.669229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:57.669229Z digest=sha256:2c9f9f207adc92a8446d66b618261cc9f959b54657ddd502927be81ebd7e8466

Observation 5ce4d7eb-816a-44b9-acb1-7698013bf443 · inbound

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding cites this paper.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:50:56.484333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:f94fc87b534d7c539621e122b144b21acbff9d8e2ea5caadd9c73c7e52a28e26

Observation 0174631a-5b6f-4a47-8001-62f592d94ac9 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 264

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.498843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:16a99473bc62ae1331a4b3a01d5f04f2fb00fbfaf2b6313bff019256a480e433

Observation 2de9cc4a-c6b4-43ab-92a2-701dbd88eb47 · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 295

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:45:47.669670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T01:16:16.834861Z digest=sha256:748fabc152409afad5ffc8aa4807857fd6c450ee8c339788fcde65bab275cdd4

Observation 3bc2e407-a963-4718-b5a0-7862c86e6e55 · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 295

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.034330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:318a1a16469692e0478dbe2c0b25e5b13102e09903077f0c5ca36c21deb632aa

Observation 31f9f1ea-8638-4cf0-88e6-79b6c128692d · inbound

Infinity-Parser2 Technical Report cites this paper.

Infinity-Parser2 Technical Report To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-10T17:07:25.769682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-10T17:02:28.089092Z digest=sha256:f448557537c84c3d76019ba85ca67707ce89eae53b2fcabf6b88bcdb9ec43cca

Observation 011eca1f-ffd9-45a6-a8a8-4b041fdf3c3c · inbound

Infinity-Parser2 Technical Report cites this paper.

Infinity-Parser2 Technical Report To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T08:03:52.315630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:03:52.315630Z digest=sha256:925e18b2a110041a2347bfdfae00b59aae2a580b26fadafb581dbff336ab3d21

Observation cb964605-00b7-4897-b627-d49e8787f565 · inbound

Twins: Learn to Predict Unified Representations with Focal Loss cites this paper.

Twins: Learn to Predict Unified Representations with Focal Loss To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 281

Resolution
unresolved
no resolver link, observed 2026-08-01T04:30:11.620048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:30:11.620048Z digest=sha256:ab9bc6a06b99539f30d3d57d7f8faddcfab5e0f3b419b317f2b8c21414531207

Observation 39244b89-6a3e-4823-8faa-1fa97e178d37 · inbound

DistMoE: Private-data Rehearsal-free Routing in Mixture-of-Experts for Distributed Instruction Tuning cites this paper.

DistMoE: Private-data Rehearsal-free Routing in Mixture-of-Experts for Distributed Instruction Tuning To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T14:28:03.481999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:28:03.481999Z digest=sha256:697d7e330805d182baa5fa110b57acadfc0ef18f691b2eb538fcdbaa26cd63c5