Pith. sign in

Paper Citation Record · LEDGER

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion

As of 7 August 2026, this Paper Citation Record lists 90 of 90 outbound references and 0 inbound Pith citation observations for arXiv:2505.18115.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18115 v1

Coverage vector

measured 90 of 90 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:37:54.579360Z

measured 90 of 90 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

90 of 90 outbound references displayed

  • verified exact0
  • verified fuzzy36
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 46b57075-b50b-4e2b-8eda-d6c7d4a00491 · outbound

This paper cites anthropic.com/news/claude-3-family, 2024.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion anthropic.com/news/claude-3-family, 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:46.891987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:46.891987Z digest=sha256:7f0b8ab77008f8a3e19cefb26a70c37da38c4c9aea539df1c1420ef74d4edc1e

Observation 782c97b6-a0ee-48a1-98b5-ddaddf9a3ffe · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:46.984855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:46.984855Z digest=sha256:7ceb3a6f54dc65067434228452f1e471ea9bb5e130338a420ee1560710c5452a

Observation 104d0814-623d-4e48-b819-5e760535fa89 · outbound

This paper cites Easyocr: Ready-to-use ocr with 80+ supported languages, 2020.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Easyocr: Ready-to-use ocr with 80+ supported languages, 2020

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.057781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.057781Z digest=sha256:b034df2376f71ca50382f34725cd751870c302a49cbfcd27aeb69404b4fa1b95

Observation 796be892-b44a-4377-ac8d-ab004995c347 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Flamingo: a Visual Language Model for Few-Shot Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.145455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.145455Z digest=sha256:d67df726ca5b7c16cbf455eaf4581a1c5d0f273ce58a9d5b1d36c3e22b38d369

Observation 651eafbd-c59b-4130-9934-46bd7538a63c · outbound

This paper cites Visual instruction tuning with polite flamingo, 2023.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Visual instruction tuning with polite flamingo, 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.204766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.204766Z digest=sha256:3aa37e00948133bfa34e3d0a49b1562eda0e3f409a6567176b2792506f1a9d78

Observation c831eee2-e73b-417e-9027-f3269ee61fb9 · outbound

This paper cites X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.318544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.318544Z digest=sha256:2d1760b18e9c09b2b2a648c83ff3bb789496a49cf98776328bb636ed81fb6e72

Observation 5cb8c7c1-401d-412e-8e3e-0e7c97190d55 · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.428346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.428346Z digest=sha256:1e390b045b8b27633a48872ae8f771f576fb81d10a4973c937ecce03f8006118

Observation b95a4dfe-e21a-45e8-b711-6b0a9c66063b · outbound

This paper cites A spatial-temporal attention- based method and a new dataset for remote sensing image change detection.Remote Sensing, 12(10), 2020.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion A spatial-temporal attention- based method and a new dataset for remote sensing image change detection.Remote Sensing, 12(10), 2020

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.525458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.525458Z digest=sha256:9c66dd140b4956e573165c89b9c485509b3d3a5e8c529084315e896738c6ebe8

Observation bd16a6c8-3983-407a-bbbc-0a74463d333d · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.643807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.643807Z digest=sha256:496abc451d9f64fa5a74e6a4f4c7e2d30759a4dbbfdbbb7fcc697df397999499

Observation b005e4b9-97e1-43e2-939f-20470ea8ce43 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.747310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.747310Z digest=sha256:c3f9928d664b8831d8f7307122e8662ea9656690b52eb9721708e16f05d51abe

Observation f9834eef-a3de-46c0-862d-9ffd3fdb35b8 · outbound

This paper cites Lawrence Zit- nick.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Lawrence Zit- nick

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.869048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.869048Z digest=sha256:93fb250047590e5573fc62f0c3d7272135d2462b760bf0888a721a8e0ca46055

Observation 9996ac57-7c7b-4b9a-87ba-f060f5621bb1 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.991699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.991699Z digest=sha256:ab9bd469ea2ff90b1a082d0c452e01d37cb3a7ef57e55e7f3e9cf6325b5ad23c

Observation 2099f931-f9a5-4ddb-ad83-46088d849220 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.067033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.067033Z digest=sha256:e9088adf448df2766ea58b96df1f124fa87ec20849142e470f3b93ca6a274242

Observation f14b0b32-e92c-413b-97d5-96e04a4e1db4 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision- Language Models with Instruction Tuning.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion InstructBLIP: Towards General-purpose Vision- Language Models with Instruction Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.157821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.157821Z digest=sha256:6f9c9625fa5c2fa8b023b14e143dc3f1681899e20b82eb678650735a77b86ac4

Observation 60b60378-128d-4a51-bfab-a54d8ec06a15 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.256692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.256692Z digest=sha256:b1152743e37a12c24d390b5f25c4cd61b1ac3bf9044e30bd22a141172bfa2306

Observation 5b0978de-f67e-4d56-95cc-d7f8b6cb2cb6 · outbound

This paper cites VILA$^2$: VILA Augmented VILA.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion VILA$^2$: VILA Augmented VILA

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.380183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.380183Z digest=sha256:9d6783967144029dce2bd715f89c34fa28271318808fff0e0b907c1748c5fd01

Observation d7ecdb0d-2492-49a6-a8cc-52313f9dbd76 · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.490128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.490128Z digest=sha256:4e2a29b5969ce5ddcfcfb66b1c5b7a3dfb96af8547feda200e10bdaf55050dde

Observation b6714ece-d05d-4107-9b94-f2872aac866a · outbound

This paper cites MultiModal-GPT: A Vision and Language Model for Dialogue with Humans.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.592607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.592607Z digest=sha256:aaf51758ae48677ae6e64a68b4c8070134cb83ac7203aa23a10e6ca7b25203d9

Observation 5f0a3843-0cf2-4425-99ad-8c4993c74386 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing, 2017.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing, 2017

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.728782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.728782Z digest=sha256:851bd6abd5c9111eda10476fb4f2831301c9acef00da76f376bf1be3e8a4cde7

Observation 67454761-f5da-407e-a1de-f6713a34c68e · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation, 2019.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Lvis: A dataset for large vocabulary instance segmentation, 2019

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.817491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.817491Z digest=sha256:408a352a373e48012751cf8b3f9a950e904e644054834a743675971f6d864e99

Observation fcce6116-1ad7-4ac3-ab54-2bfeac289962 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.902463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.902463Z digest=sha256:3147c5d8aa99ab9eb5dd832384310e0980a9ce556b574ba4ecdf862f2866f0b2

Observation 138539a2-acc1-45ab-ac6e-ffe5b680a44d · outbound

This paper cites A diagram is worth a dozen images.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion A diagram is worth a dozen images

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:49.015592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:49.015592Z digest=sha256:879cd882e7a71c95e8b737ac733699b6d66280f6c384523401eb3257f1872a0d

Observation 1701edd4-247a-4a7e-8207-88461434a3e7 · outbound

This paper cites A hierarchical approach for generating descriptive image paragraphs, 2017.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion A hierarchical approach for generating descriptive image paragraphs, 2017

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:38:00.174744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:49.108696Z digest=sha256:7d22cd8152c6ecae461d39da6c6c2e2b44f29a0470bfc0c2ccf3e3a03bfabf0c

Observation 52b62b25-d5d4-4d13-9baf-7587316d19f3 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123:32–73, 2017.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123:32–73, 2017

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:38:00.001617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:49.174900Z digest=sha256:f5f7f3d5e93d8884cc86514d0f09f37fa34d71ad9d35569c94e47a49e1dc96e3

Observation be1b42e2-5ddb-4079-abfe-66c8fdcd20f3 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:49.240559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:49.240559Z digest=sha256:6ad1985acb75b3d95b303159ac49361dd875b3bceda8014a862a4874d16619ec

Observation 371ba48b-ccb0-4cbe-afbe-c7981a230f11 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion LLaVA-OneVision: Easy Visual Task Transfer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:49.346808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:49.346808Z digest=sha256:b6385a7f5b8f9be2a7a5854a999ab0dc228d95e311fddb83b8f04d59be9cad85

Observation a77ebe23-372f-4444-aed5-cd0fa947b70b · outbound

This paper cites Llava-med: Training a large language- and-vision assistant for biomedicine in one day.Advances in Neural Information Processing Systems, 36, 2024.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Llava-med: Training a large language- and-vision assistant for biomedicine in one day.Advances in Neural Information Processing Systems, 36, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:59.903935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:49.455236Z digest=sha256:75204e6ebe627ff7327768bf49d4d20712755250b31c2dd5e9fe500be2f1a8de

Observation a8c558c5-a72f-45e9-a9d7-1f7ebbe0173a · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:49.502166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:49.502166Z digest=sha256:43916996e541c871c5faf80a604e8c28e8eab860be78173e5d8a88bb79c7225e

Observation faaad85f-8620-4213-b570-1c8ea8c4ffc9 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion VideoChat: Chat-Centric Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:49.563235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:49.563235Z digest=sha256:3bf338f9bc052d38ab9f48dce1833ba37a0e31e3f6121b7955ddb8f70fc6fb14

Observation 14856309-26db-41ac-9566-3a4537ffc053 · outbound

This paper cites M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:49.620424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:49.620424Z digest=sha256:a9b7418a968f40c3ce6b77f93cf4a71ed03c9fbb0dace21f7c729bf53366eabc

Observation c585b79d-269e-4c02-9308-fcf74de76a1e · outbound

This paper cites Vqa-e: Explaining, elaborating, and enhancing your answers for visual questions, 2018.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Vqa-e: Explaining, elaborating, and enhancing your answers for visual questions, 2018

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:59.790604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:49.686490Z digest=sha256:b0429cc09730220b507c022a436ba43687cb4f4167dbbc2a2231051cf911b1cc

Observation 94969b14-d44d-47c2-8861-5070f20e3530 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Evaluating Object Hallucination in Large Vision-Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:49.747886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:49.747886Z digest=sha256:ac47eb2689e1c8f1377bf880238226044e41746f9c3752e2e90ea65b3c44212d

Observation 7c0382da-693a-4f3a-9776-458f24d0c6e3 · outbound

This paper cites Visual spatial reasoning, 2023.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Visual spatial reasoning, 2023

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:59.666472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:49.838964Z digest=sha256:a0bba21a41ebb53b0be2294bc2bb7d9ca1cead9bc163587f8f5d482d0426acc9

Observation cab813d8-41d8-4017-b08c-62601713ff13 · outbound

This paper cites Mitigating hallucination in large multi-modal models via robust instruction tuning.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Mitigating hallucination in large multi-modal models via robust instruction tuning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:59.545183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:49.905071Z digest=sha256:d6d97512bf4910e86cb284bba4aebda554a3d7a0b0927b22e96ce10cc6837d07

Observation 215f9bce-d532-40d2-9738-92b79cd9c2c3 · outbound

This paper cites Re- moteclip: A vision language foundation model for remote sensing, 2024.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Re- moteclip: A vision language foundation model for remote sensing, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:59.390252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:50.020164Z digest=sha256:1ebd4aee3197b2147991491209a6d5728398849a3939f8ad0636f1af8a2e664f

Observation 7d2fbc4c-7672-43e1-8e6d-30159bf48853 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Improved Baselines with Visual Instruction Tuning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:50.139418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:50.139418Z digest=sha256:31d1daee28a6cb984f67f961db538a932b65ab2c36a0d1c664b534085bf605d5

Observation e31ce97b-c664-4210-859c-2a0a52e2a0a3 · outbound

This paper cites Visual Instruction Tuning.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Visual Instruction Tuning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:59.264947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:50.191652Z digest=sha256:b8827e579801e1ae46371badc96c15036560c9b4f8922a20e1e7ccb0d8a71988

Observation 06f694ee-ae57-43f3-9b51-af7dfbf81379 · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:50.260198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:50.260198Z digest=sha256:1df4a581c22057a78e1d29a3c24cf47370ea3928a9d12af0fe268f19f87ed4a0

Observation ef8d6850-20a7-4319-ab31-b52b2775110e · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521,.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:50.359568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:50.359568Z digest=sha256:dfa61a4e459782d5e7ee546d76c2c9b428bef24fe0bda48afb6392071c2254a8

Observation 1677e10a-078b-4e21-a9f7-49a49750125d · outbound

This paper cites Cheap and quick: Efficient vision- language instruction tuning for large language models.Ad- vances in Neural Information Processing Systems, 36, 2024.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Cheap and quick: Efficient vision- language instruction tuning for large language models.Ad- vances in Neural Information Processing Systems, 36, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:59.127343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:50.416318Z digest=sha256:e584a29e6995f857c3d6081c18db77dc34ddb17a7b52289d8c3f63da2ee1ec75

Observation ca9547b3-7859-44f6-81e2-abd7d2890ecc · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:59.000551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:50.498036Z digest=sha256:26d623a123c5056b13c0bdb5ed2f2a6c049393415d803bcedeee0ea068c4f75d

Observation dd0db20d-f593-4607-8163-e0c6fdd3d277 · outbound

This paper cites ChartQA: A benchmark for question answer- ing about charts with visual and logical reasoning.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion ChartQA: A benchmark for question answer- ing about charts with visual and logical reasoning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:58.871384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:50.601306Z digest=sha256:800845bcd85f252591f288e16e4194fcbd0573737799bec1d23473e60b8036cd

Observation 784ff458-d379-46e3-9ff8-67e5b2e9d52d · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Docvqa: A dataset for vqa on document images

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:58.756318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:50.652514Z digest=sha256:76a37a655ffa1cde7b191b3a8334e196b4acb0a26a7cb8557e0a424ccff6a7b3

Observation d21ae786-c348-4a0a-a84f-33f36383efbd · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Docvqa: A dataset for vqa on document images

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:58.625510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:50.686865Z digest=sha256:35fc59b67a3b993d12d486030adc8caff8e9257da585e371189823b02edb94a9

Observation 06388035-1aea-42a3-be92-4b1475ccde45 · outbound

This paper cites Infographicvqa.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Infographicvqa

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:58.498770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:50.750838Z digest=sha256:ffd11e633c5c45d537e53f409e2e5327e3661a5b682133c5b90f211be69d339a

Observation 4776cc86-5905-4a10-a79d-db2de8dde0bb · outbound

This paper cites Introducing llama 3.1: Our most capable models to date, 2024.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Introducing llama 3.1: Our most capable models to date, 2024

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:58.357548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:50.862393Z digest=sha256:2c540c861155bcbeeb34eaa5dff398d5db8ee4d7e3659a86e23f7da106edb6a6

Observation 66ace162-e509-4286-86b4-c8dd4db39ea8 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Ocr-vqa: Visual question answering by reading text in images

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:58.243549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:50.960826Z digest=sha256:2c55b53c792abbf646a3ffa0182b90ab112411b36eb218e74b34de76fbf87547

Observation a4a480c5-8733-4000-9d67-3dfc28e2529d · outbound

This paper cites GPT-4 Technical Report.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion GPT-4 Technical Report

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:51.068769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:51.068769Z digest=sha256:9495638a432d4b73263d741ffebd499907f97f90b44d909fadc8f16f2b65ba7f

Observation 7056f086-e554-422e-a905-bbf7d132200e · outbound

This paper cites Gpt-4 technical report.ArXiv, 2303:08774,.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Gpt-4 technical report.ArXiv, 2303:08774,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:58.096391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:51.149206Z digest=sha256:11879588dfb78f54a85f56a9fd8c8fab6a20210582523b0c2e1893e0ae5f65ee

Observation 5cae0bc2-df35-4cb3-a429-ba3d5a4969ae · outbound

This paper cites Im2text: Describing images using 1 million captioned pho- tographs.Advances in neural information processing sys- tems, 24, 2011.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Im2text: Describing images using 1 million captioned pho- tographs.Advances in neural information processing sys- tems, 24, 2011

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:57.959041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:51.239072Z digest=sha256:248da3eee2705786dc82e8fd4e3c97a26a2de3ed371ec1ff58bc1a69d5f6f7c1

Observation 54bf0e9f-0d87-4f8f-b08c-d3b7863163e0 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:57.800017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:51.305339Z digest=sha256:8722aa71d8f2753672bac1f4d93f8328aa924f191de1af141057ddafa088de59

Observation cd0bd085-48de-4114-8338-4ecd557293d4 · outbound

This paper cites Connecting vision and lan- guage with localized narratives.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Connecting vision and lan- guage with localized narratives

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:51.426771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:51.426771Z digest=sha256:228d64a094371202e6a231d689acffcd7971f62710f8c0e33d34ac492bf9eaf4

Observation f3b8d652-7b42-461e-b27b-47389529bc39 · outbound

This paper cites Connecting vision and lan- guage with localized narratives, 2020.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Connecting vision and lan- guage with localized narratives, 2020

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:57.640816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:51.548336Z digest=sha256:03a6480f4f7b4cbf98429f26f8dbb42e3f90b75c566a9b936d312f2dd52f3423

Observation 8d661666-9ccd-4410-9e1f-88ff7c3cc442 · outbound

This paper cites Learning Transferable Visual Models from Natural Language Supervi- sion.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Learning Transferable Visual Models from Natural Language Supervi- sion

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:57.481113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:51.644142Z digest=sha256:b63a837887a99d39078ab10c91744a6292825eca2ae9d1c70cee61767e8f48c8

Observation 91df7c65-2a47-41ee-8a80-046a9b9afc38 · outbound

This paper cites Sam 2: Segment anything in images and videos,.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Sam 2: Segment anything in images and videos,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:51.718046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:51.718046Z digest=sha256:2e5e10d1c11e8a736881edc1fbf97e85568989dc9f44db1a50b60ec6d9359c1a

Observation ef1219ee-359e-4e0b-b56a-f987b3b47ed1 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models, 2022.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Laion-5b: An open large-scale dataset for training next generation image-text models, 2022

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:57.268836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:51.814560Z digest=sha256:40a7dc917fa1ed5d5a67aec8fac7d75266b86e0bfeb7907497c90d9d12373c82

Observation d9436920-017a-44a7-afda-5d1af3ec8759 · outbound

This paper cites LAION-5b: An open large-scale dataset for training next generation image-text models.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion LAION-5b: An open large-scale dataset for training next generation image-text models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:57.125445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:51.909444Z digest=sha256:057c843b5c2ae09f26f6bd73312a655037d7314b5c73ad836702e5acf14e602a

Observation 90b898b8-7c71-46e2-95cb-1566aef7fb9e · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowl- edge, 2022.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion A-okvqa: A benchmark for visual question answering using world knowl- edge, 2022

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:56.954674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:51.987410Z digest=sha256:c9009fcf3b7122b196ce76fafaed7536796a37dc694c68656a0336ade3b9241b

Observation 3356d403-ad9b-4824-8be7-6416ca85b22c · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:52.061066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:52.061066Z digest=sha256:da4f93879921d857226603f88c9471614db8742c01175b986e9dbaa86da22c17

Observation 004772e0-2b2d-4d52-b3ea-10f23bce2144 · outbound

This paper cites Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:52.139069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:52.139069Z digest=sha256:9d3a4ca6e5ab8e311c53d5294e623e08f8b5e222c2f0ce003633f3e1a5b02cb6

Observation d92235de-80c8-4b4f-bef1-c0820066dc7c · outbound

This paper cites Textcaps: a dataset for image captioning with reading comprehension, 2020.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Textcaps: a dataset for image captioning with reading comprehension, 2020

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:56.806437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:52.225514Z digest=sha256:eb6322aca49b65aa931ddebd23e24856d2dee8ef9192c61d7edbf0e3cfdf07bd

Observation a49b2bf3-6aba-4f4c-9b1c-21bc58b08f63 · outbound

This paper cites Towards vqa models that can read.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Towards vqa models that can read

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:56.621102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:52.325073Z digest=sha256:ff262d497dd43cc1404f3d97adc12f9a86da7227a99bfde036ff65bacabfaac6

Observation fa194d73-e61a-437c-90c9-b34f3b6187a6 · outbound

This paper cites Expressing visual relationships via language,.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Expressing visual relationships via language,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:56.482837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:52.400976Z digest=sha256:4ea43c63830f2e929247233de64b53fecb5a78d418cdc16c3aa4f324dd12eae0

Observation 7869d6d7-f163-4c65-bb7d-4966d4508cb6 · outbound

This paper cites Vi- sualmrc: Machine reading comprehension on document im- ages, 2021.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Vi- sualmrc: Machine reading comprehension on document im- ages, 2021

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:56.336625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:52.451431Z digest=sha256:0d0f44ca0b6a9cf772d573150ccb2f3864a3daadcb63a63066701f1d42321918

Observation a922125a-d40a-42cd-88e0-b8e911720714 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Gemma 2: Improving Open Language Models at a Practical Size

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:52.521778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:52.521778Z digest=sha256:8051f445f6a6dec68bcbc43ac99eec1b9186dda2fd0f80f83570e429bb8cfab2

Observation 9c42d8e9-8075-4a81-9d7f-c781c5e161b8 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Gemini: A Family of Highly Capable Multimodal Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:52.582430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:52.582430Z digest=sha256:0ed0496328fe5293c8db32073d852b50f1fcbca7320ba76ff5f8ad51612c2b70

Observation 38bf7cf7-0be6-46ee-b247-97608eee72ae · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:52.700977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:52.700977Z digest=sha256:e661af822a3859b8da576e5196ee9122d0c628c23f872fbefa025941549860bb

Observation 380b9a90-ab99-4a3d-8c4d-52846bee1ac0 · outbound

This paper cites To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:52.789583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:52.789583Z digest=sha256:c135d5d1d38a8ac10467a4dc951fe33b64c8b6892679bf9941ae5e40ef007082

Observation 3b1342ee-9a8c-49d0-a595-5350d486a591 · outbound

This paper cites Caption Anything: Interactive Image Description with Diverse Multimodal Controls.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Caption Anything: Interactive Image Description with Diverse Multimodal Controls

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:52.880552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:52.880552Z digest=sha256:6466717ad101470bfb8dbef86a8dac9bb14adefded195236771e4181efd1cbbe

Observation 198dbafe-e916-4c44-9197-d8817203973c · outbound

This paper cites Visionllm: Large language model is also an open- ended decoder for vision-centric tasks.Advances in Neural Information Processing Systems, 36, 2024.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Visionllm: Large language model is also an open- ended decoder for vision-centric tasks.Advances in Neural Information Processing Systems, 36, 2024

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:56.167971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:52.996429Z digest=sha256:f14198b04d54e2ed7e284b4005503e617ae18f14a48510cff91aa65fb7e4852d

Observation 5b49823c-2299-4c5a-b696-a4f930bba88f · outbound

This paper cites MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:53.072593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:53.072593Z digest=sha256:efa349b7a1f46d62f9ef7bb3b5c9182efb570e0610062193cb3bc1cb04d9a5bd

Observation b4595a73-fa6a-4f5a-9acd-dace869b8128 · outbound

This paper cites Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:53.136267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:53.136267Z digest=sha256:12976309b442e933e52e8638ff9bb706d448dacc44814134a41403e5a93cc08b

Observation 49d01f9d-7713-4b5f-be24-b4ddca4e6e66 · outbound

This paper cites Grok-1.5 vision preview, 2024.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Grok-1.5 vision preview, 2024

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:56.031446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:53.189797Z digest=sha256:8592068a0aa683c715cff324967fa2cb169577e545aa0c44dc6c6ae46837af98

Observation 7b2e6457-62fd-42fb-a0e8-3d9c75c8c7c7 · outbound

This paper cites MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:53.282556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:53.282556Z digest=sha256:c14073086622af8dab9dc07df7c894c33c9a02f6864374694f3d473e60caf454

Observation 6942175e-fc92-4847-a539-083a5e111a60 · outbound

This paper cites Depth any- thing v2, 2024.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Depth any- thing v2, 2024

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:53.381001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:53.381001Z digest=sha256:2754af188a1494bbc3b02231d4200a1f7efa5b25d1b5f268a5c944177f6becce

Observation a3f2b3ba-ba36-44b4-85d1-72b640d28a8d · outbound

This paper cites mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:53.457998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:53.457998Z digest=sha256:65655f6e41155b0486066837546e404574253c8b66babdf900ae140aba2dd6af

Observation 0e25ae2a-754a-4ccc-9765-bee5c3134768 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:53.547479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:53.547479Z digest=sha256:f5b5c65739282873eb49a4415f6ec8b668bf1b5c2b6f97e0db190ba2f1f33308

Observation 5b4c1e2a-b493-4099-9e37-d7226f266c96 · outbound

This paper cites Berg, and Tamara L.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Berg, and Tamara L

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:55.940987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:53.643876Z digest=sha256:e15c27da0f084e8cbd86392ce9764cd9dc3b83e7b58a560a158fa37c39454a7c

Observation adab7efc-a886-42f9-bb80-f7790886c47f · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:55.792963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:53.717638Z digest=sha256:9677d00db4a1a404fac5face188b1b28a398af9df8658f2074bba9945f5fe34b

Observation ab4b3daa-e6ad-49dc-927d-d3c5aa0b623d · outbound

This paper cites an unresolved cited work.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:37:55.668988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:53.805718Z digest=sha256:cafa241a00545915327e57a061abb505ba9ff4994d966f4ad605c517a9da854f

Observation b03a1f4f-3d0b-4a85-9256-5af7f73046ce · outbound

This paper cites Rsvg: Exploring data and models for visual grounding on remote sensing data.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Rsvg: Exploring data and models for visual grounding on remote sensing data

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:55.534422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:53.898696Z digest=sha256:7f230e77d825e775f905d2af92bb9f4a4e0d237f29496d5dff7bce492134f06f

Observation a6a07919-ca3a-4726-933d-7a98911f55b3 · outbound

This paper cites Lmms- eval: Reality check on the evaluation of large multimodal models, 2024.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Lmms- eval: Reality check on the evaluation of large multimodal models, 2024

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:55.393162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:53.962404Z digest=sha256:cb99834e61c00003bdc5c78aae28699d1e972da201c4bd45a416d42fcf0e7b9f

Observation 15480673-59f4-46b2-b0ea-d123ad922208 · outbound

This paper cites InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:54.044923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:54.044923Z digest=sha256:e1b7e1b28d96b296d58fa82865ad5f6eb62dffb3ddfe72dcfb4a3c968e129b82

Observation d2f15a56-beea-495f-9df7-5032803d62a4 · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:54.116967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:54.116967Z digest=sha256:cb3ceeb34db99f789b1ec18c232e0afc98056e5ebf3b62ce862a3e982ccef3fa

Observation 1253bd24-febe-40c5-a414-cd03f10b8c9b · outbound

This paper cites Hierarchical and robust convolutional neural network for very high-resolution remote sensing object detection.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Hierarchical and robust convolutional neural network for very high-resolution remote sensing object detection

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:55.246086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:54.194802Z digest=sha256:97d8b1b362aad22e44d55b87c0db2b83b4b676e574455f8f95fbc6a596a11cae

Observation a346cf23-6cad-4a8b-867c-efcd33e76a44 · outbound

This paper cites LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:54.258922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:54.258922Z digest=sha256:76ba1a9ec35b0298a98d3190b97e2927e6bd2229778873115a13972fd716d864

Observation 05c6122e-8d8d-44a8-bd5b-e427e1374c22 · outbound

This paper cites SVIT: Scaling up Visual Instruction Tuning.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion SVIT: Scaling up Visual Instruction Tuning

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:54.331667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:54.331667Z digest=sha256:709d12c5389b664d8507c8f168203bf450f2aa96cf0833ff5183a1aead71f224

Observation f40fa1de-f74d-43db-8c29-353c44361fce · outbound

This paper cites ChatBridge: Bridging Modalities with Large Language Model as a Language Catalyst.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion ChatBridge: Bridging Modalities with Large Language Model as a Language Catalyst

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:54.397753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:54.397753Z digest=sha256:bc38cb0e9cfc1d64a71c1378c96a49b538dd07ad3aca8b4c9f5331383ba7253b

Observation 6619e153-8a2f-46be-a2bd-06e10ce65ca6 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:54.505962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:54.505962Z digest=sha256:0548748d290600a6ff6fd8e44231c7f266e189c05ca69ab372b96a7c92e42c8e

Observation a5dbdd01-7047-4333-b9db-7191c179e2e4 · outbound

This paper cites Visual7w: Grounded question answering in images, 2016.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Visual7w: Grounded question answering in images, 2016

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:55.062673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:54.579360Z digest=sha256:e8947acd05c016a72ea4072a4d6fb980819bb08180c07a02f19cc10f0592ee6d

Pith citing papers

No inbound Pith citation observations are available.