Pith. sign in

Paper Citation Record · LEDGER

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion

As of 15 August 2026, this Paper Citation Record lists 90 of 90 outbound references and 0 inbound Pith citation observations for arXiv:2505.18115.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18115 v1

Coverage vector

measured 90 of 90 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:37:54.579360Z

measured 90 of 90 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

90 of 90 outbound references displayed

  • verified exact0
  • verified fuzzy36
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 46b57075-b50b-4e2b-8eda-d6c7d4a00491 · outbound

This paper cites anthropic.com/news/claude-3-family, 2024.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion anthropic.com/news/claude-3-family, 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:46.891987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:46.891987Z digest=sha256:ba27aec9f77ca5b5520a840894c2a75f130cf32e6aba36d32a97a2d0e8210d98

Observation 782c97b6-a0ee-48a1-98b5-ddaddf9a3ffe · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:46.984855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:46.984855Z digest=sha256:7997ae6c8ee3dd53f39f78565a2412ee3fdf3eab64e71ff49da0c747b13631a2

Observation 104d0814-623d-4e48-b819-5e760535fa89 · outbound

This paper cites Easyocr: Ready-to-use ocr with 80+ supported languages, 2020.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Easyocr: Ready-to-use ocr with 80+ supported languages, 2020

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.057781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.057781Z digest=sha256:b71315901604b0d559f7e7c27db93c68bd5601c4f8fa13e8806622b90bc0c60e

Observation 796be892-b44a-4377-ac8d-ab004995c347 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Flamingo: a Visual Language Model for Few-Shot Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.145455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.145455Z digest=sha256:b67fd90a080d4eead5e97020c1bb0867919efc21cac04870706502f64a15902a

Observation 651eafbd-c59b-4130-9934-46bd7538a63c · outbound

This paper cites Visual instruction tuning with polite flamingo, 2023.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Visual instruction tuning with polite flamingo, 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.204766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.204766Z digest=sha256:871532d1d43165d30a774d72e57eaacfa81387f04064b2723fcdb756908351a9

Observation c831eee2-e73b-417e-9027-f3269ee61fb9 · outbound

This paper cites X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.318544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.318544Z digest=sha256:b59e953f98fa60bbb9cc8acb064492ded1c297f619de0e3365c717775b27519a

Observation 5cb8c7c1-401d-412e-8e3e-0e7c97190d55 · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.428346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.428346Z digest=sha256:ceea6fb9d089d5917485c56b7f9b1c06cabb07edb73c76e350380007ac95d9fb

Observation b95a4dfe-e21a-45e8-b711-6b0a9c66063b · outbound

This paper cites A spatial-temporal attention- based method and a new dataset for remote sensing image change detection.Remote Sensing, 12(10), 2020.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion A spatial-temporal attention- based method and a new dataset for remote sensing image change detection.Remote Sensing, 12(10), 2020

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.525458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.525458Z digest=sha256:e235e16ba249ebbb7dbe9f71227c3b14cf8f662627fb03a3c70534d671b4513c

Observation bd16a6c8-3983-407a-bbbc-0a74463d333d · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.643807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.643807Z digest=sha256:ccfe2bc106bd9957287350f4a6c8c8418528bf63a6b7cb2221a6557bbc59eeaa

Observation b005e4b9-97e1-43e2-939f-20470ea8ce43 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.747310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.747310Z digest=sha256:4fd5629909bb50430db629fb85f30706f81be34a511776186619a9d3751d7a54

Observation f9834eef-a3de-46c0-862d-9ffd3fdb35b8 · outbound

This paper cites Lawrence Zit- nick.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Lawrence Zit- nick

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.869048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.869048Z digest=sha256:5b67accc48663a5051f032cffa618993e90da8fc7194db1d083ad92aef9c6890

Observation 9996ac57-7c7b-4b9a-87ba-f060f5621bb1 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.991699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.991699Z digest=sha256:ab1cea43add9823974fbc71effad47fe3afc4e49c59faf75c54af0f7eb9467cd

Observation 2099f931-f9a5-4ddb-ad83-46088d849220 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.067033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.067033Z digest=sha256:0579416518126cfb2963a9f95d8b1663bd59ce60c74be1ca4b65246b7ef5fdcd

Observation f14b0b32-e92c-413b-97d5-96e04a4e1db4 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision- Language Models with Instruction Tuning.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion InstructBLIP: Towards General-purpose Vision- Language Models with Instruction Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.157821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.157821Z digest=sha256:cb4da5700ed9ceb1f40c02b5af5386c2d0c8424cfb646c4b601be173cd7a8796

Observation 60b60378-128d-4a51-bfab-a54d8ec06a15 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.256692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.256692Z digest=sha256:558672e9665b72ad7c13e3dd38c54a56fa7c6d16d925883adba9b785dc9de7dc

Observation 5b0978de-f67e-4d56-95cc-d7f8b6cb2cb6 · outbound

This paper cites VILA$^2$: VILA Augmented VILA.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion VILA$^2$: VILA Augmented VILA

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.380183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.380183Z digest=sha256:226d1e8334873c79d3f9dd4e7ccf6270a178c0259f59fff618d0b3d0a28cb6bf

Observation d7ecdb0d-2492-49a6-a8cc-52313f9dbd76 · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.490128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.490128Z digest=sha256:bfb9a75f61610c537bf5f17cdd06d558ff1799c2ff36d09c2fc6d22a3f6987c4

Observation b6714ece-d05d-4107-9b94-f2872aac866a · outbound

This paper cites MultiModal-GPT: A Vision and Language Model for Dialogue with Humans.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.592607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.592607Z digest=sha256:e8fb020ef6587a6ba13af970ef1e18fa238e4dc2754e0be1a320915e8c77ff77

Observation 5f0a3843-0cf2-4425-99ad-8c4993c74386 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing, 2017.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing, 2017

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.728782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.728782Z digest=sha256:9e7f0867e08fa4896489223ecde794bb0fcc61af8f6d3b8f8a166903032c5944

Observation 67454761-f5da-407e-a1de-f6713a34c68e · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation, 2019.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Lvis: A dataset for large vocabulary instance segmentation, 2019

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.817491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.817491Z digest=sha256:ef737f1e396f6a8cef3d44a82ab5780ad115c08ce5d1c0df3bf7895ac0f3ad15

Observation fcce6116-1ad7-4ac3-ab54-2bfeac289962 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.902463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.902463Z digest=sha256:ac844d7cef1c3f1840b52a1cc15026ee3ed4bde2effad3dc7f9d6ba889c9ac34

Observation 138539a2-acc1-45ab-ac6e-ffe5b680a44d · outbound

This paper cites A diagram is worth a dozen images.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion A diagram is worth a dozen images

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:49.015592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:49.015592Z digest=sha256:0df1127b9ad5c4fb8c0843311d245ca17fadb3caf654b3dcf63c3df4bc2b198c

Observation 1701edd4-247a-4a7e-8207-88461434a3e7 · outbound

This paper cites A hierarchical approach for generating descriptive image paragraphs, 2017.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion A hierarchical approach for generating descriptive image paragraphs, 2017

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:38:00.174744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:49.108696Z digest=sha256:e33ef21a9d4f7f30cebf470af96148427fcba072a4ecfc282a0fd1dfb72d288e

Observation 52b62b25-d5d4-4d13-9baf-7587316d19f3 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123:32–73, 2017.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123:32–73, 2017

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:38:00.001617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:49.174900Z digest=sha256:aadcd44c790d3d639c27a5f2575d91e7c45e5c6892c3a2fffb4ac96cfc4aa0e9

Observation be1b42e2-5ddb-4079-abfe-66c8fdcd20f3 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:49.240559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:49.240559Z digest=sha256:8aea977e630b2527629cfcd7d0fee5f3d096f1f81cf315aae79af2c3823e6f27

Observation 371ba48b-ccb0-4cbe-afbe-c7981a230f11 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion LLaVA-OneVision: Easy Visual Task Transfer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:49.346808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:49.346808Z digest=sha256:d364882af8d652d00948de6a068b55c119cc89cc7d12a74b7d8dadb6dd466e4a

Observation a77ebe23-372f-4444-aed5-cd0fa947b70b · outbound

This paper cites Llava-med: Training a large language- and-vision assistant for biomedicine in one day.Advances in Neural Information Processing Systems, 36, 2024.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Llava-med: Training a large language- and-vision assistant for biomedicine in one day.Advances in Neural Information Processing Systems, 36, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:59.903935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:49.455236Z digest=sha256:4874fefbae9b23cf2e9381a4c8083e379019ac9739bd25623dcd69a030379084

Observation a8c558c5-a72f-45e9-a9d7-1f7ebbe0173a · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:49.502166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:49.502166Z digest=sha256:da72c1c7285890bfb29efdb14370b517e12a67352f0eaa157af8de225adee110

Observation faaad85f-8620-4213-b570-1c8ea8c4ffc9 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion VideoChat: Chat-Centric Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:49.563235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:49.563235Z digest=sha256:25c65b04d389d741f471801b1f5bbf7eadc6258dd1d30e27f9649c5fa122159b

Observation 14856309-26db-41ac-9566-3a4537ffc053 · outbound

This paper cites M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:49.620424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:49.620424Z digest=sha256:a1c950c37bf6b8f01677e65071a11c11ea7408df860077087cbe5042a16b034d

Observation c585b79d-269e-4c02-9308-fcf74de76a1e · outbound

This paper cites Vqa-e: Explaining, elaborating, and enhancing your answers for visual questions, 2018.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Vqa-e: Explaining, elaborating, and enhancing your answers for visual questions, 2018

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:59.790604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:49.686490Z digest=sha256:cb072d7a2e3e2fa54033a9ddfc65d5c4ad52f1d90542df8b31e0e21b902a5f84

Observation 94969b14-d44d-47c2-8861-5070f20e3530 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Evaluating Object Hallucination in Large Vision-Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:49.747886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:49.747886Z digest=sha256:a783375fc916dbd1946fb0916c09c49bbe94ed801491bb74c643893aa9afadb2

Observation 7c0382da-693a-4f3a-9776-458f24d0c6e3 · outbound

This paper cites Visual spatial reasoning, 2023.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Visual spatial reasoning, 2023

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:59.666472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:49.838964Z digest=sha256:ab8d5863df6e5a2a06e4687be92c752769d25349f32c53a2a56b88e7c8bfeebe

Observation cab813d8-41d8-4017-b08c-62601713ff13 · outbound

This paper cites Mitigating hallucination in large multi-modal models via robust instruction tuning.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Mitigating hallucination in large multi-modal models via robust instruction tuning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:59.545183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:49.905071Z digest=sha256:e68dae2e38941d51cb9da46d1f174605b2698b4e32fdb0ce588fda786220b0b9

Observation 215f9bce-d532-40d2-9738-92b79cd9c2c3 · outbound

This paper cites Re- moteclip: A vision language foundation model for remote sensing, 2024.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Re- moteclip: A vision language foundation model for remote sensing, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:59.390252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:50.020164Z digest=sha256:e7b2713505c42af30775574a7bfd1233f5f5e70fe38a11760a3db6ba3b4c7f5a

Observation 7d2fbc4c-7672-43e1-8e6d-30159bf48853 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Improved Baselines with Visual Instruction Tuning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:50.139418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:50.139418Z digest=sha256:e7183ed3ea16955dcd6ae95aef5167055f8df659b4aef5a34bf7b7932acd32f5

Observation e31ce97b-c664-4210-859c-2a0a52e2a0a3 · outbound

This paper cites Visual Instruction Tuning.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Visual Instruction Tuning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:59.264947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:50.191652Z digest=sha256:b5c71565e1f28caab192f2c5075c19264f891e88c492334ba946490acf844c26

Observation 06f694ee-ae57-43f3-9b51-af7dfbf81379 · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:50.260198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:50.260198Z digest=sha256:16b3af7aae83c94749722fddc3f7cff8d3a07a27e2ec4d7d349ec089042a59f2

Observation ef8d6850-20a7-4319-ab31-b52b2775110e · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521,.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:50.359568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:50.359568Z digest=sha256:3d964a450a43401b50e715743063b8ef38fd2f40c12b7bf24708bcea387e4791

Observation 1677e10a-078b-4e21-a9f7-49a49750125d · outbound

This paper cites Cheap and quick: Efficient vision- language instruction tuning for large language models.Ad- vances in Neural Information Processing Systems, 36, 2024.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Cheap and quick: Efficient vision- language instruction tuning for large language models.Ad- vances in Neural Information Processing Systems, 36, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:59.127343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:50.416318Z digest=sha256:41861dffd890cfa15132964fa3d285c05a442ea651189fa8144b1306218f4693

Observation ca9547b3-7859-44f6-81e2-abd7d2890ecc · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:59.000551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:50.498036Z digest=sha256:f458778013f5b6955313ab68fbb0f202db4c917fb91e83247f50005c312cfe9b

Observation dd0db20d-f593-4607-8163-e0c6fdd3d277 · outbound

This paper cites ChartQA: A benchmark for question answer- ing about charts with visual and logical reasoning.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion ChartQA: A benchmark for question answer- ing about charts with visual and logical reasoning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:58.871384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:50.601306Z digest=sha256:69b239eca4bf81bfa9f74cf87f347a1c22d74dbf937eef7d124e08f838724ff8

Observation 784ff458-d379-46e3-9ff8-67e5b2e9d52d · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Docvqa: A dataset for vqa on document images

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:58.756318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:50.652514Z digest=sha256:4f5e0a4277ea3d36679589da8254d536f3a7827778f9fdea73c78067fbd7c60d

Observation d21ae786-c348-4a0a-a84f-33f36383efbd · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Docvqa: A dataset for vqa on document images

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:58.625510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:50.686865Z digest=sha256:409ad7886e95ab818224f92cfada72fd396c78833c75716d04bb050807a487d8

Observation 06388035-1aea-42a3-be92-4b1475ccde45 · outbound

This paper cites Infographicvqa.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Infographicvqa

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:58.498770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:50.750838Z digest=sha256:7bdb27e14e70a9a3696578c0702bd3403a0919a1e02136e04354082f3826a43c

Observation 4776cc86-5905-4a10-a79d-db2de8dde0bb · outbound

This paper cites Introducing llama 3.1: Our most capable models to date, 2024.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Introducing llama 3.1: Our most capable models to date, 2024

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:58.357548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:50.862393Z digest=sha256:d6cf03b80d1099eb0fe82b47827c7b70e2b503b52836613973c5b5250a78925a

Observation 66ace162-e509-4286-86b4-c8dd4db39ea8 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Ocr-vqa: Visual question answering by reading text in images

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:58.243549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:50.960826Z digest=sha256:5dbba76601670b6b1eeab1a762c0f35002741d0e1e9a05c21996cbec56d139d8

Observation a4a480c5-8733-4000-9d67-3dfc28e2529d · outbound

This paper cites GPT-4 Technical Report.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion GPT-4 Technical Report

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:51.068769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:51.068769Z digest=sha256:7feaa2cb5cfa1d6c55c4d6e460cf821ee27f465e578dd2d9a3c56917225631e8

Observation 7056f086-e554-422e-a905-bbf7d132200e · outbound

This paper cites Gpt-4 technical report.ArXiv, 2303:08774,.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Gpt-4 technical report.ArXiv, 2303:08774,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:58.096391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:51.149206Z digest=sha256:32b3f018e36e21d13c319b5e571b1d2411e4d464b5162e4c2be355651e5c096e

Observation 5cae0bc2-df35-4cb3-a429-ba3d5a4969ae · outbound

This paper cites Im2text: Describing images using 1 million captioned pho- tographs.Advances in neural information processing sys- tems, 24, 2011.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Im2text: Describing images using 1 million captioned pho- tographs.Advances in neural information processing sys- tems, 24, 2011

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:57.959041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:51.239072Z digest=sha256:9c23d53a8a35a20d4cb623cae36c0e15f562e28c9e0bd2ac7b07ca24b4ca31fc

Observation 54bf0e9f-0d87-4f8f-b08c-d3b7863163e0 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:57.800017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:51.305339Z digest=sha256:0981df0eb41d1bd942bb576323644b77ee8f7c091bd5afbf42b0cf121814be7f

Observation cd0bd085-48de-4114-8338-4ecd557293d4 · outbound

This paper cites Connecting vision and lan- guage with localized narratives.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Connecting vision and lan- guage with localized narratives

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:51.426771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:51.426771Z digest=sha256:a02b3359ab796d38e2c407ddb41bb9308829d63a73c0fe68e7500955256bf670

Observation f3b8d652-7b42-461e-b27b-47389529bc39 · outbound

This paper cites Connecting vision and lan- guage with localized narratives, 2020.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Connecting vision and lan- guage with localized narratives, 2020

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:57.640816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:51.548336Z digest=sha256:5922a4286d5bea777e68f8f24bc66028afac0e7cce5b1703e649e09aa2da89c4

Observation 8d661666-9ccd-4410-9e1f-88ff7c3cc442 · outbound

This paper cites Learning Transferable Visual Models from Natural Language Supervi- sion.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Learning Transferable Visual Models from Natural Language Supervi- sion

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:57.481113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:51.644142Z digest=sha256:6ac277136499bf7044c3f13694e4b1bbdff3b5d99b276cbb5c95d2c4d33c4ccd

Observation 91df7c65-2a47-41ee-8a80-046a9b9afc38 · outbound

This paper cites Sam 2: Segment anything in images and videos,.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Sam 2: Segment anything in images and videos,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:51.718046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:51.718046Z digest=sha256:b8b0402b72c55c1021b209f607c8bc9aeb95e9e42579d7ecc62eae6d402bc7d6

Observation ef1219ee-359e-4e0b-b56a-f987b3b47ed1 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models, 2022.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Laion-5b: An open large-scale dataset for training next generation image-text models, 2022

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:57.268836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:51.814560Z digest=sha256:afd52f9c1ee97042c458293470789572bad483feb24a7093fd0ab66a708821c6

Observation d9436920-017a-44a7-afda-5d1af3ec8759 · outbound

This paper cites LAION-5b: An open large-scale dataset for training next generation image-text models.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion LAION-5b: An open large-scale dataset for training next generation image-text models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:57.125445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:51.909444Z digest=sha256:de39ee7284b9f5529faf1ef3ef01535a0fc0caff04103dc75d83c4e2927ae6d7

Observation 90b898b8-7c71-46e2-95cb-1566aef7fb9e · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowl- edge, 2022.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion A-okvqa: A benchmark for visual question answering using world knowl- edge, 2022

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:56.954674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:51.987410Z digest=sha256:785706e086cab6c6ff2c2190f95a73801a6d8d77267257f9a12c0768a2d628e4

Observation 3356d403-ad9b-4824-8be7-6416ca85b22c · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:52.061066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:52.061066Z digest=sha256:ad26e73a39c0958d1c627ba4910caf290b9c4f8ae7c24a5bc4bf0b190c7e0e6a

Observation 004772e0-2b2d-4d52-b3ea-10f23bce2144 · outbound

This paper cites Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:52.139069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:52.139069Z digest=sha256:1daf71e8c483e982201508fbc9b7dda7576f39af262c9e5dc4030c50d81062b8

Observation d92235de-80c8-4b4f-bef1-c0820066dc7c · outbound

This paper cites Textcaps: a dataset for image captioning with reading comprehension, 2020.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Textcaps: a dataset for image captioning with reading comprehension, 2020

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:56.806437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:52.225514Z digest=sha256:441aea82048b8f4ace9e53690d40716444a51c0feb3d0f336353b935196f3903

Observation a49b2bf3-6aba-4f4c-9b1c-21bc58b08f63 · outbound

This paper cites Towards vqa models that can read.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Towards vqa models that can read

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:56.621102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:52.325073Z digest=sha256:d75db4591fd687451c89295b66f80b4e00493e2669fa96641583ee228ddce1ba

Observation fa194d73-e61a-437c-90c9-b34f3b6187a6 · outbound

This paper cites Expressing visual relationships via language,.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Expressing visual relationships via language,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:56.482837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:52.400976Z digest=sha256:9766e5b09645d24e4d63fcaad8eda882299e8640cd2887454febfacbb7998d8f

Observation 7869d6d7-f163-4c65-bb7d-4966d4508cb6 · outbound

This paper cites Vi- sualmrc: Machine reading comprehension on document im- ages, 2021.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Vi- sualmrc: Machine reading comprehension on document im- ages, 2021

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:56.336625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:52.451431Z digest=sha256:4b00706f64f50dad0d7acd2e4ba498b66eec239409919ff102570f5cf02fc910

Observation a922125a-d40a-42cd-88e0-b8e911720714 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Gemma 2: Improving Open Language Models at a Practical Size

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:52.521778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:52.521778Z digest=sha256:02b06042f9806892ea16481f035d0aaaab35c099307be33c83c55758ee9aa7d1

Observation 9c42d8e9-8075-4a81-9d7f-c781c5e161b8 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Gemini: A Family of Highly Capable Multimodal Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:52.582430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:52.582430Z digest=sha256:faed8a481769f7e9461bebd07879519d98c80de893ccec1b470f7be469892920

Observation 38bf7cf7-0be6-46ee-b247-97608eee72ae · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:52.700977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:52.700977Z digest=sha256:f1c318ec70f34b4a696e1c210602f80f8c210ec7e2b02608f50bad909c5d9a79

Observation 380b9a90-ab99-4a3d-8c4d-52846bee1ac0 · outbound

This paper cites To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:52.789583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:52.789583Z digest=sha256:44b9de281e0019f51ae38fe510644f9ac5a2594df424a2b286a728dfbc28f1fb

Observation 3b1342ee-9a8c-49d0-a595-5350d486a591 · outbound

This paper cites Caption Anything: Interactive Image Description with Diverse Multimodal Controls.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Caption Anything: Interactive Image Description with Diverse Multimodal Controls

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:52.880552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:52.880552Z digest=sha256:38cc7a236c2b306caf6afb4d7dc847b50ed66a8d02c322c239e97f3652e4f0e6

Observation 198dbafe-e916-4c44-9197-d8817203973c · outbound

This paper cites Visionllm: Large language model is also an open- ended decoder for vision-centric tasks.Advances in Neural Information Processing Systems, 36, 2024.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Visionllm: Large language model is also an open- ended decoder for vision-centric tasks.Advances in Neural Information Processing Systems, 36, 2024

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:56.167971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:52.996429Z digest=sha256:618d43552650c529d28e264c207225ccb46ac598cee9fcd9e3203bb44a256bb5

Observation 5b49823c-2299-4c5a-b696-a4f930bba88f · outbound

This paper cites MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:53.072593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:53.072593Z digest=sha256:07635aa77a8dfd335ce14b143bed84f70f542ae05abbb68d0da21f0666dabce9

Observation b4595a73-fa6a-4f5a-9acd-dace869b8128 · outbound

This paper cites Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:53.136267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:53.136267Z digest=sha256:e7c7e5cf55b7573bfbf5c3334d6abb573b4573badc03456ec364828b422327bd

Observation 49d01f9d-7713-4b5f-be24-b4ddca4e6e66 · outbound

This paper cites Grok-1.5 vision preview, 2024.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Grok-1.5 vision preview, 2024

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:56.031446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:53.189797Z digest=sha256:0c9209021a5bcdf831ca8feefb7633c26167e2b10fb49af7900f45bff77ec0b3

Observation 7b2e6457-62fd-42fb-a0e8-3d9c75c8c7c7 · outbound

This paper cites MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:53.282556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:53.282556Z digest=sha256:dc81820f00daed2cce997b300390d2546db8f4485544eabe7d1cbf9dc4d60c38

Observation 6942175e-fc92-4847-a539-083a5e111a60 · outbound

This paper cites Depth any- thing v2, 2024.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Depth any- thing v2, 2024

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:53.381001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:53.381001Z digest=sha256:68c46647cfa4a06798f6e72ddbfed018bf57ed9ad0d71442d6c546334e519b89

Observation a3f2b3ba-ba36-44b4-85d1-72b640d28a8d · outbound

This paper cites mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:53.457998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:53.457998Z digest=sha256:de057b49273895e546cf1367905bc7017aab6859e4afc36e456e1aebe1aec02e

Observation 0e25ae2a-754a-4ccc-9765-bee5c3134768 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:53.547479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:53.547479Z digest=sha256:ff7a37073fd7b9372ccd41d4c014540234b90d98aa65456b18ec88972162a115

Observation 5b4c1e2a-b493-4099-9e37-d7226f266c96 · outbound

This paper cites Berg, and Tamara L.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Berg, and Tamara L

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:55.940987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:53.643876Z digest=sha256:f889bd776f4fdb2fca519a613db38e1f685553e787c6efd932560c03e77b0b50

Observation adab7efc-a886-42f9-bb80-f7790886c47f · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:55.792963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:53.717638Z digest=sha256:ac3f4eda6e72469a5d51e505b1544c24ec4d1b9a952ba38eafce77ec39b17064

Observation ab4b3daa-e6ad-49dc-927d-d3c5aa0b623d · outbound

This paper cites an unresolved cited work.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:37:55.668988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:53.805718Z digest=sha256:80d6adbef060c3f6d590b38f88ba2f82c783446322ec816d8b3735ac0fd0ad4a

Observation b03a1f4f-3d0b-4a85-9256-5af7f73046ce · outbound

This paper cites Rsvg: Exploring data and models for visual grounding on remote sensing data.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Rsvg: Exploring data and models for visual grounding on remote sensing data

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:55.534422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:53.898696Z digest=sha256:20c44603dcb870a498c6f682359acdb4c457bee83f7ca9babe1ef1b8a7c37c07

Observation a6a07919-ca3a-4726-933d-7a98911f55b3 · outbound

This paper cites Lmms- eval: Reality check on the evaluation of large multimodal models, 2024.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Lmms- eval: Reality check on the evaluation of large multimodal models, 2024

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:55.393162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:53.962404Z digest=sha256:5f0e075f2fe379b46f1f928706b29ce509109ab6b85ffdbd7017c1d5be216d61

Observation 15480673-59f4-46b2-b0ea-d123ad922208 · outbound

This paper cites InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:54.044923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:54.044923Z digest=sha256:4f2b21a852d83150c3b59befff88521d0a1d16ffc5b4862ce0af563d53856899

Observation d2f15a56-beea-495f-9df7-5032803d62a4 · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:54.116967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:54.116967Z digest=sha256:3be401a30e31d20bbb7f91fe4bd2201b4f62784453addb0bb53de46c7ea27069

Observation 1253bd24-febe-40c5-a414-cd03f10b8c9b · outbound

This paper cites Hierarchical and robust convolutional neural network for very high-resolution remote sensing object detection.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Hierarchical and robust convolutional neural network for very high-resolution remote sensing object detection

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:55.246086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:54.194802Z digest=sha256:46773fb1b2e43a88e0c2a1dd423ad7dc8bb97f5c5d909d33929d5f177c4445d8

Observation a346cf23-6cad-4a8b-867c-efcd33e76a44 · outbound

This paper cites LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:54.258922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:54.258922Z digest=sha256:1e9577f9605056f7412c3110b33fce08bd05cd721368cfaf765960afeb17f04e

Observation 05c6122e-8d8d-44a8-bd5b-e427e1374c22 · outbound

This paper cites SVIT: Scaling up Visual Instruction Tuning.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion SVIT: Scaling up Visual Instruction Tuning

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:54.331667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:54.331667Z digest=sha256:8b487d93b225f998d43f9cdd3d2d10881dc4fc22d0423796e42a27ec571a7e4d

Observation f40fa1de-f74d-43db-8c29-353c44361fce · outbound

This paper cites ChatBridge: Bridging Modalities with Large Language Model as a Language Catalyst.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion ChatBridge: Bridging Modalities with Large Language Model as a Language Catalyst

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:54.397753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:54.397753Z digest=sha256:c0abc5a4cb5c35563ad5ec45e6da078b119d2d7b040554c005e2d1eba7e5caef

Observation 6619e153-8a2f-46be-a2bd-06e10ce65ca6 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:54.505962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:54.505962Z digest=sha256:34979f645a08812d1c99bbea72991e9e52d11c6682810b1a6990d62187c3c29f

Observation a5dbdd01-7047-4333-b9db-7191c179e2e4 · outbound

This paper cites Visual7w: Grounded question answering in images, 2016.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Visual7w: Grounded question answering in images, 2016

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:55.062673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:37:54.579360Z digest=sha256:9801759d6e3679e2eaac06e6a95f3541b2751ee3dd1a2f0dffd6fdf559eb4d5a

Pith citing papers

No inbound Pith citation observations are available.