Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:37:54.579360Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 90 of 90 outbound references and 0 inbound Pith citation observations for arXiv:2505.18115.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:37:54.579360Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
90 of 90 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 46b57075-b50b-4e2b-8eda-d6c7d4a00491 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion anthropic.com/news/claude-3-family, 2024
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 782c97b6-a0ee-48a1-98b5-ddaddf9a3ffe · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 104d0814-623d-4e48-b819-5e760535fa89 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Easyocr: Ready-to-use ocr with 80+ supported languages, 2020
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 796be892-b44a-4377-ac8d-ab004995c347 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Flamingo: a Visual Language Model for Few-Shot Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 651eafbd-c59b-4130-9934-46bd7538a63c · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Visual instruction tuning with polite flamingo, 2023
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c831eee2-e73b-417e-9027-f3269ee61fb9 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cb8c7c1-401d-412e-8e3e-0e7c97190d55 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b95a4dfe-e21a-45e8-b711-6b0a9c66063b · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion A spatial-temporal attention- based method and a new dataset for remote sensing image change detection.Remote Sensing, 12(10), 2020
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd16a6c8-3983-407a-bbbc-0a74463d333d · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b005e4b9-97e1-43e2-939f-20470ea8ce43 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9834eef-a3de-46c0-862d-9ffd3fdb35b8 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Lawrence Zit- nick
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9996ac57-7c7b-4b9a-87ba-f060f5621bb1 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Microsoft COCO Captions: Data Collection and Evaluation Server
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2099f931-f9a5-4ddb-ad83-46088d849220 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f14b0b32-e92c-413b-97d5-96e04a4e1db4 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion InstructBLIP: Towards General-purpose Vision- Language Models with Instruction Tuning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60b60378-128d-4a51-bfab-a54d8ec06a15 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b0978de-f67e-4d56-95cc-d7f8b6cb2cb6 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion VILA$^2$: VILA Augmented VILA
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7ecdb0d-2492-49a6-a8cc-52313f9dbd76 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6714ece-d05d-4107-9b94-f2872aac866a · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f0a3843-0cf2-4425-99ad-8c4993c74386 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing, 2017
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67454761-f5da-407e-a1de-f6713a34c68e · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Lvis: A dataset for large vocabulary instance segmentation, 2019
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcce6116-1ad7-4ac3-ab54-2bfeac289962 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 138539a2-acc1-45ab-ac6e-ffe5b680a44d · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion A diagram is worth a dozen images
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1701edd4-247a-4a7e-8207-88461434a3e7 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion A hierarchical approach for generating descriptive image paragraphs, 2017
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 52b62b25-d5d4-4d13-9baf-7587316d19f3 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123:32–73, 2017
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation be1b42e2-5ddb-4079-abfe-66c8fdcd20f3 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 371ba48b-ccb0-4cbe-afbe-c7981a230f11 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion LLaVA-OneVision: Easy Visual Task Transfer
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a77ebe23-372f-4444-aed5-cd0fa947b70b · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Llava-med: Training a large language- and-vision assistant for biomedicine in one day.Advances in Neural Information Processing Systems, 36, 2024
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a8c558c5-a72f-45e9-a9d7-1f7ebbe0173a · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faaad85f-8620-4213-b570-1c8ea8c4ffc9 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion VideoChat: Chat-Centric Video Understanding
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14856309-26db-41ac-9566-3a4537ffc053 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c585b79d-269e-4c02-9308-fcf74de76a1e · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Vqa-e: Explaining, elaborating, and enhancing your answers for visual questions, 2018
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 94969b14-d44d-47c2-8861-5070f20e3530 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Evaluating Object Hallucination in Large Vision-Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c0382da-693a-4f3a-9776-458f24d0c6e3 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Visual spatial reasoning, 2023
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cab813d8-41d8-4017-b08c-62601713ff13 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Mitigating hallucination in large multi-modal models via robust instruction tuning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 215f9bce-d532-40d2-9738-92b79cd9c2c3 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Re- moteclip: A vision language foundation model for remote sensing, 2024
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7d2fbc4c-7672-43e1-8e6d-30159bf48853 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Improved Baselines with Visual Instruction Tuning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e31ce97b-c664-4210-859c-2a0a52e2a0a3 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Visual Instruction Tuning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 06f694ee-ae57-43f3-9b51-af7dfbf81379 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef8d6850-20a7-4319-ab31-b52b2775110e · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1677e10a-078b-4e21-a9f7-49a49750125d · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Cheap and quick: Efficient vision- language instruction tuning for large language models.Ad- vances in Neural Information Processing Systems, 36, 2024
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ca9547b3-7859-44f6-81e2-abd7d2890ecc · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Ok-vqa: A visual question answering benchmark requiring external knowledge
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dd0db20d-f593-4607-8163-e0c6fdd3d277 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion ChartQA: A benchmark for question answer- ing about charts with visual and logical reasoning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 784ff458-d379-46e3-9ff8-67e5b2e9d52d · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Docvqa: A dataset for vqa on document images
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d21ae786-c348-4a0a-a84f-33f36383efbd · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Docvqa: A dataset for vqa on document images
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 06388035-1aea-42a3-be92-4b1475ccde45 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Infographicvqa
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4776cc86-5905-4a10-a79d-db2de8dde0bb · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Introducing llama 3.1: Our most capable models to date, 2024
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 66ace162-e509-4286-86b4-c8dd4db39ea8 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Ocr-vqa: Visual question answering by reading text in images
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a4a480c5-8733-4000-9d67-3dfc28e2529d · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion GPT-4 Technical Report
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7056f086-e554-422e-a905-bbf7d132200e · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Gpt-4 technical report.ArXiv, 2303:08774,
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5cae0bc2-df35-4cb3-a429-ba3d5a4969ae · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Im2text: Describing images using 1 million captioned pho- tographs.Advances in neural information processing sys- tems, 24, 2011
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 54bf0e9f-0d87-4f8f-b08c-d3b7863163e0 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cd0bd085-48de-4114-8338-4ecd557293d4 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Connecting vision and lan- guage with localized narratives
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3b8d652-7b42-461e-b27b-47389529bc39 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Connecting vision and lan- guage with localized narratives, 2020
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8d661666-9ccd-4410-9e1f-88ff7c3cc442 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Learning Transferable Visual Models from Natural Language Supervi- sion
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 91df7c65-2a47-41ee-8a80-046a9b9afc38 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Sam 2: Segment anything in images and videos,
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef1219ee-359e-4e0b-b56a-f987b3b47ed1 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Laion-5b: An open large-scale dataset for training next generation image-text models, 2022
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d9436920-017a-44a7-afda-5d1af3ec8759 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion LAION-5b: An open large-scale dataset for training next generation image-text models
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 90b898b8-7c71-46e2-95cb-1566aef7fb9e · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion A-okvqa: A benchmark for visual question answering using world knowl- edge, 2022
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3356d403-ad9b-4824-8be7-6416ca85b22c · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 004772e0-2b2d-4d52-b3ea-10f23bce2144 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d92235de-80c8-4b4f-bef1-c0820066dc7c · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Textcaps: a dataset for image captioning with reading comprehension, 2020
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a49b2bf3-6aba-4f4c-9b1c-21bc58b08f63 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Towards vqa models that can read
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fa194d73-e61a-437c-90c9-b34f3b6187a6 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Expressing visual relationships via language,
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7869d6d7-f163-4c65-bb7d-4966d4508cb6 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Vi- sualmrc: Machine reading comprehension on document im- ages, 2021
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a922125a-d40a-42cd-88e0-b8e911720714 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Gemma 2: Improving Open Language Models at a Practical Size
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c42d8e9-8075-4a81-9d7f-c781c5e161b8 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Gemini: A Family of Highly Capable Multimodal Models
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38bf7cf7-0be6-46ee-b247-97608eee72ae · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 380b9a90-ab99-4a3d-8c4d-52846bee1ac0 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b1342ee-9a8c-49d0-a595-5350d486a591 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Caption Anything: Interactive Image Description with Diverse Multimodal Controls
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 198dbafe-e916-4c44-9197-d8817203973c · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Visionllm: Large language model is also an open- ended decoder for vision-centric tasks.Advances in Neural Information Processing Systems, 36, 2024
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5b49823c-2299-4c5a-b696-a4f930bba88f · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4595a73-fa6a-4f5a-9acd-dace869b8128 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49d01f9d-7713-4b5f-be24-b4ddca4e6e66 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Grok-1.5 vision preview, 2024
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7b2e6457-62fd-42fb-a0e8-3d9c75c8c7c7 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6942175e-fc92-4847-a539-083a5e111a60 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Depth any- thing v2, 2024
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3f2b3ba-ba36-44b4-85d1-72b640d28a8d · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e25ae2a-754a-4ccc-9765-bee5c3134768 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b4c1e2a-b493-4099-9e37-d7226f266c96 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Berg, and Tamara L
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation adab7efc-a886-42f9-bb80-f7790886c47f · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ab4b3daa-e6ad-49dc-927d-d3c5aa0b623d · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Unresolved cited work
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b03a1f4f-3d0b-4a85-9256-5af7f73046ce · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Rsvg: Exploring data and models for visual grounding on remote sensing data
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a6a07919-ca3a-4726-933d-7a98911f55b3 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Lmms- eval: Reality check on the evaluation of large multimodal models, 2024
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 15480673-59f4-46b2-b0ea-d123ad922208 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2f15a56-beea-495f-9df7-5032803d62a4 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1253bd24-febe-40c5-a414-cd03f10b8c9b · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Hierarchical and robust convolutional neural network for very high-resolution remote sensing object detection
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a346cf23-6cad-4a8b-867c-efcd33e76a44 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05c6122e-8d8d-44a8-bd5b-e427e1374c22 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion SVIT: Scaling up Visual Instruction Tuning
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f40fa1de-f74d-43db-8c29-353c44361fce · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion ChatBridge: Bridging Modalities with Large Language Model as a Language Catalyst
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6619e153-8a2f-46be-a2bd-06e10ce65ca6 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5dbdd01-7047-4333-b9db-7191c179e2e4 · outbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Visual7w: Grounded question answering in images, 2016
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.