Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T08:23:14.354745Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2607.15299.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T08:23:14.354745Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
22 of 22 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8d6f1a2f-4de4-4484-b97b-8c267e466609 · outbound
MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation Making the V in VQA matter: Elevating the role of image understanding in visual question answering,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f7511da-4b47-4900-856e-b327a92142b8 · outbound
MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation OK- VQA: A visual question answering benchmark requiring external knowl- edge,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7589eeb-f542-4c20-acf5-6306847b39e7 · outbound
MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation A-OKVQA: A benchmark for visual question answering using world knowledge,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 070beb76-55d4-4868-8d81-0e8aaded7ad2 · outbound
MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation GQA: A new dataset for real-world visual reasoning and compositional question answering,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b26718d-569d-431e-a5ec-70b24278b39f · outbound
MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation OCR- VQA: visual question answering by reading text in images,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dda5378-58d8-4dae-b809-4bae903e56da · outbound
MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation Textcaps: A dataset for image captioning with reading comprehension,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bde9422d-61ac-439a-96db-62e1e6eab3ef · outbound
MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation Visual instruction tuning,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed94dae2-9208-426f-8483-af98f01e7dcc · outbound
MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation Sharegpt,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1b76ce2-408b-4eff-a344-2e8c7338d282 · outbound
MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation Visual genome: Connecting language and vision using crowdsourced dense image anno- tations,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02fa0fda-7c41-40b4-b63a-686b99b42b05 · outbound
MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation Generation and comprehension of unambiguous object descriptions,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7f12c14-d9fb-4f90-aa6b-c377820cccc9 · outbound
MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation Refer- itgame: Referring to objects in photographs of natural scenes,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f1f453a-b2ce-486c-b88c-74f4262453e6 · outbound
MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f5d8d4c-677f-4598-94ac-0eb6dd580342 · outbound
MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation MMBench: Is Your Multi-modal Model an All-around Player?
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67d0a9e9-c703-4375-8efc-8d5991930764 · outbound
MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b74e29df-3210-4a02-bd52-f2f67d35b039 · outbound
MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation Vizwiz grand challenge: Answering visual questions from blind people,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33885bec-54c8-416b-a3b1-db428c84249b · outbound
MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation Learn to explain: Multimodal reasoning via thought chains for science question answer- ing,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19242688-f65a-4278-b718-2d4cd6f0e09a · outbound
MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation Lora: Low-rank adaptation of large language models,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5469f98-fd0b-4730-9176-40afac55a812 · outbound
MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation Microsoft COCO: common objects in context,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57656ff4-27ec-42ad-8f40-1ffc53d41149 · outbound
MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation Modeling context in referring expressions,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23af5645-73a3-4504-9f2a-af38d4cbbf08 · outbound
MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3880485-2a00-40ba-b9e0-74d8a9587740 · outbound
MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation Visual spatial reasoning,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c43f9043-16e9-4b09-97ac-6a599b8aa2dd · outbound
MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.