Pith. sign in

Paper Citation Record · LEDGER

MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation

As of 5 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2607.15299.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.15299 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T08:23:14.354745Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8d6f1a2f-4de4-4484-b97b-8c267e466609 · outbound

This paper cites Making the V in VQA matter: Elevating the role of image understanding in visual question answering,.

MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation Making the V in VQA matter: Elevating the role of image understanding in visual question answering,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:12.297079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:12.297079Z digest=sha256:1866cceb8cc8d1ed78cb53c4e4fc1237578c1eae2ba583ce410a80dd1e709bde

Observation 5f7511da-4b47-4900-856e-b327a92142b8 · outbound

This paper cites OK- VQA: A visual question answering benchmark requiring external knowl- edge,.

MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation OK- VQA: A visual question answering benchmark requiring external knowl- edge,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:12.387098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:12.387098Z digest=sha256:a021fa0b3c29f3d7a00fd656ae3316300016c9c691919ff83ea110294c5ebfe0

Observation a7589eeb-f542-4c20-acf5-6306847b39e7 · outbound

This paper cites A-OKVQA: A benchmark for visual question answering using world knowledge,.

MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation A-OKVQA: A benchmark for visual question answering using world knowledge,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:12.539193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:12.539193Z digest=sha256:7e77fc5949528ef41cde3d6caeffbbb5107d05bc4adeeb6cef570b54cee80f26

Observation 070beb76-55d4-4868-8d81-0e8aaded7ad2 · outbound

This paper cites GQA: A new dataset for real-world visual reasoning and compositional question answering,.

MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation GQA: A new dataset for real-world visual reasoning and compositional question answering,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:12.616648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:12.616648Z digest=sha256:ea95497ab319749468e1e379c28e76b40cceb34a992f61936c42082643b25605

Observation 1b26718d-569d-431e-a5ec-70b24278b39f · outbound

This paper cites OCR- VQA: visual question answering by reading text in images,.

MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation OCR- VQA: visual question answering by reading text in images,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:12.669130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:12.669130Z digest=sha256:5dbc803034d702a08e251d4dae8e2e445828bca3f2f5187609166709701729aa

Observation 6dda5378-58d8-4dae-b809-4bae903e56da · outbound

This paper cites Textcaps: A dataset for image captioning with reading comprehension,.

MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation Textcaps: A dataset for image captioning with reading comprehension,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:12.749017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:12.749017Z digest=sha256:f31edb18436bf2f110fbafb45143cb6047d2f551f5f274901a790471ff2d9436

Observation bde9422d-61ac-439a-96db-62e1e6eab3ef · outbound

This paper cites Visual instruction tuning,.

MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation Visual instruction tuning,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:12.833347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:12.833347Z digest=sha256:d90c842779b1352ec905b5c27bf56f05535b356a08508e68133e76c4066a8a8d

Observation ed94dae2-9208-426f-8483-af98f01e7dcc · outbound

This paper cites Sharegpt,.

MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation Sharegpt,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:12.911774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:12.911774Z digest=sha256:acba1fa05deccd2140d0ea23c253ed9f77d09bffe19898e43410b9d69a1cfdbd

Observation b1b76ce2-408b-4eff-a344-2e8c7338d282 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image anno- tations,.

MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation Visual genome: Connecting language and vision using crowdsourced dense image anno- tations,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:12.971059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:12.971059Z digest=sha256:cdc294fa4dd8321d42f846ecbc7576eb083a99ce69774cbc1ac4ce7709dbd026

Observation 02fa0fda-7c41-40b4-b63a-686b99b42b05 · outbound

This paper cites Generation and comprehension of unambiguous object descriptions,.

MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation Generation and comprehension of unambiguous object descriptions,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:13.053830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:13.053830Z digest=sha256:329533ed7e3ebc2b4fa22f841c9178f32eeaa1fe9ebb9e01786e4c675e9c3811

Observation c7f12c14-d9fb-4f90-aa6b-c377820cccc9 · outbound

This paper cites Refer- itgame: Referring to objects in photographs of natural scenes,.

MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation Refer- itgame: Referring to objects in photographs of natural scenes,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:13.092521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:13.092521Z digest=sha256:d0198b8f3dbc1778b24c55d0ab88240a2771f0bf337e2a3bc128fcab5741a04c

Observation 6f1f453a-b2ce-486c-b88c-74f4262453e6 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:13.162780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:13.162780Z digest=sha256:a068b590d9ed07bede2973cba010df16e559d024cd54fc07a72afa0c3e7882a4

Observation 4f5d8d4c-677f-4598-94ac-0eb6dd580342 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation MMBench: Is Your Multi-modal Model an All-around Player?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:13.248396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:13.248396Z digest=sha256:f768392c2c66f4b47824b60c8e07dcdb92476935591ade9630ba3463173e4dd3

Observation 67d0a9e9-c703-4375-8efc-8d5991930764 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:13.395817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:13.395817Z digest=sha256:58a54fbc77ed30fdc252c96320eeb7ca91f4a1c256b4d4a3dd4bb274903f4aed

Observation b74e29df-3210-4a02-bd52-f2f67d35b039 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people,.

MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation Vizwiz grand challenge: Answering visual questions from blind people,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:13.524066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:13.524066Z digest=sha256:1f168e5e3d0f9103d7a9fb3350f2ef6be78348211f637a90e2af2e241c2d2c31

Observation 33885bec-54c8-416b-a3b1-db428c84249b · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answer- ing,.

MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation Learn to explain: Multimodal reasoning via thought chains for science question answer- ing,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:13.604750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:13.604750Z digest=sha256:b7bd6e0713fdc990790c43d13cdb321e480a5896f11cbf602adcfef74b3d115d

Observation 19242688-f65a-4278-b718-2d4cd6f0e09a · outbound

This paper cites Lora: Low-rank adaptation of large language models,.

MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation Lora: Low-rank adaptation of large language models,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:13.704757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:13.704757Z digest=sha256:8299f833400b83446c84cce2560e8ec29b13d576e9690754f981a1ea0089b599

Observation c5469f98-fd0b-4730-9176-40afac55a812 · outbound

This paper cites Microsoft COCO: common objects in context,.

MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation Microsoft COCO: common objects in context,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:13.771956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:13.771956Z digest=sha256:c62d499460de7d535b872c5918cae8dce13cbf550c9e7a9d66ff8a471f2e3c2f

Observation 57656ff4-27ec-42ad-8f40-1ffc53d41149 · outbound

This paper cites Modeling context in referring expressions,.

MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation Modeling context in referring expressions,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:13.916582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:13.916582Z digest=sha256:cd02b5242aeb241b616d27fbc0a49df38a2fff638c380ef313e16a4fc4844725

Observation 23af5645-73a3-4504-9f2a-af38d4cbbf08 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models,.

MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:14.064741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:14.064741Z digest=sha256:b1f12b12e8c41e87b71154a06af039de46f8b00a5c31daec67baf4c1c47fca04

Observation e3880485-2a00-40ba-b9e0-74d8a9587740 · outbound

This paper cites Visual spatial reasoning,.

MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation Visual spatial reasoning,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:14.194756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:14.194756Z digest=sha256:65e5d78e3d089b7e03f6680281620e03ec446f0b9bed1d981793a36f4dada452

Observation c43f9043-16e9-4b09-97ac-6a599b8aa2dd · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:14.354745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:14.354745Z digest=sha256:c5da716c91675e3c3da1ff2192816a99b677596ff9098186fda93e5543807eea

Pith citing papers

No inbound Pith citation observations are available.