Pith. sign in

Paper Citation Record · LEDGER

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping

As of 8 August 2026, this Paper Citation Record lists 100 of 141 outbound references and 0 inbound Pith citation observations for arXiv:2506.02308.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02308 v3

Coverage vector

measured 100 of 141 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:32:13.382009Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 141 outbound references displayed

  • verified exact6
  • verified fuzzy0
  • unresolved94
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4e72a6ff-7a11-48b4-ae88-cd378c6cc080 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736, 2022.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736, 2022

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:04.527558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:04.527558Z digest=sha256:0daeaaa442006976c114d7d134f1085f5cc0e90d4f9423bd70b5a788dab2d10b

Observation 5d9e4a59-fcbb-4154-a3dc-00777f240c32 · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Lawrence Zitnick, and Devi Parikh

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:04.562670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:04.562670Z digest=sha256:6942c50cf27b1e66c24a4dec6a1cfea64abda8f6fc8004e6cf567798be409b6a

Observation 9c4ecb82-6154-4acc-a4fb-6b44987a6151 · outbound

This paper cites Gated multimodal units for information fusion.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Gated multimodal units for information fusion

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:04.700976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:04.700976Z digest=sha256:024effe7537a504f1133a95742aa1ca3e8abd9397fcf1d236c5456576173b26c

Observation 6ae9519e-fc8d-4d7e-8576-94f679fb99ae · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:04.747469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:04.747469Z digest=sha256:4504ce747a2570c0acd42fb6aa44bde5e3c3e638dc53a4aabf6c1f81782a16cc

Observation cc05ce37-688a-45b6-a477-0b980139db4f · outbound

This paper cites TouchStone: Evaluating Vision-Language Models by Language Models.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping TouchStone: Evaluating Vision-Language Models by Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:04.798457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:04.798457Z digest=sha256:9216ed895813cda4dd04d12be34aa686d501995b26aa3f28d7eef552c777773e

Observation 5e354e26-95aa-4a78-a683-4a3af107b05d · outbound

This paper cites Multimodal machine learning: A survey and taxonomy.IEEE transactions on pattern analysis and machine intelligence, 41(2):423–443, 2018.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Multimodal machine learning: A survey and taxonomy.IEEE transactions on pattern analysis and machine intelligence, 41(2):423–443, 2018

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:04.843432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:04.843432Z digest=sha256:78f17b85c8d23c065a9cc77b730f824c4acce62cd6416d47c3dfa0e7c575cd19

Observation 6f29f9f5-bce4-4f82-b8e0-885e41700c97 · outbound

This paper cites Routledge, 2014.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Routledge, 2014

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:04.906604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:04.906604Z digest=sha256:bb166ac98e0b8a7827dc7857153fd951cb53788fdf5768715d7775806306ee18

Observation e89c36fe-d388-433f-bed6-d92df087b41a · outbound

This paper cites Introducing our multimodal models, 2023.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Introducing our multimodal models, 2023

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:04.936237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:04.936237Z digest=sha256:7e5d069f6dc352ff7271249d66a511850106ff01a09ff87705fb6b8ba1ef163f

Observation 6433d950-8cef-4b59-9237-4728e1b49ee0 · outbound

This paper cites Identifying beneficial task relations for multi-task learning in deep neural networks.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Identifying beneficial task relations for multi-task learning in deep neural networks

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:32:18.420739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:32:05.006588Z digest=sha256:3197e7d29cb4346d36e53ef2944b14540d5485e0917edd2993c7f693aa819b53

Observation cd280595-afce-46ff-96b7-565d47a44a47 · outbound

This paper cites VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:05.026455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:05.026455Z digest=sha256:fff0d7ed51d04fc4713dd417dab54d2d4f88a1fef2fbd8802d4f10d442025f09

Observation dcab9994-dcdf-4bb0-b60a-f1da440d338d · outbound

This paper cites Decimer—hand- drawn molecule images dataset.Journal of Cheminformatics, 14(1):1–4, 2022.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Decimer—hand- drawn molecule images dataset.Journal of Cheminformatics, 14(1):1–4, 2022

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:05.029575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:05.029575Z digest=sha256:afdef611bc86e56ff6131e8da0b76905faed05d51ae0d7a0562a3c88203b8a4d

Observation 21781edd-69dc-4733-95b0-e8b71e70acc6 · outbound

This paper cites Language models are few-shot learners.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Language models are few-shot learners

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:05.037177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:05.037177Z digest=sha256:4bba668a7a433506a71f50de211d1bd904fc95278a34c3e2e4e9297dc63e7e9d

Observation ad46f82c-7692-4788-ad1a-1e9f16dd11ad · outbound

This paper cites Multi-modal sarcasm detection in Twitter with hierarchical fusion model.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Multi-modal sarcasm detection in Twitter with hierarchical fusion model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:05.151820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:05.151820Z digest=sha256:efbb98a7e755789d340ab8efa16480a0acb02f8e877f56313de443ab8d84a59f

Observation d774d44a-756b-4322-ae24-040c575f5fd8 · outbound

This paper cites Multitask learning.Machine learning, 28:41–75, 1997.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Multitask learning.Machine learning, 28:41–75, 1997

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:05.213593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:05.213593Z digest=sha256:7fe3c02fb8f76de1811592b4c796b5b4b6c492ce32bc72d674b7bd2f8e563770

Observation 76fe4df9-9b3b-4f6e-907d-f5b74a3c32d4 · outbound

This paper cites Uniter: Universal image-text representation learning.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Uniter: Universal image-text representation learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:05.268085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:05.268085Z digest=sha256:62e1b3aa55f3550a816411e49073b9abf233c3763db9ae6c4fe76036b1fcdb01

Observation 0f224f8d-8a5b-49b5-afee-cc1d273be619 · outbound

This paper cites Remote sensing image scene classification: Benchmark and state of the art.Proceedings of the IEEE, 105(10):1865–1883, 2017.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Remote sensing image scene classification: Benchmark and state of the art.Proceedings of the IEEE, 105(10):1865–1883, 2017

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:05.304510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:05.304510Z digest=sha256:81c9cc20079a34fab1514aa53476c8e15d641717764687be8138142de003a1e5

Observation 32516e51-e7bb-4f50-b8e0-fb3821c725fa · outbound

This paper cites Scaling instruction-finetuned language models.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Scaling instruction-finetuned language models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:05.334394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:05.334394Z digest=sha256:186ec0d40de7da745d5b10cc978614fad6746b1524ee0638074a3aad052c69d0

Observation a342a713-9f91-4ec9-8763-5e592fbf8f3b · outbound

This paper cites Holistic Analysis of Hallucination in GPT-4V(ision): Bias and Interference Challenges.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Holistic Analysis of Hallucination in GPT-4V(ision): Bias and Interference Challenges

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:05.365315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:05.365315Z digest=sha256:aa0722d9adf79c0e1b62b0aec6caf09c38e1072381f927063af4cdd0a35f21c2

Observation 173bcff0-9f5c-4431-bf94-66f0292c5824 · outbound

This paper cites CLIMB: Data Foundations for Large Scale Multimodal Clinical Foundation Models.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping CLIMB: Data Foundations for Large Scale Multimodal Clinical Foundation Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:05.400551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:05.400551Z digest=sha256:fbed6f20d26d05b4d27328c1cfc855c140faa7d8dc44a02c9dc546df9e0f2f05

Observation 9902121c-7620-4247-a60c-fac86ab9ee71 · outbound

This paper cites Rico: A mobile app dataset for building data-driven design applications.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Rico: A mobile app dataset for building data-driven design applications

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:05.480514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:05.480514Z digest=sha256:1d0098c8d9eea67a9ec9539400f93e76b6a0b198e4cd18e507cab0b1e28b3a8b

Observation 1fa4f4d7-5be0-4a03-9d1c-55e166867a6e · outbound

This paper cites Multi-task learning for contextual bandits.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Multi-task learning for contextual bandits

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:05.627088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:05.627088Z digest=sha256:287aeeb92c14bd8229e33c39f92444155e53d6297a65b0e642b3f34396e6390e

Observation e9e143eb-ebcd-49c2-99d8-933fd35ba783 · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:05.719625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:05.719625Z digest=sha256:d287ad852897ae14e82ec640bbd33937cec88ce12ab26c87ba9ccd968163b21e

Observation a7a989c2-da13-49c6-a0bb-e1c9fd05f7b3 · outbound

This paper cites Regularized multi–task learning.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Regularized multi–task learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:05.877074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:05.877074Z digest=sha256:c24b2a1a27f00d67d7b72c7ebdead3aca316675902c96094bd82716173dd1c14

Observation cc85cd95-1328-4e69-a9d1-3e194e70c192 · outbound

This paper cites A survey of current datasets for vision and language research.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping A survey of current datasets for vision and language research

Reference 24

Resolution
verified exact
doi, observed 2026-08-07T11:32:16.878332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:32:05.967677Z digest=sha256:1d54bafe5e08527b3a277f63c3226732c40df74656e5a3841a6696ec52a9c59c

Observation 768b9948-d9bd-4d19-8276-08797925237a · outbound

This paper cites Efficiently identifying task groupings for multi-task learning.Advances in Neural Information Processing Systems, 34:27503–27516, 2021.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Efficiently identifying task groupings for multi-task learning.Advances in Neural Information Processing Systems, 34:27503–27516, 2021

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:06.122368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:06.122368Z digest=sha256:3c82f3ab2edd5000e82adb421892e045a57af821bf9a4a316fd8ce7149abd478

Observation ddc20d4c-2299-45cf-ae63-000c5ded65ad · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:06.278770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:06.278770Z digest=sha256:ceb38ac6af64a2f9956ecd620323ec4ff0d230bee66c09af5ee03f54b9840273

Observation 290ba769-b032-4440-9962-5770afc355c6 · outbound

This paper cites Enhancing knowledge transfer for task incremental learning with data-free subnetwork.Advances in Neural Information Processing Systems, 36: 68471–68484, 2023.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Enhancing knowledge transfer for task incremental learning with data-free subnetwork.Advances in Neural Information Processing Systems, 36: 68471–68484, 2023

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:06.390507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:06.390507Z digest=sha256:d1e8923d36a54c108cc4a16339d2721edf868372b948fd2c55805e0ebc267be9

Observation 0af10fe6-21dd-41c2-9bbd-98163be95788 · outbound

This paper cites What makes the difference? an empirical comparison of fusion strategies for multimodal language analysis.Information Fusion, 66: 184–197, 2021.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping What makes the difference? an empirical comparison of fusion strategies for multimodal language analysis.Information Fusion, 66: 184–197, 2021

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:06.461953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:06.461953Z digest=sha256:bb4db6cf55b8da2778704fbbe317c9ab8cf424f5849c2aab52037e42ca98f678

Observation 1fdc778f-7ce6-47a4-a7ae-e42df988440c · outbound

This paper cites Challenges in representation learning: A report on three machine learning contests.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Challenges in representation learning: A report on three machine learning contests

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:06.595058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:06.595058Z digest=sha256:44929c2e4e142fbf6ffb6aa9d63e70de390eb5707ad316b0f11b534fb0465b26

Observation d046d532-f878-482a-ba56-fd7bf699edb0 · outbound

This paper cites Mixture of Cluster-conditional LoRA Experts for Vision-language Instruction Tuning.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Mixture of Cluster-conditional LoRA Experts for Vision-language Instruction Tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:06.772655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:06.772655Z digest=sha256:b2f92326c1281626c49e26d3fe6b530975c6f159e2842f71a8d3ef0a9837d1d2

Observation a3f12151-0bfb-48b8-9b09-a99a9d9051e3 · outbound

This paper cites The Llama 3 Herd of Models.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping The Llama 3 Herd of Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:06.924038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:06.924038Z digest=sha256:8c8f441d4b0d1c598671a024b3113961225badf636ca6bb1c9f0218461a9f441

Observation bbf73a16-98ee-4cde-b53b-9078408175a8 · outbound

This paper cites FuseMoE: Mixture-of-Experts Transformers for Fleximodal Fusion.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping FuseMoE: Mixture-of-Experts Transformers for Fleximodal Fusion

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:07.087252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:07.087252Z digest=sha256:b861e5aaf884632bdd71cd55d071634caec263dccf7e9fe2cfd332a265bede9b

Observation 02835773-a720-4aa7-aa41-6c85c98c5c77 · outbound

This paper cites The movielens datasets: History and context.Acm transactions on interactive intelligent systems (tiis), 5(4):1–19, 2015.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping The movielens datasets: History and context.Acm transactions on interactive intelligent systems (tiis), 5(4):1–19, 2015

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:07.200221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:07.200221Z digest=sha256:f542989660cbf793d251481f699c667bbdbb58923877f92b3b85bc94c1893b5d

Observation fe890e94-9343-4d1e-a87d-64e1f16ff1e2 · outbound

This paper cites PathVQA: 30000+ Questions for Medical Visual Question Answering.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping PathVQA: 30000+ Questions for Medical Visual Question Answering

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:07.313230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:07.313230Z digest=sha256:2afe4096d05124932be1018894422090aab370065cb1d7d57d1cd73c3df79df6

Observation 5d8ce1ed-9b3e-4a1a-b565-7b8c9039a1ce · outbound

This paper cites Do Androids Laugh at Electric Sheep? Humor "Understanding" Benchmarks from The New Yorker Caption Contest.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Do Androids Laugh at Electric Sheep? Humor "Understanding" Benchmarks from The New Yorker Caption Contest

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:07.406001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:07.406001Z digest=sha256:03ed8578f93fd67c86c88f8a7ac53214212a6bf80e9d5beea976e21877e6cdb1

Observation f73ceaf9-c10b-4494-a4e0-5da38ec49ce0 · outbound

This paper cites Framing image description as a ranking task: Data, models and evaluation metrics.Journal of Artificial Intelligence Research, 47:853–899, 2013.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Framing image description as a ranking task: Data, models and evaluation metrics.Journal of Artificial Intelligence Research, 47:853–899, 2013

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:07.530750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:07.530750Z digest=sha256:6d5613a0cb3c5a95a074024449822142bf9149c8515817160ce48c94c516a408

Observation 9107c440-7526-485b-83cf-399bf2592576 · outbound

This paper cites Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:07.603688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:07.603688Z digest=sha256:3c8b2d13d7aeaa797ffb2659f226f07edcac1822192aa02427fc7a0201a627cd

Observation 79624a44-ee99-45fe-8a2f-a6ccd07ffcd1 · outbound

This paper cites Language is not all you need: Aligning perception with language models.Advances in Neural Information Processing Systems, 36, 2024.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Language is not all you need: Aligning perception with language models.Advances in Neural Information Processing Systems, 36, 2024

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:07.706753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:07.706753Z digest=sha256:6b0aa671673c2c982145ee83dc59eadb7dcf43c5ae4055b49780bdcb8b1685e9

Observation e07bd960-e60c-4927-b867-ed39c1f51aca · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:07.801704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:07.801704Z digest=sha256:cb452ec7add1c979ddd874d1d8e3c148f0974378815c001bd95ed865781fb165

Observation 9d117f85-ae8e-4a0e-ba64-07e09f48724c · outbound

This paper cites MemeCap: A Dataset for Captioning and Interpreting Memes.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping MemeCap: A Dataset for Captioning and Interpreting Memes

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:07.886072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:07.886072Z digest=sha256:1fd52664edee1e47dba67166f11faf1b5ccd3a93994bc5e918a3927d0eb167e5

Observation 3e07980a-456e-42a0-9235-cebecffaea29 · outbound

This paper cites Grounding, meaning and foundation models: Adventures in multimodal machine learning.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Grounding, meaning and foundation models: Adventures in multimodal machine learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:07.962103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:07.962103Z digest=sha256:1f7451add38db101ee674a2a3105c1c87a797a395ce3f6ca4cb7d3d7ddbd42f7

Observation 5b6b26b7-74c5-4438-b0ad-ca60c5e80416 · outbound

This paper cites The hateful memes challenge: Detecting hate speech in multimodal memes.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping The hateful memes challenge: Detecting hate speech in multimodal memes

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:08.057119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:08.057119Z digest=sha256:5a0e47be5e42a27360cd8be5112a10d7375702fc745056dff514cdd7c148f345

Observation 6ba146e0-c0e7-482f-9c0d-e1241aa6cfcc · outbound

This paper cites Grounding language models to images for multimodal inputs and outputs.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Grounding language models to images for multimodal inputs and outputs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:08.134800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:08.134800Z digest=sha256:40c8b2600f2c075011724f5af3e44780eb21779332b6272f0de462b6977c0d5b

Observation ebbdc4c9-9f27-46cd-8635-f465165eca7f · outbound

This paper cites Integrating text and image: Determining multimodal document intent in instagram posts.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Integrating text and image: Determining multimodal document intent in instagram posts

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:08.213341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:08.213341Z digest=sha256:b4de40e05681fe107257cde0f69b59cb4c3e4b4d4a6eb2a1774968bcc7356c0e

Observation 81367ec7-4bf7-4d58-8c13-4d395ffa3b38 · outbound

This paper cites A dataset of clinically generated visual questions and answers about radiology images.Scientific data, 5(1):1–10, 2018.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping A dataset of clinically generated visual questions and answers about radiology images.Scientific data, 5(1):1–10, 2018

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:08.289284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:08.289284Z digest=sha256:b2e56e63ad36876e79b7b823bd791e5ed4eda02f04c87ef65febe2d9bad0dd7c

Observation 7bf7604b-2f7c-40fc-a505-bc6a1c7de0f0 · outbound

This paper cites Visual question answering in radiology (vqa-rad), Feb 2019.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Visual question answering in radiology (vqa-rad), Feb 2019

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:08.393758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:08.393758Z digest=sha256:582219b27e18163f219e3eb44936afc1ecdc0fdb53496f398fbdcb007471089d

Observation 0e1cea30-994e-4a38-829b-4aa36efec105 · outbound

This paper cites Instruction matters: A simple yet effective task selection for optimized instruction tuning of specific tasks.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Instruction matters: A simple yet effective task selection for optimized instruction tuning of specific tasks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:08.494742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:08.494742Z digest=sha256:8fd5e8cfc8c420a5b12861bad39230e3f23c343abad5b64095135c73cd0cf14a

Observation 12000cb8-bd3f-4ebb-bdd0-71b7a6c84ed9 · outbound

This paper cites Instruction Matters: A Simple yet Effective Task Selection for Optimized Instruction Tuning of Specific Tasks.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Instruction Matters: A Simple yet Effective Task Selection for Optimized Instruction Tuning of Specific Tasks

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:08.590738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:08.590738Z digest=sha256:6bf1af850d4b0e946f5553d825a759f90fc6bb4e6842c5a87b1415e06eea369c

Observation 3339b701-7b54-4569-b8b0-1a67bf608d74 · outbound

This paper cites Holistic Evaluation of Text-To-Image Models.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Holistic Evaluation of Text-To-Image Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:08.663980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:08.663980Z digest=sha256:bda9692a4984ab86aececbffd1d52af4525f5f3d448bb50d186e377cdb7a3bac

Observation c37fffa3-e70f-4ba6-97e2-2061b0cf5c92 · outbound

This paper cites Enrico: A dataset for topic modeling of mobile ui designs.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Enrico: A dataset for topic modeling of mobile ui designs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:08.716998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:08.716998Z digest=sha256:f9853343858bd9596dad3c781b2913d1188d0f2107ac191575a8bf4229e219b2

Observation 90f00e84-b636-4dd1-947c-a54aa87792ed · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:08.822173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:08.822173Z digest=sha256:c0fbff21fde0ea9321f719cdbe52eb9b3c3634c64a7d2fa4ea24d49bf9688c27

Observation 8a5b6c77-ff9a-45c1-a1cf-45f3239321ec · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:08.910355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:08.910355Z digest=sha256:bcd306483c232ba59a3deb519391afad3d41cf2b36455d07d5ad4f8ee25c559a

Observation e19ecf89-d1a1-4719-b260-ce494f80e05b · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:09.002796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:09.002796Z digest=sha256:4f8a2f815c1a1e9c0c83ceeda8ffb165953ac6224384697577613d7858aaf6ef

Observation 8675a51e-c40c-4633-b0ee-8bd0ec2ec42a · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:09.111934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:09.111934Z digest=sha256:93ff1e0a4a10442bcdc0957e4bd6ee48a89f97446aa981ffcd0f6751567b80d5

Observation a2fa1d85-5890-4ecd-9287-59c6b632f8ef · outbound

This paper cites Identifying Task Groupings for Multi-Task Learning Using Pointwise V-Usable Information.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Identifying Task Groupings for Multi-Task Learning Using Pointwise V-Usable Information

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:32:18.162639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:32:09.209151Z digest=sha256:58053e981d32209c868dd02c0d7f7b344fb736bfa5d36cbc2a3538bbe789fcc8

Observation 7251f186-dc2b-4614-a44f-c7f06692135f · outbound

This paper cites ReForm-Eval: Evaluating Large Vision Language Models via Unified Re-Formulation of Task-Oriented Benchmarks.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping ReForm-Eval: Evaluating Large Vision Language Models via Unified Re-Formulation of Task-Oriented Benchmarks

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:32:17.935739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:32:09.315467Z digest=sha256:f0751ffaabf97c2d6d217d454f67ec3ac0e1bf693291c5e428263909b2cbdb70

Observation a3406c66-9c72-4b2f-8325-7026a08654c5 · outbound

This paper cites LoRASculpt: Sculpting LoRA for Harmonizing General and Specialized Knowledge in Multimodal Large Language Models.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping LoRASculpt: Sculpting LoRA for Harmonizing General and Specialized Knowledge in Multimodal Large Language Models

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:32:17.755225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:32:09.418694Z digest=sha256:37b28e4f742b072b4832fb92ee1bb50b5680b1bd4f6ce1cdda32861884aefab6

Observation 72297dd7-f86d-46e7-bd84-320cc9fdadf0 · outbound

This paper cites Multibench: Multiscale benchmarks for multimodal representation learning.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Multibench: Multiscale benchmarks for multimodal representation learning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:09.534609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:09.534609Z digest=sha256:8f58d92dbe2f543dd2821b49ff2d46deddc71489118808bb87b534e02b58ae66

Observation 70d246e9-cf60-4749-819d-90584ff1b2f8 · outbound

This paper cites High-Modality Multimodal Transformer: Quantifying Modality & Interaction Heterogeneity for High-Modality Representation Learning.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping High-Modality Multimodal Transformer: Quantifying Modality & Interaction Heterogeneity for High-Modality Representation Learning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:09.629828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:09.629828Z digest=sha256:868419f4fc5858ce39eb80963998604d323cd3c6191ca408da80b08334aed899

Observation 9f41e533-ab7a-411b-ab1c-e99c48761650 · outbound

This paper cites Quantifying & modeling multimodal interac- tions: An information decomposition framework.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Quantifying & modeling multimodal interac- tions: An information decomposition framework

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:09.749214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:09.749214Z digest=sha256:4feb0240fd2e639d38835c7e67e0905bb215c3f1372c2005388e171b4074b101

Observation 6b22747e-c6cf-4576-8539-59fe22518dc3 · outbound

This paper cites Foundations & trends in multimodal machine learning: Principles, challenges, and open questions.ACM Computing Surveys, 2023.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Foundations & trends in multimodal machine learning: Principles, challenges, and open questions.ACM Computing Surveys, 2023

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:09.887313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:09.887313Z digest=sha256:845465b1b1c02188de48c4ace5b0c71c6160a640d925bfc64a506d5f07c5e606

Observation 6b6ca40b-4c35-436d-9848-aaaf7681f027 · outbound

This paper cites Hemm: Holistic evaluation of multimodal foundation models.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Hemm: Holistic evaluation of multimodal foundation models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:09.983756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:09.983756Z digest=sha256:0f4c7faed26419b701c13e0f570541aef0abd55c33ef7cdb55b949df02cba766

Observation c7a293f6-c050-41ea-937e-61a38f8899be · outbound

This paper cites Microsoft coco: Common objects in context.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Microsoft coco: Common objects in context

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:10.038516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:10.038516Z digest=sha256:fa683aa4500256aa492aeb95976b430b8603cba3604780852af79c8c316563f3

Observation c5a4df77-f579-494e-a5d5-dd4f17f3f023 · outbound

This paper cites Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:10.106641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:10.106641Z digest=sha256:d7080e5958440b088fc99bb4cde9961edf34c0c5f2fa3b27b6abf1762b4a661d

Observation a369bfd1-54e3-47bf-b801-6809145c68af · outbound

This paper cites Visual Instruction Tuning.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Visual Instruction Tuning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:10.199856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:10.199856Z digest=sha256:5ce978777bf35e9c8132ffb23bf52406cc707eceaa1a54bf039d4a545fb90ddd

Observation acb3b8f2-44d4-479b-81ce-0e829631ad69 · outbound

This paper cites What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:10.270829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:10.270829Z digest=sha256:4dce8fcb7edcc5b438648e665bdfdc798337fe7eb533f06c018c47aa1181ff12

Observation 24f00f28-3cf7-493e-b113-c6dd6326232f · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping MMBench: Is Your Multi-modal Model an All-around Player?

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:10.345380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:10.345380Z digest=sha256:34a7e7c1632f63ddde1a7454e3bcb50893e9c629b314b589104dc90ddba51ab8

Observation bc85cb85-975e-4319-8afc-377612f6b484 · outbound

This paper cites NVILA: Efficient Frontier Visual Language Models.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping NVILA: Efficient Frontier Visual Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:10.410916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:10.410916Z digest=sha256:0986e78b5c1a464365477e71f127f9cf14d3989e1641fb0ecdca5e420484c564

Observation 358804ef-e94f-4b19-b305-1c4f584a089b · outbound

This paper cites The flan collection: Designing data and methods for effective instruction tuning.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping The flan collection: Designing data and methods for effective instruction tuning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:10.465143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:10.465143Z digest=sha256:75cab0a3ea51339b514a3fca37ab7e9c7df335158cd90b5e6bcc02c8181b1e56

Observation 6138fea2-d956-455b-88ca-fe92967b3a2d · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.Advances in neural information processing systems, 32, 2019.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.Advances in neural information processing systems, 32, 2019

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:10.596512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:10.596512Z digest=sha256:b2468b97e101983d746e571140bfd622a33a1a0f1bee985f8e5c0be246763101

Observation e22de362-4a5e-46c7-9914-ff50055fd6af · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:10.690316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:10.690316Z digest=sha256:ce3ef214f04a71298808ab8c0543ce963fb1de13d712496a3dded5ce4b976fec

Observation e82fa458-0e23-448a-9146-d01f3efe5fc3 · outbound

This paper cites Modeling task relationships in multi-task learning with multi-gate mixture-of-experts.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Modeling task relationships in multi-task learning with multi-gate mixture-of-experts

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:10.751352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:10.751352Z digest=sha256:73a7c8ab9333ad8c611bac2c2004d81f387c8615c57c3319527d12818b21eb37

Observation e4f8709e-9072-4134-ad22-b450a213f0ee · outbound

This paper cites A taxonomy of relationships between images and text.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping A taxonomy of relationships between images and text

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:10.833909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:10.833909Z digest=sha256:76e98a7299c012cdf0d985f44a3f532e86f47e2d727e2464bc6f1ff92aed757e

Observation 3a885e73-e46e-44de-937d-a8f1211fb382 · outbound

This paper cites Multi-Task Learning as a Bargaining Game.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Multi-Task Learning as a Bargaining Game

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:10.907531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:10.907531Z digest=sha256:4fe43bb1fa5b5fa37c84d2b8593f23ace873662c3b4d88050044929247ff52b8

Observation 38a13155-3f7d-4e6a-9d2a-f4d4255497fd · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:11.004205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:11.004205Z digest=sha256:e7f7dc8b718145dfed4b877c66585a6b96f47092ffa7073b6d7bb3c51a538ac9

Observation 50ca30bf-cb60-4064-8692-675a772afe83 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:11.089611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:11.089611Z digest=sha256:b6952fe613feba38367680c7dee65366cad432739d7ea72d42ddf1da590436b9

Observation 589ed68d-a688-407b-97c8-126f0f7d2be0 · outbound

This paper cites To Tune or Not to Tune? Adapting Pretrained Representations to Diverse Tasks.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping To Tune or Not to Tune? Adapting Pretrained Representations to Diverse Tasks

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:11.173110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:11.173110Z digest=sha256:afc22794994cd8852c75ba32d14c50c2ae51b535330a73b49fa5f925e1e0739b

Observation b16046a3-756e-449e-9dfc-c7732505f577 · outbound

This paper cites Connecting vi- sion and language with localized narratives.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Connecting vi- sion and language with localized narratives

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:11.335880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:11.335880Z digest=sha256:ade8f2afe2600d8f552c55403f3dcc7a9b681571de455c13c5dce88bf0f46982

Observation fbd1ea9b-431c-42ee-8189-12cbe4615956 · outbound

This paper cites Learning transferable visual models from natural language supervision.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Learning transferable visual models from natural language supervision

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:11.434785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:11.434785Z digest=sha256:3a629bfe7bafc066738bbb0c2c1015ee6f5aede766a759488fea3f60753dab2a

Observation 5d51db8a-ffcd-49a5-9f94-83ecae176c6c · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36:53728–53741, 2023.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36:53728–53741, 2023

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:11.524533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:11.524533Z digest=sha256:a462bc80979a81ff663af4bd967fdf6bc70c9e69850ec284fea6df1a785bcc8e

Observation 905d2572-20f2-4d9f-bc59-472f188a60e4 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:11.619517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:11.619517Z digest=sha256:c3fd968c16308a9326cf710e59120272e374b7e06d52d52b9fc9bed7129914ec

Observation 7c8118fb-a7cd-4d4f-b76a-d8fedd103c6c · outbound

This paper cites Multitask Prompted Training Enables Zero-Shot Task Generalization.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Multitask Prompted Training Enables Zero-Shot Task Generalization

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:11.690608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:11.690608Z digest=sha256:cd0be93b2cdf790f10eb5e1c60b8a499f3203708f80c0542498b8660b87c1eb0

Observation 472364db-c53c-4439-b64d-349bfecdc59d · outbound

This paper cites an unresolved cited work.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Unresolved cited work

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:11.777240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:11.777240Z digest=sha256:e3ba20b1715cd25cfdaff061a679e9394d466157728aca4ca625c4e8390a8905

Observation e61748c8-ef1e-46a9-b396-87264baf6307 · outbound

This paper cites MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:11.897837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:11.897837Z digest=sha256:c434533850ca9c89cd651ad8c3220794d7abbe55718b0dcc6bf0334d5af0f6e7

Observation 55dbad53-eca4-4db8-b5f5-5471953cdc26 · outbound

This paper cites Multimodal Instruction Tuning with Conditional Mixture of LoRA.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Multimodal Instruction Tuning with Conditional Mixture of LoRA

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:11.968805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:11.968805Z digest=sha256:880df91898a8badcd73031f677bb7a3cd6c81e7a5b939627f4a28d385e2649b2

Observation 474173e2-3c80-4f77-a82d-8fa0fbb37145 · outbound

This paper cites A Principled Approach for Learning Task Similarity in Multitask Learning.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping A Principled Approach for Learning Task Similarity in Multitask Learning

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:32:17.428702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:32:12.056052Z digest=sha256:19e266bb2cd415e55167e0ce819f3c33be579b9e068aaf29ece814a8e4f5606e

Observation f507c611-7c9e-4302-aa20-d2f80dbf868e · outbound

This paper cites Guibas, Jitendra Malik, and Silvio Savarese.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Guibas, Jitendra Malik, and Silvio Savarese

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:12.174376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:12.174376Z digest=sha256:bff7e5168442d2b3f9ce52eda44e16bd061b044d068b65ff129e3de3097d3aa5

Observation ddd338b8-b98e-4bee-9eb0-11616d61f666 · outbound

This paper cites VL-BERT: Pre-training of Generic Visual-Linguistic Representations.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping VL-BERT: Pre-training of Generic Visual-Linguistic Representations

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:12.247002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:12.247002Z digest=sha256:ce09e30dc25973de21d8a192dafba2aea8f55620eb7befbef6aeb588af02dbf0

Observation 0db03e46-08c1-4b7f-af66-5fd88c72a835 · outbound

This paper cites A corpus of natural language for visual reasoning.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping A corpus of natural language for visual reasoning

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:12.321003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:12.321003Z digest=sha256:d00d22d5e3a4771a8d324b0a9f713411b0da17fc4b4546c2fb80a338ed3b8f52

Observation f4b9e7d4-e50b-4cd9-bb20-3f36bc1f00dd · outbound

This paper cites Multimodal transformer for unaligned multimodal language sequences.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Multimodal transformer for unaligned multimodal language sequences

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:12.460206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:12.460206Z digest=sha256:6d40c1cebafad5708bf06ebea0ec2560b212439d39f48a1fbbd7bc28f67c90e2

Observation 50ce583a-c8f9-4df7-b10a-bb0b97aa119c · outbound

This paper cites Multimodal few-shot learning with frozen language models.Advances in Neural Information Processing Systems, 34: 200–212, 2021.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Multimodal few-shot learning with frozen language models.Advances in Neural Information Processing Systems, 34: 200–212, 2021

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:12.554019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:12.554019Z digest=sha256:ef917be459bfe955520be0d92cbd2acf4d15a196b745355d190486da644da983

Observation c0fd2c97-e468-422d-b884-af96401a9895 · outbound

This paper cites The inaturalist species classification and detection dataset-supplementary material.Reptilia, 32(400):1–3, 2017.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping The inaturalist species classification and detection dataset-supplementary material.Reptilia, 32(400):1–3, 2017

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:12.652313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:12.652313Z digest=sha256:1861b9c2e7406427a417af432e7d6f03bea7f74f7c58af5b7bcd3903a3021cd6

Observation cc27e696-39cc-4348-8753-55e5d6d6395a · outbound

This paper cites GIT: A Generative Image-to-text Transformer for Vision and Language.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping GIT: A Generative Image-to-text Transformer for Vision and Language

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:12.736468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:12.736468Z digest=sha256:8ddc105404c6d91021549bcffb6948f4c652eff7cfd32b5ef5024d96c3611283

Observation a06b68e7-3638-4111-af79-e9f040f8cfa8 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:12.860372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:12.860372Z digest=sha256:00c1ff52dd8ce3db081fd4c58a73ffddd6d4df948fa2f53dbdbe6799ec91108c

Observation c74420d2-fe47-45a7-8127-cbf5d76b1431 · outbound

This paper cites Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:12.947608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:12.947608Z digest=sha256:147e845f45e904e890d7cc3d16c59d479036849a5cd8e9b15fff14181d761299

Observation 5b734aa5-7731-4c40-9e97-9d368c1ba4cd · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Finetuned Language Models Are Zero-Shot Learners

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:13.042859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:13.042859Z digest=sha256:fb0e816b3b17dab11db2f5c3e3e8c37fdc330d33c848e945759bff08f058849d

Observation d905f998-4953-4734-96dc-1096ee352a4a · outbound

This paper cites On the road with gpt-4v(ision): Early explorations of visual-language model on autonomous driving, 2023.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping On the road with gpt-4v(ision): Early explorations of visual-language model on autonomous driving, 2023

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:13.109666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:13.109666Z digest=sha256:e172841f9e7eee8cae6682a674b129ca827fa0b650c26bfaafa9141e53c10c8b

Observation ea17165a-e1cd-4a8a-955a-ef3e918e3154 · outbound

This paper cites Nonnegative Decomposition of Multivariate Information.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Nonnegative Decomposition of Multivariate Information

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:13.263231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:13.263231Z digest=sha256:bfe41129c98f871590fa65df36e3f553c390e41c6d5718b648677a6d20f9d3ce

Observation c984e5dc-6606-476f-be84-b2f351141a27 · outbound

This paper cites Vision-Language Dataset Distillation.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Vision-Language Dataset Distillation

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:13.340080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:13.340080Z digest=sha256:6070830d81839894fcdea82b5f13bb07592894344e3270545440676fe9f87b53

Observation 3d9636fd-d20f-46ff-aef4-26bb12c40bef · outbound

This paper cites LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:13.382009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:13.382009Z digest=sha256:21873e2f5348e721a347cac0288789d42d7310d38eb0afc12662be853a4b1fb6

Pith citing papers

No inbound Pith citation observations are available.