Pith. sign in

Paper Citation Record · LEDGER

M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2306.04387.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.04387 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 40 of 40 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:01:14.250687Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T20:28:39.249603Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c15d0f86-cbc4-4ce8-95a1-a03c5f1703d5 · inbound

Otter: A Multi-Modal Model with In-Context Instruction Tuning cites this paper.

Otter: A Multi-Modal Model with In-Context Instruction Tuning M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:43:47.865942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T02:43:47.775691Z digest=sha256:94bc5a26b285f61efd5e2dc2fa74fbd5226265ab2b9d9735b83e04d08c2abc36

Observation d1fd2764-d939-4d45-b502-d081c97b8057 · inbound

Large Language Models are not Fair Evaluators cites this paper.

Large Language Models are not Fair Evaluators M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T12:10:42.431540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-17T12:10:42.248005Z digest=sha256:356fcfa8bf06fd8d521eac647653e674f48c195d4fb78a7ce5014382045eb408

Observation d6c9abcb-4a2e-41bc-879a-535ab6a83d07 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:42.695744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:3a78743013a84c6dbb8a5cfbdb09c1b37883a9cdd20abd022a2ad0c815505070

Observation d7316a06-7a45-4489-8abe-44752f698298 · inbound

A Comprehensive Overview of Large Language Models cites this paper.

A Comprehensive Overview of Large Language Models M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 282

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:28:39.251990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T20:28:38.900026Z digest=sha256:7531de40bf891016e5ed3870abad4a8d0b553dbc727db6e646461da931a318ce

Observation 149dc577-de54-4a50-a79e-d9dd3eb604de · inbound

Vision-Language Foundation Models as Effective Robot Imitators cites this paper.

Vision-Language Foundation Models as Effective Robot Imitators M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:44:27.666644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T21:44:27.562453Z digest=sha256:8b31a8516b4f8c420e76a8b7262621fae66dd493cd2d0725018237acc2e7cd78

Observation aacdd64c-0e17-4b4c-91b4-3f6d8bd4b716 · inbound

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI cites this paper.

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:37:41.612505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T05:37:41.401736Z digest=sha256:94cd2c8868b8cbb29aeb34d1d2f519e27b1c9553a43e0506eb07b70bfcd86d8c

Observation ffad8371-cabe-4fd4-b398-09404bcf7c7a · inbound

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark cites this paper.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.115909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:d01549fdc608a259c45dc87a2f9751d35a4e9975418bdd33d3672b22b1a67c10

Observation 2e187d83-ad34-446e-8e3e-0793fa8f2fa6 · inbound

Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations cites this paper.

Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:34:15.763258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-14T22:34:15.638114Z digest=sha256:e0b562c90098d0e1c77a226cd30e3b1a13ac7684f83cd44836b1e05a356ff7f9

Observation a9339ec0-f478-4a61-b0ed-a0e70d112f38 · inbound

Aligning Modalities in Vision Large Language Models via Preference Fine-tuning cites this paper.

Aligning Modalities in Vision Large Language Models via Preference Fine-tuning M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 162

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:58:53.407989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-17T10:58:53.215887Z digest=sha256:779caaff3d20c331308139578306aade864d59733d75e1025aa45175ff2ae7e6

Observation 0d84ef2d-5fec-424d-922f-c386ee15a067 · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.478013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:bace1a14a199fd071aca3deda598183552ceb4dcd2c97edc0b08d5ef6b4f107e

Observation c4f239d6-ed69-4c88-8393-4bc7bf26f84b · inbound

VaLiD: Mitigating the Hallucination of Large Vision Language Models by Visual Layer Fusion Contrastive Decoding cites this paper.

VaLiD: Mitigating the Hallucination of Large Vision Language Models by Visual Layer Fusion Contrastive Decoding M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T13:54:22.688860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:54:22.688860Z digest=sha256:6136b69157bc02eb9c53dea2ad40e79e1e3d88be86e16df1cd75d7753432ee79

Observation 4a1af282-6a57-49b8-a509-edcaf4ccc9fb · inbound

VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models cites this paper.

VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T12:12:30.611856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:12:30.611856Z digest=sha256:230c3a606819e08099c83e8e74058bfe335d8db39ada69806944149da0d26742

Observation f2837536-35ff-452f-816f-8af47781f5a6 · inbound

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation cites this paper.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.311788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.311788Z digest=sha256:a6bbd2c59469ddfc9d53e502cfc6a50b95e1aa2e1257ff36e51d4771824e92c2

Observation 1aed4ae9-c534-45fd-88bb-1c390699e2ef · inbound

LinVT: Empower Your Image-level Large Language Model to Understand Videos cites this paper.

LinVT: Empower Your Image-level Large Language Model to Understand Videos M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.179148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.179148Z digest=sha256:6be836a9df59058dd75d579d32d6bf871d281059342097b05efe1339920e37ac

Observation 1559abe1-c76a-49ce-9955-1af00bb34dcc · inbound

Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor cites this paper.

Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T20:24:58.550719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:24:58.550719Z digest=sha256:20080a75f5af98f7983c81b5253aa9d9673f3c1bf107c7766013abc5ee994016

Observation 0f706992-c754-4cc4-8ed3-cfa581e8a94e · inbound

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation cites this paper.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.743967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.743967Z digest=sha256:eef27dbabc706c104cfcc58e6b311c5e6e65618d2c18ecc12bf3913cf25053c6

Observation bf7444fb-3f1a-43e5-aceb-aac927384bea · inbound

Small Language Model as Data Prospector for Large Language Model cites this paper.

Small Language Model as Data Prospector for Large Language Model M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T16:33:10.459872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:33:10.459872Z digest=sha256:cb6b8075f08024212c3f0bf4041372e37c4fe32fe4d216604d0c548a6984cc09

Observation 12952629-216e-4600-8d99-dee8bdbf7ebd · inbound

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models cites this paper.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.106265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.106265Z digest=sha256:b2f701792f12fc663f2e282c3acb9b4bfd0cd5a8f109dcba43ebed3ad931d0a9

Observation 2db8fa1d-12b9-43c0-b146-e254233ba213 · inbound

Mitigating Adversarial Attacks in LLMs through Defensive Suffix Generation cites this paper.

Mitigating Adversarial Attacks in LLMs through Defensive Suffix Generation M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T12:57:06.165383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:57:06.165383Z digest=sha256:34a28971c29482781d51442d2489cd374673f998893b5e6e55554e6fbb6bd5fb

Observation 200bfab1-34d8-48de-b272-bb2615aead6c · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 237

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:02.176186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:02.176186Z digest=sha256:5da62e332adf4b8ac5b514305dc64d3adb6b3919154eafa4cdd18da364892a27

Observation 7dd9a7a5-e192-4923-9288-c98bfaceeb94 · inbound

MM-GEN: Enhancing Task Performance Through Targeted Multimodal Data Curation cites this paper.

MM-GEN: Enhancing Task Performance Through Targeted Multimodal Data Curation M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T21:45:45.940806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:45:45.940806Z digest=sha256:beda5a322603bf379d56101000f2ca356dfbf55b196023c2b440cbe21952f725

Observation ac1e12be-fce0-4996-a55a-dcd92156cbe6 · inbound

LLMs can see and hear without any training cites this paper.

LLMs can see and hear without any training M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:08.118465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:08.118465Z digest=sha256:d3a99309f744b0b484add1dadf6b91ef4bd03b095b364601a9fe3ec44c8a6861

Observation 9ad6a6c3-1526-49b5-ad6a-a51aaf86e13f · inbound

REASSEMBLE: A Multimodal Dataset for Contact-rich Robotic Assembly and Disassembly cites this paper.

REASSEMBLE: A Multimodal Dataset for Contact-rich Robotic Assembly and Disassembly M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T20:22:55.949459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:22:55.949459Z digest=sha256:c4bee999f7341dbfb73fbe22ce2e8deb4cde76152b32d27fe50a9b33bd8c59cd

Observation 1c3a06e3-66d0-465f-a8e0-54b4e09ee35c · inbound

From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs cites this paper.

From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T22:43:59.423255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:43:59.423255Z digest=sha256:bc6deb240989f08559a998b741b467e3e709107f3fe755a43e816ace03a7ef6f

Observation c8440316-70c4-4b2c-aeea-5864f1099fa8 · inbound

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types cites this paper.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.629857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.629857Z digest=sha256:a10f9a15c355a560bbde4c4352f456a220f7d788c63f8db2a13871d9e47c3a3d

Observation bc63e29f-4b9d-4d4b-a008-c6dbe6417409 · inbound

Enhancing Multimodal In-Context Learning for Image Classification through Coreset Optimization cites this paper.

Enhancing Multimodal In-Context Learning for Image Classification through Coreset Optimization M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T12:01:14.250687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:01:14.250687Z digest=sha256:21b74a3d80b65d19d016577874580f1452a05253c845ee109a8864f03846f3c7

Observation 5c99ab68-9c42-462c-9338-5a7ac2288f8c · inbound

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward cites this paper.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.862054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.862054Z digest=sha256:c562096103834c65c3c5bd6af776c973fc0c959037da8ee7945c287f002adae9

Observation e2270429-29af-40b3-9791-bcce5754e625 · inbound

A Comprehensive Analysis for Visual Object Hallucination in Large Vision-Language Models cites this paper.

A Comprehensive Analysis for Visual Object Hallucination in Large Vision-Language Models M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T04:09:09.944879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:09:09.944879Z digest=sha256:e25f0d8f3d7920f24b7b32cd55065fa17460bb7bd11be9f4bfba0572e5669906

Observation dcbfbcc7-cc48-4171-a781-de9ea4aec315 · inbound

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput cites this paper.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.519188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.519188Z digest=sha256:5781bc241a6593fdcbaddd290b8c761a7b673ac4f1fd54f8315f00fb48c6215a

Observation 14856309-26db-41ac-9566-3a4537ffc053 · inbound

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion cites this paper.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:49.620424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:49.620424Z digest=sha256:2bf7c899e2f18ae6289538389c48e3e1986e006c62468fb399a7efeb96b04531

Observation 81be836d-7602-494e-8daf-47c9171aa506 · inbound

M$^3$FinMeeting: A Multilingual, Multi-Sector, and Multi-Task Financial Meeting Understanding Evaluation Dataset cites this paper.

M$^3$FinMeeting: A Multilingual, Multi-Sector, and Multi-Task Financial Meeting Understanding Evaluation Dataset M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:16.012128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:25:16.012128Z digest=sha256:eb474a89a382e9d91f7a89978bf60c54a19b9e46f2c2cbd4bf0d7549b3457f02

Observation c26c49ee-07bd-446b-9729-c1ceebb6150b · inbound

VIP: Visual Information Protection through Adversarial Attacks on Vision-Language Models cites this paper.

VIP: Visual Information Protection through Adversarial Attacks on Vision-Language Models M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:16:29.523026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:16:29.523026Z digest=sha256:50f5dd669887df5edd11abcc2f6bc8a8a3245b0081184ed1de598bd947cf471f

Observation 6b470d80-6eba-4783-a5b4-51126b5557e2 · inbound

Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security cites this paper.

Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T12:09:39.711544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:09:39.711544Z digest=sha256:68d46ecf65552ff007153c84a530a441414f23efb7aad6c75e30aaa899671fa4

Observation 4d0bd7e3-10c6-4259-acb1-704fa9e30622 · inbound

A Survey on Video Temporal Grounding with Multimodal Large Language Model cites this paper.

A Survey on Video Temporal Grounding with Multimodal Large Language Model M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.831964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.831964Z digest=sha256:c69d13ab3c3d763ebb5565234299c2373df39c7c4a249a5485f24d5e0befcc93

Observation 07239a44-8713-4a3b-bcb2-c2b3efcf6a0a · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T20:28:42.331519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:28:42.331519Z digest=sha256:f410334451da2d62a2311edcf57aa0f5980d7002212b6df7cd832d85c443f367

Observation fe1cceee-ea0c-4c22-8a3e-da732ff23488 · inbound

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model cites this paper.

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T17:56:53.246219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:56:53.246219Z digest=sha256:fededde5053b018fe3abd2844fc81a5cf2d7cc4d42a73358f837a5e86383ef31

Observation 821da7d8-724c-49c5-a1a6-da96264ca380 · inbound

Language-Specific Layer Matters: Efficient Multilingual Enhancement for Large Vision-Language Models cites this paper.

Language-Specific Layer Matters: Efficient Multilingual Enhancement for Large Vision-Language Models M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T16:30:12.596189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:30:12.596189Z digest=sha256:1d2e00ea166278ea6f536274128e5a87a7bf2e2c2d43f24bae479076755f2d86

Observation b2e03eb9-624f-47a9-9882-fe769dadff86 · inbound

Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning cites this paper.

Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-14T22:18:56.115559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:18:56.115559Z digest=sha256:d1a0eb45f0d3068345dd0df41d43a5bd661cada9163ed3279f1b9ec2609b994d

Observation 196cf01f-10d5-4599-a8c6-801cdd4c3419 · inbound

Vision-Language Foundation Models for Comprehensive Automated Pavement Condition Assessment cites this paper.

Vision-Language Foundation Models for Comprehensive Automated Pavement Condition Assessment M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:41:21.206498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T17:30:39.410040Z digest=sha256:3c0287a8715fb9f6382c2a7ea8c0d910a1a9d16ad0cc9dc2274a6f6357c51664

Observation d6c6ff71-de57-4c21-8899-8ccce1823c8a · inbound

Test-Time Hallucination Control in Large Vision-Language Models cites this paper.

Test-Time Hallucination Control in Large Vision-Language Models M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T14:19:24.641330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:19:24.641330Z digest=sha256:5da16158dd0989875994a7641857a36c176402d4ff337b87d6731ae50764e643