Pith. sign in

Paper Citation Record · LEDGER

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning

As of 7 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 0 inbound Pith citation observations for arXiv:2507.21924.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21924 v1

Coverage vector

measured 89 of 89 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:18:41.518126Z

measured 89 of 89 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

89 of 89 outbound references displayed

  • verified exact1
  • verified fuzzy48
  • unresolved38
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5478395a-a3e2-4910-ab88-b83667be0c36 · outbound

This paper cites Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:31.323457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:31.323457Z digest=sha256:eff40ff89278ca5097d82f4c6d4067d0e5fe3543f86fd31025ce573457cffe79

Observation 0462b592-f0b1-40a1-ab79-724f553592da · outbound

This paper cites Qwen2.5-VL Technical Report.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:31.561550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:31.561550Z digest=sha256:a5a69d43a74ac2ee314a0b83840e7ea95973d3b7b093be111cc63b05dd7107a8

Observation d43985dd-d259-4ece-bddd-04e480c521b6 · outbound

This paper cites Jawahar, Ernest Valveny, and Dimos- thenis Karatzas.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Jawahar, Ernest Valveny, and Dimos- thenis Karatzas

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:31.688929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:31.688929Z digest=sha256:2a18817fe3b3b91bce63f8e3fcb8e758bf91e10c50d227ad15b5bb8e840752f5

Observation 5e75490b-058d-4a84-81f8-52a64d95e4fb · outbound

This paper cites An augmented benchmark dataset for geometric question answering through dual parallel text en- coding.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning An augmented benchmark dataset for geometric question answering through dual parallel text en- coding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:31.821625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:31.821625Z digest=sha256:b70f9ba4ba3f95a2999a4f4b70353f9f604503aceee3dd15b8749776aefb36f5

Observation 3baeeb8c-e4ef-4928-8cff-4e32bae54d62 · outbound

This paper cites FireAct: Toward Language Agent Fine-tuning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning FireAct: Toward Language Agent Fine-tuning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:31.887391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:31.887391Z digest=sha256:8dbd6bbcc47ab0ebc1ab78155895db5aa2326e746d6543424ad087b780ea07bc

Observation 463991f0-cc00-40b0-8798-d05ee1827c01 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Sharegpt4v: Improving large multi-modal models with better captions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:31.962339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:31.962339Z digest=sha256:7edd195861e40c4a8e7427ab8cf9f8cb681080dbf157be7fdc1d01271b8255d4

Observation 6705916d-5dbc-4a8f-8611-f80eeaf2b8f1 · outbound

This paper cites Are we on the right way for evaluating large vision-language models? In The Thirty-eighth Annual Con- ference on Neural Information Processing Systems, 2024.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Are we on the right way for evaluating large vision-language models? In The Thirty-eighth Annual Con- ference on Neural Information Processing Systems, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:32.056475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:32.056475Z digest=sha256:687e277ae84ce980ffc8b87b56df701e262957290182d3464212a4852b710a26

Observation c5d95906-1dc0-47af-9f83-a6d76b8c75b1 · outbound

This paper cites M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:32.202517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:32.202517Z digest=sha256:0c724a486214320ff994a3c7098db318cdfb3e6dbc6896dbaa4ee4c8171903cf

Observation b2b48a63-6335-45e0-b0b6-35b2fc97d4c6 · outbound

This paper cites Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:32.356535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:32.356535Z digest=sha256:3ba830c372179bcd237592033dbbc91886f5ade464e4c4c1306022e64335eaf4

Observation b41b9cc9-cd38-4c58-9845-d34dd9504529 · outbound

This paper cites Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:32.447571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:32.447571Z digest=sha256:baa6f400a8dc8c8efb7687574d69fe398ed8511e4a0d81fa4d45a66a4c81fa44

Observation 92c9f48c-c1af-4cba-885d-a314261a25bf · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:32.551884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:32.551884Z digest=sha256:9d8590bd7e06b1a09fdd0f5fb298fcfda7dbbf72746c1ffed8771cf401b97ce2

Observation 15d9bcc5-8cbf-46d3-96b5-c0b01a6075ea · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:32.676917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:32.676917Z digest=sha256:acd8cfe992f1439e6d1779d54cffbf486ca831fcb0ac54757f56f428e0107f65

Observation a80648dd-7b1d-458c-ab83-446322737701 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:32.810532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:32.810532Z digest=sha256:d8db0f3c6a4871334772aa35c0a764726c84df28d948cdee9f0c36fd33467c2b

Observation 971d3421-3b28-423f-9ecc-977122187998 · outbound

This paper cites Clevr-math: A dataset for compositional language, visual and mathematical reasoning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Clevr-math: A dataset for compositional language, visual and mathematical reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:32.899714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:32.899714Z digest=sha256:ba69c585a216adbcb9aac12cf73dc80b21486333865e99b82ffbe12893466468

Observation 4f5df2a3-68f0-4fb0-9d97-992800e1cd89 · outbound

This paper cites PP-OCR: A Practical Ultra Lightweight OCR System.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning PP-OCR: A Practical Ultra Lightweight OCR System

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:33.133368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:33.133368Z digest=sha256:2b41ba84170e878441c8af23c631c44c9fa6acefdd0f6a4e472f6c53b503f830

Observation b5973dad-5e46-4e7a-9168-241b722a9d96 · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:53.445321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:33.235402Z digest=sha256:7704e987f9f80d2fc74da02b8f6a423aeb92f3d2a75e84097516a0122a6655dc

Observation b764f13b-cc2a-4907-983f-74e1c22f8fe0 · outbound

This paper cites AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:33.353518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:33.353518Z digest=sha256:6f601584b0e7c795c956c25fb762a49ea75650aa6466daf13c08dfe9ffbf259a

Observation 0b4a3181-5f0a-43df-9a38-1b3f554a584c · outbound

This paper cites Multi-modal agent tuning: Building a vlm-driven agent for efficient tool usage.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Multi-modal agent tuning: Building a vlm-driven agent for efficient tool usage

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:53.080042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:33.503296Z digest=sha256:a127636d87b86457eaac94a7d8c294be17b1fc17f5aa0368f6e6bd74d3244e5e

Observation e63cb2f6-1627-44cd-8a77-77aa298bd6b9 · outbound

This paper cites Hallusionbench: an advanced diagnos- tic suite for entangled language hallucination and visual il- lusion in large vision-language models.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Hallusionbench: an advanced diagnos- tic suite for entangled language hallucination and visual il- lusion in large vision-language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:52.871053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:33.635653Z digest=sha256:af2b081803c178c5378f19ede4e80b145717f11f524541dd7c9b3375d8536f7b

Observation 1d71c6f4-9663-4f51-bedf-903cc3f88412 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Lora: Low-rank adaptation of large language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:33.799912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:33.799912Z digest=sha256:4f4dbeead3ae6837e39b7be7dee531c37685b52e4ed88716286203b9a9d9d049

Observation eae346d7-dea3-4869-a0fd-8ff980354b34 · outbound

This paper cites Icdar2019 compe- tition on scanned receipt ocr and information extraction.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Icdar2019 compe- tition on scanned receipt ocr and information extraction

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:52.676694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:33.926347Z digest=sha256:90235d5419676eb410c9496eb91a65221170531f1b1b137578d9f2eadc235ca1

Observation 3877697a-21a2-4942-8ffc-e5c49755e270 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional 9 question answering.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Gqa: A new dataset for real-world visual reasoning and compositional 9 question answering

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:52.425932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:34.070741Z digest=sha256:0ddcc8b17ed7e4949db1d55a5d6a541aabf302d4bf02e29e2b464b79b03675e4

Observation 7624ac31-6ceb-4462-b1f2-92689753961f · outbound

This paper cites GPT-4o System Card.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning GPT-4o System Card

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:34.176052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:34.176052Z digest=sha256:c41ef7670793351f65226af10a16ca43510b81f10e63e7e7d8e516768799dec8

Observation e6c32f4b-80b9-4a92-ba9d-501d5d813541 · outbound

This paper cites Lawrence Zitnick, and Ross Girshick.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Lawrence Zitnick, and Ross Girshick

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:52.105623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:34.324440Z digest=sha256:7a7d4534a841df462595c7e2406d94ed8663a6aca777784315c2debfbfce7204

Observation 3cde0bb2-3704-4693-af26-65dc37e6cc7b · outbound

This paper cites A diagram is worth a dozen images.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning A diagram is worth a dozen images

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:51.795759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:34.473275Z digest=sha256:9964222b037eb6d62ead5898b9f18f4b11c2e74803137a11fda30ff54a027646

Observation f257a9f9-e7c1-4d5c-a8b0-4517f4a3ed53 · outbound

This paper cites Are you smarter than a sixth grader? textbook question answer- ing for multimodal machine comprehension.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Are you smarter than a sixth grader? textbook question answer- ing for multimodal machine comprehension

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:51.269990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:34.736523Z digest=sha256:fccd83a2fcc259f3203e58f4c0cff23686319b2e5d60606fa835c8c08ca7205d

Observation 71b10f2c-bc23-4add-b219-882ce0430e63 · outbound

This paper cites The hateful memes challenge: Detecting hate speech in multimodal memes.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning The hateful memes challenge: Detecting hate speech in multimodal memes

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:50.997805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:34.828859Z digest=sha256:d99b6ccecd483a6c88f8bdf037bdcabbcae7c59aa45aeff26d89183fa1686edf

Observation 126b78f9-f3f7-4151-8c5e-c6987d0b1265 · outbound

This paper cites The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:50.784892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:34.949585Z digest=sha256:87862257f14f5aee3eb5b7f6c0681cf7993cdf593167c1ba8de618fb6e413d6d

Observation cb91d659-f428-42cd-9432-dd29556578f5 · outbound

This paper cites What matters when building vision-language models?.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning What matters when building vision-language models?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:35.087650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:35.087650Z digest=sha256:0a3589c6f938c5e1bc00c29c96bdc50a87de5bcf57f93a3642d4f54299c32c45

Observation 32ee5be9-6523-403c-8df5-f6171b4f14aa · outbound

This paper cites Kankanhalli.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Kankanhalli

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:50.510363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:35.219442Z digest=sha256:63a43954461fa0d94224f3a7b2c93266770c0f72db743c4e8319d6a4b17a93ef

Observation bf4c79f0-08e5-4ae4-9fb8-731fff0b5e2b · outbound

This paper cites Benchmarking Multimodal Retrieval Augmented Generation with Dynamic VQA Dataset and Self-adaptive Planning Agent.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Benchmarking Multimodal Retrieval Augmented Generation with Dynamic VQA Dataset and Self-adaptive Planning Agent

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:35.368071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:35.368071Z digest=sha256:b80a8a9c533b2c38afa7a9aeea39cb538c7f9b456a071f3f5fe9fd9effac3918

Observation e6dbb72a-e8be-4707-98fa-1a65ca8b36d9 · outbound

This paper cites Visual spatial reasoning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Visual spatial reasoning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:50.268677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:35.520950Z digest=sha256:4138ee14ab2c460174221adc6d13e8270690e9522cbd73a20ae169c944dd4e7d

Observation 150b32c7-24de-4cd7-a868-11d75292721e · outbound

This paper cites Visual instruction tuning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Visual instruction tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:35.688521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:35.688521Z digest=sha256:c59cc7af54fee5f0ce2f7d0f04501f03cd0bda561b50493a3af6b8976b868545

Observation c24c4f2e-b304-45f7-a192-054771395861 · outbound

This paper cites Improved baselines with visual instruction tuning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Improved baselines with visual instruction tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:35.794468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:35.794468Z digest=sha256:e342af35c01b5cf35da5cf43e6d1c5c984e26f2c487b241e90b6437075082b6a

Observation da4acb6d-e861-4fe5-bc60-f8c68030414d · outbound

This paper cites Llava-plus: Learning to use tools for creating multi- modal agents.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Llava-plus: Learning to use tools for creating multi- modal agents

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:50.028090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:35.884051Z digest=sha256:f3a6bd17aa7ea79712ac431b640d013a663a7313573d01c6a0791e6e3328d776

Observation 3c4785de-f488-46e6-894f-42253a8b2587 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:49.759709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:36.050970Z digest=sha256:7a4d466ca2aabaaaa6bd1f5f2b02cc34af4f5f43c9c6b66d091c76962b17dce1

Observation 17eb8dc7-fd91-4c74-a714-e93a0aae3cc0 · outbound

This paper cites On the hidden mystery of ocr in large multimodal models.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning On the hidden mystery of ocr in large multimodal models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:49.507460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:36.188259Z digest=sha256:d47c8e0a1b77cb98f04f73cca5f8f7046eef9d4fc652aa828f672d106d566d43

Observation e8520f35-a2c8-4be5-be7d-da8fdc5baeea · outbound

This paper cites Inter-gps: Interpretable geometry problem solving with formal language and sym- bolic reasoning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Inter-gps: Interpretable geometry problem solving with formal language and sym- bolic reasoning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:49.280302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:36.349318Z digest=sha256:d0703db2cc2df165f89ed4ffdef2ac948e8709d89ed949d35d7e8b60fd2bf422

Observation eec69d49-8b0a-48c7-834b-7b86943d7753 · outbound

This paper cites Iconqa: A new benchmark for abstract diagram understand- ing and visual language reasoning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Iconqa: A new benchmark for abstract diagram understand- ing and visual language reasoning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:48.930150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:36.452978Z digest=sha256:830097f11c13ef4dd3f9029675fdfb2d7989842da19dbac3c0a4dd4c3035f15b

Observation ce7ff31f-0d94-4bb9-8493-408472c1e8f2 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:48.664559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:36.570393Z digest=sha256:0bd1c4489b7cf279633f4f874a32a91dd35306b72d4a06715c073ee75ed6bf47

Observation 319aec87-0af3-4f86-bfe8-be9ab2818488 · outbound

This paper cites Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:36.719126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:36.719126Z digest=sha256:e514bda60f66238546ec160e88258ffcf4b182ac6f9ddaabed855de062e3950c

Observation 2ab389bd-043c-404e-8d5f-fce2f8f5d213 · outbound

This paper cites Mathvista: Evaluating mathe- matical reasoning of foundation models in visual contexts.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Mathvista: Evaluating mathe- matical reasoning of foundation models in visual contexts

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:48.351154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:36.825822Z digest=sha256:4a3e245f9a8ae139b816a06b7506bdd99f0767d4c02b84d8a861b123540743e5

Observation 0780e952-e4e3-4756-b3fb-50787697b212 · outbound

This paper cites ChartQA: A benchmark for question answer- ing about charts with visual and logical reasoning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning ChartQA: A benchmark for question answer- ing about charts with visual and logical reasoning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:48.243615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:36.926908Z digest=sha256:aaf3a42e6529587a2ef7faa3ba04c9acbc1f98150be404eff587d58335db3535

Observation 6f07a7d3-977d-4b42-a5e1-42ec6e1e802e · outbound

This paper cites Docvqa: A dataset for vqa on document images.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Docvqa: A dataset for vqa on document images

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:48.168702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:37.040447Z digest=sha256:a6f32fb730ab5c355c2d183e1decea9af28a11b24c1a627651a83dc8dd6dd3ec

Observation 1c296ffe-614f-4b16-b661-1b1a212f80d4 · outbound

This paper cites Infographicvqa.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Infographicvqa

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:48.004564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:37.177685Z digest=sha256:96493352c9682632671e282e01793fcea8eaeccbdc30d37eb776d62cb61c3925

Observation 6ae9bfaa-2a9c-4804-aa2a-dc2b857795ce · outbound

This paper cites Llama 3.2: Revolutionizing edge ai and vision with open, customizable models.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Llama 3.2: Revolutionizing edge ai and vision with open, customizable models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:47.920183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:37.259659Z digest=sha256:fdb356f931483be920a02e9f78cc0bd47553c3e4e3b2c20a14f2afb28964c107

Observation 71d02cb0-f2c3-4195-be9c-e0e90823e35f · outbound

This paper cites Compositional chain-of-thought prompting for large multimodal models.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Compositional chain-of-thought prompting for large multimodal models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:47.797006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:37.377593Z digest=sha256:fcdd89cbc99b9b5caeca2a74db34e5ab4596301b76503fbb48d887f8d26a80f0

Observation 89d54ed3-7d75-46e4-9c19-469c7336160e · outbound

This paper cites Compositional semantic parsing on semi-structured tables.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Compositional semantic parsing on semi-structured tables

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:47.676234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:37.497595Z digest=sha256:32f296abee73138e827987c819c6652f98de145c4b15a8ccdb80917a9a988e69

Observation 53b939d8-6681-442e-976a-2f9777000580 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:47.533383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:37.575202Z digest=sha256:f67d5703aedc887c6d08d0e0483c93653fd3436bf60aa7b955307ee771c2089e

Observation 67ba3997-6b2a-4bf6-96b9-68bb29ab71f5 · outbound

This paper cites A benchmark of facial recognition pipelines and co-usability performances of mod- ules.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning A benchmark of facial recognition pipelines and co-usability performances of mod- ules

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:47.367570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:37.702678Z digest=sha256:5bd09415cb3a516856f8d609296e562245e28c35f6efecd215e6a2bedf10ba9a

Observation 65e6ac54-3195-45f1-bb0a-adf7d0b85a17 · outbound

This paper cites Putting gpt-4o to the sword: A comprehensive evaluation of language, vision, speech, and multimodal proficiency.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Putting gpt-4o to the sword: A comprehensive evaluation of language, vision, speech, and multimodal proficiency

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:47.277223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:37.818888Z digest=sha256:09149745bd6b4098e14f334600e6ee53cb2324fa29ca67cdeacfc1650656f54b

Observation 1494792a-ecc3-4427-8057-b946dd8698eb · outbound

This paper cites Visual cot: Advancing multi-modal language models with a comprehen- sive dataset and benchmark for chain-of-thought reasoning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Visual cot: Advancing multi-modal language models with a comprehen- sive dataset and benchmark for chain-of-thought reasoning

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:47.182352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:37.933670Z digest=sha256:3e8bb73cea960e1afb1f1f9a80a4417e692752db4db4f5b23304a26708b97722

Observation be5ce8d3-57c9-458f-9ab8-cb9c930acd9a · outbound

This paper cites Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:38.079081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:38.079081Z digest=sha256:fb73b3a030cca416f2f39551cf62556512fdb6f94eb4a221ffe7335151bb8ff1

Observation d7963086-d156-4656-8cc0-45f01486735e · outbound

This paper cites Textcaps: a dataset for image caption- ing with reading comprehension.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Textcaps: a dataset for image caption- ing with reading comprehension

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:38.200474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:38.200474Z digest=sha256:22a0283faf6a2e485299d1c7ed9f714be29ead3d5891fde0e81b0916c4960f2d

Observation 031f965e-c81e-47ce-990d-88f3dc516a55 · outbound

This paper cites Towards vqa models that can read.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Towards vqa models that can read

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:46.982901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:38.315186Z digest=sha256:c092ef3b2084caac7cf77645b9d460c84c5cef2917c8e8d8f8e1b193a4a6c209

Observation 957c6707-7fcb-4195-84a1-e72f0b84820a · outbound

This paper cites Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:38.403343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:38.403343Z digest=sha256:5af527b1321b1f0d4cdc504d5c8076d85a7ec7cef34819ec6b363f97c03d31aa

Observation cc460c53-afcf-45cc-bb4e-231079fc7c88 · outbound

This paper cites Tang, Angie Boggust, and Arvind Satyanarayan.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Tang, Angie Boggust, and Arvind Satyanarayan

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:46.756525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:38.500912Z digest=sha256:56bc078b011439afd848d90f870437ad836142feec294315d8bbda473918991f

Observation e04a0778-012f-48fb-91de-e1f8e17e9509 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Gemini: A Family of Highly Capable Multimodal Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:38.622584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:38.622584Z digest=sha256:201e1c34653204e1cda07fda7945727dc1dfb386bf0153ae487bb7abf4dbb8dd

Observation 771ed215-2e01-4c95-beb8-e346183c90ee · outbound

This paper cites Document understanding dataset and evaluation (dude).

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Document understanding dataset and evaluation (dude)

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:46.522644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:38.703728Z digest=sha256:69f5eb6d2f73680a64cb5fc13f443a29a48ea8ebaaf063628434ffccad073c96

Observation 8a635456-6fae-42b1-b5e6-1c6d2232609a · outbound

This paper cites The caltech-ucsd birds-200-2011 dataset.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning The caltech-ucsd birds-200-2011 dataset

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:46.288111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:38.811721Z digest=sha256:91bfaa830cd240f75b00c713d278d5e43f20bd47710bb68d075603e3f1eb3ee6

Observation 17cea784-ef6c-4838-ba07-805935d5b36f · outbound

This paper cites Screen2words: Automatic mobile ui summarization with multimodal learning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Screen2words: Automatic mobile ui summarization with multimodal learning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:46.064600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:38.899604Z digest=sha256:1b974921e714264a78bf667866aa2c4f094b63f22c8d9c5948999b93c7536251

Observation 3d2fe1c4-0a8c-4ee2-bfd4-ebe8f53cfd67 · outbound

This paper cites LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-08-06T12:18:41.874604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:39.029360Z digest=sha256:0b175f0a4bb876dd7412ae62eaec56d8648ddbf4cd3322d0566176d709a5743a

Observation 7509c407-ae30-4486-8ab7-12cd21de8e94 · outbound

This paper cites MLLM-Tool: A Multimodal Large Language Model For Tool Agent Learning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning MLLM-Tool: A Multimodal Large Language Model For Tool Agent Learning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:39.118189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:39.118189Z digest=sha256:28a5485022133f2944965f71d85d7f7b7a59fa015a268aa8e795184330b8273a

Observation 29515936-2adf-4b96-8dcd-d245ff7c8558 · outbound

This paper cites Mea- suring multimodal mathematical reasoning with math-vision dataset.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Mea- suring multimodal mathematical reasoning with math-vision dataset

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:45.869578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:39.207676Z digest=sha256:198b0d2816ed3a458ea49066d7405dd9c11330d9dd7e43591f41eb827df8dd06

Observation 5e426e2b-2a9f-4ee5-b01b-71570168c1a5 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:39.306583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:39.306583Z digest=sha256:7728791e155b79e6562387e767f839b0cf641e3f1bb43371d646d83c0cd32cd4

Observation b2630173-c5e9-4b66-b4b2-b531121a1309 · outbound

This paper cites Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:39.399123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:39.399123Z digest=sha256:27630afad721d4c86ddd7fef88dff9d7d6aa4e29c9f544fbb9e217a08d514692

Observation 4197aa3f-1e21-4c07-b7c3-0bd0709b8881 · outbound

This paper cites Grok-1.5 vision preview.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Grok-1.5 vision preview

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:45.603526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:39.504236Z digest=sha256:47193c944eb988d588e68c45875ba84e987e7b266a3866545f30267fff1e4392

Observation 744db546-4e65-46b4-96d3-0275004cc8e8 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:39.605249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:39.605249Z digest=sha256:402f9bae7b68d37a4982b92c10921f0e406cb38506c7e4217224015b12576513

Observation 6af0f074-14a1-49a9-a077-7722867c6ca3 · outbound

This paper cites Llava-cot: Let vision language models reason step- by-step, 2024.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Llava-cot: Let vision language models reason step- by-step, 2024

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:45.387347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:39.749545Z digest=sha256:eea9ccc1c347c6d932523439910378629a2515e47457c91ff1e212c1ea52ef49

Observation 6f935308-2bd6-42db-a8bb-dbd13a588ae8 · outbound

This paper cites Gpt4tools: Teaching large language model to use tools via self-instruction.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Gpt4tools: Teaching large language model to use tools via self-instruction

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:45.106833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:39.829944Z digest=sha256:4ab66142d50acb33a15eaca94c2d3c245edd3f17b19c997fc6050df44a71c745

Observation fc6a8b88-5b9e-4155-82de-657a638a8d80 · outbound

This paper cites React: Synergizing rea- soning and acting in language models.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning React: Synergizing rea- soning and acting in language models

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:44.865434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:39.900854Z digest=sha256:8fc62c70e57deb3f65762deb2f6dd1f2e06f662dcc274d5afb92b251a1a47e89

Observation 6aed4166-fd33-4939-ae15-4e8f720b5c6a · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:40.013192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:40.013192Z digest=sha256:91e27573a10b8535783350065bb2e4d561f8fd2a0b06fc6457861d39f64217d8

Observation bed3ff95-f0d1-42b0-b8b5-0604add61aee · outbound

This paper cites Agent lumos: Unified and modular training for open-source language agents.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Agent lumos: Unified and modular training for open-source language agents

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:44.640480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:40.111923Z digest=sha256:cd517ac5c1021f4eb1e036b21cb8f3cd769c69eb0ff6c54a540ec7778a008609

Observation 09ac949a-d7ff-4981-a42b-e5acb84a3f31 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:44.437025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:40.209708Z digest=sha256:5229d0c14d8086025a1799d87a6cfaab1ae60c5b5045a3d235dfc793b264f296

Observation 565e3f9f-ecd4-4be0-a9e7-b16f64f828c3 · outbound

This paper cites AgentTuning: Enabling Generalized Agent Abilities for LLMs.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning AgentTuning: Enabling Generalized Agent Abilities for LLMs

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:40.313560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:40.313560Z digest=sha256:742c7ba3e607e3163f37cf60b8049be36ae653299e7a39423c378e562ff20954

Observation 0967ca33-3478-437a-911a-75eb29b5b5d5 · outbound

This paper cites Raven: A dataset for relational and analogical visual reasoning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Raven: A dataset for relational and analogical visual reasoning

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:44.189547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:40.418499Z digest=sha256:973da102387ccc1ca5005a5fa45caefc7786ede4846ea06f90dff0e40107bd91

Observation 0cc7d3f4-c98e-406c-8321-11f839aad7c7 · outbound

This paper cites Swift:a scal- able lightweight infrastructure for fine-tuning, 2024.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Swift:a scal- able lightweight infrastructure for fine-tuning, 2024

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:43.965856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:40.563398Z digest=sha256:4705391875e67a949df710cce84f92bdfbb516bd80195123efaacc62eeb7d3d4

Observation 4411d2fb-c374-493c-8caa-d4770086b41b · outbound

This paper cites Seq2sql: Generating structured queries from natural language using reinforcement learning, 2017.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Seq2sql: Generating structured queries from natural language using reinforcement learning, 2017

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:43.750174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:40.636458Z digest=sha256:56013a32327ef7f824d4f787583a565a8bd6f662b2d34e5ff9e23e372b909a76

Observation de5ef537-70bf-40f3-9766-c216ab3a5fd5 · outbound

This paper cites Visual7w: Grounded question answering in images.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Visual7w: Grounded question answering in images

Reference 79

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T12:18:43.524971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:40.695193Z digest=sha256:06eb0d4ab963b77d445834f3837c2e7a103e98a6561ae2bc295d7c3d8b6fe06a

Observation 5256053a-33dd-46de-8251-152ba8f8af31 · outbound

This paper cites an unresolved cited work.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Unresolved cited work

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:40.808887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:40.808887Z digest=sha256:ab391d87f348db1fa1d7e92d36564de4b19ae580d54f240604f75275ba271774

Observation c81563ba-84aa-460c-a47c-c3f38c58dd28 · outbound

This paper cites an unresolved cited work.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Unresolved cited work

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:40.922255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:40.922255Z digest=sha256:8a80adf5c3bddbba1e2381a8bc8ffb71ce706151ea1398a2b6e927f8c3e42522

Observation 8c1b6981-1d26-4138-b7c8-99894dea8acf · outbound

This paper cites objects": [.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning objects": [

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:43.246936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:41.042512Z digest=sha256:6de13bf86b19eb92be0d4d4c082959b5db1102175ecc389aa997fb98d3411b15

Observation 05128088-b15e-47bb-ac2d-8c23e84e2e5c · outbound

This paper cites an unresolved cited work.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-08-06T12:18:42.999932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:41.141614Z digest=sha256:1eee450b03fcef8bf516f08509907f5e55ca9eacecc18a5147db7b245d51566d

Observation 3f7bae2a-c3bb-48b8-98f5-9085610ac3ff · outbound

This paper cites image_caption.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning image_caption

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:42.706027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:41.259999Z digest=sha256:f86a5fb4147894a8b72daceb5cd1fc054f70a66a2a0cbdb0eb26d2d9b98c39bd

Observation d5f08797-6684-43ed-88c0-c5610b21feb7 · outbound

This paper cites needed": true,.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning needed": true,

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:42.471282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:41.390516Z digest=sha256:64877c23795ee28015d0cb78af94714768a97663126d8333c5fc376570092d7d

Observation 72e7adf1-f915-45ba-b014-9a7760bf60b4 · outbound

This paper cites continue.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning continue

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:42.265587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:41.518126Z digest=sha256:738e18e4b2fa5bb63df1181f36607c97e48bf729611e2e37b5f8b3bafd473ae1

Observation d63ce0b8-1ee8-4fc6-a71d-aea9de96110b · outbound

This paper cites an unresolved cited work.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Unresolved cited work

Reference 170

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:33.027058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:33.027058Z digest=sha256:4390c75598cb0aa72e5a5bd03b9d1a4119d0188b3ba1734e1cfe176c9bfd665a

Observation cf2204af-5146-43f9-9318-0ccc42de42dc · outbound

This paper cites an unresolved cited work.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Unresolved cited work

Reference 251

Resolution
parse uncertain
raw_fallback, observed 2026-08-06T12:18:51.528283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T12:18:34.612419Z digest=sha256:387ab805bbf53dfa302465284fbcc47063d0b1e9310c1115bc28685715633937

Observation 846a591b-e1c1-4d71-bf21-87c1e714b30f · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:31.416961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:31.416961Z digest=sha256:08df3dffcf8bec63024a38248634edce86642094d0d634d079fd90b3ea99fd01

Pith citing papers

No inbound Pith citation observations are available.