Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T12:18:41.518126Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 0 inbound Pith citation observations for arXiv:2507.21924.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T12:18:41.518126Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
89 of 89 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5478395a-a3e2-4910-ab88-b83667be0c36 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0462b592-f0b1-40a1-ab79-724f553592da · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Qwen2.5-VL Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d43985dd-d259-4ece-bddd-04e480c521b6 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Jawahar, Ernest Valveny, and Dimos- thenis Karatzas
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e75490b-058d-4a84-81f8-52a64d95e4fb · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning An augmented benchmark dataset for geometric question answering through dual parallel text en- coding
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3baeeb8c-e4ef-4928-8cff-4e32bae54d62 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning FireAct: Toward Language Agent Fine-tuning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 463991f0-cc00-40b0-8798-d05ee1827c01 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Sharegpt4v: Improving large multi-modal models with better captions
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6705916d-5dbc-4a8f-8611-f80eeaf2b8f1 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Are we on the right way for evaluating large vision-language models? In The Thirty-eighth Annual Con- ference on Neural Information Processing Systems, 2024
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5d95906-1dc0-47af-9f83-a6d76b8c75b1 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2b48a63-6335-45e0-b0b6-35b2fc97d4c6 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b41b9cc9-cd38-4c58-9845-d34dd9504529 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92c9f48c-c1af-4cba-885d-a314261a25bf · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15d9bcc5-8cbf-46d3-96b5-c0b01a6075ea · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a80648dd-7b1d-458c-ab83-446322737701 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 971d3421-3b28-423f-9ecc-977122187998 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Clevr-math: A dataset for compositional language, visual and mathematical reasoning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f5df2a3-68f0-4fb0-9d97-992800e1cd89 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning PP-OCR: A Practical Ultra Lightweight OCR System
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5973dad-5e46-4e7a-9168-241b722a9d96 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Vlmevalkit: An open-source toolkit for evaluating large multi-modality models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b764f13b-cc2a-4907-983f-74e1c22f8fe0 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b4a3181-5f0a-43df-9a38-1b3f554a584c · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Multi-modal agent tuning: Building a vlm-driven agent for efficient tool usage
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e63cb2f6-1627-44cd-8a77-77aa298bd6b9 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Hallusionbench: an advanced diagnos- tic suite for entangled language hallucination and visual il- lusion in large vision-language models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1d71c6f4-9663-4f51-bedf-903cc3f88412 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Lora: Low-rank adaptation of large language models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eae346d7-dea3-4869-a0fd-8ff980354b34 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Icdar2019 compe- tition on scanned receipt ocr and information extraction
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3877697a-21a2-4942-8ffc-e5c49755e270 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Gqa: A new dataset for real-world visual reasoning and compositional 9 question answering
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7624ac31-6ceb-4462-b1f2-92689753961f · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning GPT-4o System Card
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6c32f4b-80b9-4a92-ba9d-501d5d813541 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Lawrence Zitnick, and Ross Girshick
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3cde0bb2-3704-4693-af26-65dc37e6cc7b · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning A diagram is worth a dozen images
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f257a9f9-e7c1-4d5c-a8b0-4517f4a3ed53 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Are you smarter than a sixth grader? textbook question answer- ing for multimodal machine comprehension
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 71b10f2c-bc23-4add-b219-882ce0430e63 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning The hateful memes challenge: Detecting hate speech in multimodal memes
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 126b78f9-f3f7-4151-8c5e-c6987d0b1265 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cb91d659-f428-42cd-9432-dd29556578f5 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning What matters when building vision-language models?
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32ee5be9-6523-403c-8df5-f6171b4f14aa · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Kankanhalli
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bf4c79f0-08e5-4ae4-9fb8-731fff0b5e2b · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Benchmarking Multimodal Retrieval Augmented Generation with Dynamic VQA Dataset and Self-adaptive Planning Agent
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6dbb72a-e8be-4707-98fa-1a65ca8b36d9 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Visual spatial reasoning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 150b32c7-24de-4cd7-a868-11d75292721e · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Visual instruction tuning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c24c4f2e-b304-45f7-a192-054771395861 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Improved baselines with visual instruction tuning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da4acb6d-e861-4fe5-bc60-f8c68030414d · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Llava-plus: Learning to use tools for creating multi- modal agents
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3c4785de-f488-46e6-894f-42253a8b2587 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 17eb8dc7-fd91-4c74-a714-e93a0aae3cc0 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning On the hidden mystery of ocr in large multimodal models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e8520f35-a2c8-4be5-be7d-da8fdc5baeea · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Inter-gps: Interpretable geometry problem solving with formal language and sym- bolic reasoning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation eec69d49-8b0a-48c7-834b-7b86943d7753 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Iconqa: A new benchmark for abstract diagram understand- ing and visual language reasoning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ce7ff31f-0d94-4bb9-8493-408472c1e8f2 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 319aec87-0af3-4f86-bfe8-be9ab2818488 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ab389bd-043c-404e-8d5f-fce2f8f5d213 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Mathvista: Evaluating mathe- matical reasoning of foundation models in visual contexts
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0780e952-e4e3-4756-b3fb-50787697b212 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning ChartQA: A benchmark for question answer- ing about charts with visual and logical reasoning
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6f07a7d3-977d-4b42-a5e1-42ec6e1e802e · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Docvqa: A dataset for vqa on document images
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1c296ffe-614f-4b16-b661-1b1a212f80d4 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Infographicvqa
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6ae9bfaa-2a9c-4804-aa2a-dc2b857795ce · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Llama 3.2: Revolutionizing edge ai and vision with open, customizable models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 71d02cb0-f2c3-4195-be9c-e0e90823e35f · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Compositional chain-of-thought prompting for large multimodal models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 89d54ed3-7d75-46e4-9c19-469c7336160e · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Compositional semantic parsing on semi-structured tables
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 53b939d8-6681-442e-976a-2f9777000580 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 67ba3997-6b2a-4bf6-96b9-68bb29ab71f5 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning A benchmark of facial recognition pipelines and co-usability performances of mod- ules
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 65e6ac54-3195-45f1-bb0a-adf7d0b85a17 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Putting gpt-4o to the sword: A comprehensive evaluation of language, vision, speech, and multimodal proficiency
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1494792a-ecc3-4427-8057-b946dd8698eb · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Visual cot: Advancing multi-modal language models with a comprehen- sive dataset and benchmark for chain-of-thought reasoning
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation be5ce8d3-57c9-458f-9ab8-cb9c930acd9a · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7963086-d156-4656-8cc0-45f01486735e · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Textcaps: a dataset for image caption- ing with reading comprehension
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 031f965e-c81e-47ce-990d-88f3dc516a55 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Towards vqa models that can read
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 957c6707-7fcb-4195-84a1-e72f0b84820a · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc460c53-afcf-45cc-bb4e-231079fc7c88 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Tang, Angie Boggust, and Arvind Satyanarayan
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e04a0778-012f-48fb-91de-e1f8e17e9509 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Gemini: A Family of Highly Capable Multimodal Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 771ed215-2e01-4c95-beb8-e346183c90ee · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Document understanding dataset and evaluation (dude)
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8a635456-6fae-42b1-b5e6-1c6d2232609a · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning The caltech-ucsd birds-200-2011 dataset
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 17cea784-ef6c-4838-ba07-805935d5b36f · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Screen2words: Automatic mobile ui summarization with multimodal learning
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3d2fe1c4-0a8c-4ee2-bfd4-ebe8f53cfd67 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7509c407-ae30-4486-8ab7-12cd21de8e94 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning MLLM-Tool: A Multimodal Large Language Model For Tool Agent Learning
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29515936-2adf-4b96-8dcd-d245ff7c8558 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Mea- suring multimodal mathematical reasoning with math-vision dataset
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5e426e2b-2a9f-4ee5-b01b-71570168c1a5 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2630173-c5e9-4b66-b4b2-b531121a1309 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4197aa3f-1e21-4c07-b7c3-0bd0709b8881 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Grok-1.5 vision preview
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 744db546-4e65-46b4-96d3-0275004cc8e8 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6af0f074-14a1-49a9-a077-7722867c6ca3 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Llava-cot: Let vision language models reason step- by-step, 2024
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6f935308-2bd6-42db-a8bb-dbd13a588ae8 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Gpt4tools: Teaching large language model to use tools via self-instruction
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fc6a8b88-5b9e-4155-82de-657a638a8d80 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning React: Synergizing rea- soning and acting in language models
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6aed4166-fd33-4939-ae15-4e8f720b5c6a · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bed3ff95-f0d1-42b0-b8b5-0604add61aee · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Agent lumos: Unified and modular training for open-source language agents
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 09ac949a-d7ff-4981-a42b-e5acb84a3f31 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 565e3f9f-ecd4-4be0-a9e7-b16f64f828c3 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning AgentTuning: Enabling Generalized Agent Abilities for LLMs
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0967ca33-3478-437a-911a-75eb29b5b5d5 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Raven: A dataset for relational and analogical visual reasoning
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0cc7d3f4-c98e-406c-8321-11f839aad7c7 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Swift:a scal- able lightweight infrastructure for fine-tuning, 2024
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4411d2fb-c374-493c-8caa-d4770086b41b · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Seq2sql: Generating structured queries from natural language using reinforcement learning, 2017
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation de5ef537-70bf-40f3-9766-c216ab3a5fd5 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Visual7w: Grounded question answering in images
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5256053a-33dd-46de-8251-152ba8f8af31 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Unresolved cited work
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c81563ba-84aa-460c-a47c-c3f38c58dd28 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Unresolved cited work
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c1b6981-1d26-4138-b7c8-99894dea8acf · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning objects": [
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 05128088-b15e-47bb-ac2d-8c23e84e2e5c · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Unresolved cited work
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3f7bae2a-c3bb-48b8-98f5-9085610ac3ff · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning image_caption
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d5f08797-6684-43ed-88c0-c5610b21feb7 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning needed": true,
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 72e7adf1-f915-45ba-b014-9a7760bf60b4 · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning continue
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d63ce0b8-1ee8-4fc6-a71d-aea9de96110b · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Unresolved cited work
Reference 170
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf2204af-5146-43f9-9318-0ccc42de42dc · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Unresolved cited work
Reference 251
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 846a591b-e1c1-4d71-bf21-87c1e714b30f · outbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.