Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T05:08:18.584497Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 3 inbound Pith citation observations for arXiv:2412.18072.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T05:08:18.584497Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:45:30.706182Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-20T06:18:05.238771Z
64 of 64 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a4f5c203-505b-4adf-86d3-ef65d151a14e · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b318b47-206f-4f58-9da8-bfa1303795da · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 983816f5-d423-4c23-a836-b1c20b6bbf90 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Neural module networks
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 46b88a55-f9e9-4309-9617-c536a7a0e4ec · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks The claude 3 model family: Opus, sonnet, haiku
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 741337ec-35d3-486e-933f-4712d0f8c705 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae1473ba-30fd-4330-b683-06fac6c405a9 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Llemma: An Open Language Model For Mathematics
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9901233e-77a0-4cb9-98f1-4fff56020835 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6314d73-c738-4fc7-9268-03b3a93244db · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks 2alliance.can.ca 3https://vectorinstitute.ai/#partners Audiolm: a language modeling approach to audio genera- tion
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1d4422b2-3183-4872-b068-6c6730d3f162 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73ff7ef2-95d8-4f8a-9d90-585a70a3902a · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks LangChain, 2022
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1653d4f0-79b3-476c-a6db-990c766c1b28 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 39b13b4b-46d1-4227-8a77-cd16fb3cf9a5 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks AutoAgents: A Framework for Automatic Agent Generation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1683ec43-3ce3-4db5-b21f-6205477ce7fa · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks VideoLLM: Modeling Video Sequence with Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84e4fc5e-da9d-4de1-be9a-835ed6f73b47 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Instructblip: Towards general- purpose vision-language models with instruction tuning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4354f0ba-dcae-4d10-b5ba-6989fe90e10c · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks SpeechVerse: A Large-scale Generalizable Audio Language Model
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a2b05dd-a6b8-439e-b9e7-480947e1954d · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Improving Factuality and Reasoning in Language Models through Multiagent Debate
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 694a3193-f3c9-4de7-8b6a-d6adce4e5629 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks On Pre-training of Multimodal Language Models Customized for Chart Understanding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b0ba662-d6c9-454e-b73b-af78bec497c8 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Prompting large language models with speech recognition abilities
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2ca317ef-07e8-4b03-8e4d-b73c4f167674 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks BLINK: Multimodal Large Language Models Can See but Not Perceive
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2290030f-6860-494e-ace9-15218e44e013 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Gemini: A family of highly capa- ble multimodal models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation eb3ab64d-216d-492d-ad3a-80e99f56e846 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Github copilot, 2023
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 470c86f8-f90f-44db-8af1-1a935bf09af8 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Visual program- ming: Compositional visual reasoning without training
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 72b8b917-731e-4c93-9e43-96cebd46389d · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks ChartLlama: A Multimodal LLM for Chart Understanding and Generation
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2d58e38-d08a-4155-b724-ec49581487d9 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks 3d-llm: Inject- ing the 3d world into large language models.NeurIPS, 2023
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5f746e9-d8cc-4d06-b043-514f3c6ceeee · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks RouterBench: A Benchmark for Multi-LLM Routing System
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab77d179-93ad-4ada-936e-661e39d0a07a · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8c4865e-0b23-414e-a94b-53f6fb404155 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Mistral 7B
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cfb9508-9058-4a28-9f9b-de80d680624c · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Inferring and executing programs for visual reasoning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 02053326-1dca-4782-a625-180f14b7b47e · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks SMART-LLM: Smart Multi-Agent Robot Task Planning using Large Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10fda0a9-54f0-4947-b21e-2bf3ea596137 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Segment any- thing
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06aa363c-0278-4f65-bda5-6950f65bb378 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Seed-bench: Bench- marking multimodal large language models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 93372e29-1cdd-4f3d-a3e9-de4806047786 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Camel: Communicative agents for” mind” exploration of large language model society.NeurIPS,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 79c5602a-17c9-4f2a-972e-d0eb14d98e8e · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3861a152-e9e6-4d92-801c-b202b2f7b1e9 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Visual instruction tuning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7df78bfc-b529-4460-93eb-e96e6e96808b · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ca390be-8b4c-4ca6-93c2-04e10f0829f2 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cf7d2ef-a455-4244-ab19-7d8db1c1706f · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20ef57df-5b18-4d88-8bfb-3bb1cddc2930 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Chameleon: Plug-and-play compositional reasoning with large language models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3a333811-22ab-4f53-89ea-13bb47220807 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 693dc120-eef1-4532-aa0a-d1a7aa08b944 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks ChartAssisstant: A Universal Chart Multimodal Language Model via Chart-to-Table Pre-training and Multitask Instruction Tuning
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3879649-0653-42c3-bac8-a076c6f618c0 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks RouteLLM: Learning to Route LLMs with Preference Data
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97aeac83-7af6-4871-b09e-3498bb6ee1b3 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Gpt-4 technical report
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 31f2458a-182c-46eb-8fe9-462d45e532d2 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks MemGPT: Towards LLMs as Operating Systems
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fcfff60-fc26-4211-a69b-0171121a890c · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Gorilla: Large Language Model Connected with Massive APIs
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56f005bd-6462-4336-a141-80b99726f7ee · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09af07d0-b562-411e-b31e-828dd68f7bba · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Code Llama: Open Foundation Models for Code
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41a29bdf-f416-48bc-a48b-88829c5bac19 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Large Language Model Routing with Benchmark Datasets
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd10d932-041e-482a-b737-b58e1e3dedd3 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 721c3e8d-9974-4bcb-be31-ad77f223771e · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Vipergpt: Vi- sual inference via python execution for reasoning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c589246c-7b61-4f76-bd10-94960fad4699 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks MedAgents: Large Language Models as Collaborators for Zero-shot Medical Reasoning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ca9ce45-b44f-4cd6-800b-73ecbde2cd71 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Gemini: A Family of Highly Capable Multimodal Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51868703-5eff-4c5a-8fdf-2ae50fe43276 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e3aec71-3350-437c-9779-871d176a9006 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6340e983-61dd-4935-a9ba-3e7add5a23cc · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 252471fb-a11c-4262-80f7-2a748906c09d · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks CogVLM: Visual Expert for Pretrained Language Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21dc3783-ae4b-4acb-8d00-ab75a66c6525 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 579a8b35-8950-409a-b660-e1ba0d015d43 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks MathChat: Converse to Tackle Challenging Math Problems with LLM Agents
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54e4ad15-c8db-4373-9fed-1556437a92ab · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3abb525e-8f75-4ce3-8eb4-5334ad1e6b5f · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Depth anything: Unleashing the power of large-scale unlabeled data
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cffd7636-20ce-4ab6-a704-dd1ee3603890 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Large language models for robotics: A survey
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d87adcf-d70a-48ec-ab58-d0c48e698052 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89604d47-2af2-440d-9a69-e8f8291d8436 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa4926f1-327b-4e7a-a726-9f4e60d4615c · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7140698-038f-42a2-aa50-82dc76561d43 · outbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks 3d-vista: Pre-trained transformer for 3d vision and text alignment
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 65f1e13d-11f2-4578-9337-f6ffec9cf394 · inbound
Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification MMFactory: A Universal Solution Search Engine for Vision-Language Tasks
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdb2de89-3591-442e-b628-56222da10caf · inbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs MMFactory: A Universal Solution Search Engine for Vision-Language Tasks
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation aca3ad99-8983-48d1-a0bb-daaead6f19e8 · inbound
Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning MMFactory: A Universal Solution Search Engine for Vision-Language Tasks
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.