Pith. sign in

Paper Citation Record · LEDGER

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks

As of 17 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 3 inbound Pith citation observations for arXiv:2412.18072.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18072 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:08:18.584497Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:45:30.706182Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T06:18:05.238771Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a4f5c203-505b-4adf-86d3-ef65d151a14e · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:17.394176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:17.394176Z digest=sha256:1073659c2f4765341a2472ea3ba1541d322481d08625c05fcd338132901f5420

Observation 3b318b47-206f-4f58-9da8-bfa1303795da · outbound

This paper cites GPT-4 Technical Report.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:17.406018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:17.406018Z digest=sha256:2ba60a360433305d759b554008e318ff055b774e87eda2ced5ebd16099e449a5

Observation 983816f5-d423-4c23-a836-b1c20b6bbf90 · outbound

This paper cites Neural module networks.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Neural module networks

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:08:23.359929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:08:17.426167Z digest=sha256:0f5f5e0bbf0556c318aa8e067b201c410de5e1bbc74bbd5b8c657d84cd0fa55f

Observation 46b88a55-f9e9-4309-9617-c536a7a0e4ec · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks The claude 3 model family: Opus, sonnet, haiku

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:08:23.229639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:08:17.445014Z digest=sha256:86ce13d7fad344c5ecb7785970d1242da3e8b2b5d39806f2f77c54b2d2f059bd

Observation 741337ec-35d3-486e-933f-4712d0f8c705 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:17.461636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:17.461636Z digest=sha256:8fcb0152e0624e165b677ecc0aa8481eca5c03492b8426846930f3ca3c9cf37e

Observation ae1473ba-30fd-4330-b683-06fac6c405a9 · outbound

This paper cites Llemma: An Open Language Model For Mathematics.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Llemma: An Open Language Model For Mathematics

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:17.500113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:17.500113Z digest=sha256:9887b9f912a9355eff13365b9950dc1256c4078039b1ab79a56779e9a9f873e4

Observation 9901233e-77a0-4cb9-98f1-4fff56020835 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:17.513467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:17.513467Z digest=sha256:dbc8ab121776860df5b91b9cb2e4715ba74d1d2f8cb4ff8d60a4d2d9ac0e5a68

Observation c6314d73-c738-4fc7-9268-03b3a93244db · outbound

This paper cites 2alliance.can.ca 3https://vectorinstitute.ai/#partners Audiolm: a language modeling approach to audio genera- tion.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks 2alliance.can.ca 3https://vectorinstitute.ai/#partners Audiolm: a language modeling approach to audio genera- tion

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:08:23.133920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:08:17.526591Z digest=sha256:7f2ec3735d2dfa38cec0c90ea70331c96552313f1fcac224413028e4ac9e3c37

Observation 1d4422b2-3183-4872-b068-6c6730d3f162 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:17.537399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:17.537399Z digest=sha256:0ebe66d34bb07c7451554cc6270cafbce01be83f516cb5be48bf281bad1d93bf

Observation 73ff7ef2-95d8-4f8a-9d90-585a70a3902a · outbound

This paper cites LangChain, 2022.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks LangChain, 2022

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:08:23.040723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:08:17.551693Z digest=sha256:fc1aea63540e207dea0ee180e88eb8ca48e274bd898368b715bf2cd3dc8ef73a

Observation 1653d4f0-79b3-476c-a6db-990c766c1b28 · outbound

This paper cites Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:08:22.929265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:08:17.570667Z digest=sha256:530002c2afd19f96fc2ab5ebee4e7932f36a5c4831d0a4611de5839258ca6f6d

Observation 39b13b4b-46d1-4227-8a77-cd16fb3cf9a5 · outbound

This paper cites AutoAgents: A Framework for Automatic Agent Generation.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks AutoAgents: A Framework for Automatic Agent Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:17.587888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:17.587888Z digest=sha256:46d8899569b830b06878a4c49de6824166373fd4d96992b52240c55a0227b495

Observation 1683ec43-3ce3-4db5-b21f-6205477ce7fa · outbound

This paper cites VideoLLM: Modeling Video Sequence with Large Language Models.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks VideoLLM: Modeling Video Sequence with Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:17.601318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:17.601318Z digest=sha256:ee017ec0afc1eba7fb0926df751ad57d10c0c5689a73be796bed02ecb3342e45

Observation 84e4fc5e-da9d-4de1-be9a-835ed6f73b47 · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Instructblip: Towards general- purpose vision-language models with instruction tuning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:08:22.821793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:08:17.607869Z digest=sha256:aa046ff76c42b67158870539d6e728a26525e985967d5e2f765a493d961dd7d3

Observation 4354f0ba-dcae-4d10-b5ba-6989fe90e10c · outbound

This paper cites SpeechVerse: A Large-scale Generalizable Audio Language Model.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:17.616757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:17.616757Z digest=sha256:3df1845b9f9890fbefd6f5e417872f14d144f7fa86c7d4663063b8f0a7ad0235

Observation 9a2b05dd-a6b8-439e-b9e7-480947e1954d · outbound

This paper cites Improving Factuality and Reasoning in Language Models through Multiagent Debate.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Improving Factuality and Reasoning in Language Models through Multiagent Debate

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:17.636551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:17.636551Z digest=sha256:d65c7ba8556134139cc1eb1b58f4d4917fcfc9ca55c85a828b4082fec6e0fd09

Observation 694a3193-f3c9-4de7-8b6a-d6adce4e5629 · outbound

This paper cites On Pre-training of Multimodal Language Models Customized for Chart Understanding.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks On Pre-training of Multimodal Language Models Customized for Chart Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:17.655767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:17.655767Z digest=sha256:ec04684ca051fe2b5a713317d1b8889af692650edde51bac6e88867aa2fc0d76

Observation 7b0ba662-d6c9-454e-b73b-af78bec497c8 · outbound

This paper cites Prompting large language models with speech recognition abilities.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Prompting large language models with speech recognition abilities

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:08:22.662945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:08:17.684755Z digest=sha256:2b8f064613adc58242fa2490b87c189b11ad926b53582e32525aa1d314d590ef

Observation 2ca317ef-07e8-4b03-8e4d-b73c4f167674 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:17.696188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:17.696188Z digest=sha256:65d0349dcf55666e4d7034abede679b61cbacc41c96e1f5b732516564d0e56dc

Observation 2290030f-6860-494e-ace9-15218e44e013 · outbound

This paper cites Gemini: A family of highly capa- ble multimodal models.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Gemini: A family of highly capa- ble multimodal models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:08:22.525190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:08:17.725188Z digest=sha256:f56a4fee09ea48edf1f1038bf2253a354d8643bd077e1ba1f1fc6cb33620e8cd

Observation eb3ab64d-216d-492d-ad3a-80e99f56e846 · outbound

This paper cites Github copilot, 2023.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Github copilot, 2023

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:08:22.416960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:08:17.740931Z digest=sha256:8c909d93214797e1da67f1f5ba9d3527e157159aa0cbfc84b7a7fee3ee41237d

Observation 470c86f8-f90f-44db-8af1-1a935bf09af8 · outbound

This paper cites Visual program- ming: Compositional visual reasoning without training.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Visual program- ming: Compositional visual reasoning without training

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:08:22.309323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:08:17.750684Z digest=sha256:b9b7f62d1e3155c417bce4230c71cff01f4ff3327dc175e5eacbcfdcd14f07e7

Observation 72b8b917-731e-4c93-9e43-96cebd46389d · outbound

This paper cites ChartLlama: A Multimodal LLM for Chart Understanding and Generation.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks ChartLlama: A Multimodal LLM for Chart Understanding and Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:17.760283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:17.760283Z digest=sha256:217ed3bedec14d5ac7fbb2c05a442cc367631550c05e69f95c802e120e17f6b7

Observation d2d58e38-d08a-4155-b724-ec49581487d9 · outbound

This paper cites 3d-llm: Inject- ing the 3d world into large language models.NeurIPS, 2023.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks 3d-llm: Inject- ing the 3d world into large language models.NeurIPS, 2023

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:17.770816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:17.770816Z digest=sha256:6dc902f7670e844ae2cf78a54c38ca33321ef629a6e996a2bad0d0474dc494b8

Observation d5f746e9-d8cc-4d06-b043-514f3c6ceeee · outbound

This paper cites RouterBench: A Benchmark for Multi-LLM Routing System.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks RouterBench: A Benchmark for Multi-LLM Routing System

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:17.783683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:17.783683Z digest=sha256:e777091b26d3f591223bc5b0ddea71ed9a0acb368f1a1a624d42f1c37b52ac6e

Observation ab77d179-93ad-4ada-936e-661e39d0a07a · outbound

This paper cites Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:17.796765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:17.796765Z digest=sha256:77f6098d7c172b2c575d72a97d2615962c9fa8b45135d5d0fddf1a7a0ee31bdf

Observation e8c4865e-0b23-414e-a94b-53f6fb404155 · outbound

This paper cites Mistral 7B.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Mistral 7B

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:17.804106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:17.804106Z digest=sha256:18622f103420f010a9840cf34c0ab481fe337655bdd3c5a58040a0d9e48330fc

Observation 2cfb9508-9058-4a28-9f9b-de80d680624c · outbound

This paper cites Inferring and executing programs for visual reasoning.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Inferring and executing programs for visual reasoning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:08:22.121119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:08:17.811583Z digest=sha256:f98d6d4cf532638fc0cb32ad4ca6e8e7631e530cce14b74d7b8ba4a57dd8a7c7

Observation 02053326-1dca-4782-a625-180f14b7b47e · outbound

This paper cites SMART-LLM: Smart Multi-Agent Robot Task Planning using Large Language Models.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks SMART-LLM: Smart Multi-Agent Robot Task Planning using Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:17.819774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:17.819774Z digest=sha256:26e6ff0193d5dedff96f20bd1398d98d909e74ebc14d2c58f0f660d82b49e900

Observation 10fda0a9-54f0-4947-b21e-2bf3ea596137 · outbound

This paper cites Segment any- thing.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Segment any- thing

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:17.830533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:17.830533Z digest=sha256:46f2304b4f2f4d6b35064b6c06c975167c47da1e5852f9647c33b84093eb5eb2

Observation 06aa363c-0278-4f65-bda5-6950f65bb378 · outbound

This paper cites Seed-bench: Bench- marking multimodal large language models.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Seed-bench: Bench- marking multimodal large language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:08:21.975705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:08:17.839948Z digest=sha256:896f1ace84b38291af9cd5164e7a7866492680d55d1f337ee23d7dd62a695924

Observation 93372e29-1cdd-4f3d-a3e9-de4806047786 · outbound

This paper cites Camel: Communicative agents for” mind” exploration of large language model society.NeurIPS,.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Camel: Communicative agents for” mind” exploration of large language model society.NeurIPS,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:08:21.850513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:08:17.866599Z digest=sha256:ce0723df29ff92869f9b50fd9e67732d03316a082ec76b653397afc7cef11773

Observation 79c5602a-17c9-4f2a-972e-d0eb14d98e8e · outbound

This paper cites Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:17.884025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:17.884025Z digest=sha256:a3142fae098443b63d7fe43214eec281f1b671b5e158fee5b80e2715a8bc2792

Observation 3861a152-e9e6-4d92-801c-b202b2f7b1e9 · outbound

This paper cites Visual instruction tuning.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Visual instruction tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:17.902000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:17.902000Z digest=sha256:0cadfde4093b44122d4c99b8bbb60d4e5859231c87028da7703fce191ca72c6b

Observation 7df78bfc-b529-4460-93eb-e96e6e96808b · outbound

This paper cites LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:17.917244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:17.917244Z digest=sha256:2fd5574ed775a5faa1f2d263b43f5021d3a2b7b7dd2c17f9f5f4f747b71b1673

Observation 8ca390be-8b4c-4ca6-93c2-04e10f0829f2 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:17.930205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:17.930205Z digest=sha256:4c7cad1a96da4e051122d556c98fd3fe2190334d07a570ad8af361c38ff1ee79

Observation 6cf7d2ef-a455-4244-ab19-7d8db1c1706f · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:17.951793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:17.951793Z digest=sha256:992543359cf137f03a10861ce5d1c6828dc025650319838624c39ca498c163b0

Observation 20ef57df-5b18-4d88-8bfb-3bb1cddc2930 · outbound

This paper cites Chameleon: Plug-and-play compositional reasoning with large language models.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Chameleon: Plug-and-play compositional reasoning with large language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:08:21.710940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:08:17.969022Z digest=sha256:768f8d7fab322b6332a2d078d0fc144a2bf830b734fd1a3e9377f1c70a7f7bc7

Observation 3a333811-22ab-4f53-89ea-13bb47220807 · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:17.980106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:17.980106Z digest=sha256:f73f7c27996a32d54023df2fe64e79b900806dd138a9dd79513441bfd70721e8

Observation 693dc120-eef1-4532-aa0a-d1a7aa08b944 · outbound

This paper cites ChartAssisstant: A Universal Chart Multimodal Language Model via Chart-to-Table Pre-training and Multitask Instruction Tuning.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks ChartAssisstant: A Universal Chart Multimodal Language Model via Chart-to-Table Pre-training and Multitask Instruction Tuning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:17.988137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:17.988137Z digest=sha256:fb8f579aa4e123292476aa913f50bb1d3247ae9995e7bc66baabb3d9009c7e64

Observation c3879649-0653-42c3-bac8-a076c6f618c0 · outbound

This paper cites RouteLLM: Learning to Route LLMs with Preference Data.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks RouteLLM: Learning to Route LLMs with Preference Data

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:18.006616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:18.006616Z digest=sha256:be6501db747c8ea423173f4bc907926cc9f4653715f368ad3b8c8530d4e3a975

Observation 97aeac83-7af6-4871-b09e-3498bb6ee1b3 · outbound

This paper cites Gpt-4 technical report.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Gpt-4 technical report

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:08:21.619164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:08:18.017384Z digest=sha256:8607478293f2df10ac3dc22c5f5e04d086723622caf9919f996563b38a8feb92

Observation 31f2458a-182c-46eb-8fe9-462d45e532d2 · outbound

This paper cites MemGPT: Towards LLMs as Operating Systems.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks MemGPT: Towards LLMs as Operating Systems

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:18.038234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:18.038234Z digest=sha256:38cc06dc70c420e0cf40e5dd4fa1e3cec118eb42776fab4ae7a2818c86ae1f65

Observation 3fcfff60-fc26-4211-a69b-0171121a890c · outbound

This paper cites Gorilla: Large Language Model Connected with Massive APIs.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Gorilla: Large Language Model Connected with Massive APIs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:18.054754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:18.054754Z digest=sha256:416598a4c6606aa340ad3030026568f8a6eb9a64b5d00629c85160e14fe60b75

Observation 56f005bd-6462-4336-a141-80b99726f7ee · outbound

This paper cites ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:18.086382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:18.086382Z digest=sha256:110fdd9a95a48e8e21a77fd500a8a6b9fd6985e89168e624c4ad7955cab47ac1

Observation 09af07d0-b562-411e-b31e-828dd68f7bba · outbound

This paper cites Code Llama: Open Foundation Models for Code.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Code Llama: Open Foundation Models for Code

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:18.126322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:18.126322Z digest=sha256:a9191f75d95671a03f242b952b7f9a8f27fac0a2067ed117d2a77ecb7f4bd382

Observation 41a29bdf-f416-48bc-a48b-88829c5bac19 · outbound

This paper cites Large Language Model Routing with Benchmark Datasets.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Large Language Model Routing with Benchmark Datasets

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:18.164866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:18.164866Z digest=sha256:3f735e9cf902b59a132d15a44dae0eb5588c26875c20df19dcdd41ac7cbb4466

Observation bd10d932-041e-482a-b737-b58e1e3dedd3 · outbound

This paper cites an unresolved cited work.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:08:21.524758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:08:18.229947Z digest=sha256:9a717872cb4b899d5cbf877bb4f0dd2e387638367efd101d7d2cf73fdc9faab4

Observation 721c3e8d-9974-4bcb-be31-ad77f223771e · outbound

This paper cites Vipergpt: Vi- sual inference via python execution for reasoning.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Vipergpt: Vi- sual inference via python execution for reasoning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:18.252543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:18.252543Z digest=sha256:62e5e0345c17c2e2c2edef90a784274cd9a1f4c06d502dbe49e245076bc004e5

Observation c589246c-7b61-4f76-bd10-94960fad4699 · outbound

This paper cites MedAgents: Large Language Models as Collaborators for Zero-shot Medical Reasoning.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks MedAgents: Large Language Models as Collaborators for Zero-shot Medical Reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:18.277962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:18.277962Z digest=sha256:55b9ea55f136d01a1c86239c24c21f5b9366faa4f79881bcfb7756bb03732bb2

Observation 8ca9ce45-b44f-4cd6-800b-73ecbde2cd71 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Gemini: A Family of Highly Capable Multimodal Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:18.293949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:18.293949Z digest=sha256:fef01bcebcf92303de1b08f710b54ffb442a581c541d1edfca9f6325672acdaf

Observation 51868703-5eff-4c5a-8fdf-2ae50fe43276 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:18.304021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:18.304021Z digest=sha256:43b67d5635729125b37b482edfaa59b433dd2a826ddc85384618c0847f30169d

Observation 5e3aec71-3350-437c-9779-871d176a9006 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:18.328620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:18.328620Z digest=sha256:31acd448bb02c748c99daed03696d15bd8ec182b05bd3c738d53de75756f9fa1

Observation 6340e983-61dd-4935-a9ba-3e7add5a23cc · outbound

This paper cites MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:18.358243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:18.358243Z digest=sha256:7b0a4420712e11ae0034cae29d2c5e936284b8cda95aeda34a2a7cc07feffd0f

Observation 252471fb-a11c-4262-80f7-2a748906c09d · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks CogVLM: Visual Expert for Pretrained Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:18.392280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:18.392280Z digest=sha256:3afdc26e640d2bb9bd69e4ac852d82461a8b257f9be5a32c476650df2f47e7e3

Observation 21dc3783-ae4b-4acb-8d00-ab75a66c6525 · outbound

This paper cites AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:18.424752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:18.424752Z digest=sha256:a5c1fcbf53df656a5d259a8053b9f853a125175113ec10dfa93b877b12c0ff9f

Observation 579a8b35-8950-409a-b660-e1ba0d015d43 · outbound

This paper cites MathChat: Converse to Tackle Challenging Math Problems with LLM Agents.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks MathChat: Converse to Tackle Challenging Math Problems with LLM Agents

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:18.474752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:18.474752Z digest=sha256:4e00bb96160e94f60c53258cb161f9128962f862fc80aa144321182bdad8e137

Observation 54e4ad15-c8db-4373-9fed-1556437a92ab · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:18.494750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:18.494750Z digest=sha256:0b74cf704f88ca13034766fbecec599ad1885192333c14295ff91ce71fe0ffde

Observation 3abb525e-8f75-4ce3-8eb4-5334ad1e6b5f · outbound

This paper cites Depth anything: Unleashing the power of large-scale unlabeled data.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Depth anything: Unleashing the power of large-scale unlabeled data

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:18.504689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:18.504689Z digest=sha256:ed0250bfd0d68a095c79bcb84c5d353c93118799693e8a878ac781fa18e327e9

Observation cffd7636-20ce-4ab6-a704-dd1ee3603890 · outbound

This paper cites Large language models for robotics: A survey.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Large language models for robotics: A survey

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:18.519871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:18.519871Z digest=sha256:0cfbe74ca5fad48e93cc514660d924d2e995bffd0777fc6423201727beddb51d

Observation 7d87adcf-d70a-48ec-ab58-d0c48e698052 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:18.534753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:18.534753Z digest=sha256:0882f98ec6a63c026176ce5ba7afadbbd8f30055faaae2a44921732bc398ce20

Observation 89604d47-2af2-440d-9a69-e8f8291d8436 · outbound

This paper cites InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:18.546696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:18.546696Z digest=sha256:2096dea876638b2e4c544f637eaaac78dee04b1040e1cf7412042aa396a1bae7

Observation aa4926f1-327b-4e7a-a726-9f4e60d4615c · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:18.565674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:18.565674Z digest=sha256:1019e26e798dcbab6f843241b8496a405ef1834d83eb52dca8819d0e85e185fe

Observation f7140698-038f-42a2-aa50-82dc76561d43 · outbound

This paper cites 3d-vista: Pre-trained transformer for 3d vision and text alignment.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks 3d-vista: Pre-trained transformer for 3d vision and text alignment

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:08:21.367042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:08:18.584497Z digest=sha256:9ece8d4f62bf7d55e72988f36c91683baa855bb8bf6750b81554f0f399def99d

Pith citing papers

Observation 65f1e13d-11f2-4578-9337-f6ffec9cf394 · inbound

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification cites this paper.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification MMFactory: A Universal Solution Search Engine for Vision-Language Tasks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.706182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.706182Z digest=sha256:585d47fc8c21bd965617d7a066fc503ff0c430617f77bd3ce08dd74662df0793

Observation bdb2de89-3591-442e-b628-56222da10caf · inbound

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs cites this paper.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs MMFactory: A Universal Solution Search Engine for Vision-Language Tasks

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:26:01.993879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:c562eca0270843a6c3cc2f5477921b0f205d6ea1e925d1d85af24865af86b274

Observation aca3ad99-8983-48d1-a0bb-daaead6f19e8 · inbound

Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning cites this paper.

Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning MMFactory: A Universal Solution Search Engine for Vision-Language Tasks

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:18:05.240829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-20T06:16:47.650748Z digest=sha256:00f7fcab9fa08210d5e91c9e86d8cc9b127217562f0dc81200b182c2184b0234