Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T21:07:31.387726Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 100 of 121 outbound references and 100 inbound Pith citation observations for arXiv:2408.01800.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T21:07:31.387726Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T04:29:34.801288Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T11:37:03.161139Z
100 of 121 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6d5ebdc8-d0d4-4b72-839c-69066ad0d0b8 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7a151eb2-344f-4716-8226-a5df579b82d5 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 87f63b89-633d-43e7-ae02-462c4fbd2773 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone RealCQA: Scientific chart question answering as a test-bed for first-order logic
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f03ce59d-1456-4127-b762-010895c1e5eb · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Flamingo: A visual language model for few-shot learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9098d65f-cacd-483e-8aab-6b53ce4c3e26 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Introducing the next generation of Claude
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 88ce68f2-cc97-4cd5-a41b-e4faceca90a5 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone VQA: Visual question answering
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8726c00d-f8be-4aad-ae73-bb128016f035 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 745b3819-ff61-48f5-aa0c-b4be5368bded · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Gemma: Introducing new state-of-the-art open models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7e200340-6491-4511-9270-b894d99d8207 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Introducing our multimodal models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation efa43f8c-e91c-481e-82d1-74bedac78c2b · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone BELLE: Be everyone’s large language model engine
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ec3cc497-4dc9-4bf2-be44-88cb9918c22a · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone PaliGemma: A versatile 3B VLM for transfer
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cc4d3295-38e4-48f4-a8d1-59e85f2cce6c · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Scene text visual question answering
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ca1c7fda-b4b2-40aa-beb5-ca5a93b30bfd · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone OCR-IDL: OCR annotations for industry document library dataset
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b855ef06-e106-4bf2-ab21-b5edce1ecbcf · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Sparks of Artificial General Intelligence: Early experiments with GPT-4
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3439004c-b48c-4ddc-b07f-954109e50740 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone COYO-700M: Image-text pair dataset
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d3171b1e-821a-497a-abdc-86314ce89e5a · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone TextOCR-GPT4V
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 82d59889-eab0-4f66-8d03-5f22c6da5b96 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 56e4d9ec-0f37-4b91-a8f7-79e175ca21f7 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 93606d2e-aff9-4192-83a7-6235d9996db4 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fd8bc126-df1e-4dd1-bfe5-676d14f3fd45 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6b89ee85-9cee-4256-ba48-f8092aedb43a · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0b3ae36d-02f4-403c-8933-ec109337e893 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone TabFact: A Large-scale Dataset for Table-based Fact Verification
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c652621a-5ca1-4e1f-85fe-d51f742f87e3 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0f075659-094d-4266-bfeb-e2cddfe80577 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Are deep neural networks smarter than second graders? In CVPR, pages 10834–10844
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 609ae3d3-49e1-4bf3-95bc-b2c8c79df22a · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3211a1a1-c0b5-4310-a295-b75c02bf4afc · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone OpenCompass: A universal evaluation platform for foundation models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6a6ad91b-a52d-4f01-b805-fbb1d5cbb9a3 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone XTuner: A toolkit for efficiently fine-tuning LLM
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 51b5b339-38ff-4769-a2e3-ca0b0b0418e5 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Visual Dialog
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a846f329-e9b1-43a8-af1d-46ec64abd702 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Project Astra
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2f7e3cd1-450b-4dea-a3b3-1bd8887e5ca2 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Enhancing Chat Language Models by Scaling High-quality Instructional Conversations
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation dedd3424-8a32-42d5-8b5f-655daceeb726 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1e1aa477-1510-4eba-839a-5bcdcbce7081 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9767d1b8-3fdf-4106-9e9f-f54e9f8b517c · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1172e8f3-60a4-4156-ac72-91c7bbaab1e4 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Are you talking to a machine? dataset and methods for multilingual image question
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2fbbfae9-5780-4894-9976-efe49c0cef06 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Wukong: A 100 million large-scale Chinese cross-modal pre-training benchmark
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c3aef501-530e-4603-8906-97730d27fdf0 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone LVIS: A dataset for large vocabulary instance segmentation
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9191c5d4-7b6d-4061-ba8d-5918f928290f · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Synthetic data for text localisation in natural images
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c4cc5c2d-dbff-48dc-bd6b-5222c5cc02a5 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone VizWiz Grand Challenge: Answering visual questions from blind people
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 452d07d0-b83c-43a4-a3bf-d1f59e28ffea · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Efficient Multimodal Learning from Data-centric Perspective
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6a91c791-6dcb-4108-9cd7-332b2482e400 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Training Compute-Optimal Large Language Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6897686f-a492-4a0a-b76b-623373621975 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Large Multilingual Models Pivot Zero-Shot Multimodal Learning across Languages
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 24e9067f-6f69-44f3-81ff-721ada57ac0c · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation eb328202-6adb-41c7-ab38-2652989249c1 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Language is not all you need: Aligning perception with language models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b989c5f6-c14a-44da-85bd-01c596c8bac4 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone GQA: A new dataset for real-world visual reasoning and compositional question answering
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8d376466-b0cf-4a6e-add8-c76b5fa92038 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Phi-2: The surprising power of small language models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1e8a7373-4dde-4fde-99a0-715b65d2fd17 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4568e5b7-0b99-4622-9465-affb7d3bdcff · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone DVQA: Understanding data visualizations via question answering
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bd2c8db3-1ec4-4d64-afaa-61abe9c4482f · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone FigureQA: An Annotated Figure Dataset for Visual Reasoning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1eb88075-3e75-4ed5-84cc-b19b0d0101b8 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone A diagram is worth a dozen images
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e76bee8c-dae1-4278-ae31-d09346b64703 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone OCR-free document understanding transformer
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 90b193fa-0014-41d5-b050-5e011aac1754 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Visual Genome: Connecting language and vision using crowdsourced dense image annotations
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0a226c93-0c92-4f0c-9afa-1c63ec812ad7 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone What matters when building vision-language models?
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 35efb39b-ebae-42b7-98b6-aeec6161d9f8 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone LLaV A-NeXT: Stronger LLMs supercharge multimodal capabilities in the wild
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 595c9fd1-5b54-4283-8759-7ffa4cb2508b · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone BLIP-2: Bootstrapping language-image pre- training with frozen image encoders and large language models
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 60f2322e-4ed9-4fe8-ad74-0d9acc528876 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f39c6295-edc3-4756-815f-a84646d8a458 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 66b1e338-6978-4354-b961-c1cc452a5357 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone OpenOrca: An open dataset of GPT augmented FLAN reasoning traces
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 18747935-8c71-42d1-b77d-21cb407e6c3f · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Microsoft COCO: Common objects in context
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 80d449a9-6083-4149-9464-8edb16df5c83 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7c159ae6-3799-4fe7-8e23-f995a459016e · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone LLaV A- NeXT: Improved reasoning, OCR, and world knowledge, January 2024
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 73b1eb7f-48a0-4494-8f8a-6a90444479d9 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Visual instruction tuning
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 972b9836-c850-40e1-a6d8-e542c78caefd · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone MMBench: Is Your Multi-modal Model an All-around Player?
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 108de9e8-c92b-41a0-9cc7-4afe1ab1f30f · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d0d22a99-2dca-4228-b9e2-14d19b081665 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d30939c3-220d-4b71-877c-076d5f2ae595 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e601ca92-2c99-4e53-b93d-c3a399527b12 · outbound
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ecfee39c-28c3-4995-a159-373accc69c08 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 08aad6d0-70cf-45d3-963b-3cede61bb950 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bd8b2100-bb71-4035-9a94-90dd5cb602c0 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d7a5fec1-4360-4e16-af45-f8b29cce6e7c · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e5d880a3-0c7e-4501-b2c3-2096a3c612c8 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone OK-VQA: A visual question answering benchmark requiring external knowledge
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ed6e5273-0706-47f2-a835-e7942e9d7337 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 67102baa-ddfc-4c7b-8bf0-0dfadfe77f77 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone DocVQA: A dataset for VQA on document images
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4be4634d-1d1a-453e-84be-f7055a3b6735 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone InfographicVQA
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5e3cba1e-776f-47fa-bcee-c681ddb079c6 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e1f8cdde-7050-4609-b96f-5a41929157cc · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone OCR-VQA: Visual question answering by reading text in images
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7953d74d-16ee-486b-b36d-b37ae4009eaf · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Hello GPT-4o
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3d48d524-407d-4de5-bc89-3faaa5014c54 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Compositional Semantic Parsing on Semi-Structured Tables
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a9217093-3703-4213-bf7b-d22d82177607 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d9da07d1-7782-4963-8c3b-4695ac298cb6 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Flickr30k Entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6104f17d-8b64-49db-94de-896a36f9551f · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Direct preference optimization: Your language model is secretly a reward model
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 77c004ff-bcd5-447e-8f3f-5bb8b57d5a1d · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3de7390d-864b-452c-8809-b01ec83b88bb · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Exploring models and data for image question answering
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 85efcd41-1440-4f05-bd22-275b9eada7f5 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Object Hallucination in Image Captioning
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 12b3564d-545f-4c78-8024-5e80a4b08d17 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone LAION-5B: An open large-scale dataset for training next generation image-text models
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2bdffb64-52c6-44a3-8a61-d15a2729e7f4 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone A- OKVQA: A benchmark for visual question answering using world knowledge
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 70168559-df60-4ac7-b5d3-ce0e027ef592 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone KVQA: Knowledge-aware visual question answering
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7f3bd094-031b-4813-9b5d-0230c43ae419 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Conceptual Captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ec631bd9-21bf-42c6-b7a1-6e73cc112ba4 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone The Color of the Cat is Gray: 1 Million Full-Sentences Visual Question Answering (FSVQA)
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d2f2ca94-04eb-4fff-aad5-b6ea29ea52eb · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Towards VQA models that can read
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 26fbc4bd-1578-4bdf-ad1a-c58ece692a1e · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone WIT: Wikipedia- based image text dataset for multimodal multilingual machine learning
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 243de297-b9e1-40fc-adb5-860457ac4de1 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Kleister: Key information extraction datasets involving long documents with complex layouts
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 68dc5eae-f051-4d5c-b852-f1d588b0235e · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone A corpus of natural language for visual reasoning
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 28e659e5-ed22-443f-97cd-a2507c035fc1 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone DeepForm: Understand Structured Documents at Scale — wandb.ai
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a9c90a9f-e3c2-4438-8d2a-5a55daed2f9d · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone VisualMRC: Machine reading comprehension on document images
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 244bce56-2d6c-4971-8faf-508c6e5ac84e · outbound
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8aee351b-0802-46e4-a451-8e48a0a876b3 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone OpenHermes 2.5: An open dataset of synthetic data for generalist LLM assistants
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cbd8a619-6209-4510-84c8-8018e97c2510 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 494deac7-1827-46a9-ba61-cf7e3669be38 · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone LLaMA: Open and Efficient Foundation Language Models
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation deb21c91-904d-45f6-8645-10e8fbfcab8d · outbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5d250947-2c35-464a-94cc-17cc158247b5 · inbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 259
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cf4a475c-3692-42f6-8ae1-c03796870c41 · inbound
CogVLM2: Visual Language Models for Image and Video Understanding MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 048256bc-2905-4a64-b816-96f26b93203a · inbound
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation daf66416-f08f-4d51-8099-4bcdd218e5ec · inbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 123
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 17c4ad5f-dc70-4c37-ad5e-3797dcbb1393 · inbound
ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 64e38263-963c-4c30-bf23-819b050d7b98 · inbound
VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5b83001e-5293-432e-8119-73a7381bcc5b · inbound
LLaVA-CoT: Let Vision Language Models Reason Step-by-Step MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7e04521a-20fa-410c-9ea5-c78ad8e9b73a · inbound
Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 109
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c392ca64-216d-42c5-92cc-b81534e786c0 · inbound
NVILA: Efficient Frontier Visual Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1edd3384-b802-4f65-b7ce-d7a4f401e01d · inbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 275
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 416b2fef-0977-47fa-b7d8-10de4ec6081f · inbound
DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5fa15f62-dcc9-4bad-b5a9-9771821c0756 · inbound
GME: Improving Universal Multimodal Retrieval by Multimodal LLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c34f2ed9-59f0-4ab3-8227-9f7cff8899de · inbound
FineVQ: Fine-Grained User Generated Content Video Quality Assessment MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fa555d8-69a7-4e25-bfd6-cf2aaeecd6c6 · inbound
MM-MoralBench: A MultiModal Moral Evaluation Benchmark for Large Vision-Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2d55d716-a0c4-454c-a509-2261670d070e · inbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68a87777-4a55-48e2-9ca2-812259c9f4dc · inbound
WalkVLM:Aid Visually Impaired People Walking by Vision Language Model MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f54d04a-5df7-454b-8280-6272c16950d3 · inbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 66db1a8e-0e96-4913-b571-c44d84f7ba0f · inbound
CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0df0f12a-c6bb-45d4-98c9-88d37be09f77 · inbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9e8297b-fb45-4ee4-8ef3-3a56a86f5e7e · inbound
CultureVLM: Characterizing and Improving Cultural Understanding of Vision-Language Models for over 100 Countries MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e6cc233-c205-4650-a74a-1c5dccec854c · inbound
GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 111
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e2a0a33-88e0-4f27-80fb-af682b32c4ba · inbound
FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70e625b0-98a3-4bc0-b447-b3a6e47742e2 · inbound
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fb206a06-e5a0-4e86-87c7-d27abfecece4 · inbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f65249d-b997-44d2-a874-e22410aced12 · inbound
DriVLM: Domain Adaptation of Vision-Language Models in Autonomous Driving MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e4a0f9a-7f97-42a2-893c-ebb612184070 · inbound
Efficiently Serving Large Multimodal Models Using EPD Disaggregation MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 229b63c0-920d-4bf5-8e54-ef2870373f59 · inbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45a1ae2e-43c5-431f-93d3-48045b7b4696 · inbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de5346b8-f4ac-4498-8ad4-89cf841e88fb · inbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bf7c2dc-2048-4587-9c06-dfb079dabed5 · inbound
ChartCoder: Advancing Multimodal Large Language Model for Chart-to-Code Generation MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95c19102-a797-4ee6-907a-46d3e3319f82 · inbound
MiniRAG: Towards Extremely Simple Retrieval-Augmented Generation MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abdd08d9-1c36-4410-8b3f-a3c06f7b886e · inbound
Zero-Shot Scene Understanding for Automatic Target Recognition Using Large Vision-Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6299dc3c-d1e6-4c66-a861-d7ea2db71293 · inbound
MSTS: A Multimodal Safety Test Suite for Vision-Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 372eac7f-9528-4376-86e2-70e689fc2212 · inbound
EndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6510ff94-4e0b-4955-ab29-7a8d8f84d221 · inbound
Does Table Source Matter? Benchmarking and Improving Multimodal Scientific Table Understanding and Reasoning MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b794a00e-4b7d-40ad-af16-85c7f9576a09 · inbound
Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f1bde46-37a7-46b4-92ce-81e5ab375542 · inbound
Temporal Preference Optimization for Long-Form Video Understanding MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bf02c69-1611-4eb0-888a-7a11ed03c12f · inbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcdd8519-0379-4628-8d66-9235d10f0d7c · inbound
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 194
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a700a717-608f-4fbb-b18d-4bd2b1c7d44b · inbound
Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2aa5815-5ce0-48b4-927c-89cb709f28f6 · inbound
A Causality-aware Paradigm for Evaluating Creativity of Multimodal Large Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abd0c2c8-d85a-43af-b746-4a79d91a49a0 · inbound
Generating Negative Samples for Multi-Modal Recommendation MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54c3e1e7-ad50-4283-b2a0-586d384c5a29 · inbound
Ocean-OCR: Towards General OCR Application via a Vision-Language Model MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f972dcc8-9438-43fa-8503-ac71b087c2a5 · inbound
Can Multimodal Large Language Models be Guided to Improve Industrial Anomaly Detection? MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bd10117-1474-4938-a48d-c544a1805316 · inbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7abc8490-a37f-48ea-8e9a-50f82133dae3 · inbound
TAID: Temporally Adaptive Interpolated Distillation for Efficient Knowledge Transfer in Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6edbb00f-43ea-488c-9d79-e57a745fb3d2 · inbound
Scaling Inference-Efficient Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2996d25-593c-413e-affc-24f4c888d332 · inbound
BounTCHA: A CAPTCHA Utilizing Boundary Identification in Guided Generative AI-extended Videos MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 103
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1a9de48-f850-4970-9e30-ba437f924d71 · inbound
RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74188231-6d98-474c-82a7-f5e46f1afd07 · inbound
MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85737c1c-4f44-4ac8-b28b-184f8961b33b · inbound
VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e41759d-1964-4f45-a720-32c7192f653a · inbound
Ola: Pushing the Frontiers of Omni-Modal Language Model MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05bd489a-113c-4044-a4e5-e7baa155387c · inbound
Cached Multi-Lora Composition for Multi-Concept Image Generation MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91866d10-44df-4a43-9712-ecb46efe9795 · inbound
Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b816710-c759-41c7-b24f-d4ea4f6e995d · inbound
HCMRM: A High-Consistency Multimodal Relevance Model for Search Ads MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 831adb08-f292-4032-9ce2-52ce6f47f6c8 · inbound
Diffusion Instruction Tuning MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 028fa900-a965-4130-b355-e1a364996573 · inbound
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47909c14-4047-42c1-93b0-5f87e0fb7fe5 · inbound
Seeing is Understanding: Unlocking Causal Attention into Modality-Mutual Attention for Multimodal LLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5da04e2d-232c-48f3-bd4b-2462588b8b30 · inbound
R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3bdea0e8-8c80-4a5f-bac3-c0bf25f9f2b8 · inbound
R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 10ff2e47-624b-4e06-b7f6-4ebea279a1a0 · inbound
Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8d035bdb-0cb2-43e2-b338-88dafefc0bd9 · inbound
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b0837259-43ee-4398-a1a0-e5fd01681b30 · inbound
SmolVLM: Redefining small and efficient multimodal models MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation feae0bd9-8280-4765-92de-db8f78a2457a · inbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 136
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 497fbbe7-2723-44e2-9781-ae5f450fd426 · inbound
SkyReels-V2: Infinite-length Film Generative Model MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7318a428-6471-4495-abf4-707192b9afda · inbound
Social Human Robot Embodied Conversation (SHREC) Dataset: Benchmarking Foundational Models' Social Reasoning MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4c5624b1-b94f-4b5f-b519-a962239c7d6d · inbound
Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae9036c9-f303-4130-80b5-698248930b2a · inbound
UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1847bc0e-6b6c-46ae-b58f-46a59b596cdb · inbound
S2SBench: A Benchmark for Quantifying Intelligence Degradation in Speech-to-Speech Large Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcf3998a-0299-44c7-99b1-60537273aa35 · inbound
Clapper: Compact Learning and Video Representation in VLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 901e1fc9-dabd-4011-8231-b168b2820735 · inbound
STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fa76702-21c1-4f01-8b2a-ed31e90f6b6f · inbound
PhyX: Does Your Model Have the "Wits" for Physical Reasoning? MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1866e258-2d55-448d-b56c-dbcb34d455ad · inbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 565bbe37-6430-40e2-af36-f08aad36afa9 · inbound
Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f1abc63-c866-44a5-9634-c478b55cc4a1 · inbound
R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40a23375-22c8-4ce7-9b7a-86643a55d6ea · inbound
RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f60da05b-fea0-46ac-900d-b30b5ee232fa · inbound
From EduVisBench to EduVisAgent: A Benchmark and Multi-Agent Framework for Reasoning-Driven Pedagogical Visualization MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 602428fb-e95d-4e4b-966f-52d63434c168 · inbound
EVADE-Bench: Multimodal Benchmark for Evaluating and Enhancing Evasive Content Detection MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd735611-0f77-43ff-9c69-1e1ee27c4c29 · inbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad36b43e-9f5e-4e51-8c2c-75c1979abe93 · inbound
Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b5c142e-69e9-4709-91ce-7bed037176bb · inbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a0b16a1-b190-4077-8ad9-53da8bd22410 · inbound
MLLM-Guided VLM Fine-Tuning with Joint Inference for Zero-Shot Composed Image Retrieval MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dff226c-658e-48f7-9170-acd6bdca4573 · inbound
MLLM-Guided VLM Fine-Tuning with Joint Inference for Zero-Shot Composed Image Retrieval MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11038ea0-4bce-4a62-bb5f-b711dbd1a21b · inbound
Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84835a9a-54c1-4b11-9ecc-8ea290504b13 · inbound
Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07fee8e3-fcbb-424e-90ab-64d20723cb18 · inbound
TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d14e2908-41ad-4ccf-8550-d93ac69167d0 · inbound
Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 019551d9-c68a-4bf7-acad-7e23fd2cbdf4 · inbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5336192-b0dd-44e1-8dea-3ac387e47985 · inbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15b75bad-1466-45a4-8766-15d94b277c09 · inbound
OmniAD: Detect and Understand Industrial Anomaly via Multimodal Reasoning MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb2908f2-584b-4cca-a846-e6ff12b811b4 · inbound
VModA: An Effective Framework for Adaptive NSFW Image Moderation MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1eeb526c-3259-4c60-8d94-077fd5debf1a · inbound
Qwen Look Again: Guiding Vision-Language Reasoning Models to Re-attention Visual Information MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcb64f2c-9b35-48e2-8990-2a95b70739df · inbound
USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9d53f51-3628-4ccd-8161-bafbb69f4d9a · inbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 575783b0-dd21-4941-b5eb-464ca8c32aea · inbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1c6c84f-88c8-4d7f-876c-e230f0f6b518 · inbound
DisTime: Distribution-based Time Representation for Video Large Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e7525b8-2c27-4704-bc89-1698c2ce5cd5 · inbound
Mixpert: Mitigating Multimodal Learning Conflicts with Efficient Mixture-of-Vision-Experts MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8944cc7b-10b1-4c98-b1bf-5159dde0e1f4 · inbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd372c72-0a39-4af6-b4f3-f4bd6e731079 · inbound
EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8480938-50d6-4fad-828f-f38f3bd9edf5 · inbound
GuessBench: Sensemaking Multimodal Creativity in the Wild MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.