Pith. sign in

Paper Citation Record · LEDGER

MiniCPM-V: A GPT-4V Level MLLM on Your Phone

As of 11 August 2026, this Paper Citation Record lists 100 of 121 outbound references and 100 inbound Pith citation observations for arXiv:2408.01800.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.01800 v1

Coverage vector

measured 100 of 121 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T21:07:31.387726Z

measured 200 of 200 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 100 of 344 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:29:34.801288Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T11:37:03.161139Z

Reference resolution

100 of 121 outbound references displayed

  • verified exact41
  • verified fuzzy58
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6d5ebdc8-d0d4-4b72-839c-69066ad0d0b8 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-10T21:07:32.191837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:a70d341ad38284f81bb313b1e1f0ea2528c956b659f9529b5c4a6e938919cf7b

Observation 7a151eb2-344f-4716-8226-a5df579b82d5 · outbound

This paper cites GPT-4 Technical Report.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone GPT-4 Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-10T21:07:31.809622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:35b25aed4d5c378be827fa3657e92e856206c657a2113394b11f45a04f24baa3

Observation 87f63b89-633d-43e7-ae02-462c4fbd2773 · outbound

This paper cites RealCQA: Scientific chart question answering as a test-bed for first-order logic.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone RealCQA: Scientific chart question answering as a test-bed for first-order logic

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.307499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:1465d09c0f42bfe2412d0d1c7a8ec5bb03227a28640f123865c62a16c62e21c8

Observation f03ce59d-1456-4127-b762-010895c1e5eb · outbound

This paper cites Flamingo: A visual language model for few-shot learning.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Flamingo: A visual language model for few-shot learning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.311593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:62f0546abddc13cdb578422a63a023fd8e0f57049b3e17be8522973154779527

Observation 9098d65f-cacd-483e-8aab-6b53ce4c3e26 · outbound

This paper cites Introducing the next generation of Claude.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Introducing the next generation of Claude

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.317704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:a5af00cdbe92fbd4e12c926bcd05642675d0fc1b39c3b65058cc2277c508967e

Observation 88ce68f2-cc97-4cd5-a41b-e4faceca90a5 · outbound

This paper cites VQA: Visual question answering.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone VQA: Visual question answering

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.325104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:b6c992cf3a88acdcce2e29935628e1f42319a20ab13f8717ef47c6d323f50098

Observation 8726c00d-f8be-4aad-ae73-bb128016f035 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-10T21:07:31.778105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:4148d1f3269d3b4a74a93f5e79a4bf13d474d9b0bcad2116dffeb279c98c2161

Observation 745b3819-ff61-48f5-aa0c-b4be5368bded · outbound

This paper cites Gemma: Introducing new state-of-the-art open models.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Gemma: Introducing new state-of-the-art open models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.331255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:407a7922ba1a2393a32f231d16a49a07fe91c1de1e2ff59dcb371b8892f1ff38

Observation 7e200340-6491-4511-9270-b894d99d8207 · outbound

This paper cites Introducing our multimodal models.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Introducing our multimodal models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.335334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:92824dcd4dbf2c60233f360e6255b2fecd572a5533a238800643a18bdf5de4c7

Observation efa43f8c-e91c-481e-82d1-74bedac78c2b · outbound

This paper cites BELLE: Be everyone’s large language model engine.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone BELLE: Be everyone’s large language model engine

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.342276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:fbcd6ec2991f9a23e0a0acc357a885302b5bf1558c3496f7a8193df729473848

Observation ec3cc497-4dc9-4bf2-be44-88cb9918c22a · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone PaliGemma: A versatile 3B VLM for transfer

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:10:21.987351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:eab5ef0bf8306aef8dbb2d0d67057d67162fe8752d653a4131ca5933c339e7f2

Observation cc4d3295-38e4-48f4-a8d1-59e85f2cce6c · outbound

This paper cites Scene text visual question answering.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Scene text visual question answering

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.346596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:0aa1bf42692b19b58a367e597757b6f58caf4c4e78e7c5b21ec45edf18789ca3

Observation ca1c7fda-b4b2-40aa-beb5-ca5a93b30bfd · outbound

This paper cites OCR-IDL: OCR annotations for industry document library dataset.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone OCR-IDL: OCR annotations for industry document library dataset

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.354836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:80340af71565709e43ba33a58b9a81e7d9e1ab987147704a9ac744efe999ef08

Observation b855ef06-e106-4bf2-ab21-b5edce1ecbcf · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-10T21:07:31.676654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:4fa953d9d6a6c8c6f836ff6f8ea73b96beaf6fcf014334af04fc4792b84c9b19

Observation 3439004c-b48c-4ddc-b07f-954109e50740 · outbound

This paper cites COYO-700M: Image-text pair dataset.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone COYO-700M: Image-text pair dataset

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.359435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:e3a89d2240de586a8b55407b6907e8ba4b561c935c0447fada7508cd0a8cc1f1

Observation d3171b1e-821a-497a-abdc-86314ce89e5a · outbound

This paper cites TextOCR-GPT4V.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone TextOCR-GPT4V

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.366200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:cdc49ac15bfd2dcb0524fe00a2af14f101e2f9c76f8b64f631a529c438f9441b

Observation 82d59889-eab0-4f66-8d03-5f22c6da5b96 · outbound

This paper cites Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.381545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:d6c47fbf6f8045a5036324a568bbd7a100e5e6954754ff9143d0976c8b62f709

Observation 56e4d9ec-0f37-4b91-a8f7-79e175ca21f7 · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:20:22.000193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:78135c786a7579685230ad6fb0b190fb845fb04f7326164f3cb800e2e1d82eb1

Observation 93606d2e-aff9-4192-83a7-6235d9996db4 · outbound

This paper cites GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:07:31.981369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:c2f4e5a45ac5459530ac041de6a224e2d12c3a73e34b886897751af3f7f896ac

Observation fd8bc126-df1e-4dd1-bfe5-676d14f3fd45 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:52:36.167809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:c7d9b3d124710462d9d7d8807faa9572243d57f1a11323c58e58b8fd36648d9d

Observation 6b89ee85-9cee-4256-ba48-f8092aedb43a · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:08:13.182567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:5328101ba5decc7f7d74533b97d7f14f0a97057ddec60427e774d7d58996b8ee

Observation 0b3ae36d-02f4-403c-8933-ec109337e893 · outbound

This paper cites TabFact: A Large-scale Dataset for Table-based Fact Verification.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone TabFact: A Large-scale Dataset for Table-based Fact Verification

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T21:07:32.044962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:2d6b03071d9a63b0b779038c9384f2b5d03b9887ac5f2816383c70242d0890c8

Observation c652621a-5ca1-4e1f-85fe-d51f742f87e3 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:59.653229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:b6d3ef559673cb2a61d06f85164f59dd22878dc23eedecc6cb39ac648828a1fc

Observation 0f075659-094d-4266-bfeb-e2cddfe80577 · outbound

This paper cites Are deep neural networks smarter than second graders? In CVPR, pages 10834–10844.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Are deep neural networks smarter than second graders? In CVPR, pages 10834–10844

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.388095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:df23900f4351ab9a501dcb896f7d78a31893b2a9e3531d848edece1abec3ea38

Observation 609ae3d3-49e1-4bf3-95bc-b2c8c79df22a · outbound

This paper cites MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:35:38.347525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:bdbe98e3eb2649cf9f3d884106e5a6397661aceede35a1b3e78bb9ee054cc667

Observation 3211a1a1-c0b5-4310-a295-b75c02bf4afc · outbound

This paper cites OpenCompass: A universal evaluation platform for foundation models.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone OpenCompass: A universal evaluation platform for foundation models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.396800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:6baf7f040eb42c54653e3a714aba9cedab93e6d4843dd5f7ae083ce01b8d5ed5

Observation 6a6ad91b-a52d-4f01-b805-fbb1d5cbb9a3 · outbound

This paper cites XTuner: A toolkit for efficiently fine-tuning LLM.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone XTuner: A toolkit for efficiently fine-tuning LLM

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.407334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:05dee59c33f757effcd1b2ef9a54afae92d794c024b029d7e29452f2f381a93c

Observation 51b5b339-38ff-4769-a2e3-ca0b0b0418e5 · outbound

This paper cites Visual Dialog.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Visual Dialog

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.419182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:5339c9d1bb344a50825d634c2c89c2a1a717377ee5c5a93d3bb04f9602162f55

Observation a846f329-e9b1-43a8-af1d-46ec64abd702 · outbound

This paper cites Project Astra.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Project Astra

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.431399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:f0a5ce2e9220cfd55a1b273b9c2fa416144aa84fb2e2990243e1736b254917b3

Observation 2f7e3cd1-450b-4dea-a3b3-1bd8887e5ca2 · outbound

This paper cites Enhancing Chat Language Models by Scaling High-quality Instructional Conversations.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Enhancing Chat Language Models by Scaling High-quality Instructional Conversations

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:25:08.436185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:d8192031e2d329dec12dd8f581c10d1988f212d133b99f44c5c66ebb94163e7f

Observation dedd3424-8a32-42d5-8b5f-655daceeb726 · outbound

This paper cites InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:07:31.885344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:ddb14c341e6ba52b2d27143233489c45853e74b3bedde87260b2a6e476864d0f

Observation 1e1aa477-1510-4eba-839a-5bcdcbce7081 · outbound

This paper cites What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:07:31.925047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:f9827c964a7a83ebbc553bcf9fe72215964fb6a75ac46849c68b844a687bb7d1

Observation 9767d1b8-3fdf-4106-9e9f-f54e9f8b517c · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-10T21:07:31.933345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:8d21cc8a9f0af173c4f68170d627afac7cdd0a1d296f3b1817bf90f65a526ab5

Observation 1172e8f3-60a4-4156-ac72-91c7bbaab1e4 · outbound

This paper cites Are you talking to a machine? dataset and methods for multilingual image question.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Are you talking to a machine? dataset and methods for multilingual image question

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.440581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:717fdf16465990e2f0afd095b85ee96dc569ec2469acf45753926a5913425748

Observation 2fbbfae9-5780-4894-9976-efe49c0cef06 · outbound

This paper cites Wukong: A 100 million large-scale Chinese cross-modal pre-training benchmark.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Wukong: A 100 million large-scale Chinese cross-modal pre-training benchmark

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.450945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:18bf041ad09ffff6dc25c04f4db016fa720afbf940931c11d74cd8231afd6b5a

Observation c3aef501-530e-4603-8906-97730d27fdf0 · outbound

This paper cites LVIS: A dataset for large vocabulary instance segmentation.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone LVIS: A dataset for large vocabulary instance segmentation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.462756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:858d9e2b6614781893c68f18ca42cd03abde34836c0037a632814fb9d17ac212

Observation 9191c5d4-7b6d-4061-ba8d-5918f928290f · outbound

This paper cites Synthetic data for text localisation in natural images.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Synthetic data for text localisation in natural images

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.469780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:779f2e903380f94ef1675650d1f7b2a975909f8e9e8885ca3c6a91ba36193b28

Observation c4cc5c2d-dbff-48dc-bd6b-5222c5cc02a5 · outbound

This paper cites VizWiz Grand Challenge: Answering visual questions from blind people.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone VizWiz Grand Challenge: Answering visual questions from blind people

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.477931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:6d054ea8cf4fa90b00e61ad9e11fe3f2e644861722672442c7fd3485d99f7ab8

Observation 452d07d0-b83c-43a4-a3bf-d1f59e28ffea · outbound

This paper cites Efficient Multimodal Learning from Data-centric Perspective.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Efficient Multimodal Learning from Data-centric Perspective

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:07:32.169541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:ed381192cea9a19c1804c893901e9a613cd39ac39e0a4b4f46573a594316487c

Observation 6a91c791-6dcb-4108-9cd7-332b2482e400 · outbound

This paper cites Training Compute-Optimal Large Language Models.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Training Compute-Optimal Large Language Models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-10T21:07:31.561356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:96aa121b7b25f265ec34e748c457ddff60cc98238ca6289736d9927483af6953

Observation 6897686f-a492-4a0a-b76b-623373621975 · outbound

This paper cites Large Multilingual Models Pivot Zero-Shot Multimodal Learning across Languages.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Large Multilingual Models Pivot Zero-Shot Multimodal Learning across Languages

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:07:31.619251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:21eb6740d38b7aced9f8d30e371d892059da8822e0a355b8eea1b53ad1c0fbe1

Observation 24e9067f-6f69-44f3-81ff-721ada57ac0c · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:00:53.610624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:8cd3d0d3912b4ccb5087966baf6e58ddfeddbddccd7b0d978231c6f1e28da6dd

Observation eb328202-6adb-41c7-ab38-2652989249c1 · outbound

This paper cites Language is not all you need: Aligning perception with language models.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Language is not all you need: Aligning perception with language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.490537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:ce7062936352fa5efd76db7616d5ddae551b07be69d978c341f0c759b3cb74a6

Observation b989c5f6-c14a-44da-85bd-01c596c8bac4 · outbound

This paper cites GQA: A new dataset for real-world visual reasoning and compositional question answering.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone GQA: A new dataset for real-world visual reasoning and compositional question answering

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.496001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:67acf6531c32399620a075c013f3fbdfa5f537253dd4025cc38ee16ad649b4f4

Observation 8d376466-b0cf-4a6e-add8-c76b5fa92038 · outbound

This paper cites Phi-2: The surprising power of small language models.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Phi-2: The surprising power of small language models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.508354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:f231afed89879d932166c32b3bb8f569b2ecff9d2119e2cc836329bb27a118fe

Observation 1e8a7373-4dde-4fde-99a0-715b65d2fd17 · outbound

This paper cites CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.516256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:13f67b2ef939f812fc2e8489ff3dd0fbbfb8631258c8c163f3c3abf3cb3fc801

Observation 4568e5b7-0b99-4622-9465-affb7d3bdcff · outbound

This paper cites DVQA: Understanding data visualizations via question answering.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone DVQA: Understanding data visualizations via question answering

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.520437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:d0741e4ac9a01c6f48e69308e0949c91490b97bf38b17aefa0733e560c0d98f8

Observation bd2c8db3-1ec4-4d64-afaa-61abe9c4482f · outbound

This paper cites FigureQA: An Annotated Figure Dataset for Visual Reasoning.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone FigureQA: An Annotated Figure Dataset for Visual Reasoning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:07:31.851490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:ef920ba2087540c9941174e84c0762cfe14fd35d824326bfa0a84380be2772b5

Observation 1eb88075-3e75-4ed5-84cc-b19b0d0101b8 · outbound

This paper cites A diagram is worth a dozen images.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone A diagram is worth a dozen images

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.530256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:0e0e59299ecfb4501c7135dad335427f79520988cf1201d5c1597b7be01a44d6

Observation e76bee8c-dae1-4278-ae31-d09346b64703 · outbound

This paper cites OCR-free document understanding transformer.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone OCR-free document understanding transformer

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.534053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:3ed3501b36bf325d46225a22714410e6a497e7f33084caa8b7919602f1942848

Observation 90b193fa-0014-41d5-b050-5e011aac1754 · outbound

This paper cites Visual Genome: Connecting language and vision using crowdsourced dense image annotations.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Visual Genome: Connecting language and vision using crowdsourced dense image annotations

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.537900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:95c5f8eb6543714fd357bfe034111be317b07a0b1ce0df1127423e1fa5ab3cf9

Observation 0a226c93-0c92-4f0c-9afa-1c63ec812ad7 · outbound

This paper cites What matters when building vision-language models?.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone What matters when building vision-language models?

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:07:31.906488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:715606d7f383383a50ba6a0d6495b089e8b4520320ae4322f2ef621ecfcb1f20

Observation 35efb39b-ebae-42b7-98b6-aeec6161d9f8 · outbound

This paper cites LLaV A-NeXT: Stronger LLMs supercharge multimodal capabilities in the wild.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone LLaV A-NeXT: Stronger LLMs supercharge multimodal capabilities in the wild

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.545656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:2842dc89d1c4bbabf1b16bcbe635a7a4d350d846a3cea3451fbd3b7931320613

Observation 595c9fd1-5b54-4283-8759-7ffa4cb2508b · outbound

This paper cites BLIP-2: Bootstrapping language-image pre- training with frozen image encoders and large language models.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone BLIP-2: Bootstrapping language-image pre- training with frozen image encoders and large language models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.550042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:8ba90990b9254430ceb65da87c5570571de289d9bc3607f5e589cf772be89867

Observation 60f2322e-4ed9-4fe8-ad74-0d9acc528876 · outbound

This paper cites Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:07:31.947656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:56b6490cbaa38bdf7419cabedcec07d418875210e070d48af65e741ce9af1691

Observation f39c6295-edc3-4756-815f-a84646d8a458 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:44:47.683285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:82677640224990765035e9f005788d20e764b7c1cf3f86e40743976dfdd56399

Observation 66b1e338-6978-4354-b961-c1cc452a5357 · outbound

This paper cites OpenOrca: An open dataset of GPT augmented FLAN reasoning traces.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone OpenOrca: An open dataset of GPT augmented FLAN reasoning traces

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.554263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:041c6d8897bf1987a922b2b0b0a809c8390dc1d31084e53aa2af4330bc9ea584

Observation 18747935-8c71-42d1-b77d-21cb407e6c3f · outbound

This paper cites Microsoft COCO: Common objects in context.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Microsoft COCO: Common objects in context

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.561696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:3c8bcb73608ddd777ea4389a2fee6919464b6687019b2e1eab8d7e06fe32b494

Observation 80d449a9-6083-4149-9464-8edb16df5c83 · outbound

This paper cites Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-14T17:34:57.031001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:baf71cba52b35415320c9f7f031ac4e1c9def51ee2d4fa9a4c9fff88e142f68a

Observation 7c159ae6-3799-4fe7-8e23-f995a459016e · outbound

This paper cites LLaV A- NeXT: Improved reasoning, OCR, and world knowledge, January 2024.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone LLaV A- NeXT: Improved reasoning, OCR, and world knowledge, January 2024

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.570325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:2d8b2dbdb41928e1654987d18e485ca04ed9ba963fa7f975890f557a891000cc

Observation 73b1eb7f-48a0-4494-8f8a-6a90444479d9 · outbound

This paper cites Visual instruction tuning.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Visual instruction tuning

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.574275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:7d2ce9c7e888ead049334a3125c9cee01ebc2464a479b44c57e3902c6f203fe6

Observation 972b9836-c850-40e1-a6d8-e542c78caefd · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone MMBench: Is Your Multi-modal Model an All-around Player?

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-12T17:20:54.147488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:d0f4511e65be2873f8c7d16ef21dc5b6a07f7c18fb33c0c4ada689e2e547fd7e

Observation 108de9e8-c92b-41a0-9cc7-4afe1ab1f30f · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:55:36.035986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:76b9267577c2fd73534866462cf7a1527f1d6351b1596d6c7965e611ae10fc75

Observation d0d22a99-2dca-4228-b9e2-14d19b081665 · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:07:32.096671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:9a5d36a083850d8363bc147f3d04ff0d596ca416c6cb81850d864e3c7f706abf

Observation d30939c3-220d-4b71-877c-076d5f2ae595 · outbound

This paper cites MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:07:32.112611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:73c894b11185224778aa0f5966f6d31ee55db3bb1811b5251c3f595276be9300

Observation e601ca92-2c99-4e53-b93d-c3a399527b12 · outbound

This paper cites llama.cpp.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone llama.cpp

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.580412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:0a897b6859ef854fe033f6a17900e8d54e89a008bb843aa60fc6e57e51d73528

Observation ecfee39c-28c3-4995-a159-373accc69c08 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:58:54.896552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:d1c5b0294aeac29e2ff5ad6ea5c2971268ae6c16e67b978d67793079adb3078b

Observation 08aad6d0-70cf-45d3-963b-3cede61bb950 · outbound

This paper cites IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:07:32.144606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:6dc8b1dcf8ea49dd9032a8b6924c461e14c71c4144eb23527fef8db88e275473

Observation bd8b2100-bb71-4035-9a94-90dd5cb602c0 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.584267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:51d8122fd48e19f3ece97e36a9a854d955ba833fa3abf6ebbb1e2086d192437e

Observation d7a5fec1-4360-4e16-af45-f8b29cce6e7c · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:30:15.750849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:f6e9c05a71037cde5b4315d6e363fe9296f20e2e350aeb611b4d660f4128ed37

Observation e5d880a3-0c7e-4501-b2c3-2096a3c612c8 · outbound

This paper cites OK-VQA: A visual question answering benchmark requiring external knowledge.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone OK-VQA: A visual question answering benchmark requiring external knowledge

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.593296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:04bdaae1669f076419489285ab5815364170e639871ad62cb86b7ea1ffb6a0f6

Observation ed6e5273-0706-47f2-a835-e7942e9d7337 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:13:07.396442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:3bb9486b3381b526dad02510e24b809b9652933d4bf40dd5b375dee4cbf89ef9

Observation 67102baa-ddfc-4c7b-8bf0-0dfadfe77f77 · outbound

This paper cites DocVQA: A dataset for VQA on document images.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone DocVQA: A dataset for VQA on document images

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.597181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:2e515601f136611e28e33f1692dae2c17897444be82eb55d52bca3c5554bd57d

Observation 4be4634d-1d1a-453e-84be-f7055a3b6735 · outbound

This paper cites InfographicVQA.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone InfographicVQA

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.607246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:199180eca9e3602c67626e3b90f014657b8655c381c6025d6c1c847e548bd2c8

Observation 5e3cba1e-776f-47fa-bcee-c681ddb079c6 · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:09:36.761640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:442be789e6a08be607eca9a8daf5946bc6b67b5ec95a787299dd1a2c24b4717e

Observation e1f8cdde-7050-4609-b96f-5a41929157cc · outbound

This paper cites OCR-VQA: Visual question answering by reading text in images.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone OCR-VQA: Visual question answering by reading text in images

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.618111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:00627d31fd559b382c17da5e4fe84fdbc3a474df2167b12d95612d3a43821cb1

Observation 7953d74d-16ee-486b-b36d-b37ae4009eaf · outbound

This paper cites Hello GPT-4o.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Hello GPT-4o

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.628180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:42388fe0f8e7aac8d93932e8ecfb480daccf0b6981984560d9868f09c90d0f71

Observation 3d48d524-407d-4de5-bc89-3faaa5014c54 · outbound

This paper cites Compositional Semantic Parsing on Semi-Structured Tables.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Compositional Semantic Parsing on Semi-Structured Tables

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:07:31.719136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:b9d427d630a3061b309a9ffa6c0a138604a1e8d34a6e9eb4eca2f7b9f13eae76

Observation a9217093-3703-4213-bf7b-d22d82177607 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:19:48.115283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:1f9253d2348da7071c8bc0e90b8908ef893955ba4d66db93265ec0455531c357

Observation d9da07d1-7782-4963-8c3b-4695ac298cb6 · outbound

This paper cites Flickr30k Entities: Collecting region-to-phrase correspondences for richer image-to-sentence models.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Flickr30k Entities: Collecting region-to-phrase correspondences for richer image-to-sentence models

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.635121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:e2f2b589e86486745840394f70764380a52bdcefdd58651cacbfbbdda3e8c58d

Observation 6104f17d-8b64-49db-94de-896a36f9551f · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Direct preference optimization: Your language model is secretly a reward model

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.644717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:0287bd1c05e4dbe5306aceb2c8aae878c290b5f0ac22b7ef4ea24900a1afa669

Observation 77c004ff-bcd5-447e-8f3f-5bb8b57d5a1d · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-05-10T21:07:31.792365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:f1318458b38ec399522242e19b53efc019732c11ef9c76138a695f1d188cbdf9

Observation 3de7390d-864b-452c-8809-b01ec83b88bb · outbound

This paper cites Exploring models and data for image question answering.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Exploring models and data for image question answering

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.653121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:c4f27a9a041f13f885cfd0a4ccf00821f1d496581cbc183fa0c012c8c7c169e8

Observation 85efcd41-1440-4f05-bd22-275b9eada7f5 · outbound

This paper cites Object Hallucination in Image Captioning.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Object Hallucination in Image Captioning

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:07:31.834351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:da79d6f6f2169ee2886fe46aad6eeffa6e7a16543d4b00988b16c15ec842184a

Observation 12b3564d-545f-4c78-8024-5e80a4b08d17 · outbound

This paper cites LAION-5B: An open large-scale dataset for training next generation image-text models.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone LAION-5B: An open large-scale dataset for training next generation image-text models

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.660227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:4d0bd1530a1fcfc91569b7b1fe1a8e15daa82ad284d43c38a5149e37b9247e7b

Observation 2bdffb64-52c6-44a3-8a61-d15a2729e7f4 · outbound

This paper cites A- OKVQA: A benchmark for visual question answering using world knowledge.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone A- OKVQA: A benchmark for visual question answering using world knowledge

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.671540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:04333ffe1f84cc5c8d9e9cbd6d343cdf62ad8c29f21b3b6dcc3199ba922f1d73

Observation 70168559-df60-4ac7-b5d3-ce0e027ef592 · outbound

This paper cites KVQA: Knowledge-aware visual question answering.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone KVQA: Knowledge-aware visual question answering

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.683424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:10e328c0d71e09da59eca8041462f98303df8691a43f527975be552fe4bf8219

Observation 7f3bd094-031b-4813-9b5d-0230c43ae419 · outbound

This paper cites Conceptual Captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Conceptual Captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.690730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:8b0b65180caf45c7c46ee6daa97ba54addc1f7b5fc2b8cc8e9da8c522383be85

Observation ec631bd9-21bf-42c6-b7a1-6e73cc112ba4 · outbound

This paper cites The Color of the Cat is Gray: 1 Million Full-Sentences Visual Question Answering (FSVQA).

MiniCPM-V: A GPT-4V Level MLLM on Your Phone The Color of the Cat is Gray: 1 Million Full-Sentences Visual Question Answering (FSVQA)

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:07:31.898348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:c0495bb139a439618ef0f8075338be497b56a40422ffdd6fd36386934e448029

Observation d2f2ca94-04eb-4fff-aad5-b6ea29ea52eb · outbound

This paper cites Towards VQA models that can read.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Towards VQA models that can read

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.697129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:9484d432b8289ca4657c68c64a314dad159178eb0992e03dc65ed25139843c0e

Observation 26fbc4bd-1578-4bdf-ad1a-c58ece692a1e · outbound

This paper cites WIT: Wikipedia- based image text dataset for multimodal multilingual machine learning.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone WIT: Wikipedia- based image text dataset for multimodal multilingual machine learning

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.196533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:f7e5e976501bc42871a2d14b433426633ea105affc5aeef38f66c93e9ecff46b

Observation 243de297-b9e1-40fc-adb5-860457ac4de1 · outbound

This paper cites Kleister: Key information extraction datasets involving long documents with complex layouts.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Kleister: Key information extraction datasets involving long documents with complex layouts

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.204385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:27427969c1b439b00074f916f6e23ba44f427c9f873f02063241cf1dd3784384

Observation 68dc5eae-f051-4d5c-b852-f1d588b0235e · outbound

This paper cites A corpus of natural language for visual reasoning.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone A corpus of natural language for visual reasoning

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.208539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:49da67e9508d23baeeffe8efd63c37c131ef4c9bcf0dbc1178ea049417b08b6e

Observation 28e659e5-ed22-443f-97cd-a2507c035fc1 · outbound

This paper cites DeepForm: Understand Structured Documents at Scale — wandb.ai.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone DeepForm: Understand Structured Documents at Scale — wandb.ai

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.215102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:6321b3a1c815b81cee3a32cd45670d89eef35ee75c30905e6a107651255527aa

Observation a9c90a9f-e3c2-4438-8d2a-5a55daed2f9d · outbound

This paper cites VisualMRC: Machine reading comprehension on document images.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone VisualMRC: Machine reading comprehension on document images

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.222421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:0e0cb78739750ab0139f88f9520e6c30f1978121f77174cef44a57af245caeaa

Observation 244bce56-2d6c-4971-8faf-508c6e5ac84e · outbound

This paper cites Hashimoto.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Hashimoto

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.228083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:80f5969b987728ec553908e94fede84dec2a7c92bee32f9b6bf9e465b1163898

Observation 8aee351b-0802-46e4-a451-8e48a0a876b3 · outbound

This paper cites OpenHermes 2.5: An open dataset of synthetic data for generalist LLM assistants.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone OpenHermes 2.5: An open dataset of synthetic data for generalist LLM assistants

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T21:07:32.234390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:6eb09948cf43ed16d049d13e96935b6a67e309cc12274aa4b8b7fe3949a1678e

Observation cbd8a619-6209-4510-84c8-8018e97c2510 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:05:04.345806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:dc42f7ede7fcca9b1b357a241130f14202a5e65d032e51a889e8c1caa01a598e

Observation 494deac7-1827-46a9-ba61-cf7e3669be38 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone LLaMA: Open and Efficient Foundation Language Models

Reference 100

Resolution
verified exact
local_arxiv, observed 2026-05-10T21:07:32.013557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:1ad11dfc2b4ca4b0813b54d520bfa1b958158808a928d9bd8d94e6ddcd5e94e2

Observation deb21c91-904d-45f6-8645-10e8fbfcab8d · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 101

Resolution
verified exact
local_arxiv, observed 2026-05-10T21:07:32.019589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:0f7120d811453097d83e2b48c4367bf6c2fdf398bd083ffd172f5327a2cd1ced

Pith citing papers

Observation 5d250947-2c35-464a-94cc-17cc158247b5 · inbound

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models cites this paper.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 259

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:20:36.553203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:1b82806003a9f3dde1d10c92df0da5ac1c86eeab78f213f29c62728a55541cd1

Observation cf4a475c-3692-42f6-8ae1-c03796870c41 · inbound

CogVLM2: Visual Language Models for Image and Video Understanding cites this paper.

CogVLM2: Visual Language Models for Image and Video Understanding MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:10:27.778951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:30dc6500b5e62613a7a388fe8eb6abd1141a57eec252f53494dd4528b4a127b8

Observation 048256bc-2905-4a64-b816-96f26b93203a · inbound

MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark cites this paper.

MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-14T00:51:48.287816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-14T00:51:48.163349Z digest=sha256:07154e798009a9feb03a7e88355ef0dd3d94334446ac75ef4077c6265b4ca561

Observation daf66416-f08f-4d51-8099-4bcdd218e5ec · inbound

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models cites this paper.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 123

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:55:12.703706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:7f68ec5863f6b3ed308592dfeb2e7716e32ba0cf64e34ccd3c372d8182db2b04

Observation 17c4ad5f-dc70-4c37-ad5e-3797dcbb1393 · inbound

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection cites this paper.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.906671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:ab1a59f0914e743dd04f9a9ee243dfa1fb0e861be1fe8bc2b4b01a8f9d6e8b91

Observation 64e38263-963c-4c30-bf23-819b050d7b98 · inbound

VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents cites this paper.

VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:37:25.858911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T15:37:25.781240Z digest=sha256:72653ed382527e65da1ef40bdd12e8e9b916080667f4e4f937352eca8dbe93d9

Observation 5b83001e-5293-432e-8119-73a7381bcc5b · inbound

LLaVA-CoT: Let Vision Language Models Reason Step-by-Step cites this paper.

LLaVA-CoT: Let Vision Language Models Reason Step-by-Step MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-16T11:35:25.926654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T11:35:25.894465Z digest=sha256:34ebccc9c142c38fabb86d6d494fc5c2f1aeadbe08d3e113f3d141d0ad6034a7

Observation 7e04521a-20fa-410c-9ea5-c78ad8e9b73a · inbound

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization cites this paper.

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 109

Resolution
verified exact
local_arxiv, observed 2026-05-16T09:16:17.533490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T09:16:17.150383Z digest=sha256:a91e7c11cccaa16228f62a4a80c6c62aaf61f83221b007ff4c90484ead28b682

Observation c392ca64-216d-42c5-92cc-b81534e786c0 · inbound

NVILA: Efficient Frontier Visual Language Models cites this paper.

NVILA: Efficient Frontier Visual Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.085376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:d78cc618c5aada55197d7609692c52138ee7cacc1e2a48426e36e571a00170e1

Observation 1edd3384-b802-4f65-b7ce-d7a4f401e01d · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 275

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:07:32.698795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:b69ba6e6e5561cf67d0c664cd56d698ea79e556c302f291d8995d702d4e8af73

Observation 416b2fef-0977-47fa-b7d8-10de4ec6081f · inbound

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding cites this paper.

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 100

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:09:26.093799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T10:09:21.542356Z digest=sha256:dfd2b5f687d8a4a310df0d0c8272c36889c2567d07ff74bfb23c7e7069c0ef76

Observation 5fa15f62-dcc9-4bad-b5a9-9771821c0756 · inbound

GME: Improving Universal Multimodal Retrieval by Multimodal LLMs cites this paper.

GME: Improving Universal Multimodal Retrieval by Multimodal LLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:35:22.621781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T06:35:22.508168Z digest=sha256:5beadca629d94f8d5c9593e875b301488df6d583b361a99ad7674fc98aebf16d

Observation c34f2ed9-59f0-4ab3-8227-9f7cff8899de · inbound

FineVQ: Fine-Grained User Generated Content Video Quality Assessment cites this paper.

FineVQ: Fine-Grained User Generated Content Video Quality Assessment MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T00:53:00.900121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:53:00.900121Z digest=sha256:17041f902a726ff5bbe32fed980f4272b7f17d2ce07b91f280940d6435148bc9

Observation 7fa555d8-69a7-4e25-bfd6-cf2aaeecd6c6 · inbound

MM-MoralBench: A MultiModal Moral Evaluation Benchmark for Large Vision-Language Models cites this paper.

MM-MoralBench: A MultiModal Moral Evaluation Benchmark for Large Vision-Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:15:28.520364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T07:14:27.786918Z digest=sha256:2643b52c48f65321be470240a3a0049b7934fce12716e03a01510da9dade9a7f

Observation 2d55d716-a0c4-454c-a509-2261670d070e · inbound

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking cites this paper.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:55.085117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:55.085117Z digest=sha256:3ff34304165968f5a3e2fc219ee2233fe06b31dc73d2d90f457c96d3c9f3a3ae

Observation 68a87777-4a55-48e2-9ca2-812259c9f4dc · inbound

WalkVLM:Aid Visually Impaired People Walking by Vision Language Model cites this paper.

WalkVLM:Aid Visually Impaired People Walking by Vision Language Model MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:19.062819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:19.062819Z digest=sha256:546d045fe7c90c26830a49afdcebabfbb0a3a00fc64ba51924fc3b0a01a942d2

Observation 4f54d04a-5df7-454b-8280-6272c16950d3 · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:33:26.818046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:04552fc4c2b6a3aee9925badd375e12ff238ec520b11fb128203c9af9a9410b4

Observation 66db1a8e-0e96-4913-b571-c44d84f7ba0f · inbound

CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval cites this paper.

CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:00.992916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:00.992916Z digest=sha256:8569ae1285566d874c6824a585ac2206cadda382d648036fd862e80acf10fa2a

Observation 0df0f12a-c6bb-45d4-98c9-88d37be09f77 · inbound

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining cites this paper.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.635688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.635688Z digest=sha256:b638be1ce3cc0dd4195cc0fb2c6bd441c71f0e9259a1f39dd7d495b7bcb8857c

Observation e9e8297b-fb45-4ee4-8ef3-3a56a86f5e7e · inbound

CultureVLM: Characterizing and Improving Cultural Understanding of Vision-Language Models for over 100 Countries cites this paper.

CultureVLM: Characterizing and Improving Cultural Understanding of Vision-Language Models for over 100 Countries MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T22:36:03.847783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:36:03.847783Z digest=sha256:df4c5899bfd384442f540b4e4904f0e5f48f6105f715cada062df2da73dbf462

Observation 8e6cc233-c205-4650-a74a-1c5dccec854c · inbound

GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models cites this paper.

GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 111

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:55.650893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:55.650893Z digest=sha256:c69a69904f3f5518e22a72711e9a84f65fb1f45b3785b6dcfa57cb905e7a8403

Observation 8e2a0a33-88e0-4f27-80fb-af682b32c4ba · inbound

FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models cites this paper.

FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T23:09:25.131989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:09:25.131989Z digest=sha256:c3baf8929fa4a2c04436bedc8b2ddc1bbcdf03e9391dea208d9f166be9c77246

Observation 70e625b0-98a3-4bc0-b447-b3a6e47742e2 · inbound

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models cites this paper.

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-23T05:45:28.384333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:44:31.546843Z digest=sha256:042c152a07c7cbda7fc8820a2e2ac696b70a6aba116736c1d4309f8d6575aed9

Observation fb206a06-e5a0-4e86-87c7-d27abfecece4 · inbound

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs cites this paper.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.337149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.337149Z digest=sha256:720468f4509ad195c4f09caccae42e8f42ccfbf85e6eaef7f353574de7d82774

Observation 6f65249d-b997-44d2-a874-e22410aced12 · inbound

DriVLM: Domain Adaptation of Vision-Language Models in Autonomous Driving cites this paper.

DriVLM: Domain Adaptation of Vision-Language Models in Autonomous Driving MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T21:22:19.885529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:22:19.885529Z digest=sha256:a79d35657e0c84eecc19c9fdd23013e9ed66483251acb2322afe849af4fd9dea

Observation 5e4a0f9a-7f97-42a2-893c-ebb612184070 · inbound

Efficiently Serving Large Multimodal Models Using EPD Disaggregation cites this paper.

Efficiently Serving Large Multimodal Models Using EPD Disaggregation MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T04:29:34.801288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:29:34.801288Z digest=sha256:d184c105bbb0b2f600962bd2c34222695eacf6e037d7cbacf3313d58d7b184cb

Observation 229b63c0-920d-4bf5-8e54-ef2870373f59 · inbound

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models cites this paper.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.688619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.688619Z digest=sha256:696a45c4d4edd14bb550f601e8c4c77a6a327c2f88ece6a612a1e7fbbfabe970

Observation 45a1ae2e-43c5-431f-93d3-48045b7b4696 · inbound

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs cites this paper.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T21:11:43.373939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:11:43.373939Z digest=sha256:ae8682c451dccd98eb72159512198f527e4ade36ea8a6ba2735bdfea9bf19694

Observation de5346b8-f4ac-4498-8ad4-89cf841e88fb · inbound

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design cites this paper.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.228951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.228951Z digest=sha256:0a9865c82bd662562488bd1d73bcd9ff3b355cc5ae09b4a737ee94f5fe8f4d23

Observation 2bf7c2dc-2048-4587-9c06-dfb079dabed5 · inbound

ChartCoder: Advancing Multimodal Large Language Model for Chart-to-Code Generation cites this paper.

ChartCoder: Advancing Multimodal Large Language Model for Chart-to-Code Generation MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T21:03:45.405848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:03:45.405848Z digest=sha256:165d157030be105d8b461951cce74c6f515f8499faf0d5b14b710ca221618491

Observation 95c19102-a797-4ee6-907a-46d3e3319f82 · inbound

MiniRAG: Towards Extremely Simple Retrieval-Augmented Generation cites this paper.

MiniRAG: Towards Extremely Simple Retrieval-Augmented Generation MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T20:57:12.889393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:57:12.889393Z digest=sha256:fba1c09994c1785af7a81eb80a6dcd37e4593728bf533fa1084247c5260df4cf

Observation abdd08d9-1c36-4410-8b3f-a3c06f7b886e · inbound

Zero-Shot Scene Understanding for Automatic Target Recognition Using Large Vision-Language Models cites this paper.

Zero-Shot Scene Understanding for Automatic Target Recognition Using Large Vision-Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T20:46:17.656050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:46:17.656050Z digest=sha256:1aa6ff77cf1d901f49d390d7b80e57aaa511c3c20f17155ff4292d57f0839163

Observation 6299dc3c-d1e6-4c66-a861-d7ea2db71293 · inbound

MSTS: A Multimodal Safety Test Suite for Vision-Language Models cites this paper.

MSTS: A Multimodal Safety Test Suite for Vision-Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T19:27:42.350079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:27:42.350079Z digest=sha256:e4eca6da48b9958951302684eb1902511ea4f0af6c03f7115dcb14a7d17e52e6

Observation 372eac7f-9528-4376-86e2-70e689fc2212 · inbound

EndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery cites this paper.

EndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-10T18:24:42.306289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:24:42.306289Z digest=sha256:5c02322e2a0930aac3214f7f629d1b6fb06688454063dd305cf3a7fc47b506de

Observation 6510ff94-4e0b-4955-ab29-7a8d8f84d221 · inbound

Does Table Source Matter? Benchmarking and Improving Multimodal Scientific Table Understanding and Reasoning cites this paper.

Does Table Source Matter? Benchmarking and Improving Multimodal Scientific Table Understanding and Reasoning MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T16:34:40.479177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:34:40.479177Z digest=sha256:f47c1a39b4e61d9a9506213cb76e79a549fca8a980a2d0bd5c493a5dad144ba1

Observation b794a00e-4b7d-40ad-af16-85c7f9576a09 · inbound

Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge cites this paper.

Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T15:58:24.562366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:58:24.562366Z digest=sha256:7f0f70e8094f3342d98501990cb34724c7b7f9fcc0e29eee6591b56079190571

Observation 7f1bde46-37a7-46b4-92ce-81e5ab375542 · inbound

Temporal Preference Optimization for Long-Form Video Understanding cites this paper.

Temporal Preference Optimization for Long-Form Video Understanding MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.284542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.284542Z digest=sha256:4b14701de9a1453b49554597d0e7db647b9b9ba6e6628638acd89def36459823

Observation 5bf02c69-1611-4eb0-888a-7a11ed03c12f · inbound

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models cites this paper.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T15:18:49.949655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:18:49.949655Z digest=sha256:d06ee02202ebafdfac834ca9da1e94af5a95cdcc252153f5c1680e15768b82c8

Observation dcdd8519-0379-4628-8d66-9235d10f0d7c · inbound

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models cites this paper.

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 194

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:34.846395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:34.846395Z digest=sha256:aba4643b5101953284bb968fda24d5459e13774f1900b8885f49d81514ea2be3

Observation a700a717-608f-4fbb-b18d-4bd2b1c7d44b · inbound

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models cites this paper.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:24.523145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:24.523145Z digest=sha256:f9f5b17e918b95b12e63d39f68c8108588a583745ba9b026703348eb0ac192fb

Observation f2aa5815-5ce0-48b4-927c-89cb709f28f6 · inbound

A Causality-aware Paradigm for Evaluating Creativity of Multimodal Large Language Models cites this paper.

A Causality-aware Paradigm for Evaluating Creativity of Multimodal Large Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:25.886778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:25.886778Z digest=sha256:5a5a063289f74e3927d2640f32717202ad7c6ec0a4bbaeafa251e07bb57eda00

Observation abd0c2c8-d85a-43af-b746-4a79d91a49a0 · inbound

Generating Negative Samples for Multi-Modal Recommendation cites this paper.

Generating Negative Samples for Multi-Modal Recommendation MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T14:36:53.089817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:36:53.089817Z digest=sha256:2c05b9fea09ecfff7051da021b14e8f15b5c6b32f25f62bd3425110e743904d4

Observation 54c3e1e7-ad50-4283-b2a0-586d384c5a29 · inbound

Ocean-OCR: Towards General OCR Application via a Vision-Language Model cites this paper.

Ocean-OCR: Towards General OCR Application via a Vision-Language Model MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T14:14:54.793847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:14:54.793847Z digest=sha256:7b2409c5e19129c1d9cce84d4b683a179f43472a845a9d9970f8cae0c413f77a

Observation f972dcc8-9438-43fa-8503-ac71b087c2a5 · inbound

Can Multimodal Large Language Models be Guided to Improve Industrial Anomaly Detection? cites this paper.

Can Multimodal Large Language Models be Guided to Improve Industrial Anomaly Detection? MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T14:03:31.547592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:03:31.547592Z digest=sha256:1f23e11a67cad0eb1ee5a3adfcf741bf946d12be577535a3ec6692e2eb094281

Observation 9bd10117-1474-4938-a48d-c544a1805316 · inbound

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers cites this paper.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.875541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.875541Z digest=sha256:01303f643f688ca16dad76ecef3013ed1449626572105b8af88a7fe3a0ec7409

Observation 7abc8490-a37f-48ea-8e9a-50f82133dae3 · inbound

TAID: Temporally Adaptive Interpolated Distillation for Efficient Knowledge Transfer in Language Models cites this paper.

TAID: Temporally Adaptive Interpolated Distillation for Efficient Knowledge Transfer in Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T05:35:54.126695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:35:54.126695Z digest=sha256:9ca8c534840bd6b9cb62c92d04ee0386df98f0d98087298a1776925707823b9c

Observation 6edbb00f-43ea-488c-9d79-e57a745fb3d2 · inbound

Scaling Inference-Efficient Language Models cites this paper.

Scaling Inference-Efficient Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:29.271945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:29.271945Z digest=sha256:a9a051ce24d50850da5e30dca98fc2c41a2d4ae6808f7eeb386e1da18025e312

Observation b2996d25-593c-413e-affc-24f4c888d332 · inbound

BounTCHA: A CAPTCHA Utilizing Boundary Identification in Guided Generative AI-extended Videos cites this paper.

BounTCHA: A CAPTCHA Utilizing Boundary Identification in Guided Generative AI-extended Videos MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-09T23:03:34.685227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T23:03:34.685227Z digest=sha256:9e13d15f364933a99bc5e1e01e0ad5bdd3d6977e78c09e6aceddff398e7dd8df

Observation b1a9de48-f850-4970-9e30-ba437f924d71 · inbound

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs cites this paper.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.824073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.824073Z digest=sha256:f05e188f275b73b0f91b9455f4e0d11055b0b031c7f00373452881df92a3a93b

Observation 74188231-6d98-474c-82a7-f5e46f1afd07 · inbound

MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization cites this paper.

MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-09T19:12:56.281848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:12:56.281848Z digest=sha256:0cd9560876bed148d835abba43f9ca6e92e346bfac5453d75d4308b4f9d40767

Observation 85737c1c-4f44-4ac8-b28b-184f8961b33b · inbound

VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos cites this paper.

VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:36.909300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:36.909300Z digest=sha256:ac313ce3ae0fc3e2a66476c1ac7156971e2caab7a5d5379defb8f618b78e5f59

Observation 1e41759d-1964-4f45-a720-32c7192f653a · inbound

Ola: Pushing the Frontiers of Omni-Modal Language Model cites this paper.

Ola: Pushing the Frontiers of Omni-Modal Language Model MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.349659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.349659Z digest=sha256:d124eae215c7278ae9a647ce9f702ad1858d505ae016ba041c98348c39827cd4

Observation 05bd489a-113c-4044-a4e5-e7baa155387c · inbound

Cached Multi-Lora Composition for Multi-Concept Image Generation cites this paper.

Cached Multi-Lora Composition for Multi-Concept Image Generation MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T21:03:24.029731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T21:03:24.029731Z digest=sha256:98bcd6fe87031df1dc1a626487f00843abfa0934bb804982ec50d0003d3491c0

Observation 91866d10-44df-4a43-9712-ecb46efe9795 · inbound

Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs cites this paper.

Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T20:19:49.339201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:19:49.339201Z digest=sha256:5e3c286729067354a4ce55686cbc1ca88cfe2b5a9ad1539cdf7b5ea5cae983ac

Observation 5b816710-c759-41c7-b24f-d4ea4f6e995d · inbound

HCMRM: A High-Consistency Multimodal Relevance Model for Search Ads cites this paper.

HCMRM: A High-Consistency Multimodal Relevance Model for Search Ads MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.379507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.379507Z digest=sha256:03f9837583881dc3d55e0509eae673cd60f1019d789fd7d154a9e8517bc17ae2

Observation 831adb08-f292-4032-9ce2-52ce6f47f6c8 · inbound

Diffusion Instruction Tuning cites this paper.

Diffusion Instruction Tuning MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.930797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.930797Z digest=sha256:46bde54e1ba0598bff178e7d5bb678d9a40e7f2eb84b43f318dd1dd7515e7164

Observation 028fa900-a965-4130-b355-e1a364996573 · inbound

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering cites this paper.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.359154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.359154Z digest=sha256:2a246d5b802f7797c316f3f45df9d93ee2791f4a6ae239e0079f8751ed3e5c02

Observation 47909c14-4047-42c1-93b0-5f87e0fb7fe5 · inbound

Seeing is Understanding: Unlocking Causal Attention into Modality-Mutual Attention for Multimodal LLMs cites this paper.

Seeing is Understanding: Unlocking Causal Attention into Modality-Mutual Attention for Multimodal LLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-23T01:25:16.455129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T01:23:01.892612Z digest=sha256:d0db867b8dd4449c59edcccc2c74bbde95d71a04617fa8330a9b3b018bd914c4

Observation 5da04e2d-232c-48f3-bd4b-2462588b8b30 · inbound

R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization cites this paper.

R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-16T00:19:20.532886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T00:19:20.462455Z digest=sha256:b7303a90fd702ab10054237d99668ef945848f41601c79b8519c356cca787688

Observation 3bdea0e8-8c80-4a5f-bac3-c0bf25f9f2b8 · inbound

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization cites this paper.

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:04:22.780495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T15:04:22.690503Z digest=sha256:c986d0e6a98e0445521efe8f2fb5b2e00071ecf455c8da61700c85a202cb5f51

Observation 10ff2e47-624b-4e06-b7f6-4ebea279a1a0 · inbound

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models cites this paper.

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-23T00:02:17.774286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:58:57.819555Z digest=sha256:abfafdc797a4cc755d9253ef354eca626a09cb346021111eec03cc53f3327997

Observation 8d035bdb-0cb2-43e2-b338-88dafefc0bd9 · inbound

SpaceR: Reinforcing MLLMs in Video Spatial Reasoning cites this paper.

SpaceR: Reinforcing MLLMs in Video Spatial Reasoning MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-15T15:18:43.853359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T15:18:43.724432Z digest=sha256:5f92c80c9b4ad96f87bfcff0f73a805c9eaff98dfdaf19d7f212132c722ad625

Observation b0837259-43ee-4398-a1a0-e5fd01681b30 · inbound

SmolVLM: Redefining small and efficient multimodal models cites this paper.

SmolVLM: Redefining small and efficient multimodal models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 36

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T20:23:51.790378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T20:23:50.552549Z digest=sha256:75e895bcfd543d23f16656e6225821945c24e3283f896f11eb96646375a01904

Observation feae0bd9-8280-4765-92de-db8f78a2457a · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 136

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:07:32.698795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:5f1ed25b84d0accedb0a6bd3de2bbdb063d32c0b82ba1dbd84bc1c115b7afc7e

Observation 497fbbe7-2723-44e2-9781-ae5f450fd426 · inbound

SkyReels-V2: Infinite-length Film Generative Model cites this paper.

SkyReels-V2: Infinite-length Film Generative Model MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:23:04.160478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:23:04.022599Z digest=sha256:65873f830ec73a349b1085c91c0f8ae9b7ad1f848b9766cc1eceb003db191ed8

Observation 7318a428-6471-4495-abf4-707192b9afda · inbound

Social Human Robot Embodied Conversation (SHREC) Dataset: Benchmarking Foundational Models' Social Reasoning cites this paper.

Social Human Robot Embodied Conversation (SHREC) Dataset: Benchmarking Foundational Models' Social Reasoning MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-22T21:15:09.349994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T21:14:13.351140Z digest=sha256:824a751b23ed9596cc45ec661c7e824b1dedf091575a6f8c42314fdb9d1f360e

Observation 4c5624b1-b94f-4b5f-b519-a962239c7d6d · inbound

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting cites this paper.

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:30.405082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:30.405082Z digest=sha256:06b379ded4fe0b7c8778f67af1a7a2b7e5083faa40cc7c4571efdf9e5975ae18

Observation ae9036c9-f303-4130-80b5-698248930b2a · inbound

UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning cites this paper.

UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:56.038474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:56.038474Z digest=sha256:2378d5740b91e9eec4cfe632f0f05895b2ab6e8792b5a691ee2937abc7cdd875

Observation 1847bc0e-6b6c-46ae-b58f-46a59b596cdb · inbound

S2SBench: A Benchmark for Quantifying Intelligence Degradation in Speech-to-Speech Large Language Models cites this paper.

S2SBench: A Benchmark for Quantifying Intelligence Degradation in Speech-to-Speech Large Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:37.759717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:37.759717Z digest=sha256:7e0eef5f5e029a8285f9838853483d370a51448c7fa10906a3e404b1670b0795

Observation dcf3998a-0299-44c7-99b1-60537273aa35 · inbound

Clapper: Compact Learning and Video Representation in VLMs cites this paper.

Clapper: Compact Learning and Video Representation in VLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:47.766618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:47.766618Z digest=sha256:ab8784729a51c4e16a0d2293091c59dcbdf29368eca076ff9070607a91b8da96

Observation 901e1fc9-dabd-4011-8231-b168b2820735 · inbound

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs cites this paper.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.281622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.281622Z digest=sha256:8a4f0660a77883e05d1e675ef13d78cd70612c31c7fe5f7954126b5ea84088e7

Observation 6fa76702-21c1-4f01-8b2a-ed31e90f6b6f · inbound

PhyX: Does Your Model Have the "Wits" for Physical Reasoning? cites this paper.

PhyX: Does Your Model Have the "Wits" for Physical Reasoning? MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:14:54.840336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:14:54.840336Z digest=sha256:c3863a9fd103db3fcfb80d5cb718acaf3ee32b6e55eabea44dd05876b21bc456

Observation 1866e258-2d55-448d-b56c-dbcb34d455ad · inbound

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models cites this paper.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:21:39.814756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:4f21b61ae331432f3b946c9283ac3429cce78a9c7a5688b1c16d4f80488f294a

Observation 565bbe37-6430-40e2-af36-f08aad36afa9 · inbound

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning cites this paper.

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:26.814318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:02:26.814318Z digest=sha256:2e184837e63caa81bfc2b1bbf9575f8e79eba8f3f1d806f01e6976994a280148

Observation 6f1abc63-c866-44a5-9634-c478b55cc4a1 · inbound

R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO cites this paper.

R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:01:33.558215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:01:33.558215Z digest=sha256:011bbf8c3febbbbcf158cd11d0b3fc181c847d327c976b9b153c29bd9f5ab882

Observation 40a23375-22c8-4ce7-9b7a-86643a55d6ea · inbound

RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs cites this paper.

RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:43.703291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:43.703291Z digest=sha256:5344f2e3db1e3734b47807515db55aa2c79db309d2c7b19d9ddd739e877d68b4

Observation f60da05b-fea0-46ac-900d-b30b5ee232fa · inbound

From EduVisBench to EduVisAgent: A Benchmark and Multi-Agent Framework for Reasoning-Driven Pedagogical Visualization cites this paper.

From EduVisBench to EduVisAgent: A Benchmark and Multi-Agent Framework for Reasoning-Driven Pedagogical Visualization MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:43.483072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:57:43.483072Z digest=sha256:f497395e7020033d1bb4f13c9d3653e7ccf557c8e379f0caa5bd8c05b86c51ab

Observation 602428fb-e95d-4e4b-966f-52d63434c168 · inbound

EVADE-Bench: Multimodal Benchmark for Evaluating and Enhancing Evasive Content Detection cites this paper.

EVADE-Bench: Multimodal Benchmark for Evaluating and Enhancing Evasive Content Detection MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:05.858502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:05.858502Z digest=sha256:336d324c1a71865f44c55890387370a0bf1d46c07e36dfd6f5178973b162c699

Observation cd735611-0f77-43ff-9c69-1e1ee27c4c29 · inbound

Generative RLHF-V: Learning Principles from Multi-modal Human Preference cites this paper.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:52.093446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:52.093446Z digest=sha256:225029b4937fffd6efc95705d317a4f4146c13b1c58a55d7843df27709029733

Observation ad36b43e-9f5e-4e51-8c2c-75c1979abe93 · inbound

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models cites this paper.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:40.650777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:40.650777Z digest=sha256:daf3ad9918393be705bc43b892cfc56d0d44c67818b1757feccf019f18f535a3

Observation 0b5c142e-69e9-4709-91ce-7bed037176bb · inbound

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval cites this paper.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.015018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.015018Z digest=sha256:793aea2832e5e7550c0312421d999aaefd74ae15e7b9628c171197cd77ac38a1

Observation 8a0b16a1-b190-4077-8ad9-53da8bd22410 · inbound

MLLM-Guided VLM Fine-Tuning with Joint Inference for Zero-Shot Composed Image Retrieval cites this paper.

MLLM-Guided VLM Fine-Tuning with Joint Inference for Zero-Shot Composed Image Retrieval MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:35.622257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:35.622257Z digest=sha256:e187338809540f019f8b1689262bfa8d105e82e3e3e3560423948d8f14dbeb0e

Observation 5dff226c-658e-48f7-9170-acd6bdca4573 · inbound

MLLM-Guided VLM Fine-Tuning with Joint Inference for Zero-Shot Composed Image Retrieval cites this paper.

MLLM-Guided VLM Fine-Tuning with Joint Inference for Zero-Shot Composed Image Retrieval MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:35.711208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:35.711208Z digest=sha256:c9eca71fdf8372e2271b3a54acbff81f2b1bb40312c31daa17a66045675d4a25

Observation 11038ea0-4bce-4a62-bb5f-b711dbd1a21b · inbound

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval cites this paper.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:37.755205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:37.755205Z digest=sha256:17ab2a36391d8780e42c694b4fde571a0916cfd728ad82efa50a45d9fa50014e

Observation 84835a9a-54c1-4b11-9ecc-8ea290504b13 · inbound

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval cites this paper.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:37.813610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:37.813610Z digest=sha256:a7341b0c3d49e750a50fac24ffe673ed101ce232ccc9acb873931cfb837d3e66

Observation 07fee8e3-fcbb-424e-90ab-64d20723cb18 · inbound

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos cites this paper.

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:04.827954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:03:04.827954Z digest=sha256:7dddcc115178d319e8b026bcfa68e2bd0702c38884caebd71b6bb6fcfcfcbd8a

Observation d14e2908-41ad-4ccf-8550-d93ac69167d0 · inbound

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration cites this paper.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:51.005221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:51.005221Z digest=sha256:c9046f09c9a825923979a25ae04e32da72773996130d0213b63a0fa28880c865

Observation 019551d9-c68a-4bf7-acad-7e23fd2cbdf4 · inbound

TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs cites this paper.

TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:32.078030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:51:32.078030Z digest=sha256:02be39a96288c6d788939a5d276edb3b1ae9ff81f2c6f94fb9da66ee0b47a7b3

Observation c5336192-b0dd-44e1-8dea-3ac387e47985 · inbound

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs cites this paper.

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T13:40:25.336819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:40:25.336819Z digest=sha256:508e0b6b877210b3a633bae67af21bda981e6505cd1a3dc715bd0f6956074a6c

Observation 15b75bad-1466-45a4-8766-15d94b277c09 · inbound

OmniAD: Detect and Understand Industrial Anomaly via Multimodal Reasoning cites this paper.

OmniAD: Detect and Understand Industrial Anomaly via Multimodal Reasoning MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:03.043936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:03.043936Z digest=sha256:b97a67bf03529464dc7c57a2910eb74ab798982e3d4c9f52605b43b68282d5b6

Observation bb2908f2-584b-4cca-a846-e6ff12b811b4 · inbound

VModA: An Effective Framework for Adaptive NSFW Image Moderation cites this paper.

VModA: An Effective Framework for Adaptive NSFW Image Moderation MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:46.178092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:46.178092Z digest=sha256:e3a1fb2c64b7e8ace3aaa9f93265c6c851752585a0002d2bac9292352e59c434

Observation 1eeb526c-3259-4c60-8d94-077fd5debf1a · inbound

Qwen Look Again: Guiding Vision-Language Reasoning Models to Re-attention Visual Information cites this paper.

Qwen Look Again: Guiding Vision-Language Reasoning Models to Re-attention Visual Information MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:47:01.136361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:47:01.136361Z digest=sha256:eaf68ce843a3fbf69d4c4c559c8b132812ca75d2faa8a166f12e1bf9faa0b72e

Observation dcb64f2c-9b35-48e2-8990-2a95b70739df · inbound

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models cites this paper.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:16.297715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:16.297715Z digest=sha256:b3a50762f1fdf843fc56a504465eb0ff575e0941651a1b3cccda192bdc1da171

Observation e9d53f51-3628-4ccd-8161-bafbb69f4d9a · inbound

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders cites this paper.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:23.584263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:23.584263Z digest=sha256:1b0ddf5a5326b9692e26e3096cdf18663cdf432941b2aef8fede822fcb9d21e5

Observation 575783b0-dd21-4941-b5eb-464ca8c32aea · inbound

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models cites this paper.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:49.453135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:49.453135Z digest=sha256:5f42e8baff06761a2712959ab4ab00dc2871a8e6c850cd8515aa36eaf3981401

Observation a1c6c84f-88c8-4d7f-876c-e230f0f6b518 · inbound

DisTime: Distribution-based Time Representation for Video Large Language Models cites this paper.

DisTime: Distribution-based Time Representation for Video Large Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:54.728045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:54.728045Z digest=sha256:f49ba76317fde1efa39be231272c5a890bca84a163fb1584d496969cf805a43c

Observation 8e7525b8-2c27-4704-bc89-1698c2ce5cd5 · inbound

Mixpert: Mitigating Multimodal Learning Conflicts with Efficient Mixture-of-Vision-Experts cites this paper.

Mixpert: Mitigating Multimodal Learning Conflicts with Efficient Mixture-of-Vision-Experts MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:25:08.460726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:25:08.460726Z digest=sha256:99efa4f12559187ec8ae41bb11d925821eab5bad7352566154c9f60badf1ec83

Observation 8944cc7b-10b1-4c98-b1bf-5159dde0e1f4 · inbound

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces cites this paper.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:15.718768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:15.718768Z digest=sha256:cde915ee248469a1677268a3a8b1d75b4c0e8c5f09db2722874bdafebd4d1334

Observation cd372c72-0a39-4af6-b4f3-f4bd6e731079 · inbound

EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models cites this paper.

EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:10.058321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:09:10.058321Z digest=sha256:ad995db3cea3751a367c0efa7ceb71094bae77172831b9db0a15176bf1d28c20

Observation f8480938-50d6-4fad-828f-f38f3bd9edf5 · inbound

GuessBench: Sensemaking Multimodal Creativity in the Wild cites this paper.

GuessBench: Sensemaking Multimodal Creativity in the Wild MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:41.923083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:41.923083Z digest=sha256:cb5032821b01744796348fa80f01975983ba7af40c118a356ed6b3928c602de3