Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T16:50:05.568237Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 100 of 137 outbound references and 2 inbound Pith citation observations for arXiv:2507.12566.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T16:50:05.568237Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-13T05:12:37.339084Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-13T05:17:18.750827Z
100 of 137 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 235b395d-36c4-4a87-86a8-8c3e66eb72d8 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9774078-184d-4460-83d9-23b4956496b9 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Qwen Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0b86d26-1227-476b-8a16-ea8ab9088e16 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models InternLM2 Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c183d979-6b75-4a40-bd9a-0dafa73d6666 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Learning transferable visual models from natural language supervision,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0901f71-40b2-4641-908e-c290254308de · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Visual instruction tuning,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ae42cdb-16e2-4436-966e-725c985a3d53 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac57a490-84dc-4135-9514-b86f724a33e8 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c7033a1-d62d-4e9a-9189-9f5e541158b9 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Introducing our multimodal models,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31e4b70e-3842-4761-9ffd-58b30ea63fba · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Unveiling Encoder-Free Vision-Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c525bd1-5ec4-4b4f-9005-06ab4929f066 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models SOLO: A Single Transformer for Scalable Vision-Language Modeling
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 23b62d2c-270b-4f8c-8d6d-96db9f90abcd · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Mono-internvl: Pushing the boundaries of monolithic multimodal large language models with endogenous visual pre-training,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4dfeccd-4672-44a9-8db0-229275532d35 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df089f74-f30b-44f3-a10f-2de6df635cdf · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Investigating the Catastrophic Forgetting in Multimodal Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cffd4b2-6780-4103-aa7f-85eca601f8e6 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Delta Tuning: A Comprehensive Study of Parameter Efficient Methods for Pre-trained Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73aee5ed-f683-4926-a808-79f3ecb4aa9b · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93868ee2-dba0-452c-8fd9-2d19ae9d9976 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Lima: Less is more for alignment,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27ad4242-fc85-46d9-a397-83c596813eed · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Emu3: Next-Token Prediction is All You Need
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 901c8ced-76a7-47a9-93e0-abdc860ade0e · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c338e29-89b5-40fe-980e-4969d4907720 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Gemini: A Family of Highly Capable Multimodal Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 220026a1-5a60-44c4-8099-462afadca072 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Improved Baselines with Visual Instruction Tuning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f625dc8-cdfc-42c1-9448-9668a4a65cfd · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc75399e-c724-46e9-8338-dbd706c97401 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 718b4b77-42d6-4f16-baa8-b702afb4b3c8 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5a625a0-61af-40c6-b69e-ed10033673fa · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6ab71c8-d911-40a6-8d89-4ef5fba23175 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Instructblip: Towards general-purpose vision-language models with instruction tuning,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dab453f1-892f-41b9-9259-186e59988c55 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Llava-next: Improved reasoning, ocr, and world knowledge,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f68fdac2-7f06-4dfa-a0b8-6f5c56f104dd · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df391df8-6293-45c1-b131-2b8d35f1e3a7 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Qwen2.5-VL Technical Report
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 202f1580-fc12-439a-8af0-3029d2195b70 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 871361c5-bde1-4422-8f1b-ebb842a6a0b1 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8eba6607-61ef-454b-a467-b2be0c2d9381 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Mini-internvl: a flexible-transfer pocket multi-modal model with 5% parameters and 90% performance,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63a1e2a6-3313-4851-9e28-bdd64ec2e4a3 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models LLaMA: Open and Efficient Foundation Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 730f0235-ff65-47e0-b3e4-822e238e42da · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0244e7a9-46d2-49e3-86a4-682e461f0b46 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models An image is worth 16x16 words: Transformers for image recognition at scale,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e7d78cb-6aef-4d4a-8d55-1fea0f4b1312 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bed0b7ad-391e-4c43-bca4-ac6c35dda569 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dde65db-d26c-4439-b13c-a083fc6b3b8d · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Vlmo: Unified vision-language pre-training with mixture-of-modality-experts,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fe92c26-84cd-40aa-97f4-5e7386379e26 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e3f5af5-d0c2-4d93-89e5-8a1fe41310a5 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Scaling Vision-Language Models with Sparse Mixture of Experts
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5284dae-5a57-4d0b-9d5d-eb743de78905 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Twenty years of mixture of experts,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a5d6262-037d-431e-83b5-cee3b08144d0 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc70db5e-76e7-4e1b-af06-22a874c99c41 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Mixture-of-Depths: Dynamically allocating compute in transformer-based language models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5512c647-879c-48fd-b611-5cd6e6d1ced4 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 855a26ff-2ca3-4c63-a529-07a4d654281f · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Attention is all you need,
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb7f29c9-3b3a-430d-98aa-32245d0eae88 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Root mean square layer normalization,
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cec4baf-2c06-4943-b6c5-51583007169e · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Laion-5b: An open large-scale dataset for training next generation image-text models,
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cdd1300-3cee-46af-9e71-46005c45e6b0 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Coyo-700m: Image-text pair dataset,
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55bf65fe-a0b4-424d-a3f3-aa0da6d66bbf · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Segment Anything
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8d07cf3-fba4-4f8d-9360-55cb2ec9d4a8 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffe36186-4645-44ad-8511-f3eb6f815ab9 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Microsoft COCO Captions: Data Collection and Evaluation Server
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 712be2aa-2719-4a1b-b345-2678d7220474 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Textcaps: A dataset for image captioning with reading comprehension,
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86908a5d-546f-45a5-9e28-ddc473bd2a39 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Objects365: A large-scale, high-quality dataset for object detection,
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 776abb4f-f5de-465b-9e53-745ae1e26aa0 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models The all-seeing project: Towards panoptic visual recognition and understanding of the open world,
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba30b0f4-3e61-476c-b798-e2e44d305b27 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark,
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0b97b9e-f3a1-473d-bb14-7c57e43644a3 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Laion coco: 600m synthetic captions from laion2b-en
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ea5e301-bf72-41f3-ab76-9789c6925236 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2e3ed08-d8f3-43f3-a7ca-f6d5c4d79193 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Icdar 2019 competition on large-scale street view text with partial labeling-rrc-lsvt,
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4912a34c-18da-4fa0-97a8-d82f53569c69 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Scene text visual question answering,
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7493250-004c-4ba9-b345-df73855b3bcb · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Icdar2017 competition on reading chinese text in the wild (rctw-17),
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66f3aa88-0aee-4f2e-bd2b-d4dffbf4ddec · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Icdar 2019 robust reading challenge on reading chinese text on signboard,
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b330ee7f-0072-4b50-895e-b9ee0691b6ea · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Icdar2019 robust reading challenge on arbitrary- shaped text-rrc-art,
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3756c26d-5fd7-4b25-929a-db601c9dbe4f · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Ocr-free document understanding transformer,
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ec19cce-6c0a-4f41-be1e-fd52bd161f61 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c8d6861-01c7-4da6-9a3c-1890fc2068bc · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Chartqa: A benchmark for question answering about charts with visual and logical reasoning,
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61e6366c-974c-407f-8fcb-4b8263ade4fe · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models A large chinese text dataset in the wild,
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d043d554-ecdf-429c-88da-cf42cc6088d7 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Simple and effective multi-paragraph reading comprehension,
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b0c7140-c968-45e5-a601-f6060c883f63 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text,
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07cde5a0-af22-42d9-a8cd-9c43ab5f7c76 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Plotqa: Reasoning over scientific plots,
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c83be487-e7ae-427b-bf7a-9f520ec9f738 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Infographicvqa,
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66675f87-dc47-41fa-b6ab-4ab661215fed · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Making the V in VQA matter: Elevating the role of image understanding in visual question answering,
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 546425af-4490-472f-8619-4c156080f53f · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models GQA: A new dataset for real-world visual reasoning and compositional question answering,
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f900c643-6726-47c9-aa9f-549eb9f80e8f · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Ok-vqa: A visual question answering benchmark requiring external knowledge,
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bc39bd7-4222-4de2-9972-5cd6c418f854 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Visual spatial reasoning,
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9264504-9076-436e-8b1e-bc1716e767a3 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Visual dialog,
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eeba70c5-126f-4a87-bc5f-c70595bee6fd · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models A diagram is worth a dozen images,
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c19cef38-137e-4204-beaa-d73fa8976722 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Learn to explain: Multimodal reasoning via thought chains for science question answering,
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b270660-b3fa-4df9-b4fd-4cbb7a7e6ffd · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension,
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28539a3e-357e-43f1-a2e4-1a229e7f10a5 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Dvqa: Understanding data visualizations via question answering,
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f8cb091-aeeb-40d2-97cc-477d8c8b0a9f · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7382f2c5-4d5a-431e-bb17-b95211031b07 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models An augmented benchmark dataset for geometric question answering through dual parallel text encoding,
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8b413f7-7d5a-45e3-8655-2a7e0c828785 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33c867b7-c587-4172-a55b-d3a85925e403 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65f50570-1ed8-4434-a038-0a613d7da31b · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26aebb84-91f1-4696-a69d-57638124a7e8 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Super-clevr: A virtual benchmark to diagnose domain robustness in visual reasoning,
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3eaa7f89-402a-4e8d-bff3-c1d1320a9647 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9408ddf6-13f9-46d8-9640-f0bca390a041 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Kvqa: Knowledge- aware visual question answering,
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2063723-1016-463b-a2ca-3d5ac2691921 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models A-okvqa: A benchmark for visual question answering using world knowledge,
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 197b7174-653d-4cb5-9496-32b7ee46a3a6 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Viquae, a dataset for knowledge- based visual question answering about named entities,
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d17bb68-d896-410f-921f-b62d7d2c94aa · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cb3cda1-178f-4ead-a61f-9f8600364ab3 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Ocr-vqa: Visual question answering by reading text in images,
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b50417c2-8ed8-42df-b13a-983ad4fd2128 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Towards VQA models that can read,
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 581d54b1-e307-452a-bfe9-6335283eaa70 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Modeling context in referring expressions,
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f850fd4-aafe-4cb7-9acd-c1616efcf848 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Generation and comprehension of unambiguous object descriptions,
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed4402f5-07e4-43c4-bc10-7c2124596bd3 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Visual genome: Connecting language and vision using crowdsourced dense image annotations,
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b970ecbc-9311-4a1b-bd3c-f83feadd74b4 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 244696a9-e355-4869-9b6a-18fa1cf73075 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65664252-de60-4c8d-b58a-736bf1bdcfb6 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Gpt-4v dataset,
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d38fdb5-0f7b-4210-b379-69b149ae4558 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Judging llm-as-a-judge with mt-bench and chatbot arena,
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d86c7702-3db9-44f1-ab53-65377a19e076 · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models SVIT: Scaling up Visual Instruction Tuning
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24e6065c-6e91-4ccd-ade8-cada6d4633bf · outbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants,
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe67a2ae-8c0f-4042-abb3-5246f5c6fee4 · inbound
InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 85ce1315-55db-4859-9f18-b28e63813223 · inbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.