Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T10:49:08.294410Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 100 of 136 outbound references and 0 inbound Pith citation observations for arXiv:2412.16158.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T10:49:08.294410Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 136 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cd6c5f04-9712-41e3-be0f-31ceb933e1b2 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ba4b669-ff7f-467a-a2a1-414704687d86 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a5bb838-0f25-4d62-b1c5-f868ddb95507 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d71830f1-b684-4fb7-ac77-5feae7eb21f5 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding The claude 3 model family: Opus, sonnet, haiku
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 219df7c9-9f6a-4c83-aee1-bc117e7de5eb · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c79ec08e-c8a9-4a92-aa7a-5a39306077c5 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Introducing our multimodal models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3084111e-c871-446c-8d82-fbb750e5f9cb · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Vqa-med: Overview of the medical visual question answering task at imageclef 2019
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a365c0e-44d3-4631-8276-f6568993093e · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding PaliGemma: A versatile 3B VLM for transfer
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f20bedfa-9a3a-4066-a5d7-cfb43aac0a9d · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Scene text visual question answering
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3764b123-8dc2-4337-b946-525af5df5b4b · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Coyo-700m: Image-text pair dataset
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7496fd43-beee-455c-b42b-bcf3dacea6c2 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding InternLM2 Technical Report
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ab6b2d0-258b-4220-9d72-5ac19bca1a47 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding An augmented benchmark dataset for geometric question answering through dual parallel text encoding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ebc1937-cb48-4f23-8e8a-2ba54d57d6e3 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Textocr-gpt4v
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 553352aa-705f-41c3-a5af-9da1daecb769 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding MapQA: A Dataset for Question Answering on Choropleth Maps
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2903dba2-41f0-4bf7-b35d-9574f3f3fa5d · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a08d26df-b966-4971-bc91-a8c4295d12c0 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d90f477-d841-46ea-a979-e7e2bc1a52c2 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b1b00db-8e89-435d-aa31-4194798f1758 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding SOLO: A Single Transformer for Scalable Vision-Language Modeling
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb26d3d4-69e5-474e-9828-6aec5db85a51 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9b53d88-4422-42c2-a76f-5ee0b18f2c88 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Internvl2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16c46b68-bfd5-4e09-91ae-f253d8fe9984 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2be1e6d9-7317-492a-96cb-b4d5d3814ce7 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Complicated Table Structure Recognition
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 950614fd-409f-4756-900a-3ae088d9d0ce · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Icdar2019 robust read- ing challenge on arbitrary-shaped text-rrc-art
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e3b402d-5f3c-4763-b0f0-70d90ec52b27 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Simple and Effective Multi-Paragraph Reading Comprehension
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70e0a205-5cea-4259-94c7-09066d0dad3e · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Deep visual template-free form parsing
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09609b9d-e3e4-4981-bc3b-3ab8bd6810fa · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Unveiling Encoder-Free Vision-Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec215a64-3e95-45c2-9d6d-ca8bca9819be · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Compressing visual- linguistic model via knowledge distillation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27dc2e1d-83fc-48c1-8f66-d62f52bf04a6 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efe9e84f-84b3-4cce-9b1d-4a249d151906 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Making the v in vqa matter: El- evating the role of image understanding in visual ques- tion answering
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08586c5c-0b64-44b9-8ad0-57e846085f2d · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd9afb9a-37d4-4826-b452-1dd5b6af07a7 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c35905a8-7175-4d19-b240-07edb1c118f7 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Eaten: Entity-aware attention for sin- gle shot visual text extraction
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91afe247-306f-45d8-aeb2-ef2c6d29aee4 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Don’t stop pretraining: Adapt language models to domains and tasks
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b16ef0f-983d-4d5c-bf77-819bfe5c5843 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Icpr2018 contest on robust read- ing for multi-type web images
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 318b72b3-a1c7-4c86-a704-5608702eef7b · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding PathVQA: 30000+ Questions for Medical Visual Question Answering
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db154994-c89b-4c6e-9c64-cd612c122987 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Koniq-10k: An ecologically valid database for deep learning of blind image quality assessment
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 151511c6-132e-47b9-a5c9-77aed2643d2e · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebf20f45-12ad-4ab3-87a8-d9b266e788b7 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Medical-diff-vqa: a large-scale medical dataset for difference visual question answering on chest x-ray images, 2023
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68dde9d3-8582-44e4-99fe-dfb9e84cf9a7 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Visual program distillation: Distill- ing tools and programmatic reasoning into vision-language models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7796b487-7650-45e4-96c9-996c021f5c68 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Movienet: A holistic dataset for movie under- standing
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b13cf06-d12d-429b-8c5a-50b5583407e0 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Hires-llava: Restor- ing fragmentation input in high-resolution large vision- language models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30cc1899-fe7f-4b2e-a7a2-03cd107798c6 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Icdar2019 com- petition on scanned receipt ocr and information extraction
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbe1ef73-4e59-4946-9d6a-b0194e53a2bc · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6984120-16ce-4e69-a01b-f1b273bf083a · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Egotaskqa: Understanding human tasks in ego- centric videos
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfa1c8fa-d7c2-4353-a39a-2bf1d5c572d1 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28620bdf-2340-4720-9ac1-79731b88a07e · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Dvqa: Understanding data visualizations via ques- tion answering
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03c1d237-a189-48b7-8da5-92294ca0421a · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding FigureQA: An Annotated Figure Dataset for Visual Reasoning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0d0f8f9-24b0-4182-a9ae-f39dc7c599cf · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Chart-to-Text: A Large-Scale Benchmark for Chart Summarization
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80283ccd-36bc-4554-87e9-a114ef6a3331 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Referitgame: Referring to objects in pho- tographs of natural scenes
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a0f73ef-57e7-4010-8415-a8e8d1f296c7 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding A diagram is worth a dozen images
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c279f269-b582-41e2-b927-1c78e92357e1 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Are you smarter than a sixth grader? textbook question answer- ing for multimodal machine comprehension
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7134f05b-beec-48a2-a185-691bae37b15e · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Visual in- formation extraction in the wild: practical dataset and end- to-end solution
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaf396a2-b074-45ee-84c9-f8a36123baee · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Laion-gpt4v dataset
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18371d88-e852-4c69-aaf9-b2848e659812 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding A dataset of clinically generated visual questions and answers about radiology images
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5346708f-c4c5-40d2-a45c-d12753262f82 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Viquae, a dataset for knowledge-based visual question answering about named entities
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ed100a3-649d-4370-a71d-7b87ffceba10 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4151d93a-f332-46eb-9cb1-a6363bf78675 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbe43d26-fab1-471d-9190-0aa7d4622402 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Chemvlm: Exploring the power of multimodal large language models in chemistry area
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a2d965d-ccce-46bb-a5ad-b58730986555 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Mvbench: A comprehensive multi-modal video under- standing benchmark
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e671cbb1-0f09-43cf-a571-468d9ffa7b2a · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Evaluating object hallucination in large vision-language models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4ac0e4e-d25d-4e78-946b-ab047510633d · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Super-clevr: A virtual benchmark to diagnose do- main robustness in visual reasoning
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 010c99f5-cc3b-4171-80de-94cc76d4e775 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Mon- key: Image resolution and text label are important things for large multi-modal models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cf66104-80da-49c7-b15d-24469907c359 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Vila: On pre-training for vi- sual language models
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bda4c47d-4915-4028-8d63-e6bd1a850253 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Microsoft coco: Common objects in context
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fc4dfd6-c238-4b8c-86b9-44ddf1e760a5 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Slake: A semantically-labeled knowledge- enhanced dataset for medical visual question answering
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a2ccfc3-e863-4ee6-99ce-b02541c4840e · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Casia online and offline chinese handwriting databases
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69b4cebf-8ad5-4755-94e7-8ea27cb25f13 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Visual spa- tial reasoning
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9987486b-b6cd-4533-b172-f59f88e1e569 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Mitigating hallucination in large multi-modal models via robust instruction tuning
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e398f10-bccb-4d5f-82ac-6275bcb36b51 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d04af017-bdcf-41d4-b039-4eceb8be0b54 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Improved baselines with visual instruction tuning
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cbafedf-256f-41ba-8c7b-458a19ad09aa · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Llava-next: Im- proved reasoning, ocr, and world knowledge
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f119e9d5-9757-463c-b564-0e4f31c35e43 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Visual instruction tuning
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4da4d5f6-e510-4afe-8194-14131a1f0c54 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db7d8205-78c7-41da-91a1-0ced09a51ec2 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Mmbench: Is your multi- modal model an all-around player? In ECCV, pages 216–
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c9b2023-4620-4bd9-958a-e7b2eaacdf17 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 224d0ffb-4d93-4b43-bedf-6d21db74f05f · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95b58660-03a8-4f0c-b686-16b5209b814b · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40d830c6-af17-45df-b644-0eff71ae5228 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Learn to explain: Multimodal reason- ing via thought chains for science question answering
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f716cc29-6023-4004-ac70-5c53ffd6d5bd · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b48b5cc7-0962-44ff-ab5f-e4054d78f056 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a39fe23-80c4-4757-9ee0-5017b9b623e0 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Deepart: Learning joint representations of visual arts
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2982c736-06af-4659-b210-ceb0321ed8fc · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Ok-vqa: A visual question answering benchmark requiring external knowledge
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e44e2131-99da-4a84-aa6a-7e5a6b5fbb1d · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding The iam-database: an english sentence database for offline handwriting recognition
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23bbdee0-b510-4ad9-8d1d-5ce6f722210a · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65e13ded-6c39-4d97-9faa-2e78de22ff4f · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af3059cb-c264-425a-b98c-2fef2c7f4691 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Docvqa: A dataset for vqa on document images
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5dd9c2f-9d21-4bac-a888-6e7203069433 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Infographicvqa
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcda2877-eaa5-4991-bc3a-709b18ed9086 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f5dba71-6c98-455c-9d5f-7e7985cf1ea8 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Plotqa: Reasoning over scientific plots
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90f291d5-e8bf-4899-9895-3c61536d3e8d · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Localized sym- bolic knowledge distillation for visual commonsense mod- els
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6df6201a-e410-43c9-b86e-7f74c9d7bb18 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fa6a765-a7d3-421e-8900-87353103a565 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Efficiently scaling trans- former inference
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 81d33351-6e77-4c06-afa7-296895af5ae3 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding A dataset for movie description
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 317103cd-3ba8-4db8-817e-448820126bac · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Laion-5b: An open large-scale dataset for train- ing next generation image-text models
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 304cc45e-0549-4418-ae66-886a56839fbd · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Laion coco: 600m synthetic captions from laion2b-en
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ee3c7c9c-d14c-4a16-a9a2-810b9ada4996 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Solving geometry problems: Combining text and diagram interpretation
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 58617fa4-785f-4792-b157-0b563d1883a8 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Kvqa: Knowledge-aware visual question answering
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a3859706-8d68-47a4-83d7-113ab49bce50 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Ntu rgb+d: A large scale dataset for 3d human activity anal- ysis
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 72537121-778e-4cd3-bf3c-1b233fa6f39a · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Objects365: A large-scale, high-quality dataset for object detection
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cd7ca002-83ff-4c81-9a90-310725577f72 · outbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Towards vqa models that can read
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
No inbound Pith citation observations are available.