Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T21:27:52.652519Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 2 inbound Pith citation observations for arXiv:2412.04424.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T21:27:52.652519Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T13:19:23.116337Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-22T19:52:01.828238Z
49 of 49 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b956506a-982a-431f-8170-27aaaa06b663 · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73d66cdb-0fd9-4603-88fb-82d5e93f9db0 · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3dbb1f8-5959-44b1-bc98-6ab0b1857bba · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond, 2023
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e545521f-1398-4fcb-acaa-3c8b045d2899 · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c2634cb1-deab-4a63-af2c-963f94e4c157 · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Sharegpt4v: Improving large multi-modal models with better captions,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bbbe195-f89a-480e-8f8b-41904477a200 · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 705a2024-bc8b-4132-88a3-a692ecf4a82a · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9414ee1a-b5f6-4f34-8654-3c2d883cd8f5 · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion RedCaps: web-curated image-text data created by the people, for the people
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecdf898e-0441-4504-ada4-9e1fbcfa386e · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Davit: Dual attention vision transform- ers
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65923e19-8764-488a-8f45-b407ced021d8 · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion MouSi: Poly-Visual-Expert Vision-Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d85dddf7-2316-4c01-aef1-2e9e7880a900 · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2024
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9949e921-a402-44e7-8569-72e9ea923183 · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion 1 OCRBench ChartQA DocVQA InfoVQA Average Florence-VL 7B 41.4 24.3 44.5 29.4 34.9 OCR 40.9 22.9 44.4 29.0 34.2 (a) Ablation study on OCR features on OCR & Chart benchmark
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2a861900-2d06-4a04-aac0-5dff573460da · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1275089c-6f36-4479-b281-f93268b4aadc · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Vizwiz grand challenge: Answering visual questions from blind people
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97833a46-8d04-47f4-9ce4-7e47936c38dd · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9057387d-297f-4da6-b097-d93f084b186b · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Gqa: A new dataset for real-world visual reasoning and composi- tional question answering
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation be1c5739-1cc9-4fc8-a3d4-d63acc3e9d10 · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion https://huggingface.co/datasets/huggingfacem4/docmatix
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 25e88791-2850-4d96-abf5-cb34e4ed394a · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion BRAVE: Broadening the visual encoding of vision-language models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c81180e8-05f8-4fdf-b406-7af522edb8a9 · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion A diagram is worth a dozen images
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69d6e750-e170-4d16-b40f-58e33aa5b120 · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Segment anything
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c18cdb79-ee2e-4d68-80ad-9c318057e82e · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c185393-b516-4b8f-a95a-a7f5e050ab82 · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Evaluating Object Hallucination in Large Vision-Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7ada76e-a126-41bf-8e2d-2544309c0542 · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 585e9a9a-c40d-424c-af9c-6471ea6713dc · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Vila: On pre-training for visual language models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f82b0b6d-6db2-4632-b1e7-c07ee3860d84 · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0405be2a-9cb1-41b3-9e90-e9d964aad5f3 · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Visual instruction tuning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation df400d4d-2b78-483d-a1cd-3d232cc529a9 · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion MMBench: Is Your Multi-modal Model an All-around Player?
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66be9e77-1902-4cae-8bfa-ad0a2649ecd9 · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion On the hidden mystery of ocr in large multimodal models, 2024
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b4cc4588-c9cb-406e-a003-5ad3c972f9f1 · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7882e73a-5805-4a99-a5e8-d8f1de9d942d · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03014ae0-8ff4-4404-b100-0b1158232468 · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b97a60d-041a-4f64-84d2-d8ab627bcdff · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Docvqa: A dataset for vqa on document images
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 47f88643-2ab5-451e-a260-1f0eae8f096d · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Infographicvqa
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a398bc26-f685-401e-bcd0-73a7cbafa363 · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion DINOv2: Learning Robust Visual Features without Supervision
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4943f97-7c2e-4cbd-808e-d00c1d060b6f · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Learning transferable visual models from natural language supervi- sion
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b88d5a6c-e5f9-46c9-ace0-dc8f350875f1 · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion High-resolution image synthesis with latent diffusion models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 71e6ccaf-f75e-4a55-97a5-ec119f8c7d18 · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion High-resolution image synthesis with latent diffusion models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09010bec-878a-4b04-b1d3-177e846291a6 · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d7b7ede-4e70-42f2-9375-2c3f2253565c · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Towards vqa models that can read
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a3a990d-15cd-4897-814f-c08f30a6125b · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion From pixels to prose: A large dataset of dense image cap- tions, 2024
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ae169a02-3cd8-42c9-8da2-c956e3896f11 · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 791cc9f6-8b4e-42c2-a0d9-998d402dfeb2 · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Eyes wide shut? exploring the visual shortcomings of multimodal llms
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 539ee10d-03e0-4204-8786-9fd0a58bf643 · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Diff-eRank: A Novel Rank-Based Metric for Evaluating Large Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50447d1f-36a5-4257-859b-53a1eb9bf355 · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Grok 1.5v: The next generation of ai
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 57ac3f18-f882-475b-b822-65dc159353f6 · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Florence-2: Advancing a unified representation for a variety of vision tasks
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 297f77cb-0bbb-4f14-a6e5-a1b75c730c1d · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Vision-flan: Scaling human-labeled tasks in visual instruc- tion tuning, 2024
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 92bde076-13c4-43fb-b283-aa4ef0b83d90 · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0b8e686-718e-4e5b-9b77-3f0029a142b7 · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 981cc79f-9a2d-4088-ab3c-0344007fd948 · outbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64004760-0a21-4965-9e0b-4fc9dcc32781 · inbound
FastVLM: Efficient Vision Encoding for Vision Language Models Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 218253da-2784-4767-9d0b-c3607ba54cfd · inbound
FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.