Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T17:56:46.279061Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2412.08378.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T17:56:46.279061Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-19T20:43:16.364125Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-19T20:47:45.997187Z
29 of 29 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f813ec16-4a44-46b2-96c0-1376501fd228 · outbound
FILA: Fine-Grained Vision Language Models InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4909a288-bb1f-4aee-b11f-df4b973d623e · outbound
FILA: Fine-Grained Vision Language Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 888df4e2-6bce-4c69-945a-8a2489a81a2d · outbound
FILA: Fine-Grained Vision Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea4bc805-2138-49c8-b0d4-f514d986f7c1 · outbound
FILA: Fine-Grained Vision Language Models mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3d3e998-e05f-4e26-97dc-c4565a33a5c6 · outbound
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 428fa63d-aa58-43e8-87c2-8b460dfb63c0 · outbound
FILA: Fine-Grained Vision Language Models Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1175c106-796f-429e-8a42-686da30ca90b · outbound
FILA: Fine-Grained Vision Language Models DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ba66138-4fc2-4694-8843-2fba4401a95b · outbound
FILA: Fine-Grained Vision Language Models Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bc0d697-dd1d-4f1d-87fe-7262da9f3897 · outbound
FILA: Fine-Grained Vision Language Models ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d3c4315-3ca8-4715-82c9-2179207d2bc1 · outbound
FILA: Fine-Grained Vision Language Models Learning Transferable Visual Models From Natural Language Supervision
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0970d59-34c6-4972-a6b9-5388cb254a10 · outbound
FILA: Fine-Grained Vision Language Models Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f2e99b39-f197-4592-836e-3a76de13f774 · outbound
FILA: Fine-Grained Vision Language Models LAION-5B: An open large-scale dataset for training next generation image-text models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44f6bc17-8221-4c0d-8520-aafe88bde415 · outbound
FILA: Fine-Grained Vision Language Models LLaMA: Open and Efficient Foundation Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae29ba6e-f9d1-4903-a0f4-529a03f08af7 · outbound
FILA: Fine-Grained Vision Language Models Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55c95798-9e2b-41f5-b04c-686960fd8bd6 · outbound
FILA: Fine-Grained Vision Language Models UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a2f6583-0adc-40a9-ad94-e54e32ebd7f6 · outbound
FILA: Fine-Grained Vision Language Models Sigmoid Loss for Language Image Pre-Training
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a54fad1-509b-4f1b-a290-bca8dd6e1302 · outbound
FILA: Fine-Grained Vision Language Models OPT: Open Pre-trained Transformer Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13a0b67e-8927-45fa-b62a-1bb7f7b13556 · outbound
FILA: Fine-Grained Vision Language Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7a57c20-ba52-4629-bb46-b6dcbdfef324 · outbound
FILA: Fine-Grained Vision Language Models LIMA: Less Is More for Alignment
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97de2a77-2931-475b-ba2a-127cf6c5935f · outbound
FILA: Fine-Grained Vision Language Models We compared our model with Minigemini-HD and LLaV A-NeXT
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3cc98f9f-d5eb-4096-94be-a52d521ab81a · outbound
FILA: Fine-Grained Vision Language Models For the language model, we utilize LLaMA3-8B-Instruct (Touvron et al., 2023)
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e10d6f62-724e-42c3-a318-5986a7073dfd · outbound
FILA: Fine-Grained Vision Language Models 1 Published as a conference paper at ICLR 2025 C A LIGNMENT STRATEGY Conv StageInput Dimensions (D, H, W)Output Dimensions (D, H, W)ViT LayerViT Dimensions (D, H, W) 1 (192, 192,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 90357c10-ccc6-4179-9542-bb05f495d320 · outbound
FILA: Fine-Grained Vision Language Models DVQA: Understanding Data Visualizations via Question Answering
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c3cc7a5-af74-4f29-a919-0133eb80ee98 · outbound
FILA: Fine-Grained Vision Language Models Towards VQA Models That Can Read
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 154c7030-8c73-4199-a715-a1fa32a30540 · outbound
FILA: Fine-Grained Vision Language Models Language Models are Few-Shot Learners
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc89895b-463a-4be5-adbb-c83152f070e5 · outbound
FILA: Fine-Grained Vision Language Models Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4094e0ec-f248-4279-add7-b8db8605f171 · outbound
FILA: Fine-Grained Vision Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2f87344-0c1a-4a6f-8958-2bf4f66c849d · outbound
FILA: Fine-Grained Vision Language Models ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81823777-0b65-4849-b7b2-a4ce37034255 · outbound
FILA: Fine-Grained Vision Language Models HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c76ac201-24c0-48b3-93cf-51a6d244060f · inbound
TOBench: A Task-Oriented Omni-Modal Benchmark for Real-World Tool-Using Agents FILA: Fine-Grained Vision Language Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.