Pith. sign in

Paper Citation Record · LEDGER

FILA: Fine-Grained Vision Language Models

As of 17 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2412.08378.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.08378 v3

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:56:46.279061Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-19T20:43:16.364125Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T20:47:45.997187Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f813ec16-4a44-46b2-96c0-1376501fd228 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

FILA: Fine-Grained Vision Language Models InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.154642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.154642Z digest=sha256:ad590d4f3e4301665a7fdf63d4632685f19589b41ab442be96437f0809554b10

Observation 4909a288-bb1f-4aee-b11f-df4b973d623e · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

FILA: Fine-Grained Vision Language Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.160077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.160077Z digest=sha256:4d10c490dfe151728666dca4d5d15657ca03cdfe340a59e271db06955acf76eb

Observation 888df4e2-6bce-4c69-945a-8a2489a81a2d · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

FILA: Fine-Grained Vision Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.165232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.165232Z digest=sha256:ab8edb3dfea0d2ae13f24cf3db6a51a2f1116b618262ece63287db93cb0d16c8

Observation ea4bc805-2138-49c8-b0d4-f514d986f7c1 · outbound

This paper cites mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding.

FILA: Fine-Grained Vision Language Models mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.175229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.175229Z digest=sha256:e89449e895aceaa0e858f26c55131366bc98d811efbd529664c7712240e44279

Observation b3d3e998-e05f-4e26-97dc-c4565a33a5c6 · outbound

This paper cites Segment Anything.

FILA: Fine-Grained Vision Language Models Segment Anything

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.186496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.186496Z digest=sha256:9305904bc63feb36aec27336ac1316e3b8c7dad8e4ad1ae08e1301ba6a7a8520

Observation 428fa63d-aa58-43e8-87c2-8b460dfb63c0 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

FILA: Fine-Grained Vision Language Models Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.192029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.192029Z digest=sha256:9183dac29c74d4a2f62ec28cb8437c0dc7cc57949c6413b8e46502be34693b36

Observation 1175c106-796f-429e-8a42-686da30ca90b · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

FILA: Fine-Grained Vision Language Models DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.197387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.197387Z digest=sha256:8f43dd5eae71ca8fd0f69a6fc223a70e6616cc506f374a7c11ea90e5f53c468a

Observation 6ba66138-4fc2-4694-8843-2fba4401a95b · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

FILA: Fine-Grained Vision Language Models Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.202345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.202345Z digest=sha256:b381e4449cc932b80b206ce707da7cc54b214a462b2f85695044db44751d529c

Observation 9bc0d697-dd1d-4f1d-87fe-7262da9f3897 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

FILA: Fine-Grained Vision Language Models ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.207284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.207284Z digest=sha256:e88b3d8b3780b4cc35df0cfa900071478ef63df4bedf3155130728646e93f47e

Observation 8d3c4315-3ca8-4715-82c9-2179207d2bc1 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

FILA: Fine-Grained Vision Language Models Learning Transferable Visual Models From Natural Language Supervision

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.213549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.213549Z digest=sha256:c4f4fce99fcf4d16cdfd39b84853f90fe7f5c43bddd915deed11ddf1277b6429

Observation a0970d59-34c6-4972-a6b9-5388cb254a10 · outbound

This paper cites an unresolved cited work.

FILA: Fine-Grained Vision Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:56:46.770043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T17:56:46.218289Z digest=sha256:adf9277290cc6e9f646fd5c88c5e77c32da1d6b092aeb59961d11b70a8a2dbe5

Observation f2e99b39-f197-4592-836e-3a76de13f774 · outbound

This paper cites LAION-5B: An open large-scale dataset for training next generation image-text models.

FILA: Fine-Grained Vision Language Models LAION-5B: An open large-scale dataset for training next generation image-text models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.222949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.222949Z digest=sha256:fe2e6db883c9889068832d72062baba433c596e049e64ee1966a34075b1780ad

Observation 44f6bc17-8221-4c0d-8520-aafe88bde415 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

FILA: Fine-Grained Vision Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.233399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.233399Z digest=sha256:7449d83d2c7807352d476a217a231626379b04d7eb101e7761b7dbb62ef6d39e

Observation ae29ba6e-f9d1-4903-a0f4-529a03f08af7 · outbound

This paper cites Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models.

FILA: Fine-Grained Vision Language Models Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.238546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.238546Z digest=sha256:e34079805a6e71a97b870f64243b1194fc3d21ffffb6b2f58fd7cfc5c7c3d76a

Observation 55c95798-9e2b-41f5-b04c-686960fd8bd6 · outbound

This paper cites UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model.

FILA: Fine-Grained Vision Language Models UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.243482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.243482Z digest=sha256:bebdb6c723f71b24496958d91b3431eee23891f0d0a190415553543bdc98f3e2

Observation 8a2f6583-0adc-40a9-ad94-e54e32ebd7f6 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

FILA: Fine-Grained Vision Language Models Sigmoid Loss for Language Image Pre-Training

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.248659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.248659Z digest=sha256:bd75e798b0381ee8cd807642ff1cb2d2d1ca8db96e5343662b1baf2289d256a6

Observation 7a54fad1-509b-4f1b-a290-bca8dd6e1302 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

FILA: Fine-Grained Vision Language Models OPT: Open Pre-trained Transformer Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.253816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.253816Z digest=sha256:319daa40f314e4d7052c95cebfeda12825ecafbde96919c5ca0ce2853136ec11

Observation 13a0b67e-8927-45fa-b62a-1bb7f7b13556 · outbound

This paper cites LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding.

FILA: Fine-Grained Vision Language Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.258654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.258654Z digest=sha256:df3164bc6a7382d1e4b771c91c9f5052f2bfad6a694aa0a4c53ecbe4a41c87af

Observation d7a57c20-ba52-4629-bb46-b6dcbdfef324 · outbound

This paper cites LIMA: Less Is More for Alignment.

FILA: Fine-Grained Vision Language Models LIMA: Less Is More for Alignment

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.263567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.263567Z digest=sha256:857eb64f5167d192249f4d78a632523828fa3414640d4b0a28fa93cde5558112

Observation 97de2a77-2931-475b-ba2a-127cf6c5935f · outbound

This paper cites We compared our model with Minigemini-HD and LLaV A-NeXT.

FILA: Fine-Grained Vision Language Models We compared our model with Minigemini-HD and LLaV A-NeXT

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:56:46.753400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T17:56:46.268813Z digest=sha256:56f313bb6b16ce48b534324228d8d024945b3c06e7e7df97135f8eb416bab886

Observation 3cc98f9f-d5eb-4096-94be-a52d521ab81a · outbound

This paper cites For the language model, we utilize LLaMA3-8B-Instruct (Touvron et al., 2023).

FILA: Fine-Grained Vision Language Models For the language model, we utilize LLaMA3-8B-Instruct (Touvron et al., 2023)

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:56:46.736611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T17:56:46.274343Z digest=sha256:8af70bd4487183f444d7cb63e539bcd22905fb92b048453aea815b1d3840f517

Observation e10d6f62-724e-42c3-a318-5986a7073dfd · outbound

This paper cites 1 Published as a conference paper at ICLR 2025 C A LIGNMENT STRATEGY Conv StageInput Dimensions (D, H, W)Output Dimensions (D, H, W)ViT LayerViT Dimensions (D, H, W) 1 (192, 192,.

FILA: Fine-Grained Vision Language Models 1 Published as a conference paper at ICLR 2025 C A LIGNMENT STRATEGY Conv StageInput Dimensions (D, H, W)Output Dimensions (D, H, W)ViT LayerViT Dimensions (D, H, W) 1 (192, 192,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:56:46.720232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T17:56:46.279061Z digest=sha256:c9e592eeab4a12aecea3a3ef91c8c491a898f99fc0e3123c6fda6704b8001918

Observation 90357c10-ccc6-4179-9542-bb05f495d320 · outbound

This paper cites DVQA: Understanding Data Visualizations via Question Answering.

FILA: Fine-Grained Vision Language Models DVQA: Understanding Data Visualizations via Question Answering

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.180959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.180959Z digest=sha256:b9cbefb3083f491651ea967bc86388c457aa7108ac799cc697c0a03be928a384

Observation 4c3cc7a5-af74-4f29-a919-0133eb80ee98 · outbound

This paper cites Towards VQA Models That Can Read.

FILA: Fine-Grained Vision Language Models Towards VQA Models That Can Read

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.227473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.227473Z digest=sha256:c793215c74ed8d9d5a0ffa29c0a33f96a14128afea5829afbe7f92e7f88457f7

Observation 154c7030-8c73-4199-a715-a1fa32a30540 · outbound

This paper cites Language Models are Few-Shot Learners.

FILA: Fine-Grained Vision Language Models Language Models are Few-Shot Learners

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.136537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.136537Z digest=sha256:421bdad61d1f9c8e5c09e33126e9163fcaeb845e201a7d18783a236dc39de951

Observation cc89895b-463a-4be5-adbb-c83152f070e5 · outbound

This paper cites Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts.

FILA: Fine-Grained Vision Language Models Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.142751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.142751Z digest=sha256:80752642932fb95be755440e0dd945f3ae7e8df2e666b619545d6d1df6f48dcc

Observation 4094e0ec-f248-4279-add7-b8db8605f171 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

FILA: Fine-Grained Vision Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.130070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.130070Z digest=sha256:191e100b46d7f1075f99a0f1cc6bf4e552cc336d67545f72f481eff841f73171

Observation c2f87344-0c1a-4a6f-8958-2bf4f66c849d · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

FILA: Fine-Grained Vision Language Models ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.148931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.148931Z digest=sha256:b51872e708cc904f1c9139ea6034adc27e537554ba3c1c9720451faeb2cd2324

Observation 81823777-0b65-4849-b7b2-a4ce37034255 · outbound

This paper cites HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models.

FILA: Fine-Grained Vision Language Models HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.170245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.170245Z digest=sha256:861b2a3253139b34fbb92466566503e041a4c56644f54c625259bc0be99053b4

Pith citing papers

Observation c76ac201-24c0-48b3-93cf-51a6d244060f · inbound

TOBench: A Task-Oriented Omni-Modal Benchmark for Real-World Tool-Using Agents cites this paper.

TOBench: A Task-Oriented Omni-Modal Benchmark for Real-World Tool-Using Agents FILA: Fine-Grained Vision Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:47:45.998928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-19T20:43:16.364125Z digest=sha256:3b6cac28d93fc46c0652b48c9346ee8195cbb92b15411a9528a3e6bdadbcb1f1