Pith. sign in

Paper Citation Record · LEDGER

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion

As of 14 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 2 inbound Pith citation observations for arXiv:2412.04424.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.04424 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:27:52.652519Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:19:23.116337Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T19:52:01.828238Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b956506a-982a-431f-8170-27aaaa06b663 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.358895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.358895Z digest=sha256:ba5c64e6f214cb8c3cc85761436790b4d6013bdaf44674688039b46b8edce284

Observation 73d66cdb-0fd9-4603-88fb-82d5e93f9db0 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.365941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.365941Z digest=sha256:e90608b4c6016660712e968e79fff10538ce8f2be99e68e7a3eddeecd9c3844b

Observation e3dbb1f8-5959-44b1-bc98-6ab0b1857bba · outbound

This paper cites Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond, 2023.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond, 2023

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.373463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.373463Z digest=sha256:ea46dd48147d0dd107a39af3f39d35c4b8b2eb815c78958fd703759dd2d989b0

Observation e545521f-1398-4fcb-acaa-3c8b045d2899 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.595830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:27:52.384569Z digest=sha256:8ac3dbcc59cb3a8a5ce5e1e4deeaa8ae879f17bf64250d01ad6703e7e6e4ead4

Observation c2634cb1-deab-4a63-af2c-963f94e4c157 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions,.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Sharegpt4v: Improving large multi-modal models with better captions,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.390586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.390586Z digest=sha256:94d99c0769796e939ae3a663de45d9e4f434028ba6c702114979b2ac385c11cb

Observation 4bbbe195-f89a-480e-8f8b-41904477a200 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.396272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.396272Z digest=sha256:6fabfacff7729d0414ed9bf068c8ad5285b0f60f1b4d1cd344fc567fe3ca00a4

Observation 705a2024-bc8b-4132-88a3-a692ecf4a82a · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.402068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.402068Z digest=sha256:c3322bcfaf56f45f01e1e2f3dacd9be2074f75939718e33fc38d7364180c8a93

Observation 9414ee1a-b5f6-4f34-8654-3c2d883cd8f5 · outbound

This paper cites RedCaps: web-curated image-text data created by the people, for the people.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion RedCaps: web-curated image-text data created by the people, for the people

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.409590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.409590Z digest=sha256:179d151a799cbb2cc2fa3acd53fd1675841911edc0b092d7b8282ea4480aac9a

Observation ecdf898e-0441-4504-ada4-9e1fbcfa386e · outbound

This paper cites Davit: Dual attention vision transform- ers.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Davit: Dual attention vision transform- ers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.416001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.416001Z digest=sha256:6fb70612934450c22c222051ec6fcc7a559c4bd265ca6ffeda17a34323b41989

Observation 65923e19-8764-488a-8f45-b407ced021d8 · outbound

This paper cites MouSi: Poly-Visual-Expert Vision-Language Models.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion MouSi: Poly-Visual-Expert Vision-Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.421540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.421540Z digest=sha256:f8edc6b034538bdedd6d3d7c79a630555e9594ec532733d694e11f3f21c65b59

Observation d85dddf7-2316-4c01-aef1-2e9e7880a900 · outbound

This paper cites Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2024.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.550135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:27:52.427855Z digest=sha256:0fb07c30aa1ce5500c0d692dc0f61efa52db48c746a6afdd08f1e3ff053e2278

Observation 9949e921-a402-44e7-8569-72e9ea923183 · outbound

This paper cites 1 OCRBench ChartQA DocVQA InfoVQA Average Florence-VL 7B 41.4 24.3 44.5 29.4 34.9 OCR 40.9 22.9 44.4 29.0 34.2 (a) Ablation study on OCR features on OCR & Chart benchmark.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion 1 OCRBench ChartQA DocVQA InfoVQA Average Florence-VL 7B 41.4 24.3 44.5 29.4 34.9 OCR 40.9 22.9 44.4 29.0 34.2 (a) Ablation study on OCR features on OCR & Chart benchmark

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.530537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:27:52.433721Z digest=sha256:901350bcfc489ebda09abf81acf37aa7689ccb11023c5c09be849aa107ae59d3

Observation 2a861900-2d06-4a04-aac0-5dff573460da · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.440844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.440844Z digest=sha256:ec87fddd333af9d9a92263d1d1df63381eac2af6c921f0cf6008b5c4e8c68ffb

Observation 1275089c-6f36-4479-b281-f93268b4aadc · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Vizwiz grand challenge: Answering visual questions from blind people

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.449811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.449811Z digest=sha256:2bf0feef3813ff87bc75b3b65a860dc07811239996822d3bb47c5e903fdb4d20

Observation 97833a46-8d04-47f4-9ce4-7e47936c38dd · outbound

This paper cites Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.456702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.456702Z digest=sha256:8769eb913c5b5469bab1fd279ee1f0478c201659b6c1642bcad0c032295283b0

Observation 9057387d-297f-4da6-b097-d93f084b186b · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and composi- tional question answering.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Gqa: A new dataset for real-world visual reasoning and composi- tional question answering

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.488827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:27:52.463327Z digest=sha256:df8984b5b2c7dd9656a93a12702e4a5b7462e42f6df64d39dd92c002249bc22f

Observation be1c5739-1cc9-4fc8-a3d4-d63acc3e9d10 · outbound

This paper cites https://huggingface.co/datasets/huggingfacem4/docmatix.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion https://huggingface.co/datasets/huggingfacem4/docmatix

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.471426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:27:52.469074Z digest=sha256:5e03d31ff20cde25a54c8df9375b1c77d1aca5e711b5bb6a7c41870902362a66

Observation 25e88791-2850-4d96-abf5-cb34e4ed394a · outbound

This paper cites BRAVE: Broadening the visual encoding of vision-language models.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion BRAVE: Broadening the visual encoding of vision-language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.474048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.474048Z digest=sha256:cdb66e2a884fbba6c1f88336f812b60741accb8b1c4a98832b3b4f65a77e4d42

Observation c81180e8-05f8-4fdf-b406-7af522edb8a9 · outbound

This paper cites A diagram is worth a dozen images.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion A diagram is worth a dozen images

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.480290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.480290Z digest=sha256:4661717efc4b4a939dd35a1ddefd630c994ab6ea4de7de7ceb0342e4c69facfd

Observation 69d6e750-e170-4d16-b40f-58e33aa5b120 · outbound

This paper cites Segment anything.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Segment anything

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.443083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:27:52.485430Z digest=sha256:2200789da6ef71a2e583cc2e93e17bfb7b5e3c5e8e21db138aa25f9e1941122f

Observation c18cdb79-ee2e-4d68-80ad-9c318057e82e · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.490300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.490300Z digest=sha256:c0bd76772967b849e675423a1ef912953344f3390eed3cb15eaf05fed27a0428

Observation 1c185393-b516-4b8f-a95a-a7f5e050ab82 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Evaluating Object Hallucination in Large Vision-Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.495049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.495049Z digest=sha256:3570b39666cbef6a664e7a65d2979a2bb159a84c4a20a0010900a6364e1c3695

Observation b7ada76e-a126-41bf-8e2d-2544309c0542 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.500877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.500877Z digest=sha256:9da62a0189ea088cd6db85f89390bf9c7922d3a6fe830f7cba4ad3ea9d711ddf

Observation 585e9a9a-c40d-424c-af9c-6471ea6713dc · outbound

This paper cites Vila: On pre-training for visual language models.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Vila: On pre-training for visual language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.425426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:27:52.507812Z digest=sha256:54e115d3ba38a41c0467e52dd4502d2c909a5454fe1659f83851b5bd590d4713

Observation f82b0b6d-6db2-4632-b1e7-c07ee3860d84 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.407582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:27:52.513166Z digest=sha256:73618290f5af9d288343e733793806f0eb4245c0165b2fd5091b3540db6be3a3

Observation 0405be2a-9cb1-41b3-9e90-e9d964aad5f3 · outbound

This paper cites Visual instruction tuning.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Visual instruction tuning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.388823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:27:52.517698Z digest=sha256:519d3cb8f96ef2ed556c3b03ea593132713ea64529596a48cbb868c23c84dfae

Observation df400d4d-2b78-483d-a1cd-3d232cc529a9 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion MMBench: Is Your Multi-modal Model an All-around Player?

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.522098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.522098Z digest=sha256:b8895a7b53501dc97eac61ef483174c8b7cf7159792ad43975ace0d4aef4c86d

Observation 66be9e77-1902-4cae-8bfa-ad0a2649ecd9 · outbound

This paper cites On the hidden mystery of ocr in large multimodal models, 2024.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion On the hidden mystery of ocr in large multimodal models, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.370325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:27:52.528019Z digest=sha256:d28a9a113d45e171d3f5e4b3ca7211d03ebf5705ced6aca1fd0dc8c977532b12

Observation b4cc4588-c9cb-406e-a003-5ad3c972f9f1 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.534535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.534535Z digest=sha256:3f15049044ac48caa76420d65ef2290430b7f319b049b9c27b994263715de559

Observation 7882e73a-5805-4a99-a5e8-d8f1de9d942d · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.542222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.542222Z digest=sha256:79783b0110cab5756dd1a387e6d5e74985ab78e70471f6a583fc332d9e63729a

Observation 03014ae0-8ff4-4404-b100-0b1158232468 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.549213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.549213Z digest=sha256:b39d80a53bd625eac426de82df3f9529c6e971329e265a2748e09d9797c070a1

Observation 6b97a60d-041a-4f64-84d2-d8ab627bcdff · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Docvqa: A dataset for vqa on document images

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.340149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:27:52.555743Z digest=sha256:95e887ce73734774c0fc4193a7782f1a682b7417fec6f0900df120dc8adeba4c

Observation 47f88643-2ab5-451e-a260-1f0eae8f096d · outbound

This paper cites Infographicvqa.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Infographicvqa

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.322795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:27:52.562395Z digest=sha256:c54546c1a6a78c6fb5b0b392d3ac1cef1ad73c711d3959f355909d2cb93cdfc9

Observation a398bc26-f685-401e-bcd0-73a7cbafa363 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion DINOv2: Learning Robust Visual Features without Supervision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.567906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.567906Z digest=sha256:bc62f972387139b0cb9c5e4e86ca5b769f36bff7439bcbfe97ff51a57f747cce

Observation d4943f97-7c2e-4cbd-808e-d00c1d060b6f · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Learning transferable visual models from natural language supervi- sion

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.576254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.576254Z digest=sha256:10a58c35f4150b2998028aa790d2739458b59ab44a06156a32054ffb036b947d

Observation b88d5a6c-e5f9-46c9-ace0-dc8f350875f1 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion High-resolution image synthesis with latent diffusion models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.294408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:27:52.581690Z digest=sha256:5e35550f12d35b7b9df7d39d3d04854af41a4e0fedcfbf77e7b8afd4da7a920d

Observation 71e6ccaf-f75e-4a55-97a5-ec119f8c7d18 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion High-resolution image synthesis with latent diffusion models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.586544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.586544Z digest=sha256:c8b1bbf42bd74d207ece8d08439b101e6d3ca64f90eb879c354ea7d057cdbf8f

Observation 09010bec-878a-4b04-b1d3-177e846291a6 · outbound

This paper cites Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.592158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.592158Z digest=sha256:0247f2d24ed6d3f78407276b9921d8efd09782abe94b71ddf743adcbfbd7692e

Observation 5d7b7ede-4e70-42f2-9375-2c3f2253565c · outbound

This paper cites Towards vqa models that can read.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Towards vqa models that can read

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.597176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.597176Z digest=sha256:e5b039f94ea80ec6f2ee892c134cf7dd9a7e10bcd5036343cb12fc91706e2875

Observation 8a3a990d-15cd-4897-814f-c08f30a6125b · outbound

This paper cites From pixels to prose: A large dataset of dense image cap- tions, 2024.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion From pixels to prose: A large dataset of dense image cap- tions, 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.245545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:27:52.601645Z digest=sha256:268a09d9ba7fd0bcb6b2e66cc83d3f7e63175aacb0716de08429a09ad4499624

Observation ae169a02-3cd8-42c9-8da2-c956e3896f11 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.606243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.606243Z digest=sha256:41e38d7729d300e831cf91c3887b8b6ee4d173c73f45303cca4b2c0b47313316

Observation 791cc9f6-8b4e-42c2-a0d9-998d402dfeb2 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.223712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:27:52.611668Z digest=sha256:cf40d878f63f0acf47b8284bca2118da50397bd725d24bec79b06dc21b9e8069

Observation 539ee10d-03e0-4204-8786-9fd0a58bf643 · outbound

This paper cites Diff-eRank: A Novel Rank-Based Metric for Evaluating Large Language Models.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Diff-eRank: A Novel Rank-Based Metric for Evaluating Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.617470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.617470Z digest=sha256:12d151c04cf07bf0b7d6686ce101831748bc1ca4e1ec54131355922c898dd532

Observation 50447d1f-36a5-4257-859b-53a1eb9bf355 · outbound

This paper cites Grok 1.5v: The next generation of ai.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Grok 1.5v: The next generation of ai

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.206047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:27:52.623951Z digest=sha256:ceffc6a94f7b1835edb7060bdfabc75e4265f215d0efad133f2c24fa7dae4ba4

Observation 57ac3f18-f882-475b-b822-65dc159353f6 · outbound

This paper cites Florence-2: Advancing a unified representation for a variety of vision tasks.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Florence-2: Advancing a unified representation for a variety of vision tasks

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.189596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:27:52.630399Z digest=sha256:7028a4b139bd6ef6b1881fe57a7e65dcfb687b69ce0f069bd9cac60d8dfe1c39

Observation 297f77cb-0bbb-4f14-a6e5-a1b75c730c1d · outbound

This paper cites Vision-flan: Scaling human-labeled tasks in visual instruc- tion tuning, 2024.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Vision-flan: Scaling human-labeled tasks in visual instruc- tion tuning, 2024

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.170495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:27:52.635275Z digest=sha256:6ae08907f643cc150555b4690f0af64c6a28a9c24a48b61c6d1905a3b3d8934b

Observation 92bde076-13c4-43fb-b283-aa4ef0b83d90 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.641065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.641065Z digest=sha256:d7b03505aeac3d8c9598566f2a7e428682f28d02d87ce137265e8ae09889dd1a

Observation f0b8e686-718e-4e5b-9b77-3f0029a142b7 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.150271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:27:52.646737Z digest=sha256:3eecc73b5e439dc0f757387706f2f7503179a8d54b21b82f3cb0bea7c7e6521b

Observation 981cc79f-9a2d-4088-ab3c-0344007fd948 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.652519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.652519Z digest=sha256:95f516ec1d817821dc00887348f7539e982373f410f96db533a4efad0a9a1370

Pith citing papers

Observation 64004760-0a21-4965-9e0b-4fc9dcc32781 · inbound

FastVLM: Efficient Vision Encoding for Vision Language Models cites this paper.

FastVLM: Efficient Vision Encoding for Vision Language Models Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T13:19:23.116337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:19:23.116337Z digest=sha256:eff81fc9e9c8b060d76f2de47f25965de43a6c7669a062196e80fb56b93ec377

Observation 218253da-2784-4767-9d0b-c3607ba54cfd · inbound

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding cites this paper.

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:52:01.829787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T19:49:00.961388Z digest=sha256:9e06e77c21f32b97c20b5f5e63bcff19eae337d588cd475eb3421166d15b148e