Pith. sign in

Paper Citation Record · LEDGER

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality

As of 19 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2509.06994.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.06994 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T11:11:20.896514Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 25d29911-204a-4527-85e4-62c23e6e7ad1 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Learning Transferable Visual Models From Natural Language Supervision

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.756776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.756776Z digest=sha256:9be989b22c899b88d6eaea05f33aa4700015ccc5ba1b68befb9974c91fafac8f

Observation 2504c725-81df-4bb4-873e-cf83c4670f93 · outbound

This paper cites Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.762326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.762326Z digest=sha256:3fb8d27719bb4afbd853d31b687cc53933a7c63956f40cfad195de7757f186a2

Observation 3dc30fb9-740f-40dd-a8f9-893cf054f866 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.767242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.767242Z digest=sha256:570ea306c9456d69d8510473ea8ca2e7c4dc6d441795381a5ea8ec72e463d3bf

Observation a95b629f-7b0e-476b-863a-9eff8deca0af · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.772877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.772877Z digest=sha256:9886532336a79f419ed9f3af70e856902283ab86c3d82f78645df7a1faf5e2e4

Observation 125b24fd-37ce-4e92-8b67-0a961774bd46 · outbound

This paper cites Visual Instruction Tuning.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Visual Instruction Tuning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.778714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.778714Z digest=sha256:737119491be92e74ec4a943fb58a5d456d55479b38ace390e3cd6e39436b3d05

Observation 25f20eff-fa39-4a1d-8d7d-3b9923386019 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.784030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.784030Z digest=sha256:04728dc435869f3d663646d45a8556adb2dec93fd34351f694e30e7d69dc17be

Observation 57ce2f51-1090-450f-8057-65b228fe2d21 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.790509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.790509Z digest=sha256:94530394a77354c2eb317918f311a93aceb56dd8617899c6dbdbec4099113ce2

Observation b49cb21d-d80a-4e59-ae9f-78a74c7026e6 · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.795205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.795205Z digest=sha256:ac8cc4bce8d536f70afe2260d6e8dba33f163145eac164c5ae1a1da04cea48c9

Observation 28ca6039-9c4f-4e14-b15b-487c61ac902a · outbound

This paper cites an unresolved cited work.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-05T11:11:21.766317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T11:11:20.800168Z digest=sha256:769d449ee7ad76720de4f7472b5a723c0f334eec21b20fc506bb61ef4ca3fa2e

Observation 9ae5c169-ec7d-478d-bca4-ac87619a245b · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.804751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.804751Z digest=sha256:75fd9664b67feff7d7c9a7c3ff8c3a9294a9b3723e149623018a1300394660d8

Observation f7ab3902-57b5-4892-9090-bb475f8b8cea · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.810277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.810277Z digest=sha256:e317ee5ec504cfea9fb590dd4f58beaebc9d4579ddd5d413c0fff8a59888aeb5

Observation 81de8b30-b5c4-49e4-a8f0-2efdf15f0475 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality MMBench: Is Your Multi-modal Model an All-around Player?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.816367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.816367Z digest=sha256:1ceadf219a6312f2f7137c31aa6838196c3a0c21114772835c4beef11329d104

Observation 92a070ab-54e3-4f19-acb7-99389f5a1150 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Evaluating Object Hallucination in Large Vision-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.821484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.821484Z digest=sha256:c274c0d6afec90f705be851f052a5e9d1d25fe84dc3f6b870114dd441f131f0d

Observation 79822f8d-424e-4bce-8edd-cbd1201db123 · outbound

This paper cites Towards VQA Models That Can Read.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Towards VQA Models That Can Read

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.827420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.827420Z digest=sha256:d0a20eff22757f2d6a8f2fa4c88f4b4497bb9df9e9cbfb95adc71f98793b957e

Observation 2b5709b0-a04a-4bd6-87d9-71b46c381cc6 · outbound

This paper cites DocVQA: A Dataset for VQA on Document Images.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality DocVQA: A Dataset for VQA on Document Images

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.832097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.832097Z digest=sha256:de9bb45fc82a44decd62cdf7631ab16587c1d3c87e07cb0977527b901bd66b51

Observation f7336672-e16d-4c2b-91ac-d8155104ff79 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.836873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.836873Z digest=sha256:01c6bffbcee8771c179ce09eea5105c330aee4e50b0bdcd9aa7923153d6c2029

Observation 071e3566-2153-4815-ba45-743320f94247 · outbound

This paper cites A Diagram Is Worth A Dozen Images.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality A Diagram Is Worth A Dozen Images

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.842285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.842285Z digest=sha256:d9b262d5f7ce7f199315a8c0eb6598e823956c3ffbfc0945ac3d35ef7a7a5fae

Observation 61f6fb42-0646-4ab6-9d9d-70cffb7bfef7 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.847152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.847152Z digest=sha256:0fe664b49e0eca069509ccb933cb896bcf1dc9dec0b55895ef13ad5bc2fc5dcb

Observation 593e1ea5-8e24-414b-bfa4-a69f352623d6 · outbound

This paper cites Redmon, S.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Redmon, S

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.851633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.851633Z digest=sha256:71fa7d7402d2adc30988d3b061920f1b4b57f23dbac72a568d50d3cba64178fe

Observation 5e0164a0-d264-4499-a12d-187ea2623436 · outbound

This paper cites an unresolved cited work.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Unresolved cited work

Reference 20

Resolution
malformed identifier
no resolver link, observed 2026-08-05T11:11:20.862036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.862036Z digest=sha256:64e2e6f56c4165a69e34acb30931bb5f33b2e7ba69f315e658e962c49f731fb9

Observation fe9e5a2d-8e5d-4476-a818-28d0902e21fa · outbound

This paper cites an unresolved cited work.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Unresolved cited work

Reference 21

Resolution
malformed identifier
no resolver link, observed 2026-08-05T11:11:20.868016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.868016Z digest=sha256:be67408f3196e6a6653b9f31bd305f03b6e7a1b91f026db15e1902725a9bb641

Observation 4035c6db-6ec6-4dd3-9421-f7f623789f78 · outbound

This paper cites an unresolved cited work.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Unresolved cited work

Reference 22

Resolution
verified exact
raw_fallback, observed 2026-08-05T11:11:21.162277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T11:11:20.873379Z digest=sha256:f82b33314b83a98d927230990ea836419d9f945296256091870010dcf3a4e7a7

Observation adec65b2-bbe4-4c14-90bf-310ef96b25da · outbound

This paper cites Qwen2.5-VL Technical Report.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Qwen2.5-VL Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.878592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.878592Z digest=sha256:e7d12eae4d936374629f2bd3a8a3654656b38ca8d0df6e24554e006d186fec4d

Observation 0c8d68e1-e4a3-49ab-ae44-cdf495f6c7a2 · outbound

This paper cites MiMo-VL Technical Report.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality MiMo-VL Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.883362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.883362Z digest=sha256:af1b819d35827ab57de9670e85247ab87628d1abb1fb3ffc4b0839b7d81695d0

Observation e788a6a7-88ad-4c27-add2-a97b5fd6fea8 · outbound

This paper cites KnowsLM: A framework for evaluation of small language models for knowledge augmentation and humanised conversations.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality KnowsLM: A framework for evaluation of small language models for knowledge augmentation and humanised conversations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.890509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.890509Z digest=sha256:03dd5570b6e30bca959d71911330897efe9babd55b14c4b6a05bf8ee3311b0a0

Observation ea359d5e-3ff9-41ae-9495-eb45cfce5b40 · outbound

This paper cites Evaluating the Efficacy of Open-Source LLMs in Enterprise-Specific RAG Systems: A Comparative Study of Performance and Scalability.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Evaluating the Efficacy of Open-Source LLMs in Enterprise-Specific RAG Systems: A Comparative Study of Performance and Scalability

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.896514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.896514Z digest=sha256:da4f370bed47c6c6ff6f1b3e64c8f752d91e65ac7bf6c51f0337dc5ec27da587

Pith citing papers

No inbound Pith citation observations are available.