Pith. sign in

Paper Citation Record · LEDGER

What's in the Image? A Deep-Dive into the Vision of Vision Language Models

As of 22 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 8 inbound Pith citation observations for arXiv:2411.17491.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17491 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:08:51.495834Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:05:19.272138Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T05:54:33.575217Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation de162673-5c57-4daa-9aa8-86f1473d7bd1 · outbound

This paper cites GPT-4 Technical Report.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T12:08:51.353986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:08:51.353986Z digest=sha256:e6654372a0bce5ee9037b227221e80af7d551f44c2f0542c3e32de7600446619

Observation b5b22506-24a0-4e0f-a8ac-5e58e5b63171 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T12:08:51.356686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:08:51.356686Z digest=sha256:91f3adf6b1e64634e1cc72b43ac1167b68247a9bea6069e54011ea1d4835ec61

Observation 087802d2-6fab-4b70-b3da-6d7ddc6b265f · outbound

This paper cites Hallucination of Multimodal Large Language Models: A Survey.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Hallucination of Multimodal Large Language Models: A Survey

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T12:08:51.359108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:08:51.359108Z digest=sha256:5677b2fced74d0745fa8d6471366261c6caab9ae57ca6a262b46cc5fca20e9be

Observation c5281822-b368-42a7-a429-9a55a42a0703 · outbound

This paper cites Understanding Information Storage and Transfer in Multi-modal Large Language Models.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Understanding Information Storage and Transfer in Multi-modal Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T12:08:51.362019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:08:51.362019Z digest=sha256:08223ae10f2976df9d96f870ef0597472ad21e3641242419cfc143b6d219d251

Observation beb9b1f7-2603-43fe-9f23-28c01888eb24 · outbound

This paper cites Generic attention- model explainability for interpreting bi-modal and encoder- decoder transformers.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Generic attention- model explainability for interpreting bi-modal and encoder- decoder transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T12:08:51.365427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:08:51.365427Z digest=sha256:f6a45cfce08125eea5c891647c604413fa7689639eed41049c33f39dbb941068

Observation 892eb2f5-f6e5-4c20-85a2-2f1fd3b6fbf7 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T12:08:51.368155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:08:51.368155Z digest=sha256:3c38a93ac09acd1e2212e7fe59562d19fb9361f66522ded6cf68690ca01138df

Observation 2c5e41ca-7797-44e9-9092-f701b65660a5 · outbound

This paper cites InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:08:51.830944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:08:51.371693Z digest=sha256:42b4446791919ffce3a2c1832a6c538a5ba54e7366aff694a3cb227228e16036

Observation af800eb7-bc7a-49ae-ad87-0c72d78b3f4c · outbound

This paper cites What Does BERT Look At? An Analysis of BERT's Attention.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models What Does BERT Look At? An Analysis of BERT's Attention

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T12:08:51.374510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:08:51.374510Z digest=sha256:8d01225cac2cf43b37ae30ce1a05184def908d64ece72dd191b8e475ed9cc2de

Observation 1c8bde47-d19b-45bd-9b00-ac98c8ef9461 · outbound

This paper cites Towards automated circuit discovery for mechanistic interpretability.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Towards automated circuit discovery for mechanistic interpretability

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T12:08:51.378183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:08:51.378183Z digest=sha256:d5e32d5a08553de644e9f1ef9a26f01d3428244151e5325cc1cb4d48f81cc1ce

Observation d6ffd854-c6f7-4f43-8c74-f458c8a0410a · outbound

This paper cites The Llama 3 Herd of Models.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T12:08:51.380730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:08:51.380730Z digest=sha256:b6025dddce892f5045bf8a59e4ff5cc916437c5adb954b5832f98bcb53b2c224

Observation 11c432f5-df61-4d2d-8652-c4014c85c10b · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T12:08:51.384183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:08:51.384183Z digest=sha256:d8af0a93cf2dd6b936cf5effa2afaac0fe138810fd091ff157f353f7a838cd92

Observation 762f1bae-143f-4988-b852-fe6dfba05150 · outbound

This paper cites Transformer Feed-Forward Layers Are Key-Value Memories.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Transformer Feed-Forward Layers Are Key-Value Memories

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T12:08:51.387157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:08:51.387157Z digest=sha256:b87952f96b83a724685d303951faf559f86d43b4598abf83140da5dbf7fad058

Observation 69994d22-d342-4955-87e9-fc3bb30f59b0 · outbound

This paper cites Dissecting Recall of Factual Associations in Auto-Regressive Language Models.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Dissecting Recall of Factual Associations in Auto-Regressive Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T12:08:51.389838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:08:51.389838Z digest=sha256:3c8753df00e7c7baab79616ee919c67e099161174788bb4650d0ddedebb0aaa1

Observation e7439443-0c2a-4d36-8655-d036e87cd50b · outbound

This paper cites Chat- gpt outperforms crowd workers for text-annotation tasks.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Chat- gpt outperforms crowd workers for text-annotation tasks

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:08:51.819882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:08:51.392418Z digest=sha256:832e59b9b06a209619448bdb9c056855d0f3b0e36e7ae76e4a94924193d3a271

Observation a6fe1baa-c504-4b3e-b6e3-a84e9024d6cb · outbound

This paper cites The Unreasonable Ineffectiveness of the Deeper Layers.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models The Unreasonable Ineffectiveness of the Deeper Layers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T12:08:51.394433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:08:51.394433Z digest=sha256:d8e75605913a0c37c870d434a9bdc6e85132bad9bb182bc803da48f777c009b0

Observation 36b2e311-9696-4ac9-b50c-540d6696c5be · outbound

This paper cites Conceptgraphs: Open-vocabulary 3d scene graphs for per- ception and planning.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Conceptgraphs: Open-vocabulary 3d scene graphs for per- ception and planning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T12:08:51.396597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:08:51.396597Z digest=sha256:6562e958821cb78dcfc559beb3c9455051ae678506b9aad31a47f7f55a3e3dfd

Observation 5ee653b1-5446-478e-86a5-0f086855a170 · outbound

This paper cites Is chatgpt better than human annotators? potential and limitations of chatgpt in explaining implicit hate speech.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Is chatgpt better than human annotators? potential and limitations of chatgpt in explaining implicit hate speech

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:08:51.808786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:08:51.398506Z digest=sha256:52c7a0f97eba5e75b344a3a60a6d8fa4a70baa46ccfb93f7b7a926c6a0dba8dd

Observation 3d0eefc8-f3f7-4c04-9799-4d0f9b678951 · outbound

This paper cites Mistral 7B.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Mistral 7B

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T12:08:51.400392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:08:51.400392Z digest=sha256:5d6195a48b9ec83c7e6a4461af6c75098b721d08f3352c5b532d63fb1d090588

Observation 23d4f89a-e093-4483-952c-f7d56133c2f8 · outbound

This paper cites Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T12:08:51.404123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:08:51.404123Z digest=sha256:0cd4dce812a4658499d6e3e7b771f10b383280f9ed64be5a4226d20f457a6d01

Observation f30c3830-d93c-40f3-9e8d-a8f0597fa5d2 · outbound

This paper cites Segment any- thing.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Segment any- thing

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:08:51.801838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:08:51.407023Z digest=sha256:8c1ab993bfc6237cf20ce69b652158b256d7d033274893912d8a04e00973fac3

Observation 0d8990f3-234b-4afc-95c1-294ea1ca8b52 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T12:08:51.409875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:08:51.409875Z digest=sha256:5b6a42bc96e44c238369eca011a775e545313971c20f85ff58f6a084cdf47fdf

Observation 1b0f2264-ffa6-4b98-8eba-dfb359cab21d · outbound

This paper cites LLM-grounded Video Diffusion Models.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models LLM-grounded Video Diffusion Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T12:08:51.412603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:08:51.412603Z digest=sha256:6f9dd1ff95795212433b0900e729b55225df7ef5810879ce712d28af78d8c056

Observation 4d141e33-a30a-4f55-a70a-1c097ee82b25 · outbound

This paper cites Microsoft coco: Common objects in context.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Microsoft coco: Common objects in context

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:08:51.790775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:08:51.415312Z digest=sha256:9e9d4bdf0734e55c9346773c7b511ed4de9aef27b7de6eeea4da5732980ffba6

Observation 9accea1c-b706-46d7-8a4e-b15e7e7dcba2 · outbound

This paper cites Improved baselines with visual instruction tuning.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Improved baselines with visual instruction tuning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:08:51.783319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:08:51.417742Z digest=sha256:0b7e91d9d6fcbff1bde19afb137a3b29eb7024b40fd30275f462444788d5ea7f

Observation fe2364bd-9b3a-4fa7-a5c8-5f8c273d67e4 · outbound

This paper cites Visual instruction tuning.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Visual instruction tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T12:08:51.420681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:08:51.420681Z digest=sha256:ec26613ebd4074b57f10e3d0cf859721d5948057dc62e8eb213f1085a4df5a75

Observation 134b4812-cce6-421a-be8a-8f15f0430de6 · outbound

This paper cites Clip-driven universal model for organ segmentation and tumor detection.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Clip-driven universal model for organ segmentation and tumor detection

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T12:08:51.422975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:08:51.422975Z digest=sha256:df336a21fdc9f08c24836aad77c510dc81d2280104c0b791b5d060780eff69e1

Observation bd0fc225-4d62-4fd3-a26b-b4589b72f26b · outbound

This paper cites OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T12:08:51.426120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:08:51.426120Z digest=sha256:6999673ecbe276cab56489678a874ce9967e3322a627d55a5f05ea368a3c746a

Observation d1d4dd65-50d5-4040-a114-327c57766e75 · outbound

This paper cites Locating and editing factual associations in gpt.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Locating and editing factual associations in gpt

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:08:51.768597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:08:51.429625Z digest=sha256:84d4c803bc9f9958e2718ef3ff957c16f5a2786c477e1761704b3357def988dd

Observation 73d73144-6432-4148-9f8a-c10b50ba680b · outbound

This paper cites Towards Interpreting Visual Information Processing in Vision-Language Models.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T12:08:51.431988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:08:51.431988Z digest=sha256:33919c817f28b4e3bcae0873d5d347199bdc0c08af184c5f22f9604655e9782f

Observation 5281fb7e-973f-4dee-b8bf-c06ac3d31b2f · outbound

This paper cites Interpreting gpt: The logit lens.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Interpreting gpt: The logit lens

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:08:51.759334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:08:51.434968Z digest=sha256:120ad0521d349bb5ca618276a9b24e5022caa4b8a9c9bbe07b0e5924fff150e8

Observation b41ac9f2-1cee-4e82-b4c6-0217cac9eb25 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Learning transferable visual models from natural language supervi- sion

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T12:08:51.437050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:08:51.437050Z digest=sha256:cea89a8e3c3bfe153ccc3ef27be519322e79e37115d04ba3d4b4932e6198b6f1

Observation d516d450-f816-4a82-8e34-11933065aa28 · outbound

This paper cites A multimodal automated interpretability agent.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models A multimodal automated interpretability agent

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:08:51.747230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:08:51.439408Z digest=sha256:7de2f0bfa170409e71fb4d2834661cdcd8d6359a2e54dc6969f2bc97d244dda6

Observation d04c8566-f351-47ef-a2c9-04baecb01423 · outbound

This paper cites ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T12:08:51.441277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:08:51.441277Z digest=sha256:83383e1243887c703990030a6661309f4aaf8588320addcf572cecbafa7f10d9

Observation 73eb3bb7-fda1-4246-8e3a-b3aab52b2418 · outbound

This paper cites DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T12:08:51.443855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:08:51.443855Z digest=sha256:14311cfdb6dbe4c21ecde4d2b5b4f6ab93065b63b8fe68305b3484c28ca56a70

Observation 4f574b5b-25e9-4d99-a0f9-906b87061d33 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T12:08:51.446752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:08:51.446752Z digest=sha256:09a642958a429f177da7b658c3b4f8ac16202edab39cfcfc66e610e2dbf47db9

Observation bda0570c-9744-4e3d-b21d-977a0f286b40 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T12:08:51.449548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:08:51.449548Z digest=sha256:c21515c6c34136e5f68dfe6500850758bed6d7f2c453e49852100b94882cd91a

Observation 51702afa-4fef-46f5-88c0-487e4462ae75 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Chain-of-thought prompting elicits reasoning in large lan- guage models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T12:08:51.452004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:08:51.452004Z digest=sha256:e9b346cd97b0987cea00366cd1f8b2c6f1335cd54e496009368526ba7821afc2

Observation 0281bca4-c325-4755-a8ac-af03796a52e1 · outbound

This paper cites Qwen2 Technical Report.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Qwen2 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T12:08:51.454947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:08:51.454947Z digest=sha256:a55d98010ecacbfdb9e78a9c18a96ff4b04ffd240559b02a80b0c5860f776918

Observation 8a863dfb-10ab-4558-80f1-393dde893a8e · outbound

This paper cites EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T12:08:51.457603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:08:51.457603Z digest=sha256:768bdda2491295c87605a044f2c049b43df751c2ebe006df13993e313ff9edd0

Observation bfb880ad-6c13-4597-8f83-50af331a3d1a · outbound

This paper cites Yes” and “No.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Yes” and “No

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:08:51.731482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:08:51.460785Z digest=sha256:c38dff1af61c5e8f4504e4a7de2ec23b49c1afe8a339dcccdd16f61e9258623a

Observation f51b07f5-21bf-4cd0-808f-b3d1e1c630c7 · outbound

This paper cites Detect only tangible objects that can be interacted with.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Detect only tangible objects that can be interacted with

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:08:51.723722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:08:51.463653Z digest=sha256:415e22da3c11ae6154092b64b416f388958b71420eabb8702428df15d5ce6da1

Observation 9f2c0342-4454-4093-86ec-12f7d0279d76 · outbound

This paper cites an unresolved cited work.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:08:51.716852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:08:51.466555Z digest=sha256:5aaf4002c3d4d0bdf36acafec2f55e6e4cbde23b24ced029ba68c439fa89b906

Observation 5da0c476-abd1-4512-8179-5f2a724f5722 · outbound

This paper cites If half of the *physical objects* in the predicted caption are also in the groundtruth caption, the precision would be 0.5.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models If half of the *physical objects* in the predicted caption are also in the groundtruth caption, the precision would be 0.5

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:08:51.710171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:08:51.469117Z digest=sha256:0be9d66e8ab0d5751df6bfafbcbb362ff4db4f705fc91962eddea4423686cc5a

Observation 9d2d3c9b-50c3-4cf0-8050-b26dd699507d · outbound

This paper cites fine detail.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models fine detail

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:08:51.702409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:08:51.472159Z digest=sha256:cefce50e2d9eb94ee89e92ec01224d9a666a862e27c98fdb71a5bc3d724aad57

Observation a82c2b7e-f90a-4250-b4a5-8f04be1fbf0c · outbound

This paper cites an unresolved cited work.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:08:51.695247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:08:51.475675Z digest=sha256:bfc1bb5e03c79a62f053eef17b0f862f954cce2538c6a694e9b3d45867e69a5d

Observation 25761686-1e43-4a08-a42a-2be8d0e0fcc1 · outbound

This paper cites an unresolved cited work.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:08:51.688283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:08:51.478135Z digest=sha256:9f0de2453e4b7e8866b5a189f5a9c72f92de3c6896fdf7e10c095c08516c4e95

Observation 3653dfb8-23e9-403a-b553-f87fb5ebd7c6 · outbound

This paper cites an unresolved cited work.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:08:51.681404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:08:51.480453Z digest=sha256:e519462b3cf1969ca392e4a0088da3736e7e2c9a45bdc182ae368f06f34a9c47

Observation 73d9ce69-4a03-4532-802f-d308c8b4c901 · outbound

This paper cites an unresolved cited work.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:08:51.673315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:08:51.482458Z digest=sha256:7d1c078d14332dd5220b6f3c5c5f924c97271c7af445634f7731e3b512deaeac

Observation ac485285-edda-4d31-bf3c-fc2d54092387 · outbound

This paper cites an unresolved cited work.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:08:51.665987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:08:51.484462Z digest=sha256:90224e101ec67d6e5625ff58cecb8a678ecbb471bc96afcf0b0b95e7aad41cdc

Observation 78f7737a-3741-45b4-a557-3f4397f5e13a · outbound

This paper cites an unresolved cited work.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:08:51.659499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:08:51.487084Z digest=sha256:3fed1f06b7c1985681f698101730cf29dad9e189356d34baac71ad5a8d79a6a2

Observation 2c1ee46b-1ae9-46a4-a143-a916844f4bc0 · outbound

This paper cites an unresolved cited work.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:08:51.653036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:08:51.489343Z digest=sha256:0dfbccea7c368f2c11eb0ed1824cab02a7b6a81459b185e8a7a5f2456f907c20

Observation 458bdb31-935d-4f0f-94e9-e4bdf654f928 · outbound

This paper cites an unresolved cited work.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:08:51.645088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:08:51.491407Z digest=sha256:31b1189f8099272c41efbd19640c3c3aa3eb16d6c1fc8a084658e308f45b85b3

Observation 24a70173-6466-4293-901b-0e2ab8cb358a · outbound

This paper cites an unresolved cited work.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:08:51.637584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:08:51.493247Z digest=sha256:b390e3bb77c9e95775f80aea83a371546045925a749d90326978cb3a05a13efd

Observation 941f93b4-b6a6-430f-820b-847cf1e2ffa9 · outbound

This paper cites Please answer yes or no.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Please answer yes or no

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:08:51.629549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:08:51.495834Z digest=sha256:c8e1748a914e65c03b27d7ea1b65ead3646f7b3366c09c0a56320c8c09b686c2

Pith citing papers

Observation b7d78133-b0b0-44da-9a83-3dc8da27fa44 · inbound

Investigating Mechanisms for In-Context Vision Language Binding cites this paper.

Investigating Mechanisms for In-Context Vision Language Binding What's in the Image? A Deep-Dive into the Vision of Vision Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:29.226330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:29.226330Z digest=sha256:0e5da7ccbab9058313b525b94cd96c15526110f1de972d73df47138448e50334

Observation fd8974b9-4d49-436a-8611-a22b476e4d89 · inbound

PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning cites this paper.

PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning What's in the Image? A Deep-Dive into the Vision of Vision Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:39:01.636238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:39:01.636238Z digest=sha256:b1fa8e62744102fd500db379e9c7ca1dc8aff15a441c9bb732fbe928c5b99110

Observation 42272f09-424f-4567-976e-8fde399d364d · inbound

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling cites this paper.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling What's in the Image? A Deep-Dive into the Vision of Vision Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:23.993702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:23.993702Z digest=sha256:19e271eb9d800302a2a711cebf09a6094467cd5ef85e8b1cd57e717263d9378b

Observation fdb07771-8a63-4754-bea1-7ffeaf987dc0 · inbound

Counting to Four is still a Chore for VLMs cites this paper.

Counting to Four is still a Chore for VLMs What's in the Image? A Deep-Dive into the Vision of Vision Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:06:03.297518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T15:39:06.327888Z digest=sha256:b34fe5311b8783425d5c2875f5a2e0b8fbb41353a6ca86b8d3a475feaeceb300

Observation c3bcf70c-4763-4c29-890f-db5f2eb90ecc · inbound

Vision-Default, Prior-Override: Causal Mechanisms of Perception-Knowledge Conflict in Vision-Language Models cites this paper.

Vision-Default, Prior-Override: Causal Mechanisms of Perception-Knowledge Conflict in Vision-Language Models What's in the Image? A Deep-Dive into the Vision of Vision Language Models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T17:15:52.050405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T03:53:04.984303Z digest=sha256:6c867c38213de172befbf4e16d7aad3cdddcfff34b33f84fda201c86fee60daf

Observation f099b476-2c82-4664-bae0-8de9a7610a42 · inbound

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders cites this paper.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders What's in the Image? A Deep-Dive into the Vision of Vision Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-08T05:54:33.577235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:1f33658a19922f875d4ddff44777337164bf70672373bbd599b925564d6e74cf

Observation f90f76b3-eed1-4a3b-a13a-947ff61d0e32 · inbound

In-Context Collapse in Vision-Language Models and How to Mitigate it? cites this paper.

In-Context Collapse in Vision-Language Models and How to Mitigate it? What's in the Image? A Deep-Dive into the Vision of Vision Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T15:05:19.272138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:05:19.272138Z digest=sha256:5b49dea0d802dfb7578ef55a960ce0b2f64d67186b973819860994704f581cdb

Observation 72836731-785e-4f17-8230-0bacb9b81695 · inbound

Routing Is Least Learnable Where It Is Most Valuable: Bounds on Representation Routing for Web Agents cites this paper.

Routing Is Least Learnable Where It Is Most Valuable: Bounds on Representation Routing for Web Agents What's in the Image? A Deep-Dive into the Vision of Vision Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:52.051427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:52.051427Z digest=sha256:2f39c9ccc2a72f433fccf2b3d556a25715f899a4e83d1bab814d440b546f3903