Pith. sign in

Paper Citation Record · LEDGER

Why are Visually-Grounded Language Models Bad at Image Classification?

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2405.18415.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.18415 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:32:59.911874Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T18:26:55.252818Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a00e4c4e-a295-454d-ac90-bf1040b28d75 · inbound

Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation cites this paper.

Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation Why are Visually-Grounded Language Models Bad at Image Classification?

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:26:55.254753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T18:26:12.597756Z digest=sha256:fc42263a55cd2569762051f15a46cd095a3f7648c67534667062b411e9e1338b

Observation 4abf9151-01b7-4781-ac18-c573087a4c8d · inbound

AIGI-Holmes: Towards Explainable and Generalizable AI-Generated Image Detection via Multimodal Large Language Models cites this paper.

AIGI-Holmes: Towards Explainable and Generalizable AI-Generated Image Detection via Multimodal Large Language Models Why are Visually-Grounded Language Models Bad at Image Classification?

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-06T20:32:59.911874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:32:59.911874Z digest=sha256:4dc459d1bc136e3729947fd83c46e0156ef4d8f634b30774155d0c9efcb7b6e3

Observation 2f4deb52-f0f0-4b76-8427-1491df391ee8 · inbound

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models cites this paper.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Why are Visually-Grounded Language Models Bad at Image Classification?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:46.828449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:46.828449Z digest=sha256:18fca6fdcde45c69f3739a8ddf72246f7879df39bd5815349194e6f8759af672

Observation 5b47dfc0-6f42-40af-9456-f23975f858c7 · inbound

Filter-And-Refine: A MLLM Based Cascade System for Industrial-Scale Video Content Moderation cites this paper.

Filter-And-Refine: A MLLM Based Cascade System for Industrial-Scale Video Content Moderation Why are Visually-Grounded Language Models Bad at Image Classification?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T14:58:44.118082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:58:44.118082Z digest=sha256:1f0a9b07cc71a5bb8b99f98ab3f657da921cd37a4657cfb22237c8da0ef4b32f

Observation 9df4ce5f-c840-4476-becc-a370ef38184e · inbound

Decomposing Visual Classification: Assessing Tree-Based Reasoning in VLMs cites this paper.

Decomposing Visual Classification: Assessing Tree-Based Reasoning in VLMs Why are Visually-Grounded Language Models Bad at Image Classification?

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T20:28:22.267079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T20:28:22.267079Z digest=sha256:2e97252684d346e8a60861e420140ce8f02f8dbbe9bf2f84cfd63bde342f6874

Observation 118e2245-8dfb-4c81-b2d5-eeebd7a5a3b7 · inbound

Unpacking Hateful Memes: Presupposed Context and False Claims cites this paper.

Unpacking Hateful Memes: Presupposed Context and False Claims Why are Visually-Grounded Language Models Bad at Image Classification?

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T10:26:23.813098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:26:23.813098Z digest=sha256:5820eb2b90fcbdf1fc5124d3b4958b2ae090076bf7aaf1ec1f7a8d4acb490bbe

Observation f71b06a0-b4fc-496a-a091-da1324cc9c3b · inbound

Foundation Models for Astrophysics cites this paper.

Foundation Models for Astrophysics Why are Visually-Grounded Language Models Bad at Image Classification?

Reference 146

Resolution
unresolved
no resolver link, observed 2026-08-04T04:31:49.155307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:31:49.155307Z digest=sha256:e5cc6600c777f1d85f9c7063434528686be9d4d0ae2faeaaaee6a0a274f832f2