Pith. sign in

Paper Citation Record · LEDGER

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities

As of 11 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 4 inbound Pith citation observations for arXiv:2412.16418.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.16418 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T10:41:37.913674Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-17T05:53:26.066674Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T05:53:26.308450Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact1
  • verified fuzzy38
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 26d9fdff-b43a-4026-ad45-64adff441f29 · outbound

This paper cites Phi-3 technical report: A highly capable language model locally on your phone, 2024.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Phi-3 technical report: A highly capable language model locally on your phone, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.804660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.645085Z digest=sha256:249a5f2e7965f3114ece53f823473052d70c9c8f0858920c53c0574c9f223b24

Observation ee6d380d-4759-4055-88ad-3cf8158c5df3 · outbound

This paper cites GPT-4 Technical Report.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.649338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.649338Z digest=sha256:10f245db81876a7927a0de6ea11cbea7e97e607ef00a22ada9aeae435bcba508

Observation dd44ae15-887c-4550-ba9d-b1f3d11e1f76 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.654130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.654130Z digest=sha256:e4d5f6643be625804bf9e355682df5a0c0826d35a968c9bb1452a6f02c1c448f

Observation b147c20f-2529-43f9-a83b-fe3e3809a35d · outbound

This paper cites Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models.NeurIPS, 32, 2019.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models.NeurIPS, 32, 2019

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.791334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.659252Z digest=sha256:5ae39357238c72899c993c51fd9b8951e0b5e87742cdbb93ae4c4014697346d8

Observation 5f25b32c-41a3-4ab6-8391-260bd1192543 · outbound

This paper cites Food-101–mining discriminative components with random forests.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Food-101–mining discriminative components with random forests

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.773812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.663576Z digest=sha256:c0914d679ce46b655694129cf43281bd7984204dad6518b65250a8751fcf86c6

Observation f25dc983-710a-40ce-b637-d287d267c968 · outbound

This paper cites Are we on the right way for evaluating large vision-language models? In NeurIPS, 2024.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Are we on the right way for evaluating large vision-language models? In NeurIPS, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.757562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.667758Z digest=sha256:c5fa217d3130300690b0b594caa267f6ea0c2f888c8806e5cf09b7e19c8076f7

Observation f2bc42e5-e1e0-44fa-b969-89a1d01cacd5 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.671719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.671719Z digest=sha256:4364d6fdfec5c62bb90f3fbe31df3e68d29597b1815fbc3de00419df43f6be49

Observation 512549b8-9fb6-46fb-9c55-0f9f7350d085 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.676904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.676904Z digest=sha256:484405878edd3798fcdff52c0d144953dba515332fec10c7c8656c33fb78e825

Observation 7e37559a-9fd4-46f4-93f4-b1a016a384ff · outbound

This paper cites Reproducible scal- ing laws for contrastive language-image learning.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Reproducible scal- ing laws for contrastive language-image learning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.742658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.680980Z digest=sha256:f61541fdbe13d128253f39f19a4f2e4b6f476d63376eb3cd5aa1814c084694de

Observation 649b899d-acfc-4e74-ac00-ec9a4666e795 · outbound

This paper cites Coatnet: Marrying convolution and attention for all data sizes.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Coatnet: Marrying convolution and attention for all data sizes

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.728961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.685448Z digest=sha256:171b63d241de6d096b4d7201445561f62256fed56e8570ba254757c33dddabdb

Observation 1fae459d-ba51-44ce-bd4f-142eac21a8d3 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Imagenet: A large-scale hierarchical image database

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.712308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.691943Z digest=sha256:131e701be23e1671121b80280a07b3b91df5ba0ecb5bde0f9da819ca43f83096

Observation 2f3049f3-2cba-46e3-b277-3daad7465ce5 · outbound

This paper cites BERT: pre-training of deep bidirectional trans- formers for language understanding.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities BERT: pre-training of deep bidirectional trans- formers for language understanding

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.697360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.696993Z digest=sha256:d1e8ce5ccf2260ebdc47c58f7f8f676e5d7b6255a6abf0e089e34e01679c38c7

Observation 5f8f7f39-9aa5-4eda-9b92-0d01a523c50f · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities An image is worth 16x16 words: Transformers for image recognition at scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.703159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.703159Z digest=sha256:a9eb9d3528c789f80de4951605c3af5111e0d3928e69e148cab3518b275e3a17

Observation 6565e1be-5820-43a8-965c-4c62d5f88e7a · outbound

This paper cites Data filtering networks.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Data filtering networks

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.671134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.707016Z digest=sha256:45381ec9d5f46706ac9207b3d82cedb38de9aa265380fde2c72aa6639314c25d

Observation f5617f5b-cba8-40ea-bf28-bd89e01680a8 · outbound

This paper cites Eva-02: A visual representa- tion for neon genesis.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Eva-02: A visual representa- tion for neon genesis

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.657904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.710856Z digest=sha256:a0d1ffb5e443e293d2b8d3f0a03d90d3ab176ee0189ae57b403987e1b7ab09d0

Observation 64e592e5-588d-48fc-86b7-210724a4ffa9 · outbound

This paper cites Learning gener- ative visual models from few training examples: An incre- mental bayesian approach tested on 101 object categories.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Learning gener- ative visual models from few training examples: An incre- mental bayesian approach tested on 101 object categories

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.642239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.715175Z digest=sha256:0cb2e224cecf8d6ee0757b532ceba9db988e58b332168f55727e8e48f6662df5

Observation 135ad5c1-ab1f-4405-9072-85df1f456097 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.719841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.719841Z digest=sha256:1bc84158062f47e0b723d042e6ef6525505eb4dd09dc11de3fe86387633b2b52

Observation bb13bc77-0bc9-42b1-a4ae-ccf7ddd93ff3 · outbound

This paper cites Detect, Describe, Discriminate: Moving Beyond VQA for MLLM Evaluation.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Detect, Describe, Discriminate: Moving Beyond VQA for MLLM Evaluation

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-11T10:41:38.039197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.724041Z digest=sha256:0eb19f483909838de7ee3e336c50ec961048ca9980eb5dcaf72eb8dd186528b5

Observation a3306c69-e067-41c6-ab6a-20f523e98067 · outbound

This paper cites Deep residual learning for image recognition.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Deep residual learning for image recognition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.728454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.728454Z digest=sha256:eb688cbaad2a415b161db935627d2d256e3eef19a7585a092469855bb814102c

Observation 5a58e4a6-c8f9-4894-b8f4-02cd67bcf147 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.618633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.732817Z digest=sha256:98698d313eac96c3bdba56200d5fd76eef513dc70d88a96a88b55d35f6827031

Observation d705921b-d5a6-452f-942b-105c522a0042 · outbound

This paper cites Rwku: Benchmarking real-world knowledge unlearning for large language models.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Rwku: Benchmarking real-world knowledge unlearning for large language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.605236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.736467Z digest=sha256:dfc54ee000898a0dfa32723adeeb2874dbd4acfbbf78ef806ff537f20e81629b

Observation 6bdb4fae-00c0-4d45-9dce-5c5d20c3a406 · outbound

This paper cites A diagram is worth a dozen images.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities A diagram is worth a dozen images

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.591704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.739962Z digest=sha256:091caea88dc179e167ec5f2b36f09b99f8da4e4d69f51df455a5e7d60d3e42fd

Observation 147fb193-4c30-42cd-b226-900d5f05ed03 · outbound

This paper cites 3d object representations for fine-grained categorization.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities 3d object representations for fine-grained categorization

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.576988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.743871Z digest=sha256:83cdd62b75654d896c5f57b10b7e6108e59ebb1c92a4049dcfbc35ba0383d511

Observation 423d313b-a984-48b6-8315-49018b59265f · outbound

This paper cites Learning multiple layers of features from tiny images.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Learning multiple layers of features from tiny images

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.747756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.747756Z digest=sha256:f96dcb7e85640c268e4cc02cce0e4dc65aeb93e968af0f1c0ea7d80a2abe65df

Observation 2880705b-cf22-48c6-955e-78b6ed7c7f28 · outbound

This paper cites Imagenet classification with deep convolutional neural net- works.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Imagenet classification with deep convolutional neural net- works

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.751786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.751786Z digest=sha256:50b028e1a46f341724cba40da137e259b8858a0b50e7c392a804d8d5e466b3b9

Observation fb3ff337-0db6-4da1-8332-27cf3950accf · outbound

This paper cites Gradient-based learning applied to document recog- nition.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Gradient-based learning applied to document recog- nition

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.544978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.756024Z digest=sha256:2a7d6c4a1319cc9edaff1fd3900543c3664b6a5905aa396d06baff73627e85e9

Observation 4763ab78-4416-492a-baa7-2c9d8e3126f4 · outbound

This paper cites Seed-bench: Bench- marking multimodal large language models.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Seed-bench: Bench- marking multimodal large language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.532234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.760545Z digest=sha256:ad0ffcc9ef6096c7dc551bbc5c2de7de42175c13fde25a52f4a902e7ec7c37ac

Observation d7f22d14-e1ba-485c-a2d6-23804eac654f · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities LLaVA-OneVision: Easy Visual Task Transfer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.764851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.764851Z digest=sha256:bf9b9d2af8fa9c36d9a96bad8132fae76f3c2c2187cf9f50cd5c6289b5c83769

Observation eb5755bc-79ee-4029-99a4-ccd84213866c · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.517939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.769002Z digest=sha256:04bf99c03120e8e1a46bdbfe096d82581de5b77d6f2a36129e7135beeb963a0f

Observation 1eefded5-450b-469d-8753-2b4f87f91387 · outbound

This paper cites Visual instruction tuning.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Visual instruction tuning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.504194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.772807Z digest=sha256:41282c6e31429f35df7689cac394007574b0cea5574ccb00d3aeef9dc92485e2

Observation f05a396c-3dcf-43b6-8376-e2cc2e03942c · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In ECCV, pages 216–233.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Mmbench: Is your multi-modal model an all-around player? In ECCV, pages 216–233

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.490407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.777788Z digest=sha256:18d762f8ca4f26134c1f345533ad28ddffc5605e58f7e18c7d0fca88af63f04f

Observation b125040d-82e0-47d1-8623-32f40ce15112 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Swin transformer: Hierarchical vision transformer using shifted windows

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.781644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.781644Z digest=sha256:8f8ebb089c9707258b477097bbe1a2bad229880f7471bccff743a5bd1f9408ba

Observation 7a9eb696-d3e2-4342-b53f-63f7a724c2b7 · outbound

This paper cites Swin transformer v2: Scaling up capacity and resolution.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Swin transformer v2: Scaling up capacity and resolution

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.466547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.785576Z digest=sha256:955d2e00484ce817856a23f3e8d6cc3c03531e62615d046899dfee30ef019ff4

Observation 680c1ba2-1e15-40d8-b18e-392a34f8462c · outbound

This paper cites Deepseek-vl: Towards real-world vision- language understanding, 2024.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Deepseek-vl: Towards real-world vision- language understanding, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.452406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.789525Z digest=sha256:6407f1bb1f7de151daeb15f468006564c9ec9fc958d6af0985af95329423343f

Observation 6c907534-b710-4e45-a52d-deb2507b3eb7 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.793312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.793312Z digest=sha256:18fec4fb67371d2737e1147752c908a0c2f1fdb07b5fa70ce4b3ba935770883b

Observation ce63c050-1879-4730-b950-31b8bdf8d301 · outbound

This paper cites ChartQA: A benchmark for question answering about charts with visual and logical reasoning.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities ChartQA: A benchmark for question answering about charts with visual and logical reasoning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.428526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.797816Z digest=sha256:d60883d04d7ecd4a567218b945b3c2ed859c169c049b3f2b202d5429a0e38426

Observation fc28fc7b-fd75-4e1b-9d14-938ab88e8b51 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Docvqa: A dataset for vqa on document images

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.414259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.801625Z digest=sha256:ed73702e970f81dab315dd6780efd6e2a50864881fdf55e3d562478e40d028bf

Observation 97fc78a3-17c7-4c73-a309-f0126aae063f · outbound

This paper cites Automated flower classification over a large number of classes.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Automated flower classification over a large number of classes

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.399973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.805381Z digest=sha256:4c3f7f5043a207c39dae7bc024d5bfb2e0f0d1f153acb6d0dc8518185571d24c

Observation d84a2c1d-a71c-48d1-bc91-ec675586df64 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Learn- ing transferable visual models from natural language super- vision

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.385355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.809342Z digest=sha256:ec9259792548f1e2c0477c4522641505d66c5635d88d4d3f15ec0ce70b15582b

Observation b0dda3b1-c61c-44bd-b8c3-145b132615cd · outbound

This paper cites Do imagenet classifiers generalize to im- agenet? In ICML, pages 5389–5400.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Do imagenet classifiers generalize to im- agenet? In ICML, pages 5389–5400

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.371153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.813258Z digest=sha256:45b74588abae70f9615a49a8e507f45ba88642cd5e54f219224780dc63e7cb8a

Observation 6583b487-47db-4cab-b85e-780423ad6ef8 · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.817566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.817566Z digest=sha256:0c4639cfdd0f19c36578a51a44503006b6b21999df88b34f358b909c626c1d97

Observation c5aa90d2-8c50-4c8a-88ae-cf42f07c6516 · outbound

This paper cites Going deeper with convolutions.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Going deeper with convolutions

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.358893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.821586Z digest=sha256:e9741d48b2396b77322e9ea2e2c05a436884b68f500bbd5961e765585daf99c0

Observation c39ea7b9-18ea-49d6-916b-5f101a0d3628 · outbound

This paper cites Jn-logo: A logo database for aesthetic visual analysis.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Jn-logo: A logo database for aesthetic visual analysis

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.346202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.825220Z digest=sha256:42d5763d62d9a75609858b9f4ea68655eb0b81ca0453ad71b6e4130005e234a8

Observation 77b04fab-f437-4f57-ab40-25f814b7bbf1 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.829252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.829252Z digest=sha256:8fc76dec752b9d4e80a0a540b251d3a71a7982c41a63ed852a94f8c80a12f4df

Observation a11be65f-a96a-478e-851d-dc087973079e · outbound

This paper cites The caltech-ucsd birds-200-2011 dataset.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities The caltech-ucsd birds-200-2011 dataset

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.333684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.833118Z digest=sha256:bb06074f021570772cd97a2eb53728583d1e086c0043b56b95a32186a1e08614

Observation 79c68e05-d7ef-4e85-a426-dde92e24ba70 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.836593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.836593Z digest=sha256:ea92db4698feb05023b285ed9ac5cd7ed66f2968913934d8877270dd890e154b

Observation c9a3af94-9db8-4a0f-b0e8-afcb4f4ae00b · outbound

This paper cites Google landmarks dataset v2-a large-scale benchmark for instance-level recognition and retrieval.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Google landmarks dataset v2-a large-scale benchmark for instance-level recognition and retrieval

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.320216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.840235Z digest=sha256:d2bac9ddd369e9a3def28d48a8deb5479f111111cfe573ecc59526f71d41c23f

Observation 6dfdb39c-d2fd-4d8b-a273-3684ffda1862 · outbound

This paper cites LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.843756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.843756Z digest=sha256:c32a3a58dbaa81864e79c3f61e011fe3f5347982e4e058f962bd5247b7a30849

Observation 54f2a3cb-6db3-47cf-9fa2-0eafe9ae4a9b · outbound

This paper cites Qwen2 Technical Report.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Qwen2 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.848153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.848153Z digest=sha256:8ff6dae6bab80f0a9de8baeca52291efacc4d9e99829562e3f147e1171c93bbf

Observation 7ea62737-c138-4423-bb54-245ca0e19438 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.851906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.851906Z digest=sha256:b8199e1325db4c24fa0781c4b79cfb6d6a61e30697ef2d3f95968a73f1b3d598

Observation 9376f618-2abc-453d-ae52-23869843af28 · outbound

This paper cites Lamm: Language-assisted multi- modal instruction-tuning dataset, framework, and bench- mark.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Lamm: Language-assisted multi- modal instruction-tuning dataset, framework, and bench- mark

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.305265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.855634Z digest=sha256:ee46f7306b095f3ec81252b02db28c944758525c892905205fdfd151229982c2

Observation e22b5e0b-38f6-45c2-afc9-6075f861c1f8 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.291495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.859312Z digest=sha256:666ed6600550237af23555976a899cf203a473b6a31f347ab0a4a99ede684e5e

Observation d5484383-abc7-42e2-a928-f59e8bdeebca · outbound

This paper cites Scaling vision transformers.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Scaling vision transformers

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.279108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.862680Z digest=sha256:429eca0979202f266a0a9911ae0a9b7acfe5e98bb0dbb4033dc4bd19283792c2

Observation 7a8448cb-a3e6-4239-aeaa-ca7bb0f2b90a · outbound

This paper cites Sigmoid loss for language image pre-training.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Sigmoid loss for language image pre-training

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.266554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.866304Z digest=sha256:ca67347d9f124af4e4a533524a26ef3c9dfa429130779ba41ec71511e42262c0

Observation 31162702-5e58-482f-b745-20a7f7a518b7 · outbound

This paper cites Why are visually-grounded language models bad at image classi- fication? In NeurIPS, 2024.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Why are visually-grounded language models bad at image classi- fication? In NeurIPS, 2024

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.254538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.869890Z digest=sha256:aa05951fa95f2426a8cbb021596ebafe374985c24921d0e90fe5225f97b2a095

Observation 546b4cde-2671-4a32-9206-c001ced541ec · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.241459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.872988Z digest=sha256:5ec24cb0b2743dd0177979941675657eb7114a3cd40189578d84d412c425e17d

Observation 31a23f9b-ec1f-4a5c-bec6-052b72921c54 · outbound

This paper cites an unresolved cited work.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:41:38.214146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.881848Z digest=sha256:1687e42b1488f8c6c5855c4c829a9defd31a4db3758177c7207ae9ea186b8a13

Observation 7636536f-80a8-4b01-8360-a7b6d0483495 · outbound

This paper cites an unresolved cited work.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:41:38.201467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.887053Z digest=sha256:08cce7ec3571d0d8d07131dbc0df9c04abe25c198c56f2470c89ee35ec6a8ac2

Observation 79d88115-2d8d-4c8e-8b7a-2f93aea6b1e4 · outbound

This paper cites an unresolved cited work.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:41:38.188362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.891584Z digest=sha256:42d8002e3b66e39dfa6073593d1a5a5497a3c91f06e94c561e418ec954f83ec6

Observation 7cc51c5e-d78d-4889-a52a-4d0b5105ccad · outbound

This paper cites an unresolved cited work.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:41:38.175138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.895924Z digest=sha256:f9eb7d3fd7cc80005163c51a29aae2106cccde1a891fe8b506b4202162573fc2

Observation 7b16d058-1a34-4f6c-a60f-d65e9a096e6b · outbound

This paper cites an unresolved cited work.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:41:38.160953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.899858Z digest=sha256:6a60ae6d295eeaf8512d5b3ee4a1b999d707d37044278ac263a5e95de3284db7

Observation 4d9a79a2-fb2e-492d-af42-4d324ab77001 · outbound

This paper cites an unresolved cited work.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:41:38.148659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.904364Z digest=sha256:cd23e79d90ec73e7b548672dc0cabb1c365804da278b38a070f591b36efded9b

Observation b383dafe-c24e-4f2e-8852-d03ff902e144 · outbound

This paper cites an unresolved cited work.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:41:38.135247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.908611Z digest=sha256:cb7b81c5849b3d5a133b12f3f13756126418ad952f3e326e8ca45845acfc5a1c

Observation 7de3ba91-6a0f-4777-b72a-95f5049e0ed0 · outbound

This paper cites an unresolved cited work.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:41:38.120651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.913674Z digest=sha256:49ece47aee81531374bea55d62e08b4b455e2b4be05697fa0410c9e709816473

Observation fc92c4ea-6741-4ebb-9dc6-2d69d2c97d3c · outbound

This paper cites Then, we perform an in-depth exploration of MLLM classification evaluation (delineated in Section B), including the formulation, influence of option numbers, etc.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Then, we perform an in-depth exploration of MLLM classification evaluation (delineated in Section B), including the formulation, influence of option numbers, etc

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.227596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T10:41:37.877198Z digest=sha256:0f407ed88511be9bfccbf208a1645075666252a27e2a7fea8e20bcda28dbdd52

Pith citing papers

Observation c79c90a0-9f56-44e8-9ae0-b82a1af23dd6 · inbound

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs cites this paper.

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:53:26.314287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T05:53:26.066674Z digest=sha256:f8e7bb2c9009da2063ec4ddace79f8d94cb5f1db65129db35950b5c552409c7b

Observation cd99320d-16e6-4664-8f5c-b047690979af · inbound

Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning cites this paper.

Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:17:26.785959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T06:13:33.315525Z digest=sha256:a1a1ec2aa4c44a7bb6cde9eabc6cd4bd1a7d2c23652c80b0655974dcfc138750

Observation 6ed6e79d-0ae8-403b-81a3-9aeafa49eb0f · inbound

Specificity-aware reinforcement learning for fine-grained open-world classification cites this paper.

Specificity-aware reinforcement learning for fine-grained open-world classification Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:46:17.940206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T16:44:47.866126Z digest=sha256:ec2af1a7f96fefdc1c4bd9d32889d38477170d4fb17340382c507df76cd2fd75

Observation a966abb4-9133-462e-b8d5-fa9c4869ed82 · inbound

Bridging Coarse and Fine Recognition: A Hybrid Approach for Open-Ended Multi-Granularity Object Recognition in Interactive Educational Games cites this paper.

Bridging Coarse and Fine Recognition: A Hybrid Approach for Open-Ended Multi-Granularity Object Recognition in Interactive Educational Games Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:47:12.229585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T07:46:39.249226Z digest=sha256:cf43cdb03b3e710c5b6330d5e561261d19c5e913318b6ad929517b58cb2e10e4