Pith. sign in

Paper Citation Record · LEDGER

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities

As of 12 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 4 inbound Pith citation observations for arXiv:2412.16418.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.16418 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T10:41:37.913674Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-17T05:53:26.066674Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T05:53:26.308450Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact1
  • verified fuzzy38
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 26d9fdff-b43a-4026-ad45-64adff441f29 · outbound

This paper cites Phi-3 technical report: A highly capable language model locally on your phone, 2024.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Phi-3 technical report: A highly capable language model locally on your phone, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.804660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.645085Z digest=sha256:ad1e8eae0f56456f1ed74f484889b47e0989aacf14437550f87843ab44b4aeec

Observation ee6d380d-4759-4055-88ad-3cf8158c5df3 · outbound

This paper cites GPT-4 Technical Report.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.649338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.649338Z digest=sha256:10f245db81876a7927a0de6ea11cbea7e97e607ef00a22ada9aeae435bcba508

Observation dd44ae15-887c-4550-ba9d-b1f3d11e1f76 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.654130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.654130Z digest=sha256:e4d5f6643be625804bf9e355682df5a0c0826d35a968c9bb1452a6f02c1c448f

Observation b147c20f-2529-43f9-a83b-fe3e3809a35d · outbound

This paper cites Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models.NeurIPS, 32, 2019.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models.NeurIPS, 32, 2019

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.791334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.659252Z digest=sha256:400a0f08e603166c10eaa2409c4c233e80e2661530e89cc032ff9b18414a69f7

Observation 5f25b32c-41a3-4ab6-8391-260bd1192543 · outbound

This paper cites Food-101–mining discriminative components with random forests.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Food-101–mining discriminative components with random forests

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.773812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.663576Z digest=sha256:7a2691893f09f8c715f5fc4e63b8e723985e6b683f51be94cc0dd6ae4c91b8b5

Observation f25dc983-710a-40ce-b637-d287d267c968 · outbound

This paper cites Are we on the right way for evaluating large vision-language models? In NeurIPS, 2024.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Are we on the right way for evaluating large vision-language models? In NeurIPS, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.757562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.667758Z digest=sha256:d9c2c4299cc3bc2ae5f8d706997c8e67438d91dd8067a62ae64800236ae8703d

Observation f2bc42e5-e1e0-44fa-b969-89a1d01cacd5 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.671719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.671719Z digest=sha256:4364d6fdfec5c62bb90f3fbe31df3e68d29597b1815fbc3de00419df43f6be49

Observation 512549b8-9fb6-46fb-9c55-0f9f7350d085 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.676904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.676904Z digest=sha256:484405878edd3798fcdff52c0d144953dba515332fec10c7c8656c33fb78e825

Observation 7e37559a-9fd4-46f4-93f4-b1a016a384ff · outbound

This paper cites Reproducible scal- ing laws for contrastive language-image learning.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Reproducible scal- ing laws for contrastive language-image learning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.742658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.680980Z digest=sha256:325a6244178b566ee2ec2cc5fa956325ebfb5a814ba01957db77731c7c681f53

Observation 649b899d-acfc-4e74-ac00-ec9a4666e795 · outbound

This paper cites Coatnet: Marrying convolution and attention for all data sizes.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Coatnet: Marrying convolution and attention for all data sizes

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.728961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.685448Z digest=sha256:92f7f3bc7aac230c2d9577fe585f2870a8da6abc3107f049316d0f5fb986b7c3

Observation 1fae459d-ba51-44ce-bd4f-142eac21a8d3 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Imagenet: A large-scale hierarchical image database

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.712308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.691943Z digest=sha256:316fa0be4de659f04bedc0bbb52868bc58fb2938562b70342e371867e671ebed

Observation 2f3049f3-2cba-46e3-b277-3daad7465ce5 · outbound

This paper cites BERT: pre-training of deep bidirectional trans- formers for language understanding.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities BERT: pre-training of deep bidirectional trans- formers for language understanding

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.697360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.696993Z digest=sha256:c4b970b53860fb8fc617ae63fbc7f86c48f8ced1b7a10a6d578238ae79581154

Observation 5f8f7f39-9aa5-4eda-9b92-0d01a523c50f · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities An image is worth 16x16 words: Transformers for image recognition at scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.703159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.703159Z digest=sha256:a9eb9d3528c789f80de4951605c3af5111e0d3928e69e148cab3518b275e3a17

Observation 6565e1be-5820-43a8-965c-4c62d5f88e7a · outbound

This paper cites Data filtering networks.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Data filtering networks

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.671134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.707016Z digest=sha256:4e56d1142657c97699772a3c217130bdcb759f26fc15009cc71977f5c9b95037

Observation f5617f5b-cba8-40ea-bf28-bd89e01680a8 · outbound

This paper cites Eva-02: A visual representa- tion for neon genesis.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Eva-02: A visual representa- tion for neon genesis

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.657904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.710856Z digest=sha256:8aab3a3adca650c418c4d6b15b91e53948b018c3ac57fa15d43f9b1b47da27ad

Observation 64e592e5-588d-48fc-86b7-210724a4ffa9 · outbound

This paper cites Learning gener- ative visual models from few training examples: An incre- mental bayesian approach tested on 101 object categories.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Learning gener- ative visual models from few training examples: An incre- mental bayesian approach tested on 101 object categories

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.642239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.715175Z digest=sha256:2d9ae582267fc47d83a775c1689182f0cbf7e2ba6dfe0383943902d9d77cd41e

Observation 135ad5c1-ab1f-4405-9072-85df1f456097 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.719841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.719841Z digest=sha256:1bc84158062f47e0b723d042e6ef6525505eb4dd09dc11de3fe86387633b2b52

Observation bb13bc77-0bc9-42b1-a4ae-ccf7ddd93ff3 · outbound

This paper cites Detect, Describe, Discriminate: Moving Beyond VQA for MLLM Evaluation.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Detect, Describe, Discriminate: Moving Beyond VQA for MLLM Evaluation

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-11T10:41:38.039197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.724041Z digest=sha256:75d876f3f3b5e613e65863fb05f6c66d023bb3c3688a26224a158c424b041a3d

Observation a3306c69-e067-41c6-ab6a-20f523e98067 · outbound

This paper cites Deep residual learning for image recognition.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Deep residual learning for image recognition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.728454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.728454Z digest=sha256:eb688cbaad2a415b161db935627d2d256e3eef19a7585a092469855bb814102c

Observation 5a58e4a6-c8f9-4894-b8f4-02cd67bcf147 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.618633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.732817Z digest=sha256:272a04e5716e75dcbe0341e0fbccaa97836a4946f6eac0ac7d4a357a26ac0095

Observation d705921b-d5a6-452f-942b-105c522a0042 · outbound

This paper cites Rwku: Benchmarking real-world knowledge unlearning for large language models.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Rwku: Benchmarking real-world knowledge unlearning for large language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.605236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.736467Z digest=sha256:3153fd54b8c2ff63317f57e7236b885663bf5d601f51507db8d28f5f1bb5e3ec

Observation 6bdb4fae-00c0-4d45-9dce-5c5d20c3a406 · outbound

This paper cites A diagram is worth a dozen images.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities A diagram is worth a dozen images

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.591704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.739962Z digest=sha256:fd86ddd8f702b895dd2b7066e542de2f787c77f2c244ecaa8baef53df7e668e0

Observation 147fb193-4c30-42cd-b226-900d5f05ed03 · outbound

This paper cites 3d object representations for fine-grained categorization.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities 3d object representations for fine-grained categorization

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.576988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.743871Z digest=sha256:e03c85b99c7f84eb73c444e5014cd20ce472f6f1159a943965a8a178128ec371

Observation 423d313b-a984-48b6-8315-49018b59265f · outbound

This paper cites Learning multiple layers of features from tiny images.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Learning multiple layers of features from tiny images

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.747756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.747756Z digest=sha256:f96dcb7e85640c268e4cc02cce0e4dc65aeb93e968af0f1c0ea7d80a2abe65df

Observation 2880705b-cf22-48c6-955e-78b6ed7c7f28 · outbound

This paper cites Imagenet classification with deep convolutional neural net- works.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Imagenet classification with deep convolutional neural net- works

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.751786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.751786Z digest=sha256:50b028e1a46f341724cba40da137e259b8858a0b50e7c392a804d8d5e466b3b9

Observation fb3ff337-0db6-4da1-8332-27cf3950accf · outbound

This paper cites Gradient-based learning applied to document recog- nition.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Gradient-based learning applied to document recog- nition

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.544978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.756024Z digest=sha256:cf797f15e7f4c56c5fab431e0fef0d837eaf2e571d985e3af2f54b6da099db3d

Observation 4763ab78-4416-492a-baa7-2c9d8e3126f4 · outbound

This paper cites Seed-bench: Bench- marking multimodal large language models.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Seed-bench: Bench- marking multimodal large language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.532234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.760545Z digest=sha256:72bd69a2c51bbbd1a20a850ab7c677e9c52aed16302366cf65c2a531cf588c67

Observation d7f22d14-e1ba-485c-a2d6-23804eac654f · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities LLaVA-OneVision: Easy Visual Task Transfer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.764851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.764851Z digest=sha256:bf9b9d2af8fa9c36d9a96bad8132fae76f3c2c2187cf9f50cd5c6289b5c83769

Observation eb5755bc-79ee-4029-99a4-ccd84213866c · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.517939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.769002Z digest=sha256:aa884464e313552dbe86045e56ed416d5545a476c2937134b9af46f1a6097e26

Observation 1eefded5-450b-469d-8753-2b4f87f91387 · outbound

This paper cites Visual instruction tuning.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Visual instruction tuning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.504194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.772807Z digest=sha256:30dbee17289f1a6eb8d12cc62471966294e9e74db16f40c7e7794616bf997d54

Observation f05a396c-3dcf-43b6-8376-e2cc2e03942c · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In ECCV, pages 216–233.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Mmbench: Is your multi-modal model an all-around player? In ECCV, pages 216–233

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.490407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.777788Z digest=sha256:cad8133ff3382982e928ae175c49e4453ca2096d5f6deb04d0e75c2e525c45a2

Observation b125040d-82e0-47d1-8623-32f40ce15112 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Swin transformer: Hierarchical vision transformer using shifted windows

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.781644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.781644Z digest=sha256:8f8ebb089c9707258b477097bbe1a2bad229880f7471bccff743a5bd1f9408ba

Observation 7a9eb696-d3e2-4342-b53f-63f7a724c2b7 · outbound

This paper cites Swin transformer v2: Scaling up capacity and resolution.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Swin transformer v2: Scaling up capacity and resolution

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.466547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.785576Z digest=sha256:a1cdb4b01b3b1f68f5974552e92195bfd6511cddb603e9f3ccfe7bf714aba578

Observation 680c1ba2-1e15-40d8-b18e-392a34f8462c · outbound

This paper cites Deepseek-vl: Towards real-world vision- language understanding, 2024.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Deepseek-vl: Towards real-world vision- language understanding, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.452406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.789525Z digest=sha256:6a97693ade94d0022b696b5d6310e7c373b3a819ad3666c896299729dd9847be

Observation 6c907534-b710-4e45-a52d-deb2507b3eb7 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.793312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.793312Z digest=sha256:18fec4fb67371d2737e1147752c908a0c2f1fdb07b5fa70ce4b3ba935770883b

Observation ce63c050-1879-4730-b950-31b8bdf8d301 · outbound

This paper cites ChartQA: A benchmark for question answering about charts with visual and logical reasoning.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities ChartQA: A benchmark for question answering about charts with visual and logical reasoning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.428526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.797816Z digest=sha256:cd2cd83783e50afca29f981f062060a5559c9aaa12fbe29d87bbffe418a8ce6b

Observation fc28fc7b-fd75-4e1b-9d14-938ab88e8b51 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Docvqa: A dataset for vqa on document images

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.414259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.801625Z digest=sha256:690bf379486fee7baca2607a158ae98ac382f5be746fc88216d1ccb75575bc54

Observation 97fc78a3-17c7-4c73-a309-f0126aae063f · outbound

This paper cites Automated flower classification over a large number of classes.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Automated flower classification over a large number of classes

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.399973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.805381Z digest=sha256:89edcded0d79eed356d549dd6fda74b3f8df6afc3b492989c58b8bbba45f75a0

Observation d84a2c1d-a71c-48d1-bc91-ec675586df64 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Learn- ing transferable visual models from natural language super- vision

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.385355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.809342Z digest=sha256:cff8693b55c12ea9551f882c9a5799d1a11da3cffc6f0e557be929ed0457cac3

Observation b0dda3b1-c61c-44bd-b8c3-145b132615cd · outbound

This paper cites Do imagenet classifiers generalize to im- agenet? In ICML, pages 5389–5400.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Do imagenet classifiers generalize to im- agenet? In ICML, pages 5389–5400

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.371153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.813258Z digest=sha256:0e42297a992140c5775b764d27d8a2ed9dccc0318eabc1b8c986a9448c2ddc62

Observation 6583b487-47db-4cab-b85e-780423ad6ef8 · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.817566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.817566Z digest=sha256:0c4639cfdd0f19c36578a51a44503006b6b21999df88b34f358b909c626c1d97

Observation c5aa90d2-8c50-4c8a-88ae-cf42f07c6516 · outbound

This paper cites Going deeper with convolutions.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Going deeper with convolutions

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.358893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.821586Z digest=sha256:fdad764ccbfb9910ea55ea8849ee9bd2bd0b25076c07c7c1cf2596c98b4116e0

Observation c39ea7b9-18ea-49d6-916b-5f101a0d3628 · outbound

This paper cites Jn-logo: A logo database for aesthetic visual analysis.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Jn-logo: A logo database for aesthetic visual analysis

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.346202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.825220Z digest=sha256:92c2abd3b40baf7c2ad755ae9d6a725b14a0a2d8e42eb10bb88fafb228b6141a

Observation 77b04fab-f437-4f57-ab40-25f814b7bbf1 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.829252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.829252Z digest=sha256:8fc76dec752b9d4e80a0a540b251d3a71a7982c41a63ed852a94f8c80a12f4df

Observation a11be65f-a96a-478e-851d-dc087973079e · outbound

This paper cites The caltech-ucsd birds-200-2011 dataset.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities The caltech-ucsd birds-200-2011 dataset

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.333684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.833118Z digest=sha256:3fb6ee7c69b46e4949038dc961193068682712ec5e82e8339556432e64b7d5b7

Observation 79c68e05-d7ef-4e85-a426-dde92e24ba70 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.836593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.836593Z digest=sha256:ea92db4698feb05023b285ed9ac5cd7ed66f2968913934d8877270dd890e154b

Observation c9a3af94-9db8-4a0f-b0e8-afcb4f4ae00b · outbound

This paper cites Google landmarks dataset v2-a large-scale benchmark for instance-level recognition and retrieval.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Google landmarks dataset v2-a large-scale benchmark for instance-level recognition and retrieval

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.320216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.840235Z digest=sha256:4d9f8af72a6b248e758d801c6c8a1eee283911a30694305d01f5d1df96d55e3b

Observation 6dfdb39c-d2fd-4d8b-a273-3684ffda1862 · outbound

This paper cites LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.843756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.843756Z digest=sha256:c32a3a58dbaa81864e79c3f61e011fe3f5347982e4e058f962bd5247b7a30849

Observation 54f2a3cb-6db3-47cf-9fa2-0eafe9ae4a9b · outbound

This paper cites Qwen2 Technical Report.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Qwen2 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.848153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.848153Z digest=sha256:8ff6dae6bab80f0a9de8baeca52291efacc4d9e99829562e3f147e1171c93bbf

Observation 7ea62737-c138-4423-bb54-245ca0e19438 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.851906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.851906Z digest=sha256:b8199e1325db4c24fa0781c4b79cfb6d6a61e30697ef2d3f95968a73f1b3d598

Observation 9376f618-2abc-453d-ae52-23869843af28 · outbound

This paper cites Lamm: Language-assisted multi- modal instruction-tuning dataset, framework, and bench- mark.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Lamm: Language-assisted multi- modal instruction-tuning dataset, framework, and bench- mark

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.305265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.855634Z digest=sha256:8fd5bc95ee1ac63b67a757120424d31753cdda7654550a084f410066a2e2f40b

Observation e22b5e0b-38f6-45c2-afc9-6075f861c1f8 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.291495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.859312Z digest=sha256:cb6a05100454c1bddbb0357a0e72b3b73ab122ff5e1c0be093717feb9b0863aa

Observation d5484383-abc7-42e2-a928-f59e8bdeebca · outbound

This paper cites Scaling vision transformers.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Scaling vision transformers

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.279108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.862680Z digest=sha256:27702a6e6d10706b0e7bd9f9022ff89cd2f1293ce90145701f8feab09d40d8ab

Observation 7a8448cb-a3e6-4239-aeaa-ca7bb0f2b90a · outbound

This paper cites Sigmoid loss for language image pre-training.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Sigmoid loss for language image pre-training

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.266554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.866304Z digest=sha256:f3f439c3745653047072b8f883774695d44c25c54a05f9885b276a333d2b8b9d

Observation 31162702-5e58-482f-b745-20a7f7a518b7 · outbound

This paper cites Why are visually-grounded language models bad at image classi- fication? In NeurIPS, 2024.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Why are visually-grounded language models bad at image classi- fication? In NeurIPS, 2024

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.254538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.869890Z digest=sha256:4d8b20d77506f9d8ed131d35c32d48afcd81f77d2ad6934ec8efe404f04c9569

Observation 546b4cde-2671-4a32-9206-c001ced541ec · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.241459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.872988Z digest=sha256:efc80e266c7a824cd8d3caf9c41061f8642bf7cae8608bd6199ac3c0994d84b2

Observation 31a23f9b-ec1f-4a5c-bec6-052b72921c54 · outbound

This paper cites an unresolved cited work.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:41:38.214146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.881848Z digest=sha256:f954f5eb7da0cedc46f889afe9a65e6fbf99b24c2fffeeee8a826e5ef9d9d92f

Observation 7636536f-80a8-4b01-8360-a7b6d0483495 · outbound

This paper cites an unresolved cited work.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:41:38.201467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.887053Z digest=sha256:b7cd54d0f5aad164271385428782390bf1a389017b89ff13b8ae69c0dec89b1f

Observation 79d88115-2d8d-4c8e-8b7a-2f93aea6b1e4 · outbound

This paper cites an unresolved cited work.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:41:38.188362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.891584Z digest=sha256:f81f35cf3c8f4bbfc049738984c3b8a0b661f7c99b728f35d8bbcc17b0557aed

Observation 7cc51c5e-d78d-4889-a52a-4d0b5105ccad · outbound

This paper cites an unresolved cited work.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:41:38.175138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.895924Z digest=sha256:a35fd2a26e9634d0f70bd5be4f5a68bbb240f7b330353dc4c016d3ac90a8a797

Observation 7b16d058-1a34-4f6c-a60f-d65e9a096e6b · outbound

This paper cites an unresolved cited work.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:41:38.160953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.899858Z digest=sha256:2ab9bed2edaf9a6087fc3c9cf4cfc51c0e92fbbfed4313b34cc23acfaf2d50d5

Observation 4d9a79a2-fb2e-492d-af42-4d324ab77001 · outbound

This paper cites an unresolved cited work.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:41:38.148659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.904364Z digest=sha256:685361c3ac86bd596549d1fe4ac13e8786cdf7f061e02c8f6f6ed9d549f649a7

Observation b383dafe-c24e-4f2e-8852-d03ff902e144 · outbound

This paper cites an unresolved cited work.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:41:38.135247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.908611Z digest=sha256:bc29bc701e83b8cff25e7ea0f0bb5a6749528828b4b8e3e0f2853925382902b5

Observation 7de3ba91-6a0f-4777-b72a-95f5049e0ed0 · outbound

This paper cites an unresolved cited work.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:41:38.120651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.913674Z digest=sha256:eebfc763f3b96944b4148aed9e50f50c5493161388d1654dca32e0fe5d863abc

Observation fc92c4ea-6741-4ebb-9dc6-2d69d2c97d3c · outbound

This paper cites Then, we perform an in-depth exploration of MLLM classification evaluation (delineated in Section B), including the formulation, influence of option numbers, etc.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Then, we perform an in-depth exploration of MLLM classification evaluation (delineated in Section B), including the formulation, influence of option numbers, etc

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:38.227596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:41:37.877198Z digest=sha256:3a155378d0596d9ef1dc372d368a02ddba8f2ce4f7d1be84d397c799ea721d15

Pith citing papers

Observation c79c90a0-9f56-44e8-9ae0-b82a1af23dd6 · inbound

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs cites this paper.

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:53:26.314287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T05:53:26.066674Z digest=sha256:62a507643e88eadfed0cc3b5e554db5bf8d057f6900529ad26475f7419ddd5f8

Observation cd99320d-16e6-4664-8f5c-b047690979af · inbound

Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning cites this paper.

Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:17:26.785959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T06:13:33.315525Z digest=sha256:a801e58d4ef9285c15975e4dc74cb05af1920c5f0038028e23aa91a795e99882

Observation 6ed6e79d-0ae8-403b-81a3-9aeafa49eb0f · inbound

Specificity-aware reinforcement learning for fine-grained open-world classification cites this paper.

Specificity-aware reinforcement learning for fine-grained open-world classification Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:46:17.940206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T16:44:47.866126Z digest=sha256:baba4732189a7515e3d07e56b496b8d95c4c33c41c425a771e9d28c00a0f35df

Observation a966abb4-9133-462e-b8d5-fa9c4869ed82 · inbound

Bridging Coarse and Fine Recognition: A Hybrid Approach for Open-Ended Multi-Granularity Object Recognition in Interactive Educational Games cites this paper.

Bridging Coarse and Fine Recognition: A Hybrid Approach for Open-Ended Multi-Granularity Object Recognition in Interactive Educational Games Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:47:12.229585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:46:39.249226Z digest=sha256:1943f14563abcaf82eba021e9344d799f927c3959cb7d7cfccce08a9658b430e