Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T10:41:37.913674Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 4 inbound Pith citation observations for arXiv:2412.16418.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T10:41:37.913674Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-17T05:53:26.066674Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-17T05:53:26.308450Z
65 of 65 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 26d9fdff-b43a-4026-ad45-64adff441f29 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Phi-3 technical report: A highly capable language model locally on your phone, 2024
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ee6d380d-4759-4055-88ad-3cf8158c5df3 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd44ae15-887c-4550-ba9d-b1f3d11e1f76 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b147c20f-2529-43f9-a83b-fe3e3809a35d · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models.NeurIPS, 32, 2019
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5f25b32c-41a3-4ab6-8391-260bd1192543 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Food-101–mining discriminative components with random forests
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f25dc983-710a-40ce-b637-d287d267c968 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Are we on the right way for evaluating large vision-language models? In NeurIPS, 2024
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f2bc42e5-e1e0-44fa-b969-89a1d01cacd5 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 512549b8-9fb6-46fb-9c55-0f9f7350d085 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e37559a-9fd4-46f4-93f4-b1a016a384ff · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Reproducible scal- ing laws for contrastive language-image learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 649b899d-acfc-4e74-ac00-ec9a4666e795 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Coatnet: Marrying convolution and attention for all data sizes
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1fae459d-ba51-44ce-bd4f-142eac21a8d3 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Imagenet: A large-scale hierarchical image database
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2f3049f3-2cba-46e3-b277-3daad7465ce5 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities BERT: pre-training of deep bidirectional trans- formers for language understanding
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5f8f7f39-9aa5-4eda-9b92-0d01a523c50f · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities An image is worth 16x16 words: Transformers for image recognition at scale
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6565e1be-5820-43a8-965c-4c62d5f88e7a · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Data filtering networks
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f5617f5b-cba8-40ea-bf28-bd89e01680a8 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Eva-02: A visual representa- tion for neon genesis
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 64e592e5-588d-48fc-86b7-210724a4ffa9 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Learning gener- ative visual models from few training examples: An incre- mental bayesian approach tested on 101 object categories
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 135ad5c1-ab1f-4405-9072-85df1f456097 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb13bc77-0bc9-42b1-a4ae-ccf7ddd93ff3 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Detect, Describe, Discriminate: Moving Beyond VQA for MLLM Evaluation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a3306c69-e067-41c6-ab6a-20f523e98067 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Deep residual learning for image recognition
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a58e4a6-c8f9-4894-b8f4-02cd67bcf147 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d705921b-d5a6-452f-942b-105c522a0042 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Rwku: Benchmarking real-world knowledge unlearning for large language models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6bdb4fae-00c0-4d45-9dce-5c5d20c3a406 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities A diagram is worth a dozen images
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 147fb193-4c30-42cd-b226-900d5f05ed03 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities 3d object representations for fine-grained categorization
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 423d313b-a984-48b6-8315-49018b59265f · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Learning multiple layers of features from tiny images
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2880705b-cf22-48c6-955e-78b6ed7c7f28 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Imagenet classification with deep convolutional neural net- works
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb3ff337-0db6-4da1-8332-27cf3950accf · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Gradient-based learning applied to document recog- nition
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4763ab78-4416-492a-baa7-2c9d8e3126f4 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Seed-bench: Bench- marking multimodal large language models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d7f22d14-e1ba-485c-a2d6-23804eac654f · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities LLaVA-OneVision: Easy Visual Task Transfer
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb5755bc-79ee-4029-99a4-ccd84213866c · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1eefded5-450b-469d-8753-2b4f87f91387 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Visual instruction tuning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f05a396c-3dcf-43b6-8376-e2cc2e03942c · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Mmbench: Is your multi-modal model an all-around player? In ECCV, pages 216–233
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b125040d-82e0-47d1-8623-32f40ce15112 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Swin transformer: Hierarchical vision transformer using shifted windows
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a9eb696-d3e2-4342-b53f-63f7a724c2b7 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Swin transformer v2: Scaling up capacity and resolution
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 680c1ba2-1e15-40d8-b18e-392a34f8462c · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Deepseek-vl: Towards real-world vision- language understanding, 2024
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6c907534-b710-4e45-a52d-deb2507b3eb7 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce63c050-1879-4730-b950-31b8bdf8d301 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities ChartQA: A benchmark for question answering about charts with visual and logical reasoning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fc28fc7b-fd75-4e1b-9d14-938ab88e8b51 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Docvqa: A dataset for vqa on document images
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 97fc78a3-17c7-4c73-a309-f0126aae063f · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Automated flower classification over a large number of classes
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d84a2c1d-a71c-48d1-bc91-ec675586df64 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Learn- ing transferable visual models from natural language super- vision
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b0dda3b1-c61c-44bd-b8c3-145b132615cd · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Do imagenet classifiers generalize to im- agenet? In ICML, pages 5389–5400
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6583b487-47db-4cab-b85e-780423ad6ef8 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Very Deep Convolutional Networks for Large-Scale Image Recognition
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5aa90d2-8c50-4c8a-88ae-cf42f07c6516 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Going deeper with convolutions
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c39ea7b9-18ea-49d6-916b-5f101a0d3628 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Jn-logo: A logo database for aesthetic visual analysis
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 77b04fab-f437-4f57-ab40-25f814b7bbf1 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a11be65f-a96a-478e-851d-dc087973079e · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities The caltech-ucsd birds-200-2011 dataset
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 79c68e05-d7ef-4e85-a426-dde92e24ba70 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9a3af94-9db8-4a0f-b0e8-afcb4f4ae00b · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Google landmarks dataset v2-a large-scale benchmark for instance-level recognition and retrieval
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6dfdb39c-d2fd-4d8b-a273-3684ffda1862 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54f2a3cb-6db3-47cf-9fa2-0eafe9ae4a9b · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Qwen2 Technical Report
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ea62737-c138-4423-bb54-245ca0e19438 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9376f618-2abc-453d-ae52-23869843af28 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Lamm: Language-assisted multi- modal instruction-tuning dataset, framework, and bench- mark
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e22b5e0b-38f6-45c2-afc9-6075f861c1f8 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d5484383-abc7-42e2-a928-f59e8bdeebca · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Scaling vision transformers
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7a8448cb-a3e6-4239-aeaa-ca7bb0f2b90a · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Sigmoid loss for language image pre-training
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 31162702-5e58-482f-b745-20a7f7a518b7 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Why are visually-grounded language models bad at image classi- fication? In NeurIPS, 2024
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 546b4cde-2671-4a32-9206-c001ced541ec · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Judging llm-as-a-judge with mt-bench and chatbot arena
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 31a23f9b-ec1f-4a5c-bec6-052b72921c54 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Unresolved cited work
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7636536f-80a8-4b01-8360-a7b6d0483495 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Unresolved cited work
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 79d88115-2d8d-4c8e-8b7a-2f93aea6b1e4 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Unresolved cited work
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7cc51c5e-d78d-4889-a52a-4d0b5105ccad · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Unresolved cited work
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7b16d058-1a34-4f6c-a60f-d65e9a096e6b · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Unresolved cited work
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4d9a79a2-fb2e-492d-af42-4d324ab77001 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Unresolved cited work
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b383dafe-c24e-4f2e-8852-d03ff902e144 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Unresolved cited work
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7de3ba91-6a0f-4777-b72a-95f5049e0ed0 · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Unresolved cited work
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fc92c4ea-6741-4ebb-9dc6-2d69d2c97d3c · outbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities Then, we perform an in-depth exploration of MLLM classification evaluation (delineated in Section B), including the formulation, influence of option numbers, etc
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c79c90a0-9f56-44e8-9ae0-b82a1af23dd6 · inbound
WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cd99320d-16e6-4664-8f5c-b047690979af · inbound
Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6ed6e79d-0ae8-403b-81a3-9aeafa49eb0f · inbound
Specificity-aware reinforcement learning for fine-grained open-world classification Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a966abb4-9133-462e-b8d5-fa9c4869ed82 · inbound
Bridging Coarse and Fine Recognition: A Hybrid Approach for Open-Ended Multi-Granularity Object Recognition in Interactive Educational Games Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.