Pith. sign in

Paper Citation Record · LEDGER

Robust image classification with multi-modal large language models

As of 11 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2412.10353.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10353 v2

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:00:13.103101Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7c8f0814-7449-4a9f-af37-ed9a0d5a9160 · outbound

This paper cites Deep neural rejection against adver- sarial examples,.

Robust image classification with multi-modal large language models Deep neural rejection against adver- sarial examples,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:00:13.594571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:00:13.013540Z digest=sha256:45b43515a36246a2964025d9b50522ac103dfbb4e913240c552b39d3fca18951

Observation 84bb4307-1660-44c8-8125-eb2f712e35f7 · outbound

This paper cites Image-text Retrieval: A Survey on Recent Research and Development.

Robust image classification with multi-modal large language models Image-text Retrieval: A Survey on Recent Research and Development

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T16:00:13.040800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:00:13.040800Z digest=sha256:d6da3232f2a7a6b558199d843972596ee360a0a1e70a070b0eff5d94c157287c

Observation 60c69d0f-88ce-4201-a803-ab13abd4585b · outbound

This paper cites VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks.

Robust image classification with multi-modal large language models VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T16:00:13.052533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:00:13.052533Z digest=sha256:61769bf433bdab85b7e8046ceaafa304de0aea5ea184cd0417a366f242d0ea22

Observation f0282e34-d82e-4341-a3d1-f10c78907782 · outbound

This paper cites On Evaluating Adversarial Robustness.

Robust image classification with multi-modal large language models On Evaluating Adversarial Robustness

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T16:00:13.059287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:00:13.059287Z digest=sha256:26f3310d2d7c02a4ce64fb607df337f9c9b2e35854ea3a6b59eb998cc64b53fc

Observation e3767c66-8306-4766-8642-51167b2a8c16 · outbound

This paper cites An analysis of single-layer networks in unsupervised feature learning,.

Robust image classification with multi-modal large language models An analysis of single-layer networks in unsupervised feature learning,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:00:13.515652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:00:13.070621Z digest=sha256:45e34d9f6d9785e112671a226c8a63238858f81c71a12b4f8a00b70afb261f71

Observation 86033ae1-2dec-4428-966e-2f14dc0b4370 · outbound

This paper cites Improving robustness using generated data,.

Robust image classification with multi-modal large language models Improving robustness using generated data,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:00:13.487170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:00:13.075645Z digest=sha256:e969b4f8122ec641682d6ff5d86b2c6569ce3edb935f4d156946ec718aeac949

Observation a80416b1-e043-4b94-a4d2-fdf6bd256249 · outbound

This paper cites Adversarial robustness: From self-supervised pre-training to fine-tuning,.

Robust image classification with multi-modal large language models Adversarial robustness: From self-supervised pre-training to fine-tuning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:00:13.465157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:00:13.082583Z digest=sha256:920600845056589376e1c36d9bafd9968abefd1fe9e4cb76f69ec7fda4cf652f

Observation dc5aa641-55d7-47b6-9234-ac34f1b736b7 · outbound

This paper cites Fusionbench: A comprehensive benchmark of deep model fusion,.

Robust image classification with multi-modal large language models Fusionbench: A comprehensive benchmark of deep model fusion,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T16:00:13.088859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:00:13.088859Z digest=sha256:511278e257a0ca84e1050c2031b995cf8805afd330bc6df96306b6c10d10207b

Observation 2a39b60f-6d87-4c68-baa6-66f024898cab · outbound

This paper cites CLIPA-v2: Scaling CLIP Training with 81.1% Zero-shot ImageNet Accuracy within a \$10,000 Budget; An Extra \$4,000 Unlocks 81.8% Accuracy.

Robust image classification with multi-modal large language models CLIPA-v2: Scaling CLIP Training with 81.1% Zero-shot ImageNet Accuracy within a \$10,000 Budget; An Extra \$4,000 Unlocks 81.8% Accuracy

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T16:00:13.094737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:00:13.094737Z digest=sha256:04a130525fc22cca82eb6ad65a8f1ca87e43351e6311abeb3c34b7355be20b4b

Observation dac85a6e-423f-41a1-9f33-3bb62ac9547f · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

Robust image classification with multi-modal large language models HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T16:00:13.103101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:00:13.103101Z digest=sha256:6e7980abb9fb40ea0ba5629c82b163ae35109c2938db4ded3707eba1a3eabb65

Observation 28b0a0ed-11a8-435c-964a-785d74169182 · outbound

This paper cites Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories,.

Robust image classification with multi-modal large language models Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories,

Reference 2012

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:00:13.544798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:00:13.065711Z digest=sha256:86c74efd6ef47672bd15a95bd5aa4ed5cc6fa57760a1e741cb4870054b9df858

Observation 460ee490-2aec-4398-95b3-c508bebe017f · outbound

This paper cites Towards evaluating the robustness of neural networks,.

Robust image classification with multi-modal large language models Towards evaluating the robustness of neural networks,

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:00:13.675015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:00:12.984402Z digest=sha256:c78d603530ba047f9400c0600aa3d264a2c787d2cc948ee590b4cb1f76035004

Observation 775bfcc7-9d62-4945-ab92-753b48169323 · outbound

This paper cites Towards Deep Learning Models Resistant to Adversarial Attacks.

Robust image classification with multi-modal large language models Towards Deep Learning Models Resistant to Adversarial Attacks

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-11T16:00:13.002316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:00:13.002316Z digest=sha256:15c83d336cd9af4ccb435e228da9d067d03b26b6b50a2e23283af545c376fa45

Observation aacb442f-a626-47a5-8a97-f93dd6079843 · outbound

This paper cites Safetynet: Detecting and rejecting adversarial examples robustly,.

Robust image classification with multi-modal large language models Safetynet: Detecting and rejecting adversarial examples robustly,

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:00:13.613574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:00:13.008087Z digest=sha256:b60df42b2b4dfd414ab91dd3e4f1822dc2b881fdbf08e20f8d9db186d26024a7

Observation 970b03d7-ed8a-40c4-9025-18827242f426 · outbound

This paper cites A Survey on Multimodal Large Language Models.

Robust image classification with multi-modal large language models A Survey on Multimodal Large Language Models

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-11T16:00:13.029837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:00:13.029837Z digest=sha256:ca9657e59f54c8dac9418cbb247338788ae0daf6203322a59211bc507dcee905

Observation 8f44f540-b9c8-4e4d-968c-17fa362054f7 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Robust image classification with multi-modal large language models Learning transferable visual models from natural language supervision,

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:00:13.566416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:00:13.019822Z digest=sha256:544ca4970b8e7e219fe2322f36e5b13d284ce6ffcb85b8fe9de26273e47609b0

Observation 9ae533e3-f133-4521-a039-162509643c0a · outbound

This paper cites Evasion attacks against 8 machine learning at test time,.

Robust image classification with multi-modal large language models Evasion attacks against 8 machine learning at test time,

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:00:13.845651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:00:12.977010Z digest=sha256:f1eceb6bf964ae2fca550a6ba60cc521ed4764283b6fdfa71b4a61ed72c7f566

Observation 6e033151-2dd2-4d03-ae19-fb02cfd3de9e · outbound

This paper cites Wild patterns: Ten years after the rise of adversarial machine learning,.

Robust image classification with multi-modal large language models Wild patterns: Ten years after the rise of adversarial machine learning,

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:00:13.633314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:00:12.997383Z digest=sha256:f0f953b171e2570ae1e98d7e557242b53d604ac399968bfd0996f8cd79e769a9

Observation 4f7405e4-7fd0-4d69-acee-cdab0767910a · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

Robust image classification with multi-modal large language models VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T16:00:13.046276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:00:13.046276Z digest=sha256:cecc8a01a467332fa8d371529195fbb09ff86aae60a864df7cfc79935dafe7ff

Observation 1edc973a-b1b2-42cc-89e4-663c6096cc65 · outbound

This paper cites A Survey of Vision-Language Pre-Trained Models.

Robust image classification with multi-modal large language models A Survey of Vision-Language Pre-Trained Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T16:00:13.035351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:00:13.035351Z digest=sha256:2e739a1416ec161f7948f01d3fd076a760fd159ad4f6947ebb0398cfa511b12b

Observation 93b7f950-d061-4b8f-a711-164aab124dc5 · outbound

This paper cites Deep k-Nearest Neighbors: Towards Confident, Interpretable and Robust Deep Learning.

Robust image classification with multi-modal large language models Deep k-Nearest Neighbors: Towards Confident, Interpretable and Robust Deep Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T16:00:13.024471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:00:13.024471Z digest=sha256:0b0b87e59d254e4157975fe3cd0381190f0dc64ca511ef3d6fe78b43a09de8ef

Observation 10eed4f2-559b-4a97-a16a-7e7067022bdc · outbound

This paper cites Attackbench: Evaluating gradient-based attacks for adversarial examples,.

Robust image classification with multi-modal large language models Attackbench: Evaluating gradient-based attacks for adversarial examples,

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:00:13.654844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:00:12.992098Z digest=sha256:9f8ea8c717056d99367c8a4271d7a2153e4f752d9d7ba1cab1a1af38ef1dd5af

Pith citing papers

No inbound Pith citation observations are available.