Pith. sign in

Paper Citation Record · LEDGER

Robust image classification with multi-modal large language models

As of 12 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2412.10353.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10353 v2

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:00:13.103101Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7c8f0814-7449-4a9f-af37-ed9a0d5a9160 · outbound

This paper cites Deep neural rejection against adver- sarial examples,.

Robust image classification with multi-modal large language models Deep neural rejection against adver- sarial examples,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:00:13.594571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:00:13.013540Z digest=sha256:1c0f6071f71baa9334d5079f737bdda42f3b3c66705e00b84022b7cc01ec48bb

Observation 84bb4307-1660-44c8-8125-eb2f712e35f7 · outbound

This paper cites Image-text Retrieval: A Survey on Recent Research and Development.

Robust image classification with multi-modal large language models Image-text Retrieval: A Survey on Recent Research and Development

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T16:00:13.040800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:00:13.040800Z digest=sha256:5914ec6e3f9e408eda182826efc301d9267a87c1780696d1af5333ac6f7fd24f

Observation 60c69d0f-88ce-4201-a803-ab13abd4585b · outbound

This paper cites VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks.

Robust image classification with multi-modal large language models VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T16:00:13.052533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:00:13.052533Z digest=sha256:3a8caf3e8dd4de1febbf8a7449b5bbfc4eb4b9cb55d579f64324b811e9770020

Observation f0282e34-d82e-4341-a3d1-f10c78907782 · outbound

This paper cites On Evaluating Adversarial Robustness.

Robust image classification with multi-modal large language models On Evaluating Adversarial Robustness

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T16:00:13.059287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:00:13.059287Z digest=sha256:f4282d732883ca402fc4a922a9e51226475abdb6b2e9f311ce2e441d43fbe904

Observation e3767c66-8306-4766-8642-51167b2a8c16 · outbound

This paper cites An analysis of single-layer networks in unsupervised feature learning,.

Robust image classification with multi-modal large language models An analysis of single-layer networks in unsupervised feature learning,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:00:13.515652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:00:13.070621Z digest=sha256:d6058684669d9bd5549ff92608bd5fe029436bdf88f10b4deb532d5bc991a5dc

Observation 86033ae1-2dec-4428-966e-2f14dc0b4370 · outbound

This paper cites Improving robustness using generated data,.

Robust image classification with multi-modal large language models Improving robustness using generated data,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:00:13.487170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:00:13.075645Z digest=sha256:da144aebc8c1678408736db29fac6e52e2c7ec59023d167fb7f9c7eda7c0e505

Observation a80416b1-e043-4b94-a4d2-fdf6bd256249 · outbound

This paper cites Adversarial robustness: From self-supervised pre-training to fine-tuning,.

Robust image classification with multi-modal large language models Adversarial robustness: From self-supervised pre-training to fine-tuning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:00:13.465157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:00:13.082583Z digest=sha256:7324add08ba556dbddb61f2444cf432265d2216e9bb3945529c5897c20a14742

Observation dc5aa641-55d7-47b6-9234-ac34f1b736b7 · outbound

This paper cites Fusionbench: A comprehensive benchmark of deep model fusion,.

Robust image classification with multi-modal large language models Fusionbench: A comprehensive benchmark of deep model fusion,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T16:00:13.088859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:00:13.088859Z digest=sha256:c2b9c0938d869cf5fa9dc9be39505602d67e61bc87991e25301bd1bd3ce5de1b

Observation 2a39b60f-6d87-4c68-baa6-66f024898cab · outbound

This paper cites CLIPA-v2: Scaling CLIP Training with 81.1% Zero-shot ImageNet Accuracy within a \$10,000 Budget; An Extra \$4,000 Unlocks 81.8% Accuracy.

Robust image classification with multi-modal large language models CLIPA-v2: Scaling CLIP Training with 81.1% Zero-shot ImageNet Accuracy within a \$10,000 Budget; An Extra \$4,000 Unlocks 81.8% Accuracy

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T16:00:13.094737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:00:13.094737Z digest=sha256:99016ebc83b64d764cf9115100ff2d7b12e92cb745af847276d98660cf8edc00

Observation dac85a6e-423f-41a1-9f33-3bb62ac9547f · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

Robust image classification with multi-modal large language models HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T16:00:13.103101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:00:13.103101Z digest=sha256:96bb6aa7d3727aaa080d30221d2cb927da4ff840c296de5e11999ed90bba2ee6

Observation 28b0a0ed-11a8-435c-964a-785d74169182 · outbound

This paper cites Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories,.

Robust image classification with multi-modal large language models Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories,

Reference 2012

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:00:13.544798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:00:13.065711Z digest=sha256:336b129597c113818e4eec82242837fd145556d7ec3a8ec14d268e3979448b80

Observation 460ee490-2aec-4398-95b3-c508bebe017f · outbound

This paper cites Towards evaluating the robustness of neural networks,.

Robust image classification with multi-modal large language models Towards evaluating the robustness of neural networks,

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:00:13.675015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:00:12.984402Z digest=sha256:fa5581745f68ecb5b64df7785eabc706f66c4bd49729c3eb6cce016825c49276

Observation 775bfcc7-9d62-4945-ab92-753b48169323 · outbound

This paper cites Towards Deep Learning Models Resistant to Adversarial Attacks.

Robust image classification with multi-modal large language models Towards Deep Learning Models Resistant to Adversarial Attacks

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-11T16:00:13.002316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:00:13.002316Z digest=sha256:8cc5397a60996f17354b875c6e74296804e55af37fc1b5cdd38f1b32d718a10c

Observation aacb442f-a626-47a5-8a97-f93dd6079843 · outbound

This paper cites Safetynet: Detecting and rejecting adversarial examples robustly,.

Robust image classification with multi-modal large language models Safetynet: Detecting and rejecting adversarial examples robustly,

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:00:13.613574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:00:13.008087Z digest=sha256:9196ff452599ed52857919d709f94a0684e7dbf2a56e9c11a186589515422c04

Observation 970b03d7-ed8a-40c4-9025-18827242f426 · outbound

This paper cites A Survey on Multimodal Large Language Models.

Robust image classification with multi-modal large language models A Survey on Multimodal Large Language Models

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-11T16:00:13.029837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:00:13.029837Z digest=sha256:41aa2ebdb72a6b9ad4d57f00552fb57efacae33bd14b317ecb393f283e3fe56b

Observation 8f44f540-b9c8-4e4d-968c-17fa362054f7 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Robust image classification with multi-modal large language models Learning transferable visual models from natural language supervision,

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:00:13.566416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:00:13.019822Z digest=sha256:c45b4d2e22694093d712d58c8d9aef42b31306b188741a2711da6fe8f72a0a0c

Observation 9ae533e3-f133-4521-a039-162509643c0a · outbound

This paper cites Evasion attacks against 8 machine learning at test time,.

Robust image classification with multi-modal large language models Evasion attacks against 8 machine learning at test time,

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:00:13.845651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:00:12.977010Z digest=sha256:d861b4ceba385a7726c20e8b0754cf69bb609332bb9c7d64400ed872bab223f5

Observation 6e033151-2dd2-4d03-ae19-fb02cfd3de9e · outbound

This paper cites Wild patterns: Ten years after the rise of adversarial machine learning,.

Robust image classification with multi-modal large language models Wild patterns: Ten years after the rise of adversarial machine learning,

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:00:13.633314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:00:12.997383Z digest=sha256:0aaf6f878cec96c741dc35cfaf343cfbd45bed52517197a661f822a123d18377

Observation 4f7405e4-7fd0-4d69-acee-cdab0767910a · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

Robust image classification with multi-modal large language models VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T16:00:13.046276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:00:13.046276Z digest=sha256:5b0448a27b8a992adeb740be771a4c15465a290cda4955e1c39d9c1f2d30a365

Observation 1edc973a-b1b2-42cc-89e4-663c6096cc65 · outbound

This paper cites A Survey of Vision-Language Pre-Trained Models.

Robust image classification with multi-modal large language models A Survey of Vision-Language Pre-Trained Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T16:00:13.035351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:00:13.035351Z digest=sha256:5f023c821193ef4189773f2ed964afdbdf867ff4fefa52bbb06c279aa850e609

Observation 93b7f950-d061-4b8f-a711-164aab124dc5 · outbound

This paper cites Deep k-Nearest Neighbors: Towards Confident, Interpretable and Robust Deep Learning.

Robust image classification with multi-modal large language models Deep k-Nearest Neighbors: Towards Confident, Interpretable and Robust Deep Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T16:00:13.024471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:00:13.024471Z digest=sha256:6378ba73becccb2d46c991fa595387df2159c5e05444ede8b7e2a7fab3f29ef0

Observation 10eed4f2-559b-4a97-a16a-7e7067022bdc · outbound

This paper cites Attackbench: Evaluating gradient-based attacks for adversarial examples,.

Robust image classification with multi-modal large language models Attackbench: Evaluating gradient-based attacks for adversarial examples,

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:00:13.654844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:00:12.992098Z digest=sha256:eda9f0d150783fd29085e742e015ce3a1ae159bff7d88378a013c5f04533a62b

Pith citing papers

No inbound Pith citation observations are available.