Pith. sign in

Paper Citation Record · LEDGER

Language Model as Visual Explainer

As of 16 August 2026, this Paper Citation Record lists 100 of 104 outbound references and 0 inbound Pith citation observations for arXiv:2412.07802.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.07802 v1

Coverage vector

measured 100 of 104 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:07:35.352999Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 104 outbound references displayed

  • verified exact1
  • verified fuzzy34
  • unresolved64
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d287122b-39f3-4d6e-a10b-b317f101a046 · outbound

This paper cites Quantifying attention flow in transformers.

Language Model as Visual Explainer Quantifying attention flow in transformers

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:34.925869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:34.925869Z digest=sha256:3845799bffa89268a14c4d5280c72dc1780589b521d0c6d017ea41fd7979d5f1

Observation 39d05a9e-8ba5-4f39-b5cf-b779a25fefd2 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Language Model as Visual Explainer Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:34.931225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:34.931225Z digest=sha256:901db682e83436f96a41edbb2977a80b24a512633aaf7fde4b7731f855af6ff3

Observation 95726ec4-976d-423a-9fb5-70ad5a1eba6d · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Language Model as Visual Explainer Flamingo: a visual language model for few-shot learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:34.936341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:34.936341Z digest=sha256:9c89e702cb0a48ed3fdb3a9588b3cf5b948ad419a63994f97e0a0099505c5387

Observation 3cfbc6dc-c9dd-4010-8d77-2815001a1e52 · outbound

This paper cites Learning to Compose Neural Networks for Question Answering.

Language Model as Visual Explainer Learning to Compose Neural Networks for Question Answering

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:34.940858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:34.940858Z digest=sha256:b3cc833d6a116c1ddebfa0f3f1fb310fa2b6a212ed5ecd9f32dce2686cd23b04

Observation 8799e7bb-ec1d-4fee-b375-ed5f7e4fdd54 · outbound

This paper cites Neural module networks.

Language Model as Visual Explainer Neural module networks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:34.945558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:34.945558Z digest=sha256:72e03fedd2af5c611fda81de537659722e6229934935c48021ffc0884315318b

Observation 8b3f5dae-c2c0-492c-b403-16a6aae8ca91 · outbound

This paper cites A survey on tree edit distance and related problems.

Language Model as Visual Explainer A survey on tree edit distance and related problems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:34.950222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:34.950222Z digest=sha256:9222886301b382bcd00a987ffe8c1c09e7bc0eedb71a60f2447b0d91ab21b1d9

Observation b455b4df-91ec-465a-b011-28148c6c7fc6 · outbound

This paper cites e-snli: Natural language inference with natural language explanations.

Language Model as Visual Explainer e-snli: Natural language inference with natural language explanations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:34.954578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:34.954578Z digest=sha256:6abd35b90523b5c72f1af8b881dc6b2a8ff0309e0cb2413f281bc00db061091f

Observation 96fda153-dd83-4efd-9e5d-1afc153ea96a · outbound

This paper cites Unsuper- vised learning of visual features by contrasting cluster assignments.

Language Model as Visual Explainer Unsuper- vised learning of visual features by contrasting cluster assignments

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:34.958764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:34.958764Z digest=sha256:4995813b14b3a9aee059dfd306d0b23fa251e183956968028ab571476dec80ef

Observation 7911599c-7388-42df-82a7-220beac16e7d · outbound

This paper cites Emerging properties in self-supervised vision transformers.

Language Model as Visual Explainer Emerging properties in self-supervised vision transformers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:34.962716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:34.962716Z digest=sha256:cbbbf1e4e47c942e10dac0fe70b7131c4c1944a1773b4d21e5d8548ee9b68b6a

Observation 5eb792cf-97ab-4db0-8586-2e91e8859ada · outbound

This paper cites This looks like that: deep learning for interpretable image recognition.

Language Model as Visual Explainer This looks like that: deep learning for interpretable image recognition

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:34.966775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:34.966775Z digest=sha256:fb80d270594e6d6b8c4ced297148b1d2d7448303921fe163418c1e1e9a460d8e

Observation 59f2b850-1e45-405a-97e8-4fb377fd551a · outbound

This paper cites A simple framework for contrastive learning of visual representations.

Language Model as Visual Explainer A simple framework for contrastive learning of visual representations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:34.971195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:34.971195Z digest=sha256:d73c7ea79dddf57af6f739f875fab3075b010e26226802e1515eb4a78fb50d5e

Observation 55f551ae-40ae-4e7a-815a-05f229841883 · outbound

This paper cites Improved Baselines with Momentum Contrastive Learning.

Language Model as Visual Explainer Improved Baselines with Momentum Contrastive Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:34.975450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:34.975450Z digest=sha256:a7196181401cc3991e77a197bf56022bb6dfbfea7ba2d0c29574e014acc2df18

Observation 15fc1976-a297-4000-94ef-8463dbc12913 · outbound

This paper cites An Empirical Study of Training Self-Supervised Vision Transformers.

Language Model as Visual Explainer An Empirical Study of Training Self-Supervised Vision Transformers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:34.980032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:34.980032Z digest=sha256:94f1b5421c7d8f71ab745e9882c5aad1e7c55d2ab87ba60e61f004dc3ab9434d

Observation b6b286e1-989a-46ef-915f-63137bbc7115 · outbound

This paper cites Self-born wiring for neural trees.

Language Model as Visual Explainer Self-born wiring for neural trees

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:34.986040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:34.986040Z digest=sha256:41e524c41a3dfccfde297f4d126cff9a9c8459b623c8c5dc79b593490a8deacc

Observation a701826f-1b2c-4863-ab54-7c1ac9c81648 · outbound

This paper cites Whatever next? predictive brains, situated agents, and the future of cognitive science.

Language Model as Visual Explainer Whatever next? predictive brains, situated agents, and the future of cognitive science

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:34.990229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:34.990229Z digest=sha256:f4c50330927d6b744b3c4e183c9558c69ea0914b0128846a035f4fd47979acc4

Observation 94f491c4-af0e-42f5-b9c0-fab256aa3d4d · outbound

This paper cites Nearest neighbor pattern classification.

Language Model as Visual Explainer Nearest neighbor pattern classification

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:34.995362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:34.995362Z digest=sha256:7c202e6211e42b26ccaddaeacb54f803b1afdeaa80095c9dba16ca1fa5a37502

Observation d19211e6-aed5-496c-8bdc-3b39e9bff659 · outbound

This paper cites Extracting tree-structured representations of trained networks.

Language Model as Visual Explainer Extracting tree-structured representations of trained networks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.000185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.000185Z digest=sha256:692657495096a461b55a687167f742a91510fe32fa00e85a7e5504b534adbfe2

Observation 254d7648-fe50-4343-8479-08f52bbff4b8 · outbound

This paper cites Improved Regularization of Convolutional Neural Networks with Cutout.

Language Model as Visual Explainer Improved Regularization of Convolutional Neural Networks with Cutout

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.006591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.006591Z digest=sha256:76613e3ed3a9a311748f316fc02c075d7f0257c7195d57c1cb8a213c12c93e41

Observation e55fb3a2-fa2a-43a9-92c0-f58f3b33f9ce · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Language Model as Visual Explainer An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.011609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.011609Z digest=sha256:5dc3799a69e2b2c1bedfc8732de7a95355d65ed22609404815e867a160a3a268

Observation d183024a-fbb9-4b87-bc80-c8d6712adc51 · outbound

This paper cites wordnet: WordNet Interface, 2023.

Language Model as Visual Explainer wordnet: WordNet Interface, 2023

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.016296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.016296Z digest=sha256:a7e021f41610fc190f0dd7498b31671eff64c1b41f9cc1d25f46eb3d8a5ef341

Observation 012822fd-03b4-4dd3-a51a-23ead1a45097 · outbound

This paper cites Distilling a Neural Network Into a Soft Decision Tree.

Language Model as Visual Explainer Distilling a Neural Network Into a Soft Decision Tree

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.020566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.020566Z digest=sha256:7b04439735302f8a14b9ce6f2054cb51de20dcb8262471c76aab388829aa0c51

Observation b454b4d0-4f0d-4ba8-a340-5607468f3aa5 · outbound

This paper cites Interpreting CLIP’s image representation via text-based decomposition.

Language Model as Visual Explainer Interpreting CLIP’s image representation via text-based decomposition

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.025164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.025164Z digest=sha256:a19a68ecfef4f7ce24fe23b2c7ec59e8e3f0dcef2bc167d0eaaa3a1f9534b2a7

Observation 38a8b640-f69e-416d-bd00-8b8fbc7202c3 · outbound

This paper cites Wichmann, and Wieland Brendel.

Language Model as Visual Explainer Wichmann, and Wieland Brendel

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.029197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.029197Z digest=sha256:724a631c090a4d0d3e8111018127699aae15a2c1e20cb4ea67691559f9cd4553

Observation c05fffe8-217d-4bb1-ae20-a7539849dfd8 · outbound

This paper cites Bootstrap your own latent-a new approach to self-supervised learning.

Language Model as Visual Explainer Bootstrap your own latent-a new approach to self-supervised learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.033374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.033374Z digest=sha256:0501a7e37dd8f9ce0004e796ff511d9fbc41564d2b5938713c96554a76b636bb

Observation a467ee9f-2e36-41dd-b523-95bcc042a8fd · outbound

This paper cites Visual Programming: Compositional visual reasoning without training.

Language Model as Visual Explainer Visual Programming: Compositional visual reasoning without training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.037174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.037174Z digest=sha256:cdd5a9f75c0beb744c357cc3ae5dab344ac40d59970e0b2be81708d18fae040f

Observation 34d9927d-d889-4699-ab40-68a0690e5450 · outbound

This paper cites Deep residual learning for image recognition.

Language Model as Visual Explainer Deep residual learning for image recognition

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.041402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.041402Z digest=sha256:8150a082732666adf25f78879a9757996b256062c224f1bdfefae44ef781c8c1

Observation 1ec714f3-b8f4-4f4d-9a5f-2fc5a24b5949 · outbound

This paper cites Generating visual explanations.

Language Model as Visual Explainer Generating visual explanations

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.046241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.046241Z digest=sha256:3b5b99474f50892504a77ed74529946a269650ed427d541d18c47450b9a7b172

Observation d4bd6ceb-9a4a-4779-be0e-2d603b189fa2 · outbound

This paper cites Learning to reason: End-to-end module networks for visual question answering.

Language Model as Visual Explainer Learning to reason: End-to-end module networks for visual question answering

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.050287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.050287Z digest=sha256:1d748a357e6ba248aeb0d895b40449f717ec46b2c497ae0c8eced327835587db

Observation 2dc95556-449f-4d10-b2bc-77c303b2b2a1 · outbound

This paper cites Densely connected convolutional networks.

Language Model as Visual Explainer Densely connected convolutional networks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.054375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.054375Z digest=sha256:9f7c5370f855287a0b420705acb01f29ba0aae036acde75c356f8394a2b5dd2d

Observation 7ccf9fff-0da3-489c-a83d-de2b9d98a65b · outbound

This paper cites Prototype optimization for nearest-neighbor classification.

Language Model as Visual Explainer Prototype optimization for nearest-neighbor classification

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.058513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.058513Z digest=sha256:b2126a5763ea414155cd2e63e7b67fae7682e3df7c2dcd89038652b2cb9261ea

Observation d21f87c6-4756-40f4-90ea-1d51f7de122b · outbound

This paper cites Discovering states and transformations in image collections.

Language Model as Visual Explainer Discovering states and transformations in image collections

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.062327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.062327Z digest=sha256:c5ea62161aed366bdc490e440e8ee16c8093ee71ea3cb4bc936f458c9e54bcbf

Observation 3ded331e-ef42-4fac-86a6-d99907bf4846 · outbound

This paper cites Hierarchical mixtures of experts and the em algorithm.

Language Model as Visual Explainer Hierarchical mixtures of experts and the em algorithm

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.066381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.066381Z digest=sha256:4262590424ede42be929f0c2c357819bfc8fe4a778c9b1b96fd099348b2210bf

Observation 765281fe-6245-4c87-9969-b1c189fdd3bc · outbound

This paper cites On the approximability of the maximum common subgraph problem.

Language Model as Visual Explainer On the approximability of the maximum common subgraph problem

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:07:36.427322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.070899Z digest=sha256:8fa2a8f872f1cdb061a52d0c4b36256199c81e248932c74fd6d5f475b49ca9fa

Observation e1cc92bf-e41a-4799-bc28-a39dd6ba99bb · outbound

This paper cites Proto2proto: Can you recognize the car, the way i do? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10233–10243, 2022.

Language Model as Visual Explainer Proto2proto: Can you recognize the car, the way i do? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10233–10243, 2022

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:07:36.413594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.074800Z digest=sha256:6eddf0827cb18bdcc91a9558c0cf77d60742302a3337dcffc8a9756c95ddadc7

Observation a9717e46-6a52-4db7-95df-01a4b0de8f11 · outbound

This paper cites Textual explanations for self-driving vehicles.

Language Model as Visual Explainer Textual explanations for self-driving vehicles

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:07:36.389746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.079015Z digest=sha256:4aa4d38cc6af923b2206c189e65e85b19a2ee60c27f65cfcb505649fae8843ed

Observation 3990665d-1a60-4a5d-a354-77c68c343551 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Language Model as Visual Explainer Adam: A Method for Stochastic Optimization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.082734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.082734Z digest=sha256:7b6c1552e54940c495e4aaa9c6b70b391c410ab23119b5595b5420570bd0f32e

Observation 3ea3b2fc-549e-4109-a1b8-0b5f0728473c · outbound

This paper cites Learning vector quantization.

Language Model as Visual Explainer Learning vector quantization

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:07:36.376657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.087143Z digest=sha256:94fedad239457820a4b475d0126ab6db4a97fd1f910ff3d70923d68bc8b32e96

Observation 456c27d8-7660-46a8-b40b-dee51f2cf49e · outbound

This paper cites Deep neural decision forests.

Language Model as Visual Explainer Deep neural decision forests

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:07:36.363106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.091504Z digest=sha256:f4e4e86bceb4b8e183556618eb9eb4e8d4ee23640a3b6d523f73c497f608d1f8

Observation dc620ac5-fbfc-4293-9d18-ca252cb039fb · outbound

This paper cites Learning multiple layers of features from tiny images.

Language Model as Visual Explainer Learning multiple layers of features from tiny images

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.095714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.095714Z digest=sha256:f04a8f412111adaf14effc864889ad4d17c2f2ed3d0967df0980daf97002eb90

Observation 24a6a57f-607b-423b-84e4-5c516f71967d · outbound

This paper cites Deep learning for case-based reasoning through prototypes: A neural network that explains its predictions.

Language Model as Visual Explainer Deep learning for case-based reasoning through prototypes: A neural network that explains its predictions

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:07:36.343178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.099668Z digest=sha256:a41d8cdf07f9daded140ac1adea7671b245c710acc95d77bfe69b37f54954fb1

Observation 1613ff92-fa76-4d97-8d20-87ad7328ddc3 · outbound

This paper cites Vqa-e: Explaining, elaborating, and enhancing your answers for visual questions.

Language Model as Visual Explainer Vqa-e: Explaining, elaborating, and enhancing your answers for visual questions

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:07:36.330172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.103730Z digest=sha256:5dc485d4d29578c01c0789576c7395c5607d7a0c04c37fd74e268ef54826ae1b

Observation 482986b5-fc79-47aa-9a3b-f5ea1886c9ef · outbound

This paper cites A systematic investigation of commonsense knowledge in large language models.

Language Model as Visual Explainer A systematic investigation of commonsense knowledge in large language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:07:36.316551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.107630Z digest=sha256:e0c6f72c570bac1f50ef1723ea76141b74df8b9259592ae8c5e092647a972643

Observation d30e0b5d-2d1d-474a-8fa8-f127ab2ca0a4 · outbound

This paper cites TaskMatrix.AI: Completing Tasks by Connecting Foundation Models with Millions of APIs.

Language Model as Visual Explainer TaskMatrix.AI: Completing Tasks by Connecting Foundation Models with Millions of APIs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.112532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.112532Z digest=sha256:7b5270ff46251ee47cc165b53902a7bc1c52fafd2c0a0b8a464bc0788127227e

Observation 7aac7962-0107-4667-97c2-dc3c0ddbdc42 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

Language Model as Visual Explainer Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.116877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.116877Z digest=sha256:47f052403911978ab0d1e4ffae94c5392d0ff3eb79d4fbeed9efbe3b29b3e1ea

Observation 45362e95-15fb-4c89-8693-8ec203febe44 · outbound

This paper cites What Makes Good In-Context Examples for GPT-$3$?.

Language Model as Visual Explainer What Makes Good In-Context Examples for GPT-$3$?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.121142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.121142Z digest=sha256:f765760ccdda7779375aa98e5489022a08cdbcb90b29f1958f9484017a36ec7b

Observation 2f8bbb6b-aecb-4b68-82e5-60a2859a1b70 · outbound

This paper cites A unified approach to interpreting model predictions.

Language Model as Visual Explainer A unified approach to interpreting model predictions

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.125515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.125515Z digest=sha256:2dd92746da63cbc4390a63f140ecbe726939b845680cbfa0ef7af0b5b1749bee

Observation 2ebb5d5d-09f0-449e-9592-09ef74a9d582 · outbound

This paper cites Doubly Right Object Recognition: A Why Prompt for Visual Rationales.

Language Model as Visual Explainer Doubly Right Object Recognition: A Why Prompt for Visual Rationales

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.129760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.129760Z digest=sha256:1f560efcf54435fa4e9bf35e39694a9b5a938cc857f7dcbcda887be7e8ed8ec0

Observation ee22d62d-df0f-4d87-aa44-982c70adb750 · outbound

This paper cites Visual Classification via Description from Large Language Models.

Language Model as Visual Explainer Visual Classification via Description from Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.133678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.133678Z digest=sha256:0e473633039fc6578a5389ce39e1244e192700f0e3e053404460c7f5018b6f2c

Observation a0795a4b-191d-4a6d-b6cd-e956b2da3e05 · outbound

This paper cites Neural prototype trees for interpretable fine-grained im- age recognition.

Language Model as Visual Explainer Neural prototype trees for interpretable fine-grained im- age recognition

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:07:36.287197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.137834Z digest=sha256:50c126a5152fdbafa271650744c3f5f3ffcf4b864b9da7d01725c317a8d236d1

Observation 9eec07fe-0fe0-466c-9fb3-754417b9198a · outbound

This paper cites Coco attributes: Attributes for people, animals, and objects.

Language Model as Visual Explainer Coco attributes: Attributes for people, animals, and objects

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:07:36.274560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.141437Z digest=sha256:618a723f58c833ded7148f2771df79a670967ed7a321b7672886c3fb1558c3c0

Observation 1b7d59e7-9bb8-4097-b7cd-e17e2d047650 · outbound

This paper cites Xplainer: From X-Ray Observations to Explainable Zero-Shot Diagnosis.

Language Model as Visual Explainer Xplainer: From X-Ray Observations to Explainable Zero-Shot Diagnosis

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-11T20:07:35.547233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.145493Z digest=sha256:e3dab5645e1ab2ddb117c6a89ecebedfb457317c8fe59caa5e0df40b8cc73d27

Observation 5871c619-2400-4e97-b4cb-bed5fdccaf16 · outbound

This paper cites Learning to predict visual attributes in the wild.

Language Model as Visual Explainer Learning to predict visual attributes in the wild

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:07:36.261771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.149757Z digest=sha256:93b78cb1801a16aee61f5f8f309a70938c2bd814ee6a1208814ca800e5634df6

Observation fed01b56-a9b5-41bd-b631-c15969f2efa3 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Language Model as Visual Explainer Learning transferable visual models from natural language supervision

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.153673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.153673Z digest=sha256:52ff8d207dbfc99b8f43225913fca7663f175cbcb1379817322928d63a48d141

Observation 87a05ac6-7a06-4738-9e24-10d78c3542b7 · outbound

This paper cites Maximum common subgraph isomorphism algorithms for the matching of chemical structures.

Language Model as Visual Explainer Maximum common subgraph isomorphism algorithms for the matching of chemical structures

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:07:36.239504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.158035Z digest=sha256:db826a4715ada8f1ed1566b7f74e9b6dd6d7318dee95e52ee311881b9ab64d42

Observation 7f07c63c-e633-467a-86ff-2c4ce46c6727 · outbound

This paper cites why should I trust you?.

Language Model as Visual Explainer why should I trust you?

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:07:36.227167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.161908Z digest=sha256:eb9e3078eeb2a394f84f0d2670f8b73b06ca4ff7d0d65fe9dc5f2facbb50837f

Observation ab3ff483-422b-4ab0-ae01-c0c90f4be7d7 · outbound

This paper cites High-resolution image synthesis with latent diffusion models, 2021.

Language Model as Visual Explainer High-resolution image synthesis with latent diffusion models, 2021

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.165699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.165699Z digest=sha256:90d0c43b78528ba55dc3639522873ef4ccf2fd9456a0037a20cb3462e7a3289b

Observation 1367e644-4475-43c6-8553-4cb2f0d55192 · outbound

This paper cites Imagenet large scale visual recognition challenge.

Language Model as Visual Explainer Imagenet large scale visual recognition challenge

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:07:36.205510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.169519Z digest=sha256:f4456f7150d47d17b8a6c03d17f053f2fd6dd269b4027e43fd0af2f9e9caf9a1

Observation 0c1c8e91-b93b-4b96-aba5-6f268e79efd6 · outbound

This paper cites Protopshare: Prototypical parts sharing for similarity discovery in interpretable image classification.

Language Model as Visual Explainer Protopshare: Prototypical parts sharing for similarity discovery in interpretable image classification

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.173384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.173384Z digest=sha256:e00e6fb9657f9c35808fcf9fd7298f90d69e57ad4d8c6f1b94b7c7ed39b2962e

Observation ee74399c-0b27-40fa-8266-52c954aedcdd · outbound

This paper cites Mobilenetv2: Inverted residuals and linear bottlenecks.

Language Model as Visual Explainer Mobilenetv2: Inverted residuals and linear bottlenecks

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.178023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.178023Z digest=sha256:4ef041dfabcefbaaf717b571ef52ea5102f2b8d2ab4160c68b4114a3ac8eeb44

Observation 000088e9-98c3-4e6d-8457-398c45b882f8 · outbound

This paper cites Rule extraction from neural networks via decision tree induction.

Language Model as Visual Explainer Rule extraction from neural networks via decision tree induction

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:07:36.175606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.182506Z digest=sha256:00ea7bce66aa0ee105e210bde55c9698e94ebb1f8b907f80cf1b9937c921f341

Observation ab04397f-8eec-458d-99a3-371d72371518 · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools.

Language Model as Visual Explainer Toolformer: Language Models Can Teach Themselves to Use Tools

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.186448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.186448Z digest=sha256:1b3a381add41a408bba7a26aeade841eb2efb17ea3db3ee4ad041344a25e30fa

Observation 0dfa9e74-b922-4476-92ce-91eda4cefb69 · outbound

This paper cites Grad-cam: Visual explanations from deep networks via gradient-based localization.

Language Model as Visual Explainer Grad-cam: Visual explanations from deep networks via gradient-based localization

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.191172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.191172Z digest=sha256:db917732d8d07b5ae4fe0372004064ccab694cdd80051587ec7d6c480d9cbb95

Observation 408783c6-f5d1-429b-89b4-e392b4dc34c6 · outbound

This paper cites HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face.

Language Model as Visual Explainer HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.195048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.195048Z digest=sha256:4805fa33341300eb5fba4db571dd767beff48275bf8fc55b02025a7ef26249ff

Observation c8c8a90e-a304-461c-82b0-1225f4334e74 · outbound

This paper cites Learning important features through propagat- ing activation differences.

Language Model as Visual Explainer Learning important features through propagat- ing activation differences

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:07:36.154542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.199769Z digest=sha256:8508dea15548c40e06f13e93d79b24a8fe451cd640abcc0908b10fb52a0a9916

Observation 4f432585-f6bb-495c-8a8a-cc68b2dd27e9 · outbound

This paper cites Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps.

Language Model as Visual Explainer Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.204022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.204022Z digest=sha256:131c02bda5349239ae9d79998ca6f4ca9227716f36a74f2dc317f4c084bc611b

Observation fc6285cf-3876-41d7-b0ad-8bc420027a4e · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

Language Model as Visual Explainer Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.208656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.208656Z digest=sha256:dfde2f29c7e96169d7c7eb9978862e40cd5639778b89159f65a635c951c33844

Observation c1fa96b2-3b1d-47ee-95ce-b36cd698130d · outbound

This paper cites James Murdoch, and Bin Yu.

Language Model as Visual Explainer James Murdoch, and Bin Yu

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:07:36.142062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.213179Z digest=sha256:a61f6eda1192c159c2a2845ab72f8926f57482f3d47a20a02373a02f667b102d

Observation d5415e60-4b4e-43a4-820d-72ab50358e66 · outbound

This paper cites SmoothGrad: removing noise by adding noise.

Language Model as Visual Explainer SmoothGrad: removing noise by adding noise

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.216835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.216835Z digest=sha256:beec97f53f7adbb5d357fc9a1fd934cb1f3dc7e63a9aa84acf738289f2dfd9b1

Observation 32855f1d-5339-4a54-8be9-015b75b3c3eb · outbound

This paper cites Prototypical networks for few-shot learning.

Language Model as Visual Explainer Prototypical networks for few-shot learning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.220974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.220974Z digest=sha256:a2d78f0710bd323550b4c67e76fb9522fae750c9dfd5b07759aedcc4ab2ff2a2

Observation 2b294a20-ba62-4f81-9cba-0fcb49a3bab9 · outbound

This paper cites Neural trees-using neural nets in a tree classifier structure.

Language Model as Visual Explainer Neural trees-using neural nets in a tree classifier structure

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:07:36.119510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.224786Z digest=sha256:911fb1da974bbb8967c3e435c243219f7908938ccd3ba55690ca034f9f7a2b2c

Observation d7653afb-1897-4bf0-a6cc-bb2398cf34b6 · outbound

This paper cites Tree sequence kernel for natural language.

Language Model as Visual Explainer Tree sequence kernel for natural language

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:07:36.106327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.228574Z digest=sha256:e4deb3d000fc2758f34033e3a0b69ef05ac52cee365ae6c59e0ebf23ef5a3e68

Observation bd7b5a49-b624-4520-b810-c33b085fca85 · outbound

This paper cites Going deeper with convolutions.

Language Model as Visual Explainer Going deeper with convolutions

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.232239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.232239Z digest=sha256:ec4f53f63a9979d8d06ad51256d5e81eaa94c01aeabc643071bf3dabeb833698

Observation 223dc767-76cf-4ada-868d-df7a0ca15087 · outbound

This paper cites Rethinking the inception architecture for computer vision.

Language Model as Visual Explainer Rethinking the inception architecture for computer vision

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.236073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.236073Z digest=sha256:bbabdeebf7c548963d7259ad8098b7cb3ff465045aacef9268c403c4d1559022

Observation a0e2e079-4e5b-4db8-b637-e095d8f0c257 · outbound

This paper cites Visual correspondence-based explanations improve ai robustness and human-ai team accuracy.

Language Model as Visual Explainer Visual correspondence-based explanations improve ai robustness and human-ai team accuracy

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:07:36.075258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.239674Z digest=sha256:a6de6f6ca4173d3481c0238a60546191fe35958ce49b291a5785ce8586de9796

Observation 121342dc-954d-4434-ae29-90744e4c7f44 · outbound

This paper cites Adaptive neural trees.

Language Model as Visual Explainer Adaptive neural trees

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:07:36.060849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.243542Z digest=sha256:63f12245cf0704644097e916999a164dda19c98deb5ea8e52a268d2545e3e597

Observation 52ea45d8-eea2-4a02-93b6-1a4eaaf3138a · outbound

This paper cites an unresolved cited work.

Language Model as Visual Explainer Unresolved cited work

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.247420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.247420Z digest=sha256:e139e4a7bf13d298c6178ce62d014a88f8705b7ca4a943d65a4186f574cfce88

Observation b9446a37-73e3-4187-a810-1ef73f073e60 · outbound

This paper cites NBDT: Neural-Backed Decision Trees.

Language Model as Visual Explainer NBDT: Neural-Backed Decision Trees

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.251283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.251283Z digest=sha256:ce7e1798d513e66f932d67a668710fd2b81c994cebbcb398018ec6f82e777660

Observation f9f7b7d7-cc84-4119-baed-714e8e1753ba · outbound

This paper cites Chestx-ray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases.

Language Model as Visual Explainer Chestx-ray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.255610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.255610Z digest=sha256:7ce2663ed72da0ce169f63dff6369f35ecbcda875769f1c88a0577e0ff59e6eb

Observation ba2397b0-19e5-4056-ba89-d2f04f01093a · outbound

This paper cites A tree-based decoder for neural machine translation.

Language Model as Visual Explainer A tree-based decoder for neural machine translation

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:07:36.025760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.260080Z digest=sha256:fda1e1c8212979c44cf1582a5d75786a526345d7be521e67450bcc9d27d46722

Observation dde06ef2-8311-47b9-98ff-33916e4441a2 · outbound

This paper cites MedCLIP: Contrastive Learning from Unpaired Medical Images and Text.

Language Model as Visual Explainer MedCLIP: Contrastive Learning from Unpaired Medical Images and Text

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.264243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.264243Z digest=sha256:6f529864aaa64f89e917ffb9cccc62508aa7ba0e44058b0eedd5ca06393100bd

Observation 3ad0d013-b7ac-47e1-b349-929c8d6a8f22 · outbound

This paper cites Zero-shot learning—a compre- hensive evaluation of the good, the bad and the ugly.

Language Model as Visual Explainer Zero-shot learning—a compre- hensive evaluation of the good, the bad and the ugly

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.269179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.269179Z digest=sha256:346ba0c71ee0a3b87ae9964502e07ffc618eefd74d13a252fd1b7ea60f51767a

Observation 0bad459e-d433-470b-bb02-a03e20c937f6 · outbound

This paper cites Attribute prototype network for zero-shot learning.

Language Model as Visual Explainer Attribute prototype network for zero-shot learning

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:07:35.997304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.273865Z digest=sha256:6d3bb8cd578448642167dbacc0918c58e58c6c517f0aa05cad8968d3f279b92f

Observation 238cfb74-64f7-467d-b679-71025aad6e47 · outbound

This paper cites Factorizing knowledge in neural networks.

Language Model as Visual Explainer Factorizing knowledge in neural networks

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:07:35.981894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.278441Z digest=sha256:b48625d64dde759935639d466ac84038ad1643bd504182210d94cd7a98e1676b

Observation 019fd634-aa22-4e7d-86d5-d88f70d79f1f · outbound

This paper cites Deep model reassembly.

Language Model as Visual Explainer Deep model reassembly

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.282678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.282678Z digest=sha256:56e44e5dd43bf3441f42a32aaca0b5cf98e3ccfa150ecedb525d8413f69c79aa

Observation 9c9a57d8-6ce2-43f1-b2d4-e73d6d6d3f38 · outbound

This paper cites Deep Neural Decision Trees.

Language Model as Visual Explainer Deep Neural Decision Trees

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.288080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.288080Z digest=sha256:bf3bfe2b1490418614e5dd3f969c4f1685758e587d614d8af00d4b92fabb6c24

Observation ab3045e6-3680-4036-bae6-a84d8bb34101 · outbound

This paper cites Language in a bottle: Language model guided concept bottlenecks for interpretable image classification.

Language Model as Visual Explainer Language in a bottle: Language model guided concept bottlenecks for interpretable image classification

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:07:35.949256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.292662Z digest=sha256:04e15edf667c082a14e6c7311f3e1b0c1fa4f286da51fb46cb8ed97d42da888a

Observation 7b9ca1a5-f80f-417e-b28e-f073de009633 · outbound

This paper cites MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action.

Language Model as Visual Explainer MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.297119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.297119Z digest=sha256:bea2bdfdfdfa4371c37ef168589b65ac2314b460e0ef8851d769fb6b3a90a7cc

Observation 8ecfed26-821e-43ee-858d-9eff72c03f60 · outbound

This paper cites Representer point selection for explaining deep neural networks.

Language Model as Visual Explainer Representer point selection for explaining deep neural networks

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:07:35.935637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.301691Z digest=sha256:6230fabbf3265261b8d8ee8a02ff9cbe60056325b30767ef13bbbe11d1ec3aba

Observation 23a6f7d5-9969-42e5-9246-297f4428407e · outbound

This paper cites Neural- symbolic vqa: Disentangling reasoning from vision and language understanding.

Language Model as Visual Explainer Neural- symbolic vqa: Disentangling reasoning from vision and language understanding

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:07:35.921677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.305615Z digest=sha256:27c2ac804fa5f98528185fb4645603c84d366232ec3f1458e3a22b78ee49cefe

Observation 62beda2f-1292-4c70-8aed-56a1a4065caa · outbound

This paper cites Visualizing and understanding convolutional networks.

Language Model as Visual Explainer Visualizing and understanding convolutional networks

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.309562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.309562Z digest=sha256:54944f576ad7cf9c52e52ea080034c9bbf29aa358c63ae9c7c3f419b6fd027b8

Observation 73d66753-5129-464a-bb72-60d2d4e26ff7 · outbound

This paper cites Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language.

Language Model as Visual Explainer Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.313817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.313817Z digest=sha256:ce3c39311bc03bf1f71583fc3bd075ffa6c781f9ce8707a88a26e2380d74927f

Observation 9db7e0eb-7e5b-4419-be83-dcdc2797d95e · outbound

This paper cites Use all the labels: A hierarchical multi-label contrastive learning framework.

Language Model as Visual Explainer Use all the labels: A hierarchical multi-label contrastive learning framework

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:07:35.898108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.318375Z digest=sha256:8e08549b8ccbb79a35c15058d3799ea70dc75e955629462b8e9622609a55ec1e

Observation 55f368ce-dd45-4271-b6ef-3433280562ce · outbound

This paper cites Diagnosing and rectifying vision models using language.

Language Model as Visual Explainer Diagnosing and rectifying vision models using language

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:07:35.882378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.322670Z digest=sha256:1c8f7e6113c0cb0c422e5ba9ad3caaba5e938bc9f7598ccd8e6ac633c337522b

Observation 04b7c9f0-dc56-4db6-9cd3-2e96ebbe2408 · outbound

This paper cites Evolutionary design of neural network tree-integration of decision tree, neural network and ga.

Language Model as Visual Explainer Evolutionary design of neural network tree-integration of decision tree, neural network and ga

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:07:35.867501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.327473Z digest=sha256:b6d30f88971b11347b99b7f342916acfeca9a28aeb222aeae21563e8bfb99fff

Observation 9d66f4cc-8b04-48ad-9120-5bd7c8c4f8c1 · outbound

This paper cites Evaluating commonsense in pre-trained language models.

Language Model as Visual Explainer Evaluating commonsense in pre-trained language models

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:07:35.854722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.331577Z digest=sha256:edbd02dec89f7dc5e05f1a4818baea6d7cefb00c15a107e326b62cb507736206

Observation f0d54ef9-b2c5-4724-b7ea-c3a8cc0e0af4 · outbound

This paper cites Deepred–rule extraction from deep neural networks.

Language Model as Visual Explainer Deepred–rule extraction from deep neural networks

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:07:35.841264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.335362Z digest=sha256:6a9dc98e32d6f6ce1ed1e416038541dc08ae7ea1465428bc81fceaa0b8f488e8

Observation d0e55e4c-82d2-4c0c-93a1-8e4bbd0531b2 · outbound

This paper cites In this first step, we utilized ChatGPT 4 to generate initial attribute trees for each class with the category name, detailed Section 3.1.

Language Model as Visual Explainer In this first step, we utilized ChatGPT 4 to generate initial attribute trees for each class with the category name, detailed Section 3.1

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:07:35.829073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.339576Z digest=sha256:0fdfdf088906680d1ab378fe0e24ac444e39fae6099a651fa9367443c64b067c

Observation 14b0e640-0850-44a6-b0e3-002798340b66 · outbound

This paper cites To determine whether an attribute was present or absent in an image, we employed an ensemble of predictions from multiple CLIP [53] models 5.

Language Model as Visual Explainer To determine whether an attribute was present or absent in an image, we employed an ensemble of predictions from multiple CLIP [53] models 5

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:07:35.815519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.343498Z digest=sha256:b91f54f441f23706385bb3e92a69f71ae9b3dced03a6bc49fc07a03e6fe2a157

Observation 9af1c2fe-9c1d-4526-8a32-275d0c750f35 · outbound

This paper cites crane” could represent either a construction machine or a bird, leading to ambiguity. Furthermore, an image described as “a dog with a long tail.

Language Model as Visual Explainer crane” could represent either a construction machine or a bird, leading to ambiguity. Furthermore, an image described as “a dog with a long tail

Reference 99

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T20:07:35.801610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.348240Z digest=sha256:4b30750a5ac29c2785f067d6a8dc3d7a0d0f28563dd9e3d10ce75d2efc0d8138

Observation de991cee-695d-4af1-bf63-91fbf7ed9df8 · outbound

This paper cites an unresolved cited work.

Language Model as Visual Explainer Unresolved cited work

Reference 100

Resolution
unresolved
raw_fallback, observed 2026-08-11T20:07:35.786802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:07:35.352999Z digest=sha256:23b8e7532b31c61749b975b48eb5ff7409d467ffdfbd4e8a73ea86df152ae75b

Pith citing papers

No inbound Pith citation observations are available.