A causal audit with image interventions shows text-only models reach within 5.7 accuracy points of top multimodal VLMs on chest radiography, with some large multimodal models statistically indistinguishable from small text-only baselines.
Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization
30 Pith papers cite this work, alongside 21,555 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
representative citing papers
Signed pairwise interaction scores conflate U/R/S; Stochastic Hi-Fi uses interventional masked inference to recover per-feature uniqueness, redundancy, and synergy profiles.
A synthetic patient-free pipeline generates saccadic eye movement data to train a classifier achieving AUROC 0.76 and sensitivity 0.71 on real clinical data for normal vs abnormal saccade detection.
EEG foundation models show no single winner across failure modes, attend to correct brain regions but decode corrupted signals, and retain task information in early layers while late layers adapt during fine-tuning.
Christoffel-DPS is a distribution-free optimal sensor placement framework for diffusion posterior sampling that provides non-asymptotic recovery bounds and outperforms Gaussian baselines on non-Gaussian benchmarks.
AsmRAG detects malware at 96% F1 and attributes families at 95% F1 by retrieving functionally similar assembly code via LLM embeddings and density-weighted anchor selection, remaining robust to metamorphic obfuscation.
TRACE is a RANO 2.0-aligned concept bottleneck model for 4-class longitudinal glioblastoma response classification on 3D MRI that reports 0.4769 macro F1 on the LUMIERE dataset via 5-fold patient-wise cross-validation.
SENTRY is a plug-and-play module that replaces confidence-based memory writes with neighbor-aware cycle-consistent validation in SAM2 trackers, yielding new zero-shot SOTA results on LaSOT, GOT-10k and other benchmarks.
XAI explanations should be narratives with continuous structure, cause-effect, fluency and diversity, and new metrics are needed to evaluate this better than standard NLP scores.
UFPR-VeSV is a new real-world dataset for fine-grained vehicle classification and automatic license plate recognition collected from Brazilian police cameras, with benchmarks demonstrating its difficulty and the value of joint task use.
Introduces a unified evaluation framework for XAI using five principled metrics and the PGCA method that fuses grid perturbation with Grad-CAM++ , reporting top scores in fidelity, interpretability and fairness on ResNet-50 models across five image domains.
A dual-edge graph fuses vessel-lesion geometry and embedding-biomarker sensitivity from four aligned streams to produce interpretable DR grades on APTOS images with 0.8076 accuracy.
Soft-labelling ordinal deep learning with binomial, beta, triangular, and exponential distributions improves KL and CPPD grading over one-hot baselines on knee X-rays.
Model interpretation methods are reformulated to emphasize baselines; gradient-based methods, IG, and Taylor expansion are unified with explicit baselines identified, and a revised IG is developed for improved results from any layer.
Empirical comparison shows gradient-based explanations for GNN node similarities are actionable, consistent, and retain effects when sparsified, unlike mutual information explanations.
A CNN classifies lung cytology patches as benign or malignant at 100% sensitivity and 96.4% specificity, then routes to one of two Transformer decoders to generate findings text achieving BLEU-4 of 0.828 on 801 images.
Deep learning model on large uniform healthy MRI dataset achieves brain age MAE of 4.06 years (hold-out) and 4.21 years (independent set), with frontal lobe prominence and age gaps linked to neuropsychological scores.
Compares LIME, input perturbation and attention for explaining QA on KB+text; proposes automatic evaluation paradigm and finds input perturbation superior in both automatic and human studies.
A pathway-constrained autoencoder extended to multi-omics integration improves breast cancer stratification and provides interpretable pathway activity scores.
PEFT-MedSAM adapts MedSAM by training only its mask decoder on ISIC 2018 skin lesion data, achieving Dice 0.9411 and outperforming U-Net (0.8715) and zero-shot MedSAM (0.8997), with PH2 validation (0.9467) and 98.27% Grad-CAM pointing accuracy.
Spatial Learning Entropy Maps derived from MLP weight adaptations during spatial pixel prediction tasks highlight image points with high learning impact.
Position paper proposing Model Science as a discipline to systematically analyze AI model behavior beyond benchmarks, drawing analogies from cognitive science, neuroscience, medicine, and agriculture.
Grad-ECLIP is an equivalent but flawed variant of attention-based interpretation, with two principles proposed to ensure model explanations reflect the original model.
A three-stage framework combines dual-head CNNs, saliency attribution, neuroanatomical atlas mapping, and LLMs to generate interpretable reports for brain tumor classification on MRI images.
citing papers explorer
-
Vision-language models for chest radiography do not always need the image
A causal audit with image interventions shows text-only models reach within 5.7 accuracy points of top multimodal VLMs on chest radiography, with some large multimodal models statistically indistinguishable from small text-only baselines.
-
The Representational Limit of Scalar Interactions: An Interventional Decomposition
Signed pairwise interaction scores conflate U/R/S; Stochastic Hi-Fi uses interventional masked inference to recover per-feature uniqueness, redundancy, and synergy profiles.
-
GenEyePose: Patient-Free, Knowledge-Based Saccadic Eye Movement Modeling for Digital Neurophysiologic Biomarker Development
A synthetic patient-free pipeline generates saccadic eye movement data to train a classifier achieving AUROC 0.76 and sensitivity 0.71 on real clinical data for normal vs abnormal saccade detection.
-
Beyond Accuracy: Robustness, Interpretability and Expressiveness of EEG Foundation Models
EEG foundation models show no single winner across failure modes, attend to correct brain regions but decode corrupted signals, and retain task information in early layers while late layers adapt during fine-tuning.
-
Christoffel-DPS: Optimal sensor placement in diffusion posterior sampling for arbitrary distributions
Christoffel-DPS is a distribution-free optimal sensor placement framework for diffusion posterior sampling that provides non-asymptotic recovery bounds and outperforms Gaussian baselines on non-Gaussian benchmarks.
-
AsmRAG: LLM-Driven Malware Detection by Retrieving Functionally Similar Assembly Code
AsmRAG detects malware at 96% F1 and attributes families at 95% F1 by retrieving functionally similar assembly code via LLM embeddings and density-weighted anchor selection, remaining robust to metamorphic obfuscation.
-
TRACE: A Concept Bottleneck Model for Longitudinal 3D Glioblastoma Response Assessment
TRACE is a RANO 2.0-aligned concept bottleneck model for 4-class longitudinal glioblastoma response classification on 3D MRI that reports 0.4769 macro F1 on the LUMIERE dataset via 5-fold patient-wise cross-validation.
-
SENTRY: SAM2-Enhanced Neighbor-Aware and Temporally Reasoned Memory for Visual Tracking
SENTRY is a plug-and-play module that replaces confidence-based memory writes with neighbor-aware cycle-consistent validation in SAM2 trackers, yielding new zero-shot SOTA results on LaSOT, GOT-10k and other benchmarks.
-
On the Importance and Evaluation of Narrativity in Natural Language AI Explanations
XAI explanations should be narratives with continuous structure, cause-effect, fluency and diversity, and new metrics are needed to evaluate this better than standard NLP scores.
-
Toward Unified Fine-Grained Vehicle Classification and Automatic License Plate Recognition
UFPR-VeSV is a new real-world dataset for fine-grained vehicle classification and automatic license plate recognition collected from Brazilian police cameras, with benchmarks demonstrating its difficulty and the value of joint task use.
-
A Unified Framework for Evaluating and Enhancing the Transparency of Explainable AI Methods via Perturbation-Gradient Consensus Attribution
Introduces a unified evaluation framework for XAI using five principled metrics and the PGCA method that fuses grid perturbation with Grad-CAM++ , reporting top scores in fidelity, interpretability and fairness on ResNet-50 models across five image domains.
-
A Dual Edge Spatial Jacobian Image Graph for Interpretable Diabetic Retinopathy Grading
A dual-edge graph fuses vessel-lesion geometry and embedding-biomarker sensitivity from four aligned streams to produce interpretable DR grades on APTOS images with 0.8076 accuracy.
-
From Kellgren-Lawrence to Calcium Pyrophosphate Crystal Deposition: A Soft-Labelling Framework for Knee Osteoarthritis Assessmen
Soft-labelling ordinal deep learning with binomial, beta, triangular, and exponential distributions improves KL and CPPD grading over one-hot baselines on knee X-rays.
-
The Neglected Baseline in Model Interpretation
Model interpretation methods are reformulated to emphasize baselines; gradient-based methods, IG, and Taylor expansion are unified with explicit baselines identified, and a revised IG is developed for improved results from any layer.
-
Explaining Graph Neural Networks for Node Similarity on Graphs
Empirical comparison shows gradient-based explanations for GNN node similarities are actionable, consistent, and retain effects when sparsified, unlike mutual information explanations.
-
Automated Description Generation of Cytologic Findings for Lung Cytological Images Using a Pretrained Vision Model and Dual Text Decoders: Preliminary Study
A CNN classifies lung cytology patches as benign or malignant at 100% sensitivity and 96.4% specificity, then routes to one of two Transformer decoders to generate findings text achieving BLEU-4 of 0.828 on 801 images.
-
Estimating brain age based on a healthy population with deep learning and structural MRI
Deep learning model on large uniform healthy MRI dataset achieves brain age MAE of 4.06 years (hold-out) and 4.21 years (independent set), with frontal lobe prominence and age gaps linked to neuropsychological scores.
-
Interpretable Question Answering on Knowledge Bases and Text
Compares LIME, input perturbation and attention for explaining QA on KB+text; proposes automatic evaluation paradigm and finds input perturbation superior in both automatic and human studies.
-
Biologically Informed Deep Neural Networks for Multi-Omic Integration, Pathway Activity Inference and Risk Stratification in Cancer
A pathway-constrained autoencoder extended to multi-omics integration improves breast cancer stratification and provides interpretable pathway activity scores.
-
PEFT-MedSAM: Efficient Fine-Tuning of Medical Foundation Models for Explainable Skin Lesion Segmentation
PEFT-MedSAM adapts MedSAM by training only its mask decoder on ISIC 2018 skin lesion data, achieving Dice 0.9411 and outperforming U-Net (0.8715) and zero-shot MedSAM (0.8997), with PH2 validation (0.9467) and 98.27% Grad-CAM pointing accuracy.
-
Learning Entropy and Spatial Adaptation Dynamics of Multilayer Perceptrons for Structural Point Extraction
Spatial Learning Entropy Maps derived from MLP weight adaptations during spatial pixel prediction tasks highlight image points with high learning impact.
-
The Case for Model Science: Verify, Explore, Steer, Refine
Position paper proposing Model Science as a discipline to systematically analyze AI model behavior beyond benchmarks, drawing analogies from cognitive science, neuroscience, medicine, and agriculture.
-
Debunking Grad-ECLIP: A Comprehensive Study on Its Incorrectness and Fundamental Principles for Model Interpretation
Grad-ECLIP is an equivalent but flawed variant of attention-based interpretation, with two principles proposed to ensure model explanations reflect the original model.
-
Bridging visual saliency and large language models for explainable deep learning in medical imaging
A three-stage framework combines dual-head CNNs, saliency attribution, neuroanatomical atlas mapping, and LLMs to generate interpretable reports for brain tumor classification on MRI images.
-
Opportunistic Bone-Loss Screening from Routine Knee Radiographs Using a Multi-Task Deep Learning Framework with Sensitivity-Constrained Threshold Optimization
STR-Net achieves AUROC of 0.933 for binary bone-loss screening and 0.801 correlation for T-score estimation from knee X-rays on a held-out test set.
-
Fringe Projection Based Vision Pipeline for Autonomous Hard Drive Disassembly
An integrated fringe projection and AI pipeline delivers aligned high-accuracy 3D sensing and instance segmentation for autonomous HDD disassembly at 77.7 FPS.
-
Time-driven Survival Analysis from FDG-PET/CT in Non-Small Cell Lung Cancer
A time-aware ResNet-based model on PET/CT images improves overall survival prediction in NSCLC by incorporating temporal data, achieving 4.3% higher AUC than fixed-time baselines.
-
DB-FGA-Net: Dual Backbone Frequency Gated Attention Network for Multi-Class Brain Tumor Classification with Grad-CAM Interpretability
DB-FGA-Net fuses VGG16 and Xception backbones with a new Frequency-Gated Attention module to reach 99.24% accuracy on 4-class brain tumor classification without augmentation and generalizes to 95.77% on an independent dataset.
-
Clinical Validation of the Melanoscope AI Mobile Dermoscopy Clinical Decision Support System
Prospective single-center validation of a cascade deep learning dermoscopy CDSS found no false negatives for five malignant lesions and 88.3% specificity, with quantitative IoU assessment of attention maps.
-
Platonic Projection Structures: Operator-Induced Observability in Representation Learning
The paper introduces 'Platonic Projection Structures,' a reformulation of standard PSD operator theory applied to representation learning, with experiments that verify definitions rather than test predictions.