Pith. sign in

REVIEW 2 cited by

Mechanistic understanding and validation of large AI models with SemanticLens

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.05398 v1 pith:CP4KLXKB submitted 2025-01-09 cs.LG cs.AI

classification cs.LGcs.AI
keywords semanticlensmodelmodelsneuronsvalidationcomponentsexplanationhttps
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Unlike human-engineered systems such as aeroplanes, where each component's role and dependencies are well understood, the inner workings of AI models remain largely opaque, hindering verifiability and undermining trust. This paper introduces SemanticLens, a universal explanation method for neural networks that maps hidden knowledge encoded by components (e.g., individual neurons) into the semantically structured, multimodal space of a foundation model such as CLIP. In this space, unique operations become possible, including (i) textual search to identify neurons encoding specific concepts, (ii) systematic analysis and comparison of model representations, (iii) automated labelling of neurons and explanation of their functional roles, and (iv) audits to validate decision-making against requirements. Fully scalable and operating without human input, SemanticLens is shown to be effective for debugging and validation, summarizing model knowledge, aligning reasoning with expectations (e.g., adherence to the ABCDE-rule in melanoma classification), and detecting components tied to spurious correlations and their associated training data. By enabling component-level understanding and validation, the proposed approach helps bridge the "trust gap" between AI models and traditional engineered systems. We provide code for SemanticLens on https://github.com/jim-berend/semanticlens and a demo on https://semanticlens.hhi-research-insights.eu.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Relevance-driven Input Dropout: an Explanation-guided Regularization Technique

    cs.LG 2025-05 conditional novelty 6.0 of 10

    RelDrop, which occludes the most attribution-relevant input regions during training, improves generalization and occlusion robustness for image and point cloud classification.

  2. From What to How: Attributing CLIP's Latent Components Reveals Unexpected Semantic Reliance

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A new attribution framework combines sparse autoencoder components with gradient-based attribution to reveal how CLIP models rely on semantically unexpected concepts.

Pith tools