Pith. sign in

REVIEW 1 cited by

Crafting Large Language Models for Enhanced Interpretability

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.04307 v1 pith:NUHZW4HO submitted 2024-07-05 cs.CL cs.LG

classification cs.CLcs.LG
keywords llmslanguagecb-llminterpretabilitylargemodelsblack-boxclear
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce the Concept Bottleneck Large Language Model (CB-LLM), a pioneering approach to creating inherently interpretable Large Language Models (LLMs). Unlike traditional black-box LLMs that rely on post-hoc interpretation methods with limited neuron function insights, CB-LLM sets a new standard with its built-in interpretability, scalability, and ability to provide clear, accurate explanations. This innovation not only advances transparency in language models but also enhances their effectiveness. Our unique Automatic Concept Correction (ACC) strategy successfully narrows the performance gap with conventional black-box LLMs, positioning CB-LLM as a model that combines the high accuracy of traditional LLMs with the added benefit of clear interpretability -- a feature markedly absent in existing LLMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sparse Autoencoders for Hypothesis Generation

    cs.CL 2025-02 conditional novelty 6.0 of 10

    HypotheSAEs trains a sparse autoencoder on text embeddings, selects outcome-predictive neurons with Lasso, and asks an LLM to convert each neuron into a natural-language hypothesis.

Pith tools