Pith. sign in

REVIEW 3 cited by

Learning Interpretable Concepts: Unifying Causal Representation Learning and Foundation Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.09236 v2 pith:CEBYBTOC submitted 2024-02-14 cs.LG cs.AImath.STstat.MLstat.TH

classification cs.LGcs.AImath.STstat.MLstat.TH
keywords learningmodelsapproachbuildconceptsdataapproachescausal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

To build intelligent machine learning systems, there are two broad approaches. One approach is to build inherently interpretable models, as endeavored by the growing field of causal representation learning. The other approach is to build highly-performant foundation models and then invest efforts into understanding how they work. In this work, we relate these two approaches and study how to learn human-interpretable concepts from data. Weaving together ideas from both fields, we formally define a notion of concepts and show that they can be provably recovered from diverse data. Experiments on synthetic data and large language models show the utility of our unified approach.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures

    cs.CL 2025-06 conditional novelty 7.0 of 10

    A two-level mixture-of-experts sparse autoencoder models parent and child concepts together, improving reconstruction and reducing feature redundancy on Gemma 2-2B activations compared to flat top-k SAEs.

  2. Learning Task-Sufficient World Models by Synergizing Agentic Exploration and Structured Modeling

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Closed-loop agentic probing plus minimality/sufficiency masking recovers compact task-sufficient world-model latents that improve sample-efficient policy learning and cross-task generalization.

  3. Learning General Causal Structures with Hidden Dynamic Process for Climate Analysis

    cs.LG 2025-01 conditional novelty 6.0 of 10

    CaDRe jointly recovers latent dynamic processes and observed causal graphs from time-series data, with identifiability theory and competitive climate forecasting.

Pith tools