Pith. sign in

Eliciting latent knowledge from quirky language models

7 Pith papers cite this work. Polarity classification is still indexing.

7 Pith papers citing it

citation-role summary

background 1

citation-polarity summary

fields

cs.AI 5 cs.LG 2

years

2026 5 2025 2

roles

background 1

polarities

background 1

representative citing papers

The Impossibility of Eliciting Latent Knowledge

cs.AI · 2026-06-10 · unverdicted · novelty 7.0

Proves that no behavior-dependent feedback training strategy can guarantee an honest agent for latent knowledge even with perfect training feedback.

Deep Minds and Shallow Probes

cs.LG · 2026-05-12 · unverdicted · novelty 7.0

Symmetry under affine reparameterizations of hidden coordinates selects a unique hierarchy of shallow coordinate-stable probes and a probe-visible quotient for cross-model transfer.

Radical AI Interpretability

cs.AI · 2026-06-25 · unverdicted · novelty 6.0

A framework is proposed for solving for an AI system's beliefs and desires from its computational facts, with criteria for success tied to interpretability tests and emphasis on holistic attribution.

citing papers explorer

Showing 7 of 7 citing papers.