Pith. sign in

arXiv preprint arXiv:2401.06102 , year=

12 Pith papers cite this work, alongside 2 external citations. Polarity classification is still indexing.

12 Pith papers citing it
2 external citations · external index

citation-role summary

dataset 1

citation-polarity summary

years

2026 10 2025 2

roles

dataset 1

polarities

use dataset 1

representative citing papers

PRISM: Recovering Instruction Sets from Language Model Activations

cs.AI · 2026-06-08 · unverdicted · novelty 7.0

PRISM is a new activation-conditioned model that recovers full sets of simultaneous instructions from LLM hidden states via judge-guided GRPO training and outperforms prior activation-to-language methods on security-relevant tasks.

A framework for analyzing concept representations in neural models

cs.CL · 2026-05-02 · unverdicted · novelty 7.0

A new framework shows concept subspaces are not unique, estimator choice affects containment and disentanglement, LEACE works well but generalizes poorly, and HuBERT encodes phone info as contained and disentangled from speaker info while speaker info resists compact containment.

How Do Language Models Compose Functions?

cs.CL · 2025-10-02 · conditional · novelty 6.0

LLMs solve compositional factual recall either by computing intermediates or directly, with mechanism choice correlated to translation geometry in embedding spaces.

citing papers explorer

Showing 12 of 12 citing papers.