Pith. sign in

Are self-explanations from Large Language Models faithful?

8 Pith papers cite this work. Polarity classification is still indexing.

8 Pith papers citing it

years

2026 8

verdicts

UNVERDICTED 8

representative citing papers

iPOE: Interpretable Prompt Optimization via Explanations

cs.CL · 2026-05-18 · unverdicted · novelty 6.0

iPOE generates and optimizes annotation guidelines from explanations to produce interpretable prompts, reporting up to 39% gains over baselines on four datasets with LLM explanations substituting for human ones.

Superficial Beliefs in LLM Decision-Making

cs.AI · 2026-06-09 · unverdicted · novelty 5.0

LLMs show structured attribute-driven decisions that a behavioral model can predict, but self-reports recover those drivers only partially, indicating superficial beliefs.

citing papers explorer

Showing 8 of 8 citing papers.