Introduces BonaFide benchmark of 3,066 ground-truth labeled CoTs showing most faithfulness metrics perform near chance with biases and poor scaling to longer chains.
Title resolution pending
4 Pith papers cite this work, alongside 56 external citations. Polarity classification is still indexing.
fields
cs.CL 4verdicts
UNVERDICTED 4representative citing papers
TLRD distills tri-level rationales (instance features, dataset distributions, neighbor comparisons) from a teacher into student LLMs to close the accuracy gap with tree ensembles on tabular data while generating grounded explanations.
A single LLM rewrite of skill descriptions using false positive and negative cases matches manual optimization performance in production, with most other pipeline components adding little value.
Activation verbalization methods for LLMs largely reflect the verbalizer model's parametric knowledge rather than privileged information from the target model's activations.
citing papers explorer
-
Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
Introduces BonaFide benchmark of 3,066 ground-truth labeled CoTs showing most faithfulness metrics perform near chance with biases and poor scaling to longer chains.
-
TLRD: Teaching LLMs to Reason over Tabular Data with Tri-Level Rationale Distillation
TLRD distills tri-level rationales (instance features, dataset distributions, neighbor comparisons) from a teacher into student LLMs to close the accuracy gap with tree ensembles on tabular data while generating grounded explanations.
-
A Single Rewrite Suffices: Empirical Lessons from Production Skill Description Optimization
A single LLM rewrite of skill descriptions using false positive and negative cases matches manual optimization performance in production, with most other pipeline components adding little value.
-
Do Activation Verbalization Methods Convey Privileged Information?
Activation verbalization methods for LLMs largely reflect the verbalizer model's parametric knowledge rather than privileged information from the target model's activations.